AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine planning a multi-day outdoor adventure where your success hinges not just on planning but on the trustworthiness of your gear. You need tools that don’t just perform well in ideal conditions but stay honest and reliable under pressure. The same principle applies to artificial intelligence in business. A recent experiment by Firmulate offers a clear view into how AI models handle trust, discipline, and critical decision-making — essential qualities for AI tools that manage your company’s most vital tasks.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get travel and outdoor gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Benchmark: More Than Just Scores

Business leaders evaluating AI tools might assume that the highest scores indicate the most capable models. But the latest experiment by Firmulate reveals a more nuanced picture. Four AI models were tested in a controlled simulation that mimicked a small software company facing a tough week — with real crises, tempting manipulations, and high-stakes decisions.

In this test, every decision was auditable and consistent, designed to expose how models handle trust and discipline. Surprisingly, while all four models identified every crisis and refused every manipulation attempt, only two managed to close the deal they had analyzed and earned — a €55,000 contract. The other two, despite spotting all problems, left the offer on the table, failing to act on their own insights.

The Do-Nothing Baseline Sets the Floor at 26

One striking finding: even a ‘do-nothing’ baseline score registered 26 points out of a possible 100. This is because partial progress counts in the scoring, and the baseline includes the bare minimum — recognizing crises without acting on them. It highlights an important truth: AI models can recognize issues but may still fail to act decisively or maintain trustworthiness.

Furthermore, the experiment emphasizes that a single breach of trust — like ignoring crucial information or slipping into deception — caps the total score. No amount of good performance elsewhere can compensate for lost integrity. This principle ensures that AI models are held accountable for honesty, not just technical accuracy.

Amazon

AI ethics and trustworthiness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Trust and Discipline Are Tested in Practice

In the simulation, models faced social engineering attacks, such as fake CEO messages escalating through multiple stages and attempts to impersonate the company’s decision-makers. All models refused these manipulative tricks — a promising sign that they can resist common scams. Kimi K3, one of the top performers, explained its refusal by treating such requests as potential impersonation or approval bypass attempts.

The experiment’s real-world counterpart is a live, operating business environment where AI models support a company with 13 synthetic employees handling real money — burning €105,000 monthly against a revenue of €2,300. Every decision, every rule, is versioned and transparent, allowing watchdogs to verify performance and ethical boundaries in real-time at firmulate.com/live.

Why Partial Success Matters

The experiment also shows that models which read deeper into documents and analyze more thoroughly tend to perform better. For instance, Opus 4.8, with the most comprehensive analysis (over 80 learned rules), still left some deals on the table and slipped into process slips, like writing into a locked department instead of escalating issues. This illustrates that thoroughness alone doesn’t guarantee perfect results but is crucial for trustworthiness.

Amazon

business AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Takeaway: Trust, Discipline, and the Cost of Integrity

For business managers, the key takeaway is that AI models should be evaluated not just on their ability to identify problems but on their capacity to act ethically and reliably under pressure. The experiment underscores that a model’s failure to act decisively or breaches of trust have tangible consequences, capping potential performance and risking company integrity.

Moreover, the benchmark’s transparent scoring — with a clear floor at 26 points — emphasizes that even the most basic level of AI performance involves some recognition of issues. Full performance involves reading relevant information — such as buried facts in documents — and acting on it appropriately. The models that did this well were rewarded with closing deals at full price, translating into millions of euros in revenue.

Amazon

AI model validation and benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Travel and Outdoor Enthusiasts

Just like planning a safe, enjoyable outdoor trip depends on trusting your gear and making disciplined decisions under stress, trusting your AI tools requires confidence that they will stay honest and act reliably when it counts. The Firmulate experiment offers a glimpse into how AI models can be tested for their integrity, ensuring they’re ready to support your business — whether it’s managing customer relationships, support queues, or forecasts.

In a world where AI touches every part of your workday, understanding these benchmarks helps you choose tools that don’t just perform well in ideal conditions but uphold trust and discipline when it matters most.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI cybersecurity and scam resistance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Keep Devices Powered During Long Transit Days

Just follow these essential tips to keep your devices charged during long transit days and discover how to stay connected longer.

Expedia Group Surges In Global Coverage

Media coverage of Expedia Group has increased sharply, with 20 mentions in recent reports, indicating rising global interest in the travel platform’s activities.

Can AI Managers Make the Same Decisions? The Live Experiment That Reveals Their True Personalities

Live AI experiments reveal how different frontier models handle business crises, trustworthiness, and closing deals—crucial insights for managing real-world risks.

I Moved My Digital Stack to Europe

A tech founder relocates their entire digital stack to Europe, citing concerns over data sovereignty and jurisdictional stability. Details of the migration process and implications explained.