
Imagine hiring an assistant who never lifts a finger but still costs you money. Now, consider that an AI model, even when doing nothing, scores 26 out of 100 in a new benchmark. How is this possible? And what does it tell us about trusting AI in your home or business? In the world of AI evaluations, even inaction counts, revealing surprising truths about honesty, discipline, and reliability.
Get decor and gifts delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the AI Benchmark That Surprises Even the Experts
At first glance, you might expect a baseline AI that does nothing to score zero. After all, it’s not completing tasks, making decisions, or delivering value. But in a recent public experiment conducted by Firmulate, the do-nothing model scored a solid 26 points. This isn’t a glitch or oversight; it’s a deliberate outcome rooted in the benchmark’s design.
The experiment involved four advanced AI models tasked with running a small software company through its worst week — facing real crises, customer demands, and ethical dilemmas. Every decision was recorded, and the models were held to strict standards of honesty and discipline. The key takeaway? All models flagged every crisis and refused manipulative tactics. Yet, only two managed to deliver a signed deal, earning full marks for their diagnosis and pitch.
Why Does the Do-Nothing Baseline Score 26?
The reason is rooted in the methodology itself. The benchmark awards partial credit for honest and cautious behavior, even if no progress is made. So, simply reading the company’s files without attempting manipulation or shortcuts earns a baseline score of 26. This reflects an honest, disciplined approach — even if the AI didn’t push forward.
Partial Progress Counts and the Trust Cap
Another crucial aspect is that one breach of trust — even a small one — caps the total score at 26. This underscores a fundamental principle: no good work can outweigh dishonesty. If an AI attempts manipulation or ignores ethical boundaries, its overall score cannot improve, regardless of other achievements.
AI ethics and trust training courses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Experiment Reveals About AI Reliability
The models’ ability to recognize crises and refuse manipulation is encouraging. All four models spotted every crisis and rejected fake CEO messages and reporter tricks. Interestingly, the decisive factor in securing the deal was reading a specific, buried document in the company’s files — a detail two document references deep. Only the models that read beyond surface information succeeded in closing the deal.
This highlights a vital lesson for businesses: the depth of information an AI can access and interpret deeply impacts its effectiveness. An AI that reads only surface-level data may miss critical insights, while one that digs deeper can deliver results that truly matter.
Trust and Discipline Under Pressure
In a simulated social engineering attack, five models faced staged messages from a fake CEO escalating demands, with a reporter urging a simple yes/no response. All five refused to comply — a sign of discipline. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
This discipline is crucial for deploying AI in sensitive environments. The models’ refusal to cooperate with manipulative tactics shows they can be trusted to uphold ethical standards under pressure, an essential feature for AI systems managing real-world tasks and data.
The Real-World Application: Managing Money and Risks
The live experiment runs a small company with 13 synthetic employees, managing real money mechanics — burning €105k monthly against €2.3k in revenue. Every workday, the AI models are versioned, and their decision-making is observed in real time at firmulate.com/live. Such transparent, watchable experiments help distinguish between AI that merely ‘talks a good game’ and AI that reliably performs.
The worst performer, Opus 4.8, with the most comprehensive analysis and learned rules, still left a deal on the table, demonstrating that even in thorough models, discipline can slip. The takeaway? No AI is perfect, but disciplined, honest models consistently outperform less disciplined ones, especially in high-stakes environments.
Implications for Home and Business Users
For the average consumer or business owner, these findings might seem distant. But they carry critical lessons. When AI agents touch your CRM, support queue, or financial forecasting, the question isn’t just about how well they write or chat. It’s whether they finish what they start, read deeply, and stay honest under pressure.
Just as in home decor or gift recommendations, trustworthiness and discipline matter. An AI that refuses to cut corners — even when no one’s watching — can be a reliable partner for your daily life.
Conclusion: A Benchmark That Keeps It Real
In a landscape often dominated by flashy scores and promises, Firmulate’s transparent, auditable benchmark sets a new standard. It values honesty, discipline, and depth over superficial performance. Even a do-nothing AI scores 26, illustrating that in AI evaluation, the true measure isn’t just what’s achieved — it’s how it’s achieved.
As AI continues to embed itself into everyday life, understanding these standards helps consumers and businesses make smarter choices. Trust, discipline, and depth aren’t optional — they’re the new baseline for AI performance.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
