
Imagine buying a luxury dress and finding out it’s missing a zipper — not because it’s poorly designed, but because the seller simply chose not to add one. In the world of AI, even a ‘do-nothing’ baseline scores 26 points out of 100. That’s not a flaw; it’s a feature, revealing how trust and honesty are measured in the latest AI benchmarks.
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Heart of the Benchmark: Measuring Honest AI Performance
At first glance, you might wonder: why does a do-nothing AI, one that just sits on its hands, score 26 points? The answer is rooted in the benchmark’s methodology, which values partial progress and accountability. In essence, this score reflects the minimum baseline of honesty and decision-making the AI must demonstrate, even when it’s doing nothing.
Every AI model participating in the live experiment is tested by running the same challenging scenario: a small software company experiencing a tough week—difficult customers, crises, and the temptation to cheat or manipulate. The models are tasked with making decisions under real-world pressures, and every choice they make is tracked and verified. This rigorous approach ensures transparency and fairness.
Why Partial Progress Counts
In these tests, if an AI recognizes a crisis or identifies a critical document, it scores points—no matter if it ultimately closes the deal or not. For instance, the models that read the company’s files and spotted the buried fact, winning a deal worth over €4,583 in monthly recurring revenue, scored higher. Conversely, models that failed to find this key information or slipped up in process discipline scored lower, even if they identified some crises.
The Cap on Trust Breaches
Another unique aspect is the strict cap on trust violations. If a model attempts manipulation—say, fake CEO messages or impersonation—it is immediately capped and cannot score beyond a certain limit. This prevents overestimating a model’s integrity based on isolated successes and emphasizes the importance of consistent honesty over partial wins.
As an affiliate, we earn on qualifying purchases.
What the Results Reveal About AI and Business
The top performers, gpt-5.6-sol and Kimi K3, scored 95 and 93 respectively. They not only diagnosed issues accurately but also closed deals at full price—signaling both competence and integrity. Meanwhile, even the lowest scorer, Opus 4.8, with a score of 73, showed discipline lapsing in the final moments, leaving deals on the table.
Importantly, all models refused manipulative requests, such as escalating fake CEO messages or background approvals, demonstrating that honesty can be programmed and enforced even under pressure. Kimi K3’s on-record reasoning highlights this: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and Fashion
In fashion and luxury brands, trust is paramount. Customers expect that their data, preferences, and transactions are handled with integrity. The same principle applies to AI systems managing customer relationships, support, or forecasts.
As the live benchmark shows, AI models that are honest and disciplined tend to outperform in real-world scenarios. For example, the models that read deeper into company files closed bigger deals, showing that attention to detail and integrity translate into tangible business value.
As an affiliate, we earn on qualifying purchases.
What This Means for Your Business AI
When considering AI tools, don’t just ask if they can generate text or respond quickly. The real questions are: Will the AI finish what it starts? Will it read your files first? Will it stay honest when under pressure? And crucially, what’s the actual cost for a unit of useful work?
The Firmulate live experiment is a transparent, watchable demonstration of these principles, showing that a baseline score of 26 isn’t a flaw but a foundation—a reminder that trust and integrity are vital, even in the most advanced AI systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
