
Imagine a newcomer in fashion—fresh, daring, and untested—yet outpacing industry veterans in quality and discipline. Now, scale that story to AI models running real companies. At a recent live experiment, a new AI entrant, Kimi K3, showcased how the freshest disruptors can beat established giants, not with flashy talk but with solid, verifiable results. This is more than tech hype; it’s about trust, discipline, and measurable success in AI-driven business management.
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Breaking the League: The Live AI Wargame
In July 2026, four leading AI models faced a tough test: managing a small software company during its worst week. The challenge was real, with the same customers, crises, and temptations directed at each model. Every decision was carefully tracked and made transparently, ensuring a fair test of their managerial skills.
The Results Speak Louder Than Scores
- gpt-5.6-sol scored the highest with 95, securing the full performance in diagnosis and deal closure.
- Moonshot’s Kimi K3 was close behind at 93, demonstrating a clean discipline and strategic insight that led to winning a €55,000 deal, adding +€4,583 MRR.
- Sonnet 5 scored 88, also closing the deal but with minor slips in process discipline.
- Fable 5 and Opus 4.8 lagged behind at 77 and 73 respectively, with Opus notably leaving the final close on the table despite deep analyses.
The Hidden Power of Document Reading
“The decisive weakness was two document references deep in the company’s files, not in the immediate customer event,” the experiment revealed. Models that could read and interpret these internal documents succeeded in closing the deal at full price.
Resisting Social Engineering
The models faced a staged social engineering attack—phony CEO messages escalating in stages and a reporter trick. All five models refused to be manipulated, citing suspicion and adherence to protocols, with Kimi K3’s reasoning emphasizing a cautious approach to impersonation risks.
The Real Business Environment
Running the experiment live, the company simulated real money mechanics—burning €105k monthly against a mere €2.3k MRR—while maintaining over 680 self-learned rules to guide decisions. This setup allowed observers to watch AI decision-making unfold in real time at firmulate.com/live.
As an affiliate, we earn on qualifying purchases.
What Sets Kimi K3 Apart?
While many models demonstrated the ability to identify crises and refuse manipulation, Kimi K3 went further: it found buried internal data crucial for closing deals and did so without any signs of slipping discipline. Its approach was justified by its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Fairness and Testing Conditions
It’s noteworthy that K3 ran without an effort parameter—the API’s default—while others used a high effort setting. This shows K3’s efficiency and effectiveness without extra tuning, making its performance even more impressive.

The experiment underscores a vital insight: in AI-driven management, trustworthiness and discipline matter more than just conversational skills. The newcomer, Kimi K3, beat established models by reliably reading internal data, resisting manipulation, and closing deals—traits essential for AI to be a true business partner. As AI models are poised to touch critical business functions, choosing one that can finish what it starts and stay honest under pressure isn’t just smart; it’s essential. The league is open, and the current leaderboards highlight the importance of testing AI in real business scenarios—because without verified performance, it’s just hype.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI document reading and analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI cybersecurity and fraud prevention tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
