AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine a newcomer in fashion—fresh, daring, and untested—yet outpacing industry veterans in quality and discipline. Now, scale that story to AI models running real companies. At a recent live experiment, a new AI entrant, Kimi K3, showcased how the freshest disruptors can beat established giants, not with flashy talk but with solid, verifiable results. This is more than tech hype; it’s about trust, discipline, and measurable success in AI-driven business management.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Breaking the League: The Live AI Wargame

In July 2026, four leading AI models faced a tough test: managing a small software company during its worst week. The challenge was real, with the same customers, crises, and temptations directed at each model. Every decision was carefully tracked and made transparently, ensuring a fair test of their managerial skills.

The Results Speak Louder Than Scores

  • gpt-5.6-sol scored the highest with 95, securing the full performance in diagnosis and deal closure.
  • Moonshot’s Kimi K3 was close behind at 93, demonstrating a clean discipline and strategic insight that led to winning a €55,000 deal, adding +€4,583 MRR.
  • Sonnet 5 scored 88, also closing the deal but with minor slips in process discipline.
  • Fable 5 and Opus 4.8 lagged behind at 77 and 73 respectively, with Opus notably leaving the final close on the table despite deep analyses.

The Hidden Power of Document Reading

“The decisive weakness was two document references deep in the company’s files, not in the immediate customer event,” the experiment revealed. Models that could read and interpret these internal documents succeeded in closing the deal at full price.

Resisting Social Engineering

The models faced a staged social engineering attack—phony CEO messages escalating in stages and a reporter trick. All five models refused to be manipulated, citing suspicion and adherence to protocols, with Kimi K3’s reasoning emphasizing a cautious approach to impersonation risks.

The Real Business Environment

Running the experiment live, the company simulated real money mechanics—burning €105k monthly against a mere €2.3k MRR—while maintaining over 680 self-learned rules to guide decisions. This setup allowed observers to watch AI decision-making unfold in real time at firmulate.com/live.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Sets Kimi K3 Apart?

While many models demonstrated the ability to identify crises and refuse manipulation, Kimi K3 went further: it found buried internal data crucial for closing deals and did so without any signs of slipping discipline. Its approach was justified by its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Fairness and Testing Conditions

It’s noteworthy that K3 ran without an effort parameter—the API’s default—while others used a high effort setting. This shows K3’s efficiency and effectiveness without extra tuning, making its performance even more impressive.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The experiment underscores a vital insight: in AI-driven management, trustworthiness and discipline matter more than just conversational skills. The newcomer, Kimi K3, beat established models by reliably reading internal data, resisting manipulation, and closing deals—traits essential for AI to be a true business partner. As AI models are poised to touch critical business functions, choosing one that can finish what it starts and stay honest under pressure isn’t just smart; it’s essential. The league is open, and the current leaderboards highlight the importance of testing AI in real business scenarios—because without verified performance, it’s just hype.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making tools for companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document reading and analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI cybersecurity and fraud prevention tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Designer Tights Still Have a Place in the Luxury Conversation

The timeless appeal of designer tights in luxury fashion lies in their craftsmanship and bold innovation, leaving you curious about what makes them truly stand out today.

The AI That Worked Hard but Missed the Deal: A Lesson in Focus Over Volume

AI models excel in spotting crises and refusing manipulation, but deep reading and discipline are vital to closing deals. Prioritize focus over effort for real impact.

Why Even a Do-Nothing AI Gets a Score of 26 — And Why Trust Matters in Business Bots

Discover why even a do-nothing AI scores 26 points in the latest benchmark, highlighting the importance of honesty, trust, and thoroughness in business AI systems.

产品拓界 技术破局!2026 Intertextile Shanghai兰精演绎从纤维到产业的“价值共生” – News.tom.com

Lanxess showcased its integrated approach from fiber development to industrial applications at the 2026 Intertextile Shanghai, emphasizing ‘value co-existence’ in textile and chemical industries.