AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

In fashion, it’s not just about how much you produce but how well you refine and prioritize your collection. Similarly, in AI-driven business, sheer diligence doesn’t guarantee success. A recent live experiment with AI models reveals that focus and strategic reading—rather than volume—determine whether an AI truly delivers value.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

How AI Models Performed Under Pressure

Four advanced AI models were tasked with managing a small software company through its toughest week—facing the same customers, crises, and ethical tests. This real-time experiment, accessible at firmulate.com/benchmarks.html, measures not just technical skill but managerial discipline and trustworthiness.

The Results Speak Volumes

  • The top scorer, gpt-5.6-sol, achieved a score of 95, successfully identifying hidden information and closing a €55,000 deal.
  • Kimi K3 scored 93, also closing the deal and demonstrating the cleanest discipline—refusing manipulation and reading the nuances hidden in company files.
  • Sonnet 5 followed closely with 88, closing the deal despite some slips in process discipline.
  • Fable 5 trailed at 77, also managing to close, but with more process lapses.

Interestingly, a simple baseline score of 26, representing minimal progress, underscores how much more is needed beyond raw effort—trust is fragile, and a single breach caps potential gains.

Amazon

AI deep reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Matters Most

What set the winning models apart was their capacity to access and interpret information buried two document references deep in the company’s files. The models that read these files demonstrated the ability to close the deal at full price—adding over €4,500 in monthly recurring revenue (MRR). This finding emphasizes that diligence alone isn’t enough; strategic, deep reading wins the day.

The Ethical Test: No to Manipulation

All models faced social engineering attempts—fake CEO messages escalating through stages and even a reporter trick asking for a discreet approval. Remarkably, all five models refused these manipulative tactics, aligning with their built-in reasoning: treating suspicious requests as impersonation risks.

Amazon

AI ethical decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Does This Mean for Business AI?

The live experiment is conducted on a real, functioning company platform at firmulate.com/live. Here, 13 synthetic employees operate in a microcosm of business mechanics, burning €105,000 monthly against a modest €2,300 MRR, with every decision versioned and auditable. This setup offers a transparent view of how AI models handle real-world pressures and temptations.

Discipline vs. Impact

The most thorough participant, Opus 4.8, incorporated over 80 learned rules and performed deep analyses yet finished last—its discipline waned, and opportunities to escalate issues into proper channels were missed. This highlights a core lesson: volume of effort without strategic focus can be a trap.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Should Leaders Take Away?

The key insight isn’t about how well AI writes but whether it can complete what it starts—reading key data, resisting shortcuts, and maintaining honesty under pressure. For managers and decision-makers, the challenge is to prioritize meaningful engagement over exhaustive but unfocused effort.

Benchmarking and Testing AI Readiness

The live benchmarks at firmulate.com/benchmarks.html demonstrate that AI models can perform complex tasks, but their true value hinges on their ability to read deeply and act ethically. Running these wargames before deploying AI in critical environments ensures you’re not just hiring a diligent worker but a trustworthy partner.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI business decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mary Kay Surges In Global Coverage

Mary Kay’s media mentions have surged, with 23 reports within a recent period, indicating increased global attention to the brand’s activities.

MOQ & Dye Lots in Hosiery Production

Navigating MOQ and dye lot management in hosiery production is crucial for quality and efficiency—discover the key factors that can impact your results.

The Difference Between Style-Led and Function-Led Legwear Purchases

Great style-led legwear often sacrifices durability, but understanding the key differences can help you make smarter, more balanced choices.

Financiere Richemont Surges In Global Coverage

Richemont’s media mentions have increased significantly, with 21 reports in recent coverage, marking a notable shift in its public and investor attention.