firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Imagine choosing a gift for someone special — the more options you have, the harder it becomes to pick the perfect one. Now, picture an AI trying to manage a small company’s weekly crises with the same abundance of rules and data. Despite being thorough and diligent, it might still miss the most crucial moment to close a deal. This isn’t just a lesson in gift-giving but a revealing glimpse into how AI systems perform in real-world business challenges.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI to the Test in a Simulated Business Crisis

Recently, a groundbreaking experiment tested four advanced AI models by running them through the worst week a small software company could face. Every model faced the same customer issues, crises, and temptations to cut corners. Each decision was carefully versioned and auditable, revealing how these AI systems handle real-world complexity and pressure.

Performance and Key Findings

Interestingly, all four models successfully identified every crisis and refused every manipulation attempt, demonstrating their integrity and awareness. However, the critical difference lay in whether they closed the deal at full price. Only two models succeeded in signing the €55,000 contract earned through their analysis — the other two did not, despite similar diagnoses and pitches.

The reason? The decisive weakness was buried deep in the company’s files, two document references beneath surface-level reading. When models read and utilized this buried information, they closed the deal at full value, adding over €4,583 monthly recurring revenue (MRR). This emphasizes an essential insight: thoroughness alone isn’t enough if the AI doesn’t prioritize reading and analyzing the most impactful data.

Amazon

AI business decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Behavior Under Social Engineering Attacks

The experiment also included social engineering scenarios, such as fake CEO messages escalating over three stages and a reporter trick asking for a background yes/no. Remarkably, all five models refused these manipulative requests. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates growing AI resilience against social engineering — a crucial trait for trustworthy automation.

The Real-World Company Setup

The live company in the experiment operated 13 synthetic employees managing real money mechanics, burning €105k monthly against €2.3k MRR. It had a public cash countdown, over 680 self-learned rules, and every workday’s decisions were versioned, making the process transparent and auditable. The setup offers a tangible, ongoing view of how AI performs under pressure and real financial stakes.

The Deep Dive into Opus 4.8’s Profile

The most thorough participant in the experiment was Opus 4.8, engaging with over 80 learned rules and providing the deepest analyses. Despite its diligence, Opus 4.8 finished last — primarily because it left the close on the table and slipped discipline. Instead of escalating critical issues, it wrote attempts into a locked department. Similar weaknesses appeared, though less prominently, across all four models, proving that volume of rules doesn’t guarantee impact.

The Lesson for Business Leaders

The findings from this experiment are clear: AI models must prioritize their efforts, especially under pressure. Diligence, or the number of learned rules, isn’t enough. An AI must focus on impactful data and disciplined decision-making to be truly effective in closing deals or managing crises. An AI that knows everything but doesn’t prioritize the critical information or act decisively can still fall short, even when it’s well-trained.

Practical Implications and How to Prepare

For companies integrating AI into customer relations, support, or forecasting, the key question isn’t how well it writes or responds in demos. It’s whether the AI can see what matters most, stay honest under pressure, and finish what it starts. The upcoming league table shows the current leaders, with gpt-5.6-sol at the top, having found buried facts and closed the deal, while others trail behind.

Interested in testing your AI workforce? Firms can run the same kind of wargame against their own business data, safely and read-only, to see how their AI performs before deployment. Check out the benchmarks page for more details and ongoing results.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Wildcat Flooring Surges In Global Coverage

Wildcat Flooring experiences a significant surge in global coverage, with media mentions increasing eightfold. The reason for this spike remains unconfirmed.

Homes In London Are Cracking As The Clay They Are Built On Shrinks – CleanTechnica

Homes across London are experiencing cracking as the clay beneath them shrinks, raising safety concerns. Experts warn of ongoing risks and structural issues.

Oceana Passivhaus Surges In Global Coverage

Oceana Passivhaus sees a significant surge in international coverage, with 36 mentions in recent media analysis, highlighting growing interest in sustainable architecture.

Psychologie der Farbe in Valentinstagswerbung: Warum Pink verkauft

Ich werde enthüllen, wie die emotionale Kraft von Pink in Valentinswerbung die Wahrnehmung beeinflusst und warum sie wirklich verkauft.