firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

What Home Decor Can Teach Us About AI’s True Power

Imagine a team of highly skilled decorators tasked with transforming a home just as chaos erupts around them. Their ability to spot hidden flaws and resist shortcuts could mean the difference between a stunning makeover and a costly disaster. Similarly, in the business world, the real strength of AI isn’t just in chatting or generating text – it’s in executing complex decisions under pressure. A groundbreaking live experiment with AI models running a real, money-earning company has uncovered what really distinguishes an effective AI from one that only looks good on the surface.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing AI in the Wild: The Crucible Experiment

In a live demonstration, four of the world’s top AI models were each given the same challenging task: to run a small software company through its toughest week. This company, publicly losing money with a burn rate of €105,000 per month against a monthly revenue of just €2,300, was set up with real crises, real customers, and the temptation to cheat or cut corners. Every decision made by these models was recorded, versioned, and auditable, providing a transparent view into their decision-making process.

What the Models Could and Could Not Do

Remarkably, all four AI models identified every crisis and refused every attempt at manipulation—even fake CEO messages escalating over stages and a reporter trick asking for a quick approval. But that’s where similarity ended. Only two of the four models managed to close the €55,000 deal their own analysis had earned, signing the contract at full price. The other two, despite diagnosing the same problems and making the same pitches, left the deal on the table.

The key difference? The winning models demonstrated disciplined reading of the company’s internal files, uncovering a crucial piece of information buried two document references deep in the company’s file system. This buried fact was a decisive advantage, allowing one model to secure a significant recurring revenue boost of over €4,500 monthly.

The Invisible Weakness

Interestingly, the models that failed to close the deal all suffered from a discipline weakness: they did not escalate their findings to a human or higher authority when their analysis was incomplete or uncertain. For instance, the most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, was last place because it left the deal unexecuted, preferring to lock the decision in a department rather than escalate it. This behavior highlights a critical insight: the ability to act decisively under pressure is often hidden and cannot be measured simply by chat-based demos or superficial tests.

The Lesson for Business Decision-Makers

For industries like home decor or gifts, where the focus is often on aesthetics or customer experience, this experiment underscores a vital point: an AI’s real value lies in its capacity to finish tasks, dig into details, and stay honest under pressure. It’s not enough for AI to produce convincing dialogue or generate appealing images. Can it read critical documents? Will it follow through on commitments? These hidden skills determine whether AI will truly support or undermine your operations.

Measuring True AI Performance

The live experiment’s results—ranked in the Crucible League—show that the highest-scoring model, gpt-5.6-sol, scored 95 out of 100, successfully finding the buried fact and closing the deal. Kimi K3, the newcomer, scored 93, with the cleanest discipline, while Sonnet 5 scored 88, and Fable 5 scored 77, mainly because it failed to execute the deal despite good rule discipline. These scores reflect not just intelligence or chat quality, but the ability to deliver tangible business outcomes.

Try It Yourself

Businesses curious about testing their own decision-making processes can run simulated scenarios against their real operations, without risking any actual systems. By leveraging tools like the live ‘wargame’ at firmulate.com, you can observe how AI models handle crises, temptations, and complex decisions—providing a clear measure of their true readiness for your enterprise.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Key Takeaway

The true strength of AI in business isn’t just in chat demos or surface-level interactions. It’s in its ability to uncover hidden truths, stay disciplined under pressure, and complete tasks that matter—especially when it counts. Testing AI in real-world scenarios reveals strengths and weaknesses invisible in simple demos, helping businesses choose the right tools for lasting success.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


SUMMER

Summer Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Average 30-Year U.S. Mortgage Rate Rises To Highest Level In A Year

The average 30-year U.S. mortgage rate has increased to its highest level in a year, impacting homebuyers and the housing market.

West Elm Surges In Global Coverage

West Elm experiences a surge in international coverage, with 39 mentions in recent media monitoring, highlighting rising global interest in the brand.

Wohngeld

Germany introduces new Wohngeld reforms to increase support for low-income households, effective from next year. Details on eligibility and amounts are still emerging.

Serpar To Auction Eight Lots Across Lima Districts

Serpar announces auction of eight land lots in various districts of Lima, offering investment opportunities. Details on locations and process outlined.