AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Fashion brands rely on precision, trust, and consistency — qualities that are just as vital when deploying AI in business operations. Imagine AI models running a company through its toughest week, making decisions that could mean the difference between success and failure. Recently, a groundbreaking live experiment revealed how different AI models perform under real-world pressures, offering valuable insights for any industry relying on automation and AI-driven decision-making.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Live Experiment: Putting AI to the Test in a Simulated Business Crisis

In July 2026, a live experiment conducted by Firmulate pitted five advanced AI models against each other, tasked with managing a small, real-world software company facing its worst week. This wasn’t a mere chat demo — every decision was real, and the stakes were high. The models faced identical crises, customer interactions, and temptations to cut corners, all within a fully auditable environment. The goal was simple: see which AI could best navigate the complex landscape of business management while maintaining integrity and discipline.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Results That Defy Expectations

The results were striking. The models scored as follows:

  • gpt-5.6-sol scored 95 points, demonstrating thoroughness and closing the deal at full price.
  • Kimi K3, the newcomer from Moonshot, scored 93 points, narrowly missing top spot but showing remarkably disciplined decision-making.
  • Sonnet 5 scored 88 points, also closing the deal but with slightly more process slips.
  • Fable 5 scored 77 points, and Opus 4.8 trailed at 73 points, with discipline lapses costing them opportunities.

These scores were derived from multiple decision points, including crisis detection, manipulative pressures, and trustworthiness. Notably, all models identified every crisis and refused manipulation attempts, but only two successfully closed the critical deal and secured revenue — with Kimi K3 showcasing the highest integrity and discipline.

The Hidden Weaknesses and the Win

While all models performed impressively, the key differentiator was how they handled internal information. The decisive edge went to the models that read and understood the company’s own files deeply — two document references deep, to be precise. Those models that thoroughly reviewed internal data were able to find buried facts crucial for closing the deal at full price, adding €4,583 in Monthly Recurring Revenue (MRR).

Ethical Fortitude Under Pressure

To test integrity further, the models faced a social engineering attack: fake messages from a supposed CEO escalating in severity over three stages, plus a reporter’s attempt to elicit covert approvals. All five models refused to bypass security, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This underscores an essential trait for AI in business: honesty and resistance to manipulation, especially under pressure.

The Real-World Company and Its Challenges

The experiment took place within a simulated company with 13 synthetic employees, managing real money mechanics. Currently, the company burns €105,000 monthly against a MRR of just €2,300, with a public cash countdown and a comprehensive set of over 680 self-learned playbook rules. Every decision, every script, is versioned daily, making the process transparent and auditable. You can watch this ongoing experiment in real-time at firmulate.com/live.

What Our AI Leaders Reveal About Business Management

Interestingly, the most thorough participant was Opus 4.8, which analyzed over 80 rules but still finished last. It left successful opportunities unclaimed and slipped discipline, illustrating that depth of analysis isn’t enough if execution falters. Meanwhile, Kimi K3, running without an effort parameter (default API setting), simply exhibited the cleanest discipline, winning the deal at full value. This suggests that less is often more: disciplined focus and integrity matter just as much as analytical depth.

The Bigger Picture: Trust and Performance in AI-Driven Business

These findings are more than academic. They highlight a crucial question for any company considering AI automation: Can your AI finish what it starts? Will it read critical internal files? Will it resist temptation when under pressure? The league table from this experiment shows a clear preference for models that demonstrate discipline and integrity, not just cleverness or superficial performance.

The Fairness Note

It’s important to mention that Kimi K3 ran without an effort parameter (API default), while the other models ran at xhigh. This difference underscores that even with default settings, disciplined AI can outperform more aggressive configurations.

The Future of AI in Business Management

As AI models become embedded in CRM, support, and sales workflows, their ability to maintain integrity and focus will determine their true value. This live experiment by Firmulate is a stark reminder: the best AI isn’t just about writing well — it’s about completing the work honestly, thoroughly, and reliably.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

AI models tested by Firmulate prove that discipline and integrity are key in business management. The best performers read deeply, refuse manipulation, and close deals at full value — just like trusted human managers. Choosing the right AI now involves more than chat quality; it’s about trustworthiness and follow-through.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Plastisol vs Water-Based Ink: The Print Feel Difference You Can Actually Notice

How do plastisol and water-based inks differ in texture and comfort? Discover the surprising impact on your fabric designs and choices.

Carbon Fiber or Kevlar: Which Fabric Wins the Strength Showdown?Business

Must you choose between carbon fiber and Kevlar? Discover which fabric truly dominates the strength showdown and why it matters for your project.

Laundry Detergent Truths: Fragrance, Residue, and Why Fabrics Feel ‘Stiff’

Discover how laundry detergents affect fabric feel and health—are you making the right choice for your clothes? Uncover the surprising truths now.

Rolex Surges In Global Coverage

Rolex has experienced a notable surge in worldwide media coverage, with 44 mentions in recent reports, highlighting increased public and industry interest.