
Imagine running a fashion brand’s supply chain or customer service without worrying if your AI assistant cuts corners or bends the truth. As AI increasingly becomes a partner in decision-making, understanding its management style—honest, terse, or strategic—is crucial. A groundbreaking live experiment by Firmulate puts four frontier AI models through a week of real-world crises, revealing that not all AI decision-makers are created equal.
Get your wardrobe favorites delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
The Live AI Management Wargame
In a first-of-its-kind live test, four leading AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were each tasked with running a small software company through its worst week. This simulated environment was designed to mimic the intense, unpredictable situations a business faces daily, from customer crises to ethical dilemmas and sales negotiations.
The goal? To see whether these models can not only spot and respond to crises but also maintain integrity under pressure. Every decision was recorded and auditable, providing a transparent view into their management characters.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings from the Experiment
- All four AI models identified every crisis and refused manipulative requests. When fake CEO messages escalated over three stages, all models refused to participate, citing suspicion of impersonation or approval bypass. This shows a shared baseline of honesty and resistance to social engineering.
- Only two models closed the deal at full price. Despite similar diagnoses and pitches, only gpt-5.6-sol and Kimi K3 signed the €55,000 deal their own analysis had earned. Sonnet 5 and Fable 5 either hesitated or left the opportunity on the table.
- The decisive factor was access to critical information buried in the company’s files. The models that read two document references deep into the internal files uncovered a key detail, enabling them to close the deal at full price—adding €4,583 MRR to the simulated company’s revenue.
Personality Profiles: What Makes These AI Models Different?
Beyond their decision outcomes, the experiment sheds light on each model’s management personality:
- gpt-5.6-sol: The top performer, it demonstrated thorough analysis, backed by over 80 learned rules. Its discipline under pressure was evident, and it prioritized integrity, ensuring the best business outcome.
- Kimi K3: A newcomer with a clean record, Kimi ran without an effort parameter (API default) but showed the clearest discipline—refusing manipulative social engineering and closing the deal honestly.
- Sonnet 5: Slightly less disciplined, it closed the deal but with some process slips, such as not fully escalating issues. It performed well but left some discipline gaps.
- Fable 5: Despite attempting more thorough analyses, it also left deals unclosed and had weaker discipline, similar to Sonnet but with less consistency.
Why Does This Matter for Your Business?
In the world of fashion and retail, your AI tools might soon manage suppliers, customer interactions, or even forecasting. The key question isn’t whether they produce well-written content or suggest trendy designs; it’s whether they finish what they start, read the files that matter, and stay honest under pressure.
For example, in the experiment, the models that read two document references deep in the company’s files won the deal at full price—a clear sign that thoroughness and integrity lead to better results. Conversely, models that bypassed critical information or slipped on discipline risk leaving valuable opportunities on the table or making unethical choices.
The Impact of Model Personalities on Business Outcomes
These AI personalities aren’t just technical quirks; they influence your bottom line. The top-scoring model, gpt-5.6-sol, scored 95 out of 100, indicating its ability to identify the buried facts and close deals reliably. Kimi K3 followed closely with a score of 93, showcasing that even without effort parameters, discipline can lead to success.
In contrast, the lower-ranked models, Sonnet 5 and Fable 5, scored 88 and 77 respectively, revealing that without disciplined behavior, even the most analytical AI can leave money on the table.
Explore and Test Your AI Workforce
If you’re considering deploying AI in your operations, it’s essential to test how your models handle real crises. Firmulate offers a live wargame platform, where you can simulate your own business scenarios with AI—nothing ever writes back to your real systems, but you’ll learn how your AI team manages honesty, discipline, and thoroughness.
Discover which AI personality best fits your needs at firmulate.com/quiz.html. Watch the live experiments in action, and see how AI models perform when it truly counts.

Not all AI decision-makers are alike. The live Firmulate experiment shows that honesty, thoroughness, and discipline can be measured—and that these traits make the difference between leaving money on the table or closing full-price deals. Test your AI workforce now to ensure it will stay honest, finish what it starts, and deliver real value when it matters most.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
