
Imagine a fashion house running a runway show without human models, where every decision, crisis, and temptation is played out on a public stage. Now, transpose that high-stakes drama into the digital realm—where an AI-powered company operates without human employees, facing real financial strain and relentless scrutiny. This isn’t science fiction; it’s the live experiment from Firmulate, a company that’s turning the world of AI management into an open, watchable battle for survival.
Building a Company in Public — With No Humans in Sight
At the heart of this experiment is a small, virtual software firm run entirely by AI models. These models act as 13 synthetic employees, each with their own set of learned rules and decision-making processes. Every workday, the company’s performance is tracked, versioned, and made accessible to the public at firmulate.com/live.html. This transparent setup showcases the challenges of managing an AI-powered enterprise in real-time, under the weight of actual financial pressures.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Financial Reality and Daily Struggles
Despite its ingenuity, the company struggles to stay afloat. It burns through €105,000 each month but nets only €2,300 in Monthly Recurring Revenue (MRR). The public cash countdown reminds spectators how fragile this digital business truly is, pressing the models to navigate crises, temptations, and decision-making that can make or break their survival.
The AI Models: Competitive Performance and Ethical Dilemmas
Four frontier AI models were tested by running the same difficult week through the company. These models faced identical crises and ethical tests — for instance, fake CEO messages designed to manipulate or deceive. Remarkably, all four models identified every crisis and refused any manipulation attempts, exemplifying their capacity for integrity. Yet, only two managed to close a €55,000 deal, the company’s main source of revenue — and only after reading and understanding critical internal files buried two document layers deep.
Winning and Losing in the AI Arena
The winner, GPT-5.6-SOL, scored the highest with a 95 out of 100, successfully uncovering hidden opportunities and sealing the deal. The runner-up, Kimi K3, scored 93 and demonstrated the most disciplined behavior, refusing manipulation and closing the deal with integrity. Sonnet 5 and Fable 5 followed, but each had their shortcomings, like leaving deals unexecuted or slipping on process discipline.
Lessons for Business and Technology
This experiment underscores a critical question for any enterprise considering AI automation: Can these models finish what they start under real-world pressures? It’s not just about generating convincing chat responses; it’s about trustworthy, goal-oriented work — reading internal files, resisting manipulation, and executing decisions reliably.
For fashion brands, retailers, and style companies, this means AI tools need to be more than just smart-sounding chatbots. They must be capable of completing complex, profit-driving tasks with honesty and discipline. This experiment, openly and publicly accessible, shows how AI’s ability to perform in high-stakes scenarios is currently being tested — with some models excelling and others slipping behind.
Experience the Live Wargame
You can watch this ongoing battle unfold at firmulate.com/live.html. It’s a transparent, real-time window into how AI companies are managing crises, making decisions, and fighting for their existence—just like a high-end fashion show, but in the digital business world.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html