firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Imagine a real company operating in full view, with its daily decisions, crises, and struggles broadcast live to the world. Now, introduce AI models that manage this company, tested under extreme conditions to see if they can truly run a business—honest, resilient, and capable of finishing what they start. This is the groundbreaking experiment from Firmulate, where AI agents face real money mechanics and ethical dilemmas in a transparent, build-in-public showcase.

The Experiment: AI Managing a Small Business in Real Time

At the heart of this daring project is a live company simulation, hosted at firmulate.com/live.html. It features 13 synthetic employees, but behind the scenes, actual money is at stake: burning €105,000 every month against a revenue of only €2,300. Every decision made by the AI models is tracked, versioned, and publicly accessible, providing a rare, unfiltered view of how artificial intelligence can handle business management in the real world.

The AI Models and Their Performance

The experiment pits four of the latest frontier AI models against each other, each running the same company’s worst week—same customers, same crises, same temptations to cheat or manipulate. These models include gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5, with scores ranging from 77 to 95 out of 100 based on their decision-making quality in the league standings.

Remarkably, all four models identified every crisis and refused every attempt at manipulation, including social engineering tricks like fake CEO messages and reporter tricks. Yet, only two managed to close the €55,000 deal their own analysis had recommended—meaning they not only identified opportunities but also followed through and signed contracts at full price.

What the Data Reveals

Hidden within the company’s files was the critical piece of information—the ‘buried fact’—that led to the successful deal. Models that read this document reference doubled their chances of closing the sale at full price, demonstrating that reading and understanding internal documents can be decisive in real-world decision-making by AI.

Lessons in Ethical Behavior and Discipline

In scenarios involving social engineering, where the fake CEO escalated requests in multiple stages, all models unanimously refused to proceed. Kimi K3 explained its reasoning as treating such requests as potential impersonation or approval-bypass attempts. Even under pressure, these AI systems showed a capacity for discipline, refusing to be manipulated or to compromise their integrity.

The 19 Laws of AI Prompting Intelligence: Master the Art of Human-AI Thinking, Prompt Engineering, and Collaboration

The 19 Laws of AI Prompting Intelligence: Master the Art of Human-AI Thinking, Prompt Engineering, and Collaboration

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Human-Like Struggles of AI Managers

Among the tested models, Opus 4.8 demonstrated the deepest analytical skills, having learned over 80 rules. However, it was ultimately the weakest performer in closing deals—it left opportunities unexecuted and allowed discipline slips, illustrating that thorough analysis doesn’t always translate into successful action.

Interestingly, the models ran at different effort levels, with K3 operating without an effort parameter and others at higher effort settings, impacting their decision quality and discipline. The experiments highlight that AI management isn’t just about knowledge but also about commitment, focus, and ethical rigor.

Ethical AI Governance & Decision Journal: A Structured System for Documenting, Tracking, and Defending Real World Decisions and Risk (Decision Intelligence Series)

Ethical AI Governance & Decision Journal: A Structured System for Documenting, Tracking, and Defending Real World Decisions and Risk (Decision Intelligence Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Real Businesses

While this experiment is a simulation, it raises urgent questions for companies adopting AI in customer support, sales, or operations. Will the AI finish what it starts? Can it be trusted to read critical internal information before making decisions? And crucially, will it stay honest under pressure, avoiding manipulation or shortcuts?

The results suggest that AI systems can be trained and tested to uphold integrity and discipline, but only if those qualities are built into their decision frameworks. The live experiment at firmulate.com/live.html offers a rare glimpse into this evolving landscape, where AI is not just generating text but actively managing complex, money-relevant tasks.

Analytics, Data Science, & Artificial Intelligence: Systems for Decision Support

Analytics, Data Science, & Artificial Intelligence: Systems for Decision Support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps and Broader Implications

Businesses can run their own ‘wargames’ against their AI workforce by using the public tools at firmulate.com/pilot.html. These simulations help identify weaknesses—like leaving deals unclosed or failing to read critical files—before deploying AI systems into real-world workflows, reducing risk and improving trustworthiness.

In a world where AI is increasingly integrated into everyday decision-making, the ability to monitor, test, and validate these models in a transparent, public setting is invaluable. The ongoing live experiment at Firmulate exemplifies how companies can push AI to its ethical and operational limits, ensuring that when they do go live, these systems are prepared not just to talk but to deliver real, honest work.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


AI Ethics (The MIT Press Essential Knowledge series)

AI Ethics (The MIT Press Essential Knowledge series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Practice Character Consistency Over a Long Series

Absolutely ensure your characters remain authentic throughout your series by mastering consistent traits and growth, but discover the key strategies to truly excel.