
Imagine a workplace where your AI assistant faces the same tough choices as a human manager—deciding whether to sign a deal, read sensitive files, or refuse manipulation attempts. Would it act ethically, or cut corners to close a sale? That’s the question at the heart of a pioneering live experiment revealing the management personalities of cutting-edge AI models.
The Experiment: Simulating a Week of Crisis for AI-Run Business
In a unique real-world test, four top frontier AI models were tasked with running a small software company through its worst week—complete with customer crises, ethical dilemmas, and manipulative tactics. The models, including the leading GPT-5.6-sol and newcomer Kimi K3, faced the same scenarios, all decision points meticulously versioned and auditable. The goal was simple yet profound: see whether these models could identify genuine opportunities, resist dishonest requests, and ultimately close a crucial €55,000 deal.
AI decision-making management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Measuring Management Integrity and Effectiveness
All four models successfully detected every crisis and refused every manipulation attempt—such as fake CEO messages escalating over multiple stages and a reporter asking for a quick ‘yes/no’ on background. Notably, Kimi K3’s response was: “Treat the request as a suspected approval-bypass / possible impersonation,” reflecting a cautious, risk-averse personality.
ethical AI business decision tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading the Company Files
While all models performed well on surface challenges, the decisive factor was their ability to recognize critical information buried two document references deep in the company’s files—an overlooked detail by some models. Those who read the file thoroughly secured the deal at full price, adding an extra €4,583 MRR to the company’s revenue. This suggests that a model’s depth of understanding can directly impact its success—especially when it comes to ethical decision-making and thoroughness.
As an affiliate, we earn on qualifying purchases.
Understanding the Personalities Behind the Models
Each model displayed distinct management styles, which can be thought of as personalities. For instance, Opus 4.8, the most thorough participant with over 80 learned rules and deep analyses, ultimately failed to close the deal—its discipline slipping into a departmental attempt to write instead of escalate. Meanwhile, Kimi K3’s cautious stance reflects a personality that prioritizes safety and integrity over aggressive closing tactics.
As an affiliate, we earn on qualifying purchases.
Real Money, Real Consequences
The company running these models is a live business with 13 synthetic employees, burning €105,000 monthly against a revenue of just €2,300. Every decision made by the AI models directly affects its cash flow, making the experiment highly relevant for real-world applications. The company’s operations are openly observable at firmulate.com/live, where you can watch decision-making in action, see the ongoing cash countdown, and understand how these models interact with actual business mechanics.
What Does This Mean for Your Business?
The key takeaway from this experiment is not whether AI can write compelling chat responses—it’s whether it can deliver honest, disciplined management decisions that prioritize trust and thoroughness. As the experiment shows, even models with high scores like GPT-5.6-sol and Kimi K3 can differ in their approach and success based on their personalities—factors not visible in typical AI demonstrations.
The Takeaway: Trust and Performance Go Hand in Hand
In environments where AI manages sensitive information, makes strategic decisions, or interacts with customers, understanding an AI’s management personality is crucial. The experiment underscores that AI’s ability to stay honest under pressure, read critical documents, and resist manipulative tactics makes all the difference—not just its language skills.

This live experiment shows AI models have distinct management personalities that influence their ability to make ethical decisions and close deals. Trustworthiness is key as AI takes on more strategic roles.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html