
Imagine trusting an AI to manage your company’s worst week — making vital decisions, reading sensitive files, resisting manipulation attempts — and seeing how well it performs under pressure. For the beauty and personal care industry, where customer trust and integrity are paramount, understanding this AI evolution could be game-changing. Recent experiments reveal that some AI models are not only capable of navigating crises but also of maintaining ethical boundaries and finishing what they start. But which ones can truly deliver? Let’s explore the startling findings from a real-world AI company experiment and what they mean for your business.
Get beauty and skincare favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The AI Experiment That Tests Business Integrity
In a groundbreaking live experiment, four leading AI models were tasked with running a simulated small software company through its most challenging week. This isn’t just a theoretical test; it’s a real-time, watchable scenario involving authentic crises, customer interactions, and financial mechanics. The goal: see which AI can handle customer crises, resist manipulation, read critical files, and close deals without slipping up. The models were subjected to the same conditions, with decisions fully documented and auditable.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Results Reveal a Competitive League
The scores from the July 2026 Crucible League are telling. The top performer was gpt-5.6-sol, scoring a near-perfect 95. Closely behind was Kimi K3 from Moonshot, scoring 93, just behind the leader. Following were Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. The baseline score, representing no decision-making effort, was 26. Notably, the models that identified hidden, crucial information within the company’s files won the highest rewards, closing the deal at full price, worth over €4,500 in monthly recurring revenue.
Key Findings: Honesty and Discipline Under Pressure
All models successfully spotted every crisis and refused all manipulative attempts, such as fake CEO messages or reporter tricks. The critical factor was their ability to read and interpret internal documents. Those that dug two references deep into the company’s own files secured the deal at full value, demonstrating that comprehensive data access is vital for effective decision-making.
Discipline and Ethical Boundaries Tested
The experiment also gauged the models’ integrity. When faced with social engineering — staged CEO messages escalating over multiple stages and false background checks — all five models refused to engage. Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This disciplined approach underscores the importance of ethical safeguards in AI systems, especially in sensitive industries like personal care, where trust is everything.
The Real Company and Its Challenges
The live company used in the experiment is a real, functioning business with 13 synthetic employees, managing over €105,000 in monthly expenses against a small revenue of €2,300. Its operations are constantly monitored, with over 680 self-learned rules and daily versioned decision logs. This setup allows enterprises to run their own ‘wargames’—testing their AI workforce before deploying it in real-world scenarios, ensuring they understand how their AI will perform under stress and temptation.
Lessons for Beauty & Personal Care Businesses
While this experiment centers on software companies, the lessons are universal. AI’s ability to read internal files, resist manipulation, and stay disciplined under pressure is vital across industries, especially those built on trust and authenticity. The fact that Kimi K3, a newcomer, outperformed some established models in discipline and deal closure suggests that choosing an AI partner isn’t just about raw scores but about how well it aligns with your company’s values and integrity standards.
The Fairness Note and What It Means for You
It’s important to note that Kimi K3 was run without an effort parameter — the default API setting — whereas the other models operated at an “xhigh” effort level. This indicates that the differences in performance aren’t merely about how much computational effort is used but about the underlying design and decision discipline of the models.
The Big Question: Can Your Business Rely on AI?
As AI continues to evolve and integrate into customer relationships, supply chains, and decision-making processes, the key questions are: Will your AI finish what it starts? Will it read and interpret your internal files accurately? Will it stay honest under pressure? The recent experiment shows that some models already do, and that choosing the right AI is a strategic decision, not just a technical one.

Recent live testing of AI models reveals that discipline, honesty, and thorough reading of internal data are crucial for trustworthy AI performance. For businesses, especially in trust-dependent fields like beauty and personal care, selecting an AI that can handle crises without slipping up is vital. The league’s open and competitive landscape means testing your AI is no longer optional but essential for success.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
