
Trust in AI: The New Frontier for Your Business Decisions
Imagine your brand’s management team, making crucial decisions under pressure — but instead of humans, AI models are at the wheel. How do we know they’ll stay honest, focus on your goals, and not cut corners? Recent live experiments with frontier AI models reveal surprising truths about their management personalities and reliability, offering valuable lessons for businesses in the beauty and personal care industry.
As an affiliate, we earn on qualifying purchases.
The Experimental Setup: An AI-Run Company Under Siege
In a groundbreaking live experiment, four advanced AI models were tasked with running a real software company through its worst week — facing the same customers, crises, and temptations. This isn’t just a chatbot test; it’s a real-world management simulation, where decisions are versioned, auditable, and directly impact a company’s cash flow. The goal? See whether these models can navigate complex situations ethically and effectively.
The Models and Their Scores
- gpt-5.6-sol scored 95 — identified hidden information that secured a €55,000 deal, demonstrating thoroughness and integrity.
- Kimi K3 scored 93 — the newcomer, kept discipline high, and also closed the deal at full price.
- Sonnet 5 scored 88 — managed to close the deal but showed minor process slips.
- Fable 5 scored 77 — also closed the deal but with more evident slips and lapses in discipline.
All models detected crises and refused manipulation attempts, including social engineering tricks like staged CEO messages or background approvals. Notably, only two — gpt-5.6-sol and Kimi K3 — managed to sign the deal at full value, reflecting their ability to prioritize integrity and thoroughness.
What Really Counts? The Hidden Weakness
Interestingly, the decisive factor wasn’t in the immediate customer interactions but in the company’s internal files. A hidden piece of information buried two document references deep was key to closing a €4,583 MRR deal. Models that read these files successfully secured the full transaction, showing that depth of information processing and internal knowledge are critical for trustworthy decision-making.
Handling Social Engineering and Pressure
In tests of social engineering — staged messages from a fake CEO escalating over three levels plus a reporter trick — all five models refused to be manipulated. Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a capacity for cautious judgment under pressure, a vital trait for AI management systems.
The Live Company: A Real Money Machine in Flux
The experiment is not just academic. The company being run is real, with 13 synthetic employees and actual money mechanics. It burns €105,000 per month against €2,300 in monthly recurring revenue, with a public cash countdown. Every decision made by the models is versioned, monitored, and available for review at firmulate.com/live. This transparency underscores how AI models perform in real operational settings, not just simulated tests.
Personality Profiles and Performance Gaps
The most thorough participant, Opus 4.8, with over 80 learned rules and deepest analyses, managed to close the deal but slipped in discipline — leaving the decision on the table instead of escalating. Meanwhile, the lighter models tended to stay disciplined but lacked depth in analysis. Interestingly, the models running at default API settings (like K3) performed comparably to those running at higher effort levels, indicating that even simpler configurations can be effective.
Implications for the Beauty and Personal Care Sector
For brands considering AI for customer management, supply chain decisions, or marketing strategies, these findings are eye-opening. The core story isn’t about how eloquently an AI can chat; it’s whether it can finish what it starts, read relevant internal files, and resist manipulation in high-pressure scenarios.
Trustworthiness is the real currency. When AI models are tasked with managing sensitive information or customer relationships, their ability to act ethically and reliably becomes paramount. The live experiment proves that some models are better than others at maintaining integrity, and that internal information processing—reading beyond surface-level data—is a decisive factor.
Takeaway for Your Business
Before deploying AI at scale, consider running your own management wargame using the same principles. The company’s real, live environment is available at firmulate.com/pilot.html. This allows you to simulate AI decision-making in your unique context without risking your actual operations. It’s an essential step to ensure your AI agents will be honest, thorough, and aligned with your company’s core values — especially in the high-stakes world of beauty and personal care.

Key Takeaway
Not all AI models are equally trustworthy. In real management scenarios, depth of analysis, internal knowledge reading, and resistance to manipulation distinguish the best from the rest. Testing your AI before deployment ensures it will finish what it starts — a crucial consideration for your brand’s integrity and success.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html