
In the beauty and personal care industry, trust is everything — especially when it comes to automating customer service, supply chains, and product innovation. But how do we measure whether an AI system is truly reliable, or just good at sounding convincing? A recent public AI benchmark offers striking insights: even a ‘do-nothing’ baseline scores 26 out of 100, highlighting the importance of honesty, discipline, and thoroughness in AI-driven decision-making.
Get beauty and skincare favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Surprising Baseline Score: Why 26?
At first glance, you might expect a system doing nothing to score zero — but in the world of AI benchmarking, even minimal effort counts. In the latest experiment conducted by Firmulate, a real-world company simulation was run through its worst week, with models making decisions on customer crises, internal conflicts, and manipulative tactics. The results? Every model, including the most basic, scored at least 26 points. This baseline underscores that partial progress and cautious decision-making are valued, but it also establishes a clear floor: no AI should be trusted to do less than this minimal standard.
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Integrity Test: Trust and Breaches
One of the key lessons from the experiment is that a single breach of trust caps the entire score. Even if an AI performs well otherwise, compromising its integrity — such as signing off on a fake deal or ignoring internal rules — irreparably damages its overall assessment. For example, the models faced social engineering attempts, with fake CEO messages escalating through staged scenarios. All five models refused to be manipulated, demonstrating a shared commitment to honesty. Notably, this refusal is not just a moral stance; it’s a measurable indicator of reliability that can make or break trustworthiness in real business contexts.
Reading Between the Lines: Hidden Risks and Opportunities
Perhaps most revealing was that the decisive edge came from the models’ ability to access and interpret internal documents. Two references deep in the company’s files, rather than just surface-level customer interactions, contained the crucial information needed to close a lucrative deal. Models that read these files won at full price — worth over €4,500 monthly recurring revenue — while others missed the opportunity. This highlights a vital point: thorough information processing can be the differentiator between a good AI and a great one, especially in industries where understanding internal data is key to unlocking value.
The Discipline of Honesty and Process
Among the participants, Opus 4.8 stood out for its depth but ultimately placed last. Despite analyzing more rules and conducting comprehensive assessments, it failed to close the deal, leaving it on the table due to lapses in discipline—such as writing attempts into a locked department instead of escalating issues. This underscores that thoroughness alone isn’t enough; discipline, process adherence, and the willingness to escalate are essential traits for AI systems entrusted with business-critical decisions.
Implications for Business Leaders
For companies in beauty, personal care, or any industry, the message is clear: trust in AI isn’t just about how well it generates content or handles routine tasks. It’s about whether the AI can complete its work reliably, read and understand internal data, and resist manipulation under pressure. That’s why measuring these qualities through real, auditable experiments — like the one conducted by Firmulate — is essential for making informed decisions about deploying AI at scale.

The recent AI benchmarking by Firmulate reveals that even a passive baseline scores 26 — emphasizing that honesty, thoroughness, and discipline are non-negotiable for trustworthy AI. For beauty brands and personal care companies, this means prioritizing AI systems that can read internal data, resist manipulation, and complete tasks reliably — not just look impressive in demos.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
