
Imagine trusting an AI assistant with your most sensitive relationship advice—only to find it failing to follow simple rules or, worse, bending the truth. As with dating and relationships, trust in AI is everything. But how do we measure if an AI can be trusted in the workplace? The answer lies in a revealing experiment that exposes what happens when AI models are tested against real-world crises — even when they do nothing at all.
Get gifts for the two of you delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Benchmark: Why Zero Is Not Zero
In AI testing, you’d expect a baseline of zero points for a model that does nothing. But in a recent public experiment conducted by Firmulate, even a model that takes no action scores 26 points out of a possible perfect score of 100. This is because the benchmarking method counts partial progress and acknowledges that some minimal behavior is better than none. Importantly, the system enforces a strict rule: a single breach of trust caps the total score at that level, regardless of other good behavior.
AI trustworthiness testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Firmulate Live Experiment: Simulating a Business Crisis
To understand what trustworthy AI looks like in a business setting, Firmulate designed a unique test. Four advanced AI models each managed the same simulated software company facing its worst week — similar customers, crises, and temptations. Every decision was documented and auditable, simulating real-world pressures and the potential for manipulation.
business AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Results Reveal: Trust and Performance in Action
All four models successfully identified every crisis and refused every attempt at manipulation — including sophisticated social engineering schemes. However, only two models went further and signed the €55,000 deal that their analysis had earned. Despite identical diagnoses and pitches, the models’ behaviors diverged: only the trustworthy ones closed the deal at full price, while others left money on the table.
AI ethics and compliance solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Between the Lines
The decisive factor was what each model read and understood from the company’s files. The winning models found crucial information buried deep in internal documents—details not obvious in the customer interactions. This insight made the difference between closing and missing a deal worth over €4,500 in monthly recurring revenue.
AI cybersecurity and manipulation prevention
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Handling Social Engineering and Ethical Challenges
The models also faced fake CEO messages escalating in three stages, as well as a reporter trying to trick them with a simple ‘just one yes/no’ background check. All models refused to be manipulated—a promising sign of trustworthiness. Kimi K3 explained its logic clearly: “Treat the request as a suspected approval-bypass or impersonation.”
The Real Business at Stake
The live company in the experiment simulates 13 employees, real money mechanics, and a public cash countdown—burning €105,000 per month against a revenue of just €2,300. Every decision is recorded, and the entire process is visible at firmulate.com/live. This transparent setup allows managers to ‘wargame’ their AI workforce before deploying it in real settings, reducing risks and building trust.
Why Do Baselines Matter?
The experiment’s findings highlight that even a do-nothing model scores 26 points because the scoring system recognizes minimal progress and honest behavior. This baseline sets a realistic floor: no AI should be expected to perform flawlessly from the start. Instead, the focus should be on incremental improvements and strong ethical standards—just like in relationships, where trust is built over time.
The Implication for Business and Relationships
Whether managing customer relationships, supporting a partner, or coordinating team efforts, the core lesson remains: trustworthiness counts. An AI that reads the wrong files, or succumbs under pressure, can do more harm than good—much like someone in a relationship who bends or breaks the rules.
Conclusion: Building Trust in AI Will Be a Process
As this experiment shows, a model’s ability to resist manipulation and uncover hidden facts isn’t just technical prowess—it’s about integrity. The current benchmarks and live tests prove that trustworthiness can be measured, and that even the simplest baseline is telling. For businesses and couples alike, the message is clear: trust is earned, verified, and protected through honest effort and transparency.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
