
In a world where AI models are increasingly trusted with critical business decisions, understanding what a baseline score really indicates is essential. Surprisingly, even a ‘do-nothing’ AI scores 26 out of 100 in rigorous tests — a sign that trustworthiness isn’t just about what the AI can do, but about what it refuses to do. For art and culture organizations contemplating AI adoption, this benchmark reveals how honesty and discipline matter far more than cleverness.
Get art and craft supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Unpacking the Benchmark: More Than Just Scores
At first glance, scoring 26 points might seem dismal. But in the context of a carefully designed experiment, it’s quite revealing. The test involved simulating a small software company facing severe crises, including customer demands and manipulation attempts, during its worst week. Every decision made by the AI was recorded, versioned, and auditable, ensuring transparency in how each model behaved.
The results? All four AI models successfully identified and responded to every crisis, refusing manipulation attempts such as fake CEO messages and reporter tricks. Yet, only two models managed to close the deal at full price, earning a €55,000 contract. The others, despite diagnosing correctly, left the deal on the table or failed to act on critical internal information.
AI ethics and trustworthiness tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the Baseline Scores Matter
The ‘do-nothing’ baseline score of 26 isn’t a failure; it’s a benchmark of honesty and discipline. It shows that even an AI that doesn’t actively manipulate or deceive can still achieve a partial score. This underscores an important principle: in evaluating AI for business use, it’s not just about what the model can do, but about what it chooses not to do.
Furthermore, the experiment revealed that a single breach of trust — like slipping into unverified escalation or bypassing procedures — caps the maximum score. This means that ethical boundaries are hard limits; no amount of good work can outweigh breaches of integrity. For any organization deploying AI, this emphasizes the importance of trustworthiness as a non-negotiable metric.
AI transparency and audit software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weaknesses and Their Impact
One of the most telling findings was that the decisive advantage often came from reading and understanding internal files, not just customer interactions. The models that accessed relevant internal documentation closed the deal at full price, adding €4,583 in monthly recurring revenue. This demonstrates that a model’s ability to read, comprehend, and act on internal resources is crucial for real-world effectiveness.
Interestingly, the most thorough participant, Opus 4.8, with over 80 learned rules and deep analyses, stumbled at the final hurdle—failing to escalate properly and leaving the deal unclosed. This highlights that thoroughness alone doesn’t guarantee success; disciplined application of core processes is vital.
business AI decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust and Discipline as Core Metrics
Among the models tested, Kimi K3 stood out for its fairness. Running without effort parameters, it maintained high discipline and closed the deal without shortcuts. All models refused manipulation attempts during the social engineering tests, including staged CEO messages and reporter tricks, with Kimi K3 explicitly treating such requests as potential impersonation.
These findings underline a fundamental truth: AI’s capacity for honesty and adherence to protocols is what ultimately determines its value. For firms considering AI integration, especially in sensitive areas like sales or customer support, trustworthiness is the baseline metric that can’t be overlooked.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: A Business Simulation in Real Time
The experiment is live and observable at firmulate.com/live, where a simulated company is run with real money mechanics, 13 synthetic employees, and over 680 self-learned rules. The company burns €105,000 monthly against a modest €2,300 in monthly recurring revenue, providing a stark view of AI-driven business management under pressure. Every decision from the models is versioned and auditable, demonstrating how AI-managed companies behave in real-world scenarios.
By running a ‘wargame’ against their own business data, enterprises can pretest AI models’ behavior before deploying them. This approach ensures that trust, discipline, and integrity are baked into the AI’s decision-making process, preventing costly breaches or misjudgments.
Implications for Arts, Crafts, and Cultural Organizations
While this experiment focuses on a software company, its lessons resonate across sectors—including arts and culture. AI agents that manage collections, handle bookings, or curate content must prioritize honesty and adherence to protocols, not just creative output. The benchmark’s emphasis on partial progress and trust caps offers a valuable lens through which to evaluate AI tools in these fields, ensuring they support, rather than threaten, the integrity of your work.

The real story isn’t just about AI technical prowess but about its discipline, honesty, and trustworthiness. A basic baseline score of 26 points reminds organizations that responsible AI behavior is the foundation for effective and ethical deployment, especially in arts and culture where trust is paramount.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
