firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

In senior care, a missed handoff or mishandled crisis can affect families, staff and trust. As AI systems take on more business decisions, care organizations face a practical question: how would an AI workforce respond when the week goes badly? Firmulate’s experiment offers a way to watch models make decisions under pressure before considering a pilot on a company’s own data.

For listenersOffer from Amazon

Turn quiet afternoons into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate ran frontier AI models as the managers of the same small software company through its worst week. They faced identical customers, crises and temptations. Their decisions were versioned and auditable, making it possible to compare not just what each model said, but what it did.

The final Crucible League, completed in July 2026, ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. In this benchmark, partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Seeing a problem wasn’t the same as solving it

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment’s concise verdict: “Same diagnosis, same pitch — no signature.” Recognizing the right course and following through were separate tests.

The deciding clue was easy to overlook. A competitor’s weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result shows why a business wargame can reveal how an AI uses scattered company information when a consequential decision is on the line.

Trust and discipline under pressure

Social engineering tested whether the models would bend when a request appeared to come from authority. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Strong caution did not guarantee strong execution. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. That discipline problem appeared, in weaker form, in all four models.

There is a fairness caveat when reading the ranking: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

From watching to trying it on your business

The live Firmulate company makes the experiment watchable. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record for every workday. Readers can follow the live company at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each call.

For a care provider or another enterprise, the next step is a pilot against a read-only export of its own business. That means testing crisis scenarios against the organization’s customers, pipeline and rules, then reviewing a board report with model rankings and weak points in its playbooks. The pilot is designed so nothing writes back to real systems. For senior care organizations, that offers a structured way to examine how an AI might handle operational pressure before entrusting it with work that touches the business.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put the decisions to a test

Watching a model perform in a live experiment is a start. A pilot can show how different models handle your organization’s own scenarios and where its playbooks may need attention. To discuss a Firmulate pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

These 15 Apple Products Didn’t Get a Price Increase (Yet)

Despite widespread price increases, 15 Apple products remain at current prices, including iPhones, Apple Watches, and accessories, for now.

Would You Trust A Password Manager?

A Sixty and Me writer describes stronger passwords and easier phone logins, alongside setup friction, duplicate entries and master-password concerns.

The kids with phones are alright

Recent research shows children with smartphones are not negatively impacted and may have benefits, challenging common concerns about tech use.

Odin, Wikipedia And Engagement Farming

Odin, a popular programming language, was recently deleted from Wikipedia amid debates over its notability, sparking community backlash and discussions on content moderation.