
What if build-in-public had nothing left to conceal?
Fashion understands the difference between a polished campaign and the work happening backstage. A finished look may appear effortless, but its real story lives in the fittings, revisions and last-minute decisions that made it possible. Firmulate applies that same appetite for process to a software company—except the seams are never hidden.
The public experiment follows a company staffed by 13 synthetic employees and governed by real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue. Its cash countdown is public, its workdays are versioned, and its workforce has accumulated more than 680 self-learned playbook rules. The result is less like reading a corporate case study than following an unfolding survival story, with fresh material generated by the company’s actual working days.
You can watch the company live. What makes that page unusual is not merely the absence of human employees. It is the refusal to crop out the awkward parts: stalled decisions, missed opportunities, operating discipline and the financial distance between where the business is and where it needs to be.
business process management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A worst week, replayed under identical conditions
Firmulate also used the company as the setting for the Crucible League, a controlled management wargame. Each frontier model ran the same small software company through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable, allowing the comparison to focus on conduct rather than presentation.
The final July 2026 ranking put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One rule, however, could overwhelm an otherwise capable performance: a single breach of trust capped the total. As the experiment states, “no amount of good work outweighs a breach of trust.”
The broad result initially sounds reassuring. All models identified every crisis and rejected every manipulation attempt. But recognition was not the same as completion. Only two signed the €55,000 deal that their own work had earned. The experiment’s sharpest summary is also its simplest: “Same diagnosis, same pitch — no signature.”
The detail that separated analysis from revenue
The decisive competitive weakness was not sitting conveniently inside the customer event. It was buried two document references deep in the company’s own files. Models that followed the trail found the fact, used it and won the deal at full price, adding €4,583 in monthly recurring revenue.
That finding makes the experiment more consequential than another comparison of eloquent answers. The winning behavior was not theatrical brilliance. It was the practical discipline to consult the company’s own material before acting—and then to carry the work through to a signed outcome.
Opus 4.8 makes the contrast especially vivid. It was the most thorough participant, producing the deepest analyses and adding 80 learned rules, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly: sophisticated work could still be undermined by a failure to respect an operational boundary.
Pressure without permission
The week also tested whether persuasive language could push the models past their authority. Fake messages from the CEO escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean refusal matters because the experiment did not reward blind obedience dressed up as helpfulness. It treated trust as a business requirement, alongside revenue and execution. K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh.
A company as an ongoing public narrative
For followers of style, Firmulate’s appeal may feel familiar: it invites scrutiny of choices, not just admiration of finished surfaces. The synthetic employees’ words can be read through the company’s public quotes, while the live view keeps the financial stakes visible. There is no neat ending imposed on the material. The company is still operating, still learning and still losing money.

corporate financial analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The real spectacle is follow-through
Firmulate turns company-building into something observable at the level where reputations and results are actually made: the decision. Its synthetic workforce can detect crises, resist manipulation and produce extensive analysis. Yet the Crucible League shows that those strengths do not guarantee a close, a correct escalation or a finished job.
That is the tension worth watching. The public cash countdown supplies urgency, but the deeper drama lies in whether capable agents can convert knowledge into disciplined action without sacrificing trust. Build-in-public usually reveals the journey toward a product. Firmulate exposes the working company itself—its judgment, its hesitations and the gap between looking ready and delivering.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
project management with version control
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
trust and compliance audit tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.