What it costs (order of magnitude)
about €300 a month
servers and agent subscriptions included.
No usage-based cost: agents run on fixed-price subscriptions.
An eight-step chain. Agents at every station, a human at the three places that matter.
Scroll down: a ticket moves from station to station, and each station comes alive when you reach its text.
Step 1 of 8
Requirements written so they can be tested, acceptance criteria, architecture decisions, mockups.
Station 1 · Specification
Step 2 of 8
Three agents hunt for contradictions in the specification before anything starts. 59 findings on Merkindium, all resolved before the first ticket.
Station 2 · Pre-flight
Step 3 of 8
The project-manager agent splits it into batches and tickets, with their dependencies.
Station 3 · Splitting
Step 4 of 8
An agent takes a ticket and proposes a change (PR).
Station 4 · Workbench
Step 5 of 8
Every PR runs the automated tests, plus a real start of the application in a container.
Station 5 · Test bench
Step 6 of 8
If everything is green, the PR is merged. Otherwise a fixer agent takes it back.
Station 6 · Supervision
Step 7 of 8
The new version deploys itself to a private copy.
Station 7 · Staging showcase
Step 8 of 8
In Recette: Verified, Bug or Skip. A bug becomes a new ticket.
Station 8 · Human verification station
Three things that stay with the human.
the specification, the trade-offs when it contradicts itself, what goes to production.
payment keys, passwords, server access. Agents see none of them.
in staging, starting with what the agents doubt.
Five protections built into the chain.
ci-ok)Nothing is merged without green tests.
Each one has a daily consumption cap.
Checked automatically.
The factory builds without holding any key.
In dependency order.
GitHub for tickets and tests, two families of agents (Claude Code and Codex), twelve build machines on servers rented from OVH, Coolify for deployments.
about €300 a month
servers and agent subscriptions included.
No usage-based cost: agents run on fixed-price subscriptions.
Subscriptions have caps. When they are reached, the factory slows down or stops, and the project-manager agent alone uses a large share.
They rarely check in a real browser. They can ship code that passes the tests and is still wrong on screen. 121 of their confidence scores are 6/10 or lower, and they say so.
A real share of the work goes into fixing tests that fail at random: 40 PRs out of 440 in Merkindium.
A ticket can stall several times: the one on the home page took five attempts. When the main version breaks, the factory can stay stuck until I decide.
Deciding, checking, unblocking: that time cannot be delegated.
Collected from the GitHub API on 10 October 2026; regenerated by script.
agent PRs, out of 1,486 merged PRs across the 5 applications shown
A "PR" (pull request) is a proposed change that is tested, then merged into the software.
closed issues (tickets, bugs, dropped ones included)
average of the "Confidence: N/10" scores the agents report about themselves
agent PRs that state a confidence score
scores of 6/10 or lower (12%), and only 3 scores of 10/10
from 19 July 2026 to 10 October 2026; first repository on 19 July 2026
121 scores of 6/10 or lower (12%)3 scores of 10/10
| Application | Agent PRs / merged PRs | Closed issues | Average confidence |
|---|---|---|---|
| Chartrium | 621 / 640 | 720 | 7.5 |
| Merkindium | 440 / 445 | 525 | 7.4 |
| Bottrading | 201 / 202 | 216 | 7.7 |
| Recette | 179 / 183 | 194 | 7.6 |
| Raccourci | 16 / 16 | 26 | no score |
| Total | 1,457 / 1,486 | 1,681 | 7.5 |
A PR counts as "built by an agent" when it comes from a ticket branch (ticket/…). All of them are opened under my GitHub account, which the factory uses. The confidence score is self-reported by the agent. The rule came in along the way: Raccourci did not have it, nor did some older Chartrium and Merkindium PRs. Collected from the GitHub API on 10 October 2026; regenerated by script.
5 applications shown.