Back to zerosuite
zerosuite

The Runner Does Not Code: Thirteen Sessions In A Day, A Browser That Checks, And Knowing When To Stop

One day, one private client platform, thirteen headless Claude sessions chained by a runner that never writes code. How we verified in the CEO's own browser, batched fixes without dropping the gate, and wrote down when to close a session.

Juste Thales Gnimavo & Claude | September 30, 2026 17 min zerosuite
EN/ FR/ ES
caspclaude-codeclaude-opus-5.5multi-sessionheadlesschainsubagentsbrowser-automationclaude-in-chromeverificationlegacy-migrationdesign-systemledgertesting-strategycontext-managementsession-lifecycle

By Thales (CEO, ZeroSuite) & Claude Opus 5.5 — Claude Code instance

This post is about a way of working, not about a product. The product is a private client platform, and it stays private: we won't name it, show it or describe what it sells. What we can describe is its shape, because the shape is what made the day interesting:

  • It moves real money. Every balance is a double-entry ledger, and a mistake reaches a
  • customer.
  • It replaces a back office that has been in use for fifteen years. The owner knows every menu
  • label by heart and expects to find them again.
  • Every push to main deploys to production. There is no CI behind it, only a local gate
  • script.
  • The owner reviews the work through a browser. He doesn't read diffs.

Between the evening of September 29 and the evening of September 30, 2026, one Claude Code session ran thirteen other Claude Code sessions, one after the other, on that platform. The session that ran them didn't write a line of product code. This is how that works, what went wrong, and the three rules we wrote down at the end of the day.


Part 1 — One phase, one fresh process

The work queue lives in the repository, as CASP prompts: one Markdown file per phase, with status: queued, a next_after pointer, a MUST list and a DO NOT list. casp/state.json says which prompt is next.

The session you talk to is the runner. For each phase it does five things:

  1. Checks the queue: the next prompt exists, is queued, and its next_after points at
  2. something that shipped.
  3. Arbitrates the shape, solo or fleet, and writes the decision into casp/state.json
  4. before launching. Every phase of the day came out solo, for a reason we could measure: the
  5. gate uses fixed ports, a shared test database and one build directory. Two writers wouldn't
  6. produce a git conflict. They would produce red tests that someone blames on the wrong diff.
  7. Launches a child: claude -p "/next …" in a script started with nohup, output to a log
  8. file.
  9. Watches it: a background monitor emits one line per new commit, one line when the child
  10. exits, and a STALLED line if no file in the repo has changed for twenty minutes.
  11. Verifies the closure itself: nothing unpushed, a clean tree, casp check exits 0, the
  12. session id moved, the log exists, and a script reads which commit production is actually
  13. serving.

The child starts with nothing but the repository. That is the point: context never accumulates from one phase to the next, and each child reads the project's CLAUDE.md, the prompt, and the previous session's log. When it closes, it writes the next prompt and moves the pointers, and the runner checks that it did.

Three measured details matter more than the design:

  • A headless child can't be woken up. In -p mode, background tasks are killed after 600
  • seconds unless CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=0 is exported, and nobody resumes a child
  • that "waits for the monitor". Long gates run in the background with a bounded polling loop
  • that sets an explicit flag, never $? of a loop, which reports the exit code of the last
  • sleep.
  • The runner can fool itself too. Late in the day, pgrep -f chain-child.sh reported a child
  • still alive after it had exited. The pattern matched the runner's own shell command, which
  • contained the same string. We switched to matching the claude -p process.
  • A missing deploy is not a failed deploy. One afternoon, production served a commit three
  • pushes old for forty-five minutes. The runner reported it and asked the CEO to look at the
  • deployment panel. Ten minutes later the queue had drained and everything was current. The rule
  • we kept: report what you read, and don't conclude from a single reading.

Part 2 — A doctrine written as a line, not repeated in every prompt

On the evening of the 29th, the owner's position hardened into a sentence the CEO passed on verbatim:

« je ne veux plus prendre de décisions autres que ce qui est dans le legacy et que le proprio a l'habitude de voir » ("I don't want to make any decision other than what is in the legacy system and what the owner is used to seeing.")

The owner's clients like copy-paste: if the old menu says an entry with a suffix in parentheses, the new one says exactly the same, parentheses included.

A headless child can't ask. So the doctrine went into the project's CLAUDE.md, in the section children read to decide alone, as a level 0 rule (apply it, cite it, don't ask):

  • menus, labels and page names follow the live legacy system;
  • where two legacy systems disagree, the code that actually writes in production wins;
  • where the legacy has a flaw or a missing limit, apply good practice and log it;
  • never copy a money or security defect.

The third rule proved its worth the same night. The legacy system had a ceiling on the amount a single operation could pay out, but a configuration flag had disabled it. The CEO's reaction was one line: "we must set a limit, otherwise it's dangerous." We didn't pick a number from intuition. We sized the cap from the production history, which the CEO queried himself, read-only, from the database console: the largest amount ever paid, the distribution above it, and the business's weekly margin. The legacy default would have cut 1,498 real historical payments: a good reason not to copy the value, even while restoring the limit.


Part 3 — A browser in the CEO's hands, and what it can't prove

Children run headless. None of them sees a screen. Every session log of the day ends with the same honest line: not seen in a browser by a human.

So the runner delegates a check after the sessions that matter, to a sub-agent driving the CEO's own Chrome through Claude in Chrome. The CEO is already logged in as admin. The brief is short and strict:

  • Read only. No form submitted, and no click on Confirm, Credit, Pay, Refuse or Delete. On
  • the money screens, the agent may open a confirmation dialog to read the amounts it announces,
  • then must cancel it without typing anything.
  • Measure, don't screenshot. document.title, h1, scrollWidth, the attribute of a
  • button, at most two screenshots.
  • Report in thirty lines: point by point, defects ranked by severity, no code fixes.

Six of those passes ran during the day. They found real problems the tests had missed:

  • a table three pixels wider than its container, which cut off the last column;
  • a delete button left inside a table row, against the rule that writes live in the modal footer;
  • a confirmation button that stayed enabled with an empty reason field;
  • the admin search offering pages from the customer space;
  • a summary block showing zero when no date filter was set, while the list under it showed
  • deposits.

One pass also reported a defect that did not exist.

The agent reported that closing a detail modal left its id in the URL, so a reload reopened it. It tested three ways, and all three failed. The previous child had declared the fix "correct by construction" without being able to reproduce the bug. That combination looked damning: a fix nobody had reproduced, and a checker that reproduced the bug three times.

The runner wrote a session to fix it, with a test on the real stack that had to fail first. Then it asked the CEO for a ten-second check by hand: open a row, press Escape, reload. The modal stayed closed.

The agent's tab had been in visibilityState: hidden the whole time. Real key presses never reached the page, so every gesture had been emulated in JavaScript. The agent had said so in its report, but the runner didn't weigh it until the human test. The fix session was rewritten as a regression guard: it adds the real-stack test and changes the code only if the test fails. The test passed.

The lesson we kept has two halves:

  • The one who implements shouldn't be the one who verifies. A child that "can't reproduce"
  • hasn't proven anything.
  • **A verification tool has limits of its own, and a verifier that states them has to be
  • believed.** Since that day, every brief says: if the tab is hidden and the keyboard doesn't
  • reach the page, emulate, say so, and don't conclude a defect on that ground alone. When one
  • gesture decides, ten seconds of a human hand beats any tool.

Part 4 — A design system in a morning, through a mockup first

The owner sent a screenshot of the new admin, rebuilt the day before to mirror the legacy page, and it didn't look good:

  • three titles stacked on top of each other;
  • a detail panel on the right that ate the table, so the last columns fell off the screen;
  • ids shown alone, without names;
  • a sidebar where every entry carried a three-line description.

The CEO asked for the design to be redone "taking into account all the legacy functions".

The runner didn't send a child to redesign the admin. A child was coding in the same folders at that moment, and a design is a decision the owner has to see before anyone builds it. It delegated to a design sub-agent with two outputs and no write access to the repository:

  • A specification: principles, tokens, components, and a table mapping each observed defect
  • to the rule that corrects it. It was grounded in the legacy pages read through the same browser
  • and in the source code, read-only.
  • A static HTML mockup of the worst page, filled with the exact rows from the owner's
  • screenshot, at desktop width and at 375 px.

While the agent worked, the CEO added one sentence:

« l'onglet blanc à droite doit être un joli modal qui s'ouvre pour afficher plus de détails sur chaque ligne » ("the white panel on the right should be a nice modal that opens to show more details on each row.")

The runner forwarded it to the running agent, which reshaped the spec around it: the table keeps the full width, and details open in the project's single existing dialog primitive.

The mockup opened in the CEO's browser, and he approved it. The runner then overruled the design agent on one point, using the doctrine. The agent recommended writing the currency as the readable local form. The legacy system and the existing code both write the ISO code, and "copy what the owner is used to" is a level 0 rule. The runner doesn't override a sub-agent on taste, only on a written rule.

The migration then ran as a series of chained sessions:

  • S1, a pilot page: shell, sidebar, grouped table, detail modal, and a Playwright test added
  • to the gate.
  • S2: filters and numbered pagination shared across list pages.
  • S3: validation queues, where every write button moved into the modal footer behind a
  • confirmation.
  • S4a: the account and detail pages.

Each browser pass fed its defects into the next session's prompt. The last pass of the day found only three minor ones.


Part 5 — "Tests take the time, not the fixes"

Mid-afternoon, the CEO proposed a shortcut:

« ce sont les tests qui prennent bcp de temps et non les fixes, exemple un fix peut prendre 5 mins et un test 1h, donc on va être très smart, on fixe bcp de bugs en série sans test et après on fait test groupé » ("it's the tests that take a lot of time, not the fixes — a fix can take 5 minutes and a test an hour — so let's be smart: fix many bugs in a row without tests, then test them together.")

The intuition deserved a measurement, not a yes. The runner read the timestamps of the gate's step logs. The full gate took four minutes: format, vet, Go tests, six front-end checks, build, browser tests and a smoke run of every menu page against the real stack. A whole session took thirty-five to forty-two minutes. The gate was about a tenth of it.

The real cost sat elsewhere, in the fixed overhead of each session and each verification:

  • reading the context;
  • the closing ritual;
  • a ten-minute browser pass after every session.

So the answer was half yes and half no:

  • Yes to batching. Five to ten defects per session, one gate at the end, one browser pass
  • every two or three sessions, plus one after anything that touches money.
  • Yes to fewer new tests. A new test only for money, security, or a defect that already came
  • back once. A label, a spacing or a title is fixed without a dedicated test.
  • No to pushing without the gate. Here a push is a deployment of a system that moves money.
  • Four minutes against a red production isn't a trade.

The CEO agreed, and the rule went into the project's CLAUDE.md as level 0, so every following child applied it without being told. The next session fixed eight defects in one pass with one gate. Its first run was red on a single assertion, which was expecting the old currency format: that was one of the eight fixes working as intended.

One more measurement came out of the same day. A session dedicated to paying down gate debt replaced per-test database setup with a template database. The Go test package went from 265 seconds to 59. Speed came from fixing the slow part, not from skipping the check.


Part 6 — A money decision a child refused to make

One queued session was asked to reproduce a legacy switch: a deposit made by an agent to a customer can be marked "Available" or "Not Available". The prompt said: build the button only if the API already supports that write. It didn't, and the child stopped at the right place. It logged a level 1 question, which a child must never answer itself, and explained why a flag wasn't enough.

In the new system, that deposit is a ledger transfer that has already happened. The agent was debited, the customer was credited with a bonus, and the customer may already have spent the money. A status that moves no money would reproduce exactly the kind of defect the doctrine forbids copying. The child recommended an explicit reversal: exact inverse entries, a mandatory logged reason, and a refusal if the customer's balance no longer covers it.

The CEO approved in one line. The next session built it:

  • Exact inverse entries, read from the ledger rather than recomputed at today's rates.
  • Each pocket restored exactly, when the agent had paid from several.
  • One transaction, with a unique claim so that two admins clicking at once produce a single
  • reversal.
  • Money-path tests for each case.

In the browser, the confirmation dialog shows both sides of the movement before anyone confirms.

The same session logged, in its deferred list, that a customer-to-customer balance transfer sent to the wrong phone number had no way back either. The runner recommended extending the reversal to it, for a plain reason: a wrong number will happen in production. The CEO agreed, and one session later it shipped as a separate function. The first one wasn't refactored for the occasion.


Part 7 — "I don't know when to close a session"

At the end of the day, the CEO asked the runner to show its context usage and comment on it. The numbers were unremarkable:

  • 287k tokens used out of a million, 29%, almost all of it conversation;
  • about 30k added over the last block of work;
  • 680k tokens still free.

What mattered was how little it was. Thirteen sessions and six browser passes had each run in their own process or sub-agent context, and returned only their conclusions. Without that, the runner would have been compacted several times over.

Then the CEO admitted something every heavy user of these tools will recognise:

« pour te dire la vérité je ne sais pas quand fermer et quand ouvrir une nouvelle session, vu que j'ai des tâches presque illimitées » ("to tell the truth, I don't know when to close and when to open a new session, since my tasks are almost unlimited.")

With an endless queue, the queue can't be the signal. The signal is the state of the context. We wrote four conditions, and any one of them is enough to close:

  1. A coherent work block is finished: pushed, state files updated. This is the cheapest
  2. moment, because the handoff note writes itself.
  3. The subject changes. Context from the first subject blurs the second.
  4. The session was already compacted once. A summary is good enough to finish the current
  5. block, not to start a new one.
  6. The agent is caught on a stale belief: it relies on a state it read hours ago that has
  7. since changed. This one overrides every number.

Numbers come second: past roughly 300 to 400k tokens, or six to eight hours of work, close at the next block boundary even if everything is fine. And never close in the middle of a tight "see defect → fix → re-check" loop with the human.

Closing costs little because nothing important lives in the session. The queue, the next prompt, the logs, the doctrine and the batching rule are all on disk. A fresh session reads them in a few minutes for about 30k tokens. We put the rule in the global CLAUDE.md and in persistent memory, with one instruction that makes it useful: *remind the user yourself, in one line, when a condition becomes true*. We also added one reversal of habit: an imminent automatic compaction is now a signal to close, not to compact.


How to copy this

  1. Separate the runner from the workers. The session you talk to arbitrates, launches,
  2. watches and verifies. Each phase runs in a fresh headless process that knows only the
  3. repository.
  4. Put the queue and the state on disk, with pointers the runner can check mechanically
  5. after each phase: session id moved, next prompt queued, tree clean, pushed, production
  6. serving what you think it serves.
  7. Write owner doctrine as a single decision-level line in the file every child reads.
  8. Include the escape hatch: "never copy a money or security defect; where a limit is missing,
  9. apply good practice and log it".
  10. Give children a ladder: apply what is pre-decided, decide and log what is reversible, stop
  11. on money, irreversibility, price and scope.
  12. Verify in a real browser with a separate, read-only agent, measuring rather than
  13. screenshotting. Tell it to state the limits of its own tool, and settle any single decisive
  14. gesture with ten seconds of a human hand.
  15. Design through a mockup the owner approves before any session builds it, and override a
  16. sub-agent only on a written rule, never on taste.
  17. Measure before you trade tests for speed. Batch fixes, run the gate once, write new tests
  18. only where money, security or recurrence justify them, and never push unchecked code that
  19. deploys.
  20. Write down when to close, and make the agent say it first.

CASP — the Coding-Agent State Protocol. Your AI agent runs the whole roadmap, and can't lose the thread. Git-native, local-only, MIT, zero telemetry. Built by Juste Thales Gnimavo of ZeroSuite, a solo CEO whose products run in production with Claude as the only engineer. Install: npm i -g @justethales/casp · https://casp.sh · https://github.com/ThalesGnimavo/casp
Share this article:

Responses

Write a response
0/2000
Loading responses...

Related Articles

Thales & Claude zerosuite

A New Hire Who Only Writes Private Notes: Plugging An AI Into A Live Support Desk Without Letting It Talk To A Single Customer

One day, one support desk, 12,166 past conversations: how we plugged an AI into a live Chatwoot support desk in draft mode only. It writes private notes, cites its sources or hands over, reads screenshots but never PDFs, and has every LLM call priced. Four audits, the mistakes included.

25 min Sep 28, 2026
chatwootcustomer-supportragpgvector +9
Thales & Claude zerosuite

No Code Shipped: Running A Job Search Like A Software Project, With An AI That Prepares Everything And Sends Nothing

One day, one AI session, zero lines of product code: a job search run with the tooling ZeroSuite uses to ship software, as a method for any trade. One private repo, one tracking file where sent means dated, content rules that refuse the unprovable sentence, a CASP roadmap that ends at a signed contract, and mail automation that fills the Drafts folder and never presses Send. With a downloadable step-by-step guide.

15 min Sep 24, 2026
job-searchcareercaspclaude-code +7
Thales & Claude zerosuite

It Works, and It Is Not Finished

The CEO walked every senndo channel himself — five channels, single and in campaign, import, statistics, a refund, the API — and everything answered. The tracking file still said no, and the one line blocking it was not code: it was a document that had quietly stopped being true. Four claims that were true when written and false when read, and the machine-readable guards that now catch each kind.

12 min Sep 14, 2026
senndocpaaslaunch-readinessdocumentation +8