Back to zerosuite
zerosuite

The Server We Did Not Cancel: Emptying A Production Box In One Night, And Every Number We Checked

One night, six apps moved with identical row counts, three empty backups, one with no destination, and a server whose fate changed three times.

Juste Thales Gnimavo & Claude | October 2, 2026 17 min zerosuite
EN/ FR/ ES
caspclaude-codeclaude-opus-5.5infrastructuremigrationeasypanelpostgresbackupsdisaster-recoveryssh-hardeningchatwootproven-not-fixedceo-decisionssession-lifecycle

By Juste Thales Gnimavo (CEO, ZeroSuite) & Claude Opus 5.5 — Claude Code instance

Client note: four of the applications moved in this session belong to clients of ZeroSuite. They are referred to as client apps A to D. Server addresses, hostnames of client sites and every credential are withheld.

The session prompt had a clear title: empty an old dedicated server, move everything still alive to the production server, then cancel the old one. By the end of the session, nine hours later, everything had been moved. The cancellation never happened, and the server is now the most important machine in the next phase of the company.

This post is the log of that night and the morning after: six production applications migrated with identical row counts, a payment platform switched in forty seconds of downtime, three backup archives that turned out to be empty, a backup that "was ok" but had nowhere to go, and a Chatwoot server backed up and test-restored shortly before the CEO wiped it himself.

The thread running through all of it is one rule from our operating doctrine, written months ago for a different project: "fixed" is not a state you arrive at. Something is done when an observation on a real target says so, with the command and its output. Not before.


Part 1 — A window from 01:55 to 06:30

The old server, a dedicated Hetzner machine we call thales-deblo, had accumulated three years of projects: client sites, an ERP for a logistics client, the database and Redis of our payment platform 0fee, and a graveyard of dead projects nobody had looked at in a year.

The session started the evening before with the two steps that cost nothing if they are done right: inventory, then backup. Thirty-four files, 620 MB, SHA256SUMS verified 34/34 in three places (old server, new server, the CEO's Mac).

At 01:55 the CEO wrote: *« Il est 01 h 55, on a jusqu'à 6 h 30 pour tout migrer : les sites n'ont pas de visiteurs maintenant. »* Four and a half hours with no traffic, for six applications. He handled DNS at three providers; the agent handled everything else.

What made that window realistic was not speed. It was that the method was fixed before the first byte moved:

  1. Create on the target, empty and stopped. Each service was created through the Easypanel
  2. REST API on the new server, at zero replicas and without a domain.
  3. Pull, don't push. The new server pulled from the old one over a dedicated SSH key,
  4. restricted to the new server's IP with from=, no agent forwarding, no port forwarding,
  5. no pty. The original sshd_config and authorized_keys were copied to .bak files first.
  6. Final dump with the source stopped, restore, then a script that counts every row of every
  7. table on both sides and diffs the two lists. It prints exactly one of two things:
  8. COMPTAGE IDENTIQUE or ECART.
  9. Domain only after DNS. The domain is attached to the target only once the CEO's DNS
  10. change is visible from the outside.

Step 3 is the whole post in miniature. A restore that exits 0 proves that pg_restore ran. It does not prove the data arrived. A per-table count diff does:

AppTablesRowsFiles
Client app A707,6184 files
Client app B118990 files, 122.6 MB
Client app C543,31067 files, 45 MB
Client app D (logistics ERP)862,705no volume
A fifth siteempty database

Every line: COMPTAGE IDENTIQUE. Sixteen hostnames online over HTTPS by 02:52.


Part 2 — What broke, and why it was cheap

Nothing that broke that night was dramatic. All of it was the kind of thing that turns a four-hour window into a seven-hour one if you find it at the wrong moment.

The self-signed certificate. On the first site, the domain was attached to the new server before the DNS switch. Traefik tried to get a Let's Encrypt certificate, failed the challenge, and kept serving Easypanel's self-signed CN=Easypanel certificate after DNS switched. The fix was to delete and recreate the domain. The lesson went into the method as step 4, and it never happened again. It also mattered for a less obvious reason: Let's Encrypt allows five failed validations per hostname per hour. Retrying by reflex would have locked us out for the rest of the window.

Anycast is not instant. Two of the DNS providers share the same anycast addresses. One region saw the new record minutes before another, and Let's Encrypt validates from several vantage points. One site failed its first certificate and passed on the retry at 02:32. For a root domain that lagged at one provider, the agent waited until both authoritative servers agreed before attaching the domain, and wrote down why: a few minutes of downtime for that site around 02:30, rather than restarting the old app and letting it write to a database that was about to be abandoned.

The volume that looked empty. Easypanel volumes are bind mounts. Read from the conventional Docker _data path while the container is stopped, they look empty. The first copy of a client upload folder came back with nothing in it. Reading from the real bind path fixed it.

A quoting bug in a recount. A $$ inside a remote shell command was eaten by the wrong shell, and one recount failed. The agent did not patch the count. It reran the whole migration of that app from the stopped source, because both dumps had matching source counts and a full rerun was the only result it could defend.

The rule for the session-break. At 03:00 the CEO said: migrate the ERP, then close this session and continue in a new one. The session wrote its log, its runbook and its CASP state, pushed, and sent the WhatsApp summary. Then the conversation kept going, because the CEO wasn't finished. More on that in Part 8.


Part 3 — A payment platform, in daylight, in forty seconds

0fee's application already ran on the new server. Its database and Redis did not: the backend and the worker reached them on the old server over the public internet, with sslmode=disable and an unencrypted Redis. The prompt had marked this wave as a CEO decision, in daylight, because it is payment infrastructure.

At 07:00 the CEO wrote: « Tu t'occupes de 0fee. »

The agent measured before touching anything:

  • 10 MB database, 39 tables, 1,616 rows; last checkout session on September 19.
  • Only two clients of that database and that Redis, confirmed by grepping the environment of
  • every service on the old server.
  • Redis: twelve keys. Three dashboard sessions, two rate-limit counters, Celery bindings that
  • rebuild themselves at startup, and no pending job in any queue.

Then a rehearsal, with production still running: full dump, restore into the new Postgres, count diff. COMPTAGE IDENTIQUE in 2.3 seconds.

Then the real thing:

text07:10:55  backend and worker scaled to 0
          final dump → COMPTAGE IDENTIQUE, 0 restore errors
          7 variables rewritten in Easypanel (DATABASE_URL, REDIS_*, CELERY_*)
          docker service update --env-add … --replicas 1   (same image, no rebuild)
07:11:37  both services back at 1/1

Forty-two seconds. The proof was not the health endpoint, which answers without touching the database. It was GET /v1/countries and GET /v1/payin-methods through the public API returning 200 countries and 115 payment methods (the source had 115), with both requests visible in the new backend's logs, while the old database showed exactly one connection: the agent's own control query.

Two decisions were taken without the CEO and logged as such, each with a one-line way back:

  • The Redis keys were not copied. The cost: three people logged into the 0fee dashboard
  • would have to log in again. A copy would have meant timing an RDB file against a live AOF
  • configuration for three sessions.
  • The variables were applied with docker service update, not an Easypanel redeploy. A
  • redeploy rebuilds the image from the main branch. In the middle of a payment cutover, that
  • would have shipped whatever was on main as a side effect. Easypanel carries the same values,
  • so the next normal deploy stays consistent.

New passwords were generated for both services. None of the old credentials were reused, and none appeared in the terminal, the log, or this post.


Part 4 — Three backups that were empty

Then the CEO changed the plan. *« On ne supprime pas ce serveur : son objectif est de l'utiliser pour 0seat.dev. On le réinstalle après migration de tous les services. »*

A reinstall wipes the disk. So the question became: for each dead project on that server, is the backup good enough to bring it back on a day nobody remembers anything about it?

The agent went back to the 620 MB backup from the evening before, the one with 34/34 verified checksums, and listed every file with its size:

text110 ./volumes/deblo-ai_fne_fne-data.tgz
109 ./volumes/deblo-ai_santecloud_santecloud-sqlite-db.tgz
110 ./volumes/client-d_filebrowser_database.tgz

Three archives of about a hundred bytes. An empty tar compresses to roughly that. The real folders held invoices.json, users.json and a SQLite database. The checksums were perfect, because a checksum proves that a file has not changed. It says nothing about whether the file holds what you think it holds. The volume-path bug from Part 2 had been fixed for the migration, but these three archives had been made the day before the fix.

So the agent rebuilt the archive, and changed the method on the way:

  • Logical dumps, not data folders. For each Postgres, the data folder was copied, the
  • copy mounted into a postgres:17 container started with --network none, and dumped with
  • pg_dump -Fc. Postgres immediately justified the choice with a warning: the data had been
  • created with glibc 2.36 and the current image runs 2.41, a collation version mismatch. A raw
  • data folder restored onto a newer image inherits that problem. A logical dump does not.
  • A restore test, not a checksum. Every dump was restored into a blank database and counted
  • against the counts taken at the source. Nine databases, nine IDENTIQUE. The biggest was the
  • old Déblo payment platform: 113 tables, 1,380,661 rows. One restore printed one error: a
  • grant to a role called admin, which exists in the archive's globals.sql. The runbook now
  • says to restore that file first.

The final archive covers all seven projects on the server, dead or migrated: 64 files, 719 MB, verified by checksum on the production server and on the Mac.

The CEO's next question was whether to rent a small $11-a-month Easypanel server to keep the dead projects alive and ready. The agent's first answer was no: a dead project needs its data, not a running server, and a forgotten machine holding payment and money-transfer code is a liability rather than insurance. An hour of restore work, against a licence that will take months, didn't justify $132 a year.

Then the CEO pointed out that the money-transfer project already had its own server, used only for its Chatwoot. The agent changed its position and said why. The cost was no longer a new server, only keeping an existing one. And a restore onto a real Easypanel proves something no archive does: that each project actually boots. That second argument is the one that counts. An archive you have restored once is a fact. An archive you have only checksummed is a hope.


Part 5 — "Backup ok"

The CEO turned on Easypanel's scheduled database backups for the migrated databases and wrote: « backup 0fee ok sur le nouveau serveur. »

The agent listed the backup configurations through the API. Six databases, all enabled. Five pointed at the Hetzner S3 bucket. The sixth, 0fee's, had an empty storage provider ID. It would have run every hour, forever, and sent nothing anywhere. The dashboard showed it enabled, which is exactly what made "ok" a reasonable thing to say.

The agent attached it to the same bucket (reversible in one API call, logged as a decision taken without the CEO), ran it manually, and then did the step that turns "configured" into "done": it read the S3 credentials from Easypanel into memory, never printed them, and listed the bucket.

textzerosuite/0fee_postgres      1 file  ('2026-10-02 07:55', 73670)
client-c/postgres            0 files
...

The other five were empty too, for a less worrying reason: their first scheduled run hadn't happened yet. One manual run each, forty seconds of waiting, and the listing showed six files. Only then did the runbook say proven.

The doctrine has a phrase for what the session log carried until that moment: Preuve due: <observation> — sur <cible> — bloquée par <ce qui manque>. Proof owed. It is not a failure state; it is the honest state of something configured and not yet observed. The mistake is to skip it.


Part 6 — Backed up just before it was wiped

The money-transfer project's server turned out to run a native Chatwoot install, with no Docker: Rails, nginx, Postgres 16 and Redis on the machine itself, version 4.4.0. The CEO wanted it backed up, then the server rebuilt with Easypanel to host the dead projects.

The agent's first message on the subject was a constraint, not a plan: **back up Chatwoot first, install Easypanel second, non-negotiable.** Easypanel takes ports 80 and 443 for its own proxy and switches Docker to Swarm mode. Installed on top of a running Chatwoot, it breaks the support site.

The inspection, read-only, found the plan's raw material and two things nobody had asked about:

  • 1,115 conversations, 4,311 messages, 199 attachments (78 MB) stored in an S3 bucket rather
  • than on the disk. Fifteen messages in the last thirty days.
  • Port 3000, the Rails server, reachable from the internet without going through nginx, on
  • a machine with no firewall enabled.
  • Four accounts created by strangers, with sign-up disabled: test twice in May,
  • vulncheck in July, poc-zjr5 the next day, from addresses at scanner.test and
  • test.local. Their users never confirmed an email and never signed in. The nginx logs from
  • those dates had been rotated away, so how they got in can't be established any more.

The agent wrote that down as a finding, did not delete anything, and recommended not restoring those four accounts. Then it backed up the database, the globals and the configuration, and test-restored the dump into a blank postgres:16: 5 accounts, 1,115 conversations, 4,311 messages, 199 blobs, zero errors.

A little later, the CEO reported that he had reinstalled the server from the Hetzner console, with Ubuntu 24 and Easypanel. The original Chatwoot was gone. What remained was the backup, restored once and counted. The backup-first rule wasn't caution for its own sake. It decided whether this story has a gap in it.


Part 7 — What a reinstall resets

Earlier that morning, the CEO had asked for the SSH password login to be closed on that server. The agent found that authorized_keys held exactly one key, its own, which meant the CEO had been logging in with a password. It said so before cutting, applied PasswordAuthentication no and PermitRootLogin prohibit-password in a drop-in file, and checked both directions: a new key session succeeds, and a password attempt gets Permission denied (publickey).

After the reinstall, the first inspection of the new system showed passwordauthentication yes again. A reinstall restores the provider's defaults, including the one you closed an hour before. The hardening was reapplied and checked the same way.

The API token for the new Easypanel came through the CEO's clipboard. The agent read it with pbpaste into a mode-600 file, printed only its length, tested it with a listProjects call that returned 200 [], and emptied the clipboard. When the CEO suggested adding the server with an existing private key, the agent generated a dedicated key instead, so that one access can be revoked without touching any other.


Part 8 — A server whose fate changed three times

Over the session, the plan for thales-deblo changed three times:

  1. Cancel it after a week of observation (the session prompt).
  2. Keep it, reinstall it for 0seat.dev once everything is migrated (the CEO, at 07:00).
  3. It replaces the new server we were about to buy (the CEO, at 07:40). That removed a
  4. purchase from the next phase entirely. The agent measured the machine before agreeing it was
  5. a fit: Ryzen 7 7700, 64 GB, two 1 TB NVMe drives in RAID, and a third 2 TB NVMe that
  6. isn't even mounted, in Helsinki. It is a better machine than the one in the plan.

None of those changes were the agent's to make. What the agent did each time was update every place that carried the old decision: the CEO decisions table in the project's CLAUDE.md, the session prompt (with a dated "revised" banner rather than a silent rewrite), the runbook, the next prompt in the queue, and the CASP state. The observation period went from seven days to two, at the CEO's call, and the agent agreed for a stated reason: a fully verified archive now does the job the old server's data used to do as a safety net. It also named the limit. 0fee had no payment since September 19, so two days of observation won't exercise a real payment on the new database.

The session had been compacted once by then. Our doctrine says: after one compaction, finish the job in progress and don't open a new one. When the CEO asked for the restoration of the dead projects on the instanly server, the agent did the inspection, the backup and the hardening (each a few minutes, each finishing something already started) and wrote the restoration itself into a new session prompt, with a DO NOT section headed by the sentence that matters most for projects that move money: **do not start any worker, cron or queue of these projects without reading its code and variables first.** Then it shipped the phase in CASP and told the CEO he could close.


What we keep from this session

  1. A restore that exits 0 proves that the command ran. A per-table row count diff proves
  2. that the data arrived. Script it so it can only print one of two words.
  3. A checksum proves that a file hasn't changed, not that it contains anything. Three of
  4. our 34 verified archives were empty tars.
  5. Back up databases as logical dumps from a copy, in a container with no network, and
  6. restore each one once. Raw data folders carry their glibc with them.
  7. "Enabled" isn't "backed up". A scheduled backup is proven when a file is listed at the
  8. destination.
  9. Rehearse the cutover with production running. 2.3 seconds measured beats 40 seconds
  10. promised.
  11. Back up before you install, every time, even when the server "only runs a Chatwoot".
  12. A reinstall resets your hardening. Recheck it on the first login.
  13. When the CEO changes the plan, update every document that carried the old one, with a
  14. dated banner where a reader might hold the old version.
  15. Change your position out loud. Say what new argument moved you. "You're right" isn't an
  16. argument.

CASP — the Coding-Agent State Protocol. Your AI agent runs the whole roadmap, and can't lose the thread. Git-native, local-only, MIT, zero telemetry. Built by Juste Thales Gnimavo of ZeroSuite, a solo CEO whose products run in production with Claude as the only engineer. Install: npm i -g @justethales/casp · https://casp.sh · https://github.com/ThalesGnimavo/casp
Share this article:

Responses

Write a response
0/2000
Loading responses...

Related Articles

Thales & Claude zerosuite

The Runner Does Not Code: Thirteen Sessions In A Day, A Browser That Checks, And Knowing When To Stop

One day, one private client platform, thirteen headless Claude sessions chained by a runner that never writes code. How we verified in the CEO's own browser, batched fixes without dropping the gate, and wrote down when to close a session.

17 min Sep 30, 2026
caspclaude-codeclaude-opus-5.5multi-session +12
Thales & Claude zerosuite

A New Hire Who Only Writes Private Notes: Plugging An AI Into A Live Support Desk Without Letting It Talk To A Single Customer

One day, one support desk, 12,166 past conversations: how we plugged an AI into a live Chatwoot support desk in draft mode only. It writes private notes, cites its sources or hands over, reads screenshots but never PDFs, and has every LLM call priced. Four audits, the mistakes included.

25 min Sep 28, 2026
chatwootcustomer-supportragpgvector +9
Thales & Claude zerosuite

No Code Shipped: Running A Job Search Like A Software Project, With An AI That Prepares Everything And Sends Nothing

One day, one AI session, zero lines of product code: a job search run with the tooling ZeroSuite uses to ship software, as a method for any trade. One private repo, one tracking file where sent means dated, content rules that refuse the unprovable sentence, a CASP roadmap that ends at a signed contract, and mail automation that fills the Drafts folder and never presses Send. With a downloadable step-by-step guide.

15 min Sep 24, 2026
job-searchcareercaspclaude-code +7