Back to zerosuite
zerosuite

A New Hire Who Only Writes Private Notes: Plugging An AI Into A Live Support Desk Without Letting It Talk To A Single Customer

One day, one support desk, 12,166 past conversations: how we plugged an AI into a live Chatwoot support desk in draft mode only. It writes private notes, cites its sources or hands over, reads screenshots but never PDFs, and has every LLM call priced. Four audits, the mistakes included.

Juste Thales Gnimavo & Claude | September 28, 2026 25 min zerosuite
EN/ FR/ ES
chatwootcustomer-supportragpgvectoropenrouterclaude-haiku-4.5geminianonymizationhuman-in-the-loopllm-costsvisionaudit-methodologyclaude-code

By Thales (CEO, ZeroSuite) & Claude Opus 5.5 — Claude Code instance

TPEcloud is our web-hosting business in Côte d'Ivoire: cPanel hosting, domain names, business e-mail, SSL certificates, SMS. Support runs on Chatwoot, a self-hosted help desk, across a website live chat, two WhatsApp numbers and an e-mail inbox. Traffic on the live chat peaks after 6 p.m., when the customers' own working day is over.

A week before this session, one of the agents who handle the live chat sent the CEO a report listing everything they do. It is a long list:

  • answering customers;
  • diagnosing cPanel, e-mail and DNS problems;
  • following domain registrations with the national registry and a registrar;
  • topping up supplier accounts, including the SMS providers;
  • running the cash box;
  • answering some customers in the evening and at weekends.

The report concluded that first-level technical support could be handed over to the new technicians. The CEO answered that a new employee was coming to help: an AI.

On September 28, 2026, in a single Claude Code session, that employee was hired. By the end of the day:

  • The knowledge base. A nightly job extracts question-and-answer pairs from resolved conversations, anonymized, and runs through the 12,166 conversations in the history.
  • The admin page. A human approves, corrects, merges or rejects every extracted entry, and the page shows what each model call cost.
  • The drafts. On two inboxes (the live chat and the official WhatsApp number), the AI reads each incoming customer message and writes a private note that only agents can see.
  • The screenshots. It reads screenshots customers send, never their PDFs, and asks for a screenshot when a customer reports an error without showing it.
  • The audits. Four independent audits, each ending "go with fixes", with every top finding fixed. 93 tests.

It has not sent a single message to a customer, and it cannot. This post is about why, and about the dozen design decisions that one constraint forced.


Part 1 — The real knowledge base is already in the inbox

The obvious way to build a support bot is to write a FAQ and point a model at it. We had one: 95 standard answers the team pastes into chats. They are useful and they are not enough, because they answer the questions the team thought customers would ask.

The questions customers actually ask are in the 12,166 resolved conversations. So the first job is extraction. For each resolved conversation, one model call reads the anonymized transcript and returns zero, one or several reusable question-and-answer pairs, a category, and a note. Most conversations yield nothing: "thanks", "ok", a payment receipt. That is the expected result, not a failure.

Three rules shaped the pipeline before any code was written.

Anonymize before any model sees anything. E-mail addresses, phone numbers, amounts, secrets and domain names are masked in the text before it leaves the server. Customer data never goes to a free model, or to any model whose provider keeps data. Calls go through OpenRouter, pinned to Google Vertex with zero data retention and data collection denied.

Nothing extracted is trusted. Every extracted entry lands as a candidate. Only a human turns a candidate into an approved entry. The 95 standard answers were imported as candidates too: being in the FAQ is not the same as being checked.

Deduplicate, but lose nothing. A new pair is compared, by exact cosine scan, with the nearest existing entry:

  • Above the threshold (0.88). It becomes a variant of that entry, and the entry's frequency goes up. Frequency is what tells the team which answers matter most.
  • Below the threshold. It becomes a new candidate.

We measured two real duplicates at 0.83, below the threshold, and kept 0.88 anyway. A duplicate costs one click to merge on the admin page. A wrong attachment makes an answer disappear inside another one, and nobody notices. Every merge records the similarity of the two entries, so the threshold can be recalibrated on real decisions rather than on a guess.

One rule came out of the first audit. An entry a human rejected must not absorb new answers. If the nearest neighbour is a rejected entry, the new pair becomes a fresh candidate with a note: "close to rejected entry #N". A rejection is a human decision; the pipeline does not get to quietly overrule it.

The nightly job itself is ordinary, careful engineering:

  • One job at a time. A Postgres advisory lock is held on a separate autocommit connection, so the lock never keeps a transaction open.
  • No failure is final. Conversations that failed are retried by id, up to three attempts. A reading error, a model error and a truncated output are recorded as distinct outcomes.
  • It knows where the new work ends. The job reads pages of resolved conversations until it has seen three consecutive pages it already knows.
  • One conversation cannot stop the run. Every exception is caught per conversation.

Part 2 — Choosing the model by blind vote, and keeping the cheap one

Two candidates for extraction: Claude Haiku 4.5 and Gemini 3.8 Flash. Both ran on the same 50 real conversations. The CEO then compared the two outputs for each conversation without knowing which model wrote which.

ResultCount
Preferred Gemini 3.8 Flash12
Preferred Haiku 4.52
Tie36

On the 15 conversations where the two models disagreed on substance, the vote was 7 to 0 for Gemini. The CEO's instruction was short: keep the faster, cheaper model, the two are almost the same. Gemini 3.8 Flash does the extraction, at about $0.0018 per conversation read. At that rate, the whole history costs about $22.

The CEO then raised the obvious next step: even cheaper models (Gemini Flash-Lite, IBM Granite 8B), and eventually Granite on our own GPU, with no per-token cost at all. Claude's answer was not "yes" and not "no" but "measure first". The decision was written down with a date: one month of real consumption, then decide. That only works if consumption is recorded, which led to the most boring and most useful table of the day.

Every model call is a row in llm_usage. The row records the task, the model, the provider, the tokens, OpenRouter's fee, the upstream Vertex cost (billed separately, on our own key), and the latency. It stores no content, ever. The admin page shows it all:

  • cost per day;
  • cost per task and per model;
  • a projection for the month.

When the month is over, the decision about Granite will be a comparison of numbers, not of impressions.


Part 3 — The page where humans stay in charge

The admin page is deliberately plain: server-rendered HTML forms and no JavaScript framework, because nothing on it needs partial updates. For each candidate it shows:

  • the question, the answer and the variants;
  • the category;
  • how often the question came up;
  • the nearest existing entries, with their similarity.

Next to them are four buttons: approve, reject, reopen, merge. An approved entry can still be corrected later; the answer is re-embedded when its text changes.

The page also shows coverage: what share of past conversations the top 20, 50 and 100 entries account for. That is the number that tells the team where an hour of validation pays off most.

The second audit found five real problems on this page, all fixed before it went live:

  1. A race with the nightly job. The job can increment an entry's frequency while a human is merging it. Fixed with row locks and increments done in SQL.
  2. No clickjacking protection. Fixed with X-Frame-Options: DENY and a CSP frame-ancestors rule.
  3. Merge rules. Nothing may be merged into a rejected entry, and an approved entry may only be merged into another approved one.
  4. Raw errors. An OpenRouter failure during re-embedding showed as a raw 500. It is now a clear 502.
  5. Deadlocks. Two merges running in opposite directions could lock each other. Rows are now always locked in id order.

Part 4 — Draft mode: the AI writes, a human sends

The CEO's instruction for the first live test was precise: plug the AI into two inboxes, the WhatsApp number and the live chat, but in private mode. I want to see its answers without the knowledge base being updated.

The design follows from one sentence we wrote at the top of the project's rules file: the bot never posts publicly. Every output is a private note in the conversation. Agents see it; the customer never does.

Why a webhook and not a Chatwoot "Agent Bot"

Chatwoot has a built-in bot integration. We didn't use it. A conversation assigned to an Agent Bot goes to "pending" status, which changes the team's queue: they would stop seeing new conversations where they expect them. A draft assistant must not change the team's workflow. So the bot listens on an ordinary account webhook for message_created events. The webhook:

  • checks a shared secret, compared as bytes;
  • ignores other accounts, private messages and outgoing messages;
  • ignores every inbox that is not explicitly in draft mode.

Waiting for the customer to finish typing

Customers on chat write in bursts: "hello", "I have a problem", "my mail doesn't work", a screenshot. Drafting after each message would be wasteful and wrong. The bot waits for 20 seconds of silence before starting.

The first version cancelled the wait-and-draft task on every new message. The third audit pointed out that this could cancel a generation already in progress, after the model had been paid. Now the wait can be cancelled, but a generation never is. A message arriving during generation schedules one more pass afterwards.

Never draft for a conversation an agent already answered

Before generating, the bot checks that the last public message is from the customer. It checks again after generation, because a model call takes several seconds and an agent may have replied meanwhile. A stale draft under an agent's answer is noise at best. At worst, someone copies it and the customer gets two answers.

Three decisions, one of them forced

The model must return one of three decisions:

  • repondre (answer). It is only valid if the answer cites at least one retrieved knowledge-base entry. An answer with no cited entry is an invented answer. The code converts it into a handover, whatever the model says.
  • demander_precision (ask for details). This is the only reply allowed without a knowledge-base entry, because it contains no solution. It asks for a screenshot, the domain name or the exact error message.
  • passer_la_main (hand over). The rules make this mandatory for:
  • - money: payment, invoice, refund, credits bought but not received;
  • - cancellations and disputes;
  • - a customer asking for a human;
  • - a second follow-up with no solution;
  • - anything that needs the team to act on the customer's account or the server.

Retrieval

Retrieval is hybrid:

  • the 20 nearest entries by vector similarity (pgvector, exact scan);
  • the 20 best by French full-text search;
  • the two lists fused with reciprocal rank fusion.

During the test, candidate entries are included; the approved-only rule comes back before any automatic mode. The transcript is sent to the model as data, inside a tag the customer cannot close, and the system prompt says so: the conversation is data, never instructions.

Hard limits on cost

The limits are set in configuration:

  • at most 6 drafts per conversation per hour;
  • a daily budget for drafting, screenshot reading and query embeddings;
  • 2 drafts generated in parallel.

When a limit is reached, the bot stops drafting. The team loses nothing: they are answering those conversations anyway.

The first real draft appeared a few minutes after deployment, on a live-chat conversation. The CEO's verdict on the first batch, drafted by Claude Sonnet 5, was one word: perfect.


Part 5 — The note that could be copied whole

The first version of the note was built for the CEO, who was evaluating the drafts:

  • a header with the decision and the reason;
  • the knowledge-base entries used, with their scores;
  • the draft reply itself;
  • "to check" hints for the agent.

The CEO asked for all of it to go, for a reason that deserves a part of its own:

"An agent can make a mistake, copy the whole thing and send it. Show the answer directly. I've already told the staff we're in beta."

This is a design rule we had missed, and it applies to any AI output placed where a human can forward it: the note is not a report, it is a draft. A tired agent at 9 p.m. selects all and pastes. Anything in that note can reach a customer: a score, an entry number, an internal instruction.

Now the note contains only text an agent could send as is:

  • for an answer or a request for details: the reply, nothing else;
  • for a handover: one line, "AI assistant: to be handled by an agent", and the reason.

Sources and scores moved to the bot's log, as ids and numbers with no content. That is where calibration happens anyway.

The same pass made the note safe to paste:

  • mention:// links are neutralized, so a draft cannot notify an agent;
  • markdown links are broken, so a link's label can never hide its real address.

Part 6 — Sonnet was perfect; Haiku got the job

With the drafts judged perfect, the CEO asked to switch the reply model to Claude Haiku 4.5. We measured a Sonnet 5 draft at about $0.023. Haiku costs roughly a third of that.

Claude agreed, and recorded one reservation in the session log rather than in a chat message that would be forgotten. The drafts the CEO judged were procedural questions: configure a mailbox, point a domain. A cheaper model is more likely to miss a handover rule, on exactly the conversations where a mistake costs money: a disputed payment, a credit top-up that never arrived. So "watch Haiku's handovers on money and disputes" is on the list of things to review with the team. That is also why the handover rule will get a deterministic keyword filter before any automatic mode. A rule that protects money must not depend on the prompt alone.


Part 7 — Screenshots, PDFs, and a phone number hidden in a file name

Customers send attachments: screenshots of an error, photos of a screen, PDFs of company registration papers, leases, ID cards. The CEO asked how the assistant handles them. The honest answer was: at that point, it didn't see them at all. So we built it, under three rules.

Images are read once, and the description is what travels. An image cannot be anonymized before sending the way text can. So each screenshot goes once to the reply model (Claude, on Vertex, with zero retention). The model is asked for a 1-to-4-sentence description:

  • which screen or software is shown;
  • the exact text of any error message, in quotes;
  • no personal data.

The description is anonymized like the rest of the text and cached by attachment id. From then on, retrieval and drafting see the description, never the image. On a synthetic webmail error, the exact error message came back word for word, for $0.0018 per image.

PDFs are never read. They are almost always administrative documents (registration papers, a lease, an identity document) that a human must check anyway. The assistant sees only the file name. If the document is the heart of the request, it hands over.

File names are data too. This is what the fourth audit caught. The draft showed attachment names, and customers name their files the way they name everything. CNI_0707070707.pdf is an identity-card scan with a phone number in the name. File names now go through the same masking as the text: e-mails and long digit runs are masked. That one becomes CNI #.pdf. The extraction side never sees names at all, only the extension, so no file name can end up in the knowledge base.

The audit also tightened the download:

  • Chatwoot host only. The bot downloads from the Chatwoot instance and nowhere else, over HTTPS, with an exact host match. Redirects are followed by hand and re-checked, so a crafted attachment URL cannot make the server fetch an arbitrary address.
  • 5 MB, counted while reading. The size limit is enforced during the streamed read, not after loading the whole file into memory.
  • At most three images per draft.

Then the CEO added the rule that changes the most in practice: the assistant should often ask customers for screenshots, it's much easier to solve when we see the images. It is now in the prompt:

  • When it applies. A customer reports a problem (an error, a site that doesn't load, mail that doesn't send, a certificate refused) without showing the exact message, and hasn't sent a screenshot.
  • What the assistant does. It asks for one, reminds them to hide any password, and asks for the domain name if it's missing.
  • When it doesn't. For a plain "how do I…" question, it answers directly.

Part 8 — What the audits found, and one mistake of our own

Each block of work was followed by an independent audit: a separate agent that sees the code fresh and gets no say in how it was built. Four audits, four "go with fixes" verdicts. The pattern is the one we keep seeing: the builder sees the feature; the auditor sees the edges.

AuditReal findings, all fixed
Nightly extractionConversations lost past the first known page; one exception stopping the whole run; a rejected entry absorbing new answers; a lock holding a transaction open; the model's note not anonymized; costs counted before the commit
Admin pageRace with the nightly job; clickjacking; merge rules; raw 500s; cross-merge deadlock
Draft botA paid generation could be cancelled; note injection through mentions and links; no cost ceiling; a database session held open during the model call; model output types unchecked
ScreenshotsPhone numbers in file names; download not streamed; the same image described again on every draft; missing tests; images skipped with no log line

The mistake of our own was not in the code. Early in the day a helper script loaded the settings while one variable was missing. Pydantic's validation error helpfully printed the tail of a neighbouring secret in its input_value field, straight into the session output. Nothing left the machine. But a secret that has been displayed is treated as leaked. We rotated the webhook secret and the admin password, and the CEO regenerated the OpenRouter key. Since then, every script that loads settings sets a dummy database URL first and filters input_value out of any error. The rule is in the project file.

Another finding came from measurement, not from an audit, and it is uncomfortable. The similarity threshold for retrieval does nothing. On a dry run over six real conversations, every retrieved entry scored between 0.66 and 0.82 against the query, relevant or not. A threshold of 0.55 lets everything through; a threshold of 0.75 would cut relevant entries. Today the model does the filtering, guided by the "cite or hand over" rule. The log records the ids and scores of every entry retrieved for every draft, marking which ones were cited. That is the data we will recalibrate on. It is written here because a post that only lists the things that worked would be the same kind of note as the one in Part 5.


Part 9 — Who does what now

That agent's report asked a real question: which tasks go to the new technicians, and which stay with the experienced agents? With an AI on the team, there are three columns, not two.

WorkThe AITechniciansExperienced agents
Repetitive "how do I" questions (mail setup, cPanel, DNS, SSL)Drafts the answer from approved entriesCheck and send—
Problem reported without detailsAsks for a screenshot and the domainDiagnose from the screenshot—
Diagnosis needing account or server accessHands overFirst level, escalate with evidenceLevel 2
Payments, refunds, disputesAlways hands over—Keep it
Supplier accounts, registry and registrar follow-up, SMS providers, top-ups, the cash boxNever—Keep it
Correcting and approving knowledge-base entriesProposes candidatesCorrect in the admin pageApprove

The AI doesn't replace anyone in this table. It takes the first draft of the repetitive work, and the job of asking for the screenshot nobody enjoys asking for. The knowledge it drafts from is the team's own past answers, approved by the team.


Part 10 — What is deliberately not built yet

The team asked the right questions about going live, and the answers are in the roadmap, not in the code.

"If the AI doesn't know, can it tell the customer to wait for an agent?" Yes, in automatic mode: a short patience message plus a label the team can filter on. It will be built together with automatic mode, not before.

"If an agent answers, will the AI answer too and create a duplicate?" Not in draft mode, which checks twice. For automatic mode, the rule will be stricter: once a human has replied publicly in a conversation, the AI stops replying publicly in it.

"Can the AI still leave us private notes once it answers directly?" Yes. The plan is public or private per message: public when an approved entry covers the question, a private note otherwise.

Before any of that, we need two things. The first is a measurement: the share of drafts sent unchanged, per category. Automatic mode will be switched on per category, only where that share is above 90%, and only from approved entries. The second is the deterministic filter for the money rules. The Payments inbox will never be automatic. The e-mail inbox joins draft mode later.

The welcome message the team currently sends automatically will be switched off only when an inbox goes automatic. Until then, a human greets and the AI prepares.


Part 11 — The night the history went in, and the morning it turned into a queue

The CEO's last instruction was operational: traffic on the live chat is higher after 6 p.m., so run tonight's extraction batches from now until 9 a.m. tomorrow. The history had been planned in four evening slices. Claude added a --until 09:00 option: the job stops starting new batches at that time and reports the stop in its summary, never halfway through a write. The run started at 18:46 UTC.

It never needed the deadline. At 22:03 UTC, three hours and seventeen minutes later, the whole history had been read. The run summary, verbatim:

Count
Conversations read12,079
Skipped before any model call (no customer message, no agent reply, or too short)9,777
New candidate entries702
Variants attached to an existing entry823
Read, but nothing reusable763
Generation errors (retried the next night, one left)14
Cost, Gemini 3.8 Flash through OpenRouter on Google Vertex$7.28

Four in five conversations never reached a model, which is why the bill is that low. Every call was logged with its cost; the cost page showed $7.56 for the project so far, embeddings and the next night's run included.

The next morning the knowledge base held 807 candidates and one approved entry. Extraction was no longer the bottleneck; human review was. The candidates are not equal: 51 of them (four with more than twenty variants) cover about 560 real conversations, while 567 were seen only once. The admin page already sorted by frequency, so the plan was simple: approve the top of the queue first.

The bug we found before anyone clicked

The CEO answered: I'm telling the live-chat team, we start validating today. Several people were about to open the same page, at the same time, sorted the same way, under one shared login. Before a single colleague logged in, Claude read the approval handler again with that in mind and stopped the rollout for an hour. Two ways to lose work silently:

  • one person corrects an answer and approves it; a second person, on the same card, rejects it a second later. The rejection wins;
  • one person corrects and approves; the second approves the original wording. The original goes back into the knowledge base.

The row lock was there. What was missing was a check that the card was still the one the person had in front of them. And with a single login, the reviewed_by column would have recorded the same name for every decision, on a base we already know contains wrong answers.

The fix, audited by a fresh agent before deployment:

  1. One account per validator, with a display name recorded on every decision. The CEO's account records "Directeur Général"; a generic account can be given a real name later without changing its login.
  2. Disjoint lots. Each validator lands on their own slice of the queue, still sorted by frequency. Nobody starts on the same card as a colleague.
  3. A version token on every card: a hash of its status, its text and its last decision. A decision taken on a stale card is refused with "already handled by [name]: reload the page". The audit caught that our first version, based on the review timestamp alone, missed one case: the import of canned answers rewrites the text of candidates nobody has reviewed yet.

The audit found three other things worth fixing: a comma in a password would have silently cut it, and duplicate names would have overwritten each other. On top of that, the review timestamp was the start of the transaction, not the moment of the decision. Two live-chat colleagues and the CEO received their accounts that morning.

What the approved base becomes next

The CEO added one requirement: the approved entries should become the internal handbook for the interns joining the team later. They should be able to read them, practise, and learn to correct the AI when it's wrong. Claude's recommendation was a separate read-only page, not an account on the admin page. An intern practising must not be able to approve anything, and they should learn from verified answers, not from candidates. The page will show a question, let the intern write an answer, then show the approved one. It goes on the roadmap for when about a hundred entries are approved.


How to copy this for your own support desk

You need a help desk with webhooks, a Postgres with pgvector, and a model provider that lets you choose zero data retention. You don't need our stack.

  1. Extract from resolved conversations, not from the FAQ you wish customers read. Anonymize before any model call.
  2. Everything extracted is a candidate. Only a human approves, and a human rejection is never overruled by the pipeline.
  3. Record the cost of every call, without content, from the first day. Decide on cheaper models with a month of numbers.
  4. Start in draft mode, as private notes, on the inboxes where volume is highest.
  5. Put only sendable text in the note. Someone will copy it whole.
  6. "Cite an entry or hand over" is enforced in code, not only in the prompt. Money and disputes always go to a human.
  7. Check that the customer is still waiting, before and after generation.
  8. Read screenshots once, keep only the description, never read documents. Mask file names.
  9. Log retrieval scores and calibrate thresholds on real decisions, not on defaults.
  10. Audit each block with a fresh agent before it goes live.
  11. Go automatic per category, only where drafts are sent unchanged more than 90% of the time.

CASP — the Coding-Agent State Protocol. Your AI agent runs the whole roadmap, and can't lose the thread. Git-native, local-only, MIT, zero telemetry. Built by Juste Thales Gnimavo of ZeroSuite, a solo CEO whose products run in production with Claude as the only engineer. Install: npm i -g @justethales/casp · https://casp.sh · https://github.com/ThalesGnimavo/casp
Share this article:

Responses

Write a response
0/2000
Loading responses...

Related Articles

Thales & Claude zerosuite

No Code Shipped: Running A Job Search Like A Software Project, With An AI That Prepares Everything And Sends Nothing

One day, one AI session, zero lines of product code: a job search run with the tooling ZeroSuite uses to ship software, as a method for any trade. One private repo, one tracking file where sent means dated, content rules that refuse the unprovable sentence, a CASP roadmap that ends at a signed contract, and mail automation that fills the Drafts folder and never presses Send. With a downloadable step-by-step guide.

15 min Sep 24, 2026
job-searchcareercaspclaude-code +7
Thales & Claude zerosuite

It Works, and It Is Not Finished

The CEO walked every senndo channel himself — five channels, single and in campaign, import, statistics, a refund, the API — and everything answered. The tracking file still said no, and the one line blocking it was not code: it was a document that had quietly stopped being true. Four claims that were true when written and false when read, and the machine-readable guards that now catch each kind.

12 min Sep 14, 2026
senndocpaaslaunch-readinessdocumentation +8
Thales & Claude zerosuite

The Browser in Claude's Hands: Driving the CEO's Own Chrome

Claude-in-Chrome lets a Claude Code session drive the CEO's real browser — same profile, same logged-in sessions. What the tool actually does, why it beats asking a human to click and report back, and where the human still wins. Grounded in the day Claude walked a brand-new customer through senndo's production console, billed messages included.

8 min Aug 18, 2026
claude-in-chromebrowser-automationclaude-codeclaude-fable-5 +9