I run my engineering through one Claude Code agent. I call it Cici. It is Claude plus everything I've built around it: written protocols it must follow, a knowledge base it reads and writes every session, an MCP server I wrote in Python, a second agent that checks the first, and a record of every mistake it has made. This page shows how it works, with real sessions replayed and nothing confidential in them.
Baguio, Philippines · UTC+8 Case notes CV github.com/ralphalejandrino
The system
A git-backed Markdown knowledge base holds the live checkpoint, the task board and one log per session. Every session starts by reading it, and a hook commits it after every reply. A power cut mid-session loses almost nothing.
Deploys, desktop builds, data changes and the test bench each have a written procedure the agent loads by name. A deploy is never reconstructed from memory, because one once was, and it skipped a step.
About 550 lines of Python exposing 21 Gmail, Calendar and Drive tools, each annotated as read-only, write or destructive, so the risky ones are visible to the model and to me.
A cheap scout model does broad searches. An expensive auditor model independently re-runs tests, proves a fix's test fails without the fix, and checks claims against disk. It is never the agent that wrote the code.
When the agent gets something wrong or I correct it, the lesson is written down with the incident that taught it and loaded in every future session.
Machines sit on a private network with tag-based access. The deploy machine can reach a client's register as one user, allowed one privileged command: the deploy wrapper. Nothing sends, deploys or deletes without a yes from me in that turn.
Replayed sessions
Abridged from real session logs, with client names and addresses removed. Step through one to see what the agent did and why it did it that way.
The principle
When I asked a model to filter a long list of deadlines, it silently dropped rows. I measured it three times on the same prompt, and it dropped a 30-point item twice.
So filtering and reconciliation moved into deterministic Python, and the model relays what the tool produced. Every scan prints a coverage matrix with each source's state. A source that failed reads UNKNOWN, never "nothing found", and the command exits non-zero until every source is accounted for. The model is good at judgement and language. It is not a reliable filter, so I don't use it as one.
Skills
Type a request the way you'd say it, or tap an example, to see which protocol would load and what it guarantees.
Memory
A sample of the lessons the agent loads every session, each with the incident that produced it.
What went wrong
A connector's "create draft" tool actually sent four job applications I hadn't reviewed.
Every write is verified after the fact, and sending is gated on an explicit yes in the same turn. "It's only a draft" is checked, not assumed.
Two sources were declared in a scan but never actually read. Their cells defaulted to "n/a", so the scan reported a clean board for weeks.
A completeness gate: any unread source makes the whole run fail loudly.
The agent rebuilt a deploy from a summary in its notes instead of the runbook, and the rebuilt version omitted the step that unblocks the pull.
Runbooks became skills, and the skill is the only procedure.
It reported an "empty" system log that was really its own query with a zero-line limit, and repeated the finding in four files.
A probe must be shown to find something before its silence counts as evidence.
Guardrails