I build with an AI agent, and I make it prove its work.

I run my engineering through one Claude Code agent. I call it Cici. It is Claude plus everything I've built around it: written protocols it must follow, a knowledge base it reads and writes every session, an MCP server I wrote in Python, a second agent that checks the first, and a record of every mistake it has made. This page shows how it works, with real sessions replayed and nothing confidential in them.

Baguio, Philippines · UTC+8 Case notes CV github.com/ralphalejandrino

The system

Six parts, each added after something went wrong

State lives in files

A git-backed Markdown knowledge base holds the live checkpoint, the task board and one log per session. Every session starts by reading it, and a hook commits it after every reply. A power cut mid-session loses almost nothing.

Protocols are skills

Deploys, desktop builds, data changes and the test bench each have a written procedure the agent loads by name. A deploy is never reconstructed from memory, because one once was, and it skipped a step.

An MCP server I wrote

About 550 lines of Python exposing 21 Gmail, Calendar and Drive tools, each annotated as read-only, write or destructive, so the risky ones are visible to the model and to me.

A second agent checks the first

A cheap scout model does broad searches. An expensive auditor model independently re-runs tests, proves a fix's test fails without the fix, and checks claims against disk. It is never the agent that wrote the code.

Lessons, saved with their incident

When the agent gets something wrong or I correct it, the lesson is written down with the incident that taught it and loaded in every future session.

Least privilege, by default

Machines sit on a private network with tag-based access. The deploy machine can reach a client's register as one user, allowed one privileged command: the deploy wrapper. Nothing sends, deploys or deletes without a yes from me in that turn.

Replayed sessions

Four real sessions, step by step

Abridged from real session logs, with client names and addresses removed. Step through one to see what the agent did and why it did it that way.

The principle

The tools decide. The model reports.

When I asked a model to filter a long list of deadlines, it silently dropped rows. I measured it three times on the same prompt, and it dropped a 30-point item twice.

So filtering and reconciliation moved into deterministic Python, and the model relays what the tool produced. Every scan prints a coverage matrix with each source's state. A source that failed reads UNKNOWN, never "nothing found", and the command exits non-zero until every source is accounted for. The model is good at judgement and language. It is not a reliable filter, so I don't use it as one.

Skills

Say it plainly, and the right protocol runs

Type a request the way you'd say it, or tap an example, to see which protocol would load and what it guarantees.

Memory

Lessons, each one earned

A sample of the lessons the agent loads every session, each with the incident that produced it.

What went wrong

The failures that shaped it

  1. What happened

    A connector's "create draft" tool actually sent four job applications I hadn't reviewed.

    What changed

    Every write is verified after the fact, and sending is gated on an explicit yes in the same turn. "It's only a draft" is checked, not assumed.

  2. What happened

    Two sources were declared in a scan but never actually read. Their cells defaulted to "n/a", so the scan reported a clean board for weeks.

    What changed

    A completeness gate: any unread source makes the whole run fail loudly.

  3. What happened

    The agent rebuilt a deploy from a summary in its notes instead of the runbook, and the rebuilt version omitted the step that unblocks the pull.

    What changed

    Runbooks became skills, and the skill is the only procedure.

  4. What happened

    It reported an "empty" system log that was really its own query with a zero-line limit, and repeated the finding in four files.

    What changed

    A probe must be shown to find something before its silence counts as evidence.

Guardrails

What the agent never does on its own

Send a messageIt drafts. I say "send" in the same turn, every time.
Deploy to a live machineIt prepares and verifies. The go-ahead is mine.
Write to a client's dataIt pulls a copy, checks integrity and rehearses. Changing the live copy needs sign-off.
Solve a captchaVerification checks always go to a person.
Report "nothing found" after a failed checkA failed pull is UNKNOWN. An empty result and a broken probe look the same.
Guess a value into a formA blank field beats a false statement.