The tooling said it worked.

Five times a system I run reported success and was wrong. Each one is a real production incident, what hid it, and the check that caught it.

Baguio, Philippines · UTC+8 CV How I work with AI agents ralphmiguelalejandrino@gmail.com github.com/ralphalejandrino

What was reported, and what was true

In every case the gap below was invisible at the time. Nothing errored. That is the whole subject.

Incident Reported Actually true
A retry became three orders 1 order placed 3 orders queued
An import that lost rows 0 errors 1,604 rows dropped
A fix that did nothing normalisation applied 0 characters changed
A fabricated audit 8 migrations, 1 unapplied 22 migrations, all applied
A confident wrong diagnosis root cause found wrong subsystem, still broken

Where these come from

Locus

A multi-tenant CRM and delivery-tracking system for an LPG distribution business. Django backend, shipped to the client as a frozen Windows desktop binary. A language model sits behind a one-method provider interface and turns free-text customer messages into structured order lines.

2,800 customer records · 139,000 messages
Tested on the client's real data · source: github.com/ralphalejandrino/ProjectCRM

Tabula

Point of sale, inventory and cost-of-goods for a café, running on the register they take money with. Offline-capable web app with a versioned service worker, deployed on a watched, snapshot-first process with a rollback commit recorded before every push.

670+ automated tests · 0 unrecorded deploys
Live on client hardware · source: github.com/ralphalejandrino/ProjectPOS

01

A retry turned one customer order into three Missing idempotency key

One real order sat in the dispatch queue three times. Nobody had done anything twice. The business could have sent three deliveries and billed for one.

  1. A launch routine copied a 53 MB SQLite file inline, while the server was already taking traffic.
  2. A concurrent write hit database is locked.
  3. The intake endpoint returned 500.
  4. The upstream SMS gateway did what gateways do, and retried.
  5. The retry landed as a brand new row, because the endpoint had no idempotency key.

The fix persists the gateway's own receipt timestamp as an external key and lets a unique database constraint arbitrate. An application-level "have I seen this before?" check would race under exactly the conditions that caused the bug. Separately, the database moved to WAL so a backup stops blocking writers at all.

Proven by deleting the constraint and confirming the test went red. A fix that cannot be made to fail has not been tested. Negative control: a customer genuinely ordering the same thing twice must still produce two records.

02

An import reported success and dropped 1,604 rows Silent partial failure

There was no symptom. A 265,000-row message archive imported cleanly, no errors, no warnings. The absence of a symptom is the reason this one is here.

importer reported 263,857 rows imported, 0 errors
raw count of the source file 265,462 rows present

The parser matched one row shape and quietly skipped another: group messages, which included real customer orders. It surfaced only by ignoring the importer's own reporting and counting raw rows in the source instead. A second defect fell out of the same pass, where 27,264 dates failed to parse because one month was abbreviated with four letters where the format string expected three.

The recovered data was material. A single year's records went from 155,546 to 181,371. A clean exit is not verification. Reconcile against a count taken from the source, never from the tool doing the work.

03

The fix that would have looked correct and done nothing Plausible non-fix

A Spanish-derived character was rendering wrong throughout imported customer data, in names and in street addresses. The obvious fix is Unicode normalisation.

The obvious fix does nothing. The source contained a combining diaeresis where a tilde belonged, and Unicode has no precomposed form for that letter and diaeresis pair, so normalising is a no-op that raises no error and changes no bytes. It would have passed review, shipped, and silently accomplished nothing.

The real fix swaps the combining mark first, then normalises, scoped to the single affected letter so legitimate umlauts elsewhere survive. A test now guards the no-op case specifically, so nobody reintroduces the plausible version.

The reported scope was ten fields. It was twenty, and nine of those were street addresses, which is what a courier actually reads on the way to a delivery. Worth considerably more than the names.

04

An AI-generated audit that was fabricated end to end Unverified handoff

A session summary written by an AI assistant reported the state of a codebase. Every specific claim in it was confident, precise, and false.

the summary claimed 8 migrations, 1 destructive and unapplied, a model field missing, framework 2 majors behind
the repository 22 migrations all applied, field present, framework current

It cost five tickets deferred for no reason and a full working session spent proving nothing was broken. The response was a standing rule: any handoff's claims about the state of the code are untrusted until re-checked against the repository. Migrations listed, fields grepped, versions printed.

Roughly one minute of checking, against a session of phantom work. This is why I work AI-assisted and still verify before I act on anything a summary tells me.

05

A plausible lead shipped the wrong fix to a live machine Premature diagnosis

A client reported their register crashing during service and sent a screen recording with the report. An explanation presented itself early and fit the description well, so I shipped against it.

It was not the fault. Decoding the recording frame by frame afterwards showed a different failure entirely, and the original one was still there. The cost was a pointless deploy to a machine a business takes money with.

When a report ships evidence, decode all of it and reproduce the actual failure before writing a line of the fix. A hypothesis that explains the symptom is not a verified hypothesis. Plausibility is precisely the thing that makes a wrong diagnosis expensive.

How I work

I use AI assistants throughout and treat their output as draft work. These are the practices that make that safe rather than fast and sorry.

Prove the test can fail

After a fix passes, remove the fix and confirm the test goes red. A green test never shown to fail is decoration.

Run a negative control

Every change gets a case that must not change. Deduplication is only correct if a genuine repeat still gets through.

Run a positive control

When a query returns nothing, prove the query works before concluding the data is absent. An empty result and a broken filter look identical.

Verify claims against the code

Line numbers, field names, migration state, versions. Checked in the repository before they shape a decision.

Keep diffs small and reviewable

Scoped commits with the reasoning written down, so a reviewer follows the argument instead of reverse-engineering it.

Deploy with a way back

Snapshot first, record the rollback target before starting, verify the live artifacts independently rather than trusting the deploy script's report.