The playbook
Ten rules I apply to every system, learned from running my own in production.
- Measure before believing. Check what is actually live, not what the notes say.
- Spec, then plan, then small slices. Each slice has a release note and a list of what it does not prove.
- Tests must fail first. A new test only counts if it fails on the old version. I check this by deliberately breaking the code.
- Green is not evidence. Passing tests are not enough. I check the real thing, in a real browser, on the real database.
- Cheapest route first. A plain script, then an automation, then one AI call. A full AI agent only when nothing simpler works.
- Humans approve, agents propose. Agents draft; people approve anything sent, paid or deleted. The limits live in the system's permissions, not in the AI's instructions.
- Least privilege. An agent never holds a password or key it does not need.
- Every release has a receipt and an undo. Back up first, check nothing has drifted, then apply, and keep the way back.
- Settings, not guesses. When a rule belongs to your accountant, lawyer or team, it becomes a setting you control, with a safe default.
- Every session leaves a handoff. I build with more than one AI agent. When one stops, it writes down what changed, how it was checked and what is still unproven. The next one reviews that work before building on it, so nothing rests on memory or trust.
How I keep AI costs down
- Each task goes to the cheapest model that does it well. Premium models are only for hard reasoning.
- Routine work runs on local models where possible.
- Every agent has its own capped budget.
- One direct AI call replaces a full agent when that is enough. In my own systems, a single call used about 330 tokens, while a full agent turn used 5,000 to 147,000.
- Generate once and reuse: the same narration or summary is never paid for twice.
