The pull request
Pull request #5 in this site's repository, "AI triage of new leads with Claude", merged on September 26, 2026. It added a background job that reads each new brief, sorts it, scores how well it fits and drafts a reply, and a panel in the admin where a person reads the result. Nothing goes to the person who wrote the brief unless someone on the team sends it.
spec docs/PLAN.md §8
plan prompt triage-v1
changes 27 files · +1,619 −27
tests 15 new tests · CI green
deploy off until evals pass1. Spec
The work started from a section of the project's plan, written before any code. It said what the triage produces, when it runs, and the rules it follows:
- Structured output, validated against a schema, never free text parsed after the fact.
- The reason the model stopped is checked before anything it wrote is read.
- The model gets the brief and nothing else: never the email address.
- Each result is stored with the model that answered, the prompt version, the tokens and the cost.
- AI drafts; a person sends.
The spec is where a person makes the decisions. The agents work from it, so a decision that isn't in it gets made by whoever happens to write the code.
2. Plan
The plan for this one change named its parts: the job, the prompt, the shape of its output, where the result is stored, the admin panel and the tests. The prompt got a version, triage-v1, so every stored result says which prompt produced it.
3. Changes
Agents wrote the change: 27 files, 1,619 lines added and 27 removed, across the API, the worker and the admin. The pull request's description says what it adds, how it was checked and what wasn't checked, in the words a reviewer reads before the diff.
4. Tests
It added 15 tests. They send every model call to a stand-in server, so they run without a network or a key, and they cover what goes wrong as much as what goes right:
- The shape of the request, and that the email address never leaves.
- A refusal, an answer cut short, and output that doesn't match the schema.
- An error that should fail at once, and an overload that should go back to the queue and retry.
- Running the triage again from the admin, and what happens when no AI key is set.
Continuous integration ran them with the type checks and the linter, and the pull request merged only on green. It was also run end to end in a browser, against a local stand-in for the model.
5. Deploy: off until the evals pass
It merged switched off. The triage runs only when an AI key is set, and the plan keeps it off in production until it passes its evals. The pull request said so, and listed what it hadn't verified: a real call to the model, because there was no key where it was built.
The evals came after, as their own change: 31 graded cases and a 90% threshold. How we test an AI agent explains them.
Who did what
| Step | The person | The agents |
|---|---|---|
| Spec | Made the decisions and approved them. | Can draft it, but don't decide it. |
| Plan | Planned the change. | Worked from the plan. |
| Changes | Reviewed the description and the diff. | Wrote the code and the description. |
| Tests | Checked what the tests cover. | Wrote the tests. |
| Merge | Decided it was ready, and that it ships switched off. | Don't merge. |
That split is the method on every project: agents write most of the code, and a person answers for what ships. How we build has the rest.