L4 · Oct 9, 2026
Spec-driven development with AI coding agents
- spec-driven-development
- ai-coding-agents
github-spec-kit
- code-review
- case-study
I built a vendor management portal in about six weeks as the only engineer on the repository. AI coding agents wrote most of the code. I opened and merged all 179 pull requests.
That last part is both how the work got done and the weakest part of the process: there was no second human reviewer.
The project ran from 21 August to 4 October 2026, for a business unit in the group. It handles vendor onboarding, purchase orders, delivery acceptance, invoices with tax verification and multi-step approvals. The stack is FastAPI, PostgreSQL and Next.js, deployed to Azure App Service. By the end, it had 51 numbered specs, 21 backend domains and 110 Alembic migrations.
There are no AI features in the product. I used GitHub Copilot and Claude to implement it. My work was writing the specs, designing the architecture, reviewing the PRs and deciding what merged. That split depended on defining the work before handing it to an agent. It still let some serious bugs through.
Why specs before code
An agent can write a plausible approval engine quickly. It cannot tell me whether that engine matches the business process. I have to define what "correct" means and give myself a way to check it.
Once agents do most of the typing, that becomes the bottleneck. A long chat history is not a good home for those decisions.
My first commit was a blueprint, not code. I wrote it with GitHub Spec Kit: fifteen specs covering the domain, identity, data and deployment, plus a constitution containing rules every agent had to follow. Later feature specs each had task IDs, acceptance-criteria IDs and business-rule IDs. By the end, the repository referenced thousands of them.
Writing first also let me change the design before it became a migration problem. On day two, I replaced Cosmos DB in the blueprint with PostgreSQL because the domain needed relational integrity and constraints. I recorded the reason in a decision record.
On day four, an adversarial consistency review found 24 defects across the specs, mostly contradictions between them. I fixed those before any agent built on them. Both changes cost an afternoon of document editing instead of a migration.
The loop
Every feature went through the same path:
numbered spec (task, AC and business-rule IDs)
-> constitution + agent conventions file
-> agent implements in its own worktree
-> pull request, CI gates
-> my review, my merge
The constitution and conventions file held rules I did not want agents to reinterpret. The two-boundary rule, for example, keeps presentation and a backend-for-frontend in Next.js, with all business logic in FastAPI. Permissions must be re-derived from the database on every request, never trusted from the client.
For bigger stages, I ran several agents in parallel in separate git worktrees, then brought their branches together through fan-in PRs. CI gates made that work reviewable.
Alongside typecheck, lint, build, a bundle-size budget and ruff, CI checked that work referenced spec IDs. It also ran a drift check on the generated TypeScript API client, a secret scan, architecture tests, and an extended tier against a real PostgreSQL 16 service. Deployment ran only after CI was green.
The gates did reject work. Of the latest recorded CI runs from 1 September onward, 140 of 488 failed. I did not classify those failures, so I cannot separate agent mistakes from my own mistakes or infrastructure noise. The counts do not tell me who got what wrong.
When the workflow re-centred on purchase orders
The original model centred on delivery acceptance documents. A revised business process, approved on 11 September, moved the centre to the purchase order.
That meant one shared PO workspace for vendors and internal staff, replacing documents in place, and storing exact references instead of using "the newest document wins". On 25 September, the product owner also directed us to remove the configurable approval matrix and resolve routes in code.
I wrote new numbered specs for these changes. The purchase-order pivot landed between 9 and 16 September. Removing the approval matrix used an expand and contract migration, keeping existing data readable while the new routing took over.
The useful part, for me, was seeing a document diff before a code diff. I could identify superseded acceptance criteria and give agents bounded tasks against the new spec. The ledger guard flagged work with no requirement behind it.
That is my reasoning, not a measured time saving. I have no counterfactual.
There was a cleanup cost, too. Superseded specs left unreachable code behind. I did not remove it until a dedicated test-debt and dead-code pass on 4 October. I had told agents what to build, but not what to delete.
What got through anyway
The repository does not track whether each line came from me or an agent. Agents wrote most of it, and I approved all of it. These failures are mine to own.
CI was green because the tests did not run
The integration tier needed a database URL. If it was missing, pytest skipped the tier and CI reported green.
Running the full suite against real PostgreSQL produced 45 failures and 301 errors, of which 5 were real defects. I fixed those 5, made the database requirement explicit, and corrected what CI claimed to check.
A skipped tier looked like a passed tier. An agent iterating against that feedback had no way to know the difference.
The cookie worked until it carried real permissions
Sign-in uses a sealed HttpOnly cookie wrapping a signed token. The token carried the user's permissions.
With 58 permission grants, the cookie reached 4,234 to 4,491 bytes, beyond the roughly 4 KB browsers accept. Staff sign-ins returned 500. The fixture had passed with a 9-character fake signature.
I tried chunking the cookie, then compression, which was unavailable on the edge runtime. Neither addressed the design problem. The backend already re-derives permissions on every request; the token only needs the codes used for UI gating. Keeping those brought it to 2,083 bytes.
The browser checked the session before the transaction committed
After Entra ID sign-in, the callback returned 200. The next session check returned 401: "session does not exist".
From a second database connection, I could see that the rows did exist, just too late. The FastAPI dependency owning the transaction committed during teardown. In the pinned version, teardown ran after the response was sent. The browser beat the commit.
This affected every write endpoint, not just login. The fix commits before sending the response and fails loudly if that step is missing.
None of these was a spec gap. They came from CI configuration, a browser limit and a framework lifecycle, details the specs did not describe. The agents followed the tests I gave them. When that environment was wrong, they inherited the mistake.
What I would keep and change
I would keep the numbered specs and their task, acceptance-criteria and business-rule IDs, with an adversarial review before coding. I would also keep one constitution and conventions file for every agent, architecture rules enforced by tests, CI checks for traceability and contracts, and a human merge gate.
But I would not be the only reviewer for authentication, approvals or payment-related flows again. Those need a second human reviewer.
The other changes are specific: fail-closed CI from day one, where a skipped tier fails the build; production-shaped fixtures; and a deployed smoke test against the real identity provider before calling sign-in done. Every spec that supersedes earlier work should include deleting the code it makes obsolete.
I also need better measurement. I did not record time per spec, defects by origin, or a baseline without agents. There is no productivity claim to make from this project. Next time, I will track those from the first spec.
I now use GitHub Spec Kit with versioned constitutions across at least nine repositories. The part I would carry into another project is not "let the agents code". It is making the work explicit enough to review, and making sure the checks actually run before trusting their green result.