Track lessons
AI-Era Engineering
0 of 2 completed
- Spec-Driven Development: Write Specs AI Agents Can Execute
- Agentic TDD: Make the AI Write Failing Tests First
Other tracks
Spec-Driven Development: Write Specs AI Agents Can Execute
Coding agents move quickly but often fill gaps in vague requests with plausible assumptions. A short executable spec fixes the goal, boundaries and proof of completion before implementation begins.
When to use
- A feature spans storage, API behavior and user-facing changes.
- Several agents or pull requests will implement dependent tasks.
- The cost of scope drift is higher than the cost of a one-page spec.
Example
For password reset, define identical responses for unknown emails, a 30-minute single-use token and session revocation, then split storage and endpoints into ordered tasks.
TL;DR
Spec first, plan second, code last. The agent implements tasks from the spec, nothing else.
Read specs/password-reset.md. Implement task 1 only. Stop and show the diff.Steps
- 1
Write the spec
Capture the observable behavior, explicit exclusions, existing constraints and PR-sized tasks in one file the agent can reread.
specs/password-reset.md# Password reset ## Goal Users can reset their password by email. ## Out of scope SMS, admin resets, UI redesign. ## Acceptance criteria - [ ] POST /auth/reset sends one email; same response for unknown emails - [ ] Token expires after 30 minutes, single use - [ ] Old sessions are revoked after reset ## Constraints - Use existing mailer in src/server/mail.ts - No new dependencies ## Tasks 1. Token table + migration 2. Request endpoint + tests 3. Confirm endpoint + tests - 2
Ask for a plan, not code
Have the agent inspect the referenced code and expose unanswered questions before any implementation makes those choices expensive to reverse.
promptRead specs/password-reset.md and the code it mentions. Propose a plan per task. Do not write code yet. List open questions. - 3
Implement one task at a time
Limit the session to one task and require its checks, so the resulting diff stays small enough to review against the spec.
promptImplement task 1. Run the tests. Stop. - 4
Review against the criteria
Ask for concrete evidence per criterion after implementation; an unsupported claim should remain unmet.
promptCheck the diff against every acceptance criterion in the spec. Mark each as met or not met with evidence.
Gotchas
- "Out of scope" prevents most agent drift; never skip it.
- Acceptance criteria must be testable; avoid "works well".
- Commit the spec with the code so reviewers see intent.
- Tools like GitHub Spec Kit automate this flow; the template works without them.
Cheat sheet
| Goal | One sentence |
| Out of scope | What not to touch |
| Acceptance criteria | Testable checkboxes |
| Tasks | Small, ordered, one PR-size each |
Related
Sources
Reviewing agent output? CodeCrab reviews pull requests locally with your own AI tools.
