A handbook for software teams
Software is moving faster than you can review it.
Agents can write, test, review, and ship software at a pace that changes the old way of working. The harder problem is building the system around them: clear intent, useful context, good evidence, safe boundaries, and feedback from what actually happens in production. This handbook is about how teams are figuring that out.
01 Why this exists
PostHog went from 1,441 PRs in January to 4,725 in June. Over that same stretch, the percentage of PRs created by AI agents grew from around 20% to over 70%. The interesting question is no longer just how to make agents write code. It is how to build an engineering system that can keep up with that pace.
The teams figuring this out are changing the whole pipeline: how production is observed, how intent and context reach an agent, how a change is understood before it is merged, how evidence is gathered, and how outcomes feed back into the next decision. Humans move toward the things that still need judgment: what's actually worth building, what is safe enough to ship, how to keep systems reliable.
Most teams already have the raw tools: GitHub, CI, observability, browser tests, deployment platforms, and agent tooling. The harder problem is connecting them into a system that stays useful as the software, models, and workflows change.
There is another problem. Faster development can also mean faster accumulation of mess: stale tests, duplicated abstractions, outdated documentation, weak conventions, and architectural drift. The more we ship, the more important this maintenance becomes.
I am writing this by studying teams who are figuring this out: reading their engineering work, following what they are shipping, talking to engineers directly, and testing the patterns myself. The goal is to understand what actually works, what doesn't, and why.
02 How it's organized
The whole book maps to six stages.
The loop is simple. The hard part is building each stage well enough that the next one can rely on it.
- 1
OBSERVE
Know what is actually happening: errors, usage, traces, replays, deployments, and failures.
- 2
UNDERSTAND
Know what the change is supposed to do, what it touches, and what context the agent needs before it acts.
- 3
CHANGE
Let humans or agents make the change with the right tools, constraints, context, and durable execution.
- 4
VERIFY
Gather evidence that the intended behavior still holds. A green test suite is only one piece of that evidence.
- 5
SHIP
Move changes through safe boundaries with approvals, progressive delivery, rollback, and clear ownership.
- 6
LEARN
Turn production outcomes, failures, and human decisions into better context, evaluations, checks, and workflows.
Then back to Observe. The point is to make the next cycle better.
03 What's inside
The full table of contents
Twenty-two chapters across six parts. The structure might evolve as I learn from teams doing this in production. Each chapter is grounded in practical work, research, or both.
Part I: The New Engineering System
- 01When Writing Code Stops Being the Bottleneck
How the engineering workflow changes when implementation becomes less of a bottleneck, what happens when humans can no longer review every change, and where human judgment becomes more important.
- 02The Engineering Loop
Observe, Understand, Change, Verify, Ship, Learn, and why the stages have to work as one system.
- 03What Does Correct Actually Mean?
Intent, specifications, acceptance criteria, user journeys, behavioral contracts, invariants, and definitions of done.
- 04The Cost of Moving Fast
Review fatigue, AI-generated mess, architectural drift, stale tests and docs, duplicated abstractions, and the case for continuous maintenance.
Part II: Give the System the Right Context
- 05Make the Codebase Legible to Agents
Repository knowledge, architecture, domain context, history, ownership, retrieval, skills, and keeping context fresh.
- 06Understand the Blast Radius
Connect a change to dependencies, user journeys, ownership, historical patterns, and production behavior before it is made.
- 07Give Agents a View of Production
Errors, logs, traces, analytics, session replay, feature flags, deployments, and the signals that help explain what users actually experience.
- 08The Agent Harness
Tools, skills, memory, permissions, sandboxes, state, checkpoints, retries, and the environment around the model.
Part III: Change With Guardrails
- 09Before You Give an Agent Write Access
Repository conventions, contracts, test environments, secrets, permissions, and the boundaries that make agent work safe.
- 10Long-Running Agents
Durable execution, resumability, partial failures, checkpoints, artifacts, handoffs, and what changes when work outlives a single session.
- 11What Counts as Evidence?
Why a green check is only one signal, and how tests, browser runs, traces, replays, static checks, and human judgment fit together.
- 12Verify What Changed
Risk-based verification, test-impact analysis, targeted browser checks, historical regressions, and evidence proportional to the change.
Part IV: Trust, Autonomy & Shipping
- 13Who Should Decide?
What agents can decide alone, what needs approval, and how teams increase autonomy without losing accountability.
- 14Independent Verification
How to keep an agent from becoming its own oracle, with external checks, evidence quality, and failure-aware review.
- 15Human Review Without the Noise
Review routing, ownership, alert fatigue, AI review quality, and preserving human attention for decisions that need it.
- 16Ship Safely
Preview environments, CI/CD, progressive delivery, rollback, auditability, and the boundary between merge and production.
Part V: Learn, Evaluate, Improve
- 17Turn Production Failures Into Learning
Use incidents, user behavior, replays, and escaped defects to create regression knowledge and better future verification.
- 18Build the Evaluation Loop
Offline evals, production-derived evals, traces, golden cases, human judgments, false positives, and evaluating the evaluator.
- 19Keep the System From Decaying
Continuously maintain tests, docs, abstractions, architecture, agent skills, and the rules that keep the system legible.
- 20Measure Whether It Actually Works
Escaped defects, detection time, regression catch rate, evidence quality, intervention rate, cost, and safe engineering velocity.
Part VI: Put It Into Practice
- 21Connect the Tools You Already Have
Practical setups for GitHub, CI, preview environments, Playwright, Sentry, PostHog, Slack, agent tools, and the feedback loops between them.
- 22Case Studies, Playbooks & Research
Close reads of teams doing this in production, practical workflows you can copy, and a small research shelf kept deliberately high signal.
04 Is this for you
This is for you if
- You ship multiple times a week and a green checkmark doesn't fully reassure you anymore.
- You've started letting AI agents make meaningful changes and you're figuring out what context, evidence, and review they actually need.
- You want to understand the engineering systems emerging around agents, not another list of AI tools.
- You're interested in making software improve its own checks and workflows from what happens in production.
Probably not if
- You're looking for a plug-and-play tool. This is about the engineering system around the tools.
- You want a generic AI-agent tutorial covering prompts, models, or frameworks.
- You need a finished manual right now. I'm still learning and writing this as the field moves.
How does your team ship code when agents write a lot of it?
This handbook isn't built on theories or based on one person's opinions. It's informed by software teams figuring this out in production right now. So I'm looking for concrete workflows, failures, internal tools, and lessons from teams figuring this out in prod. If you've built something that works (or discovered where your workflow breaks), I want to learn from your experiences and feature it.
05 Who's putting this together
Hi, I'm Sourav, a product engineer. I've spent the last few years building for small teams (Paragraph, Pimlico, Gallery, RabbitHole) and working on my own things.
Right now, I'm building BeenThere, a minimal travel platform. To move faster as a solo developer, I started relying on agents to write code. I quickly learned that generating code is the easy part. The harder problem is building the system around the agent: giving it the right context, understanding what a change can affect, verifying the result, and learning from what happens after shipping.
I'm putting this handbook together because I needed it. I'm studying engineering work from teams already doing this in production, talking to engineers directly, and reading research on agentic software engineering, evaluation, observability, reliability, and self-healing systems. I want to understand what really works in practice.
Teams I'm studying:
These are the teams with useful public work or workflows worth studying. My role is to test the patterns hands-on, see what actually works, and organize what I learn. If your team is figuring this out in production, share how your team ships.