Here is a question worth sitting with before your next org review: if you gave every engineer on a team a coding agent that never sleeps and writes most of the first draft, would you make the team smaller, or would you keep it the same size and hand it a bigger charter? Most leaders answer “smaller” on reflex. That reflex is expensive, and it misreads what the two pizza team was ever for.
The two pizza team is Amazon’s most exported piece of org design, and it is about to be its most misunderstood. As agents absorb a growing share of the keyboard work, the temptation is to shrink the roster to match. But the rule was never counting mouths at the table. It was counting something harder to see.
The rule was always a proxy
The origin story is well worn. At a 2002 retreat, when managers asked for more communication, Jeff Bezos reportedly shot back that “communication is terrible” and restructured the company around small autonomous teams — small enough to feed with two pizzas. The number, five to eight people, was never the point. It was a proxy for two things that actually mattered.
The first is communication overhead. Coordination cost grows faster than linearly with headcount; every person you add multiplies the channels that must stay in sync. Capping team size caps that overhead. The second is single threaded ownership: a team small enough to hold one coherent mission in its collective head, accountable end to end for a product it can build, ship, and operate without passing work across departmental fences.
Underneath both proxies sits the real binding constraint: cognitive load. A team can only hold so much of a domain in its head before delivery slows, quality drops, and people burn out. The two pizza rule was a crude but durable way to keep one team’s cognitive budget from overflowing. Read it that way and the agent question sharpens immediately. Agents don’t change whether the rule applies. They change what the rule is measuring.
What agents actually change
At the Pragmatic Summit, Martin Fowler framed the fork precisely: “Are we seeing two pizza teams becoming one pizza teams because agents don’t eat pizza, or do we see two pizza teams staying and becoming much more effective and capable? My bet is on more effective two pizza teams.” (Hands-On Architects)
It sounds like a headcount question. It is really a question about where the load goes. Here is the mechanism that resolves it: an agent relocates cheap cognitive load and returns expensive load. The boilerplate, the glue code, the first draft implementation, the codebase spelunking — all the work that was tedious but not hard — moves to the agent. What flows back to the humans is the work that was never cheap: verification, integration, and judgment. Someone still has to decide what to build, review what came back, and own the call when the agent cannot make it.
Kent Beck, on the same stage, put the brake on the naive reading with four words: “AI is an amplifier.” An amplifier does not subtract band members; it changes what each one carries. He described pairing with an agent and, counterintuitively, praising its slowness — the three minute gaps where the humans talk about naming, about conditionals, about what they should be doing next. “The genie generates; the humans verify and steer.” That gap is not idle time. It is the irreducible work.
This is why “green tests, therefore correct” is not verification when the same agent wrote the tests. The expensive load doesn’t disappear because output got cheaper. It concentrates. Same forty hours, every line item moved: less hand writing code, far more reviewing, integrating, and keeping the specs and context current enough that next week’s drafts stay trustworthy.
Amazon is living this in real time
The most striking evidence comes from Amazon itself, which is running the experiment on its own traditions. When AWS carved out an agentic AI division, VP Swami Sivasubramanian deliberately reorganized it back to two pizza teams — the principle much of the 1.5 million person company had outgrown. His reasoning: projects that once required 30 to 40 people can now be done by teams of six to eight. (GeekWire)
The receipts are concrete. Amazon Quick — the desktop app that connects email, calendar, Slack, and documents into one AI workspace — was built by about six engineers and shipped in three months. The team wrote the classic PRFAQ after the product was already in beta, because building the demo had become faster than writing the six page narrative. Internally, a rebuild of the Bedrock inference engine was done by six engineers in 76 days, a project originally scoped at 30 developers over 12 to 18 months.
Notice what did not happen. The team did not shrink. Six to eight people is squarely a two pizza team. What changed is the charter: the same size team now owns a mission that used to need five times the roster. Sivasubramanian frames it exactly this way — the same number of people pursuing a bigger charter. And the roles inverted along the way: product managers write code, engineers make product calls. His own hard won lesson underlines where the weight moved. Rebuilding an old replication engine with an agent, he spent four frustrating nights babysitting output until he realized he had never given the agent the tools to test itself. Once he wrote the right spec and testing environment, it finished in two hours. “The bottleneck is not about the time it takes to build something. The bottleneck is about crafting the right specification and the tests.”
That is the whole argument in one executive’s mouth. The scarce resource is human judgment applied to specs, tests, and verification — not hands on keyboards.
The one pizza counterargument, and why it’s the exception
An honest post has to steelman the other fork. Mimacom argues for the one pizza team: two to three product builders orchestrating an agent layer, owning a full value stream, yielding roughly a 60% reduction in team size per product. Every goes further with the “two slice team” — one person per product. Dan Shipper runs four products with four people; Monologue, at 143,000 lines, is written almost entirely by one engineer with agents. These are real products, not weekend demos.
The counterargument is correct about something important: for greenfield products with a homogeneous stack and a clear owner, execution capacity really has stopped being the constraint, and a single builder can now own what used to take a squad. But look closely at how those organizations sustain it. Every doesn’t actually run solo teams — it surrounds them with internal agencies, floating designers, and freelance senior engineers who “dip in and out” on the hairy problems agents fumble. The load didn’t vanish. It got pushed to a shared layer around the solo builder. And the deep specialist work — the tacit, hard to verify knowledge — remains stubbornly human. The one person team is what the two pizza team looks like when you only count the person at the counter and ignore the kitchen behind them.
How to redefine the rule
Stop sizing teams by headcount. Size them by two things: bounded cognitive load and human verification capacity.
Treat agents as first class team members that expand throughput but concentrate the human work into judgment, spec writing, and review. The team number may well hold — five to eight is still a reasonable ceiling for coherent ownership — but its composition inverts: fewer hands writing code, far more expensive judgment per human.
Here is a reframe you can apply at your next org review, in four questions:
- Whose cognitive load does this team reduce, and what load flows back? Run every team through it. If nothing flows back, you have miscounted.
- Can this team pay for its returned load? A team can only get leaner if it can afford the verification, integration, and judgment the agents hand back. If it cannot, you get the same humans and less agent trust.
- Is the charter growing to match the throughput? Falling execution cost should buy a bigger mission, not just a smaller roster.
- Where does tacit specialist knowledge live? That work is the least changed by agents. Draw your boundaries to protect it, not dissolve it.
The two pizza team survives the agent era. It just stops being a rule about pizza. Do agents eat pizza? No — and that’s exactly why the number of pizzas was never the thing to measure. The right unit was always the load the team can carry and verify. Size for that, and the table sets itself.
Here is a question worth sitting with before your next org review: if you gave every engineer on a team a coding agent that never sleeps and writes most of the first draft, would you make the team smaller, or would you keep it the same size and hand it a bigger charter? Most leaders answer “smaller” on reflex. That reflex is expensive, and it misreads what the two pizza team was ever for.
The two pizza team is Amazon’s most exported piece of org design, and it is about to be its most misunderstood. As agents absorb a growing share of the keyboard work, the temptation is to shrink the roster to match. But the rule was never counting mouths at the table. It was counting something harder to see.
The rule was always a proxy
The origin story is well worn. At a 2002 retreat, when managers asked for more communication, Jeff Bezos reportedly shot back that “communication is terrible” and restructured the company around small autonomous teams — small enough to feed with two pizzas. The number, five to eight people, was never the point. It was a proxy for two things that actually mattered.
The first is communication overhead. Coordination cost grows faster than linearly with headcount; every person you add multiplies the channels that must stay in sync. Capping team size caps that overhead. The second is single threaded ownership: a team small enough to hold one coherent mission in its collective head, accountable end to end for a product it can build, ship, and operate without passing work across departmental fences.
Underneath both proxies sits the real binding constraint: cognitive load. A team can only hold so much of a domain in its head before delivery slows, quality drops, and people burn out. The two pizza rule was a crude but durable way to keep one team’s cognitive budget from overflowing. Read it that way and the agent question sharpens immediately. Agents don’t change whether the rule applies. They change what the rule is measuring.
What agents actually change
At the Pragmatic Summit, Martin Fowler framed the fork precisely: “Are we seeing two pizza teams becoming one pizza teams because agents don’t eat pizza, or do we see two pizza teams staying and becoming much more effective and capable? My bet is on more effective two pizza teams.” (Hands-On Architects)
It sounds like a headcount question. It is really a question about where the load goes. Here is the mechanism that resolves it: an agent relocates cheap cognitive load and returns expensive load. The boilerplate, the glue code, the first draft implementation, the codebase spelunking — all the work that was tedious but not hard — moves to the agent. What flows back to the humans is the work that was never cheap: verification, integration, and judgment. Someone still has to decide what to build, review what came back, and own the call when the agent cannot make it.
Kent Beck, on the same stage, put the brake on the naive reading with four words: “AI is an amplifier.” An amplifier does not subtract band members; it changes what each one carries. He described pairing with an agent and, counterintuitively, praising its slowness — the three minute gaps where the humans talk about naming, about conditionals, about what they should be doing next. “The genie generates; the humans verify and steer.” That gap is not idle time. It is the irreducible work.
This is why “green tests, therefore correct” is not verification when the same agent wrote the tests. The expensive load doesn’t disappear because output got cheaper. It concentrates. Same forty hours, every line item moved: less hand writing code, far more reviewing, integrating, and keeping the specs and context current enough that next week’s drafts stay trustworthy.
Amazon is living this in real time
The most striking evidence comes from Amazon itself, which is running the experiment on its own traditions. When AWS carved out an agentic AI division, VP Swami Sivasubramanian deliberately reorganized it back to two pizza teams — the principle much of the 1.5 million person company had outgrown. His reasoning: projects that once required 30 to 40 people can now be done by teams of six to eight. (GeekWire)
The receipts are concrete. Amazon Quick — the desktop app that connects email, calendar, Slack, and documents into one AI workspace — was built by about six engineers and shipped in three months. The team wrote the classic PRFAQ after the product was already in beta, because building the demo had become faster than writing the six page narrative. Internally, a rebuild of the Bedrock inference engine was done by six engineers in 76 days, a project originally scoped at 30 developers over 12 to 18 months.
Notice what did not happen. The team did not shrink. Six to eight people is squarely a two pizza team. What changed is the charter: the same size team now owns a mission that used to need five times the roster. Sivasubramanian frames it exactly this way — the same number of people pursuing a bigger charter. And the roles inverted along the way: product managers write code, engineers make product calls. His own hard won lesson underlines where the weight moved. Rebuilding an old replication engine with an agent, he spent four frustrating nights babysitting output until he realized he had never given the agent the tools to test itself. Once he wrote the right spec and testing environment, it finished in two hours. “The bottleneck is not about the time it takes to build something. The bottleneck is about crafting the right specification and the tests.”
That is the whole argument in one executive’s mouth. The scarce resource is human judgment applied to specs, tests, and verification — not hands on keyboards.
The one pizza counterargument, and why it’s the exception
An honest post has to steelman the other fork. Mimacom argues for the one pizza team: two to three product builders orchestrating an agent layer, owning a full value stream, yielding roughly a 60% reduction in team size per product. Every goes further with the “two slice team” — one person per product. Dan Shipper runs four products with four people; Monologue, at 143,000 lines, is written almost entirely by one engineer with agents. These are real products, not weekend demos.
The counterargument is correct about something important: for greenfield products with a homogeneous stack and a clear owner, execution capacity really has stopped being the constraint, and a single builder can now own what used to take a squad. But look closely at how those organizations sustain it. Every doesn’t actually run solo teams — it surrounds them with internal agencies, floating designers, and freelance senior engineers who “dip in and out” on the hairy problems agents fumble. The load didn’t vanish. It got pushed to a shared layer around the solo builder. And the deep specialist work — the tacit, hard to verify knowledge — remains stubbornly human. The one person team is what the two pizza team looks like when you only count the person at the counter and ignore the kitchen behind them.
How to redefine the rule
Stop sizing teams by headcount. Size them by two things: bounded cognitive load and human verification capacity.
Treat agents as first class team members that expand throughput but concentrate the human work into judgment, spec writing, and review. The team number may well hold — five to eight is still a reasonable ceiling for coherent ownership — but its composition inverts: fewer hands writing code, far more expensive judgment per human.
Here is a reframe you can apply at your next org review, in four questions:
- Whose cognitive load does this team reduce, and what load flows back? Run every team through it. If nothing flows back, you have miscounted.
- Can this team pay for its returned load? A team can only get leaner if it can afford the verification, integration, and judgment the agents hand back. If it cannot, you get the same humans and less agent trust.
- Is the charter growing to match the throughput? Falling execution cost should buy a bigger mission, not just a smaller roster.
- Where does tacit specialist knowledge live? That work is the least changed by agents. Draw your boundaries to protect it, not dissolve it.
The two pizza team survives the agent era. It just stops being a rule about pizza. Do agents eat pizza? No — and that’s exactly why the number of pizzas was never the thing to measure. The right unit was always the load the team can carry and verify. Size for that, and the table sets itself.