Test Pyramid: How Much of Each Test Type You Actually Need
The test pyramid is a shape for a test suite: many fast unit tests at the bottom, fewer integration tests in the middle, and the fewest end-to-end tests at the top. The shape isn't decorative. Each layer costs more to write, run, and maintain than the one below it, so the pyramid puts the cheap, fast layer at the base and reserves the expensive, slow layer for what only it can catch.
Get the shape backwards and the cost shows up as suite runtime and flaky builds rather than as a line item anyone budgeted for. A suite with a thousand end-to-end tests and a hundred unit tests takes an hour to run, breaks on unrelated UI changes, and tells a developer nothing about which function is wrong when it fails, just that something, somewhere in a five-minute browser session, didn't work. The same coverage built mostly from unit tests runs in seconds, points at the exact function that broke, and reserves the slow browser session for the handful of flows that actually need one.
What each layer catches
| Layer | What it tests | Speed | What breaks it |
|---|---|---|---|
| Unit | One function or class, in isolation, with dependencies mocked or stubbed | Milliseconds, thousands can run in seconds | Logic errors: wrong calculation, wrong condition, wrong return value |
| Integration | Two or more real components working together: a service and its database, an API and a queue | Seconds per test, tens to hundreds run in a normal CI job | Contract mismatches: a schema change, a wrong field name, a broken assumption between two pieces that each pass their own unit tests |
| End-to-end | A full user journey through the real (or near-real) system, usually through the UI | Seconds to minutes per test, a full suite can take significant time | Anything that only shows up when everything is wired together: routing, auth, a real database round-trip, third-party integrations |
A bug can pass every unit test and still break in production because the unit tests mocked away the exact interaction that was wrong. That's what the integration layer exists for. And a bug can pass unit and integration tests and still break because the browser renders a form differently than the test assumed, or a redirect fires in the wrong order. That's what the end-to-end layer exists for, and it's also why nobody has found a way to make it cheap: it has to exercise the real thing, slowly, to catch what the real thing does.
Free and open sourceCasely writes these cases for youAttach a spec and one file of your team's existing test cases in Claude. Casely copies your columns, names the gaps it found in the spec, and exports a single Excel file your tracker imports in one pass.Read the install docsThe classic starting ratio
A commonly cited starting point, popularized by Mike Cohn's writing on agile testing, is roughly 70% unit, 20% integration, 10% end-to-end. Treat it as a rough starting shape, not a target to hit exactly. The right ratio for a given codebase depends on where the bugs actually come from: a service with complex business logic and few external dependencies leans harder into unit tests, while a thin API gateway that mostly routes and transforms requests gets more value from integration tests than from a large unit layer testing logic that barely exists.
The number to watch isn't the ratio, it's the shape. If end-to-end tests outnumber unit tests, or a team is writing new end-to-end tests to catch bugs that a unit test could catch in a tenth of the time, the pyramid has inverted, and the suite will get slower and flakier as it grows rather than staying proportionate to the codebase.
Where it fails
A codebase with almost no internal logic doesn't need a big unit layer. A thin CRUD API that mostly passes data through to a database has few functions worth unit testing in isolation, because there's little logic to isolate. Forcing a 70% unit-test ratio onto a project like that means testing getters and setters for coverage numbers, not because those tests catch anything.
Unit tests can pass while the integrated system is broken. Mocking a dependency means trusting that the mock behaves like the real thing. When the real API changes its response shape and nobody updates the mock, every unit test built on that mock keeps passing while the actual integration is broken. That's a contract test's job, not a reason to skip unit tests, but it's a real blind spot in a suite that's almost all unit tests and almost no integration coverage.
The pyramid says nothing about what to test, only how much of each kind. A team can hit a perfect 70/20/10 split and still miss the one workflow that matters most to users, because the ratio measures test count and runtime, not risk. The pyramid is a cost model, not a coverage strategy, and it works best alongside something that decides what's worth testing at all.
What's worth automating
The layer split itself is worth watching automatically: most CI platforms can report test count and runtime by directory or tag, and a shape that's drifting toward more end-to-end tests than the codebase's logic complexity justifies is visible in that report before it becomes a slow, flaky suite. Deciding which specific behavior deserves a unit test versus an integration test versus an end-to-end check stays a judgment call about where the actual risk sits in a given change, the same judgment impact analysis and risk-based testing apply to picking what to retest at all.
Casely generates test cases from requirements, which is a separate question from which automation layer a given case belongs in once it exists. A requirement like "a returning customer sees their saved address at checkout" might become a unit test on the address-lookup function, an integration test on the checkout service, and an end-to-end test on the full flow, and picking which of those (or all three) is worth writing is still a call a person makes based on what could actually break.