Skip to content

Code Coverage vs Test Coverage: What Each Number Actually Tells You

ILIA KARPENKO4 min read

Code coverage measures which lines, branches, or functions actually executed while the tests ran, reported as a percentage of the codebase. Test coverage measures which requirements, features, or user-facing behaviors got tested at all, and it isn't a percentage of code, it's a percentage of a requirements list. One is a white-box metric produced by an instrumented test run. The other is a black-box judgment call about whether the right things were checked, and no tool can generate it automatically because it depends on a list of requirements that has to exist somewhere first.

Teams that report only code coverage tend to over-trust it, because a single clean number feels like an answer. A codebase can sit at 90% code coverage and still ship a broken checkout flow, if the tests that ran those lines only asserted that a function didn't throw, never that it returned the right total. The line executed. Nothing about that line's actual behavior was verified. Code coverage tracks whether code ran, not whether anyone checked what it did.

What each metric actually measures

Code coverageTest coverage
Question it answersWhich lines of code executed during testing?Which requirements or behaviors got tested?
Unit of measurementLines, branches, statements, or functionsFeatures, requirements, or user scenarios
Produced byA coverage tool instrumenting the test run (Istanbul, JaCoCo, coverage.py, and similar)A person mapping test cases to a requirements or feature list
Can be 100% and still miss bugsYes: a line can run with an assertion that never actually checks the outputYes: a requirement can be "tested" by a shallow case that only checks the happy path
AutomatableFully, it's a byproduct of running the existing test suitePartially: tracking which requirements have a mapped test is a spreadsheet or tool problem, deciding whether that test is any good is not

Branch coverage is the sharper version of code coverage worth knowing about specifically: it tracks whether both the true and false paths of every conditional actually ran, not just whether the line containing the condition executed. A test suite can hit 100% line coverage on a function with an if/else by only ever exercising the if branch, if that line still counts as "covered" the moment either path runs once. Branch coverage catches that gap; plain line coverage does not.

Free and open sourceCasely writes these cases for youAttach a spec and one file of your team's existing test cases in Claude. Casely copies your columns, names the gaps it found in the spec, and exports a single Excel file your tracker imports in one pass.Read the install docs

Why a high number can still hide broken software

100% code coverage with weak assertions is coverage theater: every line ran, so the number looks complete, but a test that calls a function and checks only that it didn't crash gives no signal about correctness. The line is covered. The behavior is not verified. This is also how coverage numbers get gamed under pressure, deliberately or not, when a team is told to hit a target: it's faster to write a test that touches a line than one that actually proves the line does the right thing, and the coverage tool can't tell the difference.

Test coverage has the opposite failure mode. It can look complete on paper, every requirement in the tracker has at least one linked test case, while the individual tests are shallow: one happy-path case per requirement, no edge cases, no negative tests. The requirements list says "done." The actual testing depth says otherwise, and nothing in a coverage percentage catches that gap either.

Where it fails

Neither metric measures test quality. Both are counting metrics: how much code ran, how many requirements have a linked test. Counting is not the same as verifying that the test actually asserts something meaningful, and a team that optimizes for either number in isolation will find the fastest way to move it, not the most useful one.

Code coverage says nothing about missing requirements. If a requirement was never written down, or a feature exists that nobody documented, no line of test code will ever "miss" it in a coverage report, because coverage only measures code that exists. A whole unbuilt or undocumented feature is invisible to code coverage by definition.

Test coverage depends entirely on the requirements list being complete and current. If the requirements document is stale, test coverage against it can read 100% while covering a version of the product that shipped two releases ago. The metric is only as trustworthy as the list it's measured against.

What's worth automating

Code coverage is close to free once a test suite exists: instrumentation runs automatically, the percentage updates on every build, and most CI platforms will flag a pull request that drops the number. That part needs no human judgment beyond deciding whether a drop matters for this specific change. Mapping test cases to requirements can be tracked mechanically too, a spreadsheet or a test management tool linking each case to a requirement ID isn't hard to maintain. Judging whether that test actually proves the requirement works, and whether the requirements list itself still matches the product, stays something a person has to check.

Casely generates test cases from requirements, which sets up the input side of test coverage: a documented requirement gets a mapped test case from the start, instead of coverage getting backfilled after code already shipped. It doesn't replace reading the assertions inside that test case, or confirming the requirement itself still describes what the product actually does.

Open sourceThe skill lives on GitHubMIT licensed, no account, nothing to install on your machine. Star it if it saves you an afternoon.View the repository