Skip to content

Smoke Testing vs Sanity Testing: Which One Do You Run First?

ILIA KARPENKO5 min read

Smoke testing checks whether a new build is stable enough to test at all: does the app launch, does login work, does the main workflow run without crashing. Sanity testing checks something narrower: does the specific fix or feature that just changed actually work, without touching anything else. Both are shallow on purpose and both run fast, which is exactly why they get confused with each other in standups and ticket templates.

The mix-up costs more than a naming argument. A team that runs a sanity check after every deploy and calls it a smoke test is verifying one changed module while assuming the rest of the build still boots, an assumption nobody actually tested. A team that runs a full smoke suite after every single bug fix is burning ten minutes confirming the app still launches when the only thing that changed was one conditional in a checkout form. Neither mistake shows up immediately. It shows up three weeks later, when a build that "passed smoke" turns out to have a broken signup flow nobody exercised, or when a "sanity-tested" hotfix ships alongside a regression in an unrelated screen that the narrow check never looked at.

What each one actually checks

Smoke testingSanity testing
Question it answersIs this build stable enough to test further?Does this specific change work?
ScopeBroad and shallow: core workflows onlyNarrow and shallow: the changed area only
Runs afterA new build, a merge to main, a deploy to a test environmentA bug fix, a small feature, a targeted code change
Who usually runs itDevelopers or testers, often scripted into CITesters, usually by hand, right before or after a release candidate is cut
Parent categoryA form of build verification testing, a subset of acceptance testingA form of regression testing, but deliberately partial rather than full
What a failure meansStop. Don't run the rest of the suite against this build.The fix didn't work, or it broke something nearby. Investigate before shipping.

The scope line is the one that actually matters day to day. Smoke testing asks a yes-or-no question about the whole application: does it come up, does the critical path work, is there any point running the rest of the suite. Sanity testing never asks that question, it assumes the build is already fine and asks only about the piece that just moved.

Free and open sourceCasely writes these cases for youAttach a spec and one file of your team's existing test cases in Claude. Casely copies your columns, names the gaps it found in the spec, and exports a single Excel file your tracker imports in one pass.Read the install docs

When each one runs

A typical pipeline hits both, in this order:

  1. Build compiles and deploys. Nothing has been tested yet.
  2. Smoke test runs. Core paths only: app loads, auth works, the one or two workflows that represent "the product functions at all." Five to fifteen cases, usually automated, usually under two minutes. A failure here blocks everything downstream. There's no point running a full regression suite against a build that can't complete login.
  3. Deeper testing runs, whatever that means for the team: regression, exploratory, the case set for the features actually being verified this cycle.
  4. A specific bug gets fixed or a small feature ships. Sanity testing runs against just that change, confirming the fix works and skimming the one or two areas most likely to have broken alongside it.
  5. Release candidate is cut. Sanity testing runs again, faster and narrower than the full regression pass, as a last check before sign-off.

Smoke testing is a gate at the front of the pipeline. Sanity testing is a spot check that runs anywhere a targeted change needs a fast answer before something bigger commits to it.

Where the confusion comes from

Both are unscripted or lightly scripted, both skip edge cases on purpose, and both exist because running the full suite every time is too slow to be useful. That family resemblance is exactly why the terms get swapped in casual conversation. The actual difference is what each one assumes going in. Smoke testing assumes nothing and checks the whole system's basic viability. Sanity testing assumes the system is fine and checks one change against that assumption. Reverse the assumption and you've picked the wrong tool: running a five-case smoke check after a bug fix tells you the app still boots, not whether the fix works, and running a narrow sanity check on a brand-new build tells you nothing about the ninety percent of the app that check never touched.

Where it fails

A passing smoke test says nothing about correctness. It confirms the app didn't fall over, not that any feature behaves correctly. A build can pass every smoke check and still ship a checkout flow that charges the wrong amount, because smoke testing was never designed to catch that.

A passing sanity test only covers what someone thought to check. It's unscripted by design, which means its coverage depends entirely on the tester's judgment about what's "nearby" the change. A fix to a shared utility function can break three unrelated screens that nobody thought to sanity-check because they didn't look related on the surface.

Neither replaces a real regression suite. Both exist to make a decision fast, ship this build or not, sign off this fix or not, not to find every defect. Treating either one as sufficient coverage on its own is how a team ends up debugging in production what a full regression pass would have caught before release.

What's worth automating

Smoke testing is close to fully automatable: a fixed, small set of core-path checks that should run identically on every build is exactly the kind of test that belongs in CI, triggered automatically, with no judgment call involved in which cases to run. Sanity testing resists that more, because its whole value is a human deciding, right now, which areas are actually at risk from this specific change, a judgment that shifts with every diff.

Casely generates test cases from requirements, which sets up the deeper regression and exploratory passes that come after both of these quick checks, not the checks themselves. Knowing which five cases prove a build is alive, or which two screens are worth a sanity check after a given fix, still depends on someone reading the change and deciding what's actually at risk.

Open sourceThe skill lives on GitHubMIT licensed, no account, nothing to install on your machine. Star it if it saves you an afternoon.View the repository