Skip to content

Claude Code Plan Mode: Turn the Plan Into E2E Tests

ILIA KARPENKO6 min read

Claude Code plan mode is a read-only permission mode: Claude reads the codebase, asks questions, and writes a step-by-step plan, but it cannot edit files or run changes until you approve that plan. The plan is a Markdown document describing what the feature should do and how it will be built. That makes it a usable spec for testing too, and feeding it into test case generation and then into Playwright MCP produces end-to-end tests that check what you meant to build instead of what the code happens to do.

The usual AI testing loop runs the other way. Someone finishes a feature, points an agent with browser tools at it, and asks for tests. The agent clicks through the app, sees what happens, and writes assertions that match what it saw. If the feature has a bug, the test asserts the bug. The plan you approved an hour earlier said what should happen, and nobody showed it to the agent.

How to use plan mode in Claude Code

Start a session in plan mode with claude --permission-mode plan, or press Shift+Tab during a session to cycle permission modes until the status line shows plan mode. To make every session in a project start that way, set "defaultMode": "plan" under permissions in .claude/settings.json.

In plan mode, describe the feature and let Claude read the code. It comes back with a plan and asks you to approve it. Before approving, press Ctrl+G to open the plan in your editor, change what is wrong, and save. Claude works from the edited version.

Plans are saved as Markdown files. By default they live outside the repository; set plansDirectory in .claude/settings.json to a folder inside the project, such as "plansDirectory": "./plans", and each plan becomes a file you can commit, review, and hand to the next step.

Free and open sourceCasely writes these cases for youAttach a spec and one file of your team's existing test cases in Claude. Casely copies your columns, names the gaps it found in the spec, and exports a single Excel file your tracker imports in one pass.Read the install docs

The workflow: plan, build, test cases, E2E

StepWho does itInputOutput
1. PlanYou and Claude Code in plan modeThe feature requestAn approved plan in plans/
2. BuildClaude Code in normal modeThe planThe implementation
3. Write test casesCasely, reviewed by youThe planTest cases with steps and expected results
4. Write E2E testsClaude Code with Playwright MCPThe test casesPlaywright spec files
5. Run and reviewPlaywright Test and youThe specsA passing suite you trust

1. Plan with expected behavior in it. Plans describe implementation by default: which files change, which functions get added. Ask Claude to include an "Expected behavior" section with each user-visible outcome stated as something you could observe: "After the fifth failed login the account is locked for 15 minutes and the correct password is rejected with a lockout message." Those sentences become expected results later, so make them concrete before you approve.

2. Build from the plan. Approve it and let Claude implement. If the implementation drifts, say a limit changes from 15 minutes to 30, update the plan file in the same commit. A stale plan produces test cases that fail for the wrong reason.

3. Turn the plan into test cases. Give the plan file to Casely, a free Claude skill that turns requirement documents into test cases. It reads the plan as a spec, proposes a test plan first (coverage per feature, a case count, and the gaps it found in the plan), and writes cases only after you approve that scope. Each case has preconditions, steps, test data, and an expected result taken from the plan, covering negative paths and boundaries as well as the happy path. Export them as Markdown into the repo, next to the plan.

4. Hand the cases to Playwright MCP. In Claude Code with Playwright MCP connected, give the agent one case at a time: "Implement test case LOGIN-003 from tests/cases/login.md as a Playwright test. Walk it in the browser first, use stable locators, and assert the expected result exactly as written." The agent opens the app, performs the steps, finds the locators, and writes the spec. The expected result is already decided, so the agent's job is to check it, not to invent it.

5. Run and review. Run the suite with npx playwright test. When a test fails, compare the failure to the case before changing anything: either the app is wrong and the test found a bug, or the case was wrong and the plan needs a fix. Reviewing the diff of a generated spec against its case takes a minute per test.

Why the tests come out better

The difference is where the expected result comes from.

Agent explores, then writes testsPlan, then cases, then tests
Source of expected resultsWhat the app did during the sessionWhat the approved plan says it should do
Bug present during generationGets encoded as expected behaviorFails the test
Negative and boundary pathsOnly the ones the agent stumbled intoListed in the cases before any code runs
Review effortRead every assertion and guess the intentCompare each spec to a written case
When the feature changesRegenerate and hopeUpdate the plan, regenerate the affected cases

Researchers have measured a version of this problem, which they call the misguidance effect: models writing tests from buggy code tend to assert the bug, and swapping the code for a written specification nearly doubled the number of tests that caught a real defect. A plan approved in plan mode is that written specification, and you already have it.

Where this workflow breaks

Plans that only describe code. "Add a lockoutUntil column and check it in authService.login" tells a developer what to type and tells a tester nothing. Without the expected-behavior section, the test cases have nothing to assert against beyond the happy path.

Plans nobody updates. Plan mode makes planning cheap, and that makes it easy to approve a plan, change direction halfway through, and never touch the file again. Generate test cases from a plan that no longer matches the code and every mismatch looks like a bug.

Agents that soften assertions. When the app and the case disagree, an agent asked to "make the test pass" will sometimes change the assertion to match the app. Tell it to stop and report the mismatch instead, and treat any edited expected value in a diff as a question for a person.

Flows the browser cannot see. A plan can promise a background job, an email, or a rate limit enforced by the API. Playwright MCP sees the page. Cases like these need an API or integration test, and the test cases should say so instead of forcing a browser check.

What is worth automating

Most of this pipeline is mechanical once the plan exists: turning plan sentences into structured cases, walking each case in a browser, finding locators, writing the spec. That is the part to give to Casely and to an agent with Playwright MCP, or to playwright-cli if context budget matters. Not every case needs a browser test: risk-based testing picks the cases worth the cost, and the test pyramid says most checks belong lower. The judgment stays at the two ends: writing a plan whose expected behavior is right, and deciding, when a test fails, whether the code or the plan is wrong.

Open sourceThe skill lives on GitHubMIT licensed, no account, nothing to install on your machine. Star it if it saves you an afternoon.View the repository