All guides

AI tests pass, but the feature is still broken

Passing AI-written tests can miss the real defect. Inspect mocks, protect the acceptance criteria and verify the user journey against the actual app.

Written by the Kaplira teamKaplira makes a product for this problem. The guide also says what works without it.

The short answer

Passing tests show that the assertions succeeded under the test's setup. They do not prove that the real feature works when a mock replaces the component being checked, an assertion was weakened, or test code repairs the app before checking it. Reproduce the defect, verify that a relevant test fails for that reason, and rerun the same acceptance check after the fix.

  1. A test can accidentally hide the bug

    If a test replaces the broken function with a working stand-in, it can pass without exercising the defect. The same problem occurs when a browser helper supplies missing app behavior or an assertion is changed to accept the incorrect result. Inspect the setup and assertions alongside the implementation.

    The practical question is simple: if the production defect returned, would this test notice it?

  2. Reproduce a passing test that misses the real function

    This original illustrative example uses Node.js and no additional packages. In a new temporary directory, create order.mjs with a deliberate defect: the function accepts a nonempty order without checking whether the terms were accepted.

    export function canSubmitOrder(order) {
      return order.items.length > 0;
    }
    

    Create stand-in.test.mjs. This test replaces the function it claims to check:

    import assert from 'node:assert/strict';
    const canSubmitOrder = () => false;
    assert.equal(
      canSubmitOrder({ items: ['book'], termsAccepted: false }),
      false,
    );
    

    Now create real.test.mjs, which imports the actual implementation and checks the acceptance boundary:

    import assert from 'node:assert/strict';
    import { canSubmitOrder } from './order.mjs';
    assert.equal(canSubmitOrder({ items: ['book'], termsAccepted: false }), false);
    assert.equal(canSubmitOrder({ items: [], termsAccepted: true }), false);
    assert.equal(canSubmitOrder({ items: ['book'], termsAccepted: true }), true);
    

    Run node stand-in.test.mjs, then node real.test.mjs. We executed both: the stand-in test exited successfully with code 0; the real test exited with code 1 because the first assertion received true instead of false.

    Correct order.mjs without changing the assertions:

    export function canSubmitOrder(order) {
      return order.items.length > 0 && order.termsAccepted === true;
    }
    

    Running node real.test.mjs again exited with code 0. This verifies the three function-level cases. It does not verify the checkout screen or a saved order. The difference is the test boundary: the passing stand-in test never exercised the defective implementation.

  3. Write the observable acceptance condition first

    For an illustrative broken dropdown, define what a user must be able to do before asking the agent to fix it:

    On the real form, click Country and choose Portugal.
    The selected value must appear in the form.
    Submit and verify that the saved record keeps that value.
    The test must not replace the dropdown or repair its event handlers.
    

    The saved-record check is a separate integration requirement. A component test with a mocked API can check interaction and payload shape, but cannot establish that the actual backend saved anything.

    Ask the agent to reproduce the existing defect and save the failing command or browser observation. Confirm that the failure is the intended defect, rather than a missing dependency, unavailable server or incorrect selector. After the fix, rerun that same check against the same relevant setup.

    For a new feature, establish the acceptance criteria before implementation. Do not create a fake failure just to produce a red-to-green story.

  4. Review the tests as carefully as the fix

    Inspect changed test files, helpers and runner configuration. Look for deleted assertions, skipped cases, permissive snapshots, caught errors that no longer fail the test, and a fake implementation imported instead of production code.

    In browser tests, inspect page.evaluate, injected scripts and request interception. These APIs have legitimate uses; their presence is not proof of a bad test. The concern is whether they manufacture the behavior that the test claims to verify. Seeding a fixture is different from adding the missing click handler to the app under test.

    Playwright's guidance recommends checking behavior visible to the user. Start with what the user does and observes, then decide which supporting checks you need. A test of internal helper calls alone may pass while the form remains unusable.

  5. Keep mocks, but state the boundary they exclude

    Mocks make focused tests deterministic and let you exercise errors that are difficult to trigger safely. They become misleading when the result is described as evidence for a dependency they replaced.

    For example, Playwright's API mocking documentation shows how an intercepted request can receive a supplied response without reaching the actual API. That is useful for testing a screen's response to particular data. It does not test the real endpoint.

    Separate the claims: a mocked UI check passed; an API integration check passed; a complete save-and-reload journey passed. Report only the ones you ran. Keep third-party services mocked where appropriate, and exercise your own critical integration boundaries in a controlled test environment.

    When a test seems suspicious, temporarily reintroduce a small, known defect in an isolated local checkout and confirm that the relevant check fails. Restore the correct version immediately afterwards. This is a targeted diagnostic, not a requirement to break every feature on every run.

  6. Where Kaplira fits

    Kaplira connects required checks and review evidence to a governed change. Its test-review capability can flag supported mock and production-boundary patterns in changed test source; incomplete or unsupported analysis must remain explicitly unevaluated.

    That review is a signal for inspection. It does not run every user journey, prove that every mock is appropriate, or turn a passing suite into proof of production behavior. You still need meaningful assertions, the right integration environment and human review where judgment is required.

    For the broader workflow, see preventing AI code regressions and the limits of run evidence. A useful completion report says what was observed, what passed and what remains unverified.

Try it on a project you already have.

Kaplira records decisions and known mistakes from each governed run, links them to the files they affect, and hands them to the next agent that touches those files. The desktop app is free, no signup is required, and the run record stays on your machine by default.

Kaplira does not replace your tests or human review, and it does not guarantee correct code. Read the run evidence and security boundaries to see what the record proves.