Test quality

Test quality is how much a test suite proves about behavior that matters: evidential assertions, meaningful coverage, determinism, and maintainability rather than test count or a green status.

What it means

Test quality is how much a test suite proves about the behavior that matters. Test count and a green badge measure something else: a protective suite fails when a real regression lands; a merely green one can pass with no assertions or with the paths users depend on skipped.

Martin Fowler is blunt about the usual proxy: “Test coverage is a useful tool for finding untested parts of a codebase. Test coverage is of little use as a numeric statement of how good your tests are.” Enji Guard’s audit documentation agrees: tests should “prove useful behavior, catch regressions, and remain maintainable,” and the audit “does not reward file count alone.”

How it works

Four measures separate protective suites from decorated ones.

MeasureWhat it establishesBlind spot
CoverageWhich lines or branches ran under testPIT’s mutation-testing docs: coverage “does not check that your tests are actually able to detect faults in the executed code”
Mutation scoreShare of seeded faults (mutants) a test caught; per Stryker, “the higher the percentage of mutants killed, the more effective your tests are”Runs the suite per mutant, so PIT recommends running it only against changed code
Flaky-test rateTests that pass and fail on the same code; causes per Fowler: isolation, asynchrony, remote services, time, resource leaks”Software Engineering at Google”: near 1% flakiness “the tests begin to lose value”
Decorative assertionsTests whose only check is assert true or expect(true).toBe(true); they raise counts and coverage while proving nothingNo metric flags them; someone must read the test

Why it matters for AI-written code

An agent asked to make the pipeline green can satisfy that gate with tests that execute code rather than tests that would fail on a wrong answer. Siddiq et al. evaluated Codex, GPT-3.5-Turbo, and StarCoder on JUnit generation: Codex “achieved above 80% coverage for the HumanEval dataset, but no model had more than 2% coverage for the EvoSuite SF110 benchmark.”

The generated tests also “suffered from test smells, such as Duplicated Asserts and Empty Tests.” Fowler “would be suspicious of anything like 100%”; generated tests earn the same suspicion. Weak tests also mislead the next agent, which reads them as proof that the behavior is protected.

How Enji Guard helps

The test quality audit begins by listing what makes a green suite suspect: committed skip markers, vacuous assertions, and flaky signals such as .only, sleeps, retries, and shared state. Skips are counted apart from flaky signals; “an explanation helps triage but does not make the skip healthy.”

Coverage is read project-wide: a backend number never stands in for repository coverage, a primary surface with no tests counts as 0, and a frontend without browser, end-to-end, or user-flow tests caps effective coverage at 60. Each of the 24 criteria receives 0 to 5; caps then limit the final score:

  • no tests found: 25
  • tests without a documented runnable command: 55
  • widespread decorative assertions: 65
  • committed skips that disable multiple tests, suites, or important paths: 79

The runbook’s principle behind the 65 cap: “Decorative assertions are actively misleading: assert true-style tests increase count or coverage without proving behavior.”

The paired test writing improvement adds a few focused tests for one valuable, currently testable behavior within the existing stack, without bootstrapping a framework, adding dependencies, changing runner configuration, or modifying production code; any review request it opens is never merged automatically. A rerun after the merge compares the same surfaces and skips, which holds the test area in the green zone as agents add features.

The audit is analysis-only: no issues, branches, commits, pull requests, or fixes, and no mutation testing, so its report carries no mutation score. Test cleanup, meant to repair flaky or decorative tests, remains a prototype: it neither locates such tests, edits the suite, nor opens a pull request.