J.BLOG

Verifying a verification script with deliberately broken input

Jaemyeong Jin···8 min read
한국어

Automated checks before a release can all pass while the output is still missing a required file, or holds the same content for locales that should differ. This happens with scripts that check the resources and metadata of a cross-platform app. The checks ran to the end, yet the rule they were meant to protect was never examined.

Whether a check script really catches a defect only shows when it is given input that has one. This post goes through putting a single defect into a copy of a good sample and confirming that the check fails with the expected category. Combinations of platform builds and tests are out of scope. I ran the example with a small validator on Python 3.14.5, and the result is written below the example.

What exit status 0 does not tell you

There are several ways for a check to skip the rule and still end in success.

  • A glob does not match, so there are 0 targets.
  • A missing tool or environment variable is treated as success.
  • The generating side and the checking side share the same wrong default.

In each case the exit status can be 0, so the value 0 alone cannot tell whether the rule was checked. The pytest exit codes separate these two as well. The code is 0 when all collected tests pass, and 5 when no tests were collected.

I suggest making the same distinction in a check script you write yourself. Failing to find a required output is a problem with the input contract, and a missing tool or input in a required CI job is a problem with the setup. Skipping an optional check is a separate state with the reason recorded. The candidate states are pass, validation-failure, invalid-setup, and skip. These four are my own design, not something the pytest documentation defines.

Judging a failure needs more information than the exit status.

  • category or rule_id
  • The number of inputs processed and the number of required checks
  • Exit statuses that separate environment errors from contract violations
  • A distinction between an empty target set and an allowed skip

If the only requirement is “it exited abnormally,” a syntax error or a permission problem can be miscounted as catching the defect. For a sample that breaks the image size rule, the expected category is wrong_size. An internal_error caused by failing to read the file is not evidence that the rule worked. CI compares against a stable category value, and the descriptive message is for investigating the cause.

A sample with exactly one defect

In this post, a negative control means a defective sample that has what it takes to be rejected. The term is not carried over as defined in statistics or experimental fields. Put the disallowed value dark into a copy of the good data and require invalid_value, and everything from reading the input to reporting the error is checked in one go.

A sample is made by copying a reviewed fixture to a temporary location, changing it, and discarding it afterwards. The files that actually ship are not touched. If good data is rejected, or defective data is accepted, the check has not met the conditions for being trusted.

Putting a contract violation into a good sample is what this post calls a mutation. Each run changes one rule only. Examples are inserting the disallowed value dark, deleting a required file, making the content of two locales identical, and changing the size information of an image. With several defects at once it is hard to tell which one was caught, and a tool that stops at the first failure never reaches the later rules. So one mutation is paired with one expected category.

Tools such as Stryker and PIT also use the word mutation. They change the program code and watch whether the tests react. In Stryker’s description of states, a mutant is killed when at least one test fails and survived when all tests pass. PIT’s basic concepts describe equivalent mutations, changes that do not alter behavior. Tests cannot detect those. The same holds when input is changed by hand, so the cases chosen have to really break the contract. Up to here the content is from the documentation, and carrying the idea over to checks on input files is my own reading. The idea of trusting a test because it was seen to fail is also in a Google Testing Blog post.

An example that confirms the baseline first

Before any defective sample goes in, the good sample has to pass. The exit status must be 0 with 0 failures. Without this condition, an abnormal exit caused by a defect that was already there can be counted as catching the new one.

The example assumes that the JSON result of the validator contains the category of each entry in failures, and checked. Exit status 0 is a pass, 1 is a contract violation, and anything else is a setup or internal problem. There are 4 input contracts. The first part is the function that runs the validator and reads the result, followed by the check of the good sample.

import json
import shutil
import subprocess
import tempfile
from pathlib import Path

FIXTURE = Path("validator-fixture")
PASS = 0
VALIDATION_FAILED = 1
REQUIRED_CHECKS = 4


def validate(root: Path) -> tuple[int, set[str], int]:
    result = subprocess.run(
        ["python", "validate.py", str(root), "--json"],
        check=False,
        capture_output=True,
        text=True,
    )
    report = json.loads(result.stdout)
    failures = {item["category"] for item in report["failures"]}
    return result.returncode, failures, report["checked"]


baseline_code, baseline_failures, baseline_checked = validate(FIXTURE)
if baseline_code != PASS or baseline_failures or baseline_checked != REQUIRED_CHECKS:
    raise SystemExit(
        f"invalid baseline: checked={baseline_checked}, "
        f"failures={sorted(baseline_failures)}"
    )

The next part applies each of the four mutations to a copy in a fresh temporary directory. It uses the names from the block above.

cases = [
    ("invalid-value", "invalid_value"),
    ("missing-file", "missing_file"),
    ("duplicate-content", "duplicate_content"),
    ("wrong-size", "wrong_size"),
]
if not cases:
    raise SystemExit("no mutation cases")

for mutation, expected_category in cases:
    with tempfile.TemporaryDirectory() as temp:
        root = Path(temp) / "fixture"
        shutil.copytree(FIXTURE, root)
        subprocess.run(["python", "mutate.py", mutation, str(root)], check=True)

        code, failures, checked = validate(root)
        introduced_failures = failures - baseline_failures

        if code != VALIDATION_FAILED or checked == 0 or expected_category not in introduced_failures:
            raise SystemExit(
                f"survived mutation: {mutation}, checked={checked}, "
                f"failures={sorted(failures)}"
            )

When the list of mutation cases is empty, the script fails before it enters the loop. The failures after the mutation, minus the failures of the baseline, are the newly introduced failures. A case passes when the exit status is 1, the number of checks is not 0, and the introduced failures contain the expected category. In this example the baseline failures have to be empty to get this far, so the set difference equals the set of failures after the mutation. A case still passes when other categories appear alongside the expected one.

When a condition is not met, the script raises SystemExit. It does not use assert, because Python’s -O option removes assert statements. The CI decision is made with explicit condition branches.

validate.py, mutate.py, and the fixture itself are not in this post. For the run I wrote a small separate validator with four rules. With the normal validator, all four mutations were detected. When I changed the validator to report the expected category while performing 0 checks on a mutated copy, the version without the checked == 0 condition ended as a pass, and the version with it ended with survived mutation. That condition and the check for an empty list were added to the original example. Handling of JSON that cannot be parsed, and separate handling of environment errors, are not in the example either.

Keeping the judgment apart from the generator

If the generating code and the checking code use the same decision function, or the fixture is built only from the latest generated output, the same error can be hidden on both sides. Sharing a neutral schema is fine. The criteria for allowed values, sizes, and uniqueness come from the written contract and from material a person has reviewed.

In CI, I suggest keeping one defective sample per core rule and running them when the validator or the discovery code changes. The order is this.

  1. Confirm that the tool and the fixture exist.
  2. Confirm the state of the good sample and the minimum number of checks.
  3. Apply one mutation to each copy.
  4. Confirm that the expected category was detected.

A required CI job fails when a tool or input is missing, when the number of tests is 0, when the number of checks is 0, or when there are 0 mutation cases. Only skipping an optional item is allowed, with a reason. When a defect goes undetected, look at how files are discovered, how errors are reported, and how the expected category is set. Keeping the baseline result and the expected and observed categories as CI artifacts makes that investigation easier.

A few reviewed samples are enough to start. The example compares only the set of category values, though, so it cannot tell apart several failures of the same category in different places. When the location of a failure matters, use the combination of rule_id and a normalized relative path. The CI procedure looks at whether the number of checks is at least a minimum, while the example checks that the number for the good sample equals 4. The number of checks on a mutated copy is only tested for being 0 and is not compared with a minimum. A change that does not alter behavior cannot be detected by this method either.

광고Coupang Partners

이 포스팅은 쿠팡 파트너스 활동의 일환으로, 이에 따른 일정액의 수수료를 제공받습니다.