Files
transcription/.github/skills/test-effectiveness-auditor/skill.md
T

3.5 KiB

name, description
name description
test-effectiveness-auditor Periodic reviewer for test-suite signal quality. Detects low-value or redundant tests, validates contract coverage, and recommends pruning or strengthening actions.

Test Effectiveness Auditor

Run a deterministic audit of test usefulness. Focus on whether tests catch real regressions, not whether they merely execute code.

When to Use

  • Monthly/quarterly test-health review.
  • Pre-release hardening when test count grows quickly.
  • After major AI-assisted test generation.
  • When suite runtime is increasing without clear quality gains.

Primary Objectives

  1. Identify tests that are weak, redundant, or non-diagnostic.
  2. Confirm critical contracts are guarded by meaningful assertions.
  3. Produce a prune/strengthen backlog with explicit risk and effort.

Normative References (Transcription Repo)

  1. docs/ver4/*
  2. docs/invariant/*
  3. .github/instructions/*.instructions.md
  4. tests/test_meta_contract_guards.py
  5. Contract-specific guards (tests/test_service_boundaries.py, tests/test_ui_boundaries.py, worker/evidence/media/error suites)

Deterministic Audit Checks

A. Contract Traceability

  • Each high-risk contract maps to at least one focused regression test file.
  • Missing mapping is a gap.

B. Assertion Strength

  • Flag tests that only assert status code, non-null, or “no exception” without validating state transitions or persisted outcomes.
  • Prefer assertions on domain effects: DB rows, status changes, error categories, evidence writes, or emitted payload shape.

C. Failure-Path Coverage

  • Critical paths must include negative-path tests (timeouts, provider errors, validation failures, cancellation paths, retries).
  • Happy-path-only coverage on critical modules is a gap.

D. Redundancy and Noise

  • Detect near-duplicate tests asserting the same behavior at multiple layers with no extra signal.
  • Recommend canonical location (unit/integration) and prune overlaps.

E. Mutation/Change Sensitivity

  • Prefer mutation testing for high-risk modules when practical.
  • If not run, identify tests likely to survive meaningful code mutations (low sensitivity).

F. Drift Guards

  • Verify config/doc/instruction contracts have deterministic guards and are current.
  • Ensure settings/docs synchronization checks remain active.

Evidence Standards

  • Every finding must include concrete file paths and line ranges.
  • No speculative claims.
  • Distinguish clearly between:
    • Confirmed ineffective tests
    • Likely weak tests (needs mutation/probe confirmation)

Output Format

Produce a Markdown report in docs/:

# Test Effectiveness Audit Report

## 1. Executive Verdict
- Effective / Effective with Conditions / Needs Remediation
- Top risks to confidence

## 2. Contract Coverage Matrix
| Contract | Guarding Tests | Signal Quality | Gap | Action |
| :--- | :--- | :--- | :--- | :--- |

## 3. Weak/Redundant Test Findings
| Finding ID | Location | Why Low-Signal | Risk | Recommendation |
| :--- | :--- | :--- | :--- | :--- |

## 4. Prune/Strengthen Backlog
| Task ID | Goal | Files | Acceptance Criteria | Validation |
| :--- | :--- | :--- | :--- | :--- |

## 5. Confidence Recommendation
- Go / Go with Conditions / No-Go for release confidence

Decision Rules

  • Do not recommend deleting a test unless equivalent or stronger coverage is identified.
  • Prefer strengthening assertions before adding more tests.
  • Prioritize deterministic contract guards over broad snapshot-style tests.