Files
transcription/docs/production-runbook.md
T
2026-08-20 16:35:20 -05:00

2.8 KiB

Production Runbook

This runbook is the operational checklist for releasing and monitoring the transcription system.

1. Pre-release gate checklist

  1. Run the full suite: uv run pytest
  2. Confirm contract guardrails are green:
    • uv run pytest tests/test_meta_contract_guards.py
  3. Confirm health endpoint includes worker liveness payload (/healthz returns worker.state).
  4. Confirm required runtime settings are present in deployment environment:
    • OPENROUTER_API_KEY
    • DATABASE__*
    • filesystem paths for data/logs/backups.
  5. Confirm schema contract alignment is current:
    • src/transcription/db/models.py
    • docs/schema.md

2. Release execution steps

  1. Deploy artifact/config to target environment.
  2. Validate service startup:
    • /healthz responds 200
    • worker.state is running
  3. Execute one smoke workflow:
    • create a document/job with at least one source
    • verify terminal job outcome updates
    • verify execution evidence row appended
  4. Verify log flow:
    • stdout aggregation receives events
    • file logs are written under ./data/logs

3. Rollback triggers and actions

Trigger conditions

  1. /healthz reports worker.state=failed
  2. Repeated provider timeout/error spikes beyond normal baseline
  3. Evidence write failures or DB persistence failures

Actions

  1. Roll back app artifact and config to previous release.
  2. Restart service and re-check /healthz.
  3. Re-run smoke workflow and confirm worker returns to running.
  4. Preserve incident evidence:
    • ./data/logs
    • relevant DB rows (job, job_source, execution_attempt)

4. Post-release monitoring checklist

First 24 hours

  1. Monitor /healthz periodically for worker.state.
  2. Track job terminal distribution (transcribed, partial_success, failed).
  3. Sample timeout/error categories for abnormal increase.
  4. Spot-check new execution_attempt records for append-only growth and timing metadata.

First 72 hours

  1. Re-check error/timeout trend versus 24h baseline.
  2. Verify no recurring worker-failed states.
  3. Verify storage growth and rotation behavior under ./data/logs.
  4. Confirm incident response notes are captured for any production anomalies.

5. Operator playbook for common incidents

Worker failed

  1. Check /healthz payload (error_id, error_category).
  2. Locate matching error in logs.
  3. If non-transient defect persists, roll back.

Provider timeout spike

  1. Confirm provider reachability and rate limits.
  2. Review timeout frequency and impacted job volume.
  3. If sustained, execute rollback criteria and notify stakeholders.

Partial-success increase

  1. Inspect affected job_source and execution_attempt records.
  2. Confirm failures are category-aligned (external/timeout/internal).
  3. Triage whether issue is source quality, provider, or runtime regression.