# Production Runbook This runbook is the operational checklist for releasing and monitoring the transcription system. ## 1. Pre-release gate checklist 1. Run the full suite: `uv run pytest` 2. Confirm contract guardrails are green: - `uv run pytest tests/test_meta_contract_guards.py` 3. Confirm health endpoint includes worker liveness payload (`/healthz` returns `worker.state`). 4. Confirm required runtime settings are present in deployment environment: - `OPENROUTER_API_KEY` - `DATABASE__*` - filesystem paths for data/logs/backups. 5. Confirm schema contract alignment is current: - `src/transcription/db/models.py` - `docs/schema.md` ## 2. Release execution steps 1. Deploy artifact/config to target environment. 2. Validate service startup: - `/healthz` responds `200` - `worker.state` is `running` 3. Execute one smoke workflow: - create a document/job with at least one source - verify terminal job outcome updates - verify execution evidence row appended 4. Verify log flow: - stdout aggregation receives events - file logs are written under `./data/logs` ## 3. Rollback triggers and actions ### Trigger conditions 1. `/healthz` reports `worker.state=failed` 2. Repeated provider timeout/error spikes beyond normal baseline 3. Evidence write failures or DB persistence failures ### Actions 1. Roll back app artifact and config to previous release. 2. Restart service and re-check `/healthz`. 3. Re-run smoke workflow and confirm worker returns to `running`. 4. Preserve incident evidence: - `./data/logs` - relevant DB rows (`job`, `job_source`, `execution_attempt`) ## 4. Post-release monitoring checklist ## First 24 hours 1. Monitor `/healthz` periodically for `worker.state`. 2. Track job terminal distribution (`transcribed`, `partial_success`, `failed`). 3. Sample timeout/error categories for abnormal increase. 4. Spot-check new `execution_attempt` records for append-only growth and timing metadata. ## First 72 hours 1. Re-check error/timeout trend versus 24h baseline. 2. Verify no recurring worker-failed states. 3. Verify storage growth and rotation behavior under `./data/logs`. 4. Confirm incident response notes are captured for any production anomalies. ## 5. Operator playbook for common incidents ### Worker failed 1. Check `/healthz` payload (`error_id`, `error_category`). 2. Locate matching error in logs. 3. If non-transient defect persists, roll back. ### Provider timeout spike 1. Confirm provider reachability and rate limits. 2. Review timeout frequency and impacted job volume. 3. If sustained, execute rollback criteria and notify stakeholders. ### Partial-success increase 1. Inspect affected `job_source` and `execution_attempt` records. 2. Confirm failures are category-aligned (`external`/`timeout`/`internal`). 3. Triage whether issue is source quality, provider, or runtime regression.