generated from john/python-template
gpt-5.3 codex review: Phase 7 and the addition of the new test-effectiveness-auditor skill.
Quality Gate / gate (push) Failing after 12s
Quality Gate / gate (push) Failing after 12s
This commit is contained in:
@@ -0,0 +1,84 @@
|
||||
# Production Runbook
|
||||
|
||||
This runbook is the operational checklist for releasing and monitoring the transcription system.
|
||||
|
||||
## 1. Pre-release gate checklist
|
||||
|
||||
1. Run the full suite: `uv run pytest`
|
||||
2. Confirm contract guardrails are green:
|
||||
- `uv run pytest tests/test_meta_contract_guards.py`
|
||||
3. Confirm health endpoint includes worker liveness payload (`/healthz` returns `worker.state`).
|
||||
4. Confirm required runtime settings are present in deployment environment:
|
||||
- `OPENROUTER_API_KEY`
|
||||
- `DATABASE__*`
|
||||
- filesystem paths for data/logs/backups.
|
||||
5. Confirm schema contract alignment is current:
|
||||
- `src/transcription/db/models.py`
|
||||
- `docs/ver4/schema_v4.md`
|
||||
|
||||
## 2. Release execution steps
|
||||
|
||||
1. Deploy artifact/config to target environment.
|
||||
2. Validate service startup:
|
||||
- `/healthz` responds `200`
|
||||
- `worker.state` is `running`
|
||||
3. Execute one smoke workflow:
|
||||
- create a document/job with at least one source
|
||||
- verify terminal job outcome updates
|
||||
- verify execution evidence row appended
|
||||
4. Verify log flow:
|
||||
- stdout aggregation receives events
|
||||
- file logs are written under `./data/logs`
|
||||
|
||||
## 3. Rollback triggers and actions
|
||||
|
||||
### Trigger conditions
|
||||
|
||||
1. `/healthz` reports `worker.state=failed`
|
||||
2. Repeated provider timeout/error spikes beyond normal baseline
|
||||
3. Evidence write failures or DB persistence failures
|
||||
|
||||
### Actions
|
||||
|
||||
1. Roll back app artifact and config to previous release.
|
||||
2. Restart service and re-check `/healthz`.
|
||||
3. Re-run smoke workflow and confirm worker returns to `running`.
|
||||
4. Preserve incident evidence:
|
||||
- `./data/logs`
|
||||
- relevant DB rows (`job`, `job_source`, `execution_attempt`)
|
||||
|
||||
## 4. Post-release monitoring checklist
|
||||
|
||||
## First 24 hours
|
||||
|
||||
1. Monitor `/healthz` periodically for `worker.state`.
|
||||
2. Track job terminal distribution (`transcribed`, `partial_success`, `failed`).
|
||||
3. Sample timeout/error categories for abnormal increase.
|
||||
4. Spot-check new `execution_attempt` records for append-only growth and timing metadata.
|
||||
|
||||
## First 72 hours
|
||||
|
||||
1. Re-check error/timeout trend versus 24h baseline.
|
||||
2. Verify no recurring worker-failed states.
|
||||
3. Verify storage growth and rotation behavior under `./data/logs`.
|
||||
4. Confirm incident response notes are captured for any production anomalies.
|
||||
|
||||
## 5. Operator playbook for common incidents
|
||||
|
||||
### Worker failed
|
||||
|
||||
1. Check `/healthz` payload (`error_id`, `error_category`).
|
||||
2. Locate matching error in logs.
|
||||
3. If non-transient defect persists, roll back.
|
||||
|
||||
### Provider timeout spike
|
||||
|
||||
1. Confirm provider reachability and rate limits.
|
||||
2. Review timeout frequency and impacted job volume.
|
||||
3. If sustained, execute rollback criteria and notify stakeholders.
|
||||
|
||||
### Partial-success increase
|
||||
|
||||
1. Inspect affected `job_source` and `execution_attempt` records.
|
||||
2. Confirm failures are category-aligned (`external`/`timeout`/`internal`).
|
||||
3. Triage whether issue is source quality, provider, or runtime regression.
|
||||
Reference in New Issue
Block a user