generated from john/python-template
85 lines
2.8 KiB
Markdown
85 lines
2.8 KiB
Markdown
# Production Runbook
|
|
|
|
This runbook is the operational checklist for releasing and monitoring the transcription system.
|
|
|
|
## 1. Pre-release gate checklist
|
|
|
|
1. Run the full suite: `uv run pytest`
|
|
2. Confirm contract guardrails are green:
|
|
- `uv run pytest tests/test_meta_contract_guards.py`
|
|
3. Confirm health endpoint includes worker liveness payload (`/healthz` returns `worker.state`).
|
|
4. Confirm required runtime settings are present in deployment environment:
|
|
- `OPENROUTER_API_KEY`
|
|
- `DATABASE__*`
|
|
- filesystem paths for data/logs/backups.
|
|
5. Confirm schema contract alignment is current:
|
|
- `src/transcription/db/models.py`
|
|
- `docs/schema.md`
|
|
|
|
## 2. Release execution steps
|
|
|
|
1. Deploy artifact/config to target environment.
|
|
2. Validate service startup:
|
|
- `/healthz` responds `200`
|
|
- `worker.state` is `running`
|
|
3. Execute one smoke workflow:
|
|
- create a document/job with at least one source
|
|
- verify terminal job outcome updates
|
|
- verify execution evidence row appended
|
|
4. Verify log flow:
|
|
- stdout aggregation receives events
|
|
- file logs are written under `./data/logs`
|
|
|
|
## 3. Rollback triggers and actions
|
|
|
|
### Trigger conditions
|
|
|
|
1. `/healthz` reports `worker.state=failed`
|
|
2. Repeated provider timeout/error spikes beyond normal baseline
|
|
3. Evidence write failures or DB persistence failures
|
|
|
|
### Actions
|
|
|
|
1. Roll back app artifact and config to previous release.
|
|
2. Restart service and re-check `/healthz`.
|
|
3. Re-run smoke workflow and confirm worker returns to `running`.
|
|
4. Preserve incident evidence:
|
|
- `./data/logs`
|
|
- relevant DB rows (`job`, `job_source`, `execution_attempt`)
|
|
|
|
## 4. Post-release monitoring checklist
|
|
|
|
## First 24 hours
|
|
|
|
1. Monitor `/healthz` periodically for `worker.state`.
|
|
2. Track job terminal distribution (`transcribed`, `partial_success`, `failed`).
|
|
3. Sample timeout/error categories for abnormal increase.
|
|
4. Spot-check new `execution_attempt` records for append-only growth and timing metadata.
|
|
|
|
## First 72 hours
|
|
|
|
1. Re-check error/timeout trend versus 24h baseline.
|
|
2. Verify no recurring worker-failed states.
|
|
3. Verify storage growth and rotation behavior under `./data/logs`.
|
|
4. Confirm incident response notes are captured for any production anomalies.
|
|
|
|
## 5. Operator playbook for common incidents
|
|
|
|
### Worker failed
|
|
|
|
1. Check `/healthz` payload (`error_id`, `error_category`).
|
|
2. Locate matching error in logs.
|
|
3. If non-transient defect persists, roll back.
|
|
|
|
### Provider timeout spike
|
|
|
|
1. Confirm provider reachability and rate limits.
|
|
2. Review timeout frequency and impacted job volume.
|
|
3. If sustained, execute rollback criteria and notify stakeholders.
|
|
|
|
### Partial-success increase
|
|
|
|
1. Inspect affected `job_source` and `execution_attempt` records.
|
|
2. Confirm failures are category-aligned (`external`/`timeout`/`internal`).
|
|
3. Triage whether issue is source quality, provider, or runtime regression.
|