## Step 6 Goal (from `docs/ver1/ver1.md`) Implement **minimal observability & operability** so a single operator can quickly diagnose and recover from common failures. --- ## 1) Current-State Assessment (what already exists) ### Already in place - Central startup logging initialization via `setup_logging()` and `dictConfig` (`src/transcription/config.py`, `src/transcription/app.py`). - Error taxonomy and `error_id` envelope contract (`src/transcription/errors.py`) aligned with `docs/error_handling.md`. - Error handling for API and worker includes category + error IDs in some paths (`src/transcription/api/errors.py`, `src/transcription/worker.py`). - Basic health endpoint `/healthz` (`src/transcription/api/health.py`). - UI error display already shows actionable message + error reference (`src/transcription/ui/error_presenter.py`). ### Gaps to close for Step 6 1. **Structured logging is inconsistent** (many logs are free-form text with embedded key/value; no enforced schema). 2. **Boundary coverage is incomplete** (UI/service/API/worker don’t all emit consistent operation logs). 3. `/healthz` is very basic; no lightweight readiness/startup diagnostics endpoint/reporting. 4. No concise **operator runbook** yet (start/stop, log interpretation, recovery playbooks). 5. Minimal counters/timings are not yet standardized. --- ## 2) MCP Guidance Incorporated (relevant items) From `john-stream-mcp`, these are directly applied: - **`python-logging-dictconfig`**: keep one centralized `dictConfig`, configure once at startup, named loggers in modules. - **`fastapi-async-sqlalchemy-modernization`**: include observability + health/readiness checks; explicit lifecycle and deterministic startup/shutdown checks. - **`fastapi-uv-docker`**: keep `/healthz`; add practical readiness/ops checks for deployment clarity. - **`pytesting`**: deterministic tests, concise structure, validation lanes (`collect-only`, `unit`, `not external`, full). - **`pydantic-settings`**: keep typed settings as single source for logging/health behavior flags. - **`nicegui` + `nicegui-ui-customization`**: preserve clear, actionable user-facing error feedback and non-blocking UI flows. - **`zensical-docs`**: produce focused, navigable operator docs. (Other MCP resources were reviewed but are not core to Step 6 implementation scope.) --- ## 3) Detailed Implementation Plan for Step 6 ## Workstream A — Structured Logging Contract ### A1. Define a canonical log event schema Create a project log schema (doc + code-level constants) with required keys: - `timestamp` (UTC) - `level` - `logger` - `operation` - `event` - `error_id` (when error) - `category` (when error) - `exception_type` (when error) - `job_id`, `document_id` (when relevant) - optional: `duration_ms`, `retry_count`, `status` ### A2. Standardize log emission helpers Add small logging helpers (or adapter utilities) to reduce drift: - `log_operation_start(...)` - `log_operation_success(...)` - `log_operation_error(...)` Keep this minimal and avoid heavy observability frameworks. ### A3. Update formatter to structured output Use `dictConfig` to emit either: - JSON lines (preferred for structure), or - strict key-value line format with fixed fields. **Recommendation:** JSON lines to satisfy “structured logging” unambiguously while still simple. --- ## Workstream B — Boundary-by-Boundary Instrumentation ### B1. API boundary (`src/transcription/api/*`) - Add request-level operation logs for key routes (`upload.submit`, `jobs.list`, `jobs.get`, etc.). - Ensure API exception handler logs always include `error_id`, `category`, `operation`, `exception_type`. ### B2. Service boundary (`src/transcription/services/*`) - Add operation logs around: - upload validation/persist, - transcription orchestration, - revision add/accept, - search/export. - Add timing (`duration_ms`) for high-value operations only. ### B3. Worker boundary (`src/transcription/worker.py`) - Standardize all worker log events to schema. - Ensure retry logs include: `retriable`, `retry_count`, `max_retries`, `backoff_seconds`. - Ensure terminal failure logs include error contract fields. ### B4. UI boundary (`src/transcription/ui/*`) - Keep user-safe UI messages as-is. - Add backend/UI logger events for user-triggered failures (operation + error_id + category) so UI-visible errors correlate to server logs. --- ## Workstream C — Health, Readiness, Startup Operability ### C1. Keep `/healthz` lightweight - Return “process is running” status quickly. ### C2. Add lightweight `/readyz` Include small checks: - DB connectivity ping. - Worker thread alive check. - Optional prompt directory existence check. Return structured status payload with per-check pass/fail. ### C3. Startup self-check summary log At startup, emit one concise ops summary event: - environment - schema validation result - worker started - directories checked - bootstrap/migration mode flags --- ## Workstream D — Minimal Counters & Timings Add only high-value diagnostics: 1. `worker_jobs_processed_total` 2. `worker_jobs_failed_total` 3. `worker_retries_total` 4. `transcription_duration_ms` (per job) 5. `upload_persist_duration_ms` (per upload path) Implementation can be log-derived counters (no external metrics backend required). --- ## Workstream E — Operator Runbook Create concise runbook doc (recommended: `docs/ver1/ver1-step6-operator-runbook.md`) with: 1. **Start/Stop** - local `uv` run mode - docker compose mode (if applicable) 2. **Where logs are** - stdout, docker logs commands, filtering by `error_id` / `operation`. 3. **Common failure patterns → recovery** - provider timeout - auth denied - missing prompt dir - DB unavailable - job stuck/failed with retry exhausted 4. **Recovery procedures** - restart sequence - verify health/readiness - when to requeue/re-upload 5. **Escalation artifacts** - capture timestamp + error_id + operation + job_id/document_id Also update `README.md` with short links to the runbook. --- ## Workstream F — Verification & Quality Gates ### Tests to add/update - `tests/api/test_health.py` - `/healthz` baseline - `/readyz` pass/fail behavior - `tests/api/test_error_responses.py` / `tests/api/test_routes.py` - logs include `error_id/category/operation` on failures - `tests/services/test_worker.py` - retry/failure log fields + timing presence - `tests/ui/*` - ensure UI error correlation path includes operation/ref id behavior ### Validation commands (per MCP pytest guidance) - `uv run pytest --collect-only -q` - `uv run pytest -m unit -q` - `uv run pytest -m "not external" -q` - `uv run pytest -q` --- ## 4) Traceability to Governing Docs - **`docs/ver1/ver1.md` Step 6:** all 5 implementation bullets covered. - **`docs/error_handling.md`:** logging contract fields and error taxonomy continuity enforced. - **`docs/architecture.md`:** respects modular boundaries, in-process worker model, low-complexity ops. - **`docs/requirements.md`:** - REQ-8 (startup logging/config centralization) strengthened, - REQ-5 (status visibility) improved operationally, - REQ-7 lifecycle ownership observability improved. - **`docs/intent.md`:** keeps operation simple for personal-scale archival workflow. --- ## 5) Suggested Execution Order (low risk) 1. Logging schema + formatter + helpers 2. Worker/API instrumentation (highest value) 3. Service/UI instrumentation 4. `/readyz` + startup summary check 5. Runbook + README links 6. Tests + Step 6 results artifact (`docs/ver1/ver1-step6-results.md`)