# Step 2 Implementation Plan: Error Handling & Reliability Hardening ## Purpose Implement **Ver1 Step 2** from `docs/ver1/ver1.md` by standardizing failure behavior and reliability controls so the system fails safely, predictably, and transparently across UI, API, services, worker, and provider boundaries. Primary governing docs: - `docs/error_handling.md` (authoritative contract) - `docs/requirements.md` (REQ-2, REQ-3, REQ-4, REQ-5, REQ-6) - `docs/architecture.md` (boundary ownership and worker lifecycle) - `docs/ver1/ver1.md` (Step 2 objective) --- ## MCP Skill and Guide Inputs Incorporated This plan integrates guidance from john-stream-mcp resources: 1. `resource://skills/python-logging-dictconfig/document` - centralized `dictConfig` logging - startup-only configuration - stable named loggers and boundary-level logging discipline 2. `resource://skills/pytesting/document` - deterministic, behavior-first tests - explicit marker usage and fast/slow lane discipline - integration checks for boundary behavior and error contracts 3. `resource://skills/fastapi-async-sqlalchemy-modernization/document` - classify at source boundary - explicit transaction/session behavior under failure - phased rollout with quality gates and rollback awareness 4. `resource://skills/nicegui-ui-customization/document` - explicit user-facing error feedback for each interaction - prevent duplicate actions during in-flight operations - preserve one-way dependency boundaries from UI -> services 5. `resource://skills/fastapi-uv-docker/document` (applied selectively) - lifespan-safe startup/shutdown behavior - health/readiness posture and cloud-native operational checks --- ## Current-State Gap Summary The project already has a strong baseline (`AppError`, taxonomy enum, API envelope, worker persistence), but Step 2 needs completion-level hardening: 1. **Error contract consistency** - API envelope exists, but consistency must be verified for all error pathways. 2. **Cross-boundary category normalization** - Provider/service/worker mappings exist, but require stricter policy checks and tests. 3. **Retry policy implementation depth** - Step 2 requires bounded retry policy and clear terminal behavior for retriable failures. 4. **Operational traceability** - Logging exists; Step 2 requires consistent structured fields at critical boundaries. 5. **UI failure UX consistency** - UI error handling exists; Step 2 requires explicit contract coverage and anti-duplication safeguards. --- ## Scope for Step 2 ### In scope 1. Enforce canonical error taxonomy and envelope across all boundaries. 2. Standardize logging fields and boundary-level error traceability. 3. Implement/complete bounded retry and terminal failure behavior in worker paths. 4. Improve UI/API error presentation consistency and actionable guidance. 5. Add comprehensive Step 2 test coverage and verification matrix. 6. Update documentation to reflect final Step 2 policies and behavior. ### Out of scope - Major architecture/topology changes (external queue, distributed worker) - New end-user feature expansion outside reliability/error handling - Full async ORM migration (unless required by bug fix) --- ## Target Decisions for Step 2 1. **Taxonomy stability is mandatory** - `ErrorCategory` values remain stable contract identifiers. 2. **Classification occurs at source boundary** - adapters/services normalize early; UI/API only present safely. 3. **User safety over internal detail leakage** - expose safe message + suggestion + error_id; keep sensitive detail in logs. 4. **Retry is explicit and bounded** - only retriable categories may retry; retries are capped; terminal failures persist reason. 5. **Boundary logs carry correlation fields** - include `error_id`, `category`, `operation`, and domain identifiers where available. --- ## Detailed Work Breakdown ## Phase A — Error Contract Audit and Policy Lock - [ ] **A1. Build error-path inventory** - Enumerate all failure entry points across: - `api/` - `ui/` - `services/` - `worker.py` - `providers/` - [ ] **A2. Produce taxonomy mapping table** - For each known exception path, map: - source exception type - target `ErrorCategory` - retriable flag - API status (if exposed) - [ ] **A3. Reconcile with `docs/error_handling.md`** - Resolve any mismatch in category semantics, status codes, or suggested actions. ### Deliverables - `docs/ver1/ver1-step2-audit.md` (recommended) - taxonomy mapping table ### Exit Criteria - Every known failure path has explicit category + retriable policy. --- ## Phase B — API and Service Contract Hardening - [ ] **B1. Enforce API envelope completeness** - Ensure all API errors return: - `error_id`, `category`, `message`, `suggestion`, `timestamp` - [ ] **B2. Verify category-to-status mapping consistency** - Confirm `api/errors.py` matches `docs/error_handling.md` mapping guidance. - [ ] **B3. Normalize service exceptions at boundary** - Services should raise `AppError` subclasses for known failures. - Unknown exceptions must become `internal_unexpected_error` with traceable `error_id`. - [ ] **B4. Ensure safe detail handling** - API/UI messages remain safe. - Diagnostic context remains in logs/persisted failure detail where appropriate. ### Exit Criteria - No unstructured/unclassified exception escapes core boundaries. - API responses are contract-stable for all tested failure modes. --- ## Phase C — Worker Retry and Terminal Failure Policy - [ ] **C1. Define bounded retry policy** - Add configurable retry settings (attempt limit/backoff policy). - Limit retries to retriable categories. - [ ] **C2. Implement terminal failure persistence** - On retry exhaustion, persist clear terminal reason and `error_id`. - Ensure job status transitions end deterministically at `failed`. - [ ] **C3. Add duplicate-processing safety checks** - Prevent duplicate terminal updates when job already resolved. - [ ] **C4. Validate worker lifecycle under repeated transient failures** - Ensure loop remains stable and responsive. ### Exit Criteria - Retries are bounded and policy-driven. - Exhausted retries produce deterministic failed state with evidence. --- ## Phase D — Logging and Observability Contract Enforcement - [ ] **D1. Central logging conformance check** - Confirm startup-only `dictConfig` use remains canonical. - No module-level `basicConfig` use. - [ ] **D2. Standardize error log fields** - Require at minimum when available: - `error_id`, `category`, `operation`, `exception_type`, `job_id`, `document_id` - [ ] **D3. Boundary handoff logging** - Add/normalize logs at transitions: - UI action -> service - service -> provider/db - worker pickup -> terminal state - [ ] **D4. Log noise control** - Avoid duplicate stack-trace logging across layers for same exception. ### Exit Criteria - Critical failure events are traceable end-to-end via logs and `error_id`. --- ## Phase E — UI Error UX Consistency and Interaction Hardening - [ ] **E1. Standardize user error presentation** - For upload/jobs interactions, ensure: - clear title - plain-language message - suggested action - visible error reference id - [ ] **E2. Add in-flight interaction guards** - Prevent duplicate submits/click storms during pending operations. - [ ] **E3. Ensure deterministic UI state recovery** - controls re-enable after failure - status text remains actionable - [ ] **E4. Keep UI boundary clean** - no provider/protocol details leaked into page modules ### Exit Criteria - All primary UI actions have consistent success/failure interaction behavior. --- ## Phase F — Test Expansion and Verification Apply pytesting guidance: behavior-first assertions, deterministic fixtures, strict markers. - [ ] **F1. API error contract tests** - verify envelope fields and status mapping for each category class. - [ ] **F2. Service classification tests** - verify known failures map to expected `AppError` subclasses/categories. - [ ] **F3. Worker retry policy tests** - retriable failure retries and eventual success - retriable failure exhaustion -> terminal failed - non-retriable failure -> immediate failed - [ ] **F4. UI error behavior tests** - upload/jobs actions show actionable feedback on failures - duplicate action guard behavior - [ ] **F5. Regression guard tests** - at least one test per previously observed production/real-world failure mode ### Validation Commands - `uv run pytest --collect-only -q` - `uv run pytest -m unit -q` - `uv run pytest -m "not external" -q` - `uv run pytest -q` ### Exit Criteria - All Step 2 reliability/error contract tests pass. - Existing suite remains green. --- ## Recommended Implementation Order 1. Phase A — audit and policy lock 2. Phase B — API/service contract hardening 3. Phase C — worker retry and terminal policy 4. Phase D — logging/traceability normalization 5. Phase E — UI consistency hardening 6. Phase F — test expansion and full verification This order reduces risk by locking policy first, then applying behavior changes at core boundaries before UI polish. --- ## Risks and Mitigations 1. **Risk:** Overly broad retry policy causes hidden failure loops **Mitigation:** strict category-based retry eligibility + hard cap + terminal persistence. 2. **Risk:** User-facing messages become too technical **Mitigation:** enforce safe message + suggestion contract in tests. 3. **Risk:** Logging becomes noisy/redundant **Mitigation:** boundary logging rules and single-trace ownership. 4. **Risk:** Reliability work introduces regressions in happy path **Mitigation:** run full suite continuously; preserve integration pipeline tests. --- ## Step 2 Completion Checklist - [ ] Error taxonomy mapping table completed and approved. - [ ] API envelope and HTTP status behavior verified for all relevant failure categories. - [ ] Service/provider exception normalization is consistent and tested. - [ ] Worker retry behavior is bounded, explicit, and terminal-state safe. - [ ] Structured error logging fields are present at boundary handoffs. - [ ] UI failure flows provide clear, actionable, and traceable feedback. - [ ] Full test suite passes with new Step 2 coverage included. - [ ] `docs/ver1/ver1-step2-results.md` created with evidence and residual risks. --- ## Handoff to Step 3 After Step 2 completion, Step 3 (Functional Completion by Requirement Domain) proceeds on a hardened foundation: - stable failure contracts, - predictable retries and terminal behavior, - actionable user/API error semantics, - improved diagnostic traceability.