11 KiB
Step 2 Implementation Plan: Error Handling & Reliability Hardening
Purpose
Implement Ver1 Step 2 from docs/ver1/ver1.md by standardizing failure behavior and reliability controls so the system fails safely, predictably, and transparently across UI, API, services, worker, and provider boundaries.
Primary governing docs:
docs/error_handling.md(authoritative contract)docs/requirements.md(REQ-2, REQ-3, REQ-4, REQ-5, REQ-6)docs/architecture.md(boundary ownership and worker lifecycle)docs/ver1/ver1.md(Step 2 objective)
MCP Skill and Guide Inputs Incorporated
This plan integrates guidance from john-stream-mcp resources:
-
resource://skills/python-logging-dictconfig/document- centralized
dictConfiglogging - startup-only configuration
- stable named loggers and boundary-level logging discipline
- centralized
-
resource://skills/pytesting/document- deterministic, behavior-first tests
- explicit marker usage and fast/slow lane discipline
- integration checks for boundary behavior and error contracts
-
resource://skills/fastapi-async-sqlalchemy-modernization/document- classify at source boundary
- explicit transaction/session behavior under failure
- phased rollout with quality gates and rollback awareness
-
resource://skills/nicegui-ui-customization/document- explicit user-facing error feedback for each interaction
- prevent duplicate actions during in-flight operations
- preserve one-way dependency boundaries from UI -> services
-
resource://skills/fastapi-uv-docker/document(applied selectively)- lifespan-safe startup/shutdown behavior
- health/readiness posture and cloud-native operational checks
Current-State Gap Summary
The project already has a strong baseline (AppError, taxonomy enum, API envelope, worker persistence), but Step 2 needs completion-level hardening:
- Error contract consistency
- API envelope exists, but consistency must be verified for all error pathways.
- Cross-boundary category normalization
- Provider/service/worker mappings exist, but require stricter policy checks and tests.
- Retry policy implementation depth
- Step 2 requires bounded retry policy and clear terminal behavior for retriable failures.
- Operational traceability
- Logging exists; Step 2 requires consistent structured fields at critical boundaries.
- UI failure UX consistency
- UI error handling exists; Step 2 requires explicit contract coverage and anti-duplication safeguards.
Scope for Step 2
In scope
- Enforce canonical error taxonomy and envelope across all boundaries.
- Standardize logging fields and boundary-level error traceability.
- Implement/complete bounded retry and terminal failure behavior in worker paths.
- Improve UI/API error presentation consistency and actionable guidance.
- Add comprehensive Step 2 test coverage and verification matrix.
- Update documentation to reflect final Step 2 policies and behavior.
Out of scope
- Major architecture/topology changes (external queue, distributed worker)
- New end-user feature expansion outside reliability/error handling
- Full async ORM migration (unless required by bug fix)
Target Decisions for Step 2
- Taxonomy stability is mandatory
ErrorCategoryvalues remain stable contract identifiers.
- Classification occurs at source boundary
- adapters/services normalize early; UI/API only present safely.
- User safety over internal detail leakage
- expose safe message + suggestion + error_id; keep sensitive detail in logs.
- Retry is explicit and bounded
- only retriable categories may retry; retries are capped; terminal failures persist reason.
- Boundary logs carry correlation fields
- include
error_id,category,operation, and domain identifiers where available.
- include
Detailed Work Breakdown
Phase A — Error Contract Audit and Policy Lock
-
A1. Build error-path inventory
- Enumerate all failure entry points across:
api/ui/services/worker.pyproviders/
- Enumerate all failure entry points across:
-
A2. Produce taxonomy mapping table
- For each known exception path, map:
- source exception type
- target
ErrorCategory - retriable flag
- API status (if exposed)
- For each known exception path, map:
-
A3. Reconcile with
docs/error_handling.md- Resolve any mismatch in category semantics, status codes, or suggested actions.
Deliverables
docs/ver1/ver1-step2-audit.md(recommended)- taxonomy mapping table
Exit Criteria
- Every known failure path has explicit category + retriable policy.
Phase B — API and Service Contract Hardening
-
B1. Enforce API envelope completeness
- Ensure all API errors return:
error_id,category,message,suggestion,timestamp
- Ensure all API errors return:
-
B2. Verify category-to-status mapping consistency
- Confirm
api/errors.pymatchesdocs/error_handling.mdmapping guidance.
- Confirm
-
B3. Normalize service exceptions at boundary
- Services should raise
AppErrorsubclasses for known failures. - Unknown exceptions must become
internal_unexpected_errorwith traceableerror_id.
- Services should raise
-
B4. Ensure safe detail handling
- API/UI messages remain safe.
- Diagnostic context remains in logs/persisted failure detail where appropriate.
Exit Criteria
- No unstructured/unclassified exception escapes core boundaries.
- API responses are contract-stable for all tested failure modes.
Phase C — Worker Retry and Terminal Failure Policy
-
C1. Define bounded retry policy
- Add configurable retry settings (attempt limit/backoff policy).
- Limit retries to retriable categories.
-
C2. Implement terminal failure persistence
- On retry exhaustion, persist clear terminal reason and
error_id. - Ensure job status transitions end deterministically at
failed.
- On retry exhaustion, persist clear terminal reason and
-
C3. Add duplicate-processing safety checks
- Prevent duplicate terminal updates when job already resolved.
-
C4. Validate worker lifecycle under repeated transient failures
- Ensure loop remains stable and responsive.
Exit Criteria
- Retries are bounded and policy-driven.
- Exhausted retries produce deterministic failed state with evidence.
Phase D — Logging and Observability Contract Enforcement
-
D1. Central logging conformance check
- Confirm startup-only
dictConfiguse remains canonical. - No module-level
basicConfiguse.
- Confirm startup-only
-
D2. Standardize error log fields
- Require at minimum when available:
error_id,category,operation,exception_type,job_id,document_id
- Require at minimum when available:
-
D3. Boundary handoff logging
- Add/normalize logs at transitions:
- UI action -> service
- service -> provider/db
- worker pickup -> terminal state
- Add/normalize logs at transitions:
-
D4. Log noise control
- Avoid duplicate stack-trace logging across layers for same exception.
Exit Criteria
- Critical failure events are traceable end-to-end via logs and
error_id.
Phase E — UI Error UX Consistency and Interaction Hardening
-
E1. Standardize user error presentation
- For upload/jobs interactions, ensure:
- clear title
- plain-language message
- suggested action
- visible error reference id
- For upload/jobs interactions, ensure:
-
E2. Add in-flight interaction guards
- Prevent duplicate submits/click storms during pending operations.
-
E3. Ensure deterministic UI state recovery
- controls re-enable after failure
- status text remains actionable
-
E4. Keep UI boundary clean
- no provider/protocol details leaked into page modules
Exit Criteria
- All primary UI actions have consistent success/failure interaction behavior.
Phase F — Test Expansion and Verification
Apply pytesting guidance: behavior-first assertions, deterministic fixtures, strict markers.
-
F1. API error contract tests
- verify envelope fields and status mapping for each category class.
-
F2. Service classification tests
- verify known failures map to expected
AppErrorsubclasses/categories.
- verify known failures map to expected
-
F3. Worker retry policy tests
- retriable failure retries and eventual success
- retriable failure exhaustion -> terminal failed
- non-retriable failure -> immediate failed
-
F4. UI error behavior tests
- upload/jobs actions show actionable feedback on failures
- duplicate action guard behavior
-
F5. Regression guard tests
- at least one test per previously observed production/real-world failure mode
Validation Commands
uv run pytest --collect-only -quv run pytest -m unit -quv run pytest -m "not external" -quv run pytest -q
Exit Criteria
- All Step 2 reliability/error contract tests pass.
- Existing suite remains green.
Recommended Implementation Order
- Phase A — audit and policy lock
- Phase B — API/service contract hardening
- Phase C — worker retry and terminal policy
- Phase D — logging/traceability normalization
- Phase E — UI consistency hardening
- Phase F — test expansion and full verification
This order reduces risk by locking policy first, then applying behavior changes at core boundaries before UI polish.
Risks and Mitigations
-
Risk: Overly broad retry policy causes hidden failure loops
Mitigation: strict category-based retry eligibility + hard cap + terminal persistence. -
Risk: User-facing messages become too technical
Mitigation: enforce safe message + suggestion contract in tests. -
Risk: Logging becomes noisy/redundant
Mitigation: boundary logging rules and single-trace ownership. -
Risk: Reliability work introduces regressions in happy path
Mitigation: run full suite continuously; preserve integration pipeline tests.
Step 2 Completion Checklist
- Error taxonomy mapping table completed and approved.
- API envelope and HTTP status behavior verified for all relevant failure categories.
- Service/provider exception normalization is consistent and tested.
- Worker retry behavior is bounded, explicit, and terminal-state safe.
- Structured error logging fields are present at boundary handoffs.
- UI failure flows provide clear, actionable, and traceable feedback.
- Full test suite passes with new Step 2 coverage included.
docs/ver1/ver1-step2-results.mdcreated with evidence and residual risks.
Handoff to Step 3
After Step 2 completion, Step 3 (Functional Completion by Requirement Domain) proceeds on a hardened foundation:
- stable failure contracts,
- predictable retries and terminal behavior,
- actionable user/API error semantics,
- improved diagnostic traceability.