generated from john/python-template
This commit is contained in:
@@ -103,9 +103,8 @@ Responsibilities:
|
||||
|
||||
## Status Semantics
|
||||
|
||||
- **Job statuses:** `queued`, `processing`, `transcribed`, `completed`, `partial_success`, `failed`
|
||||
- Operational success path currently resolves to `transcribed`.
|
||||
- `completed` remains a recognized legacy-compatible status value.
|
||||
- **Job statuses:** `queued`, `processing`, `transcribed`, `partial_success`, `failed`
|
||||
- Operational success path resolves to `transcribed`.
|
||||
- **JobSource statuses:** `pending`, `transcribed`, `failed`, `cancelled`
|
||||
|
||||
## Security and Path Handling Boundaries
|
||||
@@ -123,6 +122,45 @@ Responsibilities:
|
||||
- Non-retriable worker-loop faults are surfaced and stop loop spin.
|
||||
- Per-page outcomes are durably persisted before processing next page.
|
||||
|
||||
## Design Decisions and Rationale
|
||||
|
||||
### Why `transcribed` is the success terminal state
|
||||
|
||||
- The worker and job orchestration resolve successful completion to `JobStatus.TRANSCRIBED`, with mixed and failure outcomes represented by `partial_success` and `failed`.
|
||||
- This keeps terminal status vocabulary aligned with what the pipeline actually produces: transcribed page content and evidence, not a generic completion marker.
|
||||
|
||||
### Why evidence history is append-only while page text is a projection
|
||||
|
||||
- `ExecutionAttempt` stores immutable per-call evidence and preserves full attempt history across retries.
|
||||
- `Source.raw_transcription` is intentionally a mutable projection so UI and exports can show a selected current machine text without mutating historical evidence.
|
||||
- This split keeps auditability and UX both first-class: history is durable, presentation is editable.
|
||||
|
||||
### Why orchestration modules own cross-service workflows
|
||||
|
||||
- Service modules do not import each other; aggregate ownership remains local to each service.
|
||||
- Multi-aggregate writes are coordinated in orchestration modules (`store.py`, `workflows.py`) so transaction boundaries are explicit and testable.
|
||||
- This avoids circular dependencies and keeps cross-cutting workflow logic centralized.
|
||||
|
||||
### Why explicit eager loading is required
|
||||
|
||||
- ORM relationships are configured with `lazy="raise"` in key paths, so code must request needed relationships up front.
|
||||
- This prevents hidden query behavior in UI/service code and makes read shape deterministic and reviewable.
|
||||
|
||||
### Why canonical source bytes may be ingest-normalized
|
||||
|
||||
- Ingest normalization can correct orientation before persistence so provider calls, evidence hashes, and rendered processing source are consistent.
|
||||
- The canonical stored bytes, digest, and size become the durable processing identity for that source.
|
||||
|
||||
### Why media access uses controlled routes/helpers
|
||||
|
||||
- Print/export media uses record-validated API endpoints to avoid direct filesystem path exposure.
|
||||
- General UI media URLs are generated through shared resolver helpers to keep path handling consistent and centralized.
|
||||
|
||||
## Historical Context Boundary
|
||||
|
||||
Superseded V4.x scope/plan/review documents were intentionally removed from the active tree and archived at git tag `docs-v4x-archive`.
|
||||
Current architecture rules live only in `docs/ver4/*`; historical files are reference material only.
|
||||
|
||||
## Related References
|
||||
|
||||
- [System Requirements](requirements_v4.md)
|
||||
|
||||
@@ -19,6 +19,25 @@ This policy defines active V4 error taxonomy, translation boundaries, and retry
|
||||
- **Service layer:** map raw exceptions into domain-aware categories and preserve causal chain.
|
||||
- **UI/API layer:** convert category to user-safe message with contextual action guidance.
|
||||
|
||||
## Decision Context
|
||||
|
||||
### Why taxonomy is category-based (not exception-class-based)
|
||||
|
||||
- Categories encode operator-facing recovery semantics (fix input, retry later, investigate internal failure) independent of low-level exception type.
|
||||
- This keeps retry and messaging behavior consistent even when provider/client libraries change.
|
||||
|
||||
### Why page-level failure is isolated
|
||||
|
||||
- Multi-page archival documents often contain a mix of readable and degraded pages.
|
||||
- Isolating failures to page scope preserves successful results and avoids all-or-nothing loss when one page fails.
|
||||
- Aggregate job status then communicates overall outcome (`transcribed`, `partial_success`, `failed`) without hiding page detail.
|
||||
|
||||
### Why retries append evidence instead of mutating rows
|
||||
|
||||
- Retry operations are new observations, not corrections of history.
|
||||
- Appending attempts preserves forensic traceability, timing history, and provider variability analysis.
|
||||
- Projection updates remain explicit user/workflow decisions, separate from immutable evidence.
|
||||
|
||||
## Job and Page Failure Semantics
|
||||
|
||||
### Page-Level (`JobSource`)
|
||||
@@ -45,6 +64,12 @@ This policy defines active V4 error taxonomy, translation boundaries, and retry
|
||||
2. Avoid leaking stack traces or local paths into user-facing message envelopes.
|
||||
3. Preserve causal exception chains for internal diagnostics.
|
||||
|
||||
## Operator Recovery Guidance
|
||||
|
||||
- **validation/conflict:** correct input or state and retry manually.
|
||||
- **external/timeout:** allow bounded retries and keep prior attempt evidence visible.
|
||||
- **internal:** stop automatic retries, surface a safe message, and inspect diagnostics with correlation context.
|
||||
|
||||
## UI Messaging Contract
|
||||
|
||||
- User-visible errors must be actionable, bounded, and category-consistent.
|
||||
|
||||
@@ -16,7 +16,7 @@ These requirements define the active V4 contract and align to current implementa
|
||||
- **REQ-4-010 Job Creation:** The system must create `Job` records from uploaded sources and from retranscription of existing sources.
|
||||
- **REQ-4-011 Prompt Snapshotting:** Job creation must persist effective prompt and runtime settings as immutable per-job snapshots.
|
||||
- **REQ-4-012 Queue Membership:** Each `(job, source)` pair must be represented by one `JobSource` row.
|
||||
- **REQ-4-013 Job Status Lifecycle:** `Job.status` must use one of `queued`, `processing`, `transcribed`, `completed`, `partial_success`, `failed`.
|
||||
- **REQ-4-013 Job Status Lifecycle:** `Job.status` must use one of `queued`, `processing`, `transcribed`, `partial_success`, `failed`.
|
||||
- **REQ-4-014 JobSource Status Lifecycle:** `JobSource.status` must use one of `pending`, `transcribed`, `failed`, `cancelled`.
|
||||
- **REQ-4-015 Terminal Job Resolution:** Job terminal status must derive from page outcomes as `transcribed`, `partial_success`, or `failed`.
|
||||
- **REQ-4-016 Cancellation Semantics:** Job cancellation must set remaining `pending` page entries to `cancelled`.
|
||||
@@ -50,6 +50,30 @@ These requirements define the active V4 contract and align to current implementa
|
||||
- **REQ-4-104 Evidence Durability:** Attempt evidence must survive process restart once the transaction commits.
|
||||
- **REQ-4-105 Test Guardrails:** Architecture boundary tests must remain in place for services and UI boundaries.
|
||||
|
||||
## Requirement Interpretation Notes
|
||||
|
||||
### Status and lifecycle semantics
|
||||
|
||||
- `REQ-4-013` and `REQ-4-015` intentionally bind success to `transcribed`, not a generic `completed`, so docs, tests, and runtime transitions stay consistent.
|
||||
- `REQ-4-016` and `REQ-4-042` distinguish cancellation from failure at page level (`cancelled` vs `failed`) while still allowing targeted retranscription.
|
||||
|
||||
### Evidence semantics
|
||||
|
||||
- `REQ-4-020` through `REQ-4-024` separate authoritative history (`ExecutionAttempt`) from operational projection (`Source.raw_transcription`).
|
||||
- This supports immutable provenance while allowing explicit candidate promotion for operator workflows.
|
||||
|
||||
### Boundary and loading semantics
|
||||
|
||||
- `REQ-4-100` and `REQ-4-101` codify aggregate/service ownership and keep UI out of persistence concerns.
|
||||
- `REQ-4-102` exists to enforce deterministic query shape under `lazy="raise"` and avoid hidden data access in rendering callbacks.
|
||||
|
||||
## Verification Anchors
|
||||
|
||||
- Service boundary enforcement: `tests/test_service_boundaries.py`
|
||||
- UI boundary enforcement: `tests/test_ui_boundaries.py`
|
||||
- Job lifecycle reliability and terminal status behavior: `tests/services/test_workflows_reliability.py`
|
||||
- Evidence append-only and projection behavior: `tests/services/test_store.py`, `tests/services/test_transcription_service.py`
|
||||
|
||||
## Traceability Notes
|
||||
|
||||
- Source of truth for status enums:
|
||||
|
||||
+31
-1
@@ -87,7 +87,6 @@ erDiagram
|
||||
- `queued`
|
||||
- `processing`
|
||||
- `transcribed`
|
||||
- `completed` (legacy-compatible)
|
||||
- `partial_success`
|
||||
- `failed`
|
||||
|
||||
@@ -112,6 +111,31 @@ erDiagram
|
||||
4. `Job` terminal status is derived from `JobSource` outcomes.
|
||||
5. Registry semantic keys, when present, are immutable once created.
|
||||
|
||||
## Schema Design Rationale
|
||||
|
||||
### Why `ExecutionAttempt` exists alongside `JobSource`
|
||||
|
||||
- `JobSource` is the mutable queue/projection row for workflow control and current-facing page outcome state.
|
||||
- `ExecutionAttempt` is the durable, append-only evidence timeline for each provider call.
|
||||
- Keeping both avoids overloading one table with competing concerns (queue state vs immutable audit history).
|
||||
|
||||
### Why `Source.raw_transcription` remains on `Source`
|
||||
|
||||
- `Source` needs a stable, current projection for UI and print behavior.
|
||||
- Projection reads are fast and direct, while deep historical inspection remains available through attempts.
|
||||
- Explicit candidate promotion updates the projection pointer without rewriting history.
|
||||
|
||||
### Why semantic registries use UUID identity plus optional semantic keys
|
||||
|
||||
- UUIDs are durable relationship identifiers for public/domain links.
|
||||
- Optional immutable semantic keys support protected built-ins without exposing internal meaning as external API identity.
|
||||
- Labels can evolve without breaking relationships.
|
||||
|
||||
### Why status enums are narrow
|
||||
|
||||
- Restricting `JobStatus` and `JobSourceStatus` keeps lifecycle transitions explicit and testable.
|
||||
- Queue/state transitions and terminal derivation logic remain deterministic across worker and UI flows.
|
||||
|
||||
## Media Storage Semantics
|
||||
|
||||
1. `Source.storage_path` references canonical stored bytes used by processing.
|
||||
@@ -123,6 +147,12 @@ erDiagram
|
||||
- Relationship access from service/UI layers must use explicit eager loading patterns compatible with `lazy="raise"`.
|
||||
- Candidate-attempt views should select latest/selected attempts explicitly; do not rely on implicit lazy traversal.
|
||||
|
||||
## Evolution and Migration Policy
|
||||
|
||||
- Additive schema evolution is preferred for evidence-bearing records.
|
||||
- Deprecated semantics should be removed only when model enums, service logic, tests, and docs are updated together.
|
||||
- Historical V4.x schema discussions are archived under tag `docs-v4x-archive`; this file is the active contract.
|
||||
|
||||
## Cross-Reference
|
||||
|
||||
- [System Architecture](architecture_v4.md)
|
||||
|
||||
Reference in New Issue
Block a user