generated from john/python-template
189 lines
8.0 KiB
Markdown
189 lines
8.0 KiB
Markdown
# System Architecture (Version 4)
|
|
|
|
This document describes the production architecture of the document transcription system.
|
|
|
|
## Architecture Objectives
|
|
|
|
- Preserve original source material, per-execution machine output, and separate human revision.
|
|
- Support batching one or more images into ordered multi-page documents.
|
|
- Capture submission-time prompt provenance and a per-page OpenRouter SDK response snapshot.
|
|
- Execute page transcription concurrently with bounded `asyncio` workers.
|
|
- Maintain relational portability across SQLite and PostgreSQL.
|
|
- Keep operator workflows cross-platform and Python-driven.
|
|
- Support many-to-many document-person relationships with extensible roles.
|
|
- Support registry-driven document type classification.
|
|
|
|
## Core Capabilities
|
|
|
|
- Ingest one or more images into sequential `Source` pages under a `Document`.
|
|
- Execute asynchronous vision transcription with bounded worker concurrency.
|
|
- Preserve original source files with SHA-256 digests and byte sizes.
|
|
- Freeze prompt text, prompt hash, model, and explicitly configured sampling parameters on each `Job`.
|
|
- Preserve page-level machine output, normalized metadata, and an SDK-serialized OpenRouter response snapshot on `JobSource`.
|
|
- Organize historical `Person` records through many-to-many Document relationships and extensible roles.
|
|
- Classify Documents through a registry with stable type codes.
|
|
- Maintain human revision separately from machine-generated text.
|
|
- Isolate page failures so multi-page jobs can complete with partial success.
|
|
- Operate across supported platforms through Python-based application and maintenance tooling.
|
|
|
|
V4.2 extends this baseline with exact OpenRouter transport evidence and provider-neutral derived-artifact provenance. See the [V4.2 Scope Boundary](../ver4.2/scope_boundary_v4_2.md).
|
|
|
|
## Technical Stack
|
|
|
|
- **Runtime:** Python 3.12 or later.
|
|
- **Web application:** FastAPI and NiceGUI.
|
|
- **Persistence:** SQLModel and SQLAlchemy, with SQLite and PostgreSQL support.
|
|
- **Validation and settings:** Pydantic V2 and pydantic-settings.
|
|
- **Concurrency:** Python `asyncio` workers.
|
|
- **Vision integration:** OpenRouter through the application's provider adapter.
|
|
- **Testing and quality:** pytest, pytest-asyncio, Ruff, and ty.
|
|
|
|
## Runtime Topology
|
|
|
|
The runtime operates as an asynchronous Python application:
|
|
|
|
- FastAPI + NiceGUI web application process.
|
|
- In-process `asyncio` worker engine for transcription execution.
|
|
- Relational persistence via SQLModel / SQLAlchemy.
|
|
- Pydantic V2 validation across API payloads, prompt configuration, and structured metadata.
|
|
|
|
^^^mermaid
|
|
flowchart LR
|
|
U[Browser User] --> A[FastAPI + NiceGUI App]
|
|
A --> W[Asyncio Worker Engine]
|
|
A --> DB[(Relational DB)]
|
|
W --> P[Vision Provider APIs]
|
|
W --> DB
|
|
^^^
|
|
|
|
## Lifecycle Ownership
|
|
|
|
Application lifespan owns runtime setup and teardown:
|
|
|
|
- Initialize logging, settings, directories, and prompt configuration.
|
|
- Manage asynchronous database engine connection pools.
|
|
- Execute database bootstrap or migrations.
|
|
- Recover stale or interrupted jobs on startup.
|
|
- Manage graceful shutdown of active background tasks.
|
|
|
|
## Layered Module Structure
|
|
|
|
### Interface Layer
|
|
|
|
- `src/transcription/ui/**`
|
|
- `src/transcription/api/**`
|
|
|
|
Responsibilities:
|
|
|
|
- Render document, source, person, job, and classification views.
|
|
- Accept user input for uploads, editing, linking, and revisions.
|
|
- Present structured validation and conflict feedback.
|
|
|
|
### Application and Async Worker Layer
|
|
|
|
- `src/transcription/services/workflows.py`
|
|
- `src/transcription/worker.py`
|
|
|
|
Responsibilities:
|
|
|
|
- Orchestrate uploads, job creation, and status transitions.
|
|
- Execute per-page provider calls through bounded concurrency.
|
|
- Persist page-level outcomes and update aggregate job state.
|
|
|
|
### Domain and Service Layer
|
|
|
|
- `src/transcription/db/models.py`
|
|
- `src/transcription/services/documents.py`
|
|
- `src/transcription/services/sources.py`
|
|
- `src/transcription/services/jobs.py`
|
|
- `src/transcription/services/people.py`
|
|
- `src/transcription/services/workflows.py`
|
|
|
|
Responsibilities:
|
|
|
|
- Keep one primary service boundary per aggregate: Documents, Sources, Jobs, and People.
|
|
- Documents own document records and the document-type registry.
|
|
- Sources own source records, revisions, source media formats, MIME resolution, and page execution evidence.
|
|
- Jobs own job lifecycle state and transitions.
|
|
- People own person records, relationship roles, document-person links, and portrait media.
|
|
- Apply deterministic conflict handling for relationship-role writes.
|
|
- Use set-based synchronization for many-to-many relationship updates.
|
|
- Resolve and validate registry-backed document types.
|
|
|
|
### Source Media Policy
|
|
|
|
- `services/sources.py` is the single authority for accepted Source extensions and canonical MIME types.
|
|
- Storage and provider payload loading must call the same Source validation functions.
|
|
- Supported Source formats are JPEG, PNG, TIFF, and PDF.
|
|
- Upload is an interface action, not a domain aggregate. Service names, errors, and workflow variables use
|
|
`Source` terminology; compatibility aliases may remain temporarily at old import boundaries.
|
|
|
|
### Infrastructure Layer
|
|
|
|
- `src/transcription/db/**`
|
|
- `src/transcription/providers/**`
|
|
|
|
Responsibilities:
|
|
|
|
- Provide async database sessions and engine configuration.
|
|
- Provide provider adapters for vision model execution.
|
|
|
|
## Core Workflows
|
|
|
|
### 1. Multi-Page Transcription
|
|
|
|
1. User uploads one or more images for a `Document`.
|
|
2. System stores files, hashes them, creates ordered `Source` rows, and creates a `Job`.
|
|
3. Worker claims the job, marks it `processing`, and executes page calls concurrently.
|
|
4. Each page writes a `JobSource` result with machine output, normalized metadata, and an SDK-serialized OpenRouter response snapshot.
|
|
5. Aggregate status becomes `completed`, `partial_success`, or `failed`.
|
|
|
|
### 2. Document-Person Relationship Management
|
|
|
|
1. User opens a document or person edit flow.
|
|
2. UI loads existing links grouped by role.
|
|
3. User adds or removes people within one or more roles.
|
|
4. Service computes add/remove deltas rather than replacing all links blindly.
|
|
5. Conflict checks enforce uniqueness and deterministic write semantics before persistence commits.
|
|
|
|
### 3. Document Type Management
|
|
|
|
1. User selects a registry-backed document type for a document.
|
|
2. Service resolves the stable type code or id.
|
|
3. Persistence stores the `document_type_id` reference.
|
|
4. Inactive types remain valid for historical rows but are excluded from default selectors.
|
|
|
|
## V4 Domain Rules
|
|
|
|
- `JobSource.raw_transcription` preserves page output for its Job execution.
|
|
- `Source.raw_transcription` is the latest-success machine-output projection for a page.
|
|
- Human corrections occur only in `Source.revised_text`.
|
|
- Prompt and parameter provenance is frozen on `Job` at submission time.
|
|
- The SDK-serialized OpenRouter response snapshot is stored on `JobSource` for each successful page execution.
|
|
- `DocumentPerson` links are unique for `(document_id, person_id, role_id)`.
|
|
- Relationship mutations are deterministic and set-based.
|
|
- `DocumentType.code` is stable; `DocumentType.label` may evolve.
|
|
|
|
## Data Model Summary
|
|
|
|
- `Document` has one `DocumentType`, many `Source` pages, many `Job` runs, and many `Person` records through `DocumentPerson`.
|
|
- `Source` belongs to one `Document` and may participate in many `JobSource` executions.
|
|
- `Job` has many `JobSource` rows.
|
|
- `PersonRole` defines available relationship roles.
|
|
|
|
## Test Strategy
|
|
|
|
- Unit tests for models, validation, hashing, and registry resolution.
|
|
- Service tests for CRUD, set-based sync, uniqueness conflicts, and deterministic relationship writes.
|
|
- Async workflow tests for page isolation, partial failure handling, and stored evidence.
|
|
- UI integration tests for multi-page rendering, role grouping, and document type selection.
|
|
|
|
## Related Local References
|
|
|
|
- [System Overview](index_v4.md)
|
|
- [System Requirements](requirements_v4.md)
|
|
- [Data Model](schema_v4.md)
|
|
- [Error Handling Policy](error_handling_v4.md)
|
|
- [Error Handling Invariant](../invariant/error_handling.md)
|
|
- [Digital Evidence and AI Processing Provenance](../invariant/ai_evidence_and_provenance.md)
|