# System Architecture (Version 4) This document describes the production architecture of the document transcription system. ## Architecture Objectives - Preserve original source material, per-execution machine output, and separate human revision. - Support batching one or more images into ordered multi-page documents. - Capture submission-time prompt provenance and a per-page OpenRouter SDK response snapshot. - Execute page transcription concurrently with bounded `asyncio` workers. - Maintain relational portability across SQLite and PostgreSQL. - Keep operator workflows cross-platform and Python-driven. - Support many-to-many document-person relationships with extensible roles. - Support registry-driven document type classification. ## Core Capabilities - Ingest one or more images into sequential `Source` pages under a `Document`. - Execute asynchronous vision transcription with bounded worker concurrency. - Preserve original source files with SHA-256 digests and byte sizes. - Freeze prompt text, prompt hash, model, and explicitly configured sampling parameters on each `Job`. - Preserve page-level machine output, normalized metadata, and an SDK-serialized OpenRouter response snapshot on `JobSource`. - Organize historical `Person` records through many-to-many Document relationships and extensible roles. - Classify Documents through a registry with stable type codes. - Maintain human revision separately from machine-generated text. - Isolate page failures so multi-page jobs can complete with partial success. - Operate across supported platforms through Python-based application and maintenance tooling. V4.2 extends this baseline with exact OpenRouter transport evidence and provider-neutral derived-artifact provenance. See the [V4.2 Scope Boundary](../ver4.2/scope_boundary_v4_2.md). ## Technical Stack - **Runtime:** Python 3.12 or later. - **Web application:** FastAPI and NiceGUI. - **Persistence:** SQLModel and SQLAlchemy, with SQLite and PostgreSQL support. - **Validation and settings:** Pydantic V2 and pydantic-settings. - **Concurrency:** Python `asyncio` workers. - **Vision integration:** OpenRouter through the application's provider adapter. - **Testing and quality:** pytest, pytest-asyncio, Ruff, and ty. ## Runtime Topology The runtime operates as an asynchronous Python application: - FastAPI + NiceGUI web application process. - In-process `asyncio` worker engine for transcription execution. - Relational persistence via SQLModel / SQLAlchemy. - Pydantic V2 validation across API payloads, prompt configuration, and structured metadata. ^^^mermaid flowchart LR U[Browser User] --> A[FastAPI + NiceGUI App] A --> W[Asyncio Worker Engine] A --> DB[(Relational DB)] W --> P[Vision Provider APIs] W --> DB ^^^ ## Lifecycle Ownership Application lifespan owns runtime setup and teardown: - Initialize logging, settings, directories, and prompt configuration. - Manage asynchronous database engine connection pools. - Execute database bootstrap or migrations. - Recover stale or interrupted jobs on startup. - Manage graceful shutdown of active background tasks. ## Layered Module Structure ### Interface Layer - `src/transcription/ui/**` - `src/transcription/api/**` Responsibilities: - Render document, source, person, job, and classification views. - Accept user input for uploads, editing, linking, and revisions. - Present structured validation and conflict feedback. ### Application and Async Worker Layer - `src/transcription/services/workflows.py` - `src/transcription/worker.py` Responsibilities: - Orchestrate uploads, job creation, and status transitions. - Execute per-page provider calls through bounded concurrency. - Persist page-level outcomes and update aggregate job state. ### Domain and Service Layer - `src/transcription/db/models.py` - `src/transcription/services/documents.py` - `src/transcription/services/sources.py` - `src/transcription/services/jobs.py` - `src/transcription/services/people.py` - `src/transcription/services/workflows.py` Responsibilities: - Keep one primary service boundary per aggregate: Documents, Sources, Jobs, and People. - Documents own document records and the document-type registry. - Sources own source records, revisions, source media formats, MIME resolution, and page execution evidence. - Jobs own job lifecycle state and transitions. - People own person records, relationship roles, document-person links, and portrait media. - Apply deterministic conflict handling for relationship-role writes. - Use set-based synchronization for many-to-many relationship updates. - Resolve and validate registry-backed document types. ### Source Media Policy - `services/sources.py` is the single authority for accepted Source extensions and canonical MIME types. - Storage and provider payload loading must call the same Source validation functions. - Supported Source formats are JPEG, PNG, TIFF, and PDF. - Upload is an interface action, not a domain aggregate. Service names, errors, and workflow variables use `Source` terminology; compatibility aliases may remain temporarily at old import boundaries. ### Infrastructure Layer - `src/transcription/db/**` - `src/transcription/providers/**` Responsibilities: - Provide async database sessions and engine configuration. - Provide provider adapters for vision model execution. ## Core Workflows ### 1. Multi-Page Transcription 1. User uploads one or more images for a `Document`. 2. System stores files, hashes them, creates ordered `Source` rows, and creates a `Job`. 3. Worker claims the job, marks it `processing`, and executes page calls concurrently. 4. Each page writes a `JobSource` result with machine output, normalized metadata, and an SDK-serialized OpenRouter response snapshot. 5. Aggregate status becomes `completed`, `partial_success`, or `failed`. ### 2. Document-Person Relationship Management 1. User opens a document or person edit flow. 2. UI loads existing links grouped by role. 3. User adds or removes people within one or more roles. 4. Service computes add/remove deltas rather than replacing all links blindly. 5. Conflict checks enforce uniqueness and deterministic write semantics before persistence commits. ### 3. Document Type Management 1. User selects a registry-backed document type for a document. 2. Service resolves the stable type code or id. 3. Persistence stores the `document_type_id` reference. 4. Inactive types remain valid for historical rows but are excluded from default selectors. ## V4 Domain Rules - `JobSource.raw_transcription` preserves page output for its Job execution. - `Source.raw_transcription` is the latest-success machine-output projection for a page. - Human corrections occur only in `Source.revised_text`. - Prompt and parameter provenance is frozen on `Job` at submission time. - The SDK-serialized OpenRouter response snapshot is stored on `JobSource` for each successful page execution. - `DocumentPerson` links are unique for `(document_id, person_id, role_id)`. - Relationship mutations are deterministic and set-based. - `DocumentType.code` is stable; `DocumentType.label` may evolve. ## Data Model Summary - `Document` has one `DocumentType`, many `Source` pages, many `Job` runs, and many `Person` records through `DocumentPerson`. - `Source` belongs to one `Document` and may participate in many `JobSource` executions. - `Job` has many `JobSource` rows. - `PersonRole` defines available relationship roles. ## Test Strategy - Unit tests for models, validation, hashing, and registry resolution. - Service tests for CRUD, set-based sync, uniqueness conflicts, and deterministic relationship writes. - Async workflow tests for page isolation, partial failure handling, and stored evidence. - UI integration tests for multi-page rendering, role grouping, and document type selection. ## Related Local References - [System Overview](index_v4.md) - [System Requirements](requirements_v4.md) - [Data Model](schema_v4.md) - [Error Handling Policy](error_handling_v4.md) - [Error Handling Invariant](../invariant/error_handling.md) - [Digital Evidence and AI Processing Provenance](../invariant/ai_evidence_and_provenance.md)