Files
transcription/docs/ver4/architecture_v4.md
T

8.7 KiB

System Architecture (Version 4)

This document describes the production architecture of the document transcription system.

Architecture Objectives

  • Preserve original source material, per-execution machine output, and separate human revision.
  • Support batching one or more images into ordered multi-page documents.
  • Capture submission-time prompt provenance and a per-page OpenRouter SDK response snapshot.
  • Execute page transcription concurrently with bounded asyncio workers.
  • Maintain relational portability across SQLite and PostgreSQL.
  • Keep operator workflows cross-platform and Python-driven.
  • Support many-to-many document-person relationships with extensible roles.
  • Support registry-driven document type classification.

Core Capabilities

  • Ingest one or more images into sequential Source pages under a Document.
  • Execute asynchronous vision transcription with bounded worker concurrency.
  • Preserve original source files with SHA-256 digests and byte sizes.
  • Freeze prompt text, prompt hash, model, and explicitly configured sampling parameters on each Job.
  • Preserve page-level machine output, normalized metadata, and an SDK-serialized OpenRouter response snapshot on JobSource.
  • Organize historical Person records through many-to-many Document relationships and extensible roles.
  • Classify Documents through a UUID-identified registry with unique labels.
  • Maintain human revision separately from machine-generated text.
  • Isolate page failures so multi-page jobs can complete with partial success.
  • Operate across supported platforms through Python-based application and maintenance tooling.

V4.2 extends this baseline with immutable execution attempts, exact OpenRouter transport evidence, safe versioned exports, and provider-neutral derived-artifact provenance. JobSource remains the mutable queue and compatibility projection; ExecutionAttempt is the authoritative append-only processing history. See the V4.2 Scope Boundary.

Technical Stack

  • Runtime: Python 3.12 or later.
  • Web application: FastAPI and NiceGUI.
  • Persistence: SQLModel and SQLAlchemy, with SQLite and PostgreSQL support.
  • Validation and settings: Pydantic V2 and pydantic-settings.
  • Concurrency: Python asyncio workers.
  • Vision integration: OpenRouter through the application's provider adapter.
  • Testing and quality: pytest, pytest-asyncio, Ruff, and ty.

Runtime Topology

The runtime operates as an asynchronous Python application:

  • FastAPI + NiceGUI web application process.
  • In-process asyncio worker engine for transcription execution.
  • Relational persistence via SQLModel / SQLAlchemy.
  • Pydantic V2 validation across API payloads, prompt configuration, and structured metadata.

^^^mermaid flowchart LR U[Browser User] --> A[FastAPI + NiceGUI App] A --> W[Asyncio Worker Engine] A --> DB[(Relational DB)] W --> P[Vision Provider APIs] W --> DB ^^^

Lifecycle Ownership

Application lifespan owns runtime setup and teardown:

  • Initialize logging, settings, directories, and prompt configuration.
  • Manage asynchronous database engine connection pools.
  • Execute database bootstrap or migrations.
  • Recover stale or interrupted jobs on startup.
  • Manage graceful shutdown of active background tasks.

Layered Module Structure

Interface Layer

  • src/transcription/ui/**
  • src/transcription/api/**

Responsibilities:

  • Render document, source, person, job, and classification views.
  • Accept user input for uploads, editing, linking, and revisions.
  • Present structured validation and conflict feedback.

Application and Async Worker Layer

  • src/transcription/services/workflows.py
  • src/transcription/worker.py

Responsibilities:

  • Orchestrate uploads, job creation, and status transitions.
  • Execute per-page provider calls through bounded concurrency.
  • Persist page-level outcomes and update aggregate job state.

Domain and Service Layer

  • src/transcription/db/models.py
  • src/transcription/services/documents.py
  • src/transcription/services/sources.py
  • src/transcription/services/jobs.py
  • src/transcription/services/people.py
  • src/transcription/services/workflows.py

Responsibilities:

  • Keep one primary service boundary per aggregate: Documents, Sources, Jobs, and People.
  • Documents own document records and the document-type registry.
  • Sources own source records, revisions, source media formats, MIME resolution, and page execution evidence.
  • Jobs own job lifecycle state and transitions.
  • People own person records, relationship roles, document-person links, and portrait media.
  • Apply deterministic conflict handling for relationship-role writes.
  • Use set-based synchronization for many-to-many relationship updates.
  • Resolve and validate registry-backed document types by UUID.

Source Media Policy

  • services/sources.py is the single authority for accepted Source extensions and canonical MIME types.
  • Storage and provider payload loading must call the same Source validation functions.
  • Supported Source formats are JPEG, PNG, TIFF, and PDF.
  • Upload is an interface action, not a domain aggregate. Service names, errors, and workflow variables use Source terminology; compatibility aliases may remain temporarily at old import boundaries.

Infrastructure Layer

  • src/transcription/db/**
  • src/transcription/providers/**

Responsibilities:

  • Provide async database sessions and engine configuration.
  • Provide provider adapters for vision model execution.

Core Workflows

1. Multi-Page Transcription

  1. User uploads one or more images for a Document.
  2. System stores files, hashes them, creates ordered Source rows, and creates a Job.
  3. Worker claims the job, marks it processing, and executes page calls concurrently.
  4. Each provider call appends an ExecutionAttempt with its request manifest, transport evidence, SDK snapshot, normalized metadata, timing, and outcome.
  5. The linked JobSource is updated as a compatibility projection, and a successful attempt updates the Source.raw_transcription latest-success projection.
  6. Aggregate status becomes completed, partial_success, or failed.

2. Document-Person Relationship Management

  1. User opens a document or person edit flow.
  2. UI loads existing links grouped by role.
  3. User adds or removes people within one or more roles.
  4. Service computes add/remove deltas rather than replacing all links blindly.
  5. Conflict checks enforce uniqueness and deterministic write semantics before persistence commits.

3. Document Type Management

  1. User selects a registry-backed document type for a document.
  2. Service resolves the Document Type UUID.
  3. Persistence stores the document_type_id reference.
  4. Inactive types remain valid for historical rows but are excluded from default selectors.

V4 Domain Rules

  • JobSource.raw_transcription preserves page output for its Job execution.
  • Source.raw_transcription is the latest-success machine-output projection for a page.
  • Human corrections occur only in Source.revised_text.
  • Prompt and parameter provenance is frozen on Job at submission time.
  • The SDK-serialized OpenRouter response snapshot is stored on JobSource for each successful page execution.
  • Every V4.2 provider call appends a distinct ExecutionAttempt; retries never rewrite earlier attempts.
  • Exact response bytes identify the OpenRouter HTTP boundary and are not labeled as native upstream-provider JSON.
  • Generic ProcessingArtifact records use versioned schemas, digests, and one inline or external content location.
  • DocumentPerson links are unique for (document_id, person_id, role_id).
  • Relationship mutations are deterministic and set-based.
  • DocumentType.id is canonical identity; its unique label may evolve.

Data Model Summary

  • Document has one DocumentType, many Source pages, many Job runs, and many Person records through DocumentPerson.
  • Source belongs to one Document and may participate in many JobSource executions.
  • Job has many JobSource rows.
  • PersonRole defines available relationship roles.

Test Strategy

  • Unit tests for models, validation, hashing, and registry resolution.
  • Service tests for CRUD, set-based sync, uniqueness conflicts, and deterministic relationship writes.
  • Async workflow tests for page isolation, partial failure handling, and stored evidence.
  • UI integration tests for multi-page rendering, role grouping, and document type selection.