Files
transcription/docs/ver4/architecture_v4.md
T

6.2 KiB

System Architecture (Version 4)

This document describes the production architecture of the document transcription system.

Architecture Objectives

  • Preserve original source material and immutable machine transcription output.
  • Support batching one or more images into ordered multi-page documents.
  • Capture complete submission-time prompt provenance and per-page provider response evidence.
  • Execute page transcription concurrently with bounded asyncio workers.
  • Maintain relational portability across SQLite and PostgreSQL.
  • Keep operator workflows cross-platform and Python-driven.
  • Support many-to-many document-person relationships with extensible roles.
  • Support registry-driven document type classification.

Runtime Topology

The runtime operates as an asynchronous Python application:

  • FastAPI + NiceGUI web application process.
  • In-process asyncio worker engine for transcription execution.
  • Relational persistence via SQLModel / SQLAlchemy.
  • Pydantic V2 validation across API payloads, prompt configuration, and structured metadata.

^^^mermaid flowchart LR U[Browser User] --> A[FastAPI + NiceGUI App] A --> W[Asyncio Worker Engine] A --> DB[(Relational DB)] W --> P[Vision Provider APIs] W --> DB ^^^

Lifecycle Ownership

Application lifespan owns runtime setup and teardown:

  • Initialize logging, settings, directories, and prompt configuration.
  • Manage asynchronous database engine connection pools.
  • Execute database bootstrap or migrations.
  • Recover stale or interrupted jobs on startup.
  • Manage graceful shutdown of active background tasks.

Layered Module Structure

Interface Layer

  • src/transcription/ui/**
  • src/transcription/api/**

Responsibilities:

  • Render document, source, person, job, and classification views.
  • Accept user input for uploads, editing, linking, and revisions.
  • Present structured validation and conflict feedback.

Application and Async Worker Layer

  • src/transcription/services/workflows.py
  • src/transcription/worker.py

Responsibilities:

  • Orchestrate uploads, job creation, and status transitions.
  • Execute per-page provider calls through bounded concurrency.
  • Persist page-level outcomes and update aggregate job state.

Domain and Service Layer

  • src/transcription/db/models.py
  • src/transcription/services/documents.py
  • src/transcription/services/sources.py
  • src/transcription/services/jobs.py
  • src/transcription/services/people.py
  • src/transcription/services/workflows.py

Responsibilities:

  • Keep one primary service boundary per aggregate: Documents, Sources, Jobs, and People.
  • Documents own document records and the document-type registry.
  • Sources own source records, revisions, source media formats, MIME resolution, and page execution evidence.
  • Jobs own job lifecycle state and transitions.
  • People own person records, relationship roles, document-person links, and portrait media.
  • Apply deterministic conflict handling for relationship-role writes.
  • Use set-based synchronization for many-to-many relationship updates.
  • Resolve and validate registry-backed document types.

Source Media Policy

  • services/sources.py is the single authority for accepted Source extensions and canonical MIME types.
  • Storage and provider payload loading must call the same Source validation functions.
  • Supported Source formats are JPEG, PNG, TIFF, and PDF.
  • Upload is an interface action, not a domain aggregate. Service names, errors, and workflow variables use Source terminology; compatibility aliases may remain temporarily at old import boundaries.

Infrastructure Layer

  • src/transcription/db/**
  • src/transcription/providers/**

Responsibilities:

  • Provide async database sessions and engine configuration.
  • Provide provider adapters for vision model execution.

Core Workflows

1. Multi-Page Transcription

  1. User uploads one or more images for a Document.
  2. System stores files, hashes them, creates ordered Source rows, and creates a Job.
  3. Worker claims the job, marks it processing, and executes page calls concurrently.
  4. Each page writes a JobSource result with raw output, metadata, and full provider response evidence.
  5. Aggregate status becomes completed, partial_success, or failed.

2. Document-Person Relationship Management

  1. User opens a document or person edit flow.
  2. UI loads existing links grouped by role.
  3. User adds or removes people within one or more roles.
  4. Service computes add/remove deltas rather than replacing all links blindly.
  5. Conflict checks enforce uniqueness and deterministic write semantics before persistence commits.

3. Document Type Management

  1. User selects a registry-backed document type for a document.
  2. Service resolves the stable type code or id.
  3. Persistence stores the document_type_id reference.
  4. Inactive types remain valid for historical rows but are excluded from default selectors.

Domain Invariants

  • Source.raw_transcription stores immutable machine output.
  • Human corrections occur only in Source.revised_text.
  • Prompt and parameter provenance is frozen on Job at submission time.
  • Provider output evidence is stored on JobSource for each page execution.
  • DocumentPerson links are unique for (document_id, person_id, role_id).
  • Relationship mutations are deterministic and set-based.
  • DocumentType.code is stable; DocumentType.label may evolve.

Data Model Summary

  • Document has one DocumentType, many Source pages, many Job runs, and many Person records through DocumentPerson.
  • Source belongs to one Document and may participate in many JobSource executions.
  • Job has many JobSource rows.
  • PersonRole defines available relationship roles.

Test Strategy

  • Unit tests for models, validation, hashing, and registry resolution.
  • Service tests for CRUD, set-based sync, uniqueness conflicts, and deterministic relationship writes.
  • Async workflow tests for page isolation, partial failure handling, and stored evidence.
  • UI integration tests for multi-page rendering, role grouping, and document type selection.