9.8 KiB
System Architecture (Current Baseline: V6.1)
This document defines the current V6.1 architecture baseline.
Architecture Objectives
- Preserve durable archival records for Documents, Sources, People, and processing runs.
- Execute page transcription asynchronously with bounded worker behavior.
- Preserve append-only machine-attempt evidence with request/response provenance.
- Keep UI, API, service, persistence, and provider boundaries explicit and testable.
Technical Stack
- Runtime: Python 3.12+
- Web application: FastAPI + NiceGUI
- Persistence: SQLModel / SQLAlchemy — PostgreSQL in production, SQLite for local development and tests
- Validation and settings: Pydantic V2 + pydantic-settings
- Concurrency: asyncio worker loop
- Provider integration: OpenRouter adapter behind provider interface
- Deployment: Docker Compose (app, worker, PostgreSQL, Cloudflare Tunnel)
- Quality and tests: Ruff, ty, pytest, pytest-asyncio
Runtime Topology
flowchart LR
U[Browser User] --> A[FastAPI + NiceGUI App]
A --> W[Asyncio Worker]
A --> DB[(PostgreSQL / SQLite)]
W --> P[Provider Adapter]
W --> DB
The worker loop drains two queues in the same pass: queued transcription Jobs and queued
MaintenanceRun records. When neither has work, it idles.
Production deployment
Production runs as a Docker Compose stack with the app and worker as separate services, so the app
process runs with RUN_EMBEDDED_WORKER=false and the worker process owns queue draining. Local
development runs a single process with the worker embedded.
flowchart LR
I[Internet] --> CF[cloudflared tunnel + Access]
CF --> APP[app service]
APP --> PG[(postgres service)]
WK[worker service] --> PG
WK --> PROV[OpenRouter]
Deployment, rollback, and recovery procedures are in Production Runbook; backup configuration and restore are in Backup and Restore.
Layered Boundaries
Interface Layer
src/transcription/ui/**src/transcription/api/**
Responsibilities:
- Route registration, page orchestration, presentation adapters.
- Structured user messaging through shared error presenter.
- No direct persistence access from pages/components.
Service and Orchestration Layer
Aggregate services:
src/transcription/services/documents.pysrc/transcription/services/people.pysrc/transcription/services/jobs.pysrc/transcription/services/sources.pysrc/transcription/services/photos.pysrc/transcription/services/maintenance.pysrc/transcription/services/evidence.py(read/projection only)
Orchestration modules:
src/transcription/services/store.pysrc/transcription/services/workflows.py
Responsibilities:
- Aggregate ownership and invariants.
- Transaction-aware write helpers.
- Cross-service workflows in orchestration modules (
store.py,workflows.py). - Lookup-table CRUD through the generic
RegistryServicebase (registry.py), which is not an aggregate owner itself.
Module classification and per-model ownership are defined in services instructions.
Persistence Layer
src/transcription/db/**
Responsibilities:
- SQLModel definitions, async session/engine runtime, registry bootstrap.
- Loader helpers that enforce explicit eager loading with
lazy="raise"relationships.
Provider Layer
src/transcription/providers/**
Responsibilities:
- Provider API encapsulation.
- Request manifest and transport evidence capture.
- Normalized transcription result contract.
Core Domain Model
Documentowns archival metadata and links toSource,Job, andDocumentPerson.Sourceis a document page/file record with selected machine projection and human revision.Jobis an aggregate processing run with status and frozen prompt/runtime settings.JobSourceis queue/membership state for one(job, source)pair.ExecutionAttemptis append-only evidence for each provider call.Photois person imagery owned byPhotosService.MaintenanceRunis one queued or executed operational maintenance run.GenealogyPerson,GenealogyFamily,GenealogyFamilyChild, andGenealogyCitationstore imported GEDCOM genealogy data and citation provenance.DocumentTypeandPersonRoleare UUID-backed registries with optional protectedsemantic_key.Tagis a shared registry reached through both document and person tagging, linked byDocumentTagandPersonTag.
Processing and Evidence Workflow
- User creates/updates Document metadata and linked People atomically through workflow orchestration.
- User creates a Job by uploading one or more Source files or by retranscribing an existing Source.
- Source files are validated and stored; orientation normalization may be applied at ingest, and stored bytes become the canonical processing bytes.
- Worker claims queued Job, transitions to
processing, and processes pending pages in deterministic order. - Each provider call writes one immutable
ExecutionAttemptwith:- request manifest + hash
- transport evidence (when response exists)
- SDK snapshot and normalized metadata
- outcome, timing, and error details when applicable
JobSourcestatus is updated as queue/projection state;Source.raw_transcriptionis set on first successful attempt and can be explicitly re-pointed by candidate promotion.- Job terminal status resolves to
transcribed,partial_success, orfailed.
Status Semantics
- Job statuses:
queued,processing,transcribed,partial_success,failed- Operational success path resolves to
transcribed.
- Operational success path resolves to
- JobSource statuses:
pending,transcribed,failed,cancelled - MaintenanceRun statuses:
queued,processing,succeeded,failed- Maintenance uses
succeededrather thantranscribed; the transcription vocabulary does not apply to operational runs.
- Maintenance uses
- Maintenance job types:
backup,storage_reconciliation,gedcom_import
Maintenance Execution
Operational maintenance is queue-backed rather than run inline from the UI, so it survives request lifetime and is recorded:
- Settings enqueues a
MaintenanceRunwithstatus=queuedand atriggered_bymarker. - The worker claims the oldest queued run with a conditional update, moving it to
processing. backupruns the deploy backup script;storage_reconciliationcompares stored media againstDocument/Sourcerecords;gedcom_importparses the latest uploaded.gedfile and upserts genealogy records.- The run finalizes to
succeededorfailedwith summary, timing, log path, anderror_detail.
MaintenanceRun records operational history and is not evidence in the ExecutionAttempt sense;
append-only guarantees apply to transcription attempts.
Security and Path Handling Boundaries
- Print media delivery uses record-validated API route:
src/transcription/api/print_api.py
- General UI media links resolve through:
src/transcription/ui/components/media_urls.py
- Local filesystem paths must never be accepted from user input as trusted media routes.
Concurrency and Reliability Principles
- Worker loop reuses service bundle/provider resources for pooled calls.
- Provider-call timeout is explicit and bounded.
- Non-retriable worker-loop faults are surfaced and stop loop spin.
- Per-page outcomes are durably persisted before processing next page.
Design Decisions and Rationale
Why transcribed is the success terminal state
- The worker and job orchestration resolve successful completion to
JobStatus.TRANSCRIBED, with mixed and failure outcomes represented bypartial_successandfailed. - This keeps terminal status vocabulary aligned with what the pipeline actually produces: transcribed page content and evidence, not a generic completion marker.
Why evidence history is append-only while page text is a projection
ExecutionAttemptstores immutable per-call evidence and preserves full attempt history across retries.Source.raw_transcriptionis intentionally a mutable projection so UI and exports can show a selected current machine text without mutating historical evidence.- This split keeps auditability and UX both first-class: history is durable, presentation is editable.
Why orchestration modules own cross-service workflows
- Service modules do not import each other; aggregate ownership remains local to each service.
- Multi-aggregate writes are coordinated in orchestration modules (
store.py,workflows.py) so transaction boundaries are explicit and testable. - This avoids circular dependencies and keeps cross-cutting workflow logic centralized.
Why explicit eager loading is required
- ORM relationships are configured with
lazy="raise"in key paths, so code must request needed relationships up front. - This prevents hidden query behavior in UI/service code and makes read shape deterministic and reviewable.
Why canonical source bytes may be ingest-normalized
- Ingest normalization can correct orientation before persistence so provider calls, evidence hashes, and rendered processing source are consistent.
- The canonical stored bytes, digest, and size become the durable processing identity for that source.
Why media access uses controlled routes/helpers
- Print/export media uses record-validated API endpoints to avoid direct filesystem path exposure.
- General UI media URLs are generated through shared resolver helpers to keep path handling consistent and centralized.
Scope Boundary
Current architecture rules live in docs/*.