diff --git a/docs/architecture_v3.md b/docs/architecture_v3.md deleted file mode 100644 index f859732..0000000 --- a/docs/architecture_v3.md +++ /dev/null @@ -1,142 +0,0 @@ -# System Architecture (Version 3) - -This document describes the V3 production architecture of the personal historical-document transcription system. - -## Architecture Objectives - -* Preserve source material as immutable transcribed text alongside page-level spatial AI metadata and complete provider API envelopes. -* Support batching multi-image and folder uploads cleanly into sequential pages (`page_number`). -* Capture complete input prompt provenance (`system_prompt`, `user_prompt`, `prompt_hash`) and execution parameters (`temperature`, `top_p`) at submission time on `Job`. -* Leverage asynchronous worker pools (`asyncio`) for parallel single-image API execution bounded by rate limiters (`asyncio.Semaphore`). -* Maintain relational data-model portability across the supported backends by using SQLModel/SQLAlchemy and compatibility types so the same domain schema works in SQLite for local development/testing and PostgreSQL in production. -* Keep operator tooling and local maintenance workflows OS-independent by using Python or other cross-platform interfaces for canonical project automation. -* Verify image asset integrity via SHA-256 file hashing (`file_hash`) while storing binary assets on the local filesystem. -* Standardize all data validation, API parsing, and database models on **Pydantic V2** and **SQLModel**. -* Support rich historical attribution (multi-author and multi-recipient relationships via `DocumentPerson`). - -## Runtime Topology - -The V3 runtime operates as an asynchronous Python application: - -* FastAPI + NiceGUI web application process. -* In-process `asyncio` background task orchestrator for parallel API execution. -* Relational persistence via SQLModel / SQLAlchemy, using SQLite for local development/testing and PostgreSQL as the production persistence target. -* Pydantic V2 validation layer wrapping API payloads, prompt configurations, and JSON metadata schemas. -* Cross-platform operator workflows implemented in Python so core local operations run consistently on Windows, Linux, and macOS. - -^^^mermaid -flowchart LR -U[Browser User] --> A[FastAPI + NiceGUI App] -A --> W[Asyncio Worker Engine] -A --> DB[(Relational DB\nSQLite / PostgreSQL)] -W --> P[Vision Provider APIs\nOpenAI / Claude / OpenRouter] -W --> DB -^^^ - -## Lifecycle Ownership - -Application lifespan owns runtime setup/teardown: - -* Initialize environment logging, directory paths, and Pydantic configuration. -* Manage asynchronous database engine connection pools (`aiosqlite` or `asyncpg`). -* Execute database bootstrap (`SQLModel.metadata.create_all()`) or migrations. -* Recover stale or interrupted processing jobs on startup. -* Manage graceful shutdown of active `asyncio` worker pools. - -## Layered Module Structure - -### Interface Layer - -* `src/transcription/ui/**` (NiceGUI pages, multi-page renderers, person cards) -* `src/transcription/api/**` (FastAPI routes and JSON error handlers) - -### Application & Async Worker Layer - -* `src/transcription/services/workflows.py` -* `src/transcription/worker.py` - -Responsibilities: - -* Batch orchestration and status transitions (`queued` -> `processing` -> `transcribed` | `partial_success` | `failed`). -* Parallel single-image API execution using `asyncio.gather` bounded by `asyncio.Semaphore`. -* Resolve prompt configuration at submission time and persist frozen snapshot fields on `Job`. -* Pydantic schema parsing and validation prior to database storage. - -### Domain & Service Layer - -* `src/transcription/db/models.py` (SQLModel schema definitions for Document, Source, Job, JobSource, Person, DocumentPerson) -* `src/transcription/services/*.py` (Transactional operations for `Document`, `Person`, `Source`, `Job`, and `JobSource`) - -### Infrastructure Layer - -* `src/transcription/db/**` (Async database session factory, engine creation, and JSON dialect abstractions) -* `src/transcription/providers/**` (OpenAI, Anthropic, and OpenRouter Vision SDK adapters) - -## Processing Workflow - -1. User uploads a folder or batch of images for a `Document`. -2. System hashes each image file (SHA-256), writes image files to filesystem storage, and creates `Document`, `Job(status='queued')`, and ordered `Source` pages (`page_number = 1..N`). -3. Worker claims job, sets `Job.status = 'processing'`, and spawns parallel `asyncio` tasks bounded by semaphore. -4. Each task reads the frozen prompt snapshot from `Job` and calls Vision API for a **single** `Source` image. -5. On task completion: -* Writes a `JobSource` record containing `status='transcribed'`, `raw_transcription`, operational `ai_metadata`, and complete unedited `raw_api_response`. -* Caches active output text to `Source.raw_transcription`. - - -6. On page failure: -* Writes `JobSource` record with `status='failed'` and `error_detail`. - - -7. Once all page tasks resolve: -* Marks `Job.status` as `transcribed` (100% success), `partial_success` (at least 1 success, 1 failure), or `failed` (all failed). - - - -## Domain Ownership & Invariants - -* **Immutable AI Outputs:** `source.raw_transcription` and `job_source.raw_transcription` store original, point-in-time machine output and are immutable. -* **Complete Input & Output Provenance:** Every `job` stores the exact frozen input configuration sent to the model, and every `job_source` stores per-page output evidence including the complete REST response envelope returned. -* **Inlined Revisions:** Human corrections occur on `source.revised_text`. UI renders `COALESCE(revised_text, raw_transcription)`. -* **Sequential Integrity:** Multi-page documents are strictly ordered by `source.page_number ASC`. -* **Page Execution Isolation:** A failure on one page image does not invalidate successful transcriptions on sister pages in the same batch job. - -## Data Model Summary - -* `Document` has many `Source` pages, many `Job` runs, and many `Person` records via `DocumentPerson` junction (`author` or `recipient`). -* `Source` belongs to one `Document` and can be processed across many `JobSource` executions. -* `Job` has many `JobSource` execution records. -* `JobSource` holds page-level execution status, output text, and raw response JSON. - -## Test Strategy - -* Unit tests for SQLModel/Pydantic V2 models, JSON cross-dialect serialization, and file hashing functions. -* Integration tests for async database connection handling, session management, and queries. -* Async workflow tests using mock AI providers to verify `partial_success`, page-level failure isolation, and retry logic. -* UI integration tests for multi-page rendering and person attribution management. - ---- - -## Technology References - -* [FastAPI documentation](https://fastapi.tiangolo.com/) -* [NiceGUI documentation](https://nicegui.io/documentation) -* [SQLModel documentation](https://sqlmodel.tiangolo.com/) -* [SQLAlchemy Async I/O documentation](https://docs.sqlalchemy.org/en/20/orm/extensions/asyncio.html) -* [Python asyncio](https://www.google.com/search?q=https://docs.python.org/3/library/asyncio.html%23module-asyncio) -* [Pydantic Validation](https://pydantic.dev/docs/validation/latest/get-started/) - -## Related Local References - -- [System Overview](index_v3.md) -- [System Design Intent](invariant/intent.md) -- [Transcription Methodology](invariant/transcription_methodology.md) -- System Architecture (this document) -- [System Requirements](requirements_v3.md) -- [Data model](schema_v3.md) -- [Error Handling Policy](error_handling_v3.md) -- [Implementation Plan](implementation_plan_v3.md) - - - - - diff --git a/docs/error_handling_v3.md b/docs/error_handling_v3.md deleted file mode 100644 index e6f3ca5..0000000 --- a/docs/error_handling_v3.md +++ /dev/null @@ -1,87 +0,0 @@ -# Error Handling Policy (Version 3) - -This document defines the canonical error-handling policy for the v3 document transcription system. - -## Error Handling Objectives - -* Make failures visible in clear, actionable language at both the document and individual page levels. -* Support **isolated failure handling** in multi-image batches so single page errors do not crash an entire batch job. -* Preserve diagnostic detail (Pydantic validation errors, raw provider REST envelopes, exact input prompts) in generic database JSON structures for fast troubleshooting. -* Ensure consistent error envelope structure across API, UI, and async worker boundaries. - -## Scope And Authority - -Governs error behavior across NiceGUI pages, FastAPI routes, service orchestration, `asyncio` background tasks, database interactions, and AI provider adapters. - -## Error Taxonomy - -| Category | Definition | Retriable | -| --- | --- | --- | -| `validation_error` | Pydantic payload or parameter schema validation failure | no | -| `user_input_error` | Unacceptable user file (unsupported image type, corrupt file) | no | -| `not_found_error` | Requested resource (`Document`, `Source`, `Person`, `Job`) missing | no | -| `conflict_error` | Operation violates state constraints (e.g., duplicate `document_person` role) | no | -| `external_provider_error` | AI Provider API failure (rate limit, vision execution error) | yes | -| `infrastructure_transient_error` | Temporary DB connection reset or HTTP timeout | yes | -| `infrastructure_persistent_error` | Database down, missing API credentials, misconfiguration | no | -| `internal_unexpected_error` | Uncaught Python exception or logic defect | no | - -## Async Batch & Page-Level Error Behavior - -In multi-image `asyncio` batch processing: - -1. **Page Isolation:** Exceptions caught during individual page calls are trapped within the `asyncio` task wrapper. -2. **Page Record Logging:** Page failure details, along with the prompt inputs and hyperparameters attempted, are written directly to `job_source.error_detail` and `job_source.status = 'failed'`. -3. **Batch Aggregate State:** -* If **all** page tasks succeed -> `job.status = 'completed'`. -* If **some** page tasks fail -> `job.status = 'partial_success'`. -* If **all** page tasks fail -> `job.status = 'failed'`. - - -4. **Retry Strategy:** The UI exposes a "Retry Failed Pages" option for `partial_success` jobs, which spawns a new targeted `Job` containing *only* the `Source` IDs marked as `failed`. - -## API Error Response Contract - -API error responses return a structured JSON envelope: -^^^json -{ -"error_id": "err_uuid_12345", -"category": "validation_error", -"message": "The uploaded payload failed schema validation.", -"suggestion": "Check file format and metadata fields, then try again.", -"details": { -"pydantic_errors": [...] -}, -"timestamp": "2026-08-08T15:00:00Z" -} -^^^ - -HTTP Status Mappings: - -* `validation_error`, `user_input_error` -> `400` -* `not_found_error` -> `404` -* `conflict_error` -> `409` -* `external_provider_error` -> `502` / `503` -* `infrastructure_transient_error` -> `503` -* `infrastructure_persistent_error`, `internal_unexpected_error` -> `500` - ---- - -## Technology References - -* [FastAPI documentation](https://fastapi.tiangolo.com/) -* [NiceGUI documentation](https://nicegui.io/documentation) -* [SQLModel documentation](https://sqlmodel.tiangolo.com/) -* [Python asyncio](https://www.google.com/search?q=https://docs.python.org/3/library/asyncio.html%23module-asyncio) -* [Pydantic Validation](https://pydantic.dev/docs/validation/latest/get-started/) - -## Related Local References - -- [System Overview](index_v3.md) -- [System Design Intent](invariant/intent.md) -- [Transcription Methodology](invariant/transcription_methodology.md) -- [System Architecture](architecture_v3.md) -- [System Requirements](requirements_v3.md) -- [Data model](schema_v3.md) -- Error Handling Policy (this document) -- [Implementation Plan](implementation_plan_v3.md) diff --git a/docs/implementation_plan_v3.md b/docs/implementation_plan_v3.md deleted file mode 100644 index 87ce1f1..0000000 --- a/docs/implementation_plan_v3.md +++ /dev/null @@ -1,81 +0,0 @@ -# Implementation Plan (Version 3) - -## Goal - -Replace the current v2 SQLModel schema with the approved v3 schema and make sure every database operation works through the existing async SQLAlchemy/SQLModel session layer. - -Use a fresh database. There will be no migrations, data conversion, legacy compatibility shims, or parallel v2/v3 code paths. - -## Current Project Impact - -* `src/transcription/db/models.py` defines the SQLModel tables. It must be updated to match the approved v3 schema (`Document`, `Person`, `DocumentPerson`, `Source`, `Job`, `JobSource`). -* The v3 target adds frozen submission-time prompt snapshot fields (`prompt_name`, `prompt_hash`, `system_prompt`, `user_prompt`, `temperature`, `top_p`) to `Job` and full output payloads (`raw_api_response`, `ai_metadata`) to `JobSource`. -* The v3 target adds image asset verification fields (`file_hash`, `file_size_bytes`) to `Source`. -* Database operations must utilize `JSONBCompat` and the existing SQLModel/SQLAlchemy abstractions to preserve the same logical schema and JSON behavior across the supported backends, while keeping PostgreSQL as the intended production database. -* Async CRUD lives in `DocumentService`, `JobService`, `TranscriptionService`, and upload helpers. Their queries and relationship loading must be updated for v3 fields. -* Canonical operator tooling must remain OS-independent; safety workflows such as destructive-test backup and restore should run through Python or other cross-platform entry points rather than platform-specific shells. - -## Implementation - -### 1. Update the Schema and Domain Models - -* Replace the models in `src/transcription/db/models.py` with the approved v3 tables, enums, relationships, foreign keys, constraints, and indexes. -* Ensure all JSON fields use `JSONBCompat` for dialect portability across SQLite and PostgreSQL. -* Keep `SQLModel.metadata.create_all()` as the schema bootstrap for fresh databases. -* Delete `_ensure_sqlite_compat_columns()` and all legacy schema patching from `src/transcription/db/operations.py`. -* Keep the Python models and `docs/schema_v3.md` perfectly synchronized. - -### 2. Update Data Services and Async Worker Layer - -* Update job creation and worker orchestration so prompt configuration is resolved at submission and frozen onto `Job` (`prompt_name`, `prompt_hash`, `system_prompt`, `user_prompt`, `temperature`, `top_p`) before execution starts. -* Update `TranscriptionService` and provider adapters to store the complete unedited API REST response dictionary into `job_source.raw_api_response` alongside operational metrics in `job_source.ai_metadata`. -* Update upload handlers to calculate and store file metadata (`file_hash` via SHA-256, `file_size_bytes`) on `Source` records during file ingestion. -* Remove legacy single-source compatibility flows so worker paths persist per-page outcomes only through `JobSource` updates. - -### 3. Update Integration Tests and Mock AI Providers - -* Update mock provider fixtures in test suites to return realistic complete API response envelopes. -* Verify test coverage for `JSONBCompat` field writes and reads under SQLite in-memory test databases. -* Add assertions in async workflow tests to verify frozen prompt snapshot fields on `Job`, plus per-page failure isolation and output evidence on `JobSource`. - -### 4. Update the UI for the v3 Schema - -* Review the UI components and views displaying document, job, person, and source data so they reference v3 schema properties instead of v2 relationships. -* Ensure the UI correctly renders `COALESCE(revised_text, raw_transcription)` for page viewing and inline editing. -* Ensure resubmit actions only queue failed pages and preserve frozen prompt snapshot behavior on the existing `Job`. -* Consider the guidance in `docs/ui_style_guide.md` when making UI changes so updated views remain consistent with the project’s visual conventions. - -### 5. Keep Operational Tooling Portable - -* Implement destructive-test backup and restore workflows in Python so the canonical path runs on Windows, Linux, and macOS. -* Avoid making core developer or recovery procedures depend on PowerShell-only or shell-specific semantics. -* Keep operational documentation aligned with the cross-platform command path used by the repository. - -## Done When - -* A fresh database is created directly from the v3 SQLModel metadata. -* Frozen prompt input provenance is captured on `Job` for each submission, and full per-page output evidence is captured on `JobSource` for every AI execution task. -* The focused tests and full test suite pass on both SQLite and PostgreSQL backends. -* Canonical operator workflows required for development and destructive-test recovery run without a Windows-only shell dependency. - -## Out of Scope - -* Database migrations or preservation of v2 data -* Legacy compatibility code -* UI redesign, batch orchestration, worker concurrency, deployment, and operational runbooks - ---- - -## Related Local References - -- [System Overview](index_v3.md) -- [System Design Intent](invariant/intent.md) -- [Transcription Methodology](invariant/transcription_methodology.md) -- [System Architecture](architecture_v3.md) -- [System Requirements](requirements_v3.md) -- [Data model](schema_v3.md) -- [Error Handling Policy](error_handling_v3.md) -- Implementation Plan (this document) - - - diff --git a/docs/index_v3.md b/docs/index_v3.md deleted file mode 100644 index cbac3b1..0000000 --- a/docs/index_v3.md +++ /dev/null @@ -1,49 +0,0 @@ -# Document Transcription System Overview (Version 3) - -This project is a personal-scale application for transcribing, indexing, and preserving historical family documents, letters, postcards, and journals. - -## Start Here - -Read [architecture_v3.md](https://www.google.com/search?q=architecture_v3.md) first for technical overview and system design. - -## Core V3 Capabilities - -* **Folder & Multi-Image Ingestion:** Upload one or more images that map sequentially (`page_number`) under a single `Document`. -* **Parallel Async AI Vision Engine:** Concurrently process single-page image transcriptions using Python `asyncio` bounded by rate limiters. -* **Portable Relational Storage:** SQLModel and SQLAlchemy preserve a portable relational model across the supported backends, with SQLite for local development/testing and PostgreSQL as the production database target. -* **Cross-Platform Operations:** Canonical developer and recovery workflows run through Python-based, OS-independent tooling rather than platform-specific shell scripts. -* **Complete Auditability & Provenance:** Capture frozen submission-time input prompts (`system_prompt`, `user_prompt`) and hyperparameters (`temperature`, `top_p`) on `Job`, plus per-page operational metrics (`ai_metadata`) and full provider response envelopes (`raw_api_response`) on `JobSource`. -* **Asset Integrity Tracking:** Calculate and store cryptographic hashes (SHA-256) and file sizes on `Source` image records while preserving clean filesystem storage. -* **Pydantic V2 Validation:** End-to-end type safety, DB row mapping, and JSON payload validation. -* **Historical Person Management:** Track authors and recipients across documents with rich biographical entities (`Person`). -* **Page-Level Execution Auditing & Revisions:** Store immutable point-in-time machine output per run while enabling inline human corrections (`revised_text`). -* **Partial Failure Recovery:** Bounded batch execution that isolates single-page API errors (`partial_success`) for simple retries. - -## Technical Stack - -* **Application Web Framework:** FastAPI + NiceGUI -* **Persistence Engine:** SQLModel / SQLAlchemy (SQLite for development/testing, PostgreSQL for production) -* **Data Validation & Schemas:** Pydantic V2 -* **Concurrency & Workers:** Python `asyncio` worker pool with `asyncio.Semaphore` -* **Vision Providers:** OpenAI, Anthropic, and OpenRouter Vision models via native SDK adapters - ---- - -## Technology References - -* [FastAPI documentation](https://fastapi.tiangolo.com/) -* [NiceGUI documentation](https://nicegui.io/documentation) -* [SQLModel documentation](https://sqlmodel.tiangolo.com/) -* [Python asyncio](https://www.google.com/search?q=https://docs.python.org/3/library/asyncio.html%23module-asyncio) -* [Pydantic Validation](https://pydantic.dev/docs/validation/latest/get-started/) - -## Documentation Index - -- System Overview (this document) -- [System Design Intent](invariant/intent.md) -- [Transcription Methodology](invariant/transcription_methodology.md) -- [System Architecture](architecture_v3.md) -- [System Requirements](requirements_v3.md) -- [Data model](schema_v3.md) -- [Error Handling Policy](error_handling_v3.md) -- [Implementation Plan](implementation_plan_v3.md) diff --git a/docs/requirements_v3.md b/docs/requirements_v3.md deleted file mode 100644 index 9a01c7c..0000000 --- a/docs/requirements_v3.md +++ /dev/null @@ -1,45 +0,0 @@ -# Document Transcription System Requirements (Version 3) - -This document captures the **Version 3 baseline requirements** for the production implementation. - -## Requirements Model - -| ID | Category | Requirement | Verify Method | -| --- | --- | --- | --- | -| REQ-0 | System | Provide end-to-end multi-page document transcription with persistent, inspectable async job states. | demonstration | -| REQ-1 | Functional | Allow users to upload multi-image batches as sequential `Source` pages under a `Document`. | test | -| REQ-2 | Functional | Process multi-page jobs asynchronously using an `asyncio` worker pool bounded by rate limits. | test | -| REQ-3 | Functional | Persist frozen submission-time execution parameters and full input prompts (`system_prompt`, `user_prompt`, `prompt_name`, `prompt_hash`, `temperature`, `top_p`) on `Job`, and persist page-level output responses (`raw_transcription`, `ai_metadata`, `raw_api_response`) on `JobSource`. | test | -| REQ-4 | Functional | Support job states (`queued`, `processing`, `completed`, `partial_success`, `failed`) and page states (`pending`, `transcribed`, `failed`). | inspection | -| REQ-5 | Functional | Allow users to manage historical `Person` records and link multiple authors/recipients to a `Document` via `DocumentPerson`. | test | -| REQ-6 | Functional | Maintain immutable original machine output on `Source.raw_transcription` while permitting inline human edits on `Source.revised_text`. | test | -| REQ-7 | Data Constraint | Use SQLModel/SQLAlchemy to preserve a portable relational domain model and compatible data shape across the supported backends, with SQLite for local development/testing and PostgreSQL as the production system of record. | inspection | -| REQ-8 | Data Constraint | Validate all API requests, database rows, and JSON structures using Pydantic V2 schemas and SQLModel. | test | -| REQ-9 | Interface | Render multi-page transcriptions sequentially by `page_number` in the web UI with author/recipient metadata. | demonstration | -| REQ-10 | Operations | Allow operators to resubmit only failed pages for queued reprocessing while preserving the frozen prompt snapshot on the existing `Job`. | test | -| REQ-11 | Data Constraint | Calculate and store cryptographic file hashes (SHA-256) and file sizes for uploaded source images to track asset integrity. | test | -| REQ-12 | Operations Constraint | Keep core development, testing, restore, and recovery workflows OS-independent across Windows, Linux, and macOS; do not require a platform-specific shell for canonical project processes. | inspection | - -## Element Satisfaction Mapping - -* **UI (NiceGUI):** Satisfies REQ-1, REQ-5, REQ-6, REQ-9, REQ-10. -* **API (FastAPI):** Satisfies REQ-1, REQ-4, REQ-5, REQ-8. -* **WORKER (asyncio):** Satisfies REQ-2, REQ-3, REQ-4, REQ-10. -* **PERSISTENCE (SQLModel/SQLAlchemy):** Satisfies REQ-3, REQ-6, REQ-7, REQ-11. -* **MODELS (Pydantic V2 / SQLModel):** Satisfies REQ-8. -* **OPERATIONS TOOLING (Python / OS-neutral automation):** Satisfies REQ-12. - ---- - -## Related Local References - -- [System Overview](index_v3.md) -- [System Design Intent](invariant/intent.md) -- [Transcription Methodology](invariant/transcription_methodology.md) -- [System Architecture](architecture_v3.md) -- System Requirements (this document) -- [Data model](schema_v3.md) -- [Error Handling Policy](error_handling_v3.md) -- [Implementation Plan](implementation_plan_v3.md) - - diff --git a/docs/schema_v3.md b/docs/schema_v3.md deleted file mode 100644 index 5bb69cd..0000000 --- a/docs/schema_v3.md +++ /dev/null @@ -1,137 +0,0 @@ -# Database Schema (Version 3) - -This document describes the relational schema for the transcription platform. It incorporates multi-image batch orchestration, page-level execution tracking, many-to-many author/recipient attribution, submission-time prompt snapshot capture, and raw API payload evidence for archival auditing. - -The schema uses generic JSON columns compatible with SQLite in local development and PostgreSQL native JSONB/UUID types in production. - -## Entity Relationship Diagram - -```mermaid -erDiagram -PERSON { -UUID id PK -TEXT full_name -TEXT display_name -TEXT maiden_name -DATE birth_date -TEXT birth_date_raw -TEXT birth_place -DATE death_date -TEXT death_date_raw -TEXT death_place -TEXT biography -TEXT portrait_path -JSONB metadata -TIMESTAMPTZ created_at -TIMESTAMPTZ updated_at -} - -DOCUMENT { - UUID id PK - TEXT name - TEXT document_type - DATE document_date - TEXT document_date_raw - TEXT location_created - TEXT notes - TEXT archive_identifier - TIMESTAMPTZ created_at - TIMESTAMPTZ updated_at -} - -DOCUMENT_PERSON { - UUID id PK - UUID document_id FK - UUID person_id FK - VARCHAR role "author | recipient" - TIMESTAMPTZ created_at -} - -JOB { - UUID id PK - UUID document_id FK - VARCHAR status "queued | processing | transcribed | completed | partial_success | failed" - INTEGER retry_count - TEXT provider - TEXT model - TEXT prompt_name - TEXT prompt_hash - TEXT system_prompt - TEXT user_prompt - FLOAT temperature - FLOAT top_p - TIMESTAMPTZ date_created - TIMESTAMPTZ date_updated -} - -SOURCE { - UUID id PK - UUID document_id FK - INTEGER page_number - TEXT upload_name - TEXT filename - TEXT file_path - TEXT file_hash - BIGINT file_size_bytes - TEXT raw_transcription - TEXT revised_text - TIMESTAMPTZ date_uploaded - TIMESTAMPTZ date_revised -} - -JOB_SOURCE { - UUID id PK - UUID job_id FK - UUID source_id FK - VARCHAR status "pending | transcribed | failed" - TEXT raw_transcription - JSONB ai_metadata - JSONB raw_api_response - TEXT error_detail - TIMESTAMPTZ executed_at -} - -DOCUMENT ||--o{ DOCUMENT_PERSON : "has_people" -PERSON ||--o{ DOCUMENT_PERSON : "participates_in" -DOCUMENT ||--o{ JOB : "has_jobs" -DOCUMENT ||--o{ SOURCE : "contains_pages" -JOB ||--o{ JOB_SOURCE : "executes" -SOURCE ||--o{ JOB_SOURCE : "processed_in" -``` - -## Domain Invariants & Provenance Rules - -### Page-Level Execution & AI Outputs - -* **Execution Granularity:** Every single image execution attempt by an AI model produces a dedicated record in `job_source`. -* **Submission Snapshot Provenance:** Every `job` captures the frozen prompt identifier details (`prompt_name`, `prompt_hash`), full prompt text strings (`system_prompt`, `user_prompt`), and hyperparameters (`temperature`, `top_p`) at submission time. -* **Point-in-Time Output Auditability:** `job_source.raw_api_response` stores the complete, unedited provider REST response envelope for that specific image page call. `job_source.ai_metadata` stores spatial bounding boxes, normalized token usage, latency, and cost details for fast querying. -* **Active Output Caching:** Upon successful completion of an image call, `source.raw_transcription` is updated with the latest output string from `job_source.raw_transcription` for fast UI rendering. - -### Image Storage & Integrity - -* **Filesystem Storage:** Binary images are stored on disk in the local file system. The `source` table holds the relative `file_path`. -* **File Integrity Tracking:** `source` captures `file_hash` (SHA-256) and `file_size_bytes` at upload time to guarantee document file integrity and duplicate checking over long-term preservation. - -### Page Ordering & Revisions - -* **Sequential Integrity:** `source.page_number` dictates page ordering within a document. Reads assembling full documents must query `ORDER BY source.document_id, source.page_number ASC`. -* **Inlined Human Corrections:** User edits occur at the page level inside `source.revised_text`. `source.raw_transcription` remains immutable. If `source.revised_text` is non-null, application frontends must render `source.revised_text`. - -### Async Job Lifecycle & Failure Isolation - -* **Batch Orchestrator:** A job represents an overarching execution run across one or more source images belonging to a document. -* **Isolated Failures:** API requests run concurrently (e.g., using `asyncio`). A failure on page 3 does not invalidate successful transcriptions on page 1 or 2. -* **Job States:** -* `queued`: Created, awaiting worker execution. -* `processing`: Concurrent HTTP tasks actively running. -* `completed`: 100% of linked `job_source` tasks succeeded (`transcribed`). -* `partial_success`: At least one `job_source` succeeded and at least one failed. -* `failed`: All linked `job_source` tasks failed or a job-level runtime error occurred. - - - -### Attribution & Person Roles - -* **Multi-Person Roles:** Documents support zero, one, or many authors and recipients linked via `document_person`. -* **Role Uniqueness:** `(document_id, person_id, role)` must be unique to prevent duplicate role tagging.