generated from john/python-template
1.7 KiB
1.7 KiB
Document Transcription System Overview (Version 4)
This project is a personal-scale application for transcribing, organizing, and preserving historical documents, images, and related people records.
Start Here
Read architecture_v4.md first for the technical overview and system design.
Core Capabilities
- Folder and multi-image ingestion into sequential
Sourcepages under a singleDocument. - Parallel asynchronous AI vision transcription using Python
asynciobounded by rate limits. - Portable relational storage using SQLModel and SQLAlchemy across SQLite and PostgreSQL.
- Complete prompt and response provenance for every transcription job and page execution.
- File-integrity tracking through SHA-256 hashing and stored file sizes.
- Historical
Personmanagement with many-to-many document links and extensible relationship roles. - Registry-driven
DocumentTypeclassification with stable codes and controlled selection. - Inline human revision of transcribed pages while preserving immutable machine output.
- Partial-failure recovery for multi-page jobs.
- Cross-platform operational workflows driven by Python-based tooling.
Technical Stack
- Application Web Framework: FastAPI + NiceGUI
- Persistence Engine: SQLModel / SQLAlchemy
- Data Validation and Schemas: Pydantic V2
- Concurrency and Workers: Python
asyncio - Vision Providers: OpenAI, Anthropic, and OpenRouter adapters