generated from john/python-template
e61f7e75189b44d14549d76211cb23d486072f49
Transcription
Historical document transcription system for family-history documents.
The app lets you upload a document image/PDF, queues a background transcription job, and then shows job status and results in a web UI.
What the app does
- Upload document files (
.jpg,.jpeg,.png,.tif,.tiff,.pdf) - Persist document + job records in SQLite
- Process jobs in a background worker (
queued -> processing -> transcribed/failed) - Store transcript text (or failure detail)
- Show status and results in the NiceGUI interface
Quick start
1) Install dependencies
uv sync
2) Configure environment
Create a .env file in the project root (minimum required setting shown):
OPENROUTER_API_KEY=your_openrouter_api_key
Optional settings (defaults shown):
DATABASE_URL=sqlite:///./transcription.db
UPLOAD_DIR=./uploads
PROMPT_DIR=./prompts
3) Run the app
uv run uvicorn transcription.app:create_app --factory --reload
4) Open in browser
- GUI: http://[IP_ADDRESS]:8000/ui
- Health check: http://[IP_ADDRESS]:8000/healthz
How to navigate the GUI
-
Upload page (
/ui)- Select a supported file to upload.
- The app creates a queued transcription job.
- Use the View jobs link to inspect progress.
-
Jobs page (
/ui/jobs)- See all jobs and their status.
- Use Refresh to reload current states.
- Open a specific job to see details.
-
Job detail page (
/ui/jobs/{job_id})- Shows job metadata and status.
- Displays transcript text when successful.
- Displays failure detail when transcription fails.
Prompt artifacts
Prompt files are stored in prompts/ and loaded from PROMPT_DIR (default: ./prompts).
The canonical MVP prompt is:
prompts/transcribe_document.md
Description
A project to transcribe several thousand pages of family history documents
23 MiB
Languages
Python
100%