# Transcription Historical document transcription system for family-history documents. The app lets you upload a document image/PDF, queues a background transcription job, and then shows job status and results in a web UI. ## What the app does - Upload document files (`.jpg`, `.jpeg`, `.png`, `.tif`, `.tiff`, `.pdf`) - Persist document + job records in SQLite - Process jobs in a background worker (`queued -> processing -> transcribed/failed`) - Store transcript text (or failure detail) - Show status and results in the NiceGUI interface ## Quick start ### 1) Install dependencies ```bash uv sync ``` ### 2) Configure environment Create a `.env.production` file in the project root with the required OpenRouter API key: ```env OPENROUTER_API_KEY=your_openrouter_api_key ``` Settings are read from CLI arguments first, then environment variables, then `.env.production`, then the defaults below. ### Configuration Source Precedence When the same setting is provided in multiple places, the value is chosen in this order (highest priority first): 1. CLI arguments (for example `--port 8000`) 2. Settings constructor arguments (used mainly in tests) 3. Environment variables 4. `.env.production` file values 5. Model defaults in code Practical examples: - `--port 8000` overrides both `PORT=8000` in the shell and `PORT=7000` in `.env.production`. - `DATABASE__PATH=prod.db` in the shell overrides `DATABASE__PATH=dev.db` in `.env.production`. #### Server and runtime | Environment variable | Default | Description | | --- | --- | --- | | `HOST` | `0.0.0.0` | Address on which the server listens. | | `PORT` | `8000` | Server port. | | `LOG_LEVEL` | `info` | Uvicorn and application log level. | | `RELOAD` | `false` | Restart the development server when source files change. | | `ENVIRONMENT` | `development` | Runtime environment: `development`, `test`, or `production`. | | `RUN_EMBEDDED_WORKER` | `true` | Run worker loop inside web app process. Set `false` when using a dedicated worker service. | #### Provider | Environment variable | Default | Description | | --- | --- | --- | | `PROVIDER` | `openrouter` | Transcription provider. | | `OPENROUTER_API_KEY` | Required | OpenRouter API key. | | `PROVIDER_MODEL` | Provider default | Optional model override. | | `OPENROUTER_HTTP_REFERER` | Unset | Optional OpenRouter attribution URL. | | `OPENROUTER_APP_TITLE` | Unset | Optional OpenRouter attribution title. | #### Database and files Use nested env vars for database settings (recommended): ```env DATABASE__DRIVER=sqlite DATABASE__PATH=app.db # BOOTSTRAP_SCHEMA_ON_STARTUP=true SQLITE_CHECK_SAME_THREAD=false UPLOAD_DIR=./uploads PROMPT_DIR=./prompts DEFAULT_PROMPT_NAME=transcribe_document.md # TRANSCRIPTION_TEMPERATURE=0.2 # range: 0.0-2.0 # TRANSCRIPTION_TOP_P=0.9 # range: 0.0-1.0 ``` For PostgreSQL: ```env DATABASE__DRIVER=postgres DATABASE__HOST=localhost DATABASE__PORT=5432 DATABASE__DATABASE=transcription DATABASE__USER=postgres DATABASE__PASSWORD=change-me ``` This uses Pydantic nested settings (`env_nested_delimiter='__'`) and avoids JSON blobs in env files. A top-level `DATABASE={...}` JSON value is still supported as a fallback, and nested keys such as `DATABASE__PATH` take precedence over conflicting JSON keys. `BOOTSTRAP_SCHEMA_ON_STARTUP` creates missing tables when the app starts. When unset, it is enabled in `development` and `test`, and disabled in `production`; set it explicitly to override that policy. `SQLITE_CHECK_SAME_THREAD` defaults to `false`. #### Worker ```env WORKER_MAX_RETRIES=0 WORKER_RETRY_BACKOFF_SECONDS=0 WORKER_PROVIDER_TIMEOUT_SECONDS=20 WORKER_MIN_TRANSCRIPTION_CHARS=0 WORKER_MIN_TRANSCRIPTION_LINES=0 WORKER_FAIL_ON_FINISH_REASON_LENGTH=false ``` ### 3) Run the app ```bash uv run python -m transcription --port 8000 --reload --database.driver sqlite --bootstrap-schema-on-startup ``` This starts the development server with SQLite, creates missing tables, and enables automatic reload. Run `uv run python -m transcription --help` for all CLI options; CLI names use kebab case and nested database options use dot notation, such as `--database.path ./data/transcription.db`. ### 4) Open in browser - GUI: [http://localhost:8000/ui](http://localhost:8000/ui) - Health check: [http://localhost:8000/healthz](http://localhost:8000/healthz) Replace `localhost` with the server's hostname or IP address when connecting from another machine. ## Production stack (Phase 1) Use the production compose profile for split app/worker deployment with PostgreSQL and Cloudflare Tunnel: ```bash copy .env.production.example .env.production docker compose -f docker-compose.production.yml up -d --build ``` Services: - `app`: FastAPI + NiceGUI runtime (`RUN_EMBEDDED_WORKER=false`) - `worker`: standalone queue processor (`python -m transcription.worker_service`) - `postgres`: primary datastore - `cloudflared`: tunnel client using mounted ingress config + `CLOUDFLARE_TUNNEL_TOKEN` Operational defaults in the production compose file: - worker healthcheck is disabled (the worker process has no HTTP `/healthz` endpoint) - cloudflared is pinned to HTTP/2 with explicit DNS resolvers (`1.1.1.1`, `1.0.0.1`) for restricted LXC/container DNS environments Cloudflare setup files: 1. `copy deploy\cloudflared\config.yml.example deploy\cloudflared\config.yml` 2. set `CLOUDFLARE_TUNNEL_TOKEN` in `.env.production` 3. update ingress hostnames in `deploy\cloudflared\config.yml` ## How to navigate the GUI - **Upload page** (`/ui`) - Select a supported file to upload. - The app creates a queued transcription job. - Use the **View jobs** link to inspect progress. - **Jobs page** (`/ui/jobs`) - See all jobs and their status. - Use **Refresh** to reload current states. - Open a specific job to see details. - **Job detail page** (`/ui/jobs/{job_id}`) - Shows job metadata and status. - Displays transcript text when successful. - Displays failure detail when transcription fails. ## Prompt artifacts Prompt files are stored directly in `PROMPT_DIR` (default: `./prompts`). `DEFAULT_PROMPT_NAME` must be a filename, not a path. Each job snapshots the validated prompt text, SHA-256 hash, and sampling values for reproducibility. The canonical MVP prompt is: - `prompts/transcribe_document.md` ## Database migration workflow Schema upgrades use an explicit export/import rebuild flow (no runtime legacy write compatibility). See `docs/data_migration.md` for commands and cutover steps. ## Backup and restore workflow Production backup/restore (PostgreSQL + uploads + deployment config) steps are documented in `docs/backup_restore.md`. ## Destructive test procedure (with data backup) AI execution policy: before the first unit-test run in a test/fix cycle, create one backup of `./data`. Reuse that same backup for every subsequent test run in the cycle. After tests succeed, always pause and ask whether to restore now. Use the cross-platform Python wrapper below whenever an AI agent runs tests against this repository. 1. Create one backup of `./data` and mark it as the active test-cycle backup. 2. Run your test command. 3. On failure, fix the errors and run the wrapper again; it reuses the active backup and never backs up post-test data. 4. On success, always prompt whether to restore now (do not auto-restore unless explicitly approved). 5. Close the cycle only by restoring the active backup or explicitly accepting the current data. Preflight behavior: - Backup preflight is warning-only when `data/transcription.db` appears in use. - Restore preflight is blocking: the script prompts you to close conflicting applications, then type `retry` to re-check or `cancel` to skip restore. ### Run with confirmation-gated restore (default) ```bash uv run python tools/run_destructive_tests.py -- pytest tests/services/test_job_service.py tests/ui/test_jobs_page.py ``` After tests pass, the script asks whether to restore backup immediately. This is the required default mode for AI-assisted test runs because it gives time to verify and accept code changes before any restoration happens. ### Run with automatic restore (non-interactive) ```bash uv run python tools/run_destructive_tests.py --auto-restore -- pytest ``` ### Run without terminal prompt (decide restore later) ```bash uv run python tools/run_destructive_tests.py --skip-restore-prompt -- pytest ``` This keeps both the current post-test state and the backup, so restore can be decided explicitly later. Repeated wrapper invocations reuse the backup recorded in `.test-backups/.active-backup`. If that backup is missing, the wrapper stops rather than creating a replacement from potentially destructive post-test data. ### Restore later from a saved backup ```bash uv run python tools/run_destructive_tests.py --restore-from data-backup-YYYYMMDD-HHMMSS ``` To keep the current data and close the active cycle without restoring: ```bash uv run python tools/run_destructive_tests.py --accept-current-data ``` Backups are stored in `.test-backups/` and ignored by git.