zoltan57 97b3d0fd62 V4.6 Phase 4: service layer consolidation
Removes the duplicated registry CRUD, the hand-written not-found raises, and
the three divergent media writers. Behavior is preserved: every existing
Document Type and Person Role test passes unchanged, which is the primary
proof for MED-11.

[MED-11] Generic registry service
- New services/registry.py owns RegistryService[ModelT]: list, list with
  counts, create with IntegrityError -> conflict mapping, read, update,
  delete with built-in and referenced guards, is_referenced, and label
  normalization/casefold keying.
- DocumentTypeRegistry and PersonRoleRegistry declare only the model, error
  class, noun, short noun, retainer phrase, and reference columns.
- DocumentService and PeopleService keep their public method names and
  delegate. Every user-facing message, error category, and suggestion string
  is reproduced verbatim; only the noun is templated.
- Deleted _normalize_registry_label, _document_type_label_key,
  _normalize_role_label, _person_role_label_key,
  _document_type_is_referenced, and _person_role_is_referenced.

[MED-12] Shared not-found lookup
- ServiceBase._get_or_raise(model, id, *, session, error, noun, suggestion,
  options) loads by primary key or raises the caller's error type.
- documents.py: local _get_document_or_raise deleted; replaced by _read_document
  and adopted at read_document, delete_document, and set_document_type, which
  previously bypassed the helper and hand-wrote the raise.
- sources.py: 8 identical Source raises and 1 Job raise collapsed into
  _read_source / _get_or_raise.
- jobs.py and people.py already funneled through local _not_found builders and
  were left alone.

[MED-13][MED-01] Single media writer
- New services/media_storage.py owns validate -> name -> mkdir -> write ->
  wrap OSError. The write runs in asyncio.to_thread, so uploads no longer block
  the event loop.
- store_source_file, store_person_portrait, and store_homepage_image now share
  it and are async. Callers in store.py, people_page.py, and home_page.py await
  them. mkdir failures are now also translated to a domain error instead of
  escaping as a raw OSError.
- homepage_store gains HomepageStorageError so its write reports like the others.

[MED-14, partial] Service independence
- New services/source_media.py owns SOURCE_MIME_TYPES, SOURCE_EXTENSIONS,
  lookup_source_mime_type, and supported_source_formats.
- documents.py no longer imports services/sources.py. Its print projection uses
  the non-raising lookup and raises DocumentError, so DocumentService no longer
  emits a TranscriptionError.
- api/v4_print.py imports the mapping from the policy module.
- store.py and workflows.py still import sources.py; both are orchestration
  modules, which services.instructions.md:75-77 explicitly permits.
- Splitting SourceService itself remains deferred to V4.7.

[LOW-08] Query shape
- list_sources_detail filters job_id with a JOIN on JobSource instead of
  loading every Source and filtering in Python.
- read_source_navigation replaces the full ordered-id scan and .index() with
  two row-value comparisons bounded by LIMIT 1.
- list_processing_artifacts gains the limit parameter its summary sibling
  already had.
- build_evidence_export runs artifact integrity hashing and file reads through
  asyncio.to_thread.

Tests
- tests/test_service_boundaries.py: AST guard asserting no service module
  imports a sibling service module, plus a guard that the scan is non-empty.
- tests/services/test_transcription_service.py: asserts the job_id filter emits
  a JOIN, and that navigation emits exactly two LIMIT queries.
- tests/services/test_store.py: the two storage tests are now async.

Verified: 276 passed, 4 skipped; ruff check clean.
2026-08-17 16:46:15 -05:00
2026-08-10 12:34:36 -05:00
2026-08-17 15:25:52 -05:00
2026-08-11 16:42:09 -05:00
2026-06-26 19:17:18 -05:00
2026-06-22 17:32:18 -05:00
2026-06-26 19:17:18 -05:00
2026-06-26 19:17:33 -05:00

Transcription

Historical document transcription system for family-history documents.

The app lets you upload a document image/PDF, queues a background transcription job, and then shows job status and results in a web UI.

What the app does

  • Upload document files (.jpg, .jpeg, .png, .tif, .tiff, .pdf)
  • Persist document + job records in SQLite
  • Process jobs in a background worker (queued -> processing -> transcribed/failed)
  • Store transcript text (or failure detail)
  • Show status and results in the NiceGUI interface

Quick start

1) Install dependencies

uv sync

2) Configure environment

Create a .env file in the project root with the required OpenRouter API key:

OPENROUTER_API_KEY=your_openrouter_api_key

Settings are read from CLI arguments first, then environment variables, then .env, then the defaults below.

Configuration Source Precedence

When the same setting is provided in multiple places, the value is chosen in this order (highest priority first):

  1. CLI arguments (for example --port 8000)
  2. Settings constructor arguments (used mainly in tests)
  3. Environment variables
  4. .env file values
  5. Model defaults in code

Practical examples:

  • --port 8000 overrides both PORT=8000 in the shell and PORT=7000 in .env.
  • DATABASE__PATH=prod.db in the shell overrides DATABASE__PATH=dev.db in .env.

Server and runtime

Environment variable Default Description
HOST 0.0.0.0 Address on which the server listens.
PORT 8000 Server port.
LOG_LEVEL info Uvicorn and application log level.
RELOAD false Restart the development server when source files change.
ENVIRONMENT development Runtime environment: development, test, or production.

Provider

Environment variable Default Description
PROVIDER openrouter Transcription provider.
OPENROUTER_API_KEY Required OpenRouter API key.
PROVIDER_MODEL Provider default Optional model override.
OPENROUTER_HTTP_REFERER Unset Optional OpenRouter attribution URL.
OPENROUTER_APP_TITLE Unset Optional OpenRouter attribution title.

Database and files

Use nested env vars for database settings (recommended):

DATABASE__DRIVER=sqlite
DATABASE__PATH=app.db
# BOOTSTRAP_SCHEMA_ON_STARTUP=true
SQLITE_CHECK_SAME_THREAD=false
UPLOAD_DIR=./uploads
PROMPT_DIR=./prompts
DEFAULT_PROMPT_NAME=transcribe_document.md
# TRANSCRIPTION_TEMPERATURE=0.2  # range: 0.0-2.0
# TRANSCRIPTION_TOP_P=0.9       # range: 0.0-1.0

For PostgreSQL:

DATABASE__DRIVER=postgres
DATABASE__HOST=localhost
DATABASE__PORT=5432
DATABASE__DATABASE=transcription
DATABASE__USER=postgres
DATABASE__PASSWORD=change-me

This uses Pydantic nested settings (env_nested_delimiter='__') and avoids JSON blobs in .env. A top-level DATABASE={...} JSON value is still supported as a fallback, and nested keys such as DATABASE__PATH take precedence over conflicting JSON keys.

BOOTSTRAP_SCHEMA_ON_STARTUP creates missing tables when the app starts. When unset, it is enabled in development and test, and disabled in production; set it explicitly to override that policy. SQLITE_CHECK_SAME_THREAD defaults to false.

Worker

WORKER_MAX_RETRIES=0
WORKER_RETRY_BACKOFF_SECONDS=0
WORKER_PROVIDER_TIMEOUT_SECONDS=20
WORKER_MIN_TRANSCRIPTION_CHARS=0
WORKER_MIN_TRANSCRIPTION_LINES=0
WORKER_FAIL_ON_FINISH_REASON_LENGTH=false

3) Run the app

uv run python -m transcription --port 8000 --reload --database.driver sqlite --bootstrap-schema-on-startup

This starts the development server with SQLite, creates missing tables, and enables automatic reload. Run uv run python -m transcription --help for all CLI options; CLI names use kebab case and nested database options use dot notation, such as --database.path ./data/transcription.db.

4) Open in browser

Replace localhost with the server's hostname or IP address when connecting from another machine.

How to navigate the GUI

  • Upload page (/ui)

    • Select a supported file to upload.
    • The app creates a queued transcription job.
    • Use the View jobs link to inspect progress.
  • Jobs page (/ui/jobs)

    • See all jobs and their status.
    • Use Refresh to reload current states.
    • Open a specific job to see details.
  • Job detail page (/ui/jobs/{job_id})

    • Shows job metadata and status.
    • Displays transcript text when successful.
    • Displays failure detail when transcription fails.

Prompt artifacts

Prompt files are stored directly in PROMPT_DIR (default: ./prompts). DEFAULT_PROMPT_NAME must be a filename, not a path. Each job snapshots the validated prompt text, SHA-256 hash, and sampling values for reproducibility.

The canonical MVP prompt is:

  • prompts/transcribe_document.md

Destructive test procedure (with data backup)

AI execution policy: before the first unit-test run in a test/fix cycle, create one backup of ./data. Reuse that same backup for every subsequent test run in the cycle. After tests succeed, always pause and ask whether to restore now.

Use the cross-platform Python wrapper below whenever an AI agent runs tests against this repository.

  1. Create one backup of ./data and mark it as the active test-cycle backup.
  2. Run your test command.
  3. On failure, fix the errors and run the wrapper again; it reuses the active backup and never backs up post-test data.
  4. On success, always prompt whether to restore now (do not auto-restore unless explicitly approved).
  5. Close the cycle only by restoring the active backup or explicitly accepting the current data.

Preflight behavior:

  • Backup preflight is warning-only when data/transcription.db appears in use.
  • Restore preflight is blocking: the script prompts you to close conflicting applications, then type retry to re-check or cancel to skip restore.

Run with confirmation-gated restore (default)

uv run python tools/run_destructive_tests.py -- pytest tests/services/test_job_service.py tests/ui/test_jobs_page.py

After tests pass, the script asks whether to restore backup immediately.

This is the required default mode for AI-assisted test runs because it gives time to verify and accept code changes before any restoration happens.

Run with automatic restore (non-interactive)

uv run python tools/run_destructive_tests.py --auto-restore -- pytest

Run without terminal prompt (decide restore later)

uv run python tools/run_destructive_tests.py --skip-restore-prompt -- pytest

This keeps both the current post-test state and the backup, so restore can be decided explicitly later.

Repeated wrapper invocations reuse the backup recorded in .test-backups/.active-backup. If that backup is missing, the wrapper stops rather than creating a replacement from potentially destructive post-test data.

Restore later from a saved backup

uv run python tools/run_destructive_tests.py --restore-from data-backup-YYYYMMDD-HHMMSS

To keep the current data and close the active cycle without restoring:

uv run python tools/run_destructive_tests.py --accept-current-data

Backups are stored in .test-backups/ and ignored by git.

S
Description
A project to transcribe several thousand pages of family history documents
Readme
25 MiB
Languages
Python 98.4%
CSS 0.9%
Shell 0.6%
Dockerfile 0.1%