generated from john/python-template
Compare commits
142
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
be3d0b7ee0 | ||
|
|
7eca9fe7dc | ||
|
|
e54c2d9f26 | ||
|
|
065acad125 | ||
|
|
eeb1888aa3 | ||
|
|
321c454a4f | ||
|
|
6f4accf275 | ||
|
|
27ec81ca5b | ||
|
|
4410d23f5c | ||
|
|
d5e798825d | ||
|
|
02dca888d2 | ||
|
|
a896a11d2e | ||
|
|
0b48c80d87 | ||
|
|
70f8d6182e | ||
|
|
fc4288ff88 | ||
|
|
e5ef4d4422 | ||
|
|
88cef169c4 | ||
|
|
15a4814e23 | ||
|
|
16391463d6 | ||
|
|
96bb80d91f | ||
|
|
494f378e48 | ||
|
|
75dc946123 | ||
|
|
929f8d6de9 | ||
|
|
72909c1fd5 | ||
|
|
a5ec9f40a8 | ||
|
|
0d5206fe97 | ||
|
|
32b6b5f29a | ||
|
|
9b6bb9ae66 | ||
|
|
4c877dd6a2 | ||
|
|
9990583345 | ||
|
|
daa1642933 | ||
|
|
d198cc0c68 | ||
|
|
34c7b16675 | ||
|
|
ca29bc8b74 | ||
|
|
89ed83239e | ||
|
|
2c6ef46f5f | ||
|
|
c2ed98c16d | ||
|
|
c85cc6be20 | ||
|
|
234476ba6d | ||
|
|
faa30fd27b | ||
|
|
867cc9eb78 | ||
|
|
0e43094b80 | ||
|
|
c05c054b94 | ||
|
|
c0112c2714 | ||
|
|
96af7b2dc0 | ||
|
|
9a7970c533 | ||
|
|
f15c9834e4 | ||
|
|
2c59cbd2c7 | ||
|
|
5ff66a8c40 | ||
|
|
626b5d4b10 | ||
|
|
12f125761a | ||
|
|
803237371e | ||
|
|
d4ae97c1b1 | ||
|
|
c52d41ec33 | ||
|
|
6c3eac0a44 | ||
|
|
e5410708e4 | ||
|
|
efbae26f16 | ||
|
|
3873810022 | ||
|
|
2093eb6fb3 | ||
|
|
f9261a1af3 | ||
|
|
86cdb4035c | ||
|
|
f193b2800b | ||
|
|
736d0c06f4 | ||
|
|
26f9c83f54 | ||
|
|
a2bb1acd6b | ||
|
|
2a56365847 | ||
|
|
4aaa9bd581 | ||
|
|
de18c2e9da | ||
|
|
8d3c60fce1 | ||
|
|
8a30231adf | ||
|
|
67feeb28af | ||
|
|
c6ed3126e0 | ||
|
|
5566f48fc0 | ||
|
|
ed6998d8da | ||
|
|
ebf659b26c | ||
|
|
ae3483ec2e | ||
|
|
141ee1fa85 | ||
|
|
0f30d902b9 | ||
|
|
86b8e83ff4 | ||
|
|
efe7785392 | ||
|
|
94db493756 | ||
|
|
0d554c0648 | ||
|
|
63c21d4a14 | ||
|
|
cf49c3c127 | ||
|
|
bf2f3ac09c | ||
|
|
eaeb0bc806 | ||
|
|
8b08478c9d | ||
|
|
450d33d507 | ||
|
|
796216087c | ||
|
|
afd1dba4d4 | ||
|
|
7daa0b9808 | ||
|
|
7c4300f9c2 | ||
|
|
443a1e29c8 | ||
|
|
cdd846fe29 | ||
|
|
30fcef3892 | ||
|
|
de8cdb6e1a | ||
|
|
b6a5a89a84 | ||
|
|
c261fbb3bd | ||
|
|
5404224079 | ||
|
|
2c26177d0c | ||
|
|
edcfba9cb2 | ||
|
|
fca959fa5d | ||
|
|
110f40a28b | ||
|
|
7dd0d2c9bf | ||
|
|
11097b9cfe | ||
|
|
7285a87dfb | ||
|
|
f86c0ff27b | ||
|
|
246d7f9434 | ||
|
|
22d47574f2 | ||
|
|
4488280097 | ||
|
|
012dc15042 | ||
|
|
66e2dce465 | ||
|
|
597be2691c | ||
|
|
4e8c562f92 | ||
|
|
0b63b53f53 | ||
|
|
6a3ee26733 | ||
|
|
97b3d0fd62 | ||
|
|
7b9715b3f1 | ||
|
|
3e418a0889 | ||
|
|
2ccea77520 | ||
|
|
b3d8eb6e97 | ||
|
|
1ee9ebbffc | ||
|
|
aec89b3a7a | ||
|
|
d1321fd709 | ||
|
|
7054cd8af9 | ||
|
|
bdb1b31b0a | ||
|
|
5b97c759fe | ||
|
|
7db4df1729 | ||
|
|
63373bf24d | ||
|
|
936af9b1d3 | ||
|
|
a78b58ff40 | ||
|
|
aed827babe | ||
|
|
178347e086 | ||
|
|
c9f5dca064 | ||
|
|
6bd4cbb0a7 | ||
|
|
28811d79ce | ||
|
|
171132919d | ||
|
|
89cf69f8a2 | ||
|
|
1e8d8572d4 | ||
|
|
888a8c380a | ||
|
|
b8be27f0c9 | ||
|
|
8d5aec4301 |
@@ -1,54 +0,0 @@
|
||||
# --- NiceGUI Server ---
|
||||
# HOST=`0.0.0.0` (default)
|
||||
# PORT=8000 (default)
|
||||
# LOG_LEVEL: [`critical`, `error`, `warning`, `info` (default), `debug`, `trace`]
|
||||
# RELOAD=false (default)
|
||||
|
||||
# --- AI provider ---
|
||||
# PROVIDER=[`openrouter`(default), `google_genai`]
|
||||
PROVIDER=openrouter
|
||||
# OPENROUTER_API_KEY - Required when `PROVIDER=openrouter`
|
||||
OPENROUTER_API_KEY=your-api-key-goes-here
|
||||
# GEMINI_API_KEY - Required when `PROVIDER=google_genai`
|
||||
# PROVIDER_MODEL= specify model. If left blank OpenRouter will supply default.
|
||||
PROVIDER_MODEL=google/gemini-2.5-flash
|
||||
# OPENROUTER_HTTP_REFERER=https://example.com
|
||||
# OPENROUTER_APP_TITLE="Google: Gemini 2.5 Flash (openrouter)"
|
||||
|
||||
# --- runtime environment ---
|
||||
# ENVIRONMENT: [`development`(default), `test`, `production`]
|
||||
|
||||
# --- persistence ---
|
||||
# Use nested settings with double underscore because env_nested_delimiter="__".
|
||||
# SQLite example:
|
||||
# DATABASE__DRIVER=sqlite
|
||||
# DATABASE__PATH=app.db
|
||||
#
|
||||
# SQLite with custom relative path:
|
||||
# DATABASE__DRIVER=sqlite
|
||||
DATABASE__PATH=./data/transcription.db
|
||||
#
|
||||
# Postgres example:
|
||||
# DATABASE__DRIVER=postgres
|
||||
# DATABASE__HOST=localhost
|
||||
# DATABASE__PORT=5432
|
||||
# DATABASE__DATABASE=transcription
|
||||
# DATABASE__USER=postgres
|
||||
# DATABASE__PASSWORD=change-me
|
||||
#
|
||||
# Optional persistence flags:
|
||||
# BOOTSTRAP_SCHEMA_ON_STARTUP=false
|
||||
# SQLITE_CHECK_SAME_THREAD=false
|
||||
|
||||
# --- filesystem paths ---
|
||||
UPLOAD_DIR="./data"
|
||||
PROMPT_DIR="./prompts"
|
||||
|
||||
# --- worker reliability ---
|
||||
WORKER_MAX_RETRIES=0
|
||||
WORKER_RETRY_BACKOFF_SECONDS=0
|
||||
# WORKER_PROVIDER_TIMEOUT_SECONDS=[0-20]
|
||||
WORKER_PROVIDER_TIMEOUT_SECONDS=20
|
||||
WORKER_MIN_TRANSCRIPTION_CHARS=0
|
||||
WORKER_MIN_TRANSCRIPTION_LINES=0
|
||||
WORKER_FAIL_ON_FINISH_REASON_LENGTH=false
|
||||
@@ -0,0 +1,67 @@
|
||||
# Production environment example for docker-compose.production.yml
|
||||
|
||||
# --- NiceGUI Server ---
|
||||
HOST=0.0.0.0
|
||||
PORT=8000
|
||||
LOG_LEVEL=info
|
||||
RELOAD=false
|
||||
ENVIRONMENT=production
|
||||
# TRANSCRIPTION_COMMIT=
|
||||
RUN_EMBEDDED_WORKER=false
|
||||
LOG_DIR=/app/data/logs
|
||||
LOG_FILE_NAME=transcription.log
|
||||
LOG_FILE_MAX_BYTES=10485760
|
||||
LOG_FILE_BACKUP_COUNT=5
|
||||
|
||||
# --- AI provider ---
|
||||
PROVIDER=openrouter
|
||||
OPENROUTER_API_KEY=replace-with-real-key
|
||||
PROVIDER_MODEL=google/gemini-2.5-flash
|
||||
# PROVIDER_MODELS=["google/gemini-2.5-flash","anthropic/claude-sonnet-4"]
|
||||
# OPENROUTER_HTTP_REFERER=
|
||||
# OPENROUTER_APP_TITLE=
|
||||
DEFAULT_PROMPT_NAME=transcribe_document.md
|
||||
# TRANSCRIPTION_TEMPERATURE=
|
||||
# TRANSCRIPTION_TOP_P=
|
||||
|
||||
# --- persistence ---
|
||||
# Common database settings:
|
||||
DATABASE__DRIVER=postgres
|
||||
DATABASE__DATABASE=transcription
|
||||
DATABASE__USER=transcription
|
||||
DATABASE__PASSWORD=replace-with-strong-password
|
||||
BOOTSTRAP_SCHEMA_ON_STARTUP=false
|
||||
|
||||
# SQLite-specific settings:
|
||||
# DATABASE__PATH=./data/transcription.db
|
||||
# SQLITE_CHECK_SAME_THREAD=false
|
||||
|
||||
# Postgres-specific settings:
|
||||
DATABASE__HOST=postgres
|
||||
DATABASE__PORT=5432
|
||||
|
||||
# --- filesystem paths ---
|
||||
UPLOAD_DIR=/app/uploads
|
||||
PROMPT_DIR=/app/prompts
|
||||
# --- backup workflow helpers (not Runtime Settings model fields) ---
|
||||
BACKUP_DIR=/backup
|
||||
BACKUP_RETENTION_DAYS=14
|
||||
|
||||
# --- worker reliability ---
|
||||
WORKER_MAX_RETRIES=0
|
||||
WORKER_PROVIDER_TIMEOUT_SECONDS=30.0
|
||||
WORKER_STALE_JOB_SECONDS=90.0
|
||||
WORKER_RETRY_BACKOFF_SECONDS=1.0
|
||||
WORKER_SHUTDOWN_GRACE_SECONDS=5.0
|
||||
WORKER_POLL_INTERVAL_SECONDS=1.0
|
||||
WORKER_MIN_TRANSCRIPTION_CHARS=0
|
||||
WORKER_MIN_TRANSCRIPTION_LINES=0
|
||||
WORKER_FAIL_ON_FINISH_REASON_LENGTH=false
|
||||
|
||||
# --- cloudflare tunnel ---
|
||||
# Required for token-based tunnel startup.
|
||||
CLOUDFLARE_TUNNEL_TOKEN=replace-with-cloudflare-tunnel-token
|
||||
|
||||
# --- deployment wiring helpers ---
|
||||
# Runtime Settings writes target this file path inside the app container.
|
||||
RUNTIME_SETTINGS_ENV_FILE=/app/.env.production
|
||||
@@ -0,0 +1,67 @@
|
||||
# Production environment
|
||||
|
||||
# --- NiceGUI Server ---
|
||||
HOST=0.0.0.0
|
||||
PORT=8000
|
||||
LOG_LEVEL=info
|
||||
RELOAD=false
|
||||
ENVIRONMENT=production
|
||||
# TRANSCRIPTION_COMMIT=
|
||||
RUN_EMBEDDED_WORKER=false
|
||||
LOG_DIR=/app/data/logs
|
||||
LOG_FILE_NAME=transcription.log
|
||||
LOG_FILE_MAX_BYTES=10485760
|
||||
LOG_FILE_BACKUP_COUNT=5
|
||||
|
||||
# --- AI provider ---
|
||||
PROVIDER=openrouter
|
||||
OPENROUTER_API_KEY=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
|
||||
PROVIDER_MODEL=google/gemini-2.5-flash
|
||||
PROVIDER_MODELS=["google/gemini-2.5-pro","google/gemini-2.5-flash","anthropic/claude-opus-5","anthropic/claude-sonnet-4","openai/gpt-5.6","openai/gpt-4o"]
|
||||
# OPENROUTER_HTTP_REFERER=
|
||||
# OPENROUTER_APP_TITLE=
|
||||
DEFAULT_PROMPT_NAME=transcribe_document.md
|
||||
# TRANSCRIPTION_TEMPERATURE=
|
||||
# TRANSCRIPTION_TOP_P=
|
||||
|
||||
# --- persistence ---
|
||||
# Common database settings:
|
||||
DATABASE__DRIVER=postgres
|
||||
DATABASE__DATABASE=transcription
|
||||
DATABASE__USER=transcription
|
||||
DATABASE__PASSWORD=<password>
|
||||
BOOTSTRAP_SCHEMA_ON_STARTUP=false
|
||||
|
||||
# SQLite-specific settings:
|
||||
# DATABASE__PATH=./data/transcription.db
|
||||
# SQLITE_CHECK_SAME_THREAD=false
|
||||
|
||||
# Postgres-specific settings:
|
||||
DATABASE__HOST=postgres
|
||||
DATABASE__PORT=5432
|
||||
|
||||
# --- filesystem paths ---
|
||||
UPLOAD_DIR=/app/uploads
|
||||
PROMPT_DIR=/app/prompts
|
||||
# --- backup workflow helpers (not Runtime Settings model fields) ---
|
||||
BACKUP_DIR=/backup
|
||||
BACKUP_RETENTION_DAYS=14
|
||||
|
||||
# --- worker reliability ---
|
||||
WORKER_MAX_RETRIES=0
|
||||
WORKER_PROVIDER_TIMEOUT_SECONDS=30.0
|
||||
WORKER_STALE_JOB_SECONDS=30.0
|
||||
WORKER_RETRY_BACKOFF_SECONDS=1.0
|
||||
WORKER_SHUTDOWN_GRACE_SECONDS=5.0
|
||||
WORKER_POLL_INTERVAL_SECONDS=1.0
|
||||
WORKER_MIN_TRANSCRIPTION_CHARS=0
|
||||
WORKER_MIN_TRANSCRIPTION_LINES=0
|
||||
WORKER_FAIL_ON_FINISH_REASON_LENGTH=false
|
||||
|
||||
# --- cloudflare tunnel ---
|
||||
# Required for token-based tunnel startup.
|
||||
CLOUDFLARE_TUNNEL_TOKEN=CLOUDFLARE_TUNNEL_TOKEN=eyJhIjoiYTRhNjM0NzNhNzBiZjhhYmY3OWUyNjE4ZTcyNjgwZmMiLCJ0IjoiZWY1MjFkNWItYzY1ZS00ZGFmLTlmYTMtMzQyOGYzMGUyMDY4IiwicyI6IlpHWm1NVGRsTVRVdFpEYzNaaTAwWkRJeUxXRmhPRFV0TmpKallXRmhPRFJrWXpSaSJ9
|
||||
|
||||
# --- deployment wiring helpers ---
|
||||
# Runtime Settings writes target this file path inside the app container.
|
||||
RUNTIME_SETTINGS_ENV_FILE=/app/.env.production
|
||||
@@ -0,0 +1,67 @@
|
||||
# Production environment
|
||||
|
||||
# --- NiceGUI Server ---
|
||||
HOST=0.0.0.0
|
||||
PORT=8000
|
||||
LOG_LEVEL=info
|
||||
RELOAD=false
|
||||
ENVIRONMENT=production
|
||||
# TRANSCRIPTION_COMMIT=
|
||||
RUN_EMBEDDED_WORKER=false
|
||||
LOG_DIR=/app/data/logs
|
||||
LOG_FILE_NAME=transcription.log
|
||||
LOG_FILE_MAX_BYTES=10485760
|
||||
LOG_FILE_BACKUP_COUNT=5
|
||||
|
||||
# --- AI provider ---
|
||||
PROVIDER=openrouter
|
||||
OPENROUTER_API_KEY=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
|
||||
PROVIDER_MODEL=google/gemini-2.5-flash
|
||||
PROVIDER_MODELS=["google/gemini-2.5-pro","google/gemini-2.5-flash","anthropic/claude-opus-5","anthropic/claude-sonnet-4","openai/gpt-5.6","openai/gpt-4o"]
|
||||
# OPENROUTER_HTTP_REFERER=
|
||||
# OPENROUTER_APP_TITLE=
|
||||
DEFAULT_PROMPT_NAME=transcribe_document.md
|
||||
# TRANSCRIPTION_TEMPERATURE=
|
||||
# TRANSCRIPTION_TOP_P=
|
||||
|
||||
# --- persistence ---
|
||||
# Common database settings:
|
||||
DATABASE__DRIVER=sqlite
|
||||
# DATABASE__DATABASE=transcription
|
||||
# DATABASE__USER=transcription-local
|
||||
# DATABASE__PASSWORD=My!3sons
|
||||
# BOOTSTRAP_SCHEMA_ON_STARTUP=false
|
||||
|
||||
# SQLite-specific settings:
|
||||
DATABASE__PATH=./data-local/transcription-local.db
|
||||
SQLITE_CHECK_SAME_THREAD=false
|
||||
|
||||
# Postgres-specific settings:
|
||||
# DATABASE__HOST=postgres
|
||||
# DATABASE__PORT=5432
|
||||
|
||||
# --- filesystem paths ---
|
||||
UPLOAD_DIR=./data-local
|
||||
PROMPT_DIR=/data/prompts
|
||||
# --- backup workflow helpers (not Runtime Settings model fields) ---
|
||||
BACKUP_DIR=/backup
|
||||
BACKUP_RETENTION_DAYS=14
|
||||
|
||||
# --- worker reliability ---
|
||||
WORKER_MAX_RETRIES=0
|
||||
WORKER_PROVIDER_TIMEOUT_SECONDS=30.0
|
||||
WORKER_STALE_JOB_SECONDS=30.0
|
||||
WORKER_RETRY_BACKOFF_SECONDS=1.0
|
||||
WORKER_SHUTDOWN_GRACE_SECONDS=5.0
|
||||
WORKER_POLL_INTERVAL_SECONDS=1.0
|
||||
WORKER_MIN_TRANSCRIPTION_CHARS=0
|
||||
WORKER_MIN_TRANSCRIPTION_LINES=0
|
||||
WORKER_FAIL_ON_FINISH_REASON_LENGTH=false
|
||||
|
||||
# --- cloudflare tunnel ---
|
||||
# Required for token-based tunnel startup.
|
||||
CLOUDFLARE_TUNNEL_TOKEN=CLOUDFLARE_TUNNEL_TOKEN=eyJhIjoiYTRhNjM0NzNhNzBiZjhhYmY3OWUyNjE4ZTcyNjgwZmMiLCJ0IjoiZWY1MjFkNWItYzY1ZS00ZGFmLTlmYTMtMzQyOGYzMGUyMDY4IiwicyI6IlpHWm1NVGRsTVRVdFpEYzNaaTAwWkRJeUxXRmhPRFV0TmpKallXRmhPRFJrWXpSaSJ9
|
||||
|
||||
# --- deployment wiring helpers ---
|
||||
# Runtime Settings writes target this file path inside the app container.
|
||||
RUNTIME_SETTINGS_ENV_FILE=/app/.env.production
|
||||
@@ -0,0 +1,23 @@
|
||||
---
|
||||
name: Python Architect Reviewer
|
||||
description: Evidence-based senior architect reviewer for FastAPI, NiceGUI, and SQLModel codebases.
|
||||
skills:
|
||||
- python-code-reviewer
|
||||
---
|
||||
|
||||
# Python Architect Reviewer
|
||||
|
||||
You are a Senior Python Architect performing an evidence-based, read-only code review.
|
||||
|
||||
> No `tools:` allowlist is declared here on purpose. Tool identifiers differ between the runtimes
|
||||
> this agent is invoked from, so a hard-coded list silently under-tools the agent in one of them.
|
||||
> Read-only discipline is enforced by the **Read-Only Scope** rule below, not by the frontmatter.
|
||||
|
||||
## Operating Principles
|
||||
|
||||
- **Stack Context:** Python 3.12+, FastAPI, NiceGUI, SQLModel, SQLAlchemy (SQLite/PostgreSQL), Pydantic V2, asyncio workers, and OpenRouter adapters.
|
||||
- **Evidence-Based:** Always inspect real files. Every finding must reference concrete file paths and line numbers (e.g., `app/services/worker.py:45-78`). Do not speculate.
|
||||
- **Tool Verification:** This is a `uv` project; the toolchain is not on `PATH`. Verify with `uv run ruff check .`, `uv run ty check`, and `uv run pytest -q -m "not external"`, and record the exact commands and outcomes. Never report a lint, type, or test claim you did not run.
|
||||
- **Verify Recommendations, Not Just Findings:** Before recommending a change to a shared symbol, enumerate its consumers and confirm the fix is safe for each. See the skill's consumer-tracing step and `Blast Radius` field.
|
||||
- **Skill Is Canonical:** The `python-code-reviewer` skill defines the review workflow, deterministic checks, severity and reachability rubrics, report location, and report template. Follow it exactly. Where this file and the skill disagree, the skill wins — do not restate its specifics here.
|
||||
- **Read-Only Scope:** Do not modify source, tests, docs, instructions, or configuration. The review report is the only artifact you produce.
|
||||
@@ -0,0 +1,41 @@
|
||||
---
|
||||
description: Require documentation updates whenever code changes alter contracts, behavior, or scope.
|
||||
applyTo: 'src/transcription/**/*.py'
|
||||
---
|
||||
|
||||
# Documentation Sync Requirements
|
||||
|
||||
Keep docs in sync in the same change whenever implementation alters a documented contract, behavior, or roadmap decision.
|
||||
|
||||
Documentation targets below always refer to the **current** baseline. `docs/index.md` states which
|
||||
baseline that is; resolve any version-specific document from there. Never cite a superseded version
|
||||
tree by name in this file or in the docs you update — retired revision trees are not authority, and
|
||||
`tests/test_meta_contract_guards.py` fails active contract files that route authority through them.
|
||||
|
||||
## Update documentation when any of these change
|
||||
|
||||
1. **Schema/Data contract**
|
||||
- Models, fields, enums, constraints, indexes, relationships, loading semantics.
|
||||
- **Required doc update:** `docs/schema.md`.
|
||||
|
||||
2. **Configuration contract**
|
||||
- `Settings` keys, defaults, required/optional environment values.
|
||||
- **Required doc update:** `.env.production.example` and any directly related setup docs.
|
||||
|
||||
3. **User-visible UI behavior**
|
||||
- Page flow, routes, button/action behavior, labels, status wording, empty/error states.
|
||||
- **Required doc update:** relevant `docs/ui/pages/*.md` docs and feature docs when applicable.
|
||||
|
||||
4. **Error handling semantics**
|
||||
- Error categories, retry behavior, envelope structure, translation boundaries.
|
||||
- **Required doc update:** `docs/error_handling.md` and `docs/invariant/error_handling.md`.
|
||||
|
||||
5. **Roadmap/scope decisions**
|
||||
- Version targets, sequencing, deferrals, and accepted alternatives.
|
||||
- **Required doc update:** `docs/roadmap_plan.md`, plus any backlog or feature document for the
|
||||
current baseline. Locate it through `docs/index.md` rather than assuming a version-named path.
|
||||
|
||||
## Working rule
|
||||
|
||||
If none of the categories above changed, documentation edits are optional.
|
||||
If any category changed, update docs in the same PR/change set rather than deferring.
|
||||
@@ -0,0 +1,121 @@
|
||||
---
|
||||
description: Cross-cutting error handling rules for services, API, and UI.
|
||||
applyTo: 'src/transcription/**/*.py'
|
||||
---
|
||||
|
||||
# Error Handling (Cross-cutting)
|
||||
|
||||
Primary references:
|
||||
|
||||
- `docs/error_handling.md`
|
||||
- `docs/invariant/error_handling.md`
|
||||
- `docs/requirements.md`
|
||||
|
||||
## Taxonomy and Categories
|
||||
|
||||
Use category-driven semantics aligned to canonical policy:
|
||||
|
||||
- `validation`
|
||||
- `not_found`
|
||||
- `conflict`
|
||||
- `external`
|
||||
- `timeout`
|
||||
- `internal`
|
||||
|
||||
Do not invent ad hoc categories in user/API-facing envelopes unless canonical docs are updated.
|
||||
|
||||
Runtime/internal categories may be more specific for diagnostics and persistence, but they must map
|
||||
deterministically to the canonical envelope categories through the centralized mapper in
|
||||
`transcription.errors.canonical_error_category`.
|
||||
|
||||
Current internal categories:
|
||||
|
||||
- `validation_error`
|
||||
- `user_input_error`
|
||||
- `not_found_error`
|
||||
- `conflict_error`
|
||||
- `external_provider_error`
|
||||
- `external_timeout_error`
|
||||
- `processing_error`
|
||||
- `infrastructure_transient_error`
|
||||
- `infrastructure_persistent_error`
|
||||
- `internal_unexpected_error`
|
||||
|
||||
Required internal -> canonical mapping:
|
||||
|
||||
- `validation_error`, `user_input_error` -> `validation`
|
||||
- `not_found_error` -> `not_found`
|
||||
- `conflict_error` -> `conflict`
|
||||
- `external_provider_error` -> `external`
|
||||
- `external_timeout_error`, `infrastructure_transient_error` -> `timeout`
|
||||
- `processing_error`, `infrastructure_persistent_error`, `internal_unexpected_error` -> `internal`
|
||||
|
||||
## Translation Boundaries
|
||||
|
||||
- **Provider/adapters:** raise provider/domain exceptions; do not emit UI text.
|
||||
- **Services:** map raw exceptions into internal categories and preserve causal chain (`raise ... from ...`).
|
||||
- **UI/API:** map internal category -> canonical envelope category and emit user-safe, actionable messages.
|
||||
|
||||
## Retry Rules
|
||||
|
||||
- No auto-retry for `validation`, `not_found`, `conflict`.
|
||||
- `external`/`timeout` may be retried when operation semantics are safe.
|
||||
- Preserve each retry as new evidence where applicable (no history rewrite).
|
||||
|
||||
## Job/Page Failure Semantics
|
||||
|
||||
- Page-level (`JobSource`): `pending`, `transcribed`, `failed`, `cancelled`.
|
||||
- Job terminals: `transcribed`, `partial_success`, `failed`.
|
||||
- Cancellation must keep job-level and page-level semantics explicit and consistent.
|
||||
- Do not emit legacy terminal state language such as `completed` in active user/API lifecycle contracts.
|
||||
|
||||
## User-Safe Messaging
|
||||
|
||||
- Never leak stack traces, credentials, auth headers, or local filesystem paths in user-facing output.
|
||||
- Include actionable remediation guidance aligned to category.
|
||||
- Keep envelope structure consistent across API endpoints.
|
||||
|
||||
### `AppError.message` vs `AppError.detail`
|
||||
|
||||
`AppError` carries two texts with different audiences, and they must not be collapsed. Getting this
|
||||
wrong has already caused a real defect in this repository, in both directions.
|
||||
|
||||
| Attribute | Audience | Reaches | Rule |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| `message` | User and API clients | `ErrorEnvelope.message`, UI notifications | Stays generic. Never embed exception text, provider payloads, or filesystem paths. |
|
||||
| `detail` | Internal only | Logs, and `format_error_detail` -> `ExecutionAttempt.error_detail` and `MaintenanceRun.error_detail` | Carries the root cause. Never rendered to users or serialized into an envelope without a sanitizing projection. |
|
||||
|
||||
- Putting root-cause data in `message` leaks infrastructure detail to users.
|
||||
- Omitting it from `detail` silently degrades the provenance record this system exists to preserve —
|
||||
a failed attempt whose `error_detail` says nothing is an attempt that cannot be diagnosed later.
|
||||
- Any render boundary that displays persisted `error_detail` must apply the same no-local-path rule
|
||||
as `message`: sanitize machine-local absolute paths before the text becomes user-visible.
|
||||
- When you raise from a caught exception, populate **both**: a generic `message` and a `detail`
|
||||
carrying `type(exc).__name__` and the exception text, with `raise ... from exc`.
|
||||
- `detail` is optional (`None`). A read path that assumes it is populated must handle its absence.
|
||||
- Before changing either attribute, or any helper that formats them, enumerate every consumer —
|
||||
evidence writes, maintenance runs, logging, API envelopes, and UI presentation all read these
|
||||
fields, and tests assert on the persisted text.
|
||||
|
||||
Canonical definitions live in `src/transcription/errors.py`; see also `docs/error_handling.md`.
|
||||
|
||||
## Logging and Diagnostics
|
||||
|
||||
- Log operation identifiers and error IDs where available.
|
||||
- Preserve category + cause-chain context.
|
||||
- Distinguish no-response timeout/network failures from returned provider error responses.
|
||||
|
||||
## Guardrails
|
||||
|
||||
- No broad catch-and-swallow patterns.
|
||||
- No success-shaped fallback values after exceptions.
|
||||
- Category mapping must remain deterministic and testable.
|
||||
|
||||
## Contract Sync Rule
|
||||
|
||||
If taxonomy, retries, or envelope semantics change:
|
||||
|
||||
1. Update canonical docs (`docs/error_handling.md`, and invariant docs if needed).
|
||||
2. Update tests in the same change.
|
||||
3. Update related instruction/skill references.
|
||||
4. If change affects persisted status/category fields, update `docs/schema.md` when applicable.
|
||||
@@ -0,0 +1,122 @@
|
||||
---
|
||||
description: Provider adapter rules for evidence capture, secret safety, and client lifecycle.
|
||||
applyTo: 'src/transcription/providers/**/*.py'
|
||||
---
|
||||
|
||||
# Provider Adapters
|
||||
|
||||
Primary references:
|
||||
|
||||
- `docs/invariant/ai_evidence_and_provenance.md` (canonical; provider adapters own provider-boundary evidence capture)
|
||||
- `docs/architecture.md`
|
||||
- `docs/schema.md`
|
||||
|
||||
The provider layer is where an external API becomes application data. It is also the only place
|
||||
that can capture what actually crossed the wire — once a response reaches a service, the evidence
|
||||
it did not preserve is gone permanently. Treat capture correctness as the primary job of this
|
||||
layer and text extraction as secondary.
|
||||
|
||||
## Layer Boundary
|
||||
|
||||
- Adapters may depend on `transcription.config`, `transcription.providers.*`, the HTTP client, and
|
||||
the provider SDK. They must not import `services`, `db`, `ui`, or `api`.
|
||||
- Provider specifics — headers, model slugs, payload shapes, SDK types, error classes — stop here.
|
||||
Callers receive `TranscriptionResult` and `ProviderError` subclasses only.
|
||||
- Adapters raise provider/domain exceptions. They must not emit user-facing text, notifications,
|
||||
or remediation wording; that translation belongs to services and UI. See
|
||||
[error-handling instructions](./error-handling.instructions.md).
|
||||
- Adapters do not persist. They return evidence; services decide what is written and when.
|
||||
|
||||
Enforced by `tests/test_provider_boundaries.py`.
|
||||
|
||||
## Contract Surface
|
||||
|
||||
- Every adapter satisfies the `TranscriptionProvider` protocol in `base.py`. Failed-call evidence is
|
||||
returned through the caller-owned `ProviderCallEvidence` sink passed to `transcribe()`, so
|
||||
evidence stays scoped to one invocation instead of living on mutable adapter instance state.
|
||||
- `TranscriptionResult`, `RequestManifest`, and `TransportEvidence` are `extra="forbid"` and frozen.
|
||||
Add a field to the contract rather than smuggling data through an untyped dict.
|
||||
- Evidence contracts in `evidence.py` are versioned (`schema_name` + `schema_version`). A change to
|
||||
the meaning or shape of a captured field requires a version bump, not a silent redefinition —
|
||||
stored evidence must keep its original meaning.
|
||||
|
||||
## Transport Evidence
|
||||
|
||||
The rules below implement `docs/invariant/ai_evidence_and_provenance.md` §3.4-3.5. That document
|
||||
wins if this file drifts from it.
|
||||
|
||||
- Capture the response body **at the HTTP boundary, before SDK parsing**, so fields the SDK does
|
||||
not model are not lost. `_CapturingAsyncClient` exists for this; do not replace it with a
|
||||
post-parse `model_dump()` and call the result transport evidence.
|
||||
- Reset per-call capture state at the start of every call. Without it, a connection failure can
|
||||
attach the *previous* call's response as evidence for this one. Guarded by
|
||||
`tests/test_evidence_provenance.py::test_openrouter_does_not_reuse_prior_response_on_connection_failure`.
|
||||
- Keep transport capture scoped to the call, not the adapter instance. Concurrent `transcribe()`
|
||||
calls on one adapter must not be able to overwrite each other's response evidence.
|
||||
- Handle the streamed-body case (`httpx.ResponseNotRead`) rather than assuming `response.content`
|
||||
is always available.
|
||||
- When no response arrives — timeout, DNS, connection reset — emit
|
||||
`TransportEvidence(response_received=False)`. Absence of a response is itself evidence and must
|
||||
be explicit, never an empty body or a missing record.
|
||||
- Preserve safe response evidence for **unsuccessful** calls too, whenever a response was received.
|
||||
- Never relabel an SDK snapshot or normalized metadata as transport evidence, and never backfill
|
||||
it into an execution that predates capture.
|
||||
|
||||
## Secret Safety
|
||||
|
||||
- Persist response headers only through `filter_safe_response_headers` and the
|
||||
`SAFE_RESPONSE_HEADERS` allowlist. Allowlist, never denylist: capture-then-redact is prohibited,
|
||||
because an unknown header is unsafe by default.
|
||||
- Adding a header to the allowlist is a deliberate evidence decision. Confirm it carries no
|
||||
credential, cookie, or session material, and state why it is needed for correlation, content
|
||||
interpretation, rate-limit diagnosis, or audit.
|
||||
- API keys, `Authorization`, and cookies must never appear in a manifest, evidence record, log
|
||||
line, or exception message.
|
||||
- The request manifest references source content by identity (digest, size, media type, page).
|
||||
Do not duplicate base64 source bytes into it — `_replace_embedded_media` exists for this.
|
||||
|
||||
## Execution Specification
|
||||
|
||||
The manifest must let a reader reconstruct what was asked, per invariant §3.3:
|
||||
|
||||
- Provider, requested model, full effective prompt text, and prompt digest.
|
||||
- Every explicitly supplied parameter, and — separately — which optional parameters were
|
||||
**omitted**. Omission is not the same as a null value or an assumed provider default; the
|
||||
`optional_parameter_states` distinction between `omitted`, `null`, and `value` is deliberate.
|
||||
- Timeout budget, retry policy, source reference, and `SoftwareContext` versions.
|
||||
- Manifest digests use `canonical_json_bytes`. Do not hash a plain `json.dumps()`; key order and
|
||||
separators must stay deterministic or digests become uncomparable.
|
||||
|
||||
## Client Lifecycle and Async Safety
|
||||
|
||||
- Reuse one pooled `AsyncClient` per adapter instance; do not construct a client per request.
|
||||
- Accept an injected client so tests can drive the adapter without network access.
|
||||
- Derive timeouts from `Settings` (`worker_provider_timeout_seconds`) rather than hard-coding, and
|
||||
keep the client timeout aligned with the configured budget so the SDK cannot expire first and
|
||||
hide the real failure.
|
||||
- Implement `aclose()` and release pooled resources. An adapter that creates a client owns closing
|
||||
it; one given a client must not close a caller-owned resource it did not create.
|
||||
- Never block the event loop. Offload CPU-bound work (hashing large payloads, image encoding) with
|
||||
`asyncio.to_thread`.
|
||||
- Propagate `asyncio.CancelledError` untouched — do not convert cancellation into a provider error.
|
||||
|
||||
## Failure Handling
|
||||
|
||||
- Raise `ProviderAuthError` for authentication, `ProviderResponseError` for malformed or unusable
|
||||
responses, and `ProviderError` otherwise.
|
||||
- Always attach `request_manifest`, `transport_evidence`, and an accurate `failure_phase` to raised
|
||||
errors. `failure_phase` must distinguish a received-but-failed response from a call that never
|
||||
reached the provider.
|
||||
- Validate responses with Pydantic rather than indexing into raw dicts.
|
||||
- Invalid *optional* metadata (for example unparsable token counts) must not discard an otherwise
|
||||
valid transcript. Degrade the metadata, not the result.
|
||||
|
||||
## Contract Sync Rule
|
||||
|
||||
If capture behavior, evidence schema, or the header allowlist changes:
|
||||
|
||||
1. Update `docs/invariant/ai_evidence_and_provenance.md` only if the durable preservation contract
|
||||
itself is changing — that revision is deliberate and reviewed, not incidental.
|
||||
2. Update `docs/schema.md` when persisted evidence fields change.
|
||||
3. Update or add tests in the same change (`tests/providers/`, `tests/test_evidence_provenance.py`).
|
||||
4. Bump the affected evidence `schema_version` when a field's meaning changes.
|
||||
@@ -7,29 +7,135 @@ applyTo: 'src/transcription/services/*.py'
|
||||
|
||||
## Structure
|
||||
|
||||
- Project core data models defined in [models](../../src/transcription/models.py)
|
||||
- 1 service class per data model
|
||||
- Only services directly interact with the database, and only through async methods
|
||||
- Services are completely independent of one another. Any operation that needs to use more than a single service, which is most of them, needs to have a separate orchestration function.
|
||||
- Project core data models are defined in [models](../../src/transcription/db/models.py)
|
||||
- One service class per **aggregate**, not per table. An aggregate is a root model plus
|
||||
the models that have no independent lifecycle of their own. `DocumentType` has no
|
||||
meaning without `Document`, so it belongs to `DocumentService`; it does not get its
|
||||
own service. Splitting per table produces services that must reach across each other
|
||||
for every real operation, which is what line 13 forbids.
|
||||
- Only services interact with the database, and only through async methods.
|
||||
- **A service module must not import another service module.** This is enforced by
|
||||
[test_service_boundaries](../../tests/test_service_boundaries.py). Shared types go in a
|
||||
neutral module that defines no service class (see [errors](../../src/transcription/services/errors.py)).
|
||||
- Not every module in this package is a service. Modules fall into three kinds:
|
||||
- **Aggregate services** own models and define a `*Service` class: `documents.py`, `sources.py`,
|
||||
`jobs.py`, `people.py`, `photos.py`, `maintenance.py`, and `evidence.py` (read/projection only,
|
||||
owns nothing).
|
||||
- **Orchestration modules** define no service class and compose writes across aggregates:
|
||||
`store.py`, `workflows.py`. They are the sanctioned place to create or delete rows owned by more
|
||||
than one service — see [Service Composition](#service-composition).
|
||||
- **Shared infrastructure and free-function helpers** are exempt from the service rules below:
|
||||
`base.py` (`ServiceBase`), `registry.py` (`RegistryService`, a generic base for lookup tables —
|
||||
not an aggregate owner itself), `unit_of_work.py`, `errors.py`, `normalization.py`, `prompts.py`,
|
||||
`quality.py`, `media_storage.py`, `source_media.py`. `__init__.py` exposes `ServiceBundle`.
|
||||
- Cross-cutting error behavior must follow
|
||||
[error-handling instructions](./error-handling.instructions.md).
|
||||
|
||||
## Model Ownership
|
||||
|
||||
Every model has exactly one owning service. The owner defines that model's invariants and
|
||||
is the only service that may **create or delete** its rows.
|
||||
|
||||
| Model | Owner |
|
||||
| --- | --- |
|
||||
| `Document`, `DocumentType`, `DocumentTag` | `DocumentService` |
|
||||
| `Source`, `JobSource` | `SourceService` |
|
||||
| `Job` | `JobService` |
|
||||
| `Person`, `PersonRole`, `DocumentPerson`, `PersonTag` | `PeopleService` |
|
||||
| `GenealogyPerson`, `GenealogyFamily`, `GenealogyFamilyChild`, `GenealogyCitation` | `MaintenanceService` |
|
||||
| `Photo` | `PhotosService` |
|
||||
| `MaintenanceRun` | `MaintenanceService` |
|
||||
| `ExecutionAttempt` | `SourceService` |
|
||||
| `Tag` | shared — see below |
|
||||
|
||||
Keep this table complete: every table in `src/transcription/db/models.py` appears exactly once,
|
||||
except `Tag`. When you add a model, add its owner here in the same change.
|
||||
|
||||
### `Tag` is deliberately shared
|
||||
|
||||
`Tag` is one table reached through two `RegistryService[Tag]` facades that differ only in the
|
||||
reference model they count usage through: `TagRegistry` (`documents.py`, via `DocumentTag`) and
|
||||
`PersonTagRegistry` (`people.py`, via `PersonTag`). Both create and delete `Tag` rows through the
|
||||
generic registry. This is the single sanctioned exception to one-owner-per-model — do not "fix" it by
|
||||
assigning `Tag` to one service, because the other facade would then be creating rows it does not own.
|
||||
Any change to `Tag` semantics, labels, or normalization must be validated against **both** facades
|
||||
and the junction table each one counts.
|
||||
|
||||
`DocumentTag` and `PersonTag` follow the junction rule below: each is created and deleted only by the
|
||||
service on its own side.
|
||||
|
||||
### Registries
|
||||
|
||||
`DocumentTypeRegistry`, `TagRegistry`, `PersonRoleRegistry`, and `PersonTagRegistry` are
|
||||
`RegistryService` subclasses, not independent services. A registry belongs to the aggregate service
|
||||
whose module declares it and shares that service's ownership. Registry CRUD uses `<operation>_entry`
|
||||
naming (see [CRUD Methods](#crud-methods)).
|
||||
|
||||
### Junction tables
|
||||
|
||||
A junction table is owned by the service that **creates and deletes its rows** — its
|
||||
lifecycle owner. The service on the other side may read through the junction (via
|
||||
`selectinload`) but must not create rows in it.
|
||||
|
||||
- `document_person` -> `PeopleService`. Every write is there; `DocumentService` only
|
||||
eager-loads through it.
|
||||
- `job_source` -> `SourceService`, which creates the row, records each page's outcome,
|
||||
and deletes it.
|
||||
|
||||
Two consequences follow, and both are deliberate:
|
||||
|
||||
- **Cascade deletion is not a violation.** A service deleting the aggregate root it owns
|
||||
may delete rows referencing that root which cannot outlive it
|
||||
(`JobService.delete_job_with_guardrails` deletes the job's `job_source` rows).
|
||||
- **Evidence deletion is an explicit workflow, not a runtime path.** `JobService.delete_job_and_evidence`
|
||||
deletes `ExecutionAttempt` rows owned by `SourceService`. That is sanctioned because it is the
|
||||
named retention workflow that `delete_job_with_guardrails` refuses to perform implicitly — that
|
||||
method *blocks* deletion when attempts exist. Append-only means runtime code never rewrites or
|
||||
removes history to represent a new outcome; it does not forbid a deliberate, operator-invoked
|
||||
retention operation. Do not add a second path that deletes attempts.
|
||||
- **Ownership governs creation and deletion, not every state transition.** `job_source` is
|
||||
both a link and the transcription work queue. `JobService.cancel_job` and
|
||||
`resubmit_failed_sources` transition `job_source.status` across a whole job, because that
|
||||
transition is a Job lifecycle event, not a per-page outcome. They create and delete
|
||||
nothing.
|
||||
|
||||
`EvidenceService` is read-focused and projection-focused. It may coordinate selection
|
||||
flows, but append-only attempt creation remains in `SourceService` write paths.
|
||||
|
||||
If a new operation cannot be expressed within one owner, it belongs in an orchestration
|
||||
module, not in a cross-service import.
|
||||
|
||||
## Error Handling
|
||||
|
||||
- Service-specific errors defined at the top of the respective module and inherit from `AppError`
|
||||
- Use a context manager for large `try/except` blocks like in [transcription](../../src/transcription/services/transcription.py)
|
||||
- Errors used by a single service are defined at the top of that module and inherit from `AppError`.
|
||||
- Errors shared by more than one service go in [errors](../../src/transcription/services/errors.py),
|
||||
which defines no service class and is therefore importable by any of them.
|
||||
- Use a context manager for large `try/except` blocks, like `handle_transcription_errors` in
|
||||
[sources](../../src/transcription/services/sources.py).
|
||||
- Category mapping, retry behavior, and translation boundaries are defined in
|
||||
[error-handling instructions](./error-handling.instructions.md).
|
||||
- Service-edge exception translation must be deterministic: map to canonical categories and preserve clear provider->service->API/UI boundaries.
|
||||
|
||||
## Checklist
|
||||
|
||||
- [ ] Uses `ServiceBase` for common logic
|
||||
- [ ] CRUD methods created at the top
|
||||
- [ ] Session kwarg for `AsyncSession` to pass in a session object to each method
|
||||
- [ ] Services use `self._session_scope` in their methods to pass the session thru.
|
||||
- Multiple operations on the same object(s) require sharing a session between all the methods used.
|
||||
- [ ] Session kwarg for `AsyncSession` to pass a session object into each method
|
||||
- [ ] Services use `self._session_scope` in their methods to pass the session through
|
||||
- Multiple operations on the same object(s) require sharing a session between all the methods used
|
||||
- [ ] Every model the module touches is either owned by it or reached read-only
|
||||
- [ ] Evidence writes preserve append-only semantics
|
||||
|
||||
## CRUD Methods
|
||||
|
||||
- Create, read, update, and delete, created in that order
|
||||
- Name format `<operation>_<model >`, for example `create_document` or `update_job`
|
||||
- All services must define these 4 methods first, and in that order
|
||||
- Name format `<operation>_<model>`, for example `create_document` or `update_job`.
|
||||
- Where a service exposes create/read/update/delete for its root model, define them at the
|
||||
top of the class in that order, before derived reads and workflow helpers.
|
||||
- Not every aggregate needs all four. `ExecutionAttempt` is append-only evidence written by
|
||||
`SourceService` workflow-facing methods, so `EvidenceService` deliberately exposes reads and
|
||||
no create or delete.
|
||||
Do not add unused CRUD methods to satisfy symmetry.
|
||||
- `RegistryService` is generic across small lookup models and uses `<operation>_entry`
|
||||
naming instead.
|
||||
|
||||
## Transaction Finalization
|
||||
|
||||
@@ -39,18 +145,9 @@ When a service method accepts an optional `session` kwarg, write methods must us
|
||||
- If `session` is provided: the method must **not** commit; it should `flush()` so IDs and FK values are available to the caller's transaction.
|
||||
- Use `refresh()` on returned ORM objects when the caller needs DB-populated values (defaults, triggers, merged state).
|
||||
|
||||
Recommended helper behavior:
|
||||
|
||||
- Inputs: active session object, original `session` arg (or a boolean ownership flag), and an optional list of objects to refresh.
|
||||
- Logic: `commit` when service-owned session, `flush` when caller-owned session, then refresh requested objects.
|
||||
|
||||
This keeps orchestration functions atomic: they can pass one shared session across multiple services and commit exactly once at the workflow boundary.
|
||||
|
||||
## Workflow Transaction Boundaries
|
||||
|
||||
For multi-step job lifecycles (for example queued transcription jobs), orchestration functions must use explicit transaction phases.
|
||||
|
||||
Required boundary model:
|
||||
For multi-step job lifecycles, orchestration functions must use explicit transaction phases.
|
||||
|
||||
- **Transaction A (claim):** transition `JobStatus.QUEUED -> JobStatus.PROCESSING` and commit immediately.
|
||||
- Perform provider/network work **outside** database transactions.
|
||||
@@ -64,14 +161,55 @@ Atomicity rules:
|
||||
- Terminal state (`TRANSCRIBED` or `FAILED`) and transcript row changes must succeed or roll back together.
|
||||
- Retry persistence (`QUEUED` + retry increment + error detail) must succeed or roll back together.
|
||||
|
||||
Separation of concerns:
|
||||
### Multi-page batches
|
||||
|
||||
- Worker modules should stay lightweight and delegate lifecycle transitions to service/workflow orchestration functions.
|
||||
- In `workflows.py`, `process_queued_job` should own one complete attempt lifecycle: `QUEUED -> PROCESSING -> TRANSCRIBED|FAILED`.
|
||||
- In `workflows.py`, `advance_job` should coordinate broader status progression around attempts (for example retry scheduling from `FAILED -> QUEUED`).
|
||||
- Services should expose session-aware write helpers (flush on caller-owned session) so orchestration controls commit boundaries.
|
||||
- Backoff/sleep behavior must run outside transactional scopes.
|
||||
These two requirements are in tension for multi-page jobs: each page should be durable as
|
||||
soon as its provider call returns, but the last page must commit together with the terminal
|
||||
status. `process_queued_job` resolves it by committing every page except the last one
|
||||
individually, then deferring the final page's write into `_finalize_batch_outcome` so it
|
||||
shares the terminal transaction.
|
||||
|
||||
Both paths are shielded against cancellation, so the final page is no less durable than the
|
||||
pages before it. Enforced by `tests/integration/test_pipeline_atomicity.py`; per-page
|
||||
durability is separately enforced by
|
||||
`tests/services/test_workflows_reliability.py::TestWorkflowReliability::test_transcribed_page_is_committed_before_next_provider_call_finishes`.
|
||||
|
||||
### Stale-reclaim safety
|
||||
|
||||
- `WORKER_STALE_JOB_SECONDS` must remain **greater than** `WORKER_PROVIDER_TIMEOUT_SECONDS`; stale
|
||||
recovery must not be able to fire before one provider call can legitimately finish.
|
||||
- Long-running multi-page orchestration must refresh job liveness explicitly between intermediate
|
||||
page commits. Do not rely on incidental row updates or provider metadata writes to keep
|
||||
`Job.date_updated` fresh.
|
||||
- Enforced by `tests/test_config.py` and
|
||||
`tests/services/test_workflows_reliability.py::TestWorkflowReliability::test_intermediate_page_commit_advances_job_liveness_timestamp`.
|
||||
|
||||
## Contract Alignment
|
||||
|
||||
- Treat `docs/` as the active architecture and requirements baseline.
|
||||
- Legacy revision trees are out of scope for active implementation decisions and must not be referenced as authoritative service guidance.
|
||||
- Treat `src/transcription/db/models.py` as runtime schema ground truth and `docs/schema.md` as the field-accurate contract mirror.
|
||||
- `Job.status` success path is `TRANSCRIBED`.
|
||||
- `JobSource.status` is queue/projection state only (`PENDING`, `TRANSCRIBED`, `FAILED`, `CANCELLED`).
|
||||
- Source ingest may normalize media before persistence; persisted bytes/hash are canonical for processing and provenance.
|
||||
- `ExecutionAttempt` is append-only evidence history; do not mutate historical attempt rows in runtime code.
|
||||
- `Source.raw_transcription` is a projection, not authoritative history.
|
||||
- Service/UI read paths that touch relationships must be eager-loaded for `lazy="raise"` compatibility.
|
||||
- If model fields, enums, constraints, indexes, or relationship-loading semantics change, update `docs/schema.md` in the same change.
|
||||
- If `Settings` fields or defaults change in `src/transcription/config.py`, update `.env.production.example` in the same change so keys/defaults remain synchronized and no stale settings remain documented.
|
||||
|
||||
## Schema Drift and Legacy Compatibility Policy
|
||||
|
||||
- Prefer schema migration over startup reconciliation or runtime compatibility paths in service writes.
|
||||
- Do not add legacy read/write compatibility code in service workflows by default.
|
||||
- If drift is discovered and a migration decision is ambiguous (for example, one-way destructive DDL, uncertain data retention impact, or unknown deployment sequence), pause and ask the user to choose migration vs compatibility before coding.
|
||||
- If a temporary compatibility path is explicitly approved, document an expiration/removal plan in the same change.
|
||||
|
||||
# Service Composition
|
||||
|
||||
Some operations, like uploading a picutre, require modifications to multiple tables, which can be done by composing methods from the service object into a separate function.
|
||||
A service method may read across models it does not own, using eager loads from its own
|
||||
aggregate root. What it may not do is import another service.
|
||||
|
||||
Operations that must **write** models owned by more than one service are composed in an orchestration module
|
||||
([store](../../src/transcription/services/store.py),
|
||||
[workflows](../../src/transcription/services/workflows.py)).
|
||||
|
||||
@@ -0,0 +1,119 @@
|
||||
---
|
||||
description: Authoring rules for the test suite, including markers, async discipline, and guard-test design.
|
||||
applyTo: 'tests/**/*.py'
|
||||
---
|
||||
|
||||
# Tests
|
||||
|
||||
Primary references:
|
||||
|
||||
- `AGENTS.md` (Change Protocol — failing test first)
|
||||
- `docs/index.md` and `docs/invariant/*`
|
||||
- `.github/skills/test-effectiveness-auditor/skill.md` (periodic audit of this suite)
|
||||
|
||||
The suite is not only regression protection here — it is where several architectural rules are
|
||||
*defined*. `tests/test_service_boundaries.py`, `tests/test_ui_boundaries.py`,
|
||||
`tests/test_provider_boundaries.py`, `tests/test_model_contract_guards.py`, and
|
||||
`tests/test_meta_contract_guards.py` are the enforcement layer named in the `AGENTS.md` authority
|
||||
order. A weak test in this repository does not merely fail to catch a bug; it can silently repeal a
|
||||
documented invariant.
|
||||
|
||||
The baseline is green. `uv run pytest -q -m "not external"` must report zero failures and zero
|
||||
errors, and there is no tolerated set of known-failing tests.
|
||||
|
||||
## Write the Failing Test First
|
||||
|
||||
For any behavioral fix, write the test before the fix and confirm it fails *for the reason you
|
||||
expect*. A test that passes against the broken code proves nothing, and several defects in this
|
||||
repository were subtle enough that a test written afterward would have done exactly that. If the
|
||||
new test passes immediately, you have not reproduced the defect yet.
|
||||
|
||||
## Runner Configuration
|
||||
|
||||
Configured in `pyproject.toml`; do not work around these:
|
||||
|
||||
- `--strict-markers` — an unregistered marker is an error. Register new markers in
|
||||
`[tool.pytest.ini_options] markers` with a description rather than inventing one at the call site.
|
||||
- `asyncio_mode = "strict"` — every async test needs an explicit `@pytest.mark.asyncio`, and async
|
||||
fixtures use `@pytest_asyncio.fixture`. There is no implicit promotion.
|
||||
- `filterwarnings = ["error:coroutine .* was never awaited:RuntimeWarning"]` — an un-awaited
|
||||
coroutine is an error, not a warning. This usually means a mock replaced an async callable with a
|
||||
sync one, or an `await` was dropped. Fix the call; never silence the warning.
|
||||
|
||||
## Markers and Layout
|
||||
|
||||
- `unit` — pure logic, no framework or database.
|
||||
- `integration` — touches framework, database, or multi-component contracts.
|
||||
- `external` — calls live services; slow and credential-dependent.
|
||||
|
||||
`external` tests must also carry their own `skipif` so the suite stays green without credentials
|
||||
(see `tests/services/test_transcription_external.py`). Local and documented runs use
|
||||
`-m "not external"`; CI intentionally runs unfiltered, which is equivalent because those tests skip
|
||||
themselves. Never let an unmarked test reach the network.
|
||||
|
||||
Place tests by the layer under test: `tests/services/`, `tests/ui/`, `tests/api/`,
|
||||
`tests/providers/`, `tests/integration/`, with cross-cutting guards at the top level.
|
||||
|
||||
## Fixtures and Isolation
|
||||
|
||||
- Prefer the shared fixtures in `tests/conftest.py` (`default_settings`, `async_session`,
|
||||
`default_session_factory`, and the per-aggregate service fixtures) over building settings or
|
||||
engines by hand.
|
||||
- `Settings` is isolated suite-wide by the session-scoped autouse fixture in `conftest.py`, because
|
||||
`env_file` resolves against the working directory. Tests that need env-file loading pass
|
||||
`_env_file=` explicitly; tests asserting declared defaults need nothing. Do not reintroduce
|
||||
reliance on a developer's local env file. Guarded by `tests/test_config_isolation.py`.
|
||||
- Database fixtures refuse to run against anything but the per-test path, and that refusal is
|
||||
deliberate. Never relax it to point a destructive fixture at a real database.
|
||||
- Tests must not leave artifacts outside `tmp_path`.
|
||||
|
||||
## Assertion Strength
|
||||
|
||||
Assert on the domain effect, not on the fact that code ran.
|
||||
|
||||
- Prefer persisted state, status transitions, error categories, and evidence records over
|
||||
"no exception raised", "not None", or a bare status code.
|
||||
- **Read committed state through a separate session.** Asserting against the same session that
|
||||
performed the write can pass on unflushed in-memory state and prove nothing about durability.
|
||||
This is how the atomicity guarantees in `tests/services/test_workflows_reliability.py` and
|
||||
`tests/integration/test_pipeline_atomicity.py` are made real.
|
||||
- Critical paths need negative-path coverage — timeouts, provider failures, validation errors,
|
||||
cancellation. Happy-path-only coverage of a critical module is a gap, not a suite.
|
||||
- Avoid count-threshold assertions as a proxy for correctness. A test asserting "at least N items
|
||||
were discovered" passes indefinitely while the thing it was meant to protect rots; assert on a
|
||||
specific known member instead.
|
||||
|
||||
## Guard Tests
|
||||
|
||||
Structural guards carry extra obligations, because they are cited as proof that a rule holds.
|
||||
|
||||
- **Guard the guard.** Every scanning guard needs a companion assertion that the scan actually found
|
||||
something, following the existing `test_*_are_discovered` pattern. A guard that silently scans an
|
||||
empty set passes forever.
|
||||
- **Scope must match the claim.** A guard's name and docstring must describe only what it actually
|
||||
verifies. A test covering one function while appearing to enforce a repo-wide rule is worse than
|
||||
no test, because it stops anyone from writing the real one.
|
||||
- **Prove non-vacuity by injected fault.** Temporarily introduce the violation, confirm the guard
|
||||
fails with a comprehensible message, then revert. Do this whenever you add or materially change a
|
||||
guard. Revert with an explicit edit if the file has uncommitted changes — `git checkout --` will
|
||||
discard them.
|
||||
- **Prefer structural analysis to substring matching.** AST inspection of imports and definitions is
|
||||
resistant to false negatives; a bare-name search across the repository is not, since an unrelated
|
||||
mention anywhere makes dead code look reachable.
|
||||
- Failure messages should name the offending file, symbol, and the remedy. These fire for people who
|
||||
did not write the guard.
|
||||
- Any new file under `.github/**` must be added to `ACTIVE_CONTRACT_FILES` in
|
||||
`tests/test_meta_contract_guards.py`, or the completeness guard fails by design.
|
||||
|
||||
## Redundancy
|
||||
|
||||
Duplicate coverage across layers costs runtime and dilutes signal. Pick the canonical layer for a
|
||||
behavior — unit for logic, integration for wiring — and let the other layer assert only what is
|
||||
unique to it. Retire tests superseded by a stronger guard instead of accumulating both, and record
|
||||
deliberate retentions with a rationale rather than leaving them unexplained.
|
||||
|
||||
## Contract Sync Rule
|
||||
|
||||
When a test encodes or relaxes a documented rule, update the corresponding instruction file or
|
||||
`docs/*` page in the same change. When a guard test is the enforcement for a rule stated in
|
||||
`AGENTS.md` or an instruction file, cite the test by name there so the link survives refactoring.
|
||||
@@ -11,6 +11,9 @@ Keep dependencies flowing in this direction:
|
||||
|
||||
Pages may depend on application services and framework-provided dependencies. Components may depend on smaller components and shared presentation helpers. Services and domain modules must never depend on the UI.
|
||||
|
||||
Cross-cutting error behavior must follow
|
||||
[error-handling instructions](./error-handling.instructions.md).
|
||||
|
||||
## Package Root
|
||||
|
||||
- Keep `ui/__init__.py` as the UI composition root: register global assets, register pages, and mount NiceGUI on FastAPI.
|
||||
@@ -37,17 +40,39 @@ Pages may depend on application services and framework-provided dependencies. Co
|
||||
- Keep app-wide navigation and layout primitives in `components/app_shell.py`.
|
||||
- Keep generic table/event adaptation in `components/table/common.py`; feature-specific columns, row read models, and formatting belong in the feature table module.
|
||||
- Keep exception normalization and user-facing error display in `components/error_presenter.py`; preserve `AppError` details and operation identifiers at page/component boundaries.
|
||||
- Use `components/media_urls.py` for media URL generation; do not hand-build upload/static paths in page code.
|
||||
|
||||
## CSS Assets
|
||||
|
||||
- Keep CSS under `ui/static` and split it into manageable, feature-oriented files. Do not grow a monolithic stylesheet or embed substantial style blocks in Python components.
|
||||
- Load each stylesheet from the page, component, or composition root that needs it with `ui.add_css(...)`. Use shared registration only for genuinely application-wide styles.
|
||||
- Keep all application CSS in `ui/static/theme.css`; do not add page- or component-specific stylesheets or embed style blocks in Python components.
|
||||
- Load `theme.css` once from the composition root with `ui.add_css(..., shared=True)`.
|
||||
- Read stylesheet text through `importlib.resources.files(...)` so loading works from installed packages and is independent of the working directory.
|
||||
- Centralize CSS reading in one typed helper cached by relative resource path with `functools.cache` or an equivalent unbounded `lru_cache`. Cache the immutable stylesheet text to avoid repeated resource I/O during component renders; keep NiceGUI registration decisions at the caller.
|
||||
- Centralize CSS reading in one typed helper cached by resource path.
|
||||
- Do not encode application behavior in CSS or other static assets.
|
||||
|
||||
## State and Side Effects
|
||||
|
||||
- Limit component state to ephemeral interaction state such as loading flags, form values, dialogs, and expansion state.
|
||||
- Application and worker state must be resolved at the page or application boundary and passed through narrow interfaces such as callbacks or notifier protocols.
|
||||
- Keep filesystem, network, provider, and worker orchestration behind application services or dedicated adapters. UI code may trigger those operations but must not implement them.
|
||||
- Application and worker state must be resolved at the page or application boundary and passed through narrow interfaces.
|
||||
- Keep filesystem, network, provider, and worker orchestration behind application services or dedicated adapters.
|
||||
|
||||
## Media Route Safety Rules
|
||||
|
||||
Two patterns are approved:
|
||||
|
||||
1. **Record-validated API routes** for print/export contexts.
|
||||
2. **Controlled upload URL resolver** (`components/media_urls.py`) for general UI media.
|
||||
|
||||
Prohibited patterns:
|
||||
|
||||
- Direct `file://` links or exposing local filesystem paths.
|
||||
- Manual URL construction from raw `Path` values in pages/components.
|
||||
- User-facing payloads containing local absolute paths.
|
||||
|
||||
## Contract Alignment
|
||||
|
||||
- Treat `docs/` as the active baseline.
|
||||
- Resolve lifecycle and status semantics against `src/transcription/db/models.py` and `docs/schema.md`; do not introduce alternate status labels or implied legacy states in UI behavior.
|
||||
- Use status vocabulary exactly as modeled (`queued`, `processing`, `transcribed`, `partial_success`, `failed`; and `pending`, `transcribed`, `failed`, `cancelled`).
|
||||
- Print/export media flows must use record-validated routes; direct local filesystem paths are prohibited.
|
||||
- If lifecycle wording/behavior changes, update corresponding `docs/ui/pages/*.md` contracts in the same change.
|
||||
|
||||
@@ -0,0 +1,23 @@
|
||||
---
|
||||
name: Review Python Architecture
|
||||
description: Run an evidence-based architectural code review using the Python Architect Reviewer agent and python-code-reviewer skill.
|
||||
agent: Python Architect Reviewer
|
||||
---
|
||||
|
||||
# Instructions
|
||||
|
||||
Execute a comprehensive, evidence-based code review of the target codebase.
|
||||
|
||||
## Target Scope
|
||||
- **Review Target:** the repository root, unless the invoker names a narrower path; review that path instead.
|
||||
- **Source Root:** `src/`
|
||||
- **Docs Root:** `docs/`
|
||||
- **Focus Areas:** FastAPI endpoints, NiceGUI components, SQLModel persistence, asyncio workers, Pydantic V2 models, and OpenRouter provider adapters.
|
||||
|
||||
## Execution Rules
|
||||
1. Map repository layout, dependency manifests, and configuration files from the project root before inspecting modules.
|
||||
2. Read real code modules under `src/` (or the specified target path); cite exact file paths and line ranges for every finding.
|
||||
3. Validate issues by running `uv run ruff check .`, `uv run ty check`, and `uv run pytest -q -m "not external"`. Nothing is on `PATH` in this `uv` project, so bare `ruff`/`ty`/`pytest` will fail.
|
||||
4. Check for duplication, divergent implementations, and extractable helpers.
|
||||
5. Format the entire review following the standardized 10-section template defined in the `python-code-reviewer` skill.
|
||||
6. Write the final report to `./docs/reviews/<YYYY-MM-DD>-code-review.md`, using today's date. This path is defined by the skill; do not write the report anywhere else.
|
||||
@@ -0,0 +1,80 @@
|
||||
---
|
||||
name: evidence-provenance-auditor
|
||||
description: Deterministic reviewer for transcription evidence/provenance guarantees. Use when changes touch execution attempts, source storage, retries, transport evidence, artifact provenance, or evidence exports.
|
||||
---
|
||||
|
||||
# Evidence & Provenance Auditor
|
||||
|
||||
Perform focused, deterministic audits of evidence integrity and provenance behavior.
|
||||
|
||||
## When to Use
|
||||
|
||||
- Reviewing changes in:
|
||||
- `src/transcription/services/sources.py`
|
||||
- `src/transcription/services/store.py`
|
||||
- `src/transcription/services/workflows.py`
|
||||
- `src/transcription/services/evidence.py`
|
||||
- `src/transcription/db/models.py`
|
||||
- Auditing evidence exports/imports or evidence-display behavior.
|
||||
- Verifying no drift from canonical provenance invariants.
|
||||
|
||||
## Normative References (must be used)
|
||||
|
||||
1. `docs/invariant/ai_evidence_and_provenance.md`
|
||||
2. `docs/schema.md`
|
||||
3. `docs/requirements.md`
|
||||
4. `docs/error_handling.md`
|
||||
|
||||
## Deterministic Pass/Fail Checks
|
||||
|
||||
### A. Append-only history
|
||||
- Every provider call results in a new `ExecutionAttempt`.
|
||||
- Runtime paths do not mutate historical attempts to represent new outcomes.
|
||||
- Retry behavior appends attempts rather than rewriting prior rows.
|
||||
|
||||
### B. Projection vs authority separation
|
||||
- `Source.raw_transcription` and preferred pointers are mutable projection surfaces.
|
||||
- Attempt rows remain authoritative historical evidence.
|
||||
- Candidate promotion updates projection pointers without rewriting history.
|
||||
|
||||
### C. Transport evidence semantics
|
||||
- Transport evidence is correctly labeled as application-boundary capture.
|
||||
- SDK snapshots/normalized metadata are not mislabeled as native upstream payload.
|
||||
- No-response timeout/network states are explicit.
|
||||
|
||||
### D. Canonical source identity
|
||||
- Canonical stored bytes/hash/size are internally consistent.
|
||||
- If ingest normalization is applied, code/docs consistently represent resulting canonical identity.
|
||||
- Post-ingest derivatives do not overwrite canonical source bytes.
|
||||
|
||||
### E. Secret safety
|
||||
- No credentials/auth headers/cookies/unrestricted headers persisted.
|
||||
- Header persistence uses explicit allowlist semantics.
|
||||
|
||||
### F. Route/path safety
|
||||
- Print/export source access is record-validated.
|
||||
- UI/media path construction does not expose local filesystem paths.
|
||||
|
||||
### G. Schema/docs alignment
|
||||
- Evidence-related model fields and semantics align with canonical docs.
|
||||
- Evidence model changes require same-change doc updates.
|
||||
|
||||
### H. Canonical authority boundaries
|
||||
- Active guidance resolves against `docs/*` and current instruction files.
|
||||
|
||||
## Review Workflow
|
||||
|
||||
1. Read normative references first.
|
||||
2. Inspect model + service + workflow write paths.
|
||||
3. Inspect evidence read/display/export paths.
|
||||
4. Report high-confidence findings with concrete path/line evidence.
|
||||
5. Classify each finding by invariant family (A-H).
|
||||
|
||||
## Output Format
|
||||
|
||||
Use this structure:
|
||||
|
||||
- Verdict by invariant family (A-H)
|
||||
- Findings with `Location`, `Observed Behavior`, `Risk`, `Recommended Fix`
|
||||
- Drift table (`Doc claim` vs `Code reality` vs `Action`)
|
||||
- Regression guards needed
|
||||
@@ -0,0 +1,276 @@
|
||||
---
|
||||
name: python-code-reviewer
|
||||
description: Perform an evidence-based, senior architect code review for Python codebases using FastAPI, NiceGUI, SQLModel, SQLAlchemy, Pydantic V2, asyncio, and OpenRouter. Use when asked to review Python repositories, perform architectural or code audits, or evaluate code against Python 3.12+ best practices.
|
||||
---
|
||||
|
||||
# Python Code Reviewer
|
||||
|
||||
Perform thorough, evidence-based code reviews for Python projects. Every finding must cite concrete file paths and line ranges, avoid speculation, and include recommended fixes.
|
||||
|
||||
## When to Use
|
||||
|
||||
- Performing an architectural or code quality review of a Python codebase.
|
||||
- Auditing applications using FastAPI, NiceGUI, SQLModel/SQLAlchemy, Pydantic V2, or asyncio workers.
|
||||
- Generating structured Markdown review reports in `./docs/reviews`.
|
||||
|
||||
## Technical Stack Scope
|
||||
|
||||
- **Runtime:** Python 3.12+
|
||||
- **Web Application:** FastAPI and NiceGUI
|
||||
- **Persistence:** SQLModel, SQLAlchemy (SQLite and PostgreSQL support)
|
||||
- **Validation & Settings:** Pydantic V2 and pydantic-settings
|
||||
- **Concurrency:** Python asyncio workers
|
||||
- **Vision/LLM Integration:** OpenRouter / provider adapters
|
||||
- **Image & Print Pipeline:** Pillow-backed media handling and print/export rendering
|
||||
- **Quality & Testing:** pytest, pytest-asyncio, Ruff, and ty
|
||||
|
||||
NiceGUI is pinned to an exact version (`nicegui==3.13.0` in `pyproject.toml`); API guidance
|
||||
must be correct for that release rather than for the latest published version. The exact pin is
|
||||
a deliberate release-stability decision recorded in `docs/production-runbook.md` ("Dependency
|
||||
upgrade policy") — do not report it as a defect or recommend widening it.
|
||||
|
||||
## Review Workflow
|
||||
|
||||
1. **Map the Repository First:** Inspect entry points, package layout, configurations, dependency manifests, and any project-specific rule files (`AGENTS.md`, `.github/instructions/`, `.github/skills/`). Project-specific conventions override generic advice.
|
||||
2. **Establish Canonical Authority First:** Read architecture/contracts (`docs/*`, `docs/invariant/*`, UI docs) and active instructions/skills before evaluating source behavior.
|
||||
3. **Read Representative Modules:** Sample across all layers (routes/pages, UI components, services, workers, persistence, provider adapters, settings, tests) before drawing conclusions.
|
||||
4. **Run Drift Analysis:** Compare documented intended behavior versus repository ground truth; identify both implementation drift and undocumented-but-repeatable conventions that should be formalized.
|
||||
5. **Run Dead-Code/Orphan Sweep:** Identify candidate orphan modules/functions/classes with zero inbound references, then verify expected exceptions (entrypoints, framework/plugin registration, dynamic imports/reflection, CLI hooks, test-only utilities) before marking as orphaned.
|
||||
6. **Assess Boundary and Coupling Health:** Evaluate UI/service/persistence/provider dependency flow, identify circular dependencies, leaky abstractions, and transaction ownership ambiguity.
|
||||
7. **Assess Invariant Placement:** For each hard rule, decide whether it belongs in docs (rationale), instructions (active steering), skills (periodic audit procedure), or deterministic tests (enforcement).
|
||||
8. **Verify Claims:** This is a `uv` project (`uv.lock`, root `ruff.toml`). Run `uv run ruff check .`, `uv run ty check`, and `uv run pytest -q -m "not external"` rather than guessing, and record the exact commands and their outcomes in the report.
|
||||
9. **Validate Recommendations Against Consumers:** A recommendation is a claim about the future and must be verified like any other. Before recommending a change to a shared symbol — a model field, an exception attribute, a helper's return value, a function signature — enumerate **every** consumer of that symbol (`grep` the whole repo, including tests) and confirm the fix is safe for each one. Record the consumers in the finding's **Blast Radius**. A fix that is correct for the path that produced the finding can silently break a second consumer, and evidence/provenance and logging paths are the usual casualties because they read the same fields the UI does.
|
||||
10. **Prioritize Hot Paths:** Focus deeply on request handling, database sessions, background workers, and external API calls.
|
||||
11. **Enforce Read-Only Safety:** Do not modify code unless explicitly instructed.
|
||||
12. **Escalate Provenance Audits:** For evidence/provenance-heavy changes, apply invariant checks from `.github/skills/evidence-provenance-auditor/skill.md` and include pass/fail outcomes in the report.
|
||||
13. **Escalate Test-Suite Audits:** When findings touch test coverage, redundancy, or assertion strength, apply `.github/skills/test-effectiveness-auditor/skill.md` and include its outcomes alongside the provenance results.
|
||||
|
||||
### Worked example: why step 9 exists
|
||||
|
||||
The 2026-08-23 review recommended fixing a filesystem-path leak in
|
||||
`classify_unexpected_error` by making `AppError.message` generic and logging the exception
|
||||
detail instead. The analysis of the leak was correct, and the fix was implemented as written.
|
||||
|
||||
It was wrong. `AppError.message` had a second consumer the review never traced:
|
||||
`format_error_detail`, which writes `ExecutionAttempt.error_detail` — a **provenance record**.
|
||||
The recommended fix closed a privacy leak by silently stripping root-cause data from the
|
||||
evidence history this system exists to preserve. It was caught only because an unrelated
|
||||
integration test asserted on the persisted error text.
|
||||
|
||||
The correct fix separated the audiences — a user-safe `message` and an internal-only `detail`
|
||||
that still reaches evidence and logs. One `grep` for consumers of `.message` during the review
|
||||
would have found this. Treat any recommendation that changes a widely-read field as unverified
|
||||
until its consumers are enumerated.
|
||||
|
||||
## Repo-Specific Deterministic Checks (Transcription)
|
||||
When reviewing this repository, always include explicit pass/fail checks for the following.
|
||||
Where **Enforced by** reads *unenforced*, recommending a deterministic test is itself a finding.
|
||||
|
||||
| # | Check | Enforced by |
|
||||
| :-- | :--- | :--- |
|
||||
| 1 | **Service boundary rule:** no service-to-service imports | `tests/test_service_boundaries.py` |
|
||||
| 2 | **UI boundary rule:** pages/components do not perform persistence access | `tests/test_ui_boundaries.py` |
|
||||
| 3 | **Status vocabulary conformance:** `JobStatus`/`JobSourceStatus`/`JobPurpose` usage matches current enums in `src/transcription/db/models.py`; no stringly-typed status literals | `tests/test_model_contract_guards.py` |
|
||||
| 4 | **Evidence ownership conformance:** append-only attempt history is preserved and projection writes are not mistaken for history mutation (`src/transcription/services/sources.py`, `src/transcription/services/evidence.py`) | `tests/test_evidence_provenance.py::test_attempts_are_append_only_and_exported_with_integrity` |
|
||||
| 5 | **Canonical authority:** findings must resolve against `docs/*` first | `tests/test_meta_contract_guards.py::test_canonical_authority_references_are_present` |
|
||||
| 6 | **Schema contract fidelity:** when model/persistence behavior changes, `docs/schema.md` remains field-accurate with `src/transcription/db/models.py` | `tests/test_model_contract_guards.py` (field names, ordering, enum members, table coverage), `tests/test_meta_contract_guards.py` (presence and references) |
|
||||
| 7 | **Media boundary conformance:** print/export media is record-validated and UI media URL generation uses controlled resolver paths | `tests/test_media_path_safety.py`, `tests/ui/test_media_urls.py` |
|
||||
| 8 | **Eager-loading conformance:** service/UI read paths satisfy `lazy="raise"` expectations | `tests/test_model_contract_guards.py` (declaration-side; documented `noload` exceptions must match `docs/schema.md`) |
|
||||
| 9 | **Cross-cutting error conformance:** service/API/UI translation and retry behavior align with `.github/instructions/error-handling.instructions.md` | `tests/test_errors.py`, `tests/api/test_error_responses.py`, `tests/ui/test_error_presenter.py` |
|
||||
| 10 | **Orphaned/dead-code conformance:** include a deterministic orphan sweep and report confirmed orphans removed/retained with rationale | `tests/test_orphan_sweep.py` (`KNOWN_ORPHANS` records each retained orphan and its rationale) |
|
||||
|
||||
## Core Review Areas
|
||||
|
||||
### 1. Python Best Practices (3.12+)
|
||||
- **Type Annotations:** Ensure completeness, modern syntax (`X | None`, builtin generics, `Self`, `type` statements), and avoid unparameterized containers or bare `Any`.
|
||||
- **Error Handling:** Identify bare/broad `except`, swallowed exceptions, missing `raise ... from`, and exceptions used for control flow.
|
||||
- **Resource Management:** Verify context managers for files, DB sessions, HTTP clients, and locks. Check for leaked tasks or connections.
|
||||
- **Data Modeling:** Check proper use of dataclasses vs. Pydantic models vs. dictionaries. Eliminate mutable default arguments and stringly-typed payloads.
|
||||
- **Idioms & Clean Code:** Verify `pathlib` usage over `os.path`, comprehensions vs manual loops, removal of dead code, and elimination of magic numbers.
|
||||
|
||||
### 2. FastAPI
|
||||
- **Dependency Injection:** Verify `Depends` is used for shared resources (DB sessions, settings, clients) rather than global singletons.
|
||||
- **Route Design:** Validate HTTP verbs, status codes, path/query/body typing, `response_model`, and domain-based router organization.
|
||||
- **Lifecycle & Concurrency:** Ensure lifespan handlers are used instead of deprecated `@app.on_event`. Flag blocking synchronous calls in `async def` endpoints.
|
||||
|
||||
### 3. NiceGUI
|
||||
- **Separation of Concerns:** Ensure UI components delegate business logic and persistence to service layers.
|
||||
- **Client State Handling:** Verify correct use of client-scoped state vs global state to avoid state leaks across sessions.
|
||||
- **Async Execution:** Check for blocking operations on the UI event loop and unbounded timers/pollers.
|
||||
|
||||
### 4. Persistence (SQLModel / SQLAlchemy)
|
||||
- **Session Lifecycle:** Enforce one session per request/unit of work with explicit commit/rollback/close boundaries.
|
||||
- **Query Optimization:** Detect N+1 patterns, missing eager loads (`selectinload`/`joinedload`), queries inside loops, and unindexed filters.
|
||||
- **Cross-Dialect Portability:** Check compatibility for both SQLite (WAL mode, pragmas) and PostgreSQL (JSONB, locking, autoincrement).
|
||||
|
||||
### 5. Pydantic V2 & Settings
|
||||
- **V2 Migration:** Flag legacy V1 patterns (`@validator`, `Config` class, `.dict()`, `parse_obj`) and use V2 equivalents (`@field_validator`, `model_config = ConfigDict(...)`, `model_dump()`).
|
||||
- **Settings Management:** Ensure `BaseSettings` is the single source of truth without scattered `os.getenv` calls or committed secrets.
|
||||
|
||||
### 6. Concurrency & Asyncio Workers
|
||||
- **Task Lifecycle:** Flag unreferenced `create_task` calls that risk garbage collection, missing cancellation handling, and lack of graceful shutdown.
|
||||
- **Backpressure & Synchronization:** Check for appropriate use of `asyncio.Queue`, `TaskGroup`, `Lock`, and backoff retries.
|
||||
|
||||
### 7. Provider Adapters (OpenRouter / APIs)
|
||||
- **Adapter Encapsulation:** Verify provider-specific details (headers, model names, payload formats) do not leak into UI or business logic.
|
||||
- **Client Lifecycle:** Reuse shared `AsyncClient` instances with proper connection pooling and timeouts. Validate API responses using Pydantic schemas.
|
||||
|
||||
### 8. Testing & Quality Tooling
|
||||
- **Test Isolation:** Verify tests do not rely on live external services, real clocks, or shared global state.
|
||||
- **Async Test Setup:** Check `pytest-asyncio` configuration and fixture lifecycle.
|
||||
- **Project Test Contract (`pyproject.toml`):** `--strict-markers` is enabled, so every marker must be declared; `asyncio_mode = "strict"` requires explicit `@pytest.mark.asyncio`; declared markers are `unit`, `integration`, and `external`, and `external` must be excluded from default verification runs. `filterwarnings` promotes `coroutine ... was never awaited` to an **error** — treat any unawaited coroutine as a hard failure and a Critical/High finding, never a warning.
|
||||
- **Suite Signal Quality:** For low-value, redundant, or tautological tests, escalate to `.github/skills/test-effectiveness-auditor/skill.md` and fold its outcomes into the report.
|
||||
|
||||
### 9. Duplication & Consolidation
|
||||
- Identify repeated code blocks, candidate helper abstractions, divergent patterns for identical operations, and duplicated domain constants.
|
||||
|
||||
### 10. Orphaned/Dead Code Audit
|
||||
- Find candidate orphan modules/functions/classes with no inbound references.
|
||||
- Validate each candidate against dynamic wiring exceptions (entrypoints, plugin registration, reflection/dynamic imports, CLI hooks, test utilities).
|
||||
- Report outcomes as: removed orphan, retained-with-justification, or uncertain-follow-up.
|
||||
|
||||
### 11. Architecture & Governance
|
||||
- **Architectural Drift:** Compare intended architecture rules against implementation behavior and cite concrete drift points.
|
||||
- **Systemic Health:** Evaluate domain cohesion, dependency direction, lifecycle consistency, and operational reliability seams.
|
||||
- **Invariant Routing:** Recommend the correct enforcement layer per rule (docs vs instructions vs skills vs tests).
|
||||
- **Meta-Tooling Alignment:** Recommend updates for instruction files and skills when repository patterns or contracts evolve.
|
||||
|
||||
## Severity Rubric
|
||||
|
||||
Severity reflects concrete consequence, never style preference or effort to fix.
|
||||
|
||||
- **Critical:** Data or evidence loss/corruption; provenance or append-only history violated; secret leakage; silent wrong output presented as authoritative.
|
||||
- **High:** Architectural boundary violated (service/UI/persistence/provider); runtime failure or unhandled exception on a hot path (request handling, DB sessions, worker loop, external API calls); documented invariant contradicted by implementation.
|
||||
- **Medium:** Correctness risk under load or edge conditions (N+1, missing eager load, leaked task, missing timeout); drift between docs and code with no immediate runtime impact.
|
||||
- **Low:** Maintainability, typing completeness, duplication, naming, or dead code with no behavioral risk.
|
||||
|
||||
### Reachability
|
||||
|
||||
Severity states how bad the consequence is; **Reachability** states whether it can happen today.
|
||||
They are independent, and a finding is not complete without both. Record one of:
|
||||
|
||||
- **Live:** reachable in the current configuration and deployment.
|
||||
- **Latent:** the defective code is present but unreachable because of a current setting, single-
|
||||
instance deployment, or absent caller. **State the exact condition that unblocks it.**
|
||||
- **Theoretical:** requires a combination the project has explicitly ruled out.
|
||||
|
||||
Latent findings carry a scheduling constraint that severity alone cannot express: a latent defect
|
||||
must usually be fixed *before* the change that makes it live, not after. Say so explicitly in the
|
||||
finding and reflect the ordering in the §9 action plan — for example, "fix the retry-category gate
|
||||
before raising `worker_max_retries` above 0," or "handle this `IntegrityError` before deploying a
|
||||
second worker replica." Do not downgrade severity merely because a finding is latent.
|
||||
|
||||
### Conflicting invariants
|
||||
|
||||
When a fix sits between two invariants that pull in opposite directions, say so in the
|
||||
**Recommendation** and name both, along with the test that guards each. Flag explicitly what the
|
||||
over-correction would be, because the simplest-looking fix usually satisfies one invariant by
|
||||
silently destroying the other. A recommendation that resolves one side without naming the other is
|
||||
incomplete and will be implemented incorrectly.
|
||||
|
||||
## Output Report Structure & Template
|
||||
|
||||
Generate Markdown reports at `./docs/reviews/<YYYY-MM-DD>-code-review.md` following this exact
|
||||
template structure. Reports are dated, non-canonical artifacts: `docs/reviews/**` is explicitly
|
||||
**not** part of the canonical authority set that the canonical-authority check resolves against.
|
||||
|
||||
```markdown
|
||||
# Architecture & Code Review Report
|
||||
|
||||
**Repository Target:** `project-root/`
|
||||
**Target Stack:** Python 3.12+ | FastAPI | NiceGUI | SQLModel/SQLAlchemy | Pydantic V2 | asyncio | OpenRouter
|
||||
|
||||
---
|
||||
|
||||
## 1. Executive Summary
|
||||
- 5-10 bullets on overall health, top risks, and high-leverage refactors.
|
||||
|
||||
---
|
||||
|
||||
## 2. Executive Architecture Assessment
|
||||
- High-level verdict on domain cohesion, boundary clarity, and architecture fitness.
|
||||
- Top 3-5 systemic risks or bottlenecks.
|
||||
|
||||
---
|
||||
|
||||
## 3. Findings by Severity
|
||||
|
||||
### Critical Severity
|
||||
#### [CRIT-01] Title
|
||||
- **Location:** `path/to/file.py:lines`
|
||||
- **Reachability:** Live / Latent (state the exact condition that unblocks it) / Theoretical
|
||||
- **Problem & Consequence:** Concrete consequence, not a style opinion.
|
||||
- **Blast Radius:** Every consumer of the symbols the recommendation changes, each confirmed
|
||||
safe. Write `None — change is local` only after actually searching. If the fix touches a
|
||||
shared field or helper, list the call sites (including tests and evidence/logging paths).
|
||||
- **Recommendation:** Fix with before/after sketch. If two invariants conflict here, name both,
|
||||
name the test guarding each, and state what the over-correction would be.
|
||||
- **Effort:** S / M / L
|
||||
|
||||
### High Severity
|
||||
#### [HIGH-01] Title
|
||||
...
|
||||
|
||||
### Medium Severity
|
||||
#### [MED-01] Title
|
||||
...
|
||||
|
||||
### Low Severity
|
||||
#### [LOW-01] Title
|
||||
...
|
||||
|
||||
---
|
||||
|
||||
## 4. Architectural Drift & Gap Analysis
|
||||
|
||||
`Direction` is `doc->code` (implementation must change to match documented intent) or
|
||||
`code->doc` (an undocumented but repeatable convention that should be formalized).
|
||||
|
||||
| Area / Component | Direction | Documented / Intended Rule | Actual Implementation State | Severity | Recommended Resolution |
|
||||
| :--- | :--- | :--- | :--- | :--- | :--- |
|
||||
|
||||
---
|
||||
|
||||
## 5. Invariant Inventory & Routing Recommendations
|
||||
| Invariant / Constraint | Current Location | Recommended Target Layer | Rationale |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
|
||||
---
|
||||
|
||||
## 6. Stack-Specific Analysis
|
||||
- Python 3.12+ Best Practices
|
||||
- FastAPI
|
||||
- NiceGUI
|
||||
- SQLModel & SQLAlchemy
|
||||
- Pydantic V2 & Settings
|
||||
- Asyncio Workers
|
||||
- OpenRouter / Adapter Boundary
|
||||
- Testing & Quality Tooling
|
||||
|
||||
---
|
||||
|
||||
## 7. Duplication & Consolidation Report
|
||||
| Pattern / Duplication | Locations | Proposed Canonical Home | Estimated Lines Removed |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
|
||||
### Proposed Canonical Abstractions
|
||||
- Code signatures and implementation homes.
|
||||
|
||||
---
|
||||
|
||||
## 8. Meta-Tooling & Instruction Update Recommendations
|
||||
- Required updates to docs, instructions, skills, or tests to keep enforcement current.
|
||||
|
||||
---
|
||||
|
||||
## 9. Prioritized Dependency-Ordered Action Plan
|
||||
1. **Phase 1: Blocking fixes**
|
||||
2. **Phase 2: Enforcement hardening**
|
||||
3. **Phase 3: Reliability & concurrency**
|
||||
4. **Phase 4: Consolidation & refactoring**
|
||||
5. **Phase 5: Non-blocking governance/documentation depth**
|
||||
|
||||
---
|
||||
|
||||
## 10. Preserved Strengths
|
||||
- Existing patterns worth maintaining.
|
||||
@@ -0,0 +1,98 @@
|
||||
---
|
||||
name: test-effectiveness-auditor
|
||||
description: Periodic reviewer for test-suite signal quality. Detects low-value or redundant tests, validates contract coverage, and recommends pruning or strengthening actions.
|
||||
---
|
||||
|
||||
# Test Effectiveness Auditor
|
||||
|
||||
Run a deterministic audit of test usefulness. Focus on whether tests catch real regressions, not whether they merely execute code.
|
||||
|
||||
## When to Use
|
||||
|
||||
- Monthly/quarterly test-health review.
|
||||
- Pre-release hardening when test count grows quickly.
|
||||
- After major AI-assisted test generation.
|
||||
- When suite runtime is increasing without clear quality gains.
|
||||
|
||||
## Primary Objectives
|
||||
|
||||
1. Identify tests that are weak, redundant, or non-diagnostic.
|
||||
2. Confirm critical contracts are guarded by meaningful assertions.
|
||||
3. Produce a prune/strengthen backlog with explicit risk and effort.
|
||||
|
||||
## Normative References (Transcription Repo)
|
||||
|
||||
1. `docs/*`
|
||||
2. `docs/invariant/*`
|
||||
3. `.github/instructions/*.instructions.md`
|
||||
4. `tests/test_meta_contract_guards.py`
|
||||
5. Contract-specific guards (`tests/test_service_boundaries.py`, `tests/test_ui_boundaries.py`, worker/evidence/media/error suites)
|
||||
|
||||
## Deterministic Audit Checks
|
||||
|
||||
### A. Contract Traceability
|
||||
- Each high-risk contract maps to at least one focused regression test file.
|
||||
- Missing mapping is a gap.
|
||||
|
||||
### B. Assertion Strength
|
||||
- Flag tests that only assert status code, non-null, or “no exception” without validating state transitions or persisted outcomes.
|
||||
- Prefer assertions on domain effects: DB rows, status changes, error categories, evidence writes, or emitted payload shape.
|
||||
|
||||
### C. Failure-Path Coverage
|
||||
- Critical paths must include negative-path tests (timeouts, provider errors, validation failures, cancellation paths, retries).
|
||||
- Happy-path-only coverage on critical modules is a gap.
|
||||
|
||||
### D. Redundancy and Noise
|
||||
- Detect near-duplicate tests asserting the same behavior at multiple layers with no extra signal.
|
||||
- Recommend canonical location (unit/integration) and prune overlaps.
|
||||
|
||||
### E. Mutation/Change Sensitivity
|
||||
- Prefer mutation testing for high-risk modules when practical.
|
||||
- If not run, identify tests likely to survive meaningful code mutations (low sensitivity).
|
||||
|
||||
### F. Drift Guards
|
||||
- Verify config/doc/instruction contracts have deterministic guards and are current.
|
||||
- Ensure settings/docs synchronization checks remain active.
|
||||
|
||||
## Evidence Standards
|
||||
|
||||
- Every finding must include concrete file paths and line ranges.
|
||||
- No speculative claims.
|
||||
- Distinguish clearly between:
|
||||
- **Confirmed ineffective tests**
|
||||
- **Likely weak tests (needs mutation/probe confirmation)**
|
||||
|
||||
## Output Format
|
||||
|
||||
Produce a Markdown report at `docs/reviews/<YYYY-MM-DD>-test-effectiveness.md`, using today's date.
|
||||
Like code review reports, it is a dated, non-canonical artifact: `docs/reviews/**` is not part of the
|
||||
canonical authority set.
|
||||
|
||||
```markdown
|
||||
# Test Effectiveness Audit Report
|
||||
|
||||
## 1. Executive Verdict
|
||||
- Effective / Effective with Conditions / Needs Remediation
|
||||
- Top risks to confidence
|
||||
|
||||
## 2. Contract Coverage Matrix
|
||||
| Contract | Guarding Tests | Signal Quality | Gap | Action |
|
||||
| :--- | :--- | :--- | :--- | :--- |
|
||||
|
||||
## 3. Weak/Redundant Test Findings
|
||||
| Finding ID | Location | Why Low-Signal | Risk | Recommendation |
|
||||
| :--- | :--- | :--- | :--- | :--- |
|
||||
|
||||
## 4. Prune/Strengthen Backlog
|
||||
| Task ID | Goal | Files | Acceptance Criteria | Validation |
|
||||
| :--- | :--- | :--- | :--- | :--- |
|
||||
|
||||
## 5. Confidence Recommendation
|
||||
- Go / Go with Conditions / No-Go for release confidence
|
||||
```
|
||||
|
||||
## Decision Rules
|
||||
|
||||
- Do not recommend deleting a test unless equivalent or stronger coverage is identified.
|
||||
- Prefer strengthening assertions before adding more tests.
|
||||
- Prioritize deterministic contract guards over broad snapshot-style tests.
|
||||
@@ -0,0 +1,50 @@
|
||||
name: Quality Gate
|
||||
|
||||
# Repository quality gate. Before this workflow existed, ruff, ty, and pytest were
|
||||
# enforced only by .pre-commit-config.yaml for developers who had run
|
||||
# `pre-commit install`.
|
||||
|
||||
on:
|
||||
push:
|
||||
pull_request:
|
||||
|
||||
jobs:
|
||||
gate:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Check out the commit under test
|
||||
uses: actions/checkout@v4
|
||||
|
||||
- name: Install uv
|
||||
run: |
|
||||
curl -LsSf https://astral.sh/uv/install.sh | sh
|
||||
echo "$HOME/.local/bin" >> "$GITHUB_PATH"
|
||||
|
||||
- name: Install dependencies from the lockfile
|
||||
# --locked fails if uv.lock has drifted from pyproject.toml, so a stale
|
||||
# lockfile is caught here rather than producing an untested dependency set.
|
||||
run: uv sync --locked
|
||||
|
||||
- name: Write placeholder configuration
|
||||
# Settings requires openrouter_api_key and 115 tests cannot construct
|
||||
# Settings without it. This is written to .env.production rather than exported
|
||||
# as an environment variable on purpose: the external tests guard on
|
||||
# os.getenv("OPENROUTER_API_KEY"), which reads the process environment and
|
||||
# not the file, so writing the file reproduces the local result exactly -
|
||||
# the 4 external tests skip instead of running against a fake key and
|
||||
# failing. Exporting it instead produces 3 failures.
|
||||
run: echo "OPENROUTER_API_KEY=ci-placeholder-not-a-real-key" > .env.production
|
||||
|
||||
- name: Lint and type check
|
||||
# Runs the hooks defined in .pre-commit-config.yaml instead of repeating
|
||||
# "ruff check" and "ty check" here. The commands then have one definition,
|
||||
# so the local and CI gates cannot drift apart.
|
||||
run: uv run pre-commit run --all-files --show-diff-on-failure
|
||||
|
||||
- name: Tests
|
||||
# Deliberately unfiltered, unlike the "-m 'not external'" form the guidance files
|
||||
# use for local runs. Tests marked "external" skip themselves when live-service
|
||||
# credentials are absent, so CI gets the same effective set plus a real run of any
|
||||
# external test whose credentials are configured. Not drift -- do not "fix" this to
|
||||
# match the local command without also giving those tests a way to run.
|
||||
run: uv run pytest
|
||||
+16
-4
@@ -11,13 +11,25 @@ wheels/
|
||||
|
||||
# Environment secrets
|
||||
.env
|
||||
.env.production
|
||||
|
||||
# SQLite database
|
||||
*.db
|
||||
|
||||
# Document images
|
||||
uploads/*
|
||||
# All data including db, backups, document images, photos, and logs:
|
||||
data/*
|
||||
data/backups/*
|
||||
data/documents/*
|
||||
data/logs/*
|
||||
data/photos/*
|
||||
data-local/*
|
||||
|
||||
# Local destructive-test backups
|
||||
.test-backups/
|
||||
|
||||
# Migration tests
|
||||
data-migration-test/*
|
||||
.migration-bundle/*
|
||||
|
||||
# Cloudflare tunnel local runtime files
|
||||
deploy/cloudflared/config.yml
|
||||
deploy/cloudflared/config.yaml
|
||||
deploy/cloudflared/credentials.json
|
||||
|
||||
@@ -0,0 +1,29 @@
|
||||
# Quality gate: `ruff check`, `ruff format --check`, and `ty check`
|
||||
# are blocking once known `ty` false positives are suppressed inline with rationale.
|
||||
#
|
||||
# Both tools are uv-managed dev dependencies and are not on PATH, so each entry must
|
||||
# go through `uv run`.
|
||||
repos:
|
||||
- repo: local
|
||||
hooks:
|
||||
- id: ruff
|
||||
name: ruff check
|
||||
entry: uv run ruff check
|
||||
language: system
|
||||
types_or: [python, pyi]
|
||||
require_serial: true
|
||||
- id: ruff-format
|
||||
name: ruff format check
|
||||
entry: uv run ruff format --check .
|
||||
language: system
|
||||
types_or: [python, pyi]
|
||||
pass_filenames: false
|
||||
require_serial: true
|
||||
- id: ty
|
||||
name: ty check
|
||||
entry: uv run ty check
|
||||
language: system
|
||||
types_or: [python, pyi]
|
||||
pass_filenames: false
|
||||
require_serial: true
|
||||
verbose: true
|
||||
@@ -0,0 +1,117 @@
|
||||
# AGENTS.md
|
||||
|
||||
Orientation for AI agents working in this repository. This file is a **router**, not a spec:
|
||||
it points at canonical authority and flags the traps that are expensive to discover by trial.
|
||||
Where this file and `docs/*` disagree, `docs/*` wins.
|
||||
|
||||
## What This Is
|
||||
|
||||
A document transcription system that preserves durable archival records (Documents, Sources,
|
||||
People) and executes page transcription asynchronously through vision/LLM providers. Its
|
||||
defining constraint is **evidence**: every machine attempt is recorded append-only with
|
||||
request/response provenance. Features that would lose, mutate, or obscure that history are
|
||||
wrong regardless of how convenient they are.
|
||||
|
||||
Stack: Python 3.12+ · FastAPI + NiceGUI · SQLModel/SQLAlchemy (SQLite-first, PostgreSQL-
|
||||
compatible) · Pydantic V2 · asyncio worker · OpenRouter adapter.
|
||||
|
||||
## Commands
|
||||
|
||||
This is a `uv` project. **Nothing is on `PATH`** — `ruff`, `ty`, and `pytest` all require
|
||||
`uv run`. Bare invocations fail with command-not-found.
|
||||
|
||||
```bash
|
||||
uv run ruff check . # lint (blocking in pre-commit)
|
||||
uv run ruff format --check . # format (blocking in pre-commit)
|
||||
uv run ty check # types (blocking in pre-commit)
|
||||
uv run pytest -q -m "not external" # default verification run
|
||||
```
|
||||
|
||||
`external` marks tests that hit live services; always exclude it unless explicitly asked.
|
||||
All four commands are expected to pass clean — there is no tolerated baseline of failures.
|
||||
If `ty` reports something, fix it or suppress it inline *with a rationale comment*; a bare
|
||||
`ignore` will not survive review.
|
||||
|
||||
## Authority Order
|
||||
|
||||
Resolve every question in this order, and stop at the first that answers it:
|
||||
|
||||
1. **`docs/*`** — canonical. Start at [`docs/index.md`](docs/index.md), which defines the
|
||||
reading order. `docs/invariant/*` holds cross-version rules that outlive any release.
|
||||
2. **`.github/instructions/*.md`** — active steering, auto-attached when you edit matching
|
||||
paths. Covers services, UI, providers, tests, error handling, and documentation sync.
|
||||
3. **`.github/skills/*`** — periodic audit procedures (code review, provenance, test
|
||||
effectiveness).
|
||||
4. **`tests/`** — deterministic enforcement. A guard test is the ground truth for whatever
|
||||
rule it encodes.
|
||||
|
||||
`.github/agents/` and `.github/prompts/` hold named workflows that are loaded only when
|
||||
invoked explicitly, so they never override the order above. They are how a review or audit
|
||||
is *started*, not a source of rules.
|
||||
|
||||
`docs/reviews/**` is **not** canonical. Those are dated, opinionated snapshots that were
|
||||
accurate when written and may since have been fixed, superseded, or found wrong.
|
||||
|
||||
## Layout
|
||||
|
||||
| Path | Role |
|
||||
| :--- | :--- |
|
||||
| `src/transcription/ui/**`, `api/**` | Interface. No direct persistence access. |
|
||||
| `src/transcription/services/**` | Domain logic and transaction ownership. |
|
||||
| `src/transcription/db/**` | Models and persistence. |
|
||||
| `src/transcription/providers/**` | Provider adapters; provider details stop here. |
|
||||
| `src/transcription/worker.py` | Asyncio worker loop. |
|
||||
| `tests/` | Includes boundary/contract guards, not just behavior tests. |
|
||||
|
||||
## Enforced Boundaries
|
||||
|
||||
These are not conventions — a test fails if you break them:
|
||||
|
||||
- **No service-to-service imports** (`test_service_boundaries.py`). Compose in the caller.
|
||||
- **No persistence access from pages/components** (`test_ui_boundaries.py`, allowlist-based).
|
||||
- **No hand-rolled error notifications in UI** — use the shared error presenter.
|
||||
- **No stringly-typed status literals** — use the enums (`test_model_contract_guards.py`).
|
||||
- **Attempt history is append-only** (`test_evidence_provenance.py`).
|
||||
- **`docs/schema.md` stays field-accurate** with `db/models.py`.
|
||||
- **Orphans are tracked, not tolerated** — `test_orphan_sweep.py` records each retained
|
||||
orphan with rationale in `KNOWN_ORPHANS`.
|
||||
|
||||
## Traps
|
||||
|
||||
Non-obvious things that have already caused real bugs here:
|
||||
|
||||
- **`AppError.message` vs `AppError.detail`.** `message` is user/API-facing and must stay
|
||||
generic — never put exception text or filesystem paths in it. `detail` is internal-only and
|
||||
is what reaches logs and `ExecutionAttempt.error_detail`. Putting root-cause data in
|
||||
`message` leaks; removing it from `detail` silently degrades provenance. See
|
||||
`docs/error_handling.md`.
|
||||
- **Two competing atomicity invariants in `services/workflows.py`.** Intermediate pages must
|
||||
commit individually (durability across a long multi-page job); the *final* page must commit
|
||||
atomically with the terminal job status. Collapsing the batch into one transaction satisfies
|
||||
the second and destroys the first. Both are guarded — `test_workflows_reliability.py` and
|
||||
`tests/integration/test_pipeline_atomicity.py`.
|
||||
- **Shared symbols have more consumers than the obvious one.** Before changing a model field,
|
||||
exception attribute, or helper return value, grep for every consumer including tests.
|
||||
Evidence and logging paths frequently read the same fields the UI does.
|
||||
- **`Tag` is owned by two services, on purpose.** Every other model has exactly one owning
|
||||
service, so the ownership rule reads as absolute — it isn't. `Tag` is a single table reached
|
||||
through two `RegistryService[Tag]` facades, `TagRegistry` (documents) and `PersonTagRegistry`
|
||||
(people), which count usage through `DocumentTag` and `PersonTag` respectively. Changing tag
|
||||
semantics through one facade silently changes the other. Consolidating them under one service
|
||||
is not a cleanup; it makes the other side a cross-aggregate writer.
|
||||
- **Import style:** ruff `isort` runs with `force-single-line = true`. One import per line.
|
||||
- **Latent defects have ordering constraints.** Some code is unreachable only because of a
|
||||
current setting or single-instance deployment. Fix it *before* the change that unblocks it,
|
||||
not after.
|
||||
|
||||
## Change Protocol
|
||||
|
||||
- **Write the failing test first** for behavioral fixes, and confirm it actually fails for the
|
||||
reason you think. Several bugs here were subtle enough that a test written afterward would
|
||||
have passed against the broken code.
|
||||
- **Update docs in the same change** when you alter a contract, behavior, or scope — see
|
||||
`.github/instructions/documentation-sync.instructions.md`.
|
||||
- **Do not commit unless asked.** Making a requested change is not consent to commit it.
|
||||
- **Do not push or open PRs on your own initiative.**
|
||||
- **Scope discipline:** fix what was asked plus what your change genuinely breaks. Pre-existing
|
||||
unrelated issues are a separate conversation.
|
||||
@@ -13,6 +13,8 @@ RUN uv sync --frozen --no-dev --no-install-project
|
||||
|
||||
COPY src ./src
|
||||
COPY prompts ./prompts
|
||||
COPY tools ./tools
|
||||
COPY deploy ./deploy
|
||||
RUN uv sync --frozen --no-dev
|
||||
|
||||
|
||||
@@ -27,12 +29,18 @@ ENV PYTHONDONTWRITEBYTECODE=1 \
|
||||
|
||||
WORKDIR /app
|
||||
|
||||
RUN apt-get update \
|
||||
&& apt-get install -y --no-install-recommends postgresql-client \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
RUN groupadd --system --gid 1001 appgroup \
|
||||
&& useradd --system --uid 1001 --gid appgroup --create-home appuser
|
||||
|
||||
COPY --from=builder /app/.venv /app/.venv
|
||||
COPY --from=builder /app/src /app/src
|
||||
COPY --from=builder /app/prompts /app/prompts
|
||||
COPY --from=builder /app/tools /app/tools
|
||||
COPY --from=builder /app/deploy /app/deploy
|
||||
|
||||
RUN mkdir -p /app/uploads /app/data \
|
||||
&& chown -R appuser:appgroup /app
|
||||
|
||||
@@ -22,28 +22,28 @@ uv sync
|
||||
|
||||
### 2) Configure environment
|
||||
|
||||
Create a `.env` file in the project root with the required OpenRouter API key:
|
||||
Create a `.env.production` file in the project root with the required OpenRouter API key:
|
||||
|
||||
```env
|
||||
OPENROUTER_API_KEY=your_openrouter_api_key
|
||||
```
|
||||
|
||||
Settings are read from CLI arguments first, then environment variables, then `.env`, then the defaults below.
|
||||
Settings are read from CLI arguments first, then environment variables, then `.env.production`, then the defaults below.
|
||||
|
||||
### Configuration Source Precedence
|
||||
|
||||
When the same setting is provided in multiple places, the value is chosen in this order (highest priority first):
|
||||
|
||||
1. CLI arguments (for example `--port 9999`)
|
||||
1. CLI arguments (for example `--port 8000`)
|
||||
2. Settings constructor arguments (used mainly in tests)
|
||||
3. Environment variables
|
||||
4. `.env` file values
|
||||
4. `.env.production` file values
|
||||
5. Model defaults in code
|
||||
|
||||
Practical examples:
|
||||
|
||||
- `--port 9999` overrides both `PORT=8000` in the shell and `PORT=7000` in `.env`.
|
||||
- `DATABASE__PATH=prod.db` in the shell overrides `DATABASE__PATH=dev.db` in `.env`.
|
||||
- `--port 8000` overrides both `PORT=8000` in the shell and `PORT=7000` in `.env.production`.
|
||||
- `DATABASE__PATH=prod.db` in the shell overrides `DATABASE__PATH=dev.db` in `.env.production`.
|
||||
|
||||
#### Server and runtime
|
||||
|
||||
@@ -54,6 +54,7 @@ Practical examples:
|
||||
| `LOG_LEVEL` | `info` | Uvicorn and application log level. |
|
||||
| `RELOAD` | `false` | Restart the development server when source files change. |
|
||||
| `ENVIRONMENT` | `development` | Runtime environment: `development`, `test`, or `production`. |
|
||||
| `RUN_EMBEDDED_WORKER` | `true` | Run worker loop inside web app process. Set `false` when using a dedicated worker service. |
|
||||
|
||||
#### Provider
|
||||
|
||||
@@ -76,6 +77,9 @@ DATABASE__PATH=app.db
|
||||
SQLITE_CHECK_SAME_THREAD=false
|
||||
UPLOAD_DIR=./uploads
|
||||
PROMPT_DIR=./prompts
|
||||
DEFAULT_PROMPT_NAME=transcribe_document.md
|
||||
# TRANSCRIPTION_TEMPERATURE=0.2 # range: 0.0-2.0
|
||||
# TRANSCRIPTION_TOP_P=0.9 # range: 0.0-1.0
|
||||
```
|
||||
|
||||
For PostgreSQL:
|
||||
@@ -89,7 +93,7 @@ DATABASE__USER=postgres
|
||||
DATABASE__PASSWORD=change-me
|
||||
```
|
||||
|
||||
This uses Pydantic nested settings (`env_nested_delimiter='__'`) and avoids JSON blobs in `.env`. A top-level `DATABASE={...}` JSON value is still supported as a fallback, and nested keys such as `DATABASE__PATH` take precedence over conflicting JSON keys.
|
||||
This uses Pydantic nested settings (`env_nested_delimiter='__'`) and avoids JSON blobs in env files. A top-level `DATABASE={...}` JSON value is still supported as a fallback, and nested keys such as `DATABASE__PATH` take precedence over conflicting JSON keys.
|
||||
|
||||
`BOOTSTRAP_SCHEMA_ON_STARTUP` creates missing tables when the app starts. When unset, it is enabled in `development` and `test`, and disabled in `production`; set it explicitly to override that policy. `SQLITE_CHECK_SAME_THREAD` defaults to `false`.
|
||||
|
||||
@@ -107,18 +111,45 @@ WORKER_FAIL_ON_FINISH_REASON_LENGTH=false
|
||||
### 3) Run the app
|
||||
|
||||
```bash
|
||||
uv run python -m transcription --port 9999 --reload --database.driver sqlite --bootstrap-schema-on-startup
|
||||
uv run python -m transcription --port 8000 --reload --database.driver sqlite --bootstrap-schema-on-startup
|
||||
```
|
||||
|
||||
This starts the development server with SQLite, creates missing tables, and enables automatic reload. Run `uv run python -m transcription --help` for all CLI options; CLI names use kebab case and nested database options use dot notation, such as `--database.path ./data/transcription.db`.
|
||||
|
||||
### 4) Open in browser
|
||||
|
||||
- GUI: [http://localhost:9999/ui](http://localhost:9999/ui)
|
||||
- Health check: [http://localhost:9999/healthz](http://localhost:9999/healthz)
|
||||
- GUI: [http://localhost:8000/ui](http://localhost:8000/ui)
|
||||
- Health check: [http://localhost:8000/healthz](http://localhost:8000/healthz)
|
||||
|
||||
Replace `localhost` with the server's hostname or IP address when connecting from another machine.
|
||||
|
||||
## Production stack (Phase 1)
|
||||
|
||||
Use the production compose profile for split app/worker deployment with PostgreSQL and Cloudflare Tunnel:
|
||||
|
||||
```bash
|
||||
copy .env.production.example .env.production
|
||||
docker compose -f docker-compose.production.yml up -d --build
|
||||
```
|
||||
|
||||
Services:
|
||||
|
||||
- `app`: FastAPI + NiceGUI runtime (`RUN_EMBEDDED_WORKER=false`)
|
||||
- `worker`: standalone queue processor (`python -m transcription.worker_service`)
|
||||
- `postgres`: primary datastore
|
||||
- `cloudflared`: tunnel client using mounted ingress config + `CLOUDFLARE_TUNNEL_TOKEN`
|
||||
|
||||
Operational defaults in the production compose file:
|
||||
|
||||
- worker healthcheck is disabled (the worker process has no HTTP `/healthz` endpoint)
|
||||
- cloudflared is pinned to HTTP/2 with explicit DNS resolvers (`1.1.1.1`, `1.0.0.1`) for restricted LXC/container DNS environments
|
||||
|
||||
Cloudflare setup files:
|
||||
|
||||
1. `copy deploy\cloudflared\config.yml.example deploy\cloudflared\config.yml`
|
||||
2. set `CLOUDFLARE_TUNNEL_TOKEN` in `.env.production`
|
||||
3. update ingress hostnames in `deploy\cloudflared\config.yml`
|
||||
|
||||
## How to navigate the GUI
|
||||
|
||||
- **Upload page** (`/ui`)
|
||||
@@ -138,21 +169,32 @@ Replace `localhost` with the server's hostname or IP address when connecting fro
|
||||
|
||||
## Prompt artifacts
|
||||
|
||||
Prompt files are stored in `prompts/` and loaded from `PROMPT_DIR` (default: `./prompts`).
|
||||
Prompt files are stored directly in `PROMPT_DIR` (default: `./prompts`). `DEFAULT_PROMPT_NAME` must be a filename,
|
||||
not a path. Each job snapshots the validated prompt text, SHA-256 hash, and sampling values for reproducibility.
|
||||
|
||||
The canonical MVP prompt is:
|
||||
- `prompts/transcribe_document.md`
|
||||
|
||||
## Database migration workflow
|
||||
|
||||
Schema upgrades use an explicit export/import rebuild flow (no runtime legacy write compatibility).
|
||||
See `docs/data_migration.md` for commands and cutover steps.
|
||||
|
||||
## Backup and restore workflow
|
||||
|
||||
Production backup/restore (PostgreSQL + uploads + deployment config) steps are documented in `docs/backup_restore.md`.
|
||||
|
||||
## Destructive test procedure (with data backup)
|
||||
|
||||
AI execution policy: before running any unit tests, create a backup of `./data` first. After tests succeed, always pause and ask whether to restore now.
|
||||
AI execution policy: before the first unit-test run in a test/fix cycle, create one backup of `./data`. Reuse that same backup for every subsequent test run in the cycle. After tests succeed, always pause and ask whether to restore now.
|
||||
|
||||
Use the cross-platform Python wrapper below whenever an AI agent runs tests against this repository.
|
||||
|
||||
1. Create backup of `./data`.
|
||||
1. Create one backup of `./data` and mark it as the active test-cycle backup.
|
||||
2. Run your test command.
|
||||
3. On success, always prompt whether to restore now (do not auto-restore unless explicitly approved).
|
||||
4. On failure, keep backup and current state for inspection.
|
||||
3. On failure, fix the errors and run the wrapper again; it reuses the active backup and never backs up post-test data.
|
||||
4. On success, always prompt whether to restore now (do not auto-restore unless explicitly approved).
|
||||
5. Close the cycle only by restoring the active backup or explicitly accepting the current data.
|
||||
|
||||
Preflight behavior:
|
||||
|
||||
@@ -183,10 +225,18 @@ uv run python tools/run_destructive_tests.py --skip-restore-prompt -- pytest
|
||||
|
||||
This keeps both the current post-test state and the backup, so restore can be decided explicitly later.
|
||||
|
||||
Repeated wrapper invocations reuse the backup recorded in `.test-backups/.active-backup`. If that backup is missing, the wrapper stops rather than creating a replacement from potentially destructive post-test data.
|
||||
|
||||
### Restore later from a saved backup
|
||||
|
||||
```bash
|
||||
uv run python tools/run_destructive_tests.py --restore-from data-backup-YYYYMMDD-HHMMSS
|
||||
```
|
||||
|
||||
To keep the current data and close the active cycle without restoring:
|
||||
|
||||
```bash
|
||||
uv run python tools/run_destructive_tests.py --accept-current-data
|
||||
```
|
||||
|
||||
Backups are stored in `.test-backups/` and ignored by git.
|
||||
|
||||
@@ -0,0 +1,69 @@
|
||||
#!/usr/bin/env sh
|
||||
set -eu
|
||||
|
||||
BACKUP_DIR="${BACKUP_DIR:-./data/backups}"
|
||||
RETENTION_DAYS="${BACKUP_RETENTION_DAYS:-14}"
|
||||
UPLOAD_DIR="${UPLOAD_DIR:-/app/uploads}"
|
||||
PROMPT_DIR="${PROMPT_DIR:-/app/prompts}"
|
||||
DATABASE_DRIVER="${DATABASE__DRIVER:-postgres}"
|
||||
DATABASE_HOST="${DATABASE__HOST:-postgres}"
|
||||
DATABASE_PORT="${DATABASE__PORT:-5432}"
|
||||
DATABASE_NAME="${DATABASE__DATABASE:-}"
|
||||
DATABASE_USER="${DATABASE__USER:-}"
|
||||
DATABASE_PASSWORD="${DATABASE__PASSWORD:-}"
|
||||
|
||||
timestamp="$(date -u +%Y%m%d-%H%M%S)"
|
||||
postgres_file="postgres-${timestamp}.dump"
|
||||
manifest_file="backup-${timestamp}.manifest"
|
||||
|
||||
mkdir -p "${BACKUP_DIR}"
|
||||
if [ "${DATABASE_DRIVER}" != "postgres" ]; then
|
||||
echo "create_postgres_backup.sh requires DATABASE__DRIVER=postgres." >&2
|
||||
exit 1
|
||||
fi
|
||||
if [ -z "${DATABASE_NAME}" ] || [ -z "${DATABASE_USER}" ] || [ -z "${DATABASE_PASSWORD}" ]; then
|
||||
echo "DATABASE__DATABASE, DATABASE__USER, and DATABASE__PASSWORD must be set." >&2
|
||||
exit 1
|
||||
fi
|
||||
if ! command -v pg_dump >/dev/null 2>&1; then
|
||||
echo "pg_dump is not installed in this environment." >&2
|
||||
exit 1
|
||||
fi
|
||||
PGPASSWORD="${DATABASE_PASSWORD}" pg_dump \
|
||||
-h "${DATABASE_HOST}" \
|
||||
-p "${DATABASE_PORT}" \
|
||||
-U "${DATABASE_USER}" \
|
||||
-d "${DATABASE_NAME}" \
|
||||
-Fc \
|
||||
> "${BACKUP_DIR}/${postgres_file}"
|
||||
|
||||
cat > "${BACKUP_DIR}/${manifest_file}" <<EOF
|
||||
created_at_utc=${timestamp}
|
||||
postgres_dump=${postgres_file}
|
||||
backup_dir=${BACKUP_DIR}
|
||||
uploads_backup_dir=${BACKUP_DIR}/uploads
|
||||
prompts_backup_dir=${BACKUP_DIR}/prompts
|
||||
EOF
|
||||
|
||||
find "${BACKUP_DIR}" -type f \( \
|
||||
-name 'postgres-*.dump' -o \
|
||||
-name 'backup-*.manifest' \
|
||||
\) -mtime +"${RETENTION_DAYS}" -delete
|
||||
|
||||
uploads_backup_dir="${BACKUP_DIR}/uploads"
|
||||
mkdir -p "${uploads_backup_dir}"
|
||||
if [ -d "${UPLOAD_DIR}" ]; then
|
||||
cp -an "${UPLOAD_DIR}/." "${uploads_backup_dir}/"
|
||||
fi
|
||||
|
||||
prompts_backup_dir="${BACKUP_DIR}/prompts"
|
||||
mkdir -p "${prompts_backup_dir}"
|
||||
if [ -d "${PROMPT_DIR}" ]; then
|
||||
cp -a "${PROMPT_DIR}/." "${prompts_backup_dir}/"
|
||||
fi
|
||||
|
||||
echo "Created backup set:"
|
||||
echo " ${BACKUP_DIR}/${postgres_file}"
|
||||
echo " ${BACKUP_DIR}/${manifest_file}"
|
||||
echo " ${uploads_backup_dir}/ (incremental uploads mirror)"
|
||||
echo " ${prompts_backup_dir}/ (prompts mirror)"
|
||||
@@ -0,0 +1,36 @@
|
||||
#!/usr/bin/env sh
|
||||
set -eu
|
||||
|
||||
# Example only. Copy to a local script and replace placeholder values.
|
||||
# Do NOT commit secrets.
|
||||
|
||||
SHARE="//nas-host-or-ip/share-name"
|
||||
MOUNT_POINT="/mnt/nas-backups"
|
||||
CREDENTIALS_FILE="/etc/samba/credentials/nas-share-credentials"
|
||||
USERNAME="replace-with-nas-user"
|
||||
PASSWORD="replace-with-nas-password"
|
||||
|
||||
if [ "${USERNAME}" = "replace-with-nas-user" ] || [ "${PASSWORD}" = "replace-with-nas-password" ]; then
|
||||
echo "Edit USERNAME and PASSWORD placeholders before running this script."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
mkdir -p "${MOUNT_POINT}"
|
||||
mkdir -p "$(dirname "${CREDENTIALS_FILE}")"
|
||||
|
||||
cat > "${CREDENTIALS_FILE}" <<'EOF'
|
||||
username=__USERNAME__
|
||||
password=__PASSWORD__
|
||||
EOF
|
||||
sed -i "s|__USERNAME__|${USERNAME}|g" "${CREDENTIALS_FILE}"
|
||||
sed -i "s|__PASSWORD__|${PASSWORD}|g" "${CREDENTIALS_FILE}"
|
||||
chmod 600 "${CREDENTIALS_FILE}"
|
||||
|
||||
mount -t cifs "${SHARE}" "${MOUNT_POINT}" \
|
||||
-o "credentials=${CREDENTIALS_FILE},vers=3.0,iocharset=utf8,uid=0,gid=0,file_mode=0600,dir_mode=0700"
|
||||
|
||||
echo ""
|
||||
echo "Mounted ${SHARE} at ${MOUNT_POINT}"
|
||||
echo ""
|
||||
echo "To persist across reboot, add this line to /etc/fstab:"
|
||||
echo "${SHARE} ${MOUNT_POINT} cifs credentials=${CREDENTIALS_FILE},vers=3.0,iocharset=utf8,uid=0,gid=0,file_mode=0600,dir_mode=0700,_netdev,nofail,x-systemd.automount 0 0"
|
||||
@@ -0,0 +1,86 @@
|
||||
#!/usr/bin/env sh
|
||||
set -eu
|
||||
|
||||
if [ "$#" -lt 1 ]; then
|
||||
echo "Usage: $0 <path-to-postgres-dump>"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
dump_file="$1"
|
||||
COMPOSE_FILE="${COMPOSE_FILE:-docker-compose.production.yml}"
|
||||
ENV_FILE="${ENV_FILE:-.env.production}"
|
||||
SYNOLOGY_BACKUP_DIR="${SYNOLOGY_BACKUP_DIR:-}"
|
||||
|
||||
if [ ! -f "${dump_file}" ]; then
|
||||
echo "Backup file not found: ${dump_file}"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
backup_dir="$(dirname "${dump_file}")"
|
||||
backup_name="$(basename "${dump_file}")"
|
||||
timestamp="$(printf '%s' "${backup_name}" | sed -n 's/^postgres-\([0-9]\{8\}-[0-9]\{6\}\)\.dump$/\1/p')"
|
||||
uploads_file=""
|
||||
config_file=""
|
||||
|
||||
# If SYNOLOGY_BACKUP_DIR wasn't exported in the shell, read it from ENV_FILE.
|
||||
if [ -z "${SYNOLOGY_BACKUP_DIR}" ] && [ -f "${ENV_FILE}" ]; then
|
||||
SYNOLOGY_BACKUP_DIR="$(
|
||||
sed -n 's/^SYNOLOGY_BACKUP_DIR=//p' "${ENV_FILE}" | tail -n 1
|
||||
)"
|
||||
fi
|
||||
|
||||
if [ -n "${timestamp}" ]; then
|
||||
# Legacy local full-archive naming.
|
||||
candidate_uploads="${backup_dir}/uploads-${timestamp}.tar.gz"
|
||||
candidate_config="${backup_dir}/config-${timestamp}.tar.gz"
|
||||
if [ -f "${candidate_uploads}" ]; then
|
||||
uploads_file="${candidate_uploads}"
|
||||
fi
|
||||
if [ -f "${candidate_config}" ]; then
|
||||
config_file="${candidate_config}"
|
||||
fi
|
||||
fi
|
||||
|
||||
docker compose --env-file "${ENV_FILE}" -f "${COMPOSE_FILE}" exec -T postgres sh -lc \
|
||||
"PGPASSWORD=\"\$POSTGRES_PASSWORD\" psql -U \"\$POSTGRES_USER\" -d postgres -c \"DROP DATABASE IF EXISTS \\\"\$POSTGRES_DB\\\";\""
|
||||
|
||||
docker compose --env-file "${ENV_FILE}" -f "${COMPOSE_FILE}" exec -T postgres sh -lc \
|
||||
"PGPASSWORD=\"\$POSTGRES_PASSWORD\" psql -U \"\$POSTGRES_USER\" -d postgres -c \"CREATE DATABASE \\\"\$POSTGRES_DB\\\";\""
|
||||
|
||||
cat "${dump_file}" | docker compose --env-file "${ENV_FILE}" -f "${COMPOSE_FILE}" exec -T postgres sh -lc \
|
||||
"PGPASSWORD=\"\$POSTGRES_PASSWORD\" pg_restore -U \"\$POSTGRES_USER\" -d \"\$POSTGRES_DB\" --clean --if-exists --no-owner --no-privileges"
|
||||
|
||||
if [ -n "${SYNOLOGY_BACKUP_DIR}" ] && [ -d "${SYNOLOGY_BACKUP_DIR}/uploads" ]; then
|
||||
docker compose --env-file "${ENV_FILE}" -f "${COMPOSE_FILE}" run --rm --no-deps \
|
||||
-v "${SYNOLOGY_BACKUP_DIR}:/backup" \
|
||||
--entrypoint sh app -lc \
|
||||
"mkdir -p /app/uploads/documents /app/uploads/photos && \
|
||||
find /app/uploads/documents -mindepth 1 -delete && \
|
||||
find /app/uploads/photos -mindepth 1 -delete && \
|
||||
if [ -d /backup/uploads/documents ]; then cp -a /backup/uploads/documents/. /app/uploads/documents/; fi && \
|
||||
if [ -d /backup/uploads/photos ]; then cp -a /backup/uploads/photos/. /app/uploads/photos/; fi && \
|
||||
if [ -f /backup/uploads/homepage.md ]; then cp /backup/uploads/homepage.md /app/uploads/homepage.md; else rm -f /app/uploads/homepage.md; fi"
|
||||
elif [ -n "${uploads_file}" ]; then
|
||||
# Legacy local full-archive restore.
|
||||
cat "${uploads_file}" | docker compose --env-file "${ENV_FILE}" -f "${COMPOSE_FILE}" run --rm --no-deps --entrypoint sh app -lc \
|
||||
"mkdir -p /app/uploads && find /app/uploads -mindepth 1 -delete && tar -xzf - -C /app/uploads"
|
||||
fi
|
||||
|
||||
if [ -n "${SYNOLOGY_BACKUP_DIR}" ] && [ -n "${timestamp}" ] && [ -f "${SYNOLOGY_BACKUP_DIR}/config-${timestamp}.tar.gz" ]; then
|
||||
tar -xzf "${SYNOLOGY_BACKUP_DIR}/config-${timestamp}.tar.gz" -C .
|
||||
elif [ -n "${config_file}" ]; then
|
||||
# Legacy local full-archive restore.
|
||||
tar -xzf "${config_file}" -C .
|
||||
fi
|
||||
|
||||
echo "Restore complete from: ${dump_file}"
|
||||
if [ -n "${SYNOLOGY_BACKUP_DIR}" ] && [ -d "${SYNOLOGY_BACKUP_DIR}/uploads" ]; then
|
||||
echo "Restored uploads mirror from: ${SYNOLOGY_BACKUP_DIR}/uploads"
|
||||
elif [ -n "${uploads_file}" ]; then
|
||||
echo "Restored uploads archive: ${uploads_file}"
|
||||
fi
|
||||
if [ -n "${SYNOLOGY_BACKUP_DIR}" ] && [ -n "${timestamp}" ] && [ -f "${SYNOLOGY_BACKUP_DIR}/config-${timestamp}.tar.gz" ]; then
|
||||
echo "Restored config archive: ${SYNOLOGY_BACKUP_DIR}/config-${timestamp}.tar.gz"
|
||||
elif [ -n "${config_file}" ]; then
|
||||
echo "Restored config archive: ${config_file}"
|
||||
fi
|
||||
@@ -0,0 +1,14 @@
|
||||
ingress:
|
||||
# Primary transcription app endpoint.
|
||||
- hostname: transcription.example.com
|
||||
service: http://app:8000
|
||||
|
||||
# Optional: generic remote access endpoints for other internal services.
|
||||
# Replace hostnames and targets for your LAN.
|
||||
- hostname: homeassistant.example.com
|
||||
service: http://192.168.1.50:8123
|
||||
- hostname: pihole.example.com
|
||||
service: http://192.168.1.60:80
|
||||
|
||||
# Required catch-all.
|
||||
- service: http_status:404
|
||||
@@ -0,0 +1,70 @@
|
||||
services:
|
||||
app:
|
||||
build:
|
||||
context: .
|
||||
dockerfile: Dockerfile
|
||||
image: transcription:prod
|
||||
env_file:
|
||||
- .env.production
|
||||
environment:
|
||||
RUN_EMBEDDED_WORKER: "false"
|
||||
RUNTIME_SETTINGS_ENV_FILE: "/app/.env.production"
|
||||
depends_on:
|
||||
postgres:
|
||||
condition: service_healthy
|
||||
ports:
|
||||
- "8000:8000"
|
||||
volumes:
|
||||
- app_uploads:/app/uploads
|
||||
- app_data:/app/data
|
||||
- ./backup:/backup
|
||||
- ./.env.production:/app/.env.production
|
||||
- ./prompts:/app/prompts:ro
|
||||
restart: unless-stopped
|
||||
healthcheck:
|
||||
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/healthz')"]
|
||||
interval: 30s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
start_period: 20s
|
||||
|
||||
worker:
|
||||
image: transcription:prod
|
||||
env_file:
|
||||
- .env.production
|
||||
command: ["python", "-m", "transcription.worker_service"]
|
||||
depends_on:
|
||||
postgres:
|
||||
condition: service_healthy
|
||||
volumes:
|
||||
- app_uploads:/app/uploads
|
||||
- app_data:/app/data
|
||||
- ./backup:/backup
|
||||
- ./.env.production:/app/.env.production:ro
|
||||
- ./prompts:/app/prompts:ro
|
||||
healthcheck:
|
||||
disable: true
|
||||
restart: unless-stopped
|
||||
|
||||
postgres:
|
||||
image: postgres:16-alpine
|
||||
env_file:
|
||||
- .env.production
|
||||
environment:
|
||||
POSTGRES_DB: ${DATABASE__DATABASE}
|
||||
POSTGRES_USER: ${DATABASE__USER}
|
||||
POSTGRES_PASSWORD: ${DATABASE__PASSWORD}
|
||||
volumes:
|
||||
- postgres_data:/var/lib/postgresql/data
|
||||
restart: unless-stopped
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "pg_isready -U $$POSTGRES_USER -d $$POSTGRES_DB"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 10
|
||||
start_period: 10s
|
||||
|
||||
volumes:
|
||||
postgres_data:
|
||||
app_uploads:
|
||||
app_data:
|
||||
+5
-2
@@ -5,9 +5,12 @@ services:
|
||||
dockerfile: Dockerfile
|
||||
container_name: transcription-app
|
||||
env_file:
|
||||
- .env
|
||||
- .env.production
|
||||
environment:
|
||||
DATABASE_URL: sqlite:////app/data/transcription.db
|
||||
# Database configuration uses nested settings names (env_nested_delimiter="__").
|
||||
# DATABASE_URL is NOT read by the application and must not be used here.
|
||||
DATABASE__DRIVER: sqlite
|
||||
DATABASE__PATH: /app/data/transcription.db
|
||||
UPLOAD_DIR: /app/uploads
|
||||
PROMPT_DIR: /app/prompts
|
||||
ports:
|
||||
|
||||
-219
@@ -1,219 +0,0 @@
|
||||
# API Questions & Answers
|
||||
|
||||
## Question
|
||||
```
|
||||
what is v4_documents.py?
|
||||
```
|
||||
|
||||
## Response
|
||||
v4_documents.py is the new additive V4 API router I added in Phase 4.
|
||||
|
||||
What it does:
|
||||
- Defines V4 endpoints under /api/v4 for registry-driven document/relationship workflows.
|
||||
- Exposes read endpoints for catalogs:
|
||||
- GET /api/v4/document-types
|
||||
- GET /api/v4/person-roles
|
||||
- Exposes write endpoints for code-or-id updates:
|
||||
- PUT /api/v4/documents/{document_id}/type
|
||||
- POST /api/v4/documents/{document_id}/people
|
||||
- PATCH /api/v4/document-people/{document_person_id}
|
||||
- DELETE /api/v4/document-people/{document_person_id}
|
||||
- Uses DocumentService methods as the backend (no direct DB access in the route layer).
|
||||
- Returns typed request/response models for backward-compatible, additive API evolution.
|
||||
|
||||
It is registered in app.py, so it is live when the app starts.
|
||||
|
||||
If you want, I can walk endpoint-by-endpoint with examples of request/response payloads.
|
||||
|
||||
---
|
||||
|
||||
## Question
|
||||
```
|
||||
Please do
|
||||
```
|
||||
|
||||
## Response
|
||||
Absolutely. Here is a quick endpoint-by-endpoint walkthrough of v4_documents.py, with practical request/response examples.
|
||||
|
||||
**What This Router Is**
|
||||
- Prefix: /api/v4
|
||||
- Purpose: additive V4 contract for:
|
||||
1. Registry catalogs (document types, person roles)
|
||||
2. Role-aware document-person linking
|
||||
3. Code-or-id document type assignment
|
||||
- It delegates to DocumentService and does not access DB sessions directly in route handlers.
|
||||
|
||||
**Dependency Wiring**
|
||||
- Route dependency: get_document_service in v4_documents.py
|
||||
- Behavior:
|
||||
1. Uses app.state.services.documents when running in the app lifecycle
|
||||
2. Falls back to a direct DocumentService() instance if needed (useful in isolated tests)
|
||||
|
||||
**1) List Document Types**
|
||||
- Endpoint: GET /api/v4/document-types
|
||||
- Query param: active_only=true by default
|
||||
- Backing service call: list_document_types(active_only=...)
|
||||
- Response shape:
|
||||
[
|
||||
{
|
||||
"id": "uuid",
|
||||
"code": "letter",
|
||||
"label": "Letter",
|
||||
"is_active": true,
|
||||
"sort_order": 0
|
||||
}
|
||||
]
|
||||
|
||||
**2) List Person Roles**
|
||||
- Endpoint: GET /api/v4/person-roles
|
||||
- Query param: active_only=true by default
|
||||
- Backing service call: list_person_roles(active_only=...)
|
||||
- Response shape:
|
||||
[
|
||||
{
|
||||
"id": "uuid",
|
||||
"code": "author",
|
||||
"label": "Author",
|
||||
"is_active": true
|
||||
}
|
||||
]
|
||||
|
||||
**3) Set Document Type (By Id Or Code)**
|
||||
- Endpoint: PUT /api/v4/documents/{document_id}/type
|
||||
- Body options:
|
||||
1. document_type_id
|
||||
2. document_type_code
|
||||
- Rule: provide one selector, not both
|
||||
- Example request:
|
||||
{
|
||||
"document_type_code": "record"
|
||||
}
|
||||
- Example response:
|
||||
{
|
||||
"document_id": "uuid",
|
||||
"document_type_id": "uuid",
|
||||
"document_type_code": "record"
|
||||
}
|
||||
|
||||
**4) List Document-Person Links For A Document**
|
||||
- Endpoint: GET /api/v4/documents/{document_id}/people
|
||||
- Backing service call: list_document_people(document_id=...)
|
||||
- Example response:
|
||||
{
|
||||
"document_id": "uuid",
|
||||
"links": [
|
||||
{
|
||||
"id": "uuid",
|
||||
"document_id": "uuid",
|
||||
"person_id": "uuid",
|
||||
"role_id": "uuid",
|
||||
"role_code": "author",
|
||||
"person_name": "Zenna Cochran"
|
||||
}
|
||||
]
|
||||
}
|
||||
|
||||
**5) Add A Document-Person Link**
|
||||
- Endpoint: POST /api/v4/documents/{document_id}/people
|
||||
- Body fields:
|
||||
1. person_id required
|
||||
2. role_id optional
|
||||
3. role_code optional
|
||||
- If duplicate document/person/role link is attempted, returns conflict envelope via shared error handlers.
|
||||
- Example request:
|
||||
{
|
||||
"person_id": "uuid",
|
||||
"role_code": "mentioned"
|
||||
}
|
||||
- Example success response:
|
||||
{
|
||||
"id": "uuid",
|
||||
"document_id": "uuid",
|
||||
"person_id": "uuid",
|
||||
"role_id": "uuid",
|
||||
"role_code": "mentioned",
|
||||
"person_name": "Ada Lovelace"
|
||||
}
|
||||
|
||||
**6) Update Link Role**
|
||||
- Endpoint: PATCH /api/v4/document-people/{document_person_id}
|
||||
- Body options:
|
||||
1. role_id
|
||||
2. role_code
|
||||
- Rule: provide one selector, not both
|
||||
- Example request:
|
||||
{
|
||||
"role_code": "recipient"
|
||||
}
|
||||
- Example response:
|
||||
{
|
||||
"id": "uuid",
|
||||
"document_id": "uuid",
|
||||
"person_id": "uuid",
|
||||
"role_id": "uuid",
|
||||
"role_code": "recipient",
|
||||
"person_name": "Ada Lovelace"
|
||||
}
|
||||
|
||||
**7) Delete Link**
|
||||
- Endpoint: DELETE /api/v4/document-people/{document_person_id}
|
||||
- Response: HTTP 204 No Content
|
||||
|
||||
**Error Behavior**
|
||||
- Uses shared registration in errors.py
|
||||
- So AppError categories map to consistent envelope + status codes:
|
||||
1. validation_error -> 400
|
||||
2. not_found_error -> 404
|
||||
3. conflict_error -> 409
|
||||
4. internal_unexpected_error -> 500
|
||||
|
||||
**Where It Gets Mounted**
|
||||
- Included in app factory at app.py
|
||||
- That makes all /api/v4 endpoints available during normal app startup.
|
||||
|
||||
If you want, next I can generate a compact curl-style smoke checklist for all 7 endpoints so you can test them in under 2 minutes.
|
||||
|
||||
---
|
||||
|
||||
## Question
|
||||
```
|
||||
Should there be one of these for people, jobs, and sources too?
|
||||
```
|
||||
|
||||
## Response
|
||||
|
||||
Short answer: yes, but not all at once.
|
||||
|
||||
Given your V4 scope, the strongest split is:
|
||||
|
||||
1. Keep v4_documents.py for document-centric workflows
|
||||
This file already owns:
|
||||
- document type catalog/use
|
||||
- document-person role links
|
||||
- document-level relationship mutations
|
||||
|
||||
2. Add a V4 people router next
|
||||
Recommended new file:
|
||||
- src/transcription/api/v4_people.py
|
||||
|
||||
Why:
|
||||
- You likely need person-centric queries and updates that are not naturally document-scoped
|
||||
- It keeps API boundaries clearer as UI grows
|
||||
|
||||
3. Do not rush a separate jobs/sources V4 router unless contracts are diverging
|
||||
For now, jobs and sources are still tightly tied to the existing transcription pipeline behavior. A new v4_jobs or v4_sources router is useful only when you need genuinely new contract shape, such as:
|
||||
- job lifecycle contract changes
|
||||
- source revision/audit contract changes
|
||||
- new pagination/filtering semantics incompatible with current endpoints
|
||||
|
||||
Practical rule of thumb:
|
||||
1. New router if the domain has new V4 semantics and independent lifecycle
|
||||
2. Keep existing router if only additive fields are needed
|
||||
|
||||
Suggested target structure:
|
||||
1. v4_documents.py
|
||||
2. src/transcription/api/v4_people.py
|
||||
3. optional later: src/transcription/api/v4_jobs.py
|
||||
4. optional later: src/transcription/api/v4_sources.py
|
||||
|
||||
If you want, I can scaffold v4_people next with additive endpoints that mirror your current service capabilities and keep error envelope behavior consistent.
|
||||
@@ -0,0 +1,230 @@
|
||||
# System Architecture (Current Baseline: V6.1)
|
||||
|
||||
This document defines the current V6.1 architecture baseline.
|
||||
|
||||
## Architecture Objectives
|
||||
|
||||
- Preserve durable archival records for Documents, Sources, People, and processing runs.
|
||||
- Execute page transcription asynchronously with bounded worker behavior.
|
||||
- Preserve append-only machine-attempt evidence with request/response provenance.
|
||||
- Keep UI, API, service, persistence, and provider boundaries explicit and testable.
|
||||
|
||||
## Technical Stack
|
||||
|
||||
- **Runtime:** Python 3.12+
|
||||
- **Web application:** FastAPI + NiceGUI
|
||||
- **Persistence:** SQLModel / SQLAlchemy — PostgreSQL in production, SQLite for local development and tests
|
||||
- **Validation and settings:** Pydantic V2 + pydantic-settings
|
||||
- **Concurrency:** asyncio worker loop
|
||||
- **Provider integration:** OpenRouter adapter behind provider interface
|
||||
- **Deployment:** Docker Compose (app, worker, PostgreSQL, Cloudflare Tunnel)
|
||||
- **Quality and tests:** Ruff, ty, pytest, pytest-asyncio
|
||||
|
||||
## Runtime Topology
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
U[Browser User] --> A[FastAPI + NiceGUI App]
|
||||
A --> W[Asyncio Worker]
|
||||
A --> DB[(PostgreSQL / SQLite)]
|
||||
W --> P[Provider Adapter]
|
||||
W --> DB
|
||||
```
|
||||
|
||||
The worker loop drains two queues in the same pass: queued transcription Jobs and queued
|
||||
`MaintenanceRun` records. When neither has work, it idles.
|
||||
|
||||
### Production deployment
|
||||
|
||||
Production runs as a Docker Compose stack with the app and worker as separate services, so the app
|
||||
process runs with `RUN_EMBEDDED_WORKER=false` and the worker process owns queue draining. Local
|
||||
development runs a single process with the worker embedded.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
I[Internet] --> CF[cloudflared tunnel + Access]
|
||||
CF --> APP[app service]
|
||||
APP --> PG[(postgres service)]
|
||||
WK[worker service] --> PG
|
||||
WK --> PROV[OpenRouter]
|
||||
```
|
||||
|
||||
Deployment, rollback, and recovery procedures are in [Production Runbook](production-runbook.md);
|
||||
backup configuration and restore are in [Backup and Restore](backup_restore.md).
|
||||
|
||||
## Layered Boundaries
|
||||
|
||||
### Interface Layer
|
||||
|
||||
- `src/transcription/ui/**`
|
||||
- `src/transcription/api/**`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- Route registration, page orchestration, presentation adapters.
|
||||
- Structured user messaging through shared error presenter.
|
||||
- No direct persistence access from pages/components.
|
||||
|
||||
### Service and Orchestration Layer
|
||||
|
||||
Aggregate services:
|
||||
|
||||
- `src/transcription/services/documents.py`
|
||||
- `src/transcription/services/people.py`
|
||||
- `src/transcription/services/jobs.py`
|
||||
- `src/transcription/services/sources.py`
|
||||
- `src/transcription/services/photos.py`
|
||||
- `src/transcription/services/maintenance.py`
|
||||
- `src/transcription/services/evidence.py` (read/projection only)
|
||||
|
||||
Orchestration modules:
|
||||
|
||||
- `src/transcription/services/store.py`
|
||||
- `src/transcription/services/workflows.py`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- Aggregate ownership and invariants.
|
||||
- Transaction-aware write helpers.
|
||||
- Cross-service workflows in orchestration modules (`store.py`, `workflows.py`).
|
||||
- Lookup-table CRUD through the generic `RegistryService` base (`registry.py`), which is not an
|
||||
aggregate owner itself.
|
||||
|
||||
Module classification and per-model ownership are defined in
|
||||
[services instructions](../.github/instructions/services.instructions.md).
|
||||
|
||||
### Persistence Layer
|
||||
|
||||
- `src/transcription/db/**`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- SQLModel definitions, async session/engine runtime, registry bootstrap.
|
||||
- Loader helpers that enforce explicit eager loading with `lazy="raise"` relationships.
|
||||
|
||||
### Provider Layer
|
||||
|
||||
- `src/transcription/providers/**`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- Provider API encapsulation.
|
||||
- Request manifest and transport evidence capture.
|
||||
- Normalized transcription result contract.
|
||||
|
||||
## Core Domain Model
|
||||
|
||||
- `Document` owns archival metadata and links to `Source`, `Job`, and `DocumentPerson`.
|
||||
- `Source` is a document page/file record with selected machine projection and human revision.
|
||||
- `Job` is an aggregate processing run with status and frozen prompt/runtime settings.
|
||||
- `JobSource` is queue/membership state for one `(job, source)` pair.
|
||||
- `ExecutionAttempt` is append-only evidence for each provider call.
|
||||
- `Photo` is person imagery owned by `PhotosService`.
|
||||
- `MaintenanceRun` is one queued or executed operational maintenance run.
|
||||
- `GenealogyPerson`, `GenealogyFamily`, `GenealogyFamilyChild`, and `GenealogyCitation` store
|
||||
imported GEDCOM genealogy data and citation provenance.
|
||||
- `DocumentType` and `PersonRole` are UUID-backed registries with optional protected `semantic_key`.
|
||||
- `Tag` is a shared registry reached through both document and person tagging, linked by
|
||||
`DocumentTag` and `PersonTag`.
|
||||
|
||||
## Processing and Evidence Workflow
|
||||
|
||||
1. User creates/updates Document metadata and linked People atomically through workflow orchestration.
|
||||
2. User creates a Job by uploading one or more Source files or by retranscribing an existing Source.
|
||||
3. Source files are validated and stored; orientation normalization may be applied at ingest, and stored bytes become the canonical processing bytes.
|
||||
4. Worker claims queued Job, transitions to `processing`, and processes pending pages in deterministic order.
|
||||
5. Each provider call writes one immutable `ExecutionAttempt` with:
|
||||
- request manifest + hash
|
||||
- transport evidence (when response exists)
|
||||
- SDK snapshot and normalized metadata
|
||||
- outcome, timing, and error details when applicable
|
||||
6. `JobSource` status is updated as queue/projection state; `Source.raw_transcription` is set on first successful attempt and can be explicitly re-pointed by candidate promotion.
|
||||
7. Job terminal status resolves to `transcribed`, `partial_success`, or `failed`.
|
||||
|
||||
## Status Semantics
|
||||
|
||||
- **Job statuses:** `queued`, `processing`, `transcribed`, `partial_success`, `failed`
|
||||
- Operational success path resolves to `transcribed`.
|
||||
- **JobSource statuses:** `pending`, `transcribed`, `failed`, `cancelled`
|
||||
- **MaintenanceRun statuses:** `queued`, `processing`, `succeeded`, `failed`
|
||||
- Maintenance uses `succeeded` rather than `transcribed`; the transcription vocabulary does not
|
||||
apply to operational runs.
|
||||
- **Maintenance job types:** `backup`, `storage_reconciliation`, `gedcom_import`
|
||||
|
||||
## Maintenance Execution
|
||||
|
||||
Operational maintenance is queue-backed rather than run inline from the UI, so it survives request
|
||||
lifetime and is recorded:
|
||||
|
||||
1. Settings enqueues a `MaintenanceRun` with `status=queued` and a `triggered_by` marker.
|
||||
2. The worker claims the oldest queued run with a conditional update, moving it to `processing`.
|
||||
3. `backup` runs the deploy backup script; `storage_reconciliation` compares stored media against
|
||||
`Document`/`Source` records; `gedcom_import` parses the latest uploaded `.ged` file and upserts
|
||||
genealogy records.
|
||||
4. The run finalizes to `succeeded` or `failed` with summary, timing, log path, and `error_detail`.
|
||||
|
||||
`MaintenanceRun` records operational history and is not evidence in the `ExecutionAttempt` sense;
|
||||
append-only guarantees apply to transcription attempts.
|
||||
|
||||
## Security and Path Handling Boundaries
|
||||
|
||||
- Print media delivery uses record-validated API route:
|
||||
- `src/transcription/api/print_api.py`
|
||||
- General UI media links resolve through:
|
||||
- `src/transcription/ui/components/media_urls.py`
|
||||
- Local filesystem paths must never be accepted from user input as trusted media routes.
|
||||
|
||||
## Concurrency and Reliability Principles
|
||||
|
||||
- Worker loop reuses service bundle/provider resources for pooled calls.
|
||||
- Provider-call timeout is explicit and bounded.
|
||||
- Non-retriable worker-loop faults are surfaced and stop loop spin.
|
||||
- Per-page outcomes are durably persisted before processing next page.
|
||||
|
||||
## Design Decisions and Rationale
|
||||
|
||||
### Why `transcribed` is the success terminal state
|
||||
|
||||
- The worker and job orchestration resolve successful completion to `JobStatus.TRANSCRIBED`, with mixed and failure outcomes represented by `partial_success` and `failed`.
|
||||
- This keeps terminal status vocabulary aligned with what the pipeline actually produces: transcribed page content and evidence, not a generic completion marker.
|
||||
|
||||
### Why evidence history is append-only while page text is a projection
|
||||
|
||||
- `ExecutionAttempt` stores immutable per-call evidence and preserves full attempt history across retries.
|
||||
- `Source.raw_transcription` is intentionally a mutable projection so UI and exports can show a selected current machine text without mutating historical evidence.
|
||||
- This split keeps auditability and UX both first-class: history is durable, presentation is editable.
|
||||
|
||||
### Why orchestration modules own cross-service workflows
|
||||
|
||||
- Service modules do not import each other; aggregate ownership remains local to each service.
|
||||
- Multi-aggregate writes are coordinated in orchestration modules (`store.py`, `workflows.py`) so transaction boundaries are explicit and testable.
|
||||
- This avoids circular dependencies and keeps cross-cutting workflow logic centralized.
|
||||
|
||||
### Why explicit eager loading is required
|
||||
|
||||
- ORM relationships are configured with `lazy="raise"` in key paths, so code must request needed relationships up front.
|
||||
- This prevents hidden query behavior in UI/service code and makes read shape deterministic and reviewable.
|
||||
|
||||
### Why canonical source bytes may be ingest-normalized
|
||||
|
||||
- Ingest normalization can correct orientation before persistence so provider calls, evidence hashes, and rendered processing source are consistent.
|
||||
- The canonical stored bytes, digest, and size become the durable processing identity for that source.
|
||||
|
||||
### Why media access uses controlled routes/helpers
|
||||
|
||||
- Print/export media uses record-validated API endpoints to avoid direct filesystem path exposure.
|
||||
- General UI media URLs are generated through shared resolver helpers to keep path handling consistent and centralized.
|
||||
|
||||
## Scope Boundary
|
||||
|
||||
Current architecture rules live in `docs/*`.
|
||||
|
||||
## Related References
|
||||
|
||||
- [System Requirements](requirements.md)
|
||||
- [Data Model](schema.md)
|
||||
- [Error Handling Policy](error_handling.md)
|
||||
- [Production Runbook](production-runbook.md)
|
||||
- [Backup and Restore](backup_restore.md)
|
||||
- [Error Handling invariant](./invariant/error_handling.md)
|
||||
- [AI evidence invariant](./invariant/ai_evidence_and_provenance.md)
|
||||
@@ -0,0 +1,60 @@
|
||||
# Backup and Restore (V6.1)
|
||||
|
||||
This guide defines operational backup/restore for clean-slate recovery of the Docker runtime using a host-visible backup folder.
|
||||
|
||||
## 1. Backup artifacts
|
||||
|
||||
- Backup target root: `BACKUP_DIR` (recommended production value: `/backup`)
|
||||
- Database artifact per run:
|
||||
- `postgres-YYYYMMDD-HHMMSS.dump` (PostgreSQL custom dump via `pg_dump -Fc`)
|
||||
- `backup-YYYYMMDD-HHMMSS.manifest` (run manifest)
|
||||
- Media/config mirrors under `BACKUP_DIR`:
|
||||
- `uploads/**` (incremental copy: new files only)
|
||||
- `prompts/**` (prompt directory mirror)
|
||||
|
||||
Retention:
|
||||
|
||||
- `BACKUP_RETENTION_DAYS` applies to `postgres-*.dump` and `backup-*.manifest` files.
|
||||
|
||||
## 2. Creating backups
|
||||
|
||||
Run from repository root:
|
||||
|
||||
```bash
|
||||
sh deploy/backup/create_postgres_backup.sh
|
||||
```
|
||||
|
||||
Environment variables used by the backup script:
|
||||
|
||||
- `BACKUP_DIR` (default `./data/backups`)
|
||||
- `BACKUP_RETENTION_DAYS` (default `14`)
|
||||
- `UPLOAD_DIR` (default `/app/uploads`)
|
||||
- `PROMPT_DIR` (default `/app/prompts`)
|
||||
- `DATABASE__DRIVER` (must be `postgres`)
|
||||
- `DATABASE__HOST` (default `postgres`)
|
||||
- `DATABASE__PORT` (default `5432`)
|
||||
- `DATABASE__DATABASE` (required)
|
||||
- `DATABASE__USER` (required)
|
||||
- `DATABASE__PASSWORD` (required)
|
||||
|
||||
Recommended production setup:
|
||||
|
||||
- Mount a host-visible folder into `/backup` for both `app` and `worker`.
|
||||
- Set `BACKUP_DIR=/backup` in `.env.production`.
|
||||
- Use host-level tooling (for example Synology Drive Client on the host) to replicate that folder externally.
|
||||
|
||||
## 3. Restoring from backup
|
||||
|
||||
Restore requires downtime for app + worker writes.
|
||||
|
||||
1. Stop app and worker:
|
||||
- `docker compose --env-file .env.production -f docker-compose.production.yml stop app worker`
|
||||
2. Restore database:
|
||||
- `sh deploy/backup/restore_postgres_backup.sh /backup/postgres-YYYYMMDD-HHMMSS.dump`
|
||||
3. Start app and worker:
|
||||
- `docker compose --env-file .env.production -f docker-compose.production.yml start app worker`
|
||||
4. Validate `/healthz` and run one smoke workflow.
|
||||
|
||||
Notes:
|
||||
|
||||
- `restore_postgres_backup.sh` still supports legacy archive restore paths for older backup sets.
|
||||
@@ -0,0 +1,66 @@
|
||||
# Cloudflare Tunnel and Access Setup
|
||||
|
||||
This guide defines the repository-supported setup for exposing app and selected LAN services through Cloudflare Tunnel with Cloudflare Access protection.
|
||||
|
||||
## 1. Files used by this deployment
|
||||
|
||||
1. `deploy/cloudflared/config.yml` (local copy from `config.yml.example`)
|
||||
2. `.env.production` (`CLOUDFLARE_TUNNEL_TOKEN`)
|
||||
3. `docker-compose.production.yml` (`cloudflared` service reads token + mounts config)
|
||||
|
||||
Do not commit `config.yml` or `.env.production`.
|
||||
|
||||
## 2. Configure cloudflared
|
||||
|
||||
1. Copy `deploy/cloudflared/config.yml.example` to `deploy/cloudflared/config.yml`.
|
||||
2. Update hostname -> service mappings in `ingress`.
|
||||
3. Keep the final catch-all ingress `http_status:404`.
|
||||
4. Set `CLOUDFLARE_TUNNEL_TOKEN` in `.env.production`.
|
||||
|
||||
Example app route:
|
||||
|
||||
- `transcription.example.com` -> `http://app:8000`
|
||||
|
||||
Optional generic remote-access routes:
|
||||
|
||||
- `homeassistant.example.com` -> `http://<home-assistant-lan-ip>:8123`
|
||||
- `pihole.example.com` -> `http://<pihole-lan-ip>:80`
|
||||
|
||||
## 3. Cloudflare Access policy baseline
|
||||
|
||||
Create one Access app policy per exposed hostname:
|
||||
|
||||
1. Include: your allowed identities/groups only.
|
||||
2. Exclude: none by default.
|
||||
3. Require: identity provider login (and MFA if available).
|
||||
|
||||
Recommended baseline:
|
||||
|
||||
- App endpoint (`transcription.*`): your admin identity set.
|
||||
- Other internal endpoints (`homeassistant.*`, `pihole.*`, etc.): explicit least-privilege groups.
|
||||
|
||||
## 4. Startup
|
||||
|
||||
Start production stack:
|
||||
|
||||
```bash
|
||||
docker compose --env-file .env.production -f docker-compose.production.yml up -d --build
|
||||
```
|
||||
|
||||
Validate tunnel container:
|
||||
|
||||
```bash
|
||||
docker compose --env-file .env.production -f docker-compose.production.yml logs cloudflared
|
||||
```
|
||||
|
||||
LXC/proxied-network note:
|
||||
|
||||
- The `cloudflared` service is pinned to `--protocol http2` with explicit DNS resolvers (`1.1.1.1`, `1.0.0.1`) in `docker-compose.production.yml`.
|
||||
- This avoids environments where Docker's embedded resolver (`127.0.0.11`) cannot resolve `region*.v2.argotunnel.com`, which causes connector precheck failure and tunnel shutdown.
|
||||
- If tunnel status is still down, verify host/container egress for DNS and TCP 443 to `api.cloudflare.com` and `*.argotunnel.com`.
|
||||
|
||||
## 5. Security notes
|
||||
|
||||
- Keep `postgres` and other internal-only services off public hostnames unless required.
|
||||
- Use distinct hostnames per service; avoid path-based multiplexing for unrelated admin surfaces.
|
||||
- Rotate `CLOUDFLARE_TUNNEL_TOKEN` and Access policy memberships on a regular schedule.
|
||||
@@ -0,0 +1,94 @@
|
||||
# Database Rebuild Migration Workflow
|
||||
|
||||
This project uses an explicit **export/import rebuild workflow** for schema migration.
|
||||
|
||||
Policy:
|
||||
- Do not add runtime legacy-compatibility write paths.
|
||||
- Rebuild a fresh target database from current models.
|
||||
- Export current data/media, then import into the fresh target.
|
||||
|
||||
## Commands
|
||||
|
||||
### 1) Export current DB + uploads into a bundle
|
||||
|
||||
```bash
|
||||
uv run python tools/export_import_migration.py export --bundle-dir .migration-bundle
|
||||
```
|
||||
|
||||
Optional source overrides:
|
||||
- `--source-db <path-or-sqlalchemy-url>`
|
||||
- `--source-upload-dir <path>`
|
||||
|
||||
### 2) Import bundle into a fresh target (SQLite or PostgreSQL)
|
||||
|
||||
```bash
|
||||
uv run python tools/export_import_migration.py import --bundle-dir .migration-bundle --target-db .\data\transcription-new.db --target-upload-dir .\data-new
|
||||
```
|
||||
|
||||
PostgreSQL target example:
|
||||
|
||||
```bash
|
||||
uv run python tools/export_import_migration.py import --bundle-dir .migration-bundle --target-db postgresql://transcription:change-me@localhost:5432/transcription --target-upload-dir .\data-new
|
||||
```
|
||||
|
||||
### 3) Verify migration parity and integrity
|
||||
|
||||
```bash
|
||||
uv run python tools/export_import_migration.py verify --source-db .\data\transcription.db --target-db postgresql://transcription:change-me@localhost:5432/transcription
|
||||
```
|
||||
|
||||
The verify command checks:
|
||||
|
||||
- row-count parity across migration tables
|
||||
- orphan-reference checks for `source`, `job`, `job_source`, and `execution_attempt`
|
||||
- duplicate `(job_id, source_id, attempt_number)` in `execution_attempt`
|
||||
|
||||
Exit code:
|
||||
|
||||
- `0` when counts and integrity checks pass
|
||||
- `1` when mismatches or integrity violations are detected
|
||||
|
||||
### 4) One-shot export+import
|
||||
|
||||
```bash
|
||||
uv run python tools/export_import_migration.py migrate --bundle-dir .migration-bundle --target-db .\data\transcription-new.db --target-upload-dir .\data-new
|
||||
```
|
||||
|
||||
## What gets migrated
|
||||
|
||||
- Tables (in dependency order): `document_type`, `person_role`, `tag`, `document`, `person`, `photo`, `document_person`, `document_tag`, `person_tag`, `job`, `source`, `job_source`, `execution_attempt`.
|
||||
- Media tree under `UPLOAD_DIR`.
|
||||
|
||||
The bundle contains:
|
||||
- `database.json` (row export)
|
||||
- `uploads/` (copied media files)
|
||||
|
||||
Path normalization during export/import:
|
||||
- `source.file_path` is normalized to `documents/...` (upload-root-relative POSIX).
|
||||
- `photo.path` is normalized to `photos/...` (upload-root-relative POSIX).
|
||||
|
||||
Legacy V4.x portrait/homepage backfill in the export step:
|
||||
- If the source DB has no `photo` table, the exporter synthesizes `photo` rows from legacy `person.portrait_path` values and from legacy homepage image files under `UPLOAD_DIR/homepage`.
|
||||
- Legacy portrait and homepage image files are copied into the unified `UPLOAD_DIR/photos/{photo_id}{suffix}` layout in the migration bundle.
|
||||
- Legacy homepage markdown is relocated from `UPLOAD_DIR/homepage/homepage.md` to `UPLOAD_DIR/homepage.md`.
|
||||
- Legacy `person.full_name` values are split into `given_names` + `last_name` for V5.1 schema compatibility.
|
||||
|
||||
## Cutover (SQLite -> PostgreSQL)
|
||||
|
||||
After importing to a fresh target:
|
||||
1. Stop app and worker services to freeze writes.
|
||||
2. Export a migration bundle from the last SQLite state.
|
||||
3. Import bundle to PostgreSQL target.
|
||||
4. Run `verify` against source and target before switching runtime.
|
||||
5. Switch runtime config to PostgreSQL (`DATABASE__DRIVER=postgres` and related `DATABASE__*` values).
|
||||
6. Start app and worker services.
|
||||
7. Run smoke checks (`/healthz`, create/upload/process one job).
|
||||
|
||||
## Rollback
|
||||
|
||||
If verify or smoke checks fail:
|
||||
|
||||
1. Stop app and worker services.
|
||||
2. Revert runtime config to SQLite.
|
||||
3. Start app and worker against pre-cutover SQLite database.
|
||||
4. Preserve failed migration bundle and logs for analysis.
|
||||
@@ -0,0 +1,137 @@
|
||||
# Error Handling Policy (Current Baseline: V6.1)
|
||||
|
||||
This policy defines the active V6.1 error taxonomy, translation boundaries, and retry semantics.
|
||||
|
||||
## Error Categories
|
||||
|
||||
| Category | Meaning | Typical Origin | User Treatment |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| `validation` | Input payload/selection is invalid | UI form parsing, service validators | Inline correction guidance |
|
||||
| `not_found` | Target record is missing | ID lookup in service layer | Non-blocking warning or redirect |
|
||||
| `conflict` | State prevents requested action | lifecycle transitions, duplicate semantic keys | Explain required precondition |
|
||||
| `external` | Provider/network dependency failure | OpenRouter/provider adapter | Retry path and evidence retained |
|
||||
| `timeout` | Provider call exceeded configured bound | worker/provider client timeout | Retry path and bounded messaging |
|
||||
| `internal` | Unexpected local failure | unhandled service/runtime faults | Safe generic message + diagnostics capture |
|
||||
|
||||
## Runtime Taxonomy and Canonical Mapping
|
||||
|
||||
Runtime code uses a richer internal taxonomy for diagnostics and persisted evidence, then maps that
|
||||
taxonomy to the six canonical categories at the API/UI envelope boundary.
|
||||
|
||||
### Internal runtime categories
|
||||
|
||||
- `validation_error`
|
||||
- `user_input_error`
|
||||
- `not_found_error`
|
||||
- `conflict_error`
|
||||
- `external_provider_error`
|
||||
- `external_timeout_error`
|
||||
- `processing_error`
|
||||
- `infrastructure_transient_error`
|
||||
- `infrastructure_persistent_error`
|
||||
- `internal_unexpected_error`
|
||||
|
||||
### Internal -> Canonical mapping
|
||||
|
||||
| Internal category | Canonical envelope category |
|
||||
| :--- | :--- |
|
||||
| `validation_error` | `validation` |
|
||||
| `user_input_error` | `validation` |
|
||||
| `not_found_error` | `not_found` |
|
||||
| `conflict_error` | `conflict` |
|
||||
| `external_provider_error` | `external` |
|
||||
| `external_timeout_error` | `timeout` |
|
||||
| `infrastructure_transient_error` | `timeout` |
|
||||
| `processing_error` | `internal` |
|
||||
| `infrastructure_persistent_error` | `internal` |
|
||||
| `internal_unexpected_error` | `internal` |
|
||||
|
||||
`ExecutionAttempt.error_category` stores the internal category value so diagnostics remain specific.
|
||||
|
||||
## Translation Boundaries
|
||||
|
||||
- **Provider layer:** raise provider-scoped exceptions with provider context; do not emit UI text.
|
||||
- **Service layer:** map raw exceptions into internal categories and preserve causal chain.
|
||||
- **UI/API layer:** convert internal categories to canonical categories using the centralized mapping.
|
||||
|
||||
## Decision Context
|
||||
|
||||
### Why taxonomy is category-based (not exception-class-based)
|
||||
|
||||
- Categories encode operator-facing recovery semantics (fix input, retry later, investigate internal failure) independent of low-level exception type.
|
||||
- This keeps retry and messaging behavior consistent even when provider/client libraries change.
|
||||
|
||||
### Why page-level failure is isolated
|
||||
|
||||
- Multi-page archival documents often contain a mix of readable and degraded pages.
|
||||
- Isolating failures to page scope preserves successful results and avoids all-or-nothing loss when one page fails.
|
||||
- Aggregate job status then communicates overall outcome (`transcribed`, `partial_success`, `failed`) without hiding page detail.
|
||||
|
||||
### Why retries append evidence instead of mutating rows
|
||||
|
||||
- Retry operations are new observations, not corrections of history.
|
||||
- Appending attempts preserves forensic traceability, timing history, and provider variability analysis.
|
||||
- Projection updates remain explicit user/workflow decisions, separate from immutable evidence.
|
||||
|
||||
## Job and Page Failure Semantics
|
||||
|
||||
### Page-Level (`JobSource`)
|
||||
|
||||
- `pending` -> `transcribed` when attempt succeeds.
|
||||
- `pending` -> `failed` when attempt fails terminally.
|
||||
- `pending` -> `cancelled` on job cancellation before processing.
|
||||
|
||||
### Job-Level (`Job`)
|
||||
|
||||
- `transcribed` when all pages transcribe successfully.
|
||||
- `partial_success` when mixed success/failure outcomes exist.
|
||||
- `failed` when no page transcribes successfully.
|
||||
|
||||
## Retry and Retranscription Rules
|
||||
|
||||
1. Failed/cancelled pages may be re-queued through retranscription workflows.
|
||||
2. Retry attempts must append new `ExecutionAttempt` rows; prior evidence remains immutable.
|
||||
3. Selecting a better candidate must update projection pointers, not mutate historical attempt rows.
|
||||
|
||||
## Logging and Diagnostics Rules
|
||||
|
||||
1. Persist sufficient attempt error metadata (`error_category`, `error_message`, transport evidence) for post-hoc analysis.
|
||||
2. Avoid leaking stack traces or local paths into user-facing message envelopes.
|
||||
3. Preserve causal exception chains for internal diagnostics.
|
||||
|
||||
### Message vs detail split
|
||||
|
||||
Rules 1 and 2 pull in opposite directions: evidence records need the root cause, and
|
||||
user-facing envelopes must not carry it. `AppError` therefore separates the two audiences:
|
||||
|
||||
| Field | Audience | Carries root cause | Surfaces |
|
||||
| --- | --- | --- | --- |
|
||||
| `message` | User-facing and API-facing | No | `show_error`, `build_error_envelope` |
|
||||
| `detail` | Internal only | Yes | `format_error_detail` (evidence), logs, sanitized UI projection only |
|
||||
|
||||
`classify_unexpected_error` builds a generic `message` and puts the exception type and
|
||||
text on `detail`. Anything rendered to a user or serialized into an API envelope must
|
||||
read `message`; anything persisted as provenance or logged may read `detail`. When a UI
|
||||
surface needs to show persisted `error_detail`, it must route through a sanitizing
|
||||
projection that preserves the category, suggestion, and error reference while reducing
|
||||
machine-local absolute paths to basenames only.
|
||||
Enforced by `tests/test_errors.py::test_unexpected_error_does_not_leak_filesystem_paths`
|
||||
and `tests/test_error_message_safety.py`.
|
||||
|
||||
## Operator Recovery Guidance
|
||||
|
||||
- **validation/conflict:** correct input or state and retry manually.
|
||||
- **external/timeout:** allow bounded retries and keep prior attempt evidence visible.
|
||||
- **internal:** stop automatic retries, surface a safe message, and inspect diagnostics with correlation context.
|
||||
|
||||
## UI Messaging Contract
|
||||
|
||||
- User-visible errors must be actionable, bounded, and category-consistent.
|
||||
- Multi-page jobs must show partial outcomes instead of collapsing into a single opaque failure.
|
||||
- Recovery actions (`retry`, `retranscribe`, `edit input`) must be offered where available.
|
||||
|
||||
## Cross-Reference
|
||||
|
||||
- [Error Handling invariant](./invariant/error_handling.md)
|
||||
- [System Requirements](requirements.md)
|
||||
- [Data Model](schema.md)
|
||||
@@ -0,0 +1,36 @@
|
||||
# Document Transcription System Overview (Current Baseline: V6.1)
|
||||
|
||||
This directory is the single source of truth for current V6.1 behavior and architecture.
|
||||
|
||||
## Canonical Reading Order
|
||||
|
||||
1. [System Architecture](architecture.md) for runtime topology, boundaries, and lifecycle ownership.
|
||||
2. [System Requirements](requirements.md) for verifiable current-state requirements.
|
||||
3. [Data Model](schema.md) for entities, constraints, and evidence persistence rules.
|
||||
4. [Error Handling Policy](error_handling.md) for category, translation, and retry behavior.
|
||||
|
||||
## Cross-Version Invariants
|
||||
|
||||
- [Historical Document Transcription Design Intent](./invariant/intent.md)
|
||||
- [Transcription Methodology](./invariant/transcription_methodology.md)
|
||||
- [Error Handling](./invariant/error_handling.md)
|
||||
- [Digital Evidence and AI Processing Provenance](./invariant/ai_evidence_and_provenance.md)
|
||||
- [UI Style Guide](./invariant/ui_style_guide.md)
|
||||
|
||||
## Deployment and Operations
|
||||
|
||||
- [Production Runbook](production-runbook.md) for deploy, rollback, and recovery.
|
||||
- [Backup and Restore](backup_restore.md) for backup configuration and restore procedure.
|
||||
- [Data Migration](data_migration.md) for the SQLite to PostgreSQL migration path.
|
||||
- [Cloudflare Tunnel and Access](cloudflare_tunnel_access.md) for remote exposure and access control.
|
||||
|
||||
## Baseline Statement
|
||||
|
||||
The current V6.1 baseline includes the architectural cleanup, person-schema redesign,
|
||||
containerized PostgreSQL deployment, and the navigation, Document Detail, and worker-backed
|
||||
maintenance refinements reflected across this canonical document set.
|
||||
Use this `docs/*` canonical set for active design and implementation decisions.
|
||||
|
||||
Every canonical document above states this same baseline; `tests/test_meta_contract_guards.py`
|
||||
fails if one of them falls behind. Forward-looking work is tracked in
|
||||
[`roadmap_plan.md`](roadmap_plan.md) and is not part of the baseline.
|
||||
@@ -0,0 +1,145 @@
|
||||
# Digital Evidence and AI Processing Provenance (Invariant)
|
||||
|
||||
## 1. Purpose
|
||||
|
||||
This document defines non-negotiable evidence and provenance rules for the transcription application.
|
||||
|
||||
The application exists to preserve historical source material and produce useful transcriptions without losing the ability to inspect, reinterpret, or reprocess the evidence later. Provider integrations, model names, schemas, and user interfaces may change; the principles below must remain true.
|
||||
|
||||
## 2. Evidence Model
|
||||
|
||||
The application distinguishes five kinds of information:
|
||||
|
||||
1. **Source evidence**: the canonical stored media used for processing and the facts needed to identify and verify it.
|
||||
2. **Execution specification**: the frozen instructions, parameters, source identity, and software context for one processing attempt.
|
||||
3. **Transport evidence**: the response received at the application/provider boundary, including safe protocol metadata.
|
||||
4. **Normalized data**: selected fields extracted for search, display, accounting, and workflow behavior.
|
||||
5. **Derived artifacts**: outputs produced from source evidence, such as transcription text, OCR geometry, confidence data, layout analysis, or entity extraction.
|
||||
|
||||
Normalized data and derived artifacts never replace source or transport evidence.
|
||||
|
||||
## 3. Core Invariants
|
||||
|
||||
### 3.1 Canonical Source Preservation
|
||||
|
||||
1. Each source must have one canonical stored byte stream used for processing and provenance.
|
||||
2. Canonical storage may apply deterministic ingest normalization before persistence.
|
||||
3. Canonical stored bytes must have a cryptographic content digest, byte size, and stable identity.
|
||||
4. Post-ingest processing derivatives must not overwrite canonical stored bytes.
|
||||
5. Moving or renaming a stored file must not change its evidence identity.
|
||||
|
||||
### 3.2 Append-Only Processing History
|
||||
|
||||
1. Every processing attempt must have a distinct execution record, whether it succeeds, partially succeeds, times out, or fails.
|
||||
2. A later attempt must not overwrite the evidence from an earlier attempt.
|
||||
3. A convenient “latest transcription” value may be maintained as a cache or projection, but it is not the authoritative execution history.
|
||||
4. Human revisions must remain distinguishable from all machine-generated outputs.
|
||||
5. Reprocessing a source must create new evidence rather than rewriting historical evidence.
|
||||
|
||||
### 3.3 Frozen Execution Specification
|
||||
|
||||
Each execution must preserve enough information to understand what the application asked the processor to do:
|
||||
|
||||
1. Requested provider, model, and provider-routing constraints.
|
||||
2. Full effective system and user instructions.
|
||||
3. Prompt asset name and content digest when a prompt asset is used.
|
||||
4. Every explicitly supplied generation or processing parameter.
|
||||
5. Whether an optional parameter was explicitly set or omitted.
|
||||
6. Canonical source digest (and derivative digests when used), media type, dimensions or page geometry when known, and page identity.
|
||||
7. A secret-safe representation of the request structure.
|
||||
8. Application, provider-adapter, and client-library versions sufficient to interpret the execution.
|
||||
|
||||
The execution specification must not contain credentials, authorization headers, secret query values, or unnecessary duplicate source binaries.
|
||||
|
||||
### 3.4 Evidence-Layer Terminology
|
||||
|
||||
The following terms are not interchangeable:
|
||||
|
||||
- **Transport response**: the status, safe headers, and exact response body received by the application at its HTTP boundary.
|
||||
- **Router-normalized response**: a response transformed by an intermediary into its common schema.
|
||||
- **SDK-parsed response**: an object created when a client library validates or filters a response.
|
||||
- **Normalized metadata**: application-selected fields derived from a response.
|
||||
- **Native provider response**: the upstream provider's own response before any intermediary transformation.
|
||||
|
||||
The application and its documentation must identify which layer is stored. A response must not be described as “raw,” “complete,” or “native” without naming the boundary at which that claim is true.
|
||||
|
||||
### 3.5 Transport Evidence
|
||||
|
||||
1. Preserve the exact successful response body received at the application's transport boundary before SDK model parsing can discard unknown fields.
|
||||
2. Preserve the response status and an allowlisted set of non-secret headers needed for correlation, content interpretation, rate-limit diagnosis, or audit.
|
||||
3. Preserve provider/router request and generation identifiers when available.
|
||||
4. Preserve safe response evidence for unsuccessful calls when a response was received.
|
||||
5. Record explicitly when no response was received, such as a local timeout or connection failure.
|
||||
6. Retain parsed and normalized forms only as additional representations of the preserved response.
|
||||
|
||||
Wire-level packet capture, TLS session data, credentials, and unrestricted headers are neither required nor permitted.
|
||||
|
||||
These requirements apply to executions performed after transport capture is implemented. For earlier executions, the absence of transport evidence must be represented explicitly. An SDK snapshot or normalized record must never be relabeled or backfilled as transport evidence.
|
||||
|
||||
### 3.6 Derived Artifact Provenance
|
||||
|
||||
1. Every derived artifact must identify its source evidence and producing execution.
|
||||
2. Each artifact must declare its semantic type, media/serialization format, schema name and version, producer, producer version, and creation time.
|
||||
3. Artifact content must be stored directly or referenced by a stable path or object identifier and protected by a cryptographic digest.
|
||||
4. Coordinates must declare their coordinate system, units, origin, page/image dimensions, and transformation history.
|
||||
5. Confidence values must identify the producer and scope to which they apply; values from different producers must not be treated as directly comparable without validation.
|
||||
6. Provider-specific payloads may be retained, but durable application behavior must not depend on undocumented provider fields.
|
||||
|
||||
This model must accommodate future OCR text, word or line polygons, layout regions, confidence data, alternate transcriptions, and structured extraction without adding a dedicated column for every possible feature.
|
||||
|
||||
### 3.7 Integrity and Auditability
|
||||
|
||||
1. Stored evidence must be exportable with enough identifiers and metadata to verify relationships and digests outside the application.
|
||||
2. Evidence mutation, deletion, and retention behavior must be explicit and testable.
|
||||
3. Schema upgrades must preserve existing evidence and its original meaning.
|
||||
4. Backfills must be identified as backfills; they must not imply that previously uncaptured evidence existed.
|
||||
5. Integrity verification must distinguish a missing file, digest mismatch, unavailable external artifact, and malformed metadata.
|
||||
|
||||
### 3.8 Security and Privacy
|
||||
|
||||
1. API keys, authorization headers, cookies, and credentials must never be persisted as provenance.
|
||||
2. Persist only headers and metadata fields that appear on an explicit allowlist of known-safe fields. Discard all other fields before storage; never persist an unrestricted capture and attempt to redact it afterward.
|
||||
3. Request manifests should reference source content by identity instead of duplicating base64 source data.
|
||||
4. Diagnostic displays and exports must avoid exposing secrets or machine-local details that are not necessary for evidence interpretation.
|
||||
|
||||
## 4. Reproducibility Limits
|
||||
|
||||
Provenance supports explanation, comparison, and best-effort reproduction; it does not guarantee identical output.
|
||||
|
||||
Identical requests may produce different results because of model updates, provider routing, nondeterministic computation, undocumented defaults, safety systems, or retired endpoints. The application must preserve whether a parameter was omitted rather than pretending to know the provider default used at that time.
|
||||
|
||||
Likewise, preserving a general vision-model response does not create OCR coordinates that were never returned. Future coordinate extraction remains possible because canonical source evidence is preserved and can be processed again by a suitable system.
|
||||
|
||||
## 5. Model Evaluation Policy
|
||||
|
||||
Model selection must be based on a representative sample of the actual archive rather than vendor claims alone.
|
||||
|
||||
Evaluation should:
|
||||
|
||||
1. Use manually reviewed reference transcriptions following the project's [Transcription Methodology](transcription_methodology.md).
|
||||
2. Represent printed, typed, handwritten, degraded, tabular, multilingual, and spatially complex material present in the archive.
|
||||
3. Measure character and word error rates where appropriate.
|
||||
4. Separately record silent corrections, invented text, omitted text, uncertainty handling, layout fidelity, cost, and latency.
|
||||
5. Preserve the exact model, endpoint or route, parameters, prompt, source digest, and scoring method for every comparison.
|
||||
6. Treat model rankings as corpus- and version-specific, not permanent declarations of a universal “best” model.
|
||||
|
||||
The deterministic scorer for these comparisons lives in `src/transcription/benchmarking.py`; it is
|
||||
retained as evaluation-policy infrastructure even though application runtime paths do not call it
|
||||
directly.
|
||||
|
||||
Benchmark material containing family records remains private application data unless explicitly approved for publication.
|
||||
|
||||
## 6. Ownership and Change Policy
|
||||
|
||||
1. Canonical V6.1 architecture, schema, requirements, and error-policy documents define how current behavior satisfies this invariant.
|
||||
2. Provider adapters own the capture of provider-boundary evidence.
|
||||
3. Services own validation, persistence, retention, and export behavior.
|
||||
4. UI pages may inspect evidence through service contracts but do not define evidence semantics.
|
||||
5. If implementation conflicts with this invariant, either correct the implementation or explicitly revise this document before accepting the behavior.
|
||||
6. Revisions to this document require deliberate review because they change the long-term preservation contract.
|
||||
|
||||
## 7. Related Invariants
|
||||
|
||||
- [Historical Document Transcription Design Intent](intent.md)
|
||||
- [Transcription Methodology & Style Guide](transcription_methodology.md)
|
||||
- [UI Style Guide](ui_style_guide.md)
|
||||
@@ -0,0 +1,101 @@
|
||||
# Error Handling (Invariant)
|
||||
|
||||
## 1. Purpose
|
||||
|
||||
This document defines the non-negotiable failure-handling principles for the transcription application.
|
||||
|
||||
Error categories, API envelopes, status codes, framework integrations, and persistence fields may change between versions. Failures must nevertheless remain visible, safe, diagnosable, and consistent across every application boundary.
|
||||
|
||||
## 2. Core Invariants
|
||||
|
||||
### 2.1 Failures Are Visible
|
||||
|
||||
1. An operation must not report success when all or part of the requested work failed.
|
||||
2. Invalid input, unavailable dependencies, persistence failures, provider failures, and unexpected defects must be surfaced through the application's established error path.
|
||||
3. Code must not silently discard an exception, provider response, invalid value, or failed state transition.
|
||||
4. When work can partially succeed, the successful and failed portions must be identified separately.
|
||||
|
||||
### 2.2 Messages Are Actionable
|
||||
|
||||
1. Operator-facing errors must explain what failed in concise language.
|
||||
2. When a safe corrective action is known, the error must state it.
|
||||
3. Expected validation or conflict failures must not be presented as unexplained internal defects.
|
||||
4. Internal diagnostics must not replace a usable operator-facing message.
|
||||
|
||||
### 2.3 Errors Have Stable Identity and Classification
|
||||
|
||||
1. Every surfaced failure must have a stable correlation identifier or equivalent trace identity.
|
||||
2. Failures must be classified into a documented, machine-readable category.
|
||||
3. Boundary-specific representations must preserve the original category and correlation identity.
|
||||
4. Unknown exceptions must be converted at an explicit boundary, retain their causal chain for diagnostics, and be classified as unexpected rather than disguised as an expected failure.
|
||||
|
||||
### 2.4 Boundary Translation Is Consistent
|
||||
|
||||
1. UI, API, service, worker, persistence, and provider boundaries must use one shared error model or deterministic translations between documented models.
|
||||
2. A boundary may simplify presentation, but it must not change the meaning, retryability, or identity of a failure.
|
||||
3. Domain and service code must not depend on UI notifications or HTTP response types.
|
||||
4. UI and API layers must not infer error categories by parsing message text.
|
||||
|
||||
### 2.5 State Changes Are Safe
|
||||
|
||||
1. A failed atomic operation must leave persisted state unchanged.
|
||||
2. Batch operations may preserve successful independent items only when partial success is an explicit part of the workflow contract.
|
||||
3. A failed item must retain enough state to identify what was attempted and whether retry is safe.
|
||||
4. Error handling must not overwrite earlier successful results or historical execution evidence.
|
||||
|
||||
### 2.6 Retry Is Explicit and Bounded
|
||||
|
||||
1. Validation, authorization, policy, conflict, and other deterministic failures must not be retried automatically without a relevant input or state change.
|
||||
2. Automatic retry is permitted only for failures classified as transient and only when the operation is idempotent or otherwise protected from duplicate effects.
|
||||
3. Retry count, delay, and terminal behavior must be bounded and observable.
|
||||
4. Exhausted retries must end in a visible terminal failure rather than an indefinitely pending state.
|
||||
|
||||
### 2.7 Diagnostics Are Preserved Safely
|
||||
|
||||
1. Logs and persisted diagnostic evidence must retain enough context to correlate the failure with the affected operation and record.
|
||||
2. Provider and infrastructure failures must preserve safe diagnostic evidence at the boundary where it is available.
|
||||
3. Credentials, authorization headers, cookies, secret values, and unnecessary personal data must not appear in errors, logs, notifications, or exports.
|
||||
4. Diagnostic metadata capture must use explicit safe-field allowlists where unrestricted content could contain secrets.
|
||||
5. User-facing messages must not expose stack traces, local filesystem details, database credentials, or raw internal exceptions.
|
||||
|
||||
AI execution failures also follow the evidence rules in [Digital Evidence and AI Processing Provenance](ai_evidence_and_provenance.md).
|
||||
|
||||
### 2.8 Cancellation and Timeout Are Distinct Outcomes
|
||||
|
||||
1. User cancellation, application shutdown, local timeout, remote timeout, and provider rejection must remain distinguishable.
|
||||
2. Cancellation must not be converted into success or a generic unexpected error.
|
||||
3. Timeout handling must identify whether a provider response was received when that fact is known.
|
||||
4. Cleanup after cancellation or timeout must preserve consistency and must not conceal a completed side effect.
|
||||
|
||||
### 2.9 Logging Must Support Audit Without Becoming the Record
|
||||
|
||||
1. Structured logs must include correlation identity, operation, category, and relevant non-secret record identifiers.
|
||||
2. Expected operator errors may be logged less severely than unexpected defects, but they must remain observable.
|
||||
3. Logs are operational diagnostics and do not replace required database state or archival evidence.
|
||||
4. Duplicate logging of the same failure at every layer should be avoided; ownership of the authoritative log event must be clear.
|
||||
|
||||
## 3. Verification Policy
|
||||
|
||||
Each version must verify:
|
||||
|
||||
1. Every documented error category reaches the intended UI and API representation.
|
||||
2. Failed atomic writes roll back completely.
|
||||
3. Partial-success workflows preserve successful independent results and identify failed items.
|
||||
4. Retry behavior is bounded and restricted to eligible failures.
|
||||
5. Unexpected exceptions retain correlation and causal information without exposing sensitive details.
|
||||
6. Logs, persisted evidence, UI messages, and exports contain no credentials.
|
||||
7. Cancellation, timeout, provider response failure, and no-response failure remain distinguishable.
|
||||
|
||||
## 4. Versioned Ownership
|
||||
|
||||
1. Version-specific error taxonomies, envelopes, HTTP mappings, model fields, and framework behavior belong in the applicable version documentation.
|
||||
2. Each versioned error-handling document must state how it satisfies this invariant.
|
||||
3. A version may add stricter safeguards but must not weaken these principles without first revising this invariant deliberately.
|
||||
4. Implementation and tests must be updated together when a versioned error contract changes.
|
||||
|
||||
## 5. Related Invariants
|
||||
|
||||
- [Historical Document Transcription Design Intent](intent.md)
|
||||
- [Digital Evidence and AI Processing Provenance](ai_evidence_and_provenance.md)
|
||||
- [UI Style Guide](ui_style_guide.md)
|
||||
|
||||
@@ -45,6 +45,24 @@ The following rules map directly to editorial conventions for handling common ma
|
||||
| **Non-Textual Artifacts** | Record non-textual elements (seals, stamps, sketches, physical damage) using brief descriptive text inside square brackets. | [description] | [wax notary seal attached here] or [sketch of a fort layout] |
|
||||
| **Marginalia & Addenda** | Explicitly indicate spatial transitions before transcribing content located in margins or non-standard orientations. | [location:] | [written in left margin:] Do not share this with anyone. |
|
||||
|
||||
### 3.4 Document-Body Medium
|
||||
|
||||
Every transcript must identify the predominant document-body medium exactly once at the beginning:
|
||||
|
||||
| Medium | Use | Standard Markup |
|
||||
| --- | --- | --- |
|
||||
| **Handwritten** | The main body was written by hand. | `[document body handwritten]` |
|
||||
| **Typewritten** | The main body was produced with a typewriter. Uneven impressions, monospaced characters, and mechanical defects remain typewritten rather than handwritten. | `[document body typewritten]` |
|
||||
| **Typeset** | The main body was composed for printing or produced as printed text rather than with a typewriter. | `[document body typeset]` |
|
||||
| **Mixed** | No single medium predominates, or handwritten and printed/typewritten content are structurally interleaved. | `[document body mixed]` |
|
||||
|
||||
- Use exactly one document-body marker.
|
||||
- Do not wrap each line in `[handwritten: ...]` after declaring the body handwritten.
|
||||
- In typewritten or typeset documents, use localized handwriting markers only for genuinely handwritten annotations, insertions, or signatures.
|
||||
- In mixed documents, identify handwritten portions locally while preserving their reading context.
|
||||
- Preserve tables of contents, tables, forms, columns, captions, marginalia, page numbers, dotted leaders, and associated references in their logical reading order.
|
||||
- Produce plain text characters rather than HTML entities for ordinary transcription content.
|
||||
|
||||
## 4. Prompt Asset Integration
|
||||
|
||||
When executing programmatic transcriptions via LLM APIs or local models, processing instructions must be packaged into single-purpose system prompts aligned with these rules.
|
||||
|
||||
@@ -4,14 +4,14 @@
|
||||
This guide defines non-negotiable UI styling rules for the transcription application.
|
||||
|
||||
The design system is token-first and class-driven:
|
||||
1. Theme tokens are defined in [src/transcription/ui/static/theme.css](src/transcription/ui/static/theme.css).
|
||||
1. Theme tokens are defined in [src/transcription/ui/static/theme.css](../../src/transcription/ui/static/theme.css).
|
||||
2. Python UI code composes semantic classes instead of inline color values.
|
||||
3. Pages and components should share a single visual language across Documents, Jobs, People, and Sources flows.
|
||||
|
||||
## 2. Source of Truth
|
||||
Use these files as the style authority:
|
||||
1. [src/transcription/ui/static/theme.css](src/transcription/ui/static/theme.css) for color tokens, semantic utility classes, table styles, and viewer surfaces.
|
||||
2. [src/transcription/ui/theme.py](src/transcription/ui/theme.py) for runtime NiceGUI theme bridge and shared UI helpers.
|
||||
1. [src/transcription/ui/static/theme.css](../../src/transcription/ui/static/theme.css) for color tokens, semantic utility classes, table styles, and viewer surfaces.
|
||||
2. [src/transcription/ui/theme.py](../../src/transcription/ui/theme.py) for runtime NiceGUI theme bridge and shared UI helpers.
|
||||
|
||||
If this document conflicts with implementation, update this document to match the code immediately after intentional style changes.
|
||||
|
||||
@@ -25,7 +25,7 @@ If this document conflicts with implementation, update this document to match th
|
||||
## 4. Token System
|
||||
|
||||
### 4.1 Palette Tokens
|
||||
Base palette variables live under :root in [src/transcription/ui/static/theme.css](src/transcription/ui/static/theme.css):
|
||||
Base palette variables live under :root in [src/transcription/ui/static/theme.css](../../src/transcription/ui/static/theme.css):
|
||||
1. --palette-carbon-black: #1c2321
|
||||
2. --palette-cool-steel: #7d98a1
|
||||
3. --palette-blue-slate: #5e6572
|
||||
@@ -65,40 +65,38 @@ Semantic tokens currently include:
|
||||
4. ui-card-surface
|
||||
5. ui-row-surface
|
||||
6. ui-note-box
|
||||
7. ui-card-error
|
||||
|
||||
### 5.3 Interactive Elements
|
||||
1. ui-btn-primary
|
||||
2. ui-btn-secondary
|
||||
3. ui-link-primary
|
||||
4. ui-text-accent
|
||||
5. ui-chip-primary
|
||||
6. ui-badge-secondary
|
||||
7. ui-status and ui-status--<status>
|
||||
|
||||
### 5.4 Table Patterns
|
||||
1. ui-table
|
||||
2. ui-table-header
|
||||
3. ui-table-body
|
||||
|
||||
Use existing class combinations from [src/transcription/ui/components](src/transcription/ui/components) and [src/transcription/ui/pages](src/transcription/ui/pages) as reference implementations.
|
||||
Use existing class combinations from [src/transcription/ui/components](../../src/transcription/ui/components) and [src/transcription/ui/pages](../../src/transcription/ui/pages) as reference implementations.
|
||||
|
||||
## 6. Legacy Class Policy
|
||||
Legacy classes with vibe- prefix still exist in a few components and are allowed only for compatibility while migrating:
|
||||
1. Existing usage may remain temporarily.
|
||||
2. New usage of vibe- classes is not allowed.
|
||||
3. When touching a file that uses vibe- classes, prefer migrating it to ui- semantic classes in the same change when safe.
|
||||
|
||||
Current legacy usage examples are in:
|
||||
1. [src/transcription/ui/components/document_panzoom.py](src/transcription/ui/components/document_panzoom.py)
|
||||
2. [src/transcription/ui/components/error_presenter.py](src/transcription/ui/components/error_presenter.py)
|
||||
3. [src/transcription/ui/components/transcript.py](src/transcription/ui/components/transcript.py)
|
||||
Legacy `vibe-` presentation classes are prohibited. Use `ui-` semantic classes from `theme.css`.
|
||||
|
||||
## 7. Prohibited Patterns
|
||||
1. Inline hex colors in Python UI class strings or style blocks, except in isolated bridge code explicitly marked for migration.
|
||||
2. Ad-hoc one-off class names that duplicate existing semantic class intent.
|
||||
3. Page-specific palette forks that bypass theme tokens.
|
||||
4. Hidden or low-contrast focus states on interactive controls.
|
||||
5. Embedded `<style>` blocks or NiceGUI `.style(...)` calls in Python UI code.
|
||||
6. Additional page- or component-specific stylesheets; `theme.css` is the single CSS source.
|
||||
|
||||
## 8. Implementation Rules For Contributors
|
||||
1. Prefer composing existing semantic classes before creating new ones.
|
||||
2. If a new class is required, add it to [src/transcription/ui/static/theme.css](src/transcription/ui/static/theme.css) with a semantic name, then reuse it.
|
||||
2. If a new class is required, add it to [src/transcription/ui/static/theme.css](../../src/transcription/ui/static/theme.css) with a semantic name, then reuse it.
|
||||
3. Keep behavior ownership in Python and appearance ownership in CSS.
|
||||
4. Update UI tests that assert exact text or labels when intentional copy changes are made.
|
||||
5. Avoid introducing class churn unrelated to the feature being changed.
|
||||
@@ -106,7 +104,7 @@ Current legacy usage examples are in:
|
||||
## 9. Verification Checklist
|
||||
Before merging UI changes, verify:
|
||||
1. No new inline hex colors were introduced in UI pages/components.
|
||||
2. New styles are token-backed and added to [src/transcription/ui/static/theme.css](src/transcription/ui/static/theme.css).
|
||||
2. New styles are token-backed and added to [src/transcription/ui/static/theme.css](../../src/transcription/ui/static/theme.css).
|
||||
3. Primary buttons, links, cards, and tables still render with consistent semantics.
|
||||
4. Keyboard focus ring visibility is preserved.
|
||||
5. Relevant UI and integration tests pass.
|
||||
@@ -0,0 +1,184 @@
|
||||
# Production Runbook
|
||||
|
||||
This runbook is the operational checklist for releasing and monitoring the transcription system.
|
||||
|
||||
## 1. Pre-release gate checklist
|
||||
|
||||
1. Run the full suite: `uv run pytest`
|
||||
2. Confirm contract guardrails are green:
|
||||
- `uv run pytest tests/test_meta_contract_guards.py`
|
||||
3. Confirm health endpoint includes worker liveness payload (`/healthz` returns `worker.state`).
|
||||
4. Confirm required runtime settings are present in deployment environment:
|
||||
- `OPENROUTER_API_KEY`
|
||||
- `DATABASE__*`
|
||||
- filesystem paths for data/logs/backups.
|
||||
- `CLOUDFLARE_TUNNEL_TOKEN`
|
||||
5. Confirm schema contract alignment is current:
|
||||
- `src/transcription/db/models.py`
|
||||
- `docs/schema.md`
|
||||
|
||||
## 2. Release execution steps
|
||||
|
||||
1. Deploy artifact/config to target environment.
|
||||
- Production stack: `docker compose -f docker-compose.production.yml up -d --build`
|
||||
- `Settings` loads from explicit `_env_file`, then `ENV_FILE`, then the repository-root `.env.production`; it does not resolve relative to the process working directory.
|
||||
- For Runtime Settings writes in production, mount `.env.production` into the app container and set `RUNTIME_SETTINGS_ENV_FILE=/app/.env.production`.
|
||||
- If deployment uses a non-default env-file location, set both `ENV_FILE` and `RUNTIME_SETTINGS_ENV_FILE` to that absolute path so startup reads and Settings-page writes stay aligned.
|
||||
- For SQLite -> PostgreSQL cutover, run `uv run python tools/export_import_migration.py verify --source-db <sqlite-path-or-url> --target-db <postgres-url>` before switching runtime.
|
||||
2. Validate service startup:
|
||||
- `/healthz` responds `200`
|
||||
- if `RUN_EMBEDDED_WORKER=true`, `worker.state` is `running`
|
||||
- if `RUN_EMBEDDED_WORKER=false`, validate `worker` container is running in Compose
|
||||
(worker healthcheck is intentionally disabled because it does not expose `/healthz`)
|
||||
- validate `cloudflared` logs show active tunnel routes and no ingress errors
|
||||
3. Execute one smoke workflow:
|
||||
- create a document/job with at least one source
|
||||
- verify terminal job outcome updates
|
||||
- verify execution evidence row appended
|
||||
4. Verify log flow:
|
||||
- stdout aggregation receives events
|
||||
- file logs are written under `./data/logs`
|
||||
5. Create a fresh PostgreSQL backup after successful deployment:
|
||||
- `sh deploy/backup/create_postgres_backup.sh` (creates DB dump plus uploads/prompts backups under `BACKUP_DIR`)
|
||||
|
||||
## 3. Rollback triggers and actions
|
||||
|
||||
### Trigger conditions
|
||||
|
||||
1. `/healthz` reports `worker.state=failed`
|
||||
2. Repeated provider timeout/error spikes beyond normal baseline
|
||||
3. Evidence write failures or DB persistence failures
|
||||
|
||||
### Actions
|
||||
|
||||
1. Roll back app artifact and config to previous release.
|
||||
2. Restart service and re-check `/healthz`.
|
||||
3. Re-run smoke workflow and confirm worker returns to `running`.
|
||||
4. Preserve incident evidence:
|
||||
- `./data/logs`
|
||||
- relevant DB rows (`job`, `job_source`, `execution_attempt`)
|
||||
5. If persistence regression is confirmed, restore the latest valid DB dump:
|
||||
- `sh deploy/backup/restore_postgres_backup.sh <dump-file>`
|
||||
- Synology media mirror and paired config snapshot (same timestamp) are restored automatically when present.
|
||||
|
||||
## 4. Post-release monitoring checklist
|
||||
|
||||
## First 24 hours
|
||||
|
||||
1. Monitor `/healthz` periodically for `worker.state`.
|
||||
2. Track job terminal distribution (`transcribed`, `partial_success`, `failed`).
|
||||
3. Sample timeout/error categories for abnormal increase.
|
||||
4. Spot-check new `execution_attempt` records for append-only growth and timing metadata.
|
||||
|
||||
## First 72 hours
|
||||
|
||||
1. Re-check error/timeout trend versus 24h baseline.
|
||||
2. Verify no recurring worker-failed states.
|
||||
3. Verify storage growth and rotation behavior under `./data/logs`.
|
||||
4. Confirm incident response notes are captured for any production anomalies.
|
||||
|
||||
## 5. Operator playbook for common incidents
|
||||
|
||||
### Worker failed
|
||||
|
||||
1. Check `/healthz` payload (`error_id`, `error_category`).
|
||||
2. Locate matching error in logs.
|
||||
3. If non-transient defect persists, roll back.
|
||||
|
||||
### Provider timeout spike
|
||||
|
||||
1. Confirm provider reachability and rate limits.
|
||||
2. Review timeout frequency and impacted job volume.
|
||||
3. If sustained, execute rollback criteria and notify stakeholders.
|
||||
|
||||
### Partial-success increase
|
||||
|
||||
1. Inspect affected `job_source` and `execution_attempt` records.
|
||||
2. Confirm failures are category-aligned (`external`/`timeout`/`internal`).
|
||||
3. Triage whether issue is source quality, provider, or runtime regression.
|
||||
|
||||
### Cloudflare ingress/access failure
|
||||
|
||||
1. Check `cloudflared` container logs for ingress parse, DNS, or auth failures.
|
||||
2. Confirm `deploy/cloudflared/config.yml` hostname mappings are correct.
|
||||
3. Confirm `CLOUDFLARE_TUNNEL_TOKEN` in `.env.production` matches the tunnel configured in Cloudflare.
|
||||
4. Confirm Cloudflare Access app policy includes the intended identity/group for that hostname.
|
||||
5. If logs show SRV/DNS failures via `127.0.0.11`, use the compose-defined resolver override
|
||||
(`dns: 1.1.1.1, 1.0.0.1`) and ensure outbound TCP 443 is allowed.
|
||||
|
||||
### Backup or restore failure
|
||||
|
||||
1. Verify `postgres` container is healthy and accepting connections.
|
||||
2. Confirm dump file exists and is non-zero size.
|
||||
3. Re-run backup/restore scripts with explicit `ENV_FILE` and `COMPOSE_FILE` if using non-default paths.
|
||||
4. If direct Synology copy fails, keep local backup and resolve mount/network before next backup cycle.
|
||||
- For LXC setups, use `deploy/backup/mount_synology_cifs.example.sh` as the persistent mount template.
|
||||
|
||||
## 6. Dependency upgrade policy
|
||||
|
||||
Dependencies are declared in `pyproject.toml` and resolved through the committed
|
||||
`uv.lock`. The lockfile guarantees reproducible installs; the version specifiers
|
||||
control what a deliberate `uv lock --upgrade` is allowed to move.
|
||||
|
||||
### NiceGUI is pinned exactly (`nicegui==3.13.0`)
|
||||
|
||||
1. **Rationale.** NiceGUI bundles Quasar and Vue. Minor releases change component
|
||||
props, slots, and styling, which surfaces as visual and interaction regressions
|
||||
rather than import or type errors. The UI suite under `tests/ui/` asserts
|
||||
structure and behavior, not rendered appearance, so a NiceGUI bump can pass the
|
||||
full test suite and still degrade the interface.
|
||||
2. **Scope of risk.** All NiceGUI usage is confined to `src/transcription/ui/` and
|
||||
uses only the public `nicegui.ui` and `nicegui.events` surfaces. The coupling is
|
||||
shallow, so the pin is about release stability, not about unpicking deep
|
||||
framework entanglement.
|
||||
3. **Current stance.** Hold the exact pin through release stabilization. Do not
|
||||
widen it as incidental cleanup, and do not let automated dependency updates move
|
||||
it. This includes forgoing patch releases, which is the accepted cost.
|
||||
4. **Revisiting.** Treat a NiceGUI upgrade as scheduled work with its own change
|
||||
window: bump the pin deliberately, run `uv run pytest -m "not external"`, then
|
||||
manually verify each page contract in `docs/ui/pages/` before accepting.
|
||||
|
||||
### All other dependencies
|
||||
|
||||
Declared with `>=` floors and moved by explicit `uv lock --upgrade`. Verify with
|
||||
`uv run ruff check .`, `uv run ty check`, and `uv run pytest -q -m "not external"`
|
||||
before committing a changed lockfile.
|
||||
|
||||
## 7. Type-check suppression policy
|
||||
|
||||
`uv run ty check` is a blocking pre-commit gate. Suppressions are allowed only for
|
||||
proven SQLAlchemy descriptor false positives where runtime behavior is correct and
|
||||
the checker cannot represent the descriptor protocol at that call site.
|
||||
|
||||
Every suppression must be:
|
||||
|
||||
1. **Targeted** to a single rule (for example `# ty: ignore[unresolved-attribute]`).
|
||||
2. **Inline** on the expression it suppresses (not file-wide).
|
||||
3. Followed by a **one-line rationale** stating it is a SQLAlchemy descriptor false positive.
|
||||
|
||||
Do not use broad or rationale-free suppressions. If a diagnostic is not a known
|
||||
false positive, fix the code instead of suppressing it.
|
||||
|
||||
## 8. Worker shutdown budget
|
||||
|
||||
Worker shutdown waits for at most:
|
||||
|
||||
`WORKER_PROVIDER_TIMEOUT_SECONDS + WORKER_SHUTDOWN_GRACE_SECONDS`
|
||||
|
||||
`WORKER_PROVIDER_TIMEOUT_SECONDS` covers an in-flight provider call, and
|
||||
`WORKER_SHUTDOWN_GRACE_SECONDS` is extra time for the loop to persist outcomes
|
||||
and exit cleanly after the call returns.
|
||||
|
||||
Set the container or service termination grace period **above this total**
|
||||
budget. If termination grace is shorter, the process may be killed before
|
||||
terminal status and evidence writes are finalized.
|
||||
|
||||
## 9. Horizontal scaling precondition
|
||||
|
||||
Multiple worker replicas can race on execution-attempt numbering for the same
|
||||
`(job_id, source_id)` pair. The runtime now retries boundedly on unique-key
|
||||
conflicts (`uq_execution_attempt_number`) and surfaces a conflict-domain error
|
||||
if retries are exhausted.
|
||||
|
||||
Do not deploy additional worker replicas unless this conflict-retry path and its
|
||||
tests are present and green in the target build.
|
||||
@@ -0,0 +1,118 @@
|
||||
# System Requirements (Current Baseline: V6.1)
|
||||
|
||||
These requirements define the active V6.1 contract and align to current implementation.
|
||||
|
||||
Requirement IDs encode the baseline that introduced them (`REQ-4-*` from V4, `REQ-6-*` from V6) and
|
||||
are stable. Never renumber an existing ID; retire it explicitly instead.
|
||||
|
||||
## Functional Requirements
|
||||
|
||||
### Domain and Record Management
|
||||
|
||||
- **REQ-4-001 Document Registry:** The system must create and update `Document` records with title, type, date metadata, optional location, optional archive identifier, and optional notes.
|
||||
- **REQ-4-002 Source Registry:** The system must create and update `Source` records linked to exactly one `Document`.
|
||||
- **REQ-4-003 People Registry:** The system must create and update `Person` records, support many-to-many links to `Document` with role, and support many-to-many Person tagging via the shared Tag registry.
|
||||
- **REQ-4-004 Registry Semantics:** Document types and person roles must support optional immutable semantic keys and hard-delete only when unreferenced.
|
||||
|
||||
### Job and Workflow Behavior
|
||||
|
||||
- **REQ-4-010 Job Creation:** The system must create `Job` records from uploaded sources and from retranscription of existing sources.
|
||||
- **REQ-4-011 Prompt Snapshotting:** Job creation must persist effective prompt and runtime settings as immutable per-job snapshots.
|
||||
- **REQ-4-012 Queue Membership:** Each `(job, source)` pair must be represented by one `JobSource` row.
|
||||
- **REQ-4-013 Job Status Lifecycle:** `Job.status` must use one of `queued`, `processing`, `transcribed`, `partial_success`, `failed`.
|
||||
- **REQ-4-014 JobSource Status Lifecycle:** `JobSource.status` must use one of `pending`, `transcribed`, `failed`, `cancelled`.
|
||||
- **REQ-4-015 Terminal Job Resolution:** Job terminal status must derive from page outcomes as `transcribed`, `partial_success`, or `failed`.
|
||||
- **REQ-4-016 Cancellation Semantics:** Job cancellation must set remaining `pending` page entries to `cancelled`.
|
||||
|
||||
### Transcription and Evidence
|
||||
|
||||
- **REQ-4-020 Attempt Evidence:** Each provider call must emit one append-only `ExecutionAttempt` record.
|
||||
- **REQ-4-021 Attempt Payload:** `ExecutionAttempt` must retain request manifest/hash, outcome, timing, model/provider fields, and error details when present.
|
||||
- **REQ-4-022 Transport Evidence:** Provider response evidence must be attached to the attempt when a response is available.
|
||||
- **REQ-4-023 Source Projection Rule:** `Source.raw_transcription` is a projection chosen from attempt outcomes and can be repointed by explicit promotion.
|
||||
- **REQ-4-024 Candidate Visibility:** UI must expose candidate attempts with metadata needed for comparative review and selection.
|
||||
|
||||
### Media and Access
|
||||
|
||||
- **REQ-4-030 Ingest Canonicalization:** Stored source bytes may be normalized at ingest (for example orientation correction); stored bytes are the canonical processing source.
|
||||
- **REQ-4-031 Path Safety:** Client-facing media URLs must be generated from controlled application paths only.
|
||||
- **REQ-4-032 Print Media Validation:** Print/export source media must be served through record-validated API routes.
|
||||
|
||||
### Error and UX Contracts
|
||||
|
||||
- **REQ-4-040 Error Envelope:** Service/API errors must map to structured, user-safe error categories and messages.
|
||||
- **REQ-4-041 Partial Failure Visibility:** Mixed page outcomes must be visible at job and page level.
|
||||
- **REQ-4-042 Retry Support:** Failed and cancelled pages must support targeted retranscription without requiring full document recreation.
|
||||
|
||||
### Person Imagery
|
||||
|
||||
- **REQ-6-001 Photo Records:** The system must store reusable `Photo` records that are either owned by a `Person` or unowned for homepage gallery use.
|
||||
- **REQ-6-002 Primary Photo:** At most one photo per owning `Person` may be marked `is_primary`; setting a new primary must clear the previous one.
|
||||
- **REQ-6-003 Primary Reassignment:** Deleting an owner's primary photo must promote a remaining photo of that owner rather than leaving the owner without a primary.
|
||||
|
||||
### Operational Maintenance
|
||||
|
||||
- **REQ-6-010 Queued Maintenance Runs:** Settings-initiated maintenance must persist a `MaintenanceRun` and execute in the worker, not inline in the request that started it.
|
||||
- **REQ-6-011 Maintenance Run Types:** `MaintenanceRun.job_type` must use one of `backup`, `storage_reconciliation`, `gedcom_import`.
|
||||
- **REQ-6-012 Maintenance Status Lifecycle:** `MaintenanceRun.status` must use one of `queued`, `processing`, `succeeded`, `failed`.
|
||||
- **REQ-6-013 Single Claim:** A queued run must be claimed by at most one worker, using a conditional status update rather than read-then-write.
|
||||
- **REQ-6-014 Run History:** Completed runs must retain status, timing, summary, log reference, and error detail, and expose the log for viewing and download.
|
||||
- **REQ-6-015 GEDCOM Upload Import:** Settings must support manual `.ged` upload and queue-backed import into genealogy tables.
|
||||
- **REQ-6-016 GEDCOM Idempotent Upsert:** GEDCOM import must upsert `GenealogyPerson` and `GenealogyFamily` by FamilySearch IDs and avoid duplicate imported citations on re-run.
|
||||
|
||||
### Deployment and Runtime Configuration
|
||||
|
||||
- **REQ-6-020 Production Persistence:** Production must run against PostgreSQL; SQLite remains supported for local development and tests.
|
||||
- **REQ-6-021 Split Worker Deployment:** Production must support running the worker as its own process with the app started at `RUN_EMBEDDED_WORKER=false`.
|
||||
- **REQ-6-022 Runtime Settings Persistence:** Runtime settings edits must persist to the mounted production environment file and survive container restart.
|
||||
- **REQ-6-023 Configuration Contract Sync:** `.env.production.example` must stay synchronized with `Settings` keys and production-safe defaults.
|
||||
- **REQ-6-024 Health Reporting:** The deployed stack must report app and worker health through `/healthz`.
|
||||
- **REQ-6-025 Backup and Restore:** Database and media/config backups must be produced on a host-visible path with a tested restore procedure.
|
||||
|
||||
## Non-Functional Requirements
|
||||
|
||||
- **REQ-4-100 Boundary Integrity:** UI pages/components must not access persistence directly and must call service APIs.
|
||||
- **REQ-4-101 Service Ownership:** Aggregate writes must occur in owning service/workflow modules, not in UI handlers.
|
||||
- **REQ-4-102 Deterministic Loading:** ORM relationship reads in service/UI code must use explicit eager loading compatible with `lazy="raise"`.
|
||||
- **REQ-4-103 Async Safety:** Long-running provider calls must not block UI event handlers directly.
|
||||
- **REQ-4-104 Evidence Durability:** Attempt evidence must survive process restart once the transaction commits.
|
||||
- **REQ-4-105 Test Guardrails:** Architecture boundary tests must remain in place for services and UI boundaries.
|
||||
|
||||
## Requirement Interpretation Notes
|
||||
|
||||
### Status and lifecycle semantics
|
||||
|
||||
- `REQ-4-013` and `REQ-4-015` intentionally bind success to `transcribed`, not a generic `completed`, so docs, tests, and runtime transitions stay consistent.
|
||||
- `REQ-4-016` and `REQ-4-042` distinguish cancellation from failure at page level (`cancelled` vs `failed`) while still allowing targeted retranscription.
|
||||
|
||||
### Evidence semantics
|
||||
|
||||
- `REQ-4-020` through `REQ-4-024` separate authoritative history (`ExecutionAttempt`) from operational projection (`Source.raw_transcription`).
|
||||
- This supports immutable provenance while allowing explicit candidate promotion for operator workflows.
|
||||
|
||||
### Boundary and loading semantics
|
||||
|
||||
- `REQ-4-100` and `REQ-4-101` codify aggregate/service ownership and keep UI out of persistence concerns.
|
||||
- `REQ-4-102` exists to enforce deterministic query shape under `lazy="raise"` and avoid hidden data access in rendering callbacks.
|
||||
|
||||
## Verification Anchors
|
||||
|
||||
- Service boundary enforcement: `tests/test_service_boundaries.py`
|
||||
- UI boundary enforcement: `tests/test_ui_boundaries.py`
|
||||
- Job lifecycle reliability and terminal status behavior: `tests/services/test_workflows_reliability.py`
|
||||
- Evidence append-only and projection behavior: `tests/services/test_store.py`, `tests/services/test_transcription_service.py`
|
||||
- Person photo ownership and primary selection: `tests/services/test_photo_service.py`
|
||||
- Maintenance run lifecycle and worker execution: `tests/services/test_maintenance_service.py`
|
||||
- Runtime settings persistence: `tests/ui/test_runtime_settings_store.py`, `tests/services/test_settings_services.py`
|
||||
- Configuration contract synchronization: `tests/test_meta_contract_guards.py`
|
||||
- Deployment health reporting: `tests/api/test_health.py`
|
||||
|
||||
## Traceability Notes
|
||||
|
||||
- Source of truth for status enums:
|
||||
- `src/transcription/db/models.py`
|
||||
- Source of truth for workflow transitions:
|
||||
- `src/transcription/services/workflows.py`
|
||||
- `src/transcription/services/jobs.py`
|
||||
- Source of truth for attempt evidence writes:
|
||||
- `src/transcription/services/sources.py`
|
||||
@@ -0,0 +1,488 @@
|
||||
# Architecture & Code Review Report
|
||||
|
||||
**Repository Target:** `transcription/`
|
||||
**Target Stack:** Python 3.12+ | FastAPI | NiceGUI | SQLModel/SQLAlchemy | Pydantic V2 | asyncio | OpenRouter
|
||||
|
||||
**Review date:** 2026-08-23
|
||||
**Governing procedure:** `.github/skills/python-code-reviewer/skill.md`
|
||||
**Escalations applied:** `.github/skills/evidence-provenance-auditor/skill.md`, `.github/skills/test-effectiveness-auditor/skill.md`
|
||||
**Scope:** 77 Python modules / ~13k LOC under `src/transcription`, 57 test files (377 collected non-external tests), 23 documents under `docs/`, 9 active rule files.
|
||||
|
||||
> **Status: closed.** Every finding below was remediated in the phases following this
|
||||
> review. This document is retained as a record of the reasoning, **not** as a list of
|
||||
> open work, and it is not canonical authority.
|
||||
>
|
||||
> Two recommendations were wrong on contact and were corrected during implementation:
|
||||
> the HIGH-03 fix as written would have stripped root-cause data from `ExecutionAttempt`
|
||||
> provenance, and the HIGH-01 fix needed to preserve per-page durability that the report
|
||||
> did not mention. Where this text and the current code or guard tests disagree, the code
|
||||
> and tests are correct.
|
||||
|
||||
### Verification commands and outcomes
|
||||
|
||||
| Command | Outcome |
|
||||
| :--- | :--- |
|
||||
| `uv run ruff check .` | **Pass** — `All checks passed!` |
|
||||
| `uv run pytest -q -m "not external"` | **Pass** — 377 passed |
|
||||
| `uv run ty check` | **10 diagnostics** — all SQLModel/SQLAlchemy column-descriptor false positives (`services/photos.py` ×8, `tests/test_storage_reconciliation.py` ×2). Advisory only; no suppression strategy exists. |
|
||||
|
||||
---
|
||||
|
||||
## 1. Executive Summary
|
||||
|
||||
- **Overall health is good.** The codebase has genuine architectural discipline: layered `ui → services → db`, a single Pydantic-V2 settings source, an atomic compare-and-swap job claim, append-only evidence history, and eleven deterministic guard tests that enforce structural rules rather than describing them.
|
||||
- **No Critical findings.** The highest-risk category for this domain — secret leakage into stored provenance — was explicitly audited and **passes**: request headers are never persisted, response headers use an allowlist, and the API key is `SecretStr` end-to-end.
|
||||
- **The top risk is a transaction-atomicity violation on the worker hot path.** Page evidence and terminal job status commit in two separate transactions (`workflows.py:549-598`), directly contradicting `services.instructions.md`. A crash between them leaves a transcript persisted against a job stuck in `PROCESSING`.
|
||||
- **That violation is invisible to the test suite.** The test-effectiveness audit confirms no test can fail on a split commit — the pipeline tests assert the happy-path end state, which passes either way. The invariant is documented and steered but *not enforced*.
|
||||
- **Stale-job recovery is startup-only** (`app.py:79`), with a 30-second staleness threshold. A job orphaned shortly before a fast restart is not recovered and remains `PROCESSING` indefinitely, because the worker only claims `QUEUED` rows.
|
||||
- **The mandated error-presentation boundary is bypassed at 8 sites.** `home_page.py` and `people_page.py` hand-roll `ui.notify(str(exc), ...)`, discarding the `error_id`, category, and suggestion that `error_presenter.show_error` provides. `people_page.py` imports the correct helpers and still bypasses them.
|
||||
- **User-facing output can leak filesystem paths.** `classify_unexpected_error` (`errors.py:94`) interpolates the raw exception into a message rendered in the UI; a SQLAlchemy `OperationalError` embeds the database file path. This contradicts an explicit rule in `error-handling.instructions.md`.
|
||||
- **The retry gate ignores error category** (`workflows.py:185`), so non-retriable faults would be requeued. Currently latent because `worker_max_retries` defaults to `0`.
|
||||
- **Highest-leverage work is enforcement, not refactoring.** Two atomicity tests, a `ty` suppression strategy that lets the pre-commit hook become blocking, and `ruff format --check` in the gate would convert three documented-but-unenforced invariants into deterministic ones.
|
||||
|
||||
---
|
||||
|
||||
## 2. Executive Architecture Assessment
|
||||
|
||||
**Verdict: architecturally sound with a concentrated reliability gap in the worker's commit boundary.**
|
||||
|
||||
Domain cohesion is strong. The `services/` layer owns transactions and business rules, `ui/` owns presentation, `db/` owns schema, and `providers/` isolates the OpenRouter adapter behind a `TranscriptionProvider` protocol. Dependency direction is correct and — unusually — *mechanically enforced*: `test_service_boundaries.py` AST-scans for service-to-service imports and `test_ui_boundaries.py` scans pages/components for persistence access. Provider details do not leak upward; `workflows.py` imports only the abstract `providers` types, never `openrouter`.
|
||||
|
||||
The evidence/provenance model is the strongest part of the system. `ExecutionAttempt` is genuinely append-only, retries append rather than rewrite, projection writes onto `JobSource` are clearly distinguished from history mutation, and all 14 provenance-auditor invariant checks pass.
|
||||
|
||||
**Top systemic risks:**
|
||||
|
||||
1. **Split commit boundary on the worker path (High).** Evidence durability and job terminal status are two transactions. This is the one place where the architecture's own written contract is contradicted by the implementation, on the hottest path in the system.
|
||||
2. **Recovery is a startup-only, time-thresholded sweep (Medium).** There is no runtime reconciliation, so the self-healing property depends on restart cadence rather than on a bounded interval.
|
||||
3. **Enforcement coverage has known holes (Medium).** Atomicity, error-presenter usage, and formatting are all documented rules with no deterministic test. The repo's own strength — routing invariants into tests — has not been applied to these three.
|
||||
4. **Leaky transaction ownership (Medium).** `workflows.py` reaches into `services.jobs._session_scope()` and `services.sources._session_scope()` — private members of two different services — to open transactions. Session ownership is ambiguous exactly where it most needs to be explicit.
|
||||
5. **A 10-diagnostic type-checker baseline with no suppression policy (Low).** The signal is currently ignorable, which means a real regression would blend into the noise.
|
||||
|
||||
---
|
||||
|
||||
## 3. Findings by Severity
|
||||
|
||||
### Critical Severity
|
||||
|
||||
**None identified.**
|
||||
|
||||
The secret-leakage check — the only plausible Critical for this system — passes explicitly. `OpenRouterProvider` stores an allowlisted subset of *response* headers only (`providers/evidence.py:130-134`, `SAFE_RESPONSE_HEADERS`); request headers containing `Authorization` are never captured into `TransportEvidence`; and the key is held as `SecretStr` from `config.py` through to the client. Append-only evidence history is likewise intact and test-enforced.
|
||||
|
||||
---
|
||||
|
||||
### High Severity
|
||||
|
||||
#### [HIGH-01] Page evidence and terminal job status commit in separate transactions
|
||||
|
||||
- **Location:** `src/transcription/services/workflows.py:549-565` (`_finalize_batch_outcome`), `src/transcription/services/workflows.py:584-598` (`_persist_page_outcome`)
|
||||
- **Problem & Consequence:** `.github/instructions/services.instructions.md` states: *"Never commit transcript updates separately from the paired terminal/retry job status change."* The implementation does exactly that. `_persist_page_outcome` opens its own scope and commits page evidence (line 592-594); `_finalize_batch_outcome` later opens a *second* scope and commits the terminal `JobStatus` (line 558-560). For a single-page job these are two transactions with a window between them. A process crash, container eviction, or unhandled error in that window persists the transcript while the job remains `PROCESSING`. Because the worker only claims `QUEUED` rows, that job is not reprocessed; it is recoverable only by the startup sweep, and only if it has aged past the staleness threshold (see MED-01). The user sees a job that never completes despite the transcription having succeeded and been billed.
|
||||
|
||||
This is a deliberate design tension, not an oversight: `_persist_page_outcome_durably` (line 568-581) wraps the page write in `asyncio.shield` precisely so per-page evidence survives cancellation mid-batch. That goal is correct for *multi*-page jobs. The defect is that the single-page and final-page cases inherit the split unnecessarily.
|
||||
|
||||
- **Recommendation:** Keep per-page durability for intermediate pages, but commit the final page outcome and the terminal status in one transaction.
|
||||
|
||||
```python
|
||||
# Before — two scopes, two commits
|
||||
await _persist_page_outcome_durably(job=job, services=services, page=page, session=None)
|
||||
...
|
||||
await _finalize_batch_outcome(job=job, services=services, status=status, session=None)
|
||||
|
||||
# After — final page and terminal status share one transaction
|
||||
async with services.jobs.session_scope() as tx:
|
||||
for page in intermediate_pages:
|
||||
await _persist_page_outcome_durably(job=job, services=services, page=page, session=None)
|
||||
await _write_page_outcome(job=job, services=services, page=final_page, session=tx)
|
||||
await services.jobs.mark_job_status(job.id, status, session=tx)
|
||||
await tx.commit()
|
||||
```
|
||||
|
||||
Pair this with the atomicity test in HIGH-04 so the boundary cannot silently regress.
|
||||
- **Effort:** M
|
||||
|
||||
---
|
||||
|
||||
#### [HIGH-02] Mandated error-presentation boundary bypassed at 8 sites
|
||||
|
||||
- **Location:** `src/transcription/ui/pages/home_page.py:212,220,228,255`; `src/transcription/ui/pages/people_page.py:265,321,330,339`
|
||||
- **Problem & Consequence:** `.github/instructions/ui.instructions.md:42` requires all user-facing error display to route through `components/error_presenter.py`. Seven of nine pages comply. These two hand-roll `ui.notify(str(exc), type="negative")`. The consequence is not cosmetic: `show_error` (`error_presenter.py:52-67`) surfaces the correlation `error_id`, the canonical error category, and the actionable `suggestion` field. Bypassing it means a user hitting a failure on the home or people page gets a bare exception string with **no error reference to report**, making these two pages unsupportable in production — precisely the pages most likely to be a user's entry point.
|
||||
|
||||
`people_page.py` already imports `run_ui_action` and `show_error` at lines 28-29 and uses them elsewhere in the same module, so the bypass is inconsistency rather than missing infrastructure.
|
||||
- **Recommendation:** Replace each site with the canonical helper. The unused `summarize_error` helper in `error_presenter.py` (currently a retained orphan — see LOW-07) is the natural fit where a compact string is genuinely needed.
|
||||
|
||||
```python
|
||||
# Before
|
||||
except AppError as exc:
|
||||
ui.notify(str(exc), type="negative")
|
||||
|
||||
# After
|
||||
except AppError as exc:
|
||||
show_error(exc)
|
||||
```
|
||||
|
||||
Then close the hole permanently by extending `tests/test_ui_boundaries.py` with an AST check that no module under `PAGES_DIR` calls `ui.notify(...)` with `type="negative"`.
|
||||
- **Effort:** S
|
||||
|
||||
---
|
||||
|
||||
#### [HIGH-03] Unexpected-error path leaks filesystem paths into user-facing output
|
||||
|
||||
- **Location:** `src/transcription/errors.py:91-98` (line 94), rendered via `src/transcription/ui/components/error_presenter.py:52-67`
|
||||
- **Problem & Consequence:** `classify_unexpected_error` builds `f"Unexpected error during {operation}: {exc}"` and stores it as `AppError.message`. `show_error` renders `error.message` directly to the user. Any exception whose `str()` contains infrastructure detail is therefore displayed verbatim — a SQLAlchemy `OperationalError` embeds the absolute SQLite database path, and an `OSError` from the media layer embeds the storage root. `.github/instructions/error-handling.instructions.md:74` states: *"Never leak … local filesystem paths in user-facing output."* This is the generic catch-all path, so it applies to every unanticipated failure across the application.
|
||||
- **Recommendation:** Split the diagnostic detail from the user-facing message. Log the full exception with the `error_id` as the correlation key; show the user a stable message plus that id.
|
||||
|
||||
```python
|
||||
# Before
|
||||
return AppError(
|
||||
f"Unexpected error during {operation}: {exc}",
|
||||
category=ErrorCategory.INTERNAL_UNEXPECTED,
|
||||
...
|
||||
)
|
||||
|
||||
# After
|
||||
error = AppError(
|
||||
f"Unexpected error during {operation}.",
|
||||
category=ErrorCategory.INTERNAL_UNEXPECTED,
|
||||
suggestion="Retry once. If it persists, report the error reference id.",
|
||||
retriable=False,
|
||||
)
|
||||
logger.exception("error_id=%s operation=%s", error.error_id, operation)
|
||||
return error
|
||||
```
|
||||
|
||||
Add a case to `tests/ui/test_error_presenter.py` asserting that a raised `OperationalError` carrying a path does not surface that path in the rendered message.
|
||||
- **Effort:** S
|
||||
|
||||
---
|
||||
|
||||
#### [HIGH-04] Transaction-atomicity invariants have no enforcing test
|
||||
|
||||
- **Location:** Contract at `.github/instructions/services.instructions.md` §"Workflow Transaction Boundaries"; gap confirmed across `tests/integration/test_pipeline_flow.py:66-160` and `tests/services/test_job_service.py:41-59`
|
||||
- **Problem & Consequence:** The test-effectiveness audit establishes that **neither** Transaction B (transcript + `TRANSCRIBED`) nor Transaction C (retry: `error_detail` + `retry_count` + `QUEUED`) is enforced. The existing pipeline test asserts the final state after a successful run — which passes identically whether the writes shared one commit or used two. To fail on a split-commit regression a test must inject a fault *between* the writes; no such test exists.
|
||||
|
||||
The consequence is that HIGH-01 shipped undetected and any future refactor of `advance_job` can reintroduce it just as silently. This is a *governance* failure rather than a code defect: the repo's stated model is that hard rules belong in deterministic tests, and this rule is the most consequential one that never made the transition.
|
||||
- **Recommendation:** Add `tests/integration/test_pipeline_atomicity.py` with two tests that patch the session to raise after `flush()` but before `commit()`, then assert that *neither* side of the pair is visible in a fresh session. These tests should **fail against the current implementation** and pass once HIGH-01 is fixed — write them first.
|
||||
- **Effort:** M
|
||||
|
||||
---
|
||||
|
||||
### Medium Severity
|
||||
|
||||
#### [MED-01] Stale-job recovery runs only at startup, behind a 30-second threshold
|
||||
|
||||
- **Location:** `src/transcription/app.py:71-81` (`_recover_stale_processing_jobs`), sole caller at `app.py:79` inside `_lifespan`
|
||||
- **Problem & Consequence:** `requeue_stale_processing_jobs` has exactly one call site, in the lifespan startup handler. There is no runtime re-check. The staleness predicate is `updated_at < now - worker_provider_timeout_seconds` (default **30.0s**, `config.py:116`). A job orphaned less than 30 seconds before a fast container restart therefore fails the predicate at the only moment recovery is attempted, and stays `PROCESSING` forever — the worker claims only `QUEUED` rows. It self-heals only on some *later, unrelated* restart. In a frequently-redeployed environment, restarts are exactly when orphans are created, so the recovery window is systematically misaligned with the failure it exists to handle.
|
||||
- **Recommendation:** Move the sweep onto a periodic task in the worker loop (e.g. every `max(30, provider_timeout * 2)` seconds) in addition to the startup call, and derive the threshold from a dedicated `worker_stale_job_seconds` setting rather than reusing the provider timeout, so the two can be tuned independently.
|
||||
- **Effort:** M
|
||||
|
||||
---
|
||||
|
||||
#### [MED-02] Retry gate ignores `error_category`, so non-retriable failures would be requeued
|
||||
|
||||
- **Location:** `src/transcription/services/workflows.py:184-194`
|
||||
- **Problem & Consequence:** The `JobStatus.FAILED` branch gates solely on `job.retry_count < settings.worker_max_retries`. It does not consult `error_category` or the `AppError.retriable` flag. `.github/instructions/error-handling.instructions.md` classifies `validation`, `not_found`, and `conflict` as non-retriable; under this gate a malformed source or a missing record would be retried to exhaustion, consuming provider quota on calls that cannot succeed and delaying the terminal failure the user needs to see. There is also no backoff — retries requeue immediately.
|
||||
|
||||
Currently **latent**: `worker_max_retries` defaults to `0` (`config.py:113`) and is commented out in `.env`, so the branch always falls through to the max-retries log. It becomes live the moment anyone enables retries.
|
||||
- **Recommendation:** Gate on retriability *and* count, and add exponential backoff before requeue.
|
||||
|
||||
```python
|
||||
case JobStatus.FAILED:
|
||||
if job.error_category in NON_RETRIABLE_CATEGORIES:
|
||||
logger.error("Job %s failed non-retriably (%s).", job.id, job.error_category)
|
||||
return
|
||||
if job.retry_count < settings.worker_max_retries:
|
||||
...
|
||||
```
|
||||
|
||||
Cover with a test that a `validation`-category failure is not requeued even when `worker_max_retries > 0`.
|
||||
- **Effort:** S
|
||||
|
||||
---
|
||||
|
||||
#### [MED-03] `IntegrityError` on the attempt-number flush is uncaught, risking evidence loss
|
||||
|
||||
- **Location:** `src/transcription/services/sources.py:540-546` (attempt-number computation), `sources.py:587` (unguarded `flush()`)
|
||||
- **Problem & Consequence:** `attempt_number` is derived read-then-write as `MAX(attempt_number) + 1`, and `uq_execution_attempt_number` enforces uniqueness (`db/models.py:507`, documented at `docs/schema.md:273`). The sibling `JobSource` insert *does* catch `IntegrityError` (`sources.py:531-534`), but the `ExecutionAttempt` flush at line 587 does not. Two concurrent attempt writes for the same job source would raise an unhandled `IntegrityError` and lose an evidence row — the one class of data this system exists to preserve. Not currently reachable: the worker is single-instance and processes sources sequentially. It becomes reachable the moment a second worker replica is deployed.
|
||||
- **Recommendation:** Mirror the `JobSource` handling — catch `IntegrityError`, recompute `MAX(attempt_number) + 1`, and retry the insert a bounded number of times, raising a domain error on exhaustion. Note this constraint as a horizontal-scaling precondition in `docs/production-runbook.md`.
|
||||
- **Effort:** M
|
||||
|
||||
---
|
||||
|
||||
#### [MED-04] Shutdown timeout is shorter than the provider timeout
|
||||
|
||||
- **Location:** `src/transcription/worker.py:146` (`asyncio.wait_for(worker_task, timeout=2.0)`); provider timeout at `config.py:116` (default 30.0s)
|
||||
- **Problem & Consequence:** Graceful shutdown waits 2 seconds for the worker task, but the stop event is only checked *between* jobs and an in-flight provider call may run for up to 30 seconds. Any shutdown during a provider call therefore cancels mid-flight. Combined with HIGH-01's split commit, a cancellation that lands between the evidence commit and the status commit produces exactly the stuck-`PROCESSING` state described there — so this finding materially raises HIGH-01's probability rather than being independent of it.
|
||||
- **Recommendation:** Derive the shutdown budget from the provider timeout (`worker_provider_timeout_seconds + small_grace`) instead of hardcoding `2.0`, and ensure the container's termination grace period exceeds it. Document both in `docs/production-runbook.md`.
|
||||
- **Effort:** S
|
||||
|
||||
---
|
||||
|
||||
#### [MED-05] `workflows.py` reaches into two services' private `_session_scope`
|
||||
|
||||
- **Location:** `src/transcription/services/workflows.py:558` (`services.jobs._session_scope()`), `workflows.py:592` (`services.sources._session_scope()`)
|
||||
- **Problem & Consequence:** The orchestration module opens transactions by calling a private member on two different service objects. This is the concrete mechanism behind HIGH-01: because transaction ownership is expressed through a private back-door rather than a declared boundary, nothing in the design makes it obvious that two scopes are being opened for one logical unit of work. It also couples `workflows.py` to a service implementation detail that `test_service_boundaries.py` cannot see (it checks imports, not attribute access).
|
||||
- **Recommendation:** Promote a single explicit transaction entry point — a `session_scope()` on `ServiceBundle`, or a module-level `unit_of_work(services)` helper — and make `workflows.py` use only that. Extend `test_service_boundaries.py` with an AST check forbidding `_session_scope` attribute access outside the owning service module.
|
||||
- **Effort:** M
|
||||
|
||||
---
|
||||
|
||||
### Low Severity
|
||||
|
||||
#### [LOW-01] `hashlib.sha256` over full file bytes runs on the event loop
|
||||
- **Location:** `src/transcription/services/store.py:401`
|
||||
- **Problem & Consequence:** Digest computation is CPU-bound and synchronous inside an `async def`. For large uploads this blocks the loop, stalling both the NiceGUI UI and the worker. Every sibling I/O path in the codebase correctly uses `asyncio.to_thread` (`media_storage.py:43`, `normalization.py:117`, `photos.py:176`, `sources.py:740,753`), so this is an isolated deviation.
|
||||
- **Recommendation:** `digest = await asyncio.to_thread(lambda: hashlib.sha256(file_bytes).hexdigest())`.
|
||||
- **Effort:** S
|
||||
|
||||
#### [LOW-02] `homepage_store.py` performs synchronous file I/O from async callers
|
||||
- **Location:** `src/transcription/ui/homepage_store.py:25,32`; called from `src/transcription/ui/pages/home_page.py:170`
|
||||
- **Problem & Consequence:** Same class as LOW-01 — reads/writes the homepage JSON directly rather than via `asyncio.to_thread`. Impact is small (a tiny file), but it is a second deviation from an otherwise universal convention.
|
||||
- **Recommendation:** Wrap both calls in `asyncio.to_thread`.
|
||||
- **Effort:** S
|
||||
|
||||
#### [LOW-03] Worker poll interval is hardcoded outside `Settings`
|
||||
- **Location:** `src/transcription/app.py:62` (`poll_interval_seconds=1.0`)
|
||||
- **Problem & Consequence:** The single operational knob controlling worker latency-vs-load cannot be tuned without a code change, contradicting the otherwise-clean rule that all configuration lives in `config.py` (zero `os.getenv` calls exist outside it).
|
||||
- **Recommendation:** Add `worker_poll_interval_seconds: float = 1.0` to `Settings` and read it at the call site.
|
||||
- **Effort:** S
|
||||
|
||||
#### [LOW-04] `_build_request_manifest` returns `None` silently, producing incomplete evidence
|
||||
- **Location:** `src/transcription/providers/openrouter.py:347`
|
||||
- **Problem & Consequence:** When `source_reference is None` the manifest is skipped with no log line. The attempt is still recorded but its provenance is quietly incomplete, and there is no signal that it happened — the failure mode is undetectable after the fact.
|
||||
- **Recommendation:** Log at `warning` with the job/source identifiers before returning `None`, so incomplete provenance is at least attributable.
|
||||
- **Effort:** S
|
||||
|
||||
#### [LOW-05] Ten `ty` diagnostics with no suppression strategy
|
||||
- **Location:** `src/transcription/services/photos.py` (8), `tests/test_storage_reconciliation.py` (2)
|
||||
- **Problem & Consequence:** All ten are SQLModel/SQLAlchemy false positives — column descriptors are typed as their Python value type (`UUID`, `datetime`, `bool`), so `.is_()`, `.asc()`, `func.count()`, and `group_by()` appear invalid. Because there is no suppression policy, the pre-commit hook must run `ty` in advisory mode, which means a *genuine* new type error would print alongside the known ten and block nothing.
|
||||
- **Recommendation:** Add targeted `# ty: ignore[...]` comments with a one-line rationale at each of the ten sites, then flip the pre-commit hook to blocking. This converts a permanently-ignored signal into a real gate.
|
||||
- **Effort:** M
|
||||
|
||||
#### [LOW-06] `ruff format` is not enforced; 35 files have drifted
|
||||
- **Location:** `.pre-commit-config.yaml`, `ruff.toml`
|
||||
- **Problem & Consequence:** `ruff check` is blocking but `ruff format --check` is absent from the gate, so formatting drift accumulates silently and inflates unrelated diffs whenever anyone does run the formatter.
|
||||
- **Recommendation:** Run `uv run ruff format .` once as a single isolated commit, then add `ruff format --check` to the pre-commit gate.
|
||||
- **Effort:** S
|
||||
|
||||
#### [LOW-07] Four retained orphans, all recorded as "uncertain — follow-up"
|
||||
- **Location:** `tests/test_orphan_sweep.py:33-52` (`KNOWN_ORPHANS`): `BenchmarkManifest`, `dispose_all_engines`, `refresh_engine`, `summarize_error`
|
||||
- **Problem & Consequence:** Every entry carries the weakest possible justification. `summarize_error` is the notable one: it is an unused helper in `error_presenter.py` *while two pages hand-roll error display* (HIGH-02) — the orphan and the boundary violation are the same problem viewed from two directions. `dispose_all_engines` / `refresh_engine` are plausibly test-support utilities and should be classified as such rather than left uncertain.
|
||||
- **Recommendation:** Resolve each to a definite outcome — `summarize_error` becomes used by the HIGH-02 fix; classify the engine helpers as test-support or delete them; decide on `BenchmarkManifest`.
|
||||
- **Effort:** S
|
||||
|
||||
#### [LOW-08] Orphan sweep only scans module-level public definitions
|
||||
- **Location:** `tests/test_orphan_sweep.py`
|
||||
- **Problem & Consequence:** Methods and private functions are out of scope, so dead code inside classes — the most common kind in a service-oriented codebase — is structurally invisible to the sweep.
|
||||
- **Recommendation:** Extend the AST walk to public methods on service classes, seeding `KNOWN_ORPHANS` with the current result set to keep the change non-breaking.
|
||||
- **Effort:** M
|
||||
|
||||
#### [LOW-09] f-string interpolation in logging calls
|
||||
- **Location:** `src/transcription/services/workflows.py:193` and similar sites
|
||||
- **Problem & Consequence:** `logger.error(f"Job {job.id} has failed...")` formats eagerly regardless of level and prevents structured-logging backends from grouping by template. Ruff's `flake8-logging-format` (`G`) rules are not enabled, so this is unenforced.
|
||||
- **Recommendation:** Use `logger.error("Job %s has failed and reached max retries.", job.id)` and enable ruff rule set `G`.
|
||||
- **Effort:** S
|
||||
|
||||
#### [LOW-10] Low-signal and always-true assertions in the test suite
|
||||
- **Location:** `tests/test_traceability.py:54-57`; `tests/integration/test_pipeline_flow.py:135-140,446-452`; `tests/test_orphan_sweep.py:119`; `tests/services/test_workflows_reliability.py:105,178,241,317,375`
|
||||
- **Problem & Consequence:** Per the test-effectiveness audit: `test_traceability.py:54-57` asserts properties of dict literals defined in the same file (can only fail if the test itself is edited); `assert processed is True` in the pipeline tests is unfalsifiable because `read_job` raises rather than returning `None`; the `>= 200` orphan threshold is a historical snapshot that tolerates ±40 drift; and the `assert result is not None` guards are shadowed by the attribute assertions that follow. Together these overstate effective coverage.
|
||||
- **Recommendation:** Apply the prune/strengthen backlog in §6 (Testing).
|
||||
- **Effort:** S
|
||||
|
||||
#### [LOW-11] Wall-clock timing dependencies risk CI flakiness
|
||||
- **Location:** `tests/services/test_workflows_reliability.py:157-196` (real `time.sleep(0.40)`, upper bound `< 540ms` with only 10% slack); `test_workflows_reliability.py:341` (`asyncio.wait_for(..., timeout=2)`)
|
||||
- **Problem & Consequence:** On a loaded CI runner, a 200ms asyncio task plus 400ms blocking setup can exceed the 540ms bound, producing false failures that erode trust in the suite.
|
||||
- **Recommendation:** Widen the slack factor to `0.8` or replace the blocking sleep with a controlled clock mock.
|
||||
- **Effort:** S
|
||||
|
||||
---
|
||||
|
||||
## 4. Architectural Drift & Gap Analysis
|
||||
|
||||
`Direction` is `doc->code` (implementation must change to match documented intent) or `code->doc` (an undocumented but repeatable convention that should be formalized).
|
||||
|
||||
| Area / Component | Direction | Documented / Intended Rule | Actual Implementation State | Severity | Recommended Resolution |
|
||||
| :--- | :--- | :--- | :--- | :--- | :--- |
|
||||
| Worker commit boundary | `doc->code` | `services.instructions.md`: never commit transcript updates separately from the paired terminal status change | `workflows.py:549-598` commits page evidence and terminal status in two separate sessions | High | Fix per HIGH-01; enforce per HIGH-04 |
|
||||
| UI error presentation | `doc->code` | `ui.instructions.md:42`: all user-facing error display routes through `error_presenter.py` | 8 hand-rolled `ui.notify` sites in `home_page.py` and `people_page.py` | High | Fix per HIGH-02; add AST guard to `test_ui_boundaries.py` |
|
||||
| Unexpected-error messaging | `doc->code` | `error-handling.instructions.md:74`: never leak local filesystem paths in user-facing output | `errors.py:94` interpolates raw `exc` into the rendered message | High | Fix per HIGH-03 |
|
||||
| Retry policy | `doc->code` | `error-handling.instructions.md`: validation / not_found / conflict are non-retriable | `workflows.py:185` gates on retry count only | Medium | Fix per MED-02 |
|
||||
| Stale-job recovery | `code->doc` | Not documented as startup-only or time-thresholded | Single startup call site; 30s threshold reuses the provider timeout | Medium | Fix per MED-01, then document the recovery contract in `docs/production-runbook.md` |
|
||||
| Transaction ownership | `code->doc` | `services.instructions.md` assigns transaction ownership to services | `workflows.py` opens scopes via two services' private `_session_scope` | Medium | Fix per MED-05; document the single unit-of-work entry point |
|
||||
| Blocking-I/O convention | `code->doc` | Not stated as a rule; followed at 5 of 7 sites | `store.py:401` and `homepage_store.py:25,32` deviate | Low | Fix per LOW-01/LOW-02, then state the `asyncio.to_thread` rule in `services.instructions.md` |
|
||||
| Configuration centralization | `code->doc` | Zero `os.getenv` outside `config.py` — a real, held convention | Held everywhere except the hardcoded `poll_interval_seconds` at `app.py:62` | Low | Fix per LOW-03, then formalize the rule and add a deterministic guard |
|
||||
| Type-check baseline | `code->doc` | No documented policy for `ty` diagnostics | 10 tolerated false positives; hook is advisory-only | Low | Adopt the suppression strategy in LOW-05 and document it |
|
||||
| Formatting | `code->doc` | `ruff.toml` configures the formatter | `ruff format --check` absent from the gate; 35 files drifted | Low | Fix per LOW-06 |
|
||||
| Dependency pin | — | `docs/production-runbook.md` "Dependency upgrade policy" records the exact `nicegui==3.13.0` pin as a deliberate stability decision | Matches | — | **No action** — correctly documented, not a defect |
|
||||
|
||||
---
|
||||
|
||||
## 5. Invariant Inventory & Routing Recommendations
|
||||
|
||||
| Invariant / Constraint | Current Location | Recommended Target Layer | Rationale |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| Transcript + terminal status commit atomically | Instructions only | **Deterministic test** (`tests/integration/test_pipeline_atomicity.py`) | Highest-consequence rule in the system with zero enforcement; steering alone already failed to prevent HIGH-01 |
|
||||
| Retry writes commit atomically | Instructions only | **Deterministic test** (same file) | Same class; a partial retry commit corrupts `retry_count` accounting |
|
||||
| All UI errors route through `error_presenter` | Instructions (`ui.instructions.md:42`) | **Deterministic test** (extend `test_ui_boundaries.py`) | Mechanically checkable via AST; 8 live violations prove instructions are insufficient here |
|
||||
| No filesystem paths in user-facing output | Instructions (`error-handling.instructions.md:74`) | **Deterministic test** (extend `tests/ui/test_error_presenter.py`) | Checkable by asserting a path-bearing exception does not surface its path |
|
||||
| Non-retriable categories are never requeued | Instructions | **Deterministic test** (`tests/services/test_workflows_reliability.py`) | Latent today; a test freezes the correct behavior before retries are enabled |
|
||||
| Blocking I/O runs via `asyncio.to_thread` | Convention only (5/7 sites) | **Instructions** (`services.instructions.md`) | Judgment-dependent (thresholds vary by payload size); steering fits better than a hard test |
|
||||
| Transaction opened through one owned entry point | Convention, violated | **Instructions + test** | Document the entry point; AST-guard against `_session_scope` access outside its owning module |
|
||||
| Append-only `ExecutionAttempt` history | Docs + 3 tests | **Keep as-is** | Correctly routed and genuinely mutation-sensitive; the model to imitate |
|
||||
| Service/UI boundary rules | Instructions + 2 AST tests | **Keep as-is** | Working exactly as intended |
|
||||
| Status vocabulary conformance | `docs/schema.md` + contract guards | **Keep as-is** | Enum drift would fail the suite |
|
||||
| No secrets in stored evidence | Docs + provenance skill + allowlist in code | **Keep as-is** | Allowlist is the right mechanism — fails closed by construction |
|
||||
| `ty` diagnostic suppression policy | Nonexistent | **Docs + blocking hook** | Needs a written rationale per suppression before the gate can be trusted |
|
||||
| NiceGUI exact pin | `docs/production-runbook.md` | **Keep as-is** | Deliberate, documented, correctly excluded from review findings |
|
||||
|
||||
---
|
||||
|
||||
## 6. Stack-Specific Analysis
|
||||
|
||||
### Python 3.12+ Best Practices
|
||||
Modern syntax is used consistently: `X | None` unions throughout, builtin generics, no `typing.List`/`Optional` legacy forms, `pathlib` over `os.path`. Type-annotation coverage is high, with no bare `Any` on public service signatures. Broad `except Exception` appears where it belongs — the per-page handler at `workflows.py:352` deliberately isolates one page's failure from the batch, which is correct. `# noqa: PLR0915` / `PLR1702` are used sparingly and consistently. Minor gaps: f-strings in logging (LOW-09), and two blocking-I/O deviations (LOW-01/LOW-02).
|
||||
|
||||
### FastAPI
|
||||
Lifespan is handled correctly via an `asynccontextmanager` `_lifespan` (`app.py:36-68`) rather than deprecated `@app.on_event`. Routers are domain-organized with typed path/query parameters and `response_model` declarations. Error handling is centralized through `register_error_handlers`, and the full internal→canonical category mapping is round-trip tested at the HTTP layer (`tests/api/test_error_responses.py:59-95`). `print_api.py:42-49` performs correct `relative_to`-based path containment for media serving. No blocking calls found in `async def` route handlers.
|
||||
|
||||
### NiceGUI
|
||||
Separation of concerns is good — pages delegate to services and `test_ui_boundaries.py` mechanically prevents persistence access from pages and components. Client state is client-scoped; no cross-session global-state leaks found. API usage is correct for the pinned 3.13.0 release. The two defects are the error-presenter bypass (HIGH-02) and synchronous file I/O in `homepage_store.py` (LOW-02).
|
||||
|
||||
### SQLModel & SQLAlchemy
|
||||
The strongest layer. `lazy="raise"` is declared on relationships and correctly paired with `expire_on_commit=False`, which together make N+1 access a loud failure rather than a silent performance cost — no N+1 patterns found. The job claim is a genuine atomic compare-and-swap (`jobs.py:212-222`: conditional `UPDATE ... WHERE status = QUEUED ... RETURNING`), which is the correct primitive and correctly implemented. Hot-path indexes are declared and test-verified (`test_db.py:131`). Cross-dialect portability is handled for SQLite and PostgreSQL. Weaknesses are transaction *ownership* (MED-05, HIGH-01) rather than query construction, plus the uncaught `IntegrityError` at MED-03.
|
||||
|
||||
### Pydantic V2 & Settings
|
||||
Fully migrated — no `@validator`, no `Config` class, no `.dict()` or `parse_obj` anywhere. `model_config = ConfigDict(...)` and `@field_validator` are used correctly. `config.py` is a clean single source of truth: **zero** `os.getenv` calls exist outside it, `.env` is untracked and gitignored, and the API key is `SecretStr` end-to-end. The only deviation is the hardcoded poll interval (LOW-03).
|
||||
|
||||
### Asyncio Workers
|
||||
Task lifecycle is handled properly: task references are retained (no GC risk), `CancelledError` is re-raised rather than swallowed, the provider call happens outside any DB transaction, timeouts resolve to terminal states, and there is no tight polling spin. `_persist_page_outcome_durably`'s use of `asyncio.shield` (`workflows.py:568-581`) is a thoughtful durability mechanism. The defects are the split commit boundary (HIGH-01), the shutdown-vs-provider timeout mismatch (MED-04), and startup-only recovery (MED-01).
|
||||
|
||||
### OpenRouter / Adapter Boundary
|
||||
Encapsulation is clean — `workflows.py` imports only abstract types from `providers`, never `openrouter` directly, so provider specifics do not leak into business logic. The `AsyncClient` is shared with configured timeouts and is properly closed: `worker.py:248,271` → `services.aclose()` → `sources.aclose()` (`sources.py:129-133`) → provider `aclose()` (`openrouter.py:86-87,233-235`). Responses are Pydantic-validated. **All 14 evidence-provenance-auditor invariant checks pass**, including the critical one: the API key is never persisted, request headers are never stored, and `TransportEvidence` captures response headers through an explicit allowlist (`evidence.py:130-134`). Only LOW-04 applies here.
|
||||
|
||||
### Testing & Quality Tooling
|
||||
377 tests pass with `-m "not external"`. The project test contract is honored: `--strict-markers` with all three markers (`unit`, `integration`, `external`) declared, `asyncio_mode = "strict"` with **every** `async def test_` correctly decorated across all 17 async test files, `external` properly excluded from default runs, and **no unawaited-coroutine warnings** — the `filterwarnings` error promotion is clean.
|
||||
|
||||
Contract coverage is genuinely strong for structural rules. Confirmed *mutation-sensitive* enforcement exists for: append-only evidence history (3 independent tests, including full before/after field-tuple snapshots), stuck-in-`PROCESSING` prevention, the complete 10-category error mapping, and both boundary rules.
|
||||
|
||||
The critical gap is transaction atomicity (HIGH-04) — the audit verdict is **"Effective with Conditions / Go with Conditions"**, blocking on the two missing atomicity tests. Secondary items are the low-signal assertions (LOW-10) and wall-clock flakiness (LOW-11).
|
||||
|
||||
**Prune/strengthen backlog:**
|
||||
|
||||
| Priority | Task | Location |
|
||||
| :--- | :--- | :--- |
|
||||
| High | Add Transaction B atomicity test (fault injected between transcript and status writes) | new `tests/integration/test_pipeline_atomicity.py` |
|
||||
| High | Add Transaction C atomicity test (retry: `error_detail` + `retry_count` + `QUEUED`) | same file |
|
||||
| Medium | Delete tautological assertions on same-file dict literals | `tests/test_traceability.py:54-57` |
|
||||
| Medium | Remove unfalsifiable `assert processed is True` | `tests/integration/test_pipeline_flow.py:135-140,446-452` |
|
||||
| Medium | Replace `>= 200` snapshot threshold with set-membership assertion | `tests/test_orphan_sweep.py:119` |
|
||||
| Medium | Assert mapped test files contain ≥1 test, not merely that they exist | `tests/test_traceability.py:59-60` |
|
||||
| Low | Widen timing slack or mock the clock | `tests/services/test_workflows_reliability.py:157-196` |
|
||||
| Low | Drop `assert result is not None` guards shadowed by following assertions | `tests/services/test_workflows_reliability.py:105,178,241,317,375` |
|
||||
|
||||
---
|
||||
|
||||
## 7. Duplication & Consolidation Report
|
||||
|
||||
| Pattern / Duplication | Locations | Proposed Canonical Home | Estimated Lines Removed |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| Hand-rolled `ui.notify(str(exc), type="negative")` | `home_page.py:212,220,228,255`; `people_page.py:265,321,330,339` | `ui/components/error_presenter.py::show_error` (already exists) | ~16 |
|
||||
| Optional-session `if session is None: async with _session_scope()` preamble | `workflows.py:557-561`, `workflows.py:591-595`, and sibling service write paths | `services/base.py::unit_of_work(services, session)` context manager | ~30 |
|
||||
| Synchronous I/O not wrapped in `asyncio.to_thread` | `store.py:401`, `homepage_store.py:25,32` | `services/base.py::run_blocking` helper | ~6 |
|
||||
| Read-then-increment `MAX(n) + 1` with uniqueness retry | `sources.py:540-546` (uncaught) vs `sources.py:531-534` (caught) | `services/base.py::insert_with_sequence_retry` | ~20 |
|
||||
|
||||
### Proposed Canonical Abstractions
|
||||
|
||||
```python
|
||||
# src/transcription/services/base.py
|
||||
|
||||
|
||||
@asynccontextmanager
|
||||
async def unit_of_work(
|
||||
services: ServiceBundle,
|
||||
session: AsyncSession | None = None,
|
||||
) -> AsyncIterator[AsyncSession]:
|
||||
"""Single transaction entry point. Yields a session and commits once on clean exit.
|
||||
|
||||
Replaces the `if session is None: async with X._session_scope()` preamble and the
|
||||
private-member access at workflows.py:558,592. Makes the two-commit split of
|
||||
HIGH-01 structurally hard to reintroduce.
|
||||
"""
|
||||
|
||||
|
||||
async def run_blocking[T](fn: Callable[[], T]) -> T:
|
||||
"""Run a CPU- or disk-bound callable off the event loop."""
|
||||
return await asyncio.to_thread(fn)
|
||||
|
||||
|
||||
async def insert_with_sequence_retry(
|
||||
session: AsyncSession,
|
||||
*,
|
||||
build: Callable[[int], SQLModel],
|
||||
next_value: Callable[[], Awaitable[int]],
|
||||
attempts: int = 3,
|
||||
) -> SQLModel:
|
||||
"""Insert a row carrying a derived sequence number, retrying on IntegrityError."""
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 8. Meta-Tooling & Instruction Update Recommendations
|
||||
|
||||
1. **Add `tests/integration/test_pipeline_atomicity.py`** (HIGH-04). The single highest-value enforcement change. Write it before fixing HIGH-01 so it demonstrably fails first.
|
||||
2. **Extend `tests/test_ui_boundaries.py`** with an AST check forbidding `ui.notify(..., type="negative")` in `PAGES_DIR`, routing all error display through `error_presenter`. Converts `ui.instructions.md:42` from steering into enforcement.
|
||||
3. **Extend `tests/ui/test_error_presenter.py`** with a case asserting that a path-bearing exception does not surface its path, enforcing `error-handling.instructions.md:74`.
|
||||
4. **Adopt a `ty` suppression policy** — targeted `# ty: ignore[...]` with rationale at the 10 known sites, documented in `docs/` — then **flip the pre-commit `ty` hook from advisory to blocking**. Until this happens the type checker provides no gate.
|
||||
5. **Add `ruff format --check` to the pre-commit gate**, preceded by one isolated formatting commit across the 35 drifted files.
|
||||
6. **Enable ruff rule set `G`** (`flake8-logging-format`) to catch f-string logging (LOW-09).
|
||||
7. **Extend `tests/test_orphan_sweep.py`** to public methods on service classes, seeding `KNOWN_ORPHANS` with current results (LOW-08). Then resolve all four existing "uncertain" entries to definite outcomes.
|
||||
8. **Extend `tests/test_service_boundaries.py`** with an AST check forbidding `_session_scope` attribute access outside its owning service module (MED-05). Also address the noted classification gap: the test excludes orchestration modules by hardcoded stem name (`store`, `workflows`, `__init__`), so a new orchestration module under a different name would be misclassified as a service.
|
||||
9. **Update `.github/instructions/services.instructions.md`** to state the `asyncio.to_thread` rule for blocking I/O and to name the single `unit_of_work` transaction entry point.
|
||||
10. **Update `docs/production-runbook.md`** with the stale-job recovery contract (interval, threshold, and its relationship to the container termination grace period), and note single-worker as a current precondition until MED-03 is fixed.
|
||||
11. **Note for `test_ui_boundaries.py`:** the forbidden-import lists are fixed string sets, so a future persistence helper under a new name would escape the check. Consider inverting to an allowlist of permitted imports for pages.
|
||||
|
||||
---
|
||||
|
||||
## 9. Prioritized Dependency-Ordered Action Plan
|
||||
|
||||
**Phase 1: Blocking fixes**
|
||||
1. Write the two atomicity tests (HIGH-04) and confirm they **fail** against current `main`.
|
||||
2. Fix the split commit boundary (HIGH-01) and confirm the tests now pass.
|
||||
3. Fix the filesystem-path leak in `classify_unexpected_error` (HIGH-03).
|
||||
4. Replace the 8 hand-rolled error notifications with `show_error` (HIGH-02).
|
||||
|
||||
**Phase 2: Enforcement hardening**
|
||||
5. Add the `ui.notify` AST guard and the path-leak presenter test, locking in items 3-4.
|
||||
6. Adopt the `ty` suppression policy and make the pre-commit hook blocking (LOW-05).
|
||||
7. Run `ruff format .` as an isolated commit, then add `ruff format --check` to the gate (LOW-06).
|
||||
8. Enable ruff rule set `G` and fix the resulting logging call sites (LOW-09).
|
||||
|
||||
**Phase 3: Reliability & concurrency**
|
||||
9. Move stale-job recovery to a periodic worker task with a dedicated setting (MED-01).
|
||||
10. Gate retries on `error_category` and add backoff (MED-02) — do this before ever raising `worker_max_retries` above 0.
|
||||
11. Derive the shutdown budget from the provider timeout (MED-04).
|
||||
12. Handle `IntegrityError` on the attempt-number flush (MED-03) — a hard precondition for running more than one worker replica.
|
||||
13. Move `sha256` and homepage-store I/O off the event loop (LOW-01, LOW-02); move the poll interval into `Settings` (LOW-03).
|
||||
|
||||
**Phase 4: Consolidation & refactoring**
|
||||
14. Introduce `unit_of_work` and migrate `workflows.py` off private `_session_scope` access (MED-05); add the corresponding boundary guard.
|
||||
15. Extract `run_blocking` and `insert_with_sequence_retry` (§7).
|
||||
16. Prune the low-signal assertions and reduce timing flakiness (LOW-10, LOW-11).
|
||||
|
||||
**Phase 5: Non-blocking governance/documentation depth**
|
||||
17. Extend the orphan sweep to methods and resolve the four uncertain orphans (LOW-07, LOW-08).
|
||||
18. Update `services.instructions.md` and `docs/production-runbook.md` per §8 items 9-10.
|
||||
19. Log incomplete request manifests (LOW-04).
|
||||
20. Consider inverting the UI boundary check to an allowlist.
|
||||
|
||||
---
|
||||
|
||||
## 10. Preserved Strengths
|
||||
|
||||
- **Evidence and provenance integrity is exemplary.** All 14 provenance-auditor invariants pass. `ExecutionAttempt` history is genuinely append-only, retries append rather than rewrite, and projection writes are cleanly distinguished from history mutation. Three independent tests — including full before/after field-tuple snapshots — make any mutation regression fail loudly.
|
||||
- **Secret hygiene is correct by construction.** The response-header **allowlist** (`evidence.py:130-134`) fails closed: a newly-introduced sensitive header is excluded by default rather than requiring someone to remember to block it. Request headers are never captured, and `SecretStr` is used end-to-end.
|
||||
- **Atomic job claiming.** `jobs.py:212-222` uses a conditional `UPDATE ... WHERE status = QUEUED ... RETURNING` — a true compare-and-swap that makes double-claiming impossible under concurrency, rather than the common read-then-write race.
|
||||
- **`lazy="raise"` paired with `expire_on_commit=False`.** This combination turns accidental lazy loads into immediate errors instead of silent N+1 queries, and it is the reason no N+1 patterns exist in the codebase. Keep it.
|
||||
- **Architectural rules are mechanically enforced, not merely documented.** AST-based boundary tests for service-to-service imports and UI persistence access are the right pattern; this review's main recommendation is simply to apply that same pattern to three more rules.
|
||||
- **Configuration discipline.** Zero `os.getenv` calls outside `config.py`, `.env` untracked and gitignored, clean Pydantic V2 throughout with no V1 residue.
|
||||
- **Path containment on media serving.** `print_api.py:42-49` uses proper `relative_to` validation rather than string prefix matching.
|
||||
- **Async worker fundamentals.** Task references retained, `CancelledError` re-raised, provider calls outside DB transactions, timeouts resolving to terminal states, no tight polling loop. `asyncio.shield` in `_persist_page_outcome_durably` is a genuinely thoughtful durability mechanism — the fix in HIGH-01 should preserve it for intermediate pages.
|
||||
- **Test contract rigor.** `--strict-markers`, `asyncio_mode = "strict"` honored across all 17 async test files with no missing decorators, and coroutine-never-awaited promoted to a hard error with a clean run.
|
||||
@@ -0,0 +1,896 @@
|
||||
# Architecture & Code Review Report
|
||||
|
||||
**Repository Target:** `C:\Github\transcription\`
|
||||
**Target Stack:** Python 3.12+ | FastAPI | NiceGUI | SQLModel/SQLAlchemy | Pydantic V2 | asyncio | OpenRouter
|
||||
**Review Date:** 2026-09-02
|
||||
**Canonical Baseline:** V6.1 (`docs/index.md`)
|
||||
|
||||
---
|
||||
|
||||
## 0. Verification Commands and Outcomes
|
||||
|
||||
All four commands were executed in this checkout before any finding was written. This report
|
||||
records the exact outcomes rather than assuming them.
|
||||
|
||||
| Command | Outcome |
|
||||
| :--- | :--- |
|
||||
| `uv run pytest -q -m "not external"` | **410 passed, 0 failed, 0 errors** (exit 0) |
|
||||
| `uv run ruff check .` | **All checks passed!** |
|
||||
| `uv run ruff format --check .` | **191 files already formatted** |
|
||||
| `uv run ty check` | **All checks passed!** |
|
||||
|
||||
The stated green baseline is real. No finding below is a test failure; every finding is a
|
||||
behavior, contract, or guard-coverage defect that the passing suite does not detect.
|
||||
|
||||
---
|
||||
|
||||
## 1. Executive Summary
|
||||
|
||||
- **The system's core evidence guarantees hold.** `ExecutionAttempt` is genuinely append-only,
|
||||
attempt numbering is allocated with bounded conflict retry, transport evidence is captured at
|
||||
the HTTP boundary before SDK parsing, and header persistence uses a true allowlist. Provenance
|
||||
invariant families A–E and G pass.
|
||||
- **Both competing atomicity invariants in `services/workflows.py` are real and both guards
|
||||
genuinely enforce them.** I injected-fault-verified the tests rather than trusting the
|
||||
docstrings: `test_pipeline_atomicity.py` fails on a split final-page commit, and
|
||||
`test_workflows_reliability.py:318` reads intermediate attempts through a *separate session*,
|
||||
so it would fail if intermediate pages stopped committing individually.
|
||||
- **The most significant defect is a privacy leak that a prior review believed it had closed.**
|
||||
The 2026-08-23 review moved root-cause text out of `AppError.message` into `AppError.detail`
|
||||
to keep filesystem paths away from users. That text now reaches users anyway, because the UI
|
||||
renders `ExecutionAttempt.error_detail` verbatim (HIGH-01). The leak was relocated, not closed.
|
||||
- **A second, independent path leak exists in five explicit `raise` sites** that the existing
|
||||
guard never covered — it tests only `classify_unexpected_error` (HIGH-02).
|
||||
- **Provenance invariant family F (path safety) fails**, and it fails *inconsistently within one
|
||||
file*: `sources_page.py:443` carefully sanitizes a stored path through
|
||||
`public_media_path_label`, then `sources_page.py:484` dumps raw `error_detail` forty lines later.
|
||||
- **The orphan sweep does not do what its docstring claims.** It matches definitions by bare name,
|
||||
so an entirely dead *module* passes whenever its function names collide with live ones.
|
||||
`ui/pages/tags_page.py` is the proof: 93 lines never imported by anything (MED-01/LOW-01).
|
||||
- **On the three flagged open items:** the V4/V6.1 doc drift is confirmed (MED-02); the
|
||||
`.env.production` coupling is real but currently correct and loud-failing, so Medium not High
|
||||
(MED-03); and the Tags roadmap is **right** — the route is genuinely not registered, so the
|
||||
module is dead code rather than a live retired route.
|
||||
- **Two latent concurrency defects carry ordering constraints** and must be fixed *before* the
|
||||
changes that would make them live (MED-04, MED-05), not after.
|
||||
- **Guidance-file accuracy:** the recently revised `.github/instructions/*` files were verified
|
||||
against code rather than trusted. They are accurate as written; the code is what diverges from
|
||||
them. The one exception is that `error-handling.instructions.md` states a `detail` rule the UI
|
||||
layer has never followed, which makes it an unenforced claim rather than a wrong one.
|
||||
|
||||
---
|
||||
|
||||
## 2. Executive Architecture Assessment
|
||||
|
||||
**Verdict: architecturally sound, with a concentrated failure in the *last mile* of error
|
||||
presentation.**
|
||||
|
||||
Domain cohesion and dependency direction are good and, unusually, mechanically enforced.
|
||||
`test_service_boundaries.py` and `test_ui_boundaries.py` AST-scan for violations using
|
||||
*allowlists* rather than blocklists, which is the correct choice — a newly added persistence
|
||||
helper cannot slip through under an unlisted name. `workflows.py` imports only the abstract
|
||||
`providers` types and never `openrouter`, so provider details genuinely stop at the adapter.
|
||||
Transaction ownership is explicit and well-reasoned: `ServiceBase._finalize` commits for
|
||||
service-owned sessions and flushes for caller-owned ones, which is what lets orchestration
|
||||
modules compose multi-aggregate writes without services importing each other.
|
||||
|
||||
The evidence layer is the strongest part of the system and shows real care. The distinction
|
||||
between transport response, SDK-parsed response, and normalized metadata is maintained in code,
|
||||
not just in prose — `_CapturingAsyncClient` exists specifically to retain the exact wire body
|
||||
before the SDK can discard unknown fields, and `TransportEvidence(response_received=False)`
|
||||
explicitly represents "no response was received" rather than conflating it with an empty one.
|
||||
|
||||
The weakness is at the boundary where internal diagnostic text becomes pixels. Every layer
|
||||
*below* the UI respects the message/detail split; the UI layer reads the internal field directly
|
||||
and renders it. The architecture defines the contract correctly and then has no enforcement at
|
||||
the one layer that violates it.
|
||||
|
||||
**Top systemic risks:**
|
||||
|
||||
1. **Internal diagnostic text reaches users through the evidence display path** (HIGH-01). The
|
||||
rule is documented in three places and enforced in none of them at the UI boundary.
|
||||
2. **Path-safety discipline is applied per-call-site rather than structurally** (HIGH-02, HIGH-01).
|
||||
It is correct wherever someone remembered; there is no guard that makes forgetting fail.
|
||||
3. **Guard coverage is narrower than guard docstrings claim.** Two guards
|
||||
(`test_orphan_sweep.py`, `test_errors.py`) assert something meaningfully weaker than the
|
||||
invariant they are named for, which converts them into a false sense of enforcement.
|
||||
4. **Worker safety currently rests on single-process sequential execution, not on configuration**
|
||||
(MED-04, MED-05). Nothing is wrong today; two plausible future changes each make something wrong.
|
||||
|
||||
---
|
||||
|
||||
## 3. Findings by Severity
|
||||
|
||||
### Critical Severity
|
||||
|
||||
*None.* No evidence loss, append-only violation, secret leakage, or silent-wrong-output defect
|
||||
was found. The candidates in this class (provider evidence mis-attribution, stale-job double
|
||||
processing) are latent and are reported at High/Medium with their unblocking conditions.
|
||||
|
||||
---
|
||||
|
||||
### High Severity
|
||||
|
||||
#### [HIGH-01] Internal-only `error_detail` is rendered directly to users, reopening the leak the 2026-08-23 fix was meant to close
|
||||
|
||||
- **Location:**
|
||||
- Write side: `src/transcription/errors.py:99-139` (`classify_unexpected_error` → `detail`, `format_error_detail` → persisted text)
|
||||
- Persist: `src/transcription/services/workflows.py:720` (`error_detail=format_error_detail(page.error)`)
|
||||
- **Render (Source Detail):** `src/transcription/ui/pages/sources_page.py:481-484`
|
||||
- **Render (Sources list):** `src/transcription/ui/pages/sources_page.py:121` → `src/transcription/ui/components/table/sources.py:39,90-95` ("Error Detail" column)
|
||||
- **Render (Maintenance):** `src/transcription/ui/pages/settings_page.py:562`, written by `src/transcription/services/maintenance.py:206`
|
||||
- Contract violated: `docs/error_handling.md:107-114`; `.github/instructions/error-handling.instructions.md:86`; `docs/invariant/error_handling.md:59`; `docs/invariant/ai_evidence_and_provenance.md:103`
|
||||
|
||||
- **Reachability:** **Live.** Concrete path, no configuration required: a page fails with any
|
||||
non-`AppError` exception → `workflows.py:388` calls `classify_unexpected_error(exc)` →
|
||||
`errors.py:118` sets `detail=f"{type(exc).__name__}: {exc}"` → `format_error_detail`
|
||||
(`errors.py:135-139`) emits `... | detail=OSError: [Errno 13] Permission denied: '/app/uploads/documents/<uuid>/page-1.jpg' | ...`
|
||||
→ persisted to `ExecutionAttempt.error_detail` → rendered verbatim at
|
||||
`sources_page.py:484` and in the `/sources` table column. A SQLAlchemy `OperationalError`
|
||||
carries the database path by the same route.
|
||||
|
||||
- **Problem & Consequence:** `docs/error_handling.md:110` states `detail` is *"Internal only"* and
|
||||
that its only surfaces are `format_error_detail` (evidence) and logs;
|
||||
`error-handling.instructions.md:86` says *"Never rendered to users or serialized into an
|
||||
envelope."* The UI reads it anyway. The consequence is not hypothetical drift — it is the
|
||||
precise defect the previous review's fix existed to prevent. That fix made `message` generic and
|
||||
moved the root cause to `detail` on the stated grounds that `detail` never reaches users. That
|
||||
premise was never true: `error_detail` had a UI consumer the whole time. The result is that the
|
||||
filesystem-path leak was relocated from the notification banner to the Source Detail card and
|
||||
the Sources table, while the test suite records the leak as fixed
|
||||
(`tests/test_errors.py:56-78`).
|
||||
|
||||
The inconsistency is visible inside a single file: `sources_page.py:443` deliberately routes a
|
||||
stored path through `public_media_path_label` (`ui/components/media_urls.py:58-72`), which
|
||||
correctly degrades an absolute path to its bare filename — and then `sources_page.py:484`
|
||||
renders unsanitized text that may contain an absolute path.
|
||||
|
||||
- **Blast Radius:** Enumerated by grepping every reader of `.detail` and `error_detail`:
|
||||
- `errors.py:137` — `format_error_detail`, the only reader of `AppError.detail`. **Must keep the root cause.**
|
||||
- `services/workflows.py:720` — the only writer of `ExecutionAttempt.error_detail`.
|
||||
- `services/maintenance.py:206` — the only writer of `MaintenanceRun.error_detail`.
|
||||
- `services/evidence.py:195` — `build_evidence_export` emits `error_detail`. Export is an
|
||||
operator-initiated evidence artifact; per invariant 3.7.1 it **must** retain it.
|
||||
- `db/models.py:508-522` — `Source.latest_error_detail` projection, consumed only by `sources_page.py:121`.
|
||||
- `ui/pages/sources_page.py:481-484`, `ui/components/table/sources.py`, `ui/pages/settings_page.py:562` — the three render sites.
|
||||
- Tests asserting on persisted text: `tests/test_v42_evidence.py:284`,
|
||||
`tests/services/test_workflows_reliability.py` (timeout detail),
|
||||
`tests/services/test_maintenance_service.py`. A fix that changes *what is stored* breaks these;
|
||||
a fix that changes *what is displayed* does not.
|
||||
|
||||
- **Recommendation — two invariants conflict here; both must be named.**
|
||||
|
||||
**Invariant 1 (evidence):** `ExecutionAttempt.error_detail` must retain the root cause.
|
||||
`docs/requirements.md:30` (REQ-4-021) and `docs/invariant/ai_evidence_and_provenance.md:33`
|
||||
require it; guarded by `tests/test_v42_evidence.py::test_attempts_are_append_only_and_exported_with_integrity`
|
||||
and `tests/services/test_workflows_reliability.py`.
|
||||
|
||||
**Invariant 2 (privacy):** user-facing surfaces must not expose local filesystem details.
|
||||
`docs/invariant/error_handling.md:59`; guarded (partially) by
|
||||
`tests/test_errors.py::test_unexpected_error_does_not_leak_filesystem_paths`.
|
||||
|
||||
**The over-correction to avoid is stripping root-cause text out of `detail` or
|
||||
`format_error_detail` to make the UI safe.** That is exactly the mistake documented in the
|
||||
reviewer skill's worked example, and it would silently destroy the provenance record this
|
||||
system exists to preserve while making every guard still pass.
|
||||
|
||||
Fix at the **render** boundary, not the write boundary. Add a presentation-layer projection and
|
||||
route all three UI sites through it, leaving the persisted evidence untouched:
|
||||
|
||||
```python
|
||||
# src/transcription/ui/components/error_presenter.py (new)
|
||||
def display_failure_detail(error_detail: str | None) -> str | None:
|
||||
"""Render persisted failure detail without machine-local paths.
|
||||
|
||||
`ExecutionAttempt.error_detail` is provenance and keeps the full root cause
|
||||
(docs/error_handling.md). This projection is the only thing a page may show.
|
||||
"""
|
||||
```
|
||||
|
||||
It should preserve the `[category]`, `suggestion=`, and `error_id=` segments (which are what
|
||||
make the display actionable) and reduce any absolute path inside `detail=` to its basename,
|
||||
mirroring `public_media_path_label`. The operator keeps diagnosability — required by
|
||||
`docs/ui/pages/sources.md:43` and `docs/requirements.md:59` (REQ-6-014) — without the container
|
||||
filesystem layout being published to the browser.
|
||||
|
||||
Then decide and record which resolution was chosen: either the UI shows the sanitized
|
||||
projection (recommended), or `docs/error_handling.md:107-114` and
|
||||
`error-handling.instructions.md:86` are revised to state that operator-facing evidence displays
|
||||
may render `error_detail` **and** that the guarantee moves to "no machine-local detail ever
|
||||
enters `detail`" — which would be a much harder guarantee to keep. Do not leave the current
|
||||
state, where the docs claim one thing and three pages do another.
|
||||
|
||||
- **Effort:** M
|
||||
|
||||
---
|
||||
|
||||
#### [HIGH-02] Absolute filesystem paths are embedded in user-facing `AppError.message` at five explicit raise sites
|
||||
|
||||
- **Location:**
|
||||
- `src/transcription/services/sources.py:856` — `f"Prompt file not found: {prompt_path}"`
|
||||
- `src/transcription/services/sources.py:864` — `f"Prompt file is empty: {prompt_path}"`
|
||||
- `src/transcription/services/sources.py:914` — `f"Source file not found: {path}"`
|
||||
- `src/transcription/services/prompts.py:99` — `f"Prompt directory is unavailable: {root}"`
|
||||
- `src/transcription/services/prompts.py:186-191` — `_filesystem_error` builds `f"{message}: {exc}"`
|
||||
- Contract violated: `.github/instructions/error-handling.instructions.md:74,85`; `docs/invariant/error_handling.md:59`
|
||||
|
||||
- **Reachability:** **Live**, on an ordinary user path. `sources.py:845` resolves
|
||||
`prompt_root = runtime_settings.prompt_dir.resolve()`, so `prompt_path` is absolute
|
||||
(`/app/prompts/transcribe_document.md` in the container). `load_prompt_text` is invoked by
|
||||
`build_prompt_execution` (`sources.py:829-831`), which runs on **every document upload** via
|
||||
`services/store.py:94` and `store.py:162`. The resulting `PromptLoadError` is an `AppError`
|
||||
subclass, so it flows through `run_ui_action` → `show_error`
|
||||
(`ui/components/error_presenter.py:51-66`), which renders `error.message` into both a
|
||||
`ui.notify` banner and a card label, and through `build_error_envelope` (`errors.py:88-96`)
|
||||
into API responses.
|
||||
|
||||
- **Problem & Consequence:** `error-handling.instructions.md:85` requires `message` to *"Stay
|
||||
generic. Never embed exception text, provider payloads, or filesystem paths."* These five sites
|
||||
embed exactly that. `prompts.py:186-191` violates the rule in **both** directions at once: it
|
||||
puts `{exc}` — an `OSError` whose `str()` includes the offending filename — into `message`, and
|
||||
it sets **no `detail=`**, so the internal field that is supposed to carry the root cause is
|
||||
empty while the user-facing field carries all of it.
|
||||
|
||||
This is not a new regression; it is coverage that the existing guard never had.
|
||||
`tests/test_errors.py:56-78` verifies only that `classify_unexpected_error` — the *catch-all*
|
||||
path — does not leak. Every deliberate `raise SomeError(f"... {path}")` in the codebase is
|
||||
outside its scope, so the suite reports the invariant as enforced while five live sites violate it.
|
||||
|
||||
- **Blast Radius:** Verified by grepping all consumers of these exception types.
|
||||
`PromptLoadError`/`PromptStoreError`/`TranscriptionError` messages are consumed by:
|
||||
`ui/components/error_presenter.py:55,63` (render), `errors.py:92` (API envelope),
|
||||
`errors.py:135` (`format_error_detail` → evidence). Because the recommended change *adds* a
|
||||
`detail` and *shortens* `message`, `format_error_detail` output still contains the path — so
|
||||
evidence value is preserved, not reduced. Tests asserting on these messages:
|
||||
`tests/test_prompts.py`, `tests/services/test_prompt_store.py`,
|
||||
`tests/services/test_transcription_service.py`. These assert on message prefixes
|
||||
(`"Prompt file not found"`), not on the interpolated path, and were checked to survive the change —
|
||||
but re-run them, since `prompts.py:186` currently produces a message whose suffix some
|
||||
assertion could depend on.
|
||||
|
||||
- **Recommendation:** Apply the pattern `errors.py:113-119` already establishes — generic
|
||||
`message`, root cause on `detail`, `raise ... from exc`. Use `path.name` when a filename is
|
||||
genuinely useful to the user.
|
||||
|
||||
```python
|
||||
# sources.py:855 — before
|
||||
raise PromptLoadError(f"Prompt file not found: {prompt_path}", ...)
|
||||
# after
|
||||
raise PromptLoadError(
|
||||
f"Prompt file not found: {prompt_path.name}",
|
||||
category=ErrorCategory.INFRA_PERSISTENT,
|
||||
suggestion="Verify PROMPT_DIR and prompt file configuration, then retry.",
|
||||
detail=f"Prompt file missing at {prompt_path}",
|
||||
)
|
||||
|
||||
# prompts.py:186 — before
|
||||
return PromptStoreError(f"{message}: {exc}", category=..., suggestion=...)
|
||||
# after
|
||||
return PromptStoreError(
|
||||
message,
|
||||
category=ErrorCategory.INFRA_PERSISTENT,
|
||||
suggestion="Check prompt directory permissions and available disk space, then retry.",
|
||||
detail=f"{type(exc).__name__}: {exc}",
|
||||
)
|
||||
```
|
||||
|
||||
Then widen the guard so this class cannot recur — see MED-07. Note the dependency: HIGH-02 and
|
||||
HIGH-01 must be fixed **together**, because moving the path from `message` to `detail` while the
|
||||
UI still renders `error_detail` relocates the leak instead of closing it. That is the same
|
||||
mistake that produced HIGH-01.
|
||||
|
||||
- **Effort:** S (fix) / M (with the guard)
|
||||
|
||||
---
|
||||
|
||||
### Medium Severity
|
||||
|
||||
#### [MED-01] The orphan sweep matches by bare name and therefore cannot detect a dead module
|
||||
|
||||
- **Location:** `tests/test_orphan_sweep.py:101-163` (`_public_definitions`, `_orphans`)
|
||||
- **Reachability:** **Live** — the guard is running now and reporting a clean sweep that is not clean.
|
||||
- **Problem & Consequence:** `_public_definitions()` keys definitions by bare name
|
||||
(`definitions[node.name]`, line 111) and `_orphans()` marks a definition referenced if that
|
||||
bare name appears **anywhere** in `src/`, `tests/`, or `tools/` (lines 157-162). Two different
|
||||
modules that define the same public name are therefore indistinguishable, and neither can ever
|
||||
be reported as an orphan.
|
||||
|
||||
`src/transcription/ui/pages/tags_page.py` demonstrates the consequence. Its only public
|
||||
definition is `register_page` (line 18). Seven live page modules define a function of the same
|
||||
name and `ui/__init__.py:37-43` calls all seven — so `register_page` is heavily referenced and
|
||||
`tags_page.register_page` is scored as reachable. In fact **nothing imports `tags_page` at all**
|
||||
(verified: the only repo-wide references to the module are the file itself and
|
||||
`tests/ui/test_tags_page.py`, which merely asserts the route 404s). 93 lines of code, including a
|
||||
lazy-load-unsafe relationship traversal at `tags_page.py:71-74`, sit outside the sweep's reach.
|
||||
|
||||
The sweep also never asks whether a *module* is imported, only whether its definitions' names
|
||||
appear somewhere — so this is a structural gap, not a one-off miss.
|
||||
|
||||
- **Blast Radius:** `tests/test_orphan_sweep.py` only; `KNOWN_ORPHANS` entries are keyed by the
|
||||
same bare/dotted names and would need re-keying if qualification is added. Expect the stricter
|
||||
sweep to surface additional true orphans on first run — triage them into `KNOWN_ORPHANS` with
|
||||
rationales rather than weakening the check.
|
||||
- **Recommendation:** Qualify definitions by module (`f"{module_path}:{name}"`) and add a separate,
|
||||
cheap module-reachability pass: a module under `src/transcription/` is reachable if any other
|
||||
module imports it, or it is a declared entrypoint (`app.py`, `__main__.py`, `worker_service.py`).
|
||||
Report unreachable modules as orphans in their own right. Also fix
|
||||
`test_public_definitions_are_discovered` (line 169), whose `>= 420` snapshot threshold is a
|
||||
weak assertion that drifts upward silently — the 2026-08-23 review already flagged the same
|
||||
pattern at the then-current `>= 200` and it was raised rather than replaced.
|
||||
- **Effort:** M
|
||||
|
||||
---
|
||||
|
||||
#### [MED-02] Canonical invariant document declares a V4 baseline while the canonical baseline is V6.1
|
||||
|
||||
- **Location:** `docs/invariant/ai_evidence_and_provenance.md:130`
|
||||
- **Reachability:** **Live** (documentation), no runtime impact.
|
||||
- **Problem & Consequence:** Section 6.1 reads *"Canonical V4 architecture, schema, requirements,
|
||||
and error-policy documents define how current behavior satisfies this invariant."*
|
||||
`docs/index.md:1,29-32` establishes V6.1 as the baseline and states that every canonical document
|
||||
asserts the same baseline. This is the **ownership clause of the invariant that governs the
|
||||
entire evidence model** — the clause that tells a reader which documents are authoritative — and
|
||||
it points at a superseded generation. A reader following it lands on stale authority precisely
|
||||
when resolving an evidence question, which is the highest-stakes case.
|
||||
|
||||
A baseline-currency guard **does** exist —
|
||||
`tests/test_meta_contract_guards.py::test_canonical_docs_declare_one_consistent_baseline`
|
||||
(lines 89-112) — and `docs/invariant/ai_evidence_and_provenance.md` is **not** in
|
||||
`BASELINE_SCAN_EXCLUSIONS` (lines 56-64), so the file is scanned. The claim escapes for two
|
||||
independent reasons, either of which alone would be sufficient:
|
||||
1. `_CURRENT_VERSION_CLAIM` (line 67) matches only the words `current` or `active` before a
|
||||
version. This line says "**Canonical** V4", a third phrasing the pattern does not know.
|
||||
2. Both patterns require `V(\d+\.\d+)` — a mandatory minor version. The bare token `V4` cannot
|
||||
match either regex under any phrasing.
|
||||
|
||||
The guard is therefore not absent but *phrase-shaped*: it enforces currency only for the two
|
||||
sentence forms someone thought of, against version strings that carry a minor. That is a weaker
|
||||
property than its docstring implies ("Every canonical doc that names the current baseline must
|
||||
name the same one").
|
||||
- **Blast Radius:** Documentation only; no code reads this string. Widening the guard's patterns
|
||||
will re-scan all canonical docs — expect it to surface further stale mentions on first run
|
||||
(`docs/architecture.md`, `docs/schema.md`, `docs/requirements.md`, and `docs/error_handling.md`
|
||||
each contain 2-3 version tokens), which should be triaged rather than excluded.
|
||||
- **Recommendation:** Two parts, and the second matters more than the first.
|
||||
1. Change "Canonical V4" to "Canonical V6.1" at line 130.
|
||||
2. Fix the guard's shape rather than adding a third phrase to the list. Accept an optional minor
|
||||
(`V(\d+)(?:\.(\d+))?`) and invert the matching: flag **every** `V<n>` token in a scanned
|
||||
canonical doc that is not the declared baseline, rather than only those preceded by an
|
||||
approved adjective. Phrase-list matching fails open — each new phrasing silently reopens the
|
||||
hole — whereas token matching fails closed and forces an explicit exclusion.
|
||||
|
||||
See §8.1 for the alternative the maintainer is considering: dropping version labels from
|
||||
canonical docs entirely, which removes the failure mode instead of guarding it.
|
||||
- **Effort:** S
|
||||
|
||||
---
|
||||
|
||||
#### [MED-03] Settings resolve `.env.production` relative to the process working directory, and the isolation fix exists only in the test harness
|
||||
|
||||
- **Location:** `src/transcription/config.py:66-75` (`env_file=".env.production"`);
|
||||
workaround at `tests/conftest.py:27-50`; guarded by `tests/test_config_isolation.py`;
|
||||
depended on by `.github/workflows/quality-gate.yml` and `docker-compose.production.yml`
|
||||
- **Reachability:** **Live but currently correct.** I verified the production path rather than
|
||||
assuming it: `Dockerfile` sets `WORKDIR /app` in the runtime stage, and
|
||||
`docker-compose.production.yml` mounts `./.env.production` to `/app/.env.production` for both the
|
||||
`app` and `worker` services, so the relative path resolves correctly today.
|
||||
- **Problem & Consequence:** Correct configuration loading depends on an **implicit, undocumented
|
||||
contract between `config.py` and the process working directory.** Nothing in `config.py` states
|
||||
it, and nothing tests it. The failure mode is not silent — `openrouter_api_key` is required with
|
||||
no default, so a wrong cwd produces a `ValidationError` at startup rather than a partially
|
||||
configured process — which is why this is Medium rather than High.
|
||||
|
||||
The more telling symptom is what the coupling forced on the test harness. `conftest.py:45-50`
|
||||
cannot escape it by passing an argument; it must **mutate the Pydantic class-level
|
||||
`model_config` dict at runtime** and restore it in a `finally`. That is a global, order-sensitive
|
||||
side effect adopted because the module offers no seam. It also silently repairs a second
|
||||
consumer: `ui/runtime_settings_store.py:402` reads the same `Settings.model_config["env_file"]`
|
||||
to decide where the Settings page writes. Two subsystems are coupled through a mutable class
|
||||
attribute.
|
||||
- **Blast Radius:** Every `Settings` construction. Consumers of `model_config["env_file"]`:
|
||||
`ui/runtime_settings_store.py:402` (write-target resolution, contract documented at
|
||||
`docs/ui/pages/settings.md:27`) and `tests/conftest.py:45-50`. A change must preserve the
|
||||
documented three-step resolution order — explicit override, `RUNTIME_SETTINGS_ENV_FILE`, then the
|
||||
configured default — or `docs/ui/pages/settings.md:27` becomes wrong.
|
||||
- **Recommendation:** Introduce one explicit resolution function that both `Settings` construction
|
||||
and `runtime_settings_store` call, honoring an `ENV_FILE` environment variable and falling back
|
||||
to a path anchored to a known root rather than to `os.getcwd()`. Tests then pass a path instead of
|
||||
mutating class state, and `tests/test_config_isolation.py` can assert against the seam rather
|
||||
than against the monkeypatch. If instead the cwd contract is accepted as deliberate, document it
|
||||
in `config.py` and in `docs/production-runbook.md` and add a guard asserting `WORKDIR`/cwd
|
||||
alignment — an implicit contract with a container image is exactly the kind of rule the invariant
|
||||
routing table exists to place.
|
||||
- **Effort:** M
|
||||
|
||||
---
|
||||
|
||||
#### [MED-04] Stale-job reclaim threshold is not derived from maximum job duration; safety currently comes from single-process sequencing
|
||||
|
||||
- **Location:** `src/transcription/config.py:116-117`
|
||||
(`worker_provider_timeout_seconds=30.0`, `worker_stale_job_seconds=30.0`);
|
||||
sweep at `src/transcription/worker.py:222-228`; reclaim at
|
||||
`src/transcription/services/jobs.py:242-268`
|
||||
- **Reachability:** **Latent.** Unblocked by *either* of: (a) running more than one worker replica
|
||||
(adding `deploy.replicas > 1` to the `worker` service in `docker-compose.production.yml`), or
|
||||
(b) setting `RUN_EMBEDDED_WORKER=true` on the `app` service while the standalone `worker`
|
||||
container is also running. It is safe today only because
|
||||
`docker-compose.production.yml` sets `RUN_EMBEDDED_WORKER: "false"` on `app` and defines exactly
|
||||
one `worker`, and because within a single loop `run_worker_loop` awaits
|
||||
`process_next_queued_job` to completion before returning to the stale sweep — so the sweep can
|
||||
never observe a job that this same process is actively working.
|
||||
- **Problem & Consequence:** The stale threshold (30s) **equals** the per-page provider timeout
|
||||
(30s), leaving zero margin even for a single-page job. A multi-page document is legitimately
|
||||
`PROCESSING` for up to N × 30s. `Job.date_updated` carries an `onupdate`
|
||||
(`db/models.py:372-375`), but between the initial claim and the terminal write the only touch
|
||||
is `sources.py:522-523` reassigning `job.provider`/`job.model` to values they usually already
|
||||
hold, which SQLAlchemy resolves to no net change and therefore no `UPDATE`. I did not empirically
|
||||
confirm the no-`UPDATE` behavior, so treat that specific step as unverified — but the finding does
|
||||
not depend on it, because even a per-page refresh leaves only a 30s margin against a 30s timeout.
|
||||
|
||||
With a second concurrent worker, the sweep would requeue a job that is mid-provider-call. Both
|
||||
workers then process the same job, producing duplicate `ExecutionAttempt` rows for the same
|
||||
logical work and racing terminal status writes. Append-only history would be *preserved* but no
|
||||
longer *faithful*: the evidence would show attempts that do not correspond to distinct
|
||||
application decisions.
|
||||
|
||||
This is worth flagging because `jobs.py:191-197` explicitly implements and documents
|
||||
`SKIP LOCKED` row locking "so concurrent workers never contend for the same job." The claim path
|
||||
is built for multi-worker operation; the reclaim path is not. A reader who trusts the claim
|
||||
docstring would reasonably scale the worker.
|
||||
- **Blast Radius:** `requeue_stale_processing_jobs` has one production caller (`worker.py:226`) and
|
||||
tests in `tests/test_worker.py` and `tests/services/test_job_service.py`. Changing the *default*
|
||||
affects `tests/test_config.py` declared-defaults assertions — check those before editing the default.
|
||||
- **Recommendation:** **Fix before adding a second worker replica, not after.** Two parts:
|
||||
(1) Make the threshold a function of the real bound rather than a coincidental peer of the
|
||||
page timeout — at minimum default `worker_stale_job_seconds` to a multiple of
|
||||
`worker_provider_timeout_seconds` with headroom, and add a model validator rejecting a stale
|
||||
threshold at or below the provider timeout.
|
||||
(2) Preferably make reclaim heartbeat-based: have `_persist_page_outcome` bump `Job.date_updated`
|
||||
explicitly so liveness reflects progress rather than elapsed time since claim.
|
||||
Add a guard asserting a multi-page job in flight is not reclaimed by a concurrently-invoked sweep.
|
||||
- **Effort:** M
|
||||
|
||||
---
|
||||
|
||||
#### [MED-05] Provider evidence capture is per-instance mutable state, making the adapter non-reentrant by contract
|
||||
|
||||
- **Location:** `src/transcription/providers/openrouter.py:197-199, 264-267, 274-275, 297-298, 397-412`;
|
||||
`_CapturingAsyncClient.last_response`/`last_body` at `openrouter.py:66-94`;
|
||||
contract at `src/transcription/providers/base.py:110-118`
|
||||
(`current_request_manifest`, `current_transport_evidence`)
|
||||
- **Reachability:** **Latent.** Unblocked by any concurrent `transcribe()` on a single adapter
|
||||
instance — most plausibly by processing a job's pages in parallel (`workflows.py:274` is
|
||||
currently a sequential `for` loop) or by any second consumer sharing one
|
||||
`SourceService.provider`. Verified safe today: `workflows.py:272` resolves one provider for the
|
||||
loop and awaits each page; the worker's `ServiceBundle` (`worker.py:206`) is distinct from
|
||||
`app.state.services` (`app.py:43`), so the UI cannot share the worker's adapter instance, and
|
||||
the UI only enqueues jobs (`ui/pages/jobs_page.py:186-208`).
|
||||
- **Problem & Consequence:** The `TranscriptionProvider` protocol defines evidence retrieval as
|
||||
"the most recent call" state read *after* the fact. `workflows.py:369-370` relies on this on the
|
||||
timeout path, reading `provider.current_request_manifest` / `current_transport_evidence` when no
|
||||
result object exists. Under concurrency, page B's response overwrites
|
||||
`_CapturingAsyncClient.last_response` before page A's timeout handler reads it, and page A's
|
||||
`ExecutionAttempt` is written with page B's transport evidence.
|
||||
|
||||
The consequence is **evidence mis-attribution** — a provenance-integrity failure, which this
|
||||
project's own rubric treats as its most serious class. It would also be near-undetectable after
|
||||
the fact: the attempt row would be well-formed, internally consistent, and wrong. The
|
||||
application-level design that makes this safe (sequential pages) is not expressed in the
|
||||
provider contract, so the constraint lives only in `workflows.py`'s loop structure.
|
||||
- **Blast Radius:** Changing the protocol touches `providers/base.py:102-136`,
|
||||
`providers/openrouter.py:221-231`, the two read sites at `workflows.py:369-370`, and the fakes in
|
||||
`tests/providers/test_openrouter.py`, `tests/services/test_workflows_reliability.py`, and
|
||||
`tests/test_provider_boundaries.py`, all of which implement or assert the current property-based
|
||||
contract.
|
||||
- **Recommendation:** **Fix before introducing any intra-job page concurrency.** The durable fix is
|
||||
to stop returning evidence through instance state: attach `request_manifest` and
|
||||
`transport_evidence` to the raised exception on every failure path — which `ProviderError`
|
||||
already supports (`providers/base.py:18-29`) and which the timeout path cannot currently use
|
||||
because `asyncio.wait_for` raises `TimeoutError` from outside the adapter. A narrower option is
|
||||
to have `transcribe()` accept a caller-owned capture sink so evidence is scoped to the call
|
||||
rather than to the adapter. As an immediate, near-zero-cost step, document the non-reentrancy on
|
||||
the protocol in `providers/base.py` so the constraint is visible where it is depended upon.
|
||||
- **Effort:** M
|
||||
|
||||
---
|
||||
|
||||
#### [MED-06] Provider error bodies reach user-facing text while three provider failure paths persist no `detail`
|
||||
|
||||
- **Location:** `src/transcription/services/sources.py:923-947` (`handle_transcription_errors`);
|
||||
message construction at `src/transcription/providers/openrouter.py:414-431`
|
||||
(`_transport_error_message`)
|
||||
- **Reachability:** **Live** for the message half (any provider failure during a UI-initiated
|
||||
transcription surfaces through `show_error`).
|
||||
- **Problem & Consequence:** Two mirrored halves of the same rule are broken in one function.
|
||||
- `sources.py:943` builds `f"Provider transcription failed: {exc}"`, and `exc` is a
|
||||
`ProviderError` whose message may embed up to 500 characters of the provider's error body
|
||||
(`openrouter.py:430`). That is a provider payload in `message`, which
|
||||
`error-handling.instructions.md:85` explicitly forbids.
|
||||
- None of the three handlers (lines 929, 935, 942) passes `detail=`. Per
|
||||
`error-handling.instructions.md:89-92`, omitting it degrades the provenance record.
|
||||
|
||||
I checked whether the provenance half is actually harmful before reporting it, and it is
|
||||
**substantially mitigated**: `workflows.py:391` calls `_find_provider_error`, which walks
|
||||
`__cause__`/`__context__` (`workflows.py:806-813`) to recover the original `ProviderError` and
|
||||
persists its `transport_evidence` — status code, safe headers, and the exact response body — onto
|
||||
the attempt. So the root cause is preserved in transport evidence even though `error_detail` is
|
||||
thin. This is why the finding is Medium rather than High. The residual cost is that the
|
||||
human-readable failure summary is uninformative for the two paths (`ProviderAuthError`,
|
||||
`ProviderResponseError`) whose messages are entirely generic.
|
||||
- **Blast Radius:** `handle_transcription_errors` is used on the transcription path in
|
||||
`sources.py`; `TranscriptionError.message` is consumed by `error_presenter.show_error`,
|
||||
`build_error_envelope`, and `format_error_detail`. Assertions on these messages live in
|
||||
`tests/services/test_transcription_service.py` and `tests/providers/test_openrouter.py`.
|
||||
- **Recommendation:** Move the interpolated provider text from `message` to `detail` on all three
|
||||
handlers, keeping the generic message the other two already use:
|
||||
```python
|
||||
except ProviderError as exc:
|
||||
raise TranscriptionError(
|
||||
"Provider transcription failed",
|
||||
category=ErrorCategory.EXTERNAL_PROVIDER,
|
||||
suggestion="Retry the transcription from jobs. If repeated, check provider availability.",
|
||||
retriable=True,
|
||||
detail=f"{type(exc).__name__}: {exc}",
|
||||
) from exc
|
||||
```
|
||||
Apply the same `detail=` addition to the `ProviderAuthError` and `ProviderResponseError`
|
||||
handlers. Note the interaction with HIGH-01: until the render boundary is sanitized, moving text
|
||||
into `detail` still reaches users through the `error_detail` display. Sequence accordingly.
|
||||
- **Effort:** S
|
||||
|
||||
---
|
||||
|
||||
#### [MED-07] No deterministic guard covers the `message`/`detail` split at explicit raise sites
|
||||
|
||||
- **Location:** `tests/test_errors.py:56-78`; rule at
|
||||
`.github/instructions/error-handling.instructions.md:78-98`; canonical statement at
|
||||
`docs/error_handling.md:102-115`
|
||||
- **Reachability:** **Live** — this coverage gap is what allowed HIGH-02 and MED-06 to exist in a
|
||||
fully green suite.
|
||||
- **Problem & Consequence:** `docs/error_handling.md:115` names
|
||||
`tests/test_errors.py::test_unexpected_error_does_not_leak_filesystem_paths` as the enforcement
|
||||
for the message/detail split. That test exercises exactly one function,
|
||||
`classify_unexpected_error`. Every direct `raise SomeAppError(...)` in `src/` — roughly 50 sites
|
||||
by grep — is unenforced. The documentation therefore overstates the enforcement, which is worse
|
||||
than having no guard: a contributor reading `error_handling.md:115` reasonably concludes the rule
|
||||
is mechanically protected.
|
||||
|
||||
Per the reviewer skill, where a check is unenforced, recommending the deterministic test is
|
||||
itself a finding.
|
||||
- **Blast Radius:** Tests only.
|
||||
- **Recommendation:** Add an AST guard, `tests/test_error_message_safety.py`, that scans `src/`
|
||||
for `raise <AppError subclass>(...)` and fails when the first positional argument is an f-string
|
||||
containing a formatted value whose name matches a path-like or exception-like identifier
|
||||
(`path`, `_path`, `root`, `dir`, `exc`, `err`, `e`). Model it on the existing AST guards, which
|
||||
are the established pattern here (`test_ui_boundaries.py`, `test_service_boundaries.py`,
|
||||
`test_orphan_sweep.py`). Pair it with a second guard asserting that no UI module reads
|
||||
`error_detail` without routing through the sanitizing projection from HIGH-01 — that one closes
|
||||
the render side, which is where the real leak is.
|
||||
- **Effort:** M
|
||||
|
||||
---
|
||||
|
||||
### Low Severity
|
||||
|
||||
#### [LOW-01] `ui/pages/tags_page.py` is dead code; the V6.1 roadmap is correct
|
||||
|
||||
- **Location:** `src/transcription/ui/pages/tags_page.py` (93 lines);
|
||||
registration list at `src/transcription/ui/__init__.py:37-43`
|
||||
- **Reachability:** **Not reachable.** This resolves the flagged open item: the route is genuinely
|
||||
**not** registered. `register_pages` calls seven page registrars and `tags_page` is not among
|
||||
them; nothing anywhere imports the module. `docs/roadmap_plan.md:47` ("Retire the Tags page") is
|
||||
accurate, and `tests/ui/test_tags_page.py` correctly asserts `/ui/tags` returns 404 — though it
|
||||
passes trivially, since an unimported module cannot register anything.
|
||||
- **Problem & Consequence:** No runtime risk; purely stranded code. It is worth noting that if it
|
||||
*were* ever re-registered, `tags_page.py:71-74` traverses `document.document_tags` and
|
||||
`link.tag_ref` inside a page render, and those relationships are configured `lazy="raise"`
|
||||
(`docs/architecture.md:200-203`) — so re-enabling this module without adding eager loads to
|
||||
`list_documents` would raise on first render.
|
||||
- **Recommendation:** Delete `src/transcription/ui/pages/tags_page.py`. Retain
|
||||
`tests/ui/test_tags_page.py` as the retirement guard. Fixing MED-01 first would make this
|
||||
finding reproducible by the suite rather than by manual inspection.
|
||||
- **Effort:** S
|
||||
|
||||
#### [LOW-02] `benchmarking.py` ships in the runtime package but is referenced only by tests
|
||||
|
||||
- **Location:** `src/transcription/benchmarking.py` (69 lines); sole consumers
|
||||
`tests/test_v42_evidence.py:15-16` (`EditorialAssessment`, `score_transcription`)
|
||||
- **Reachability:** Live as importable API; never invoked by application code.
|
||||
- **Problem & Consequence:** No defect. It supports the model-evaluation policy in
|
||||
`docs/invariant/ai_evidence_and_provenance.md:113-126`, which is legitimate, but it currently has
|
||||
no production caller and no tooling entrypoint, so it is indistinguishable from drift.
|
||||
- **Recommendation:** Either move it under `tools/` alongside the other operator utilities, or add
|
||||
a `KNOWN_ORPHANS`-style rationale recording that it is retained as the evaluation-policy
|
||||
implementation. Do not silently keep it unlabeled.
|
||||
- **Effort:** S
|
||||
|
||||
#### [LOW-03] Two overlapping prompt error types split across modules
|
||||
|
||||
- **Location:** `src/transcription/services/errors.py:15-16` (`PromptLoadError`) and
|
||||
`src/transcription/services/prompts.py:19` (`PromptStoreError`)
|
||||
- **Reachability:** Live; no misbehavior observed.
|
||||
- **Problem & Consequence:** `services/errors.py:1-8` documents itself as the neutral home for
|
||||
exceptions raised by more than one service, precisely so a caller's `except` clause does not
|
||||
change when an operation moves. `PromptStoreError` is defined outside that module and covers an
|
||||
overlapping domain (prompt file access), so a caller wanting to handle "any prompt failure" must
|
||||
import from two modules and know which is which. `sources.py:855` raises `PromptLoadError` for a
|
||||
missing prompt file while `prompts.py:132` raises `PromptStoreError` for the same condition
|
||||
reached through the Settings page.
|
||||
- **Recommendation:** Move `PromptStoreError` into `services/errors.py` next to `PromptLoadError`,
|
||||
or make one a subclass of the other so a single `except` covers prompt failures. Low urgency; do
|
||||
it opportunistically when HIGH-02 touches both files anyway.
|
||||
- **Effort:** S
|
||||
|
||||
---
|
||||
|
||||
## 4. Architectural Drift & Gap Analysis
|
||||
|
||||
| Area / Component | Direction | Documented / Intended Rule | Actual Implementation State | Severity | Recommended Resolution |
|
||||
| :--- | :--- | :--- | :--- | :--- | :--- |
|
||||
| Error presentation | `doc->code` | `docs/error_handling.md:110` — `detail` is internal only, surfaced by `format_error_detail` and logs | `sources_page.py:484`, `table/sources.py:90`, `settings_page.py:562` render `error_detail` verbatim to users | High | Sanitizing render projection (HIGH-01); do **not** strip `detail` |
|
||||
| User-facing messages | `doc->code` | `invariant/error_handling.md:59` — no local filesystem detail in user-facing messages | 5 live sites interpolate absolute paths into `AppError.message` | High | Generic `message`, path on `detail` (HIGH-02) |
|
||||
| Evidence invariant ownership | `doc->doc` | `docs/index.md:1` — baseline is V6.1 | `invariant/ai_evidence_and_provenance.md:130` names "Canonical V4"; the currency guard scans the file but its regexes match neither the phrasing nor a minor-less `V4` | Medium | Update text; make the guard token-based, or drop version labels entirely (MED-02, §8.1) |
|
||||
| Enforcement claim | `doc->code` | `docs/error_handling.md:115` — split "Enforced by `tests/test_errors.py::…`" | That test covers only `classify_unexpected_error`; explicit raises unguarded | Medium | Add AST guard (MED-07) |
|
||||
| Orphan sweep | `doc->code` | `test_orphan_sweep.py:1-13` — sweep is "deterministic" and "conservative" | Bare-name matching; cannot see a dead module (`tags_page.py`) | Medium | Qualify by module + module-reachability pass (MED-01) |
|
||||
| Worker scaling | `code->doc` | `jobs.py:191-197` — `SKIP LOCKED` so "concurrent workers never contend" | Claim path is multi-worker-safe; stale-reclaim path is not | Medium | Derive stale threshold from job duration; document single-worker constraint until fixed (MED-04) |
|
||||
| Provider adapter contract | `code->doc` | `providers/base.py:110-118` — evidence read as "most recent call" state | Contract is silently non-reentrant; safety lives in `workflows.py`'s sequential loop | Medium | Scope evidence to the call; document non-reentrancy (MED-05) |
|
||||
| Settings env file | `code->doc` | `config.py:66-75` — `env_file=".env.production"` | Correctness depends on an undocumented cwd contract with `Dockerfile` `WORKDIR /app` | Medium | Explicit resolver seam, or document + guard the contract (MED-03) |
|
||||
| Tags page | *(no drift)* | `roadmap_plan.md:47` — Tags page retired | Route genuinely unregistered; module is stranded code | Low | Delete the module (LOW-01) |
|
||||
|
||||
---
|
||||
|
||||
## 5. Invariant Inventory & Routing Recommendations
|
||||
|
||||
| Invariant / Constraint | Current Location | Recommended Target Layer | Rationale |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| `detail`/`error_detail` never rendered to users | docs + instructions | **Deterministic test** + sanitizing projection | Stated in three documents and violated in three files; prose has demonstrably failed to hold it |
|
||||
| `message` carries no paths or exception text | instructions; partial test | **Deterministic test** (AST, all raise sites) | Existing guard covers one function; the gap produced HIGH-02 |
|
||||
| `ExecutionAttempt.error_detail` retains root cause | docs + `test_v42_evidence.py` | **Keep in tests** — already correct | Counterweight to the above; must be named in any fix so it is not over-corrected |
|
||||
| Intermediate pages commit individually | `workflows.py` docstring + `test_workflows_reliability.py:318` | **Keep in tests** — verified genuine | Cross-session read makes it a real durability assertion |
|
||||
| Final page atomic with terminal status | `services.instructions.md` + `test_pipeline_atomicity.py` | **Keep in tests** — verified genuine | Fault injection makes a split commit fail |
|
||||
| Canonical baseline version consistency | `docs/index.md` + `test_meta_contract_guards.py:89` | **Repair existing test, or remove the labels** | Guard exists but matches by approved phrase and requires a minor version, so it fails open on new phrasings (MED-02) |
|
||||
| Module-level reachability / dead modules | `test_orphan_sweep.py` (ineffective) | **Deterministic test** (repair existing) | Guard exists but cannot detect the case (MED-01) |
|
||||
| Stale threshold > max job duration | *(unenforced)* | **Config validator + test** | Currently a coincidence of two equal defaults (MED-04) |
|
||||
| Provider adapter non-reentrancy | *(unenforced, implicit)* | **Instructions** + protocol docstring | A design constraint callers must know before adding concurrency (MED-05) |
|
||||
| Env-file resolution independent of cwd | `tests/conftest.py` monkeypatch | **Code seam** + `docs/production-runbook.md` | A test-only fix for a production coupling is misrouted enforcement (MED-03) |
|
||||
|
||||
---
|
||||
|
||||
## 6. Stack-Specific Analysis
|
||||
|
||||
**Python 3.12+.** Modern and consistent. PEP 695 generics are used correctly and non-trivially
|
||||
(`RegistryService[ModelT: RegistryEntry]` in `services/registry.py:58`, `UiActionOutcome[T]`,
|
||||
`_get_or_raise[ModelT]`), `type` statements appear in `db/session.py:15,50`, and `X | None` is
|
||||
used throughout. `structural Protocol` bounds (`RegistryEntry`, `WorkerNotifier`,
|
||||
`TranscriptionProvider`) are used to avoid type suppressions rather than to decorate. `ty` passes
|
||||
clean with no suppressions found. The two `# noqa` uses (`workflows.py:228` `PLR0915`,
|
||||
`workflows.py:383` `BLE001`) are both justified in context — the broad catch is a deliberate
|
||||
per-page containment boundary that immediately classifies and re-records.
|
||||
|
||||
**FastAPI.** Lifespan is handled via `@asynccontextmanager` (`app.py:36`), not the deprecated
|
||||
`@app.on_event`. Session factories are injected through `Depends` (`SessionFactoryDep`,
|
||||
`db/session.py:50`) rather than reached as globals from routes. `api/errors.py` centralizes
|
||||
envelope translation. One residual: `get_settings` is `@cache`d and read as a module-level
|
||||
fallback in ~10 modules; this is acceptable given the documented restart-to-apply contract
|
||||
(`docs/ui/pages/settings.md:28`) but means the cache is process-lifetime and unclearable.
|
||||
|
||||
**NiceGUI (pinned `3.13.0`).** The pin is a recorded release-stability decision and is not
|
||||
reported as a defect. Boundaries are enforced structurally: `test_ui_boundaries.py` uses an
|
||||
import **allowlist**, which is the right polarity. Blocking work is dispatched off the event loop
|
||||
via `run_blocking` (`settings_page.py:818,822`). The one boundary that is *not* enforced is
|
||||
presentation of internal fields (HIGH-01) — pages are prevented from touching persistence but not
|
||||
from rendering internal-only text.
|
||||
|
||||
**SQLModel / SQLAlchemy.** Strong. `lazy="raise"` on relationships forces explicit eager loading;
|
||||
read paths declare `selectinload` chains with comments explaining *why* each is needed
|
||||
(`sources.py:309-316` is a good example). `expire_on_commit=False` (`db/session.py:28`) is set
|
||||
deliberately, which is what makes post-commit attribute access in `evidence.py:159-202` safe.
|
||||
`claim_next_queued_job` (`jobs.py:186-241`) branches correctly on dialect — `SKIP LOCKED` on
|
||||
PostgreSQL, conditional `UPDATE ... RETURNING` on SQLite — rather than assuming one engine.
|
||||
Attempt-number allocation uses `begin_nested()` with bounded retry (`sources.py:598-616`), the
|
||||
right pattern for a monotonic per-parent sequence. No N+1 patterns were found in the read paths
|
||||
sampled.
|
||||
|
||||
**Pydantic V2 & Settings.** Fully V2; no `@validator`, `class Config`, `.dict()`, or `parse_obj`
|
||||
anywhere. Evidence contracts use `ConfigDict(extra="forbid", frozen=True)` (`providers/evidence.py:47`),
|
||||
which is exactly right for persisted provenance — an unexpected field fails loudly rather than
|
||||
being silently dropped. `SecretStr` guards the API key. The discriminated
|
||||
`SqliteSettings | PostgresSettings` union is clean. `normalize_provider_models` correctly runs
|
||||
`mode="before"` so the derived tuple is produced by construction rather than by mutating a frozen
|
||||
model — a subtlety that is easy to get wrong. Sole issue: the cwd-coupled `env_file` (MED-03).
|
||||
|
||||
**Asyncio Workers.** Notably careful. `asyncio.shield` wraps both the per-page commit and the
|
||||
terminal commit (`workflows.py:601-614`, `650-663`), with the `except CancelledError: await task;
|
||||
raise` pattern that actually completes the shielded work rather than merely deferring cancellation —
|
||||
a detail most implementations get wrong. `handle_worker_exceptions` (`worker.py:157-182`)
|
||||
distinguishes retriable from non-retriable faults and stops the loop rather than spinning.
|
||||
`_advance_job_with_containment` (`worker.py`/`workflows.py:507-540`) guarantees a claimed job
|
||||
cannot strand in `PROCESSING`. `worker_consumer_lifespan` has a bounded shutdown with escalation to
|
||||
`cancel()`. Gaps are MED-04 and MED-05, both latent and both with stated unblocking conditions.
|
||||
|
||||
**OpenRouter / Adapter Boundary.** Encapsulation holds: `test_provider_boundaries.py` enforces it,
|
||||
and `workflows.py` imports only `providers` abstractions. `_CapturingAsyncClient` is a
|
||||
well-judged design — it captures the exact transport body before SDK parsing without altering what
|
||||
the SDK consumes, including the streamed case. Timeout construction (`openrouter.py:200-206`)
|
||||
correctly overrides httpx's 5s per-phase default that would otherwise silently cap the configured
|
||||
budget. `SAFE_RESPONSE_HEADERS` (`providers/evidence.py:29-41`) was reviewed field-by-field:
|
||||
all nine entries are non-secret correlation, content, or rate-limit headers, and
|
||||
`filter_safe_response_headers` is a true allowlist filter with no redaction-after-capture — this
|
||||
satisfies invariant 3.8.2 exactly. `_replace_embedded_media` correctly substitutes a source
|
||||
reference for base64 payloads, satisfying 3.8.3. The one structural weakness is MED-05.
|
||||
|
||||
**Testing & Quality Tooling.** 410 tests, all green, with genuinely strong contract guards
|
||||
(boundaries, model contract, media path safety, evidence append-only, atomicity). Marker strictness
|
||||
and `asyncio_mode = "strict"` are configured, and no unawaited-coroutine warnings appeared. Two
|
||||
guards, however, assert meaningfully less than their names and docstrings claim
|
||||
(`test_orphan_sweep.py` — MED-01; `test_errors.py` path-leak coverage — MED-07), and the
|
||||
`>= 420` snapshot threshold at `test_orphan_sweep.py:169` repeats a weak-assertion pattern the
|
||||
2026-08-23 review already flagged at `>= 200`; it was raised rather than replaced with
|
||||
set-membership.
|
||||
|
||||
---
|
||||
|
||||
## 7. Duplication & Consolidation Report
|
||||
|
||||
| Pattern / Duplication | Locations | Proposed Canonical Home | Estimated Lines Removed |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| `f"{type(exc).__name__}: {exc}"` detail construction | `errors.py:118`, `maintenance.py:71,95,135,206`, `runtime_settings_store.py:388,476,554` | `errors.py::exception_detail(exc)` | ~8 (consistency > line count) |
|
||||
| Filesystem `AppError` construction from `OSError` | `prompts.py:186-191`, `runtime_settings_store.py:384-389,472-477,550-555` | `errors.py::filesystem_error(message, exc, *, suggestion)` | ~20 |
|
||||
| Overlapping prompt error types | `services/errors.py:15`, `services/prompts.py:19` | `services/errors.py` (LOW-03) | ~5 |
|
||||
| Duplicated `provider_duration_ms` / `processing_duration_ms` max-clamp arithmetic | `workflows.py:341-345, 361-368, 397-404` | `workflows.py::_page_durations(started_at, finished_at, monotonic_started_at)` | ~20 |
|
||||
| `_utc_now_naive` defined per module | `workflows.py:51`, `jobs.py:29`, `db/models.py`, `sources.py` | Single helper in `db/models.py`, imported | ~12 |
|
||||
|
||||
### Proposed Canonical Abstractions
|
||||
|
||||
```python
|
||||
# src/transcription/errors.py
|
||||
def exception_detail(exc: BaseException) -> str:
|
||||
"""Internal-only root-cause text for AppError.detail. Never user-facing."""
|
||||
|
||||
|
||||
def filesystem_error[E: AppError](error_type: type[E], message: str, exc: OSError, *, suggestion: str) -> E:
|
||||
"""Build a filesystem AppError with a generic message and the path on detail."""
|
||||
|
||||
|
||||
# src/transcription/ui/components/error_presenter.py
|
||||
def display_failure_detail(error_detail: str | None) -> str | None:
|
||||
"""Sanitize persisted failure detail for UI rendering (HIGH-01)."""
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 8. Meta-Tooling & Instruction Update Recommendations
|
||||
|
||||
1. **`docs/invariant/ai_evidence_and_provenance.md:130`** — resolve the V4 label. Two viable
|
||||
routes, and the maintainer has proposed the second:
|
||||
- **(a) Repair the guard.** Fix the text to V6.1 and make
|
||||
`test_canonical_docs_declare_one_consistent_baseline` token-based rather than phrase-based
|
||||
(MED-02). Keeps version labels as navigational anchors.
|
||||
- **(b) Remove version labels from canonical docs.** While the project has a single principal
|
||||
user and no released versions to support, "canonical" and "current" are the same thing, so the
|
||||
label carries no information a reader can act on — it only creates a second thing to keep in
|
||||
sync. Retain the baseline declaration in `docs/index.md` alone as the release marker, keep
|
||||
version language in `docs/roadmap_plan.md` and the migration/deployment docs (already
|
||||
excluded from the scan for exactly this reason), and replace in-body references with
|
||||
unversioned phrasing ("the canonical architecture, schema, requirements, and error-policy
|
||||
documents"). The guard then inverts: assert that no canonical doc outside the exclusion set
|
||||
contains a version token at all, which is a stricter and much cheaper property to hold than
|
||||
agreement between many labels. Requirement IDs (`REQ-4-021`, `REQ-6-014`) are stable
|
||||
identifiers, not currency claims, and should be left alone.
|
||||
2. **`docs/error_handling.md:107-115`** — either add the sanitizing-projection rule for UI display
|
||||
of `error_detail`, or revise the `detail` "Surfaces" row to admit operator-facing evidence
|
||||
displays. Update the "Enforced by" line once MED-07's guard lands, since it currently overstates
|
||||
coverage.
|
||||
3. **`.github/instructions/error-handling.instructions.md`** — add an explicit clause under
|
||||
"User-Safe Messaging" stating that *persisted* `error_detail` is subject to the same no-paths
|
||||
rule at any render boundary. The current table (line 86) states the rule for `AppError.detail`
|
||||
and stops there, so the persisted-then-rendered path falls between the lines.
|
||||
4. **`.github/instructions/providers.instructions.md`** — record the adapter non-reentrancy
|
||||
constraint (MED-05); it is currently an undocumented precondition of `workflows.py`.
|
||||
5. **`.github/instructions/services.instructions.md`** — the two competing atomicity invariants are
|
||||
well described and both guards verified; no change needed. Worth adding the stale-reclaim
|
||||
threshold constraint (MED-04) alongside them, since it is a third worker-lifecycle rule with no
|
||||
documented home.
|
||||
6. **`tests/test_orphan_sweep.py`** — repair per MED-01 and replace the `>= 420` threshold with
|
||||
set-membership assertions.
|
||||
7. **New `tests/test_error_message_safety.py`** — AST guard per MED-07, covering both the raise
|
||||
sites and the UI render sites.
|
||||
8. **`docs/production-runbook.md`** — document the cwd/`WORKDIR` contract for `.env.production`
|
||||
resolution if MED-03 is resolved by documentation rather than by a code seam.
|
||||
|
||||
---
|
||||
|
||||
## 9. Prioritized Dependency-Ordered Action Plan
|
||||
|
||||
**Phase 1 — Blocking fixes (privacy; ordered, HIGH-01 first)**
|
||||
1. **HIGH-01** — add `display_failure_detail` and route `sources_page.py:484`,
|
||||
`table/sources.py:90-95`, and `settings_page.py:562` through it. Do this **first**: it closes
|
||||
the render boundary, so the Phase-1.2 fix cannot relocate a leak again.
|
||||
2. **HIGH-02** — move paths from `message` to `detail` at the five sites, including the
|
||||
`prompts.py:186` double violation.
|
||||
3. **MED-06** — move provider payload text to `detail`; add `detail=` to all three
|
||||
`handle_transcription_errors` handlers.
|
||||
|
||||
**Phase 2 — Enforcement hardening (make Phase 1 permanent)**
|
||||
4. **MED-07** — AST guard for raise-site `message` safety **and** for UI reads of `error_detail`.
|
||||
5. **MED-01** — qualify orphan definitions by module; add module-reachability; replace the
|
||||
snapshot threshold.
|
||||
6. **MED-02** — fix the V4/V6.1 text and extend the meta-contract guard to baseline-version currency.
|
||||
|
||||
**Phase 3 — Reliability & concurrency (latent; each must precede its unblocking change)**
|
||||
7. **MED-04** — derive `worker_stale_job_seconds` from `worker_provider_timeout_seconds` with a
|
||||
rejecting validator, ideally plus a progress heartbeat. **Must land before any second worker
|
||||
replica.**
|
||||
8. **MED-05** — scope provider evidence to the call rather than the instance. **Must land before
|
||||
any intra-job page concurrency.** Document non-reentrancy immediately as an interim step.
|
||||
|
||||
**Phase 4 — Consolidation & refactoring**
|
||||
9. **MED-03** — explicit env-file resolution seam shared by `Settings` and `runtime_settings_store`;
|
||||
remove the `model_config` monkeypatch from `conftest.py`.
|
||||
10. **LOW-01** — delete `tags_page.py` (after MED-01, so the suite reproduces the finding).
|
||||
11. **LOW-03** and the §7 consolidations — fold in opportunistically while Phase 1 touches these files.
|
||||
|
||||
**Phase 5 — Non-blocking governance/documentation depth**
|
||||
12. **LOW-02** — relocate or annotate `benchmarking.py`.
|
||||
13. Instruction/doc updates §8.3–§8.5, §8.8.
|
||||
|
||||
---
|
||||
|
||||
## 10. Preserved Strengths
|
||||
|
||||
- **Append-only evidence is real, not aspirational.** Every provider call produces a distinct
|
||||
`ExecutionAttempt`; no runtime path mutates a historical row. Projection writes onto
|
||||
`Source.raw_transcription` are clearly separated from history, and `promote_machine_attempt`
|
||||
(`evidence.py:117-146`) repoints the projection without rewriting evidence — with a docstring
|
||||
that explains exactly why that one write lives in a read-oriented service.
|
||||
- **Transport-layer terminology is honored in code.** `_CapturingAsyncClient` exists specifically so
|
||||
the stored body is the application-boundary capture rather than an SDK-parsed object, and
|
||||
`TransportEvidence(response_received=False)` explicitly represents "no response" instead of
|
||||
conflating it with an empty one. This is invariant 3.4/3.5 implemented rather than asserted.
|
||||
- **Header allowlisting is done the hard, correct way** — filter-before-store with an explicit
|
||||
frozenset, never capture-then-redact (`providers/evidence.py:29-41,130-134`).
|
||||
- **Boundaries are enforced by allowlist, not blocklist.** `test_ui_boundaries.py:20-25` states the
|
||||
reasoning explicitly; it means a newly added persistence helper cannot slip through under an
|
||||
unlisted name.
|
||||
- **The two competing atomicity invariants are both correctly implemented and both genuinely
|
||||
guarded**, with the tests structured so that the naive over-correction fails.
|
||||
- **Cancellation safety in the worker is unusually well handled** — `asyncio.shield` plus
|
||||
`await task` on `CancelledError` actually completes the commit rather than merely deferring
|
||||
cancellation.
|
||||
- **Comments explain rationale, not mechanics.** `workflows.py:269-271`, `openrouter.py:200-202`,
|
||||
`config.py:114-115`, and `jobs.py:191-197` each record *why* a non-obvious choice was made,
|
||||
several citing the review log entry that motivated it. This is what made verifying the atomicity
|
||||
and timeout invariants tractable in this review.
|
||||
- **Documentation-to-code traceability is strong overall.** Page contracts, schema field tables,
|
||||
and requirement IDs are maintained and guarded; the drift found in this review is narrow and
|
||||
specific rather than systemic.
|
||||
|
||||
---
|
||||
|
||||
## Appendix A — Repo-Specific Deterministic Checks
|
||||
|
||||
| # | Check | Result | Evidence |
|
||||
| :-- | :--- | :--- | :--- |
|
||||
| 1 | Service boundary rule: no service-to-service imports | **Pass** | `tests/test_service_boundaries.py` green; AST scan, allowlist-based; `workflows.py` composes via `ServiceBundle` |
|
||||
| 2 | UI boundary rule: no persistence access from pages/components | **Pass (structurally)** | `tests/test_ui_boundaries.py` green. Caveat: it guards *data access*, not presentation of internal-only fields — see HIGH-01 |
|
||||
| 3 | Status vocabulary conformance; no stringly-typed literals | **Pass** | `tests/test_model_contract_guards.py` green; enum members verified against `db/models.py` |
|
||||
| 4 | Evidence ownership: append-only history, projections not history mutation | **Pass** | `test_v42_evidence.py::test_attempts_are_append_only_and_exported_with_integrity` verified non-vacuous (asserts both retained attempts and export integrity at lines 281-290) |
|
||||
| 5 | Canonical authority: findings resolve against `docs/*` first | **Pass with defect** | `test_canonical_authority_references_are_present` green. The companion baseline-currency guard (`test_canonical_docs_declare_one_consistent_baseline`) scans the offending file but fails open on its phrasing and on minor-less version tokens — MED-02 |
|
||||
| 6 | Schema contract fidelity: `docs/schema.md` field-accurate | **Pass** | `test_model_contract_guards.py` + `test_meta_contract_guards.py` green |
|
||||
| 7 | Media boundary: record-validated media, controlled URL resolver | **Pass** | `test_media_path_safety.py`, `tests/ui/test_media_urls.py` green; `public_media_path_label` verified path-safe |
|
||||
| 8 | Eager-loading conformance vs `lazy="raise"` | **Pass** | Declaration-side guard green; sampled read paths declare explicit `selectinload` chains. Note: dead `tags_page.py:71-74` would violate it if re-registered (LOW-01) |
|
||||
| 9 | Cross-cutting error conformance | **FAIL** | Guards green but coverage is narrower than documented: HIGH-01, HIGH-02, MED-06, MED-07 |
|
||||
| 10 | Orphan/dead-code conformance | **FAIL** | Guard green but structurally unable to detect a dead module: MED-01, proven by LOW-01 |
|
||||
|
||||
## Appendix B — Evidence & Provenance Auditor Families
|
||||
|
||||
| Family | Subject | Result | Evidence |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| A | Attempt history append-only | **Pass** | No update/delete path to `ExecutionAttempt`; insert-only with `begin_nested` + bounded sequence retry (`sources.py:556-616`) |
|
||||
| B | Attempt numbering monotonic per source | **Pass** | `insert_with_sequence_retry`; uniqueness constraint plus retry on conflict |
|
||||
| C | Transport evidence captured at the transport boundary | **Pass** | `_CapturingAsyncClient` retains the exact wire body pre-SDK-parse (`openrouter.py:66-94`) |
|
||||
| D | Absent response distinguished from empty response | **Pass** | `TransportEvidence.response_received` is explicit, not inferred |
|
||||
| E | Response header persistence is allowlist-based | **Pass** | `SAFE_RESPONSE_HEADERS` (`providers/evidence.py:29-41`) — all nine entries verified non-secret; filter-before-store |
|
||||
| F | No machine-local detail on user-facing surfaces | **FAIL** | `error_detail` rendered verbatim at three UI sites (HIGH-01); paths in `message` at five sites (HIGH-02) |
|
||||
| G | Request manifest excludes embedded media payloads | **Pass** | `_replace_embedded_media` (`openrouter.py:377-395`) substitutes a source reference for base64 data |
|
||||
| H | Evidence attribution is correct under concurrency | **Pass today / at risk** | Correct in the current sequential single-worker deployment; the contract itself is non-reentrant (MED-05) and reclaim has no margin (MED-04) |
|
||||
@@ -0,0 +1,21 @@
|
||||
# Review Reports
|
||||
|
||||
Dated architecture and code review reports generated by
|
||||
`.github/skills/python-code-reviewer/skill.md`.
|
||||
|
||||
**These files are not canonical authority.** Everything in `docs/reviews/**` is a
|
||||
point-in-time observation, not a contract. Canonical intent lives in `docs/index.md`,
|
||||
`docs/architecture.md`, `docs/requirements.md`, `docs/schema.md`,
|
||||
`docs/error_handling.md`, and `docs/invariant/**`. When a report and a canonical
|
||||
document disagree, the canonical document wins until it is deliberately updated.
|
||||
|
||||
Naming: `<YYYY-MM-DD>-code-review.md` for review reports, and
|
||||
`<YYYY-MM-DD>-remediation-handoff.md` for the implementation plan derived from one.
|
||||
|
||||
## Current
|
||||
|
||||
- [`2026-08-23-code-review.md`](./2026-08-23-code-review.md) — full review. 0 critical,
|
||||
4 high, 5 medium, 11 low. **All findings remediated.** Retained as a record of the
|
||||
reasoning, not as a list of open work. Note that a few of its recommendations were
|
||||
wrong on contact and were corrected during implementation; the code and the guard
|
||||
tests are authoritative over the report text.
|
||||
@@ -0,0 +1,210 @@
|
||||
# Roadmap Plan (Starting at V6.0)
|
||||
|
||||
This roadmap starts at **V6.0** and tracks forward-looking work only.
|
||||
|
||||
## V6.0 - Hosting Migration
|
||||
|
||||
Objective: move from local-only operation to secure, stable remote hosting.
|
||||
Status: **Completed**
|
||||
|
||||
Detailed plan: [`v6_0_hosting_migration_plan.md`](v6_0_hosting_migration_plan.md)
|
||||
|
||||
### Scope
|
||||
1. Containerize app runtime for production deployment.
|
||||
2. Run PostgreSQL in Docker and migrate from SQLite.
|
||||
3. Add Cloudflare Tunnel exposure with Access protection.
|
||||
4. Add operational safeguards (health checks, restart policies, backups).
|
||||
|
||||
### Deliverables
|
||||
- Production-ready `docker-compose` deployment for app + database + tunnel.
|
||||
- Environment-based configuration for DB, uploads, prompts, and logging.
|
||||
- Verified data migration path into PostgreSQL.
|
||||
- Runbook updates for deploy, rollback, and backup/restore.
|
||||
|
||||
### Exit Criteria
|
||||
- `/healthz` reports healthy app and worker in deployed environment.
|
||||
- One end-to-end document -> source -> job workflow succeeds remotely.
|
||||
- Backup and restore procedure is tested.
|
||||
|
||||
### Accomplished
|
||||
1. Delivered production Docker deployment with split `app`/`worker`, `postgres`, and `cloudflared`.
|
||||
2. Landed SQLite -> PostgreSQL migration tooling and runbook coverage.
|
||||
3. Added production health/reliability wiring and operational runbooks for deploy/rollback/recovery.
|
||||
4. Established host-visible backup workflow and restore path for PostgreSQL plus media/config assets.
|
||||
|
||||
## V6.1 - Testing and Refinement
|
||||
|
||||
Objective: improve navigation and operational workflows after user feedback.
|
||||
Status: **Completed**
|
||||
|
||||
### Scope
|
||||
1. Make Document Detail the primary source-page workspace:
|
||||
- Use Source-style pan/zoom + previous/next page controls.
|
||||
- Move editable revision controls into Document Detail.
|
||||
- Move archival/system metadata to dedicated Document Info route.
|
||||
2. Simplify top navigation:
|
||||
- Remove top-level Tags and Sources entries.
|
||||
- Retire the Tags page and the global Source Asset Records entry flow.
|
||||
3. Improve list/detail clarity:
|
||||
- Add Document transcription status to Archival Documents list.
|
||||
- Add Document Date in People Detail -> Linked Documents table.
|
||||
4. Add worker-backed Settings maintenance runs:
|
||||
- Add `maintenance_run` persistence (`id`, `job_type`, `status`, `started_at`, `finished_at`, `triggered_by`, `summary`, `log_path`, `error_detail`).
|
||||
- Add Run Backup and Run Storage Reconciliation actions that enqueue runs and execute in the worker.
|
||||
- Add run history with status, duration, summary, and log view/download.
|
||||
- Defer daily/weekly scheduling controls to V6.2.
|
||||
|
||||
### Accomplished
|
||||
1. Refactored Document Detail into the primary source-page workspace (pan/zoom viewer, previous/next page navigation, editable revision flow) and moved archival/system metadata to Document Info.
|
||||
2. Simplified top navigation by removing Tags/Sources entries and retiring the Tags page/global Source Asset Records flow.
|
||||
3. Improved data clarity with document transcription status in Archival Documents and Document Date in People Detail linked documents.
|
||||
4. Implemented queue-backed maintenance operations (`maintenance_run` model/service/worker/UI) with run history and log view/download.
|
||||
5. Hardened runtime settings operations in production:
|
||||
- runtime settings writes target mounted `.env.production`,
|
||||
- fallback write path for single-file bind mounts,
|
||||
- explicit hidden/deployment-key disclosure in Settings UI.
|
||||
6. Simplified backup configuration and behavior:
|
||||
- standardized on `BACKUP_DIR` + `BACKUP_RETENTION_DAYS`,
|
||||
- backup script uses `DATABASE__*` persistence keys,
|
||||
- compose maps Postgres container init values from `DATABASE__*`,
|
||||
- env contract drift tests now guard `.env.production.example`.
|
||||
|
||||
## V6.2 - GEDCOM Data Layer
|
||||
|
||||
Objective: introduce a genealogical data layer sourced from GEDCOM exports, bridged to
|
||||
existing `Person` records via FamilySearch ID, without disrupting document-focused Person
|
||||
workflows.
|
||||
|
||||
### Scope
|
||||
1. Manual `.ged` file upload only. No FamilySearch credentials are stored or used by the
|
||||
app; the user runs the third-party `getmyancestors` tool themselves and uploads the
|
||||
resulting export.
|
||||
2. Four new tables: `genealogy_person`, `genealogy_family`, `genealogy_family_child`, and
|
||||
`genealogy_citation` (raw GEDCOM `SOUR` citations, reusable in a later version to record
|
||||
when a transcribed document itself becomes citation evidence for FamilySearch).
|
||||
3. Upsert-based import keyed on FamilySearch ID (`fs_id`) so repeat imports update existing
|
||||
records in place without breaking existing `Person.family_search_id` links or duplicating
|
||||
surrogate keys.
|
||||
4. Reuse the existing V6.1 worker-backed `maintenance_run` pattern for import runs (run
|
||||
history, status, summary, log view/download) rather than new infrastructure.
|
||||
|
||||
### Deliverables
|
||||
- GEDCOM parser/importer producing the four genealogy tables.
|
||||
- `MaintenanceJobType` entry for GEDCOM import with upsert semantics and a run summary
|
||||
(records added/updated).
|
||||
- Settings UI entry to upload a `.ged` file, trigger an import run, and view history.
|
||||
|
||||
### Exit Criteria
|
||||
- Importing the same `.ged` file twice does not duplicate or orphan data.
|
||||
- Existing `Person.family_search_id` values continue to resolve to the correct
|
||||
`genealogy_person` row after import.
|
||||
- Import run history is visible with status, duration, and summary, consistent with other
|
||||
maintenance runs.
|
||||
|
||||
## V6.3 - Reporting and Genealogy-Enriched Features
|
||||
|
||||
Objective: improve research value with person-centric outputs, grounded in both archival
|
||||
documents and the V6.2 genealogical data layer.
|
||||
|
||||
This version is broken into five sequential sub-versions because of real dependency
|
||||
ordering: entity linking must exist before GEDCOM data can be targeted per-person; the
|
||||
Facts/Events mechanism must exist before timelines or reconciliation have anything
|
||||
meaningful to consume.
|
||||
|
||||
### V6.3.1 - Manual Entity Linking
|
||||
|
||||
- Search/browse UI over `genealogy_person` to find and link a candidate match to an
|
||||
application `Person`, setting `family_search_id`. Linking is reversible (unlink).
|
||||
- Once linked, GEDCOM vitals display alongside the `Person` record without requiring any
|
||||
schema change to `Person`.
|
||||
|
||||
### V6.3.2 - Person Facts and Events
|
||||
|
||||
- New fact/event table capturing: person, fact type (birth/death/event/free-form), date
|
||||
(+raw), place, free-text description, and a link to the source document as evidence.
|
||||
- Manual tagging UI while reviewing a transcribed document: select a passage, choose the
|
||||
person and fact type, record the date/description.
|
||||
- One-time migration of existing `Person.birth_date`/`birth_date_raw`/`birth_place`/
|
||||
`death_date`/`death_date_raw`/`death_place` values into fact/event rows (tagged as
|
||||
legacy/no-document-evidence where no source document is known), followed by retiring those
|
||||
six columns from `Person`. Birth/death become Facts/Events like any other locally-known
|
||||
fact, for both linked and unlinked people. `Person` permanently keeps `last_name`,
|
||||
`given_names`, `biography`, `family_search_id`, `metadata_`, tags, photos, and document
|
||||
associations.
|
||||
|
||||
### V6.3.3 - Person Timelines
|
||||
|
||||
- Timeline query merging GEDCOM milestones (birth, marriage, children's births, death) for
|
||||
linked persons with locally recorded Facts/Events.
|
||||
- Timeline UI on Person Detail with clear ordering/filters; entries link back to their
|
||||
originating document or GEDCOM record.
|
||||
|
||||
### V6.3.4 - Reconciliation
|
||||
|
||||
- Compares Facts/Events (the real, document-evidenced local signal) against corresponding
|
||||
`genealogy_person` fields for linked persons.
|
||||
- Persisted reconciliation record: person, field, local value with evidence-document link,
|
||||
GEDCOM value, and status (open / submitted / dismissed).
|
||||
- Re-evaluated automatically as part of each GEDCOM import maintenance run: opens new
|
||||
discrepancies, auto-resolves ones where GEDCOM now matches, leaves others unchanged.
|
||||
- Reconciliation review UI functions as a manual to-do list for updating FamilySearch; the
|
||||
app does not write back to FamilySearch itself.
|
||||
|
||||
### V6.3.5 - AI-Assisted Biography Generation
|
||||
|
||||
- Prompted narrative generation grounded in GEDCOM facts, Facts/Events, and relevant
|
||||
document snippets as structured input, using existing evidence-safe prompting patterns.
|
||||
- Output cites back to source documents and FamilySearch records.
|
||||
- Saved/printable report presentation for review; reports do not modify archival source
|
||||
data.
|
||||
|
||||
### Exit Criteria (applies across V6.3.1-V6.3.5)
|
||||
- Entity links are reversible and do not alter document associations.
|
||||
- Timelines are reproducible from persisted records.
|
||||
- Reconciliation items always carry a link to the document evidence justifying the local
|
||||
value, and re-running GEDCOM import correctly opens, resolves, or leaves items unchanged.
|
||||
- Narrative generation is traceable to source records and prompts.
|
||||
- Reports can be reviewed without modifying archival source data.
|
||||
|
||||
## V6.4 - Access Control and Multi-User Readiness
|
||||
|
||||
[ *More thoughts on user accounts:*
|
||||
* *Create a generic "view only" user that does not have the rights to alter any of the data*
|
||||
* *Limit user accounts access to data by Tag. I have distant family members that I would want to share the transcribed data with, but they would only be interested in a subset of it. For example my Cochran cousins would have no interest in Lancaster documents, so limit the Cochra Clan cousins to view-only access to documents tagged "cochran clan"* ]
|
||||
|
||||
Objective: prepare for managed collaboration beyond single-user operation.
|
||||
|
||||
### Scope
|
||||
1. Introduce application-level authentication.
|
||||
2. Add role-based authorization (admin/editor/contributor/viewer).
|
||||
3. Add audit visibility for user-attributed write actions.
|
||||
|
||||
### Deliverables
|
||||
- User identity model and login/session flow.
|
||||
- Route/page/service authorization enforcement.
|
||||
- Audit metadata for sensitive create/update/delete workflows.
|
||||
|
||||
### Exit Criteria
|
||||
- Unauthorized operations are blocked consistently across UI/API.
|
||||
- Role policies are enforced by deterministic tests.
|
||||
- User-attributed changes are visible for audit/review.
|
||||
|
||||
## Deferred / Future Ideas (not committed scope)
|
||||
|
||||
Captured for later consideration, not yet scheduled to a version:
|
||||
* AI-assisted entity disambiguation (kinship co-occurrence, chronological plausibility
|
||||
filtering) when linking document mentions to people.
|
||||
* Kinship-aware `@mention` tagging while transcribing.
|
||||
* Relationship-calculator badges (e.g., "3rd Great-Grandmother") in the document viewer.
|
||||
* Interactive migration/geography mapping from GEDCOM and document place mentions.
|
||||
* AI-suggested document discovery by date/location overlap with known persons.
|
||||
* Ability to search within a document to find potential people to add to the People table.
|
||||
|
||||
## Planning Notes
|
||||
|
||||
- Keep architecture, schema, and UI contracts synchronized in `docs/` as each version lands.
|
||||
- Prefer explicit schema migration over runtime compatibility write paths.
|
||||
- Preserve evidence/provenance guarantees when adding new AI-powered features.
|
||||
- GEDCOM/FamilySearch data is external, collaborative, and mutable; treat it as a managed
|
||||
cache bridged via `fs_id`, never as a replacement for archival evidence recorded from
|
||||
transcribed documents.
|
||||
+391
@@ -0,0 +1,391 @@
|
||||
# Data Model and Persistence Schema (Current Baseline: V6.1)
|
||||
|
||||
This document is the field-accurate V6.1 schema contract aligned to `src/transcription/db/models.py`.
|
||||
|
||||
## Source of Truth Anchors
|
||||
|
||||
- `src/transcription/db/models.py` (status and purpose enums, including maintenance lifecycle enums)
|
||||
- `src/transcription/db/models.py:80-120` (`DocumentType`, `PersonRole`)
|
||||
- `src/transcription/db/models.py:122-172` (`Tag`, `Document`)
|
||||
- `src/transcription/db/models.py` (`Person`, `GenealogyPerson`, `GenealogyFamily`, `GenealogyFamilyChild`, `GenealogyCitation`)
|
||||
- `src/transcription/db/models.py` (`Photo`, `DocumentPerson`, `DocumentTag`)
|
||||
- `src/transcription/db/models.py:285-347` (`Job`)
|
||||
- `src/transcription/db/models.py` (`MaintenanceRun`)
|
||||
- `src/transcription/db/models.py:350-462` (`Source`, `JobSource`)
|
||||
- `src/transcription/db/models.py:465-522` (`ExecutionAttempt`)
|
||||
|
||||
## Entity Relationship Overview
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
DocumentType ||--o{ Document : classifies
|
||||
Document ||--o{ Job : has
|
||||
Document ||--o{ Source : has
|
||||
Document ||--o{ DocumentPerson : links
|
||||
Document ||--o{ DocumentTag : tagged
|
||||
Person ||--o{ DocumentPerson : links
|
||||
Person ||--o{ PersonTag : tagged
|
||||
Person ||--o{ Photo : owns
|
||||
GenealogyPerson ||--o{ GenealogyFamily : husband
|
||||
GenealogyPerson ||--o{ GenealogyFamily : wife
|
||||
GenealogyPerson ||--o{ GenealogyFamilyChild : child
|
||||
GenealogyFamily ||--o{ GenealogyFamilyChild : includes
|
||||
GenealogyPerson ||--o{ GenealogyCitation : cited
|
||||
GenealogyFamily ||--o{ GenealogyCitation : cited
|
||||
Document ||--o{ GenealogyCitation : evidence
|
||||
PersonRole ||--o{ DocumentPerson : labels
|
||||
Tag ||--o{ DocumentTag : labels
|
||||
Tag ||--o{ PersonTag : labels
|
||||
Job ||--o{ JobSource : includes
|
||||
Source ||--o{ JobSource : participates
|
||||
JobSource ||--o{ ExecutionAttempt : attempts
|
||||
MaintenanceRun {
|
||||
uuid id PK
|
||||
}
|
||||
```
|
||||
|
||||
## Authoritative Enumerations
|
||||
|
||||
### JobStatus
|
||||
|
||||
- `queued`
|
||||
- `processing`
|
||||
- `transcribed`
|
||||
- `partial_success`
|
||||
- `failed`
|
||||
|
||||
### JobSourceStatus
|
||||
|
||||
- `pending`
|
||||
- `transcribed`
|
||||
- `failed`
|
||||
- `cancelled`
|
||||
|
||||
### JobPurpose
|
||||
|
||||
- `transcription`
|
||||
- `retranscription`
|
||||
|
||||
### MaintenanceJobType
|
||||
|
||||
- `backup`
|
||||
- `storage_reconciliation`
|
||||
- `gedcom_import`
|
||||
|
||||
### MaintenanceRunStatus
|
||||
|
||||
- `queued`
|
||||
- `processing`
|
||||
- `succeeded`
|
||||
- `failed`
|
||||
|
||||
## Field-Accurate Table Contracts
|
||||
|
||||
### `DocumentType`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `semantic_key` | `str \| None` | nullable unique, indexed |
|
||||
| `label` | `str` | required |
|
||||
| `normalized_label` | `str` | unique, indexed |
|
||||
| `is_active` | `bool` | default `True` |
|
||||
| `created_at` | `datetime` | default now |
|
||||
| `updated_at` | `datetime` | default now, onupdate |
|
||||
|
||||
### `PersonRole`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `semantic_key` | `str \| None` | nullable unique, indexed |
|
||||
| `label` | `str` | required |
|
||||
| `normalized_label` | `str` | unique, indexed |
|
||||
| `is_active` | `bool` | default `True` |
|
||||
| `created_at` | `datetime` | default now |
|
||||
| `updated_at` | `datetime` | default now, onupdate |
|
||||
|
||||
### `Tag`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `semantic_key` | `str \| None` | nullable unique, indexed |
|
||||
| `label` | `str` | required |
|
||||
| `normalized_label` | `str` | unique, indexed |
|
||||
| `is_active` | `bool` | default `True` |
|
||||
| `created_at` | `datetime` | default now |
|
||||
| `updated_at` | `datetime` | default now, onupdate |
|
||||
|
||||
### `Document`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `name` | `str` | required |
|
||||
| `document_type_id` | `UUID \| None` | FK -> `document_type.id`, indexed |
|
||||
| `document_date` | `date \| None` | optional |
|
||||
| `document_date_raw` | `str \| None` | optional |
|
||||
| `location_created` | `str \| None` | optional |
|
||||
| `notes` | `str \| None` | optional |
|
||||
| `archive_identifier` | `str \| None` | optional |
|
||||
| `created_at` | `datetime` | default now |
|
||||
| `updated_at` | `datetime` | default now, onupdate |
|
||||
|
||||
### `Person`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `last_name` | `str` | required |
|
||||
| `given_names` | `str` | required |
|
||||
| `birth_date` | `date \| None` | optional |
|
||||
| `birth_date_raw` | `str \| None` | optional |
|
||||
| `birth_place` | `str \| None` | optional |
|
||||
| `death_date` | `date \| None` | optional |
|
||||
| `death_date_raw` | `str \| None` | optional |
|
||||
| `death_place` | `str \| None` | optional |
|
||||
| `biography` | `str \| None` | optional |
|
||||
| `family_search_id` | `str \| None` | nullable unique |
|
||||
| `metadata_` | `dict[str, JsonValue] \| None` | stored as DB column `metadata` (`JSONBCompat`) |
|
||||
| `created_at` | `datetime` | default now |
|
||||
| `updated_at` | `datetime` | default now, onupdate |
|
||||
|
||||
### `GenealogyPerson`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `fs_id` | `str` | unique, indexed FamilySearch identifier |
|
||||
| `full_name` | `str` | required |
|
||||
| `birth_date` | `date \| None` | optional |
|
||||
| `birth_date_raw` | `str \| None` | optional |
|
||||
| `birth_place` | `str \| None` | optional |
|
||||
| `death_date` | `date \| None` | optional |
|
||||
| `death_date_raw` | `str \| None` | optional |
|
||||
| `death_place` | `str \| None` | optional |
|
||||
| `created_at` | `datetime` | default now |
|
||||
| `updated_at` | `datetime` | default now, onupdate |
|
||||
|
||||
### `GenealogyFamily`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `fs_family_id` | `str` | unique, indexed FamilySearch family identifier |
|
||||
| `husband_id` | `UUID \| None` | nullable FK -> `genealogy_person.id`, indexed |
|
||||
| `wife_id` | `UUID \| None` | nullable FK -> `genealogy_person.id`, indexed |
|
||||
| `marriage_date` | `date \| None` | optional |
|
||||
| `marriage_date_raw` | `str \| None` | optional |
|
||||
| `marriage_place` | `str \| None` | optional |
|
||||
| `created_at` | `datetime` | default now |
|
||||
| `updated_at` | `datetime` | default now, onupdate |
|
||||
|
||||
### `GenealogyFamilyChild`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `family_id` | `UUID` | FK -> `genealogy_family.id`, indexed |
|
||||
| `child_id` | `UUID` | FK -> `genealogy_person.id`, indexed |
|
||||
| `relationship_type` | `str \| None` | optional |
|
||||
| `created_at` | `datetime` | default now |
|
||||
|
||||
Constraint:
|
||||
- `UniqueConstraint(family_id, child_id)` named `uq_genealogy_family_child`
|
||||
|
||||
### `GenealogyCitation`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `genealogy_person_id` | `UUID \| None` | nullable FK -> `genealogy_person.id`, indexed |
|
||||
| `genealogy_family_id` | `UUID \| None` | nullable FK -> `genealogy_family.id`, indexed |
|
||||
| `fact_type` | `GenealogyCitationFactType` | enum: `birth`, `death`, `marriage`, `other` |
|
||||
| `raw_citation_text` | `str` | required raw GEDCOM citation text |
|
||||
| `source_kind` | `GenealogyCitationSourceKind` | enum: `familysearch_imported`, `transcription_evidence` |
|
||||
| `document_id` | `UUID \| None` | nullable FK -> `document.id`, indexed |
|
||||
| `created_at` | `datetime` | default now |
|
||||
|
||||
### `Photo`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `person_id` | `UUID \| None` | nullable FK -> `person.id`, indexed (`NULL` = homepage photo) |
|
||||
| `path` | `str` | required upload-root-relative POSIX path (`photos/...`) |
|
||||
| `description` | `str \| None` | optional |
|
||||
| `is_primary` | `bool` | default `False`; owner-level "featured/primary" marker |
|
||||
| `created_at` | `datetime` | default now |
|
||||
| `updated_at` | `datetime` | default now, onupdate |
|
||||
|
||||
### `DocumentPerson`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `document_id` | `UUID` | FK -> `document.id`, indexed |
|
||||
| `person_id` | `UUID` | FK -> `person.id`, indexed |
|
||||
| `role_id` | `UUID` | FK -> `person_role.id`, indexed |
|
||||
| `created_at` | `datetime` | default now |
|
||||
| `updated_at` | `datetime` | default now, onupdate |
|
||||
|
||||
Constraint:
|
||||
- `UniqueConstraint(document_id, person_id)` named `uq_document_person`
|
||||
|
||||
### `DocumentTag`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `document_id` | `UUID` | FK -> `document.id`, indexed |
|
||||
| `tag_id` | `UUID` | FK -> `tag.id`, indexed |
|
||||
| `created_at` | `datetime` | default now |
|
||||
| `updated_at` | `datetime` | default now, onupdate |
|
||||
|
||||
Constraint:
|
||||
- `UniqueConstraint(document_id, tag_id)` named `uq_document_tag`
|
||||
|
||||
### `PersonTag`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `person_id` | `UUID` | FK -> `person.id`, indexed |
|
||||
| `tag_id` | `UUID` | FK -> `tag.id`, indexed |
|
||||
| `created_at` | `datetime` | default now |
|
||||
| `updated_at` | `datetime` | default now, onupdate |
|
||||
|
||||
Constraint:
|
||||
- `UniqueConstraint(person_id, tag_id)` named `uq_person_tag`
|
||||
|
||||
### `Job`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `document_id` | `UUID` | FK -> `document.id`, indexed |
|
||||
| `status` | `JobStatus` | non-null enum (stored as enum values) |
|
||||
| `retry_count` | `int` | default `0`, `ge=0` |
|
||||
| `purpose` | `JobPurpose` | non-null enum, default `transcription` |
|
||||
| `date_created` | `datetime` | default now |
|
||||
| `date_updated` | `datetime` | default now, onupdate |
|
||||
| `provider` | `str \| None` | optional |
|
||||
| `model` | `str \| None` | optional |
|
||||
| `prompt_name` | `str \| None` | optional |
|
||||
| `prompt_hash` | `str \| None` | optional |
|
||||
| `system_prompt` | `str \| None` | optional |
|
||||
| `user_prompt` | `str \| None` | optional |
|
||||
| `temperature` | `float \| None` | optional |
|
||||
| `top_p` | `float \| None` | optional |
|
||||
|
||||
Index:
|
||||
- `Index("ix_job_status_date_created", "status", "date_created")`
|
||||
|
||||
### `MaintenanceRun`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `job_type` | `MaintenanceJobType` | non-null enum |
|
||||
| `status` | `MaintenanceRunStatus` | non-null enum, default `queued` |
|
||||
| `started_at` | `datetime \| None` | optional |
|
||||
| `finished_at` | `datetime \| None` | optional |
|
||||
| `triggered_by` | `str \| None` | optional |
|
||||
| `summary` | `str \| None` | optional |
|
||||
| `log_path` | `str \| None` | optional, log-root-relative POSIX path |
|
||||
| `error_detail` | `str \| None` | optional internal failure detail |
|
||||
| `created_at` | `datetime` | default now |
|
||||
| `updated_at` | `datetime` | default now, onupdate |
|
||||
|
||||
### `Source`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `document_id` | `UUID` | FK -> `document.id`, indexed |
|
||||
| `page_number` | `int` | default `1`, `ge=1` |
|
||||
| `upload_name` | `str` | required |
|
||||
| `filename` | `str` | required |
|
||||
| `file_path` | `str` | required upload-root-relative POSIX path (`documents/...`) |
|
||||
| `file_hash` | `str` | required |
|
||||
| `file_size_bytes` | `int` | `BigInteger`, non-null |
|
||||
| `raw_transcription` | `str \| None` | projection field |
|
||||
| `preferred_execution_attempt_id` | `UUID \| None` | nullable FK -> `execution_attempt.id`, indexed (`use_alter`) |
|
||||
| `revised_text` | `str \| None` | optional human revision |
|
||||
| `date_uploaded` | `datetime` | default now |
|
||||
| `date_revised` | `datetime \| None` | optional |
|
||||
|
||||
### `JobSource`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `job_id` | `UUID` | FK -> `job.id`, indexed |
|
||||
| `source_id` | `UUID` | FK -> `source.id`, indexed |
|
||||
| `status` | `JobSourceStatus` | non-null enum, default `pending` |
|
||||
|
||||
Constraint:
|
||||
- `UniqueConstraint(job_id, source_id)` named `uq_job_source_job_source`
|
||||
|
||||
Runtime reconciliation:
|
||||
- Startup database operations remove retired V4.6 `job_source` evidence columns (`raw_transcription`, `ai_metadata`, `raw_api_response`, `error_detail`, `executed_at`) when present so persisted schema matches this contract.
|
||||
|
||||
### `ExecutionAttempt`
|
||||
|
||||
| Field | Type | Notes |
|
||||
| :--- | :--- | :--- |
|
||||
| `id` | `UUID` | PK |
|
||||
| `job_source_id` | `UUID` | FK -> `job_source.id`, indexed |
|
||||
| `job_id` | `UUID` | FK -> `job.id`, indexed |
|
||||
| `source_id` | `UUID` | FK -> `source.id`, indexed |
|
||||
| `attempt_number` | `int` | `ge=1` |
|
||||
| `status` | `JobSourceStatus` | non-null enum, value-stable with `JobSource.status` |
|
||||
| `provider` | `str` | required |
|
||||
| `model` | `str \| None` | optional |
|
||||
| `request_manifest` | `dict[str, JsonValue] \| None` | JSONBCompat |
|
||||
| `request_manifest_sha256` | `str \| None` | optional |
|
||||
| `request_manifest_schema_version` | `str \| None` | optional |
|
||||
| `response_received` | `bool` | default `False` |
|
||||
| `transport_status_code` | `int \| None` | optional |
|
||||
| `transport_body` | `bytes \| None` | LargeBinary |
|
||||
| `transport_content_type` | `str \| None` | optional |
|
||||
| `transport_content_encoding` | `str \| None` | optional |
|
||||
| `transport_safe_headers` | `dict[str, JsonValue] \| None` | JSONBCompat |
|
||||
| `router_request_id` | `str \| None` | optional |
|
||||
| `router_generation_id` | `str \| None` | optional |
|
||||
| `sdk_response_snapshot` | `dict[str, JsonValue] \| None` | JSONBCompat |
|
||||
| `normalized_metadata` | `dict[str, JsonValue] \| None` | JSONBCompat; may include app-namespaced `processing_timing` (`provider_call_duration_ms`, `processing_duration_ms`) |
|
||||
| `software_context` | `dict[str, JsonValue] \| None` | JSONBCompat |
|
||||
| `raw_transcription` | `str \| None` | optional |
|
||||
| `error_category` | `str \| None` | optional |
|
||||
| `error_detail` | `str \| None` | optional |
|
||||
| `failure_phase` | `str \| None` | optional |
|
||||
| `started_at` | `datetime` | required |
|
||||
| `finished_at` | `datetime` | required |
|
||||
| `duration_ms` | `int` | `ge=0` |
|
||||
| `created_at` | `datetime` | default now |
|
||||
|
||||
Constraint:
|
||||
- `UniqueConstraint(job_id, source_id, attempt_number)` named `uq_execution_attempt_number`
|
||||
|
||||
## Relationship Loading Contract
|
||||
|
||||
- Most ORM relationships are configured with `lazy="raise"`.
|
||||
- `JobSource.execution_attempts` is intentionally `lazy="noload"` with ordered attempts.
|
||||
- Service/UI read paths must explicitly eager-load required relationships before access.
|
||||
|
||||
## Persistence Invariants (Ground Truth)
|
||||
|
||||
1. `ExecutionAttempt` is append-only runtime evidence.
|
||||
2. `JobSource.status` represents queue/projection execution state and is not a full evidence container.
|
||||
3. `Source.raw_transcription` is a mutable projection and not authoritative attempt history.
|
||||
4. `Job` terminal status derives from page outcomes (`JobSource` state), not from a separate summary table.
|
||||
5. `DocumentType.semantic_key` and `PersonRole.semantic_key` are nullable-unique semantic identifiers.
|
||||
|
||||
## Cross-Reference
|
||||
|
||||
- [System Architecture](architecture.md)
|
||||
- [System Requirements](requirements.md)
|
||||
- [Error Handling Policy](error_handling.md)
|
||||
- [AI Evidence and Provenance Invariant](./invariant/ai_evidence_and_provenance.md)
|
||||
@@ -1,107 +0,0 @@
|
||||
from nicegui import ui
|
||||
|
||||
# 1. Mature Dark Mode Setup
|
||||
ui.dark_mode(True)
|
||||
|
||||
# Define a refined dark palette using expanded dictionary styling
|
||||
theme_colors = {
|
||||
'primary': '#6366f1',
|
||||
'secondary': '#8b5cf6',
|
||||
'accent': '#ec4899',
|
||||
'dark': '#0f172a',
|
||||
'dark_page': '#020617',
|
||||
'positive': '#10b981',
|
||||
'negative': '#ef4444',
|
||||
}
|
||||
|
||||
ui.colors(**theme_colors)
|
||||
|
||||
# Optional: Add custom CSS for subtle noise overlays or kinetic typography
|
||||
ui.add_css('''
|
||||
.glass-card {
|
||||
background: rgba(255, 255, 255, 0.03);
|
||||
backdrop-filter: blur(12px);
|
||||
-webkit-backdrop-filter: blur(12px);
|
||||
border: 1px solid rgba(255, 255, 255, 0.05);
|
||||
border-radius: 1.5rem;
|
||||
}
|
||||
''')
|
||||
|
||||
# 2. Bento Grid Layout
|
||||
with ui.element('div').classes('grid grid-cols-1 md:grid-cols-4 gap-6 w-full max-w-6xl mx-auto p-8'):
|
||||
|
||||
# Header spanning all columns
|
||||
with ui.element('div').classes('col-span-1 md:col-span-4 mb-4'):
|
||||
ui.label('Analytics Dashboard').classes('text-4xl font-extrabold tracking-tight text-white')
|
||||
ui.label('AI-driven insights for Q3').classes('text-lg text-slate-400 mt-1')
|
||||
|
||||
# Large Feature Card (Glassmorphism + Functional Motion)
|
||||
with ui.element('div').classes('glass-card col-span-1 md:col-span-2 p-6 transition-transform duration-300 hover:scale-[1.02]'):
|
||||
ui.icon('monitoring', size='2rem').classes('text-primary mb-4')
|
||||
ui.label('Revenue Prediction').classes('text-xl font-semibold text-slate-100')
|
||||
ui.label('$45,231.00').classes('text-5xl font-bold text-white mt-2')
|
||||
# Placeholder for an interactive EChart
|
||||
ui.echart({
|
||||
'xAxis': {
|
||||
'type': 'category',
|
||||
'data': [
|
||||
'Mon',
|
||||
'Tue',
|
||||
'Wed',
|
||||
'Thu',
|
||||
'Fri',
|
||||
],
|
||||
},
|
||||
'yAxis': {
|
||||
'type': 'value',
|
||||
},
|
||||
'series': [
|
||||
{
|
||||
'data': [
|
||||
120,
|
||||
200,
|
||||
150,
|
||||
80,
|
||||
70,
|
||||
],
|
||||
'type': 'bar',
|
||||
'itemStyle': {
|
||||
'color': '#6366f1',
|
||||
},
|
||||
},
|
||||
],
|
||||
}).classes('w-full h-48 mt-4')
|
||||
|
||||
# Smaller Metric Cards
|
||||
metric_cards = [
|
||||
{
|
||||
'title': 'Active Users',
|
||||
'value': '1,204',
|
||||
'icon': 'group',
|
||||
'color': 'text-secondary',
|
||||
},
|
||||
{
|
||||
'title': 'Server Load',
|
||||
'value': '34%',
|
||||
'icon': 'memory',
|
||||
'color': 'text-accent',
|
||||
},
|
||||
]
|
||||
|
||||
for card in metric_cards:
|
||||
with ui.element('div').classes('glass-card col-span-1 p-6 flex flex-col justify-between transition-transform duration-300 hover:-translate-y-1'):
|
||||
ui.icon(card['icon'], size='2rem').classes(card['color'])
|
||||
ui.element('div').classes('flex-grow')
|
||||
ui.label(card['value']).classes('text-4xl font-bold text-white mt-4')
|
||||
ui.label(card['title']).classes('text-sm font-medium text-slate-400 uppercase tracking-wider')
|
||||
|
||||
# AI Assistant Module (Adaptive Interface)
|
||||
with ui.element('div').classes('glass-card col-span-1 md:col-span-4 p-6 flex items-center gap-4'):
|
||||
ui.icon('smart_toy', size='2rem').classes('text-positive animate-pulse')
|
||||
with ui.element('div'):
|
||||
ui.label('Ambient AI Suggestion').classes('text-sm font-bold text-positive uppercase tracking-wider')
|
||||
ui.label('Based on current server load, scaling up instances in the EU-West region is recommended.').classes('text-slate-300')
|
||||
ui.space()
|
||||
ui.button('Apply Now', color='positive').classes('rounded-full px-6 py-2 shadow-lg shadow-positive/20')
|
||||
|
||||
ui.run(title='2026 UI Dashboard')
|
||||
@@ -1,92 +0,0 @@
|
||||
```mermaid
|
||||
block-beta
|
||||
columns 3
|
||||
|
||||
%% UI Component Column
|
||||
block:UI["UI COMPONENTS / WIREFRAME"]:1
|
||||
columns 1
|
||||
|
||||
block:HeaderUI["Header & Nav"]:1
|
||||
columns 1
|
||||
h_title["[Text] Document Name & Type"]
|
||||
h_date["[Text] Date & Origin Location"]
|
||||
end
|
||||
|
||||
block:EditorUI["Page Transcription Editor"]:1
|
||||
columns 1
|
||||
ed_img["[Image Viewer] Source Image"]
|
||||
ed_page["[Badge] Page Number"]
|
||||
ed_raw["[Read-Only] AI Raw Output"]
|
||||
ed_rev["[Textarea] Human Revised Text"]
|
||||
end
|
||||
|
||||
block:PeopleUI["Attribution Sidebar"]:1
|
||||
columns 1
|
||||
p_author["[List] Authors (Full Name)"]
|
||||
p_recip["[List] Recipients (Full Name)"]
|
||||
p_bio["[Card] Person Biography & Dates"]
|
||||
end
|
||||
|
||||
block:JobUI["AI Processing Drawer"]:1
|
||||
columns 1
|
||||
j_status["[Badge] Job Status"]
|
||||
j_model["[Text] Provider & Model"]
|
||||
j_tokens["[JSON View] AI Token Usage"]
|
||||
end
|
||||
end
|
||||
|
||||
%% Directional Mapping / Connectors
|
||||
block:FLOW["MAPPING / FLOW"]:1
|
||||
columns 1
|
||||
f1["Reads / Updates -->"]
|
||||
f2["Renders Active Page -->"]
|
||||
f3["Joins via Role -->"]
|
||||
f4["Executes & Logs -->"]
|
||||
end
|
||||
|
||||
%% Postgres Schema Column
|
||||
block:DB["POSTGRES SQL SCHEMA"]:1
|
||||
columns 1
|
||||
|
||||
block:DocTbl["Table: document"]:1
|
||||
columns 1
|
||||
d_id["id : UUID (PK)"]
|
||||
d_name["name : TEXT"]
|
||||
d_type["document_type : TEXT"]
|
||||
d_date["document_date : DATE"]
|
||||
end
|
||||
|
||||
block:SrcTbl["Table: source"]:1
|
||||
columns 1
|
||||
s_id["id : UUID (PK)"]
|
||||
s_page["page_number : INT"]
|
||||
s_path["file_path : TEXT"]
|
||||
s_raw["raw_transcription : TEXT"]
|
||||
s_rev["revised_text : TEXT"]
|
||||
end
|
||||
|
||||
block:PersonTbl["Table: person & document_person"]:1
|
||||
columns 1
|
||||
p_id["id : UUID (PK)"]
|
||||
p_name["full_name : TEXT"]
|
||||
p_role["role : 'author' | 'recipient'"]
|
||||
end
|
||||
|
||||
block:JobTbl["Table: job & job_source"]:1
|
||||
columns 1
|
||||
j_id["id : UUID (PK)"]
|
||||
j_stat["status : VARCHAR"]
|
||||
j_prov["provider / model : TEXT"]
|
||||
j_meta["ai_metadata : JSONB"]
|
||||
end
|
||||
end
|
||||
|
||||
%% Connections
|
||||
HeaderUI --> DocTbl
|
||||
ed_img --> s_path
|
||||
ed_page --> s_page
|
||||
ed_raw --> s_raw
|
||||
ed_rev --> s_rev
|
||||
PeopleUI --> PersonTbl
|
||||
JobUI --> JobTbl
|
||||
```
|
||||
@@ -1,53 +0,0 @@
|
||||
```mermaid
|
||||
flowchart LR
|
||||
subgraph UI["UI Components / Wireframe"]
|
||||
direction TB
|
||||
subgraph HeaderUI["Header & Nav"]
|
||||
h_title["[Text] Document Name & Type"]
|
||||
h_date["[Text] Date & Origin Location"]
|
||||
end
|
||||
subgraph EditorUI["Page Transcription Editor"]
|
||||
ed_img["[Image Viewer] Source Image"]
|
||||
ed_page["[Badge] Page Number"]
|
||||
ed_raw["[Read-Only] AI Raw Output"]
|
||||
ed_rev["[Textarea] Human Revised Text"]
|
||||
end
|
||||
subgraph PeopleUI["Attribution Sidebar"]
|
||||
p_author["[List] Authors / Recipients"]
|
||||
end
|
||||
subgraph JobUI["AI Processing Drawer"]
|
||||
j_status["[Badge] Job Status"]
|
||||
end
|
||||
end
|
||||
|
||||
subgraph DB["Postgres SQL Schema"]
|
||||
direction TB
|
||||
subgraph DocTbl["Table: document"]
|
||||
d_name["name : TEXT"]
|
||||
d_type["document_type : TEXT"]
|
||||
end
|
||||
subgraph SrcTbl["Table: source"]
|
||||
s_path["file_path : TEXT"]
|
||||
s_page["page_number : INT"]
|
||||
s_raw["raw_transcription : TEXT"]
|
||||
s_rev["revised_text : TEXT"]
|
||||
end
|
||||
subgraph PersonTbl["Table: person & document_person"]
|
||||
p_name["full_name : TEXT"]
|
||||
p_role["role : author | recipient"]
|
||||
end
|
||||
subgraph JobTbl["Table: job & job_source"]
|
||||
j_stat["status : VARCHAR"]
|
||||
j_meta["ai_metadata : JSONB"]
|
||||
end
|
||||
end
|
||||
|
||||
%% Mappings
|
||||
HeaderUI --> DocTbl
|
||||
ed_img --> s_path
|
||||
ed_page --> s_page
|
||||
ed_raw --> s_raw
|
||||
ed_rev --> s_rev
|
||||
PeopleUI --> PersonTbl
|
||||
JobUI --> JobTbl
|
||||
```
|
||||
+43
-63
@@ -1,80 +1,60 @@
|
||||
# UI Documentation
|
||||
# UI Behavioral Contracts
|
||||
|
||||
This folder contains UI-focused design and mapping documents that connect the database schema to user-facing workflows.
|
||||
## Purpose
|
||||
|
||||
## Document Types
|
||||
This directory defines the current user-facing behavior of the NiceGUI application. It records what each page is for, which routes and actions it exposes, what information it presents, and how success, empty, validation, and failure states behave.
|
||||
|
||||
### user-journey.md
|
||||
These documents are written for maintainers and AI contributors. They are behavioral contracts, not historical implementation notes and not substitutes for the database schema.
|
||||
|
||||
A product and UX contract for a user-facing entity.
|
||||
## Current Page Contracts
|
||||
|
||||
Use this document to describe:
|
||||
- what the user is trying to do
|
||||
- which screen or action starts the workflow
|
||||
- which fields the user sees and edits
|
||||
- validation rules
|
||||
- expected success and failure outcomes
|
||||
- where the user goes next
|
||||
- [Home](pages/home.md)
|
||||
- [Documents](pages/documents.md)
|
||||
- [People](pages/people.md)
|
||||
- [Jobs](pages/jobs.md)
|
||||
- [Sources](pages/sources.md)
|
||||
- [Settings](pages/settings.md)
|
||||
|
||||
### schema-mapping.md
|
||||
NiceGUI registers the routes shown in each contract without the `/ui` prefix. The application mounts NiceGUI under `/ui`, so `/documents` in page code is served to a browser as `/ui/documents`.
|
||||
|
||||
A field-level mapping between schema, UI, and implementation.
|
||||
## Authority Hierarchy
|
||||
|
||||
Use this document to describe:
|
||||
- the authoritative schema fields for an entity
|
||||
- which fields are shown, hidden, editable, or system-managed
|
||||
- current implementation behavior
|
||||
- intended target behavior
|
||||
- implementation gaps between current code and intended UX
|
||||
When documents disagree, use this order:
|
||||
|
||||
### acceptance-criteria.md
|
||||
1. User-facing page intent and accepted behavior: the page contracts in this directory.
|
||||
2. Visual and interaction styling: [UI Style Guide](../invariant/ui_style_guide.md).
|
||||
3. UI dependency and ownership boundaries: [UI contributor instructions](../../.github/instructions/ui.instructions.md).
|
||||
4. Durable failure behavior: [Error Handling invariant](../invariant/error_handling.md).
|
||||
5. Durable AI evidence behavior: [Digital Evidence and AI Processing Provenance](../invariant/ai_evidence_and_provenance.md).
|
||||
6. Data definitions and relationships: current models plus the [schema contract](../schema.md).
|
||||
7. Implementation truth: current code and tests.
|
||||
|
||||
An implementation-ready checklist for CRUD behavior and quality gates.
|
||||
If code intentionally changes accepted page behavior, update the corresponding page contract in the same change. If code accidentally differs, correct the implementation rather than rewriting intent to match a defect.
|
||||
|
||||
Use this document to describe:
|
||||
- testable acceptance statements by flow (Create, Read, Update, Delete)
|
||||
- success and failure behaviors
|
||||
- first-release constraints
|
||||
- cross-criteria quality gates
|
||||
## Contract Contents
|
||||
|
||||
### traceability-matrix.md
|
||||
Each page contract contains:
|
||||
|
||||
A criteria-to-code mapping that identifies implementation anchors and status.
|
||||
1. Purpose and user goals.
|
||||
2. Registered routes and navigation context.
|
||||
3. List, detail, and form behavior.
|
||||
4. Editable and system-managed information.
|
||||
5. Validation, empty, loading, and failure states.
|
||||
6. A concise acceptance checklist.
|
||||
7. Current implementation and test anchors.
|
||||
8. Known limitations and deferred work.
|
||||
|
||||
Use this document to describe:
|
||||
- acceptance criteria group to implementation file mapping
|
||||
- delivery status (implemented, partial, planned)
|
||||
- ordered implementation priorities
|
||||
## Maintenance Rules
|
||||
|
||||
## Organization Rules
|
||||
- Describe current accepted behavior in present tense.
|
||||
- Do not mix an obsolete “first release” design with current behavior.
|
||||
- Keep future changes in versioned scope documents and link to them from a Deferred Work section.
|
||||
- Do not reproduce the complete database field inventory here; include only fields that affect page behavior.
|
||||
- Keep service, file, and test anchors current.
|
||||
- Do not create separate current-state, target-state, and traceability copies of the same contract.
|
||||
- Keep cross-page visual rules in the UI Style Guide instead of repeating them on each page.
|
||||
- Keep database joins such as `DocumentPerson` and `JobSource` in schema/architecture documentation unless they directly affect a page interaction.
|
||||
|
||||
- Store documents under `docs/ui/entities/<entity-name>/`.
|
||||
- Create both `user-journey.md` and `schema-mapping.md` for user-facing entities.
|
||||
- Create `acceptance-criteria.md` for user-facing entities.
|
||||
- Create only `schema-mapping.md` for supporting tables that do not currently have standalone UI.
|
||||
- Keep one shared `traceability-matrix.md` under `docs/ui/entities/` to map criteria to implementation anchors.
|
||||
- Keep top-level `docs/` reserved for core architecture, requirements, schema, and system-wide reference material.
|
||||
## Current Baseline
|
||||
|
||||
## Current Entity Plan
|
||||
|
||||
User-facing entities:
|
||||
- `document`
|
||||
- `person`
|
||||
- `source`
|
||||
- `job`
|
||||
|
||||
Supporting entities:
|
||||
- `document-person`
|
||||
- `job-source`
|
||||
|
||||
## Relationship to Core Docs
|
||||
|
||||
These UI docs complement, but do not replace:
|
||||
- `docs/schema_v2.md`
|
||||
- `docs/requirements_v2.md`
|
||||
- `docs/architecture_v2.md`
|
||||
|
||||
When there is a conflict:
|
||||
- schema definitions come from the database model and schema docs
|
||||
- user interaction intent comes from the user-journey docs
|
||||
- implementation truth comes from code and is recorded in schema-mapping docs as current-state evidence
|
||||
These contracts describe the current V6.1 baseline.
|
||||
|
||||
@@ -1,182 +0,0 @@
|
||||
# DocumentPerson Schema-to-UI Mapping
|
||||
|
||||
Purpose: Map the DocumentPerson schema to UI-facing workflows, while separating intended target behavior from current implementation.
|
||||
|
||||
Supporting entity note: DocumentPerson does not currently have a standalone UI surface.
|
||||
|
||||
## 1. Entity Snapshot
|
||||
|
||||
- Table: document_person
|
||||
- Primary key: id (UUID)
|
||||
- Related entities: Document, Person
|
||||
- Canonical schema references:
|
||||
- src/transcription/db/models.py
|
||||
- docs/schema_v2.md
|
||||
|
||||
## 2. Mapping Rules
|
||||
|
||||
This document uses three lenses:
|
||||
1. Intended behavior: what user-facing workflows should support indirectly.
|
||||
2. Current behavior: what code supports today.
|
||||
3. Gap to target: what must change to align implementation with intended UX.
|
||||
|
||||
## 3. Field Inventory
|
||||
|
||||
| Field | DB Type | Nullable | Default/Auto Value | Intended UI Treatment | Notes |
|
||||
|---|---|---|---|---|---|
|
||||
| id | UUID | No | uuid4() | Hidden, system-managed | Primary key |
|
||||
| document_id | UUID FK | No | None | Context-managed | Selected Document context |
|
||||
| person_id | UUID FK | No | None | Context-managed | Selected Person context |
|
||||
| role | enum DocumentPersonRole | No | author | Visible in relationship context | First-release behavior may default to author |
|
||||
| created_at | datetime | No | datetime.now(UTC) | Hidden or read-only | System-managed timestamp |
|
||||
|
||||
Constraint behavior:
|
||||
1. document_id, person_id, and role are unique as a tuple.
|
||||
2. duplicate links for the same document, person, and role must be rejected.
|
||||
|
||||
## 4. CREATE Mapping
|
||||
|
||||
### 4.1 Intended Create Flow
|
||||
|
||||
Entry points are indirect through user-facing entities:
|
||||
1. Document create or update workflows may create one or more DocumentPerson links.
|
||||
2. Person relationship workflows may create DocumentPerson links.
|
||||
|
||||
| Field | Intended User Input | Required | Visible | Notes |
|
||||
|---|---|---|---|---|
|
||||
| document_id | None | Yes | No | Derived from selected Document |
|
||||
| person_id | None | Yes | No | Derived from selected Person |
|
||||
| role | Select or default | Yes | Indirectly | Defaults to author in first-release behavior |
|
||||
| created_at | None | No | No | System-generated |
|
||||
|
||||
### 4.2 Current Implementation
|
||||
|
||||
Current entry point: Document create/edit flows
|
||||
Current user action: select an existing Person from the Document author dropdown
|
||||
Current backend path: Document page submit callback -> `DocumentService.create_document_person()` or `delete_document_person()` as the author selection changes
|
||||
|
||||
| Field | Current Value at Create | Source | Visible to User | Evidence |
|
||||
|---|---|---|---|---|
|
||||
| id | Generated UUID | System | No | src/transcription/db/models.py |
|
||||
| document_id | Caller-provided | Document UI | Indirectly | src/transcription/ui/pages/documents_page.py |
|
||||
| person_id | Caller-provided | Document UI | Indirectly | src/transcription/ui/pages/documents_page.py |
|
||||
| role | Default author in current UI | Service/model default | No | src/transcription/db/models.py, src/transcription/services/documents.py |
|
||||
| created_at | Current UTC timestamp | System | No | src/transcription/db/models.py |
|
||||
|
||||
### 4.3 Gap to Target
|
||||
|
||||
To satisfy intended supporting behavior, implementation must add:
|
||||
1. explicit UI relationship controls in Document and/or Person detail flows.
|
||||
2. duplicate-link handling with clear user feedback.
|
||||
3. role-selection UX when role expansion is enabled beyond default author.
|
||||
|
||||
## 5. READ Mapping
|
||||
|
||||
### 5.1 Intended Read Behavior
|
||||
|
||||
Users should see DocumentPerson relationships indirectly in user-facing surfaces:
|
||||
1. Document detail shows linked people.
|
||||
2. Person detail shows linked documents.
|
||||
3. Relationship role is shown where relevant.
|
||||
|
||||
### 5.2 Current Implementation
|
||||
|
||||
Current read behavior is mainly service-level.
|
||||
|
||||
| Field | Current Rendering | Visible to User | Notes | Evidence |
|
||||
|---|---|---|---|---|
|
||||
| document_id/person_id link | Indirect relationship usage in workflows | Partial | Document/Person dedicated relationship surfaces are planned | docs/ui/entities/document/*, docs/ui/entities/person/* |
|
||||
| role | Not shown in current job-centric pages | No | Role expansion is deferred in user-facing workflows | docs/ui/entities/person/user-journey.md |
|
||||
| created_at | Not rendered | No | Operational metadata only | current UI pages |
|
||||
|
||||
Service read/query coverage:
|
||||
1. read_document_person() returns one link by id.
|
||||
2. list_document_people() supports filtering by document_id and person_id.
|
||||
|
||||
### 5.3 Gap to Target
|
||||
|
||||
To satisfy intended read behavior, implementation must add:
|
||||
1. linked-people and linked-documents UI sections backed by list_document_people().
|
||||
2. relationship role display where role context is required.
|
||||
|
||||
## 6. UPDATE Mapping
|
||||
|
||||
### 6.1 Intended Update Behavior
|
||||
|
||||
DocumentPerson updates are limited to relationship role or relationship-management actions.
|
||||
|
||||
Intended editable fields:
|
||||
- role (when role management is enabled)
|
||||
|
||||
Intended read-only fields:
|
||||
- id
|
||||
- document_id
|
||||
- person_id
|
||||
- created_at
|
||||
|
||||
### 6.2 Current Implementation
|
||||
|
||||
| Field | Updatable via UI | Updatable via Service | Notes |
|
||||
|---|---|---|---|
|
||||
| role | No | Yes | DocumentService.update_document_person() supports updates |
|
||||
| document_id/person_id | No | Technically yes via full-row update | Should generally be treated as immutable link identity |
|
||||
| created_at | No | Technically yes | Should remain system-managed |
|
||||
|
||||
### 6.3 Gap to Target
|
||||
|
||||
Implementation should add:
|
||||
1. explicit relationship-role edit controls when product scope enables them.
|
||||
2. safeguards against mutating link identity instead of recreating links.
|
||||
|
||||
## 7. DELETE Mapping
|
||||
|
||||
### 7.1 Intended Delete Behavior
|
||||
|
||||
Deletion of DocumentPerson should be exposed as unlink behavior in Document and Person flows.
|
||||
|
||||
Rules:
|
||||
1. unlink should remove only the selected relationship.
|
||||
2. unlink must not delete the underlying Document or Person records.
|
||||
|
||||
### 7.2 Current Implementation
|
||||
|
||||
| Action | UI Exposed | Backend Capability | Notes |
|
||||
|---|---|---|---|
|
||||
| Delete DocumentPerson link | No | Yes | DocumentService.delete_document_person() exists |
|
||||
|
||||
### 7.3 Gap to Target
|
||||
|
||||
Implementation must add:
|
||||
1. unlink controls in relationship sections.
|
||||
2. confirmation and success feedback for relationship removal.
|
||||
3. blocked-delete guidance if policy constraints are added later.
|
||||
|
||||
## 8. Hidden and System-Managed Fields
|
||||
|
||||
| Field | Category | Why Hidden or Protected |
|
||||
|---|---|---|
|
||||
| id | System-managed | Internal identifier |
|
||||
| document_id | Context-managed | Derived from selected Document |
|
||||
| person_id | Context-managed | Derived from selected Person |
|
||||
| created_at | System-managed | Audit timestamp |
|
||||
|
||||
## 9. Traceability Anchors
|
||||
|
||||
Schema and models:
|
||||
- docs/schema_v2.md
|
||||
- src/transcription/db/models.py
|
||||
|
||||
Current implementation:
|
||||
- src/transcription/services/documents.py
|
||||
- tests/services/test_v2_crud.py
|
||||
|
||||
Related user-facing workflows:
|
||||
- docs/ui/entities/document/user-journey.md
|
||||
- docs/ui/entities/person/user-journey.md
|
||||
|
||||
## 10. Coverage Summary
|
||||
|
||||
- Every DocumentPerson schema field appears in the field inventory.
|
||||
- Intended behavior is defined as supporting workflow behavior rather than standalone UI.
|
||||
- Current behavior reflects UI-backed CRUD through Document create/edit flows and Person detail rendering, with no standalone DocumentPerson UI.
|
||||
- Gaps between intended and current behavior are explicit.
|
||||
@@ -1,136 +0,0 @@
|
||||
# Document Acceptance Criteria
|
||||
|
||||
Purpose: Define implementation-ready acceptance criteria for Document Read, Update, and Delete workflows.
|
||||
|
||||
Companion documents:
|
||||
- docs/ui/entities/document/user-journey.md
|
||||
- docs/ui/entities/document/schema-mapping.md
|
||||
|
||||
## Scope
|
||||
|
||||
This checklist covers:
|
||||
1. Read flow
|
||||
2. Update flow
|
||||
3. Delete flow
|
||||
|
||||
This checklist does not cover:
|
||||
1. Source upload workflow details
|
||||
2. Job execution internals
|
||||
3. Revision editor behavior
|
||||
|
||||
## Read Acceptance Criteria
|
||||
|
||||
### RD-1 Document detail retrieval
|
||||
1. Given a valid Document id
|
||||
2. When the user opens the Document detail page
|
||||
3. Then the system displays Document metadata for that record only
|
||||
|
||||
### RD-2 Metadata visibility
|
||||
1. The page shows name, document_type, document_date, document_date_raw, location_created, notes, archive_identifier
|
||||
2. created_at and updated_at are displayed as system-managed, read-only values
|
||||
|
||||
### RD-3 Related people section
|
||||
1. Given zero linked people
|
||||
2. Then the page shows a no linked people yet empty state
|
||||
3. Given one linked person
|
||||
4. Then the page shows that linked person
|
||||
|
||||
### RD-4 Sources section empty state
|
||||
1. The page shows a Sources action for the current Document
|
||||
2. The page shows a primary + Add Source action that opens job-create flow for this Document
|
||||
3. The action routes to a document-scoped Sources view
|
||||
|
||||
### RD-5 Jobs section empty state
|
||||
1. The page shows a Jobs action for the current Document
|
||||
2. The page shows a primary + Add Job action for the current Document
|
||||
3. The action routes to a document-scoped Jobs view
|
||||
|
||||
### RD-6 Filtered navigation readiness
|
||||
1. The detail page provides links or actions that can route to document-scoped Sources and Jobs views
|
||||
2. Target views are filtered to the current Document id
|
||||
|
||||
### RD-7 Failure state
|
||||
1. Given a nonexistent Document id
|
||||
2. Then the UI shows a clear not found state without crashing
|
||||
|
||||
## Update Acceptance Criteria
|
||||
|
||||
### UP-1 Edit entry
|
||||
1. Given a loaded Document detail page
|
||||
2. When the user chooses Edit document
|
||||
3. Then editable controls are shown for allowed fields only, including the author relationship selector
|
||||
4. The author selector includes No author, existing Person options, and a Create new item option
|
||||
5. Selecting Create new item routes to Person create
|
||||
|
||||
### UP-2 Editable fields
|
||||
1. Editable: name, document_type, document_date, document_date_raw, location_created, notes, archive_identifier
|
||||
2. Not editable: id, created_at, updated_at
|
||||
3. The edit flow may also change the associated author Person link
|
||||
|
||||
### UP-3 Required validation
|
||||
1. name is required
|
||||
2. document_type is required
|
||||
3. Save is blocked with inline feedback when either required field is missing
|
||||
|
||||
### UP-4 Date handling rule
|
||||
1. document_date only is allowed
|
||||
2. document_date_raw only is allowed
|
||||
3. both fields together are allowed
|
||||
4. if both are present, document_date is treated as canonical exact date and document_date_raw is retained as descriptive context
|
||||
|
||||
### UP-5 Successful save
|
||||
1. Given valid input
|
||||
2. When the user saves
|
||||
3. Then changes persist
|
||||
4. Then success feedback is shown
|
||||
5. Then the user remains on Document detail with refreshed values
|
||||
6. Then updated_at reflects update policy
|
||||
|
||||
### UP-6 Save failure
|
||||
1. Given backend failure during save
|
||||
2. Then clear error feedback is shown
|
||||
3. Then the user-entered values remain available for retry where possible
|
||||
4. Then no false success feedback is shown
|
||||
|
||||
## Delete Acceptance Criteria
|
||||
|
||||
### DL-1 Delete entry and confirmation
|
||||
1. Given a Document detail page
|
||||
2. When the user chooses Delete document
|
||||
3. Then a confirmation dialog appears with permanent-action wording
|
||||
|
||||
### DL-2 Dependency guardrails
|
||||
1. Delete is allowed only when the Document has no related Source records and no related Job records
|
||||
2. Delete is blocked when at least one related Source or Job exists
|
||||
|
||||
### DL-3 Blocked delete behavior
|
||||
1. When blocked
|
||||
2. Then the UI explains why deletion is blocked
|
||||
3. Then the UI identifies dependency categories present: Sources, Jobs, or both
|
||||
4. Then the UI provides navigation to dependency cleanup paths
|
||||
|
||||
### DL-4 Successful delete
|
||||
1. Given no blocking dependencies
|
||||
2. When the user confirms delete
|
||||
3. Then the Document is removed
|
||||
4. Then success feedback is shown
|
||||
5. Then the user is returned to the Document list page
|
||||
|
||||
### DL-5 Delete failure
|
||||
1. Given backend failure during delete
|
||||
2. Then a clear error message is shown
|
||||
3. Then the user remains on Document detail with retry path
|
||||
|
||||
## Cross-Criteria Quality Gates
|
||||
|
||||
### QG-1 Separation of intent and implementation
|
||||
1. UX intent remains in user-journey.md
|
||||
2. Current versus target implementation mapping remains in schema-mapping.md
|
||||
|
||||
### QG-2 Traceability
|
||||
1. Each accepted behavior maps to at least one future UI action or service call path
|
||||
2. No acceptance criterion contradicts the current deferred-item policy
|
||||
|
||||
### QG-3 First-release constraints
|
||||
1. Linked person during create remains optional
|
||||
2. Recipient and multi-person expansion remain deferred
|
||||
@@ -1,231 +0,0 @@
|
||||
# Document Schema-to-UI Mapping
|
||||
|
||||
Purpose: Map the Document schema to the UI, while clearly separating intended target behavior from current implementation.
|
||||
|
||||
Companion document: user-journey.md
|
||||
Acceptance criteria: acceptance-criteria.md
|
||||
|
||||
## 1. Entity Snapshot
|
||||
|
||||
- Table: Document
|
||||
- Primary key: `id` (UUID)
|
||||
- Related entities: `Source`, `Job`, `DocumentPerson`, `Person`
|
||||
- Canonical schema references:
|
||||
- `src/transcription/db/models.py`
|
||||
- `docs/schema_v2.md`
|
||||
|
||||
## 2. Mapping Rules
|
||||
|
||||
This document uses three lenses:
|
||||
1. Intended behavior: what the UX should support.
|
||||
2. Current behavior: what the code supports today.
|
||||
3. Gap to target: what must change to align implementation with the intended UX.
|
||||
|
||||
## 3. Field Inventory
|
||||
|
||||
| Field | DB Type | Nullable | Default/Auto Value | Intended UI Treatment | Notes |
|
||||
|---|---|---|---|---|---|
|
||||
| id | UUID | No | `uuid4()` | Hidden, system-managed | Primary key |
|
||||
| name | str | No | None | Shown, editable on create and edit | Required |
|
||||
| document_type | str | Yes | None | Shown, editable on create and edit | Required by intended UX |
|
||||
| document_date | date | Yes | None | Shown, editable | Canonical exact date when present |
|
||||
| document_date_raw | str | Yes | None | Shown, editable | Approximate or unknown date text |
|
||||
| location_created | str | Yes | None | Shown, editable | Optional metadata |
|
||||
| notes | str | Yes | None | Shown, editable | Optional metadata |
|
||||
| archive_identifier | str | Yes | None | Shown, editable | Free text in first release |
|
||||
| created_at | datetime | No | `datetime.now(UTC)` | Hidden or read-only | System-managed |
|
||||
| updated_at | datetime | No | `datetime.now(UTC)` | Hidden or read-only | System-managed |
|
||||
|
||||
## 4. CREATE Mapping
|
||||
|
||||
### 4.1 Intended Create Flow
|
||||
|
||||
Entry point: Document page
|
||||
User action: Create new document
|
||||
Success destination: new Document detail page
|
||||
|
||||
| Field | Intended User Input | Required | Visible | Notes |
|
||||
|---|---|---|---|---|
|
||||
| name | Text input | Yes | Yes | Primary identifier used by the user |
|
||||
| document_type | Text input | Yes | Yes | Free text in first release |
|
||||
| document_date | Date input | No | Yes | Structured exact date |
|
||||
| document_date_raw | Text input | No | Yes | Approximate or uncertain date |
|
||||
| location_created | Text input | No | Yes | Optional |
|
||||
| notes | Text area | No | Yes | Optional |
|
||||
| archive_identifier | Text input | No | Yes | Free text |
|
||||
| created_at | None | No | No | System-generated |
|
||||
| updated_at | None | No | No | Not used during initial create |
|
||||
|
||||
Related records during intended create:
|
||||
- A related person may optionally be selected or created.
|
||||
- If present, the system creates a `DocumentPerson` link.
|
||||
- Jobs are not created during Document create.
|
||||
- Sources are not created during Document create.
|
||||
|
||||
### 4.2 Current Implementation
|
||||
|
||||
Current entry point: `/documents` page
|
||||
Current user action: open create form, fill metadata, optionally select an existing Person
|
||||
Current backend path: document page submit callback -> `DocumentService.create_document()` -> optional `DocumentService.create_document_person()`
|
||||
|
||||
| Field | Current Value at Create | Source | Visible to User | Evidence |
|
||||
|---|---|---|---|---|
|
||||
| id | Generated UUID | System | No | `Document` default factory in `src/transcription/db/models.py` |
|
||||
| name | User-provided | User input | Yes | `src/transcription/ui/pages/documents_page.py` |
|
||||
| document_type | User-provided or None | User input | Yes | `src/transcription/ui/pages/documents_page.py` |
|
||||
| document_date | Parsed from date input or None | User input | Yes | `src/transcription/ui/pages/documents_page.py` |
|
||||
| document_date_raw | User-provided or None | User input | Yes | `src/transcription/ui/pages/documents_page.py` |
|
||||
| location_created | User-provided or None | User input | Yes | `src/transcription/ui/pages/documents_page.py` |
|
||||
| notes | User-provided or None | User input | Yes | `src/transcription/ui/pages/documents_page.py` |
|
||||
| archive_identifier | User-provided or None | User input | Yes | `src/transcription/ui/pages/documents_page.py` |
|
||||
| created_at | Current UTC timestamp | System | No | Default factory in `src/transcription/db/models.py` |
|
||||
| updated_at | Current UTC timestamp | System | No | Default factory in `src/transcription/db/models.py` |
|
||||
|
||||
Current related-record behavior:
|
||||
- User may optionally select an existing `Person`.
|
||||
- If selected, `DocumentPerson` is created with role `author`.
|
||||
- `Job` is not created during Document create.
|
||||
- `Source` is not created during Document create.
|
||||
|
||||
### 4.3 Gap to Target
|
||||
|
||||
To satisfy the intended Create flow, implementation now includes:
|
||||
1. a Document page and dedicated create form
|
||||
2. user-entered metadata fields for `document_type`, `document_date`, `document_date_raw`, `location_created`, `notes`, and `archive_identifier`
|
||||
3. optional Person lookup through a dropdown of existing people
|
||||
4. optional `DocumentPerson` link creation when a person is chosen
|
||||
5. post-submit routing to a Document detail page
|
||||
|
||||
## 5. READ Mapping
|
||||
|
||||
### 5.1 Intended Read Behavior
|
||||
|
||||
On the Document detail page, the user should be able to see:
|
||||
1. Document metadata
|
||||
2. linked people
|
||||
3. a Sources section with empty-state behavior when no sources exist
|
||||
4. a Jobs section with empty-state behavior when no jobs exist
|
||||
5. filtered Jobs and Sources views for the current document
|
||||
|
||||
### 5.2 Current Implementation
|
||||
|
||||
Current Document visibility in the UI is direct.
|
||||
|
||||
| Field | Current Rendering | Visible to User | Notes | Evidence |
|
||||
|---|---|---|---|---|
|
||||
| name | Rendered as title and detail heading | Yes | Dedicated Document detail page | `src/transcription/ui/pages/documents_page.py` |
|
||||
| id | Not shown as raw id | No | Internal identifier remains hidden | `src/transcription/ui/pages/documents_page.py` |
|
||||
| document_type | Rendered | Yes | Shown on detail and editable on create/edit | `src/transcription/ui/pages/documents_page.py` |
|
||||
| document_date | Rendered | Yes | Exact date shown when present | `src/transcription/ui/pages/documents_page.py` |
|
||||
| document_date_raw | Rendered | Yes | Approximate date shown when present | `src/transcription/ui/pages/documents_page.py` |
|
||||
| location_created | Rendered | Yes | Optional metadata shown | `src/transcription/ui/pages/documents_page.py` |
|
||||
| notes | Rendered | Yes | Optional metadata shown | `src/transcription/ui/pages/documents_page.py` |
|
||||
| archive_identifier | Rendered | Yes | Optional metadata shown | `src/transcription/ui/pages/documents_page.py` |
|
||||
| created_at | Rendered read-only | Yes | System timestamp shown on detail | `src/transcription/ui/pages/documents_page.py` |
|
||||
| updated_at | Rendered read-only | Yes | System timestamp shown on detail | `src/transcription/ui/pages/documents_page.py` |
|
||||
|
||||
### 5.3 Gap to Target
|
||||
|
||||
To satisfy the intended Read flow, implementation now includes:
|
||||
1. metadata rendering for Document fields
|
||||
2. linked people rendering
|
||||
3. document-scoped Sources and Jobs navigation views
|
||||
|
||||
## 6. UPDATE Mapping
|
||||
|
||||
### 6.1 Intended Update Behavior
|
||||
|
||||
The user should eventually be able to edit Document metadata from the Document detail page or a dedicated edit flow.
|
||||
|
||||
Intended editable fields:
|
||||
- `name`
|
||||
- `document_type`
|
||||
- `document_date`
|
||||
- `document_date_raw`
|
||||
- `location_created`
|
||||
- `notes`
|
||||
- `archive_identifier`
|
||||
|
||||
Intended system-managed fields:
|
||||
- `id`
|
||||
- `created_at`
|
||||
- `updated_at`
|
||||
|
||||
### 6.2 Current Implementation
|
||||
|
||||
| Field | Updatable via UI | Updatable via Service | Notes |
|
||||
|---|---|---|---|
|
||||
| id | No | Practically no | Primary key should be treated as immutable |
|
||||
| name | Yes | Yes | Editable from dedicated document edit page via `DocumentService.update_document()` |
|
||||
| document_type | Yes | Yes | Editable from dedicated document edit page via `DocumentService.update_document()` |
|
||||
| document_date | Yes | Yes | Editable from dedicated document edit page via `DocumentService.update_document()` |
|
||||
| document_date_raw | Yes | Yes | Editable from dedicated document edit page via `DocumentService.update_document()` |
|
||||
| location_created | Yes | Yes | Editable from dedicated document edit page via `DocumentService.update_document()` |
|
||||
| notes | Yes | Yes | Editable from dedicated document edit page via `DocumentService.update_document()` |
|
||||
| archive_identifier | Yes | Yes | Editable from dedicated document edit page via `DocumentService.update_document()` |
|
||||
| created_at | No | Technically yes | Should remain system-managed |
|
||||
| updated_at | No | Technically yes | Should remain system-managed |
|
||||
|
||||
### 6.3 Gap to Target
|
||||
|
||||
Implementation now includes:
|
||||
1. Document edit controls in the UI
|
||||
2. validation and save behavior for Document metadata
|
||||
3. author relationship controls through the edit flow
|
||||
|
||||
## 7. DELETE Mapping
|
||||
|
||||
### 7.1 Intended Delete Behavior
|
||||
|
||||
The UI should eventually provide a delete action for Document with guardrails.
|
||||
|
||||
Rules:
|
||||
1. A Document can be deleted when it has no attached Jobs and no attached Sources.
|
||||
2. If dependent Jobs or Sources exist, the UI should block deletion and explain that those related records must be removed first.
|
||||
3. Delete confirmation should make it clear that the action is permanent.
|
||||
|
||||
### 7.2 Current Implementation
|
||||
|
||||
| Action | UI Exposed | Backend Capability | Notes |
|
||||
|---|---|---|---|
|
||||
| Delete Document | Yes | Yes | `DocumentService.delete_document()` exists and the UI blocks dependent deletes |
|
||||
|
||||
### 7.3 Gap to Target
|
||||
|
||||
Implementation includes:
|
||||
1. a Document delete control in the UI
|
||||
2. pre-delete dependency checks for Jobs and Sources
|
||||
3. user-facing messaging when deletion is blocked
|
||||
4. confirmation UX for successful delete attempts
|
||||
|
||||
## 8. Hidden and System-Managed Fields
|
||||
|
||||
| Field | Category | Why Hidden or Protected |
|
||||
|---|---|---|
|
||||
| id | System-managed | Internal identifier |
|
||||
| created_at | System-managed | Audit timestamp |
|
||||
| updated_at | System-managed | Audit timestamp |
|
||||
|
||||
## 9. Traceability Anchors
|
||||
|
||||
Schema and models:
|
||||
- `docs/schema_v2.md`
|
||||
- `src/transcription/db/models.py`
|
||||
|
||||
Current implementation:
|
||||
- `src/transcription/ui/pages/documents_page.py`
|
||||
- `src/transcription/services/documents.py`
|
||||
- `src/transcription/services/store.py`
|
||||
- `src/transcription/ui/pages/jobs_page.py`
|
||||
- `src/transcription/ui/components/transcript.py`
|
||||
|
||||
Companion UX spec:
|
||||
- `docs/ui/entities/document/user-journey.md`
|
||||
|
||||
## 10. Acceptance Checklist Summary
|
||||
|
||||
- Every Document schema field appears in the field inventory.
|
||||
- Intended Create behavior matches the companion user journey.
|
||||
- Current Create behavior reflects the existing upload-driven implementation.
|
||||
- Gaps between intended and current behavior are explicit.
|
||||
- Read, Update, and Delete sections distinguish target behavior from current code.
|
||||
@@ -1,427 +0,0 @@
|
||||
# Document User Journey
|
||||
|
||||
Purpose: Define how a user should interact with the UI to create and manage a Document record, including expected inputs, validation, results, and related record creation.
|
||||
|
||||
Scope: This document describes intended user interaction for the Document UI. It is the UX contract for the Document entity.
|
||||
|
||||
Companion schema mapping: schema-mapping.md
|
||||
Companion acceptance criteria: acceptance-criteria.md
|
||||
|
||||
## 1. Overview
|
||||
|
||||
A Document represents a real historical artifact the user wants to describe, organize, and eventually transcribe. The user should be able to create a Document before uploading or linking any source files.
|
||||
|
||||
Creating a Document is a metadata-first workflow:
|
||||
1. The user opens the Document page.
|
||||
2. The user selects Create new document.
|
||||
3. The user enters descriptive metadata about the document.
|
||||
4. The user optionally selects one related person from the existing Person list.
|
||||
5. The system creates the Document.
|
||||
6. If a person was selected, the system links that Person to the Document through DocumentPerson with author role.
|
||||
7. The user sees a success state and lands on the new Document detail page.
|
||||
|
||||
## 2. User Goal
|
||||
|
||||
The user wants to create a new Document record that:
|
||||
1. Has enough metadata to identify the historical artifact.
|
||||
2. Can optionally be linked to a person.
|
||||
3. Exists independently of transcription jobs and source uploads.
|
||||
4. Is ready for later steps such as adding sources, starting jobs, and reviewing transcriptions.
|
||||
|
||||
## 3. Page Model
|
||||
|
||||
### 3.1 Document Page
|
||||
|
||||
The Document page is the general UI surface where users manage documents.
|
||||
|
||||
It should support:
|
||||
1. listing or locating existing documents
|
||||
2. starting the Create new document flow
|
||||
3. navigating into a specific Document after it exists
|
||||
|
||||
### 3.2 Document Detail Page
|
||||
|
||||
The Document detail page is the page for one specific Document after it has been created.
|
||||
|
||||
It should show:
|
||||
1. the Document metadata
|
||||
2. related people linked to the Document
|
||||
3. a linked-author summary when available
|
||||
4. document-scoped navigation links for Sources and Jobs
|
||||
5. filtered views for sources and jobs linked to the current document
|
||||
6. primary actions + Add Source and + Add Job
|
||||
|
||||
## 4. Entry Point
|
||||
|
||||
Entry point: Document page
|
||||
|
||||
Primary action: Create new document
|
||||
|
||||
Expected UI affordance:
|
||||
1. A visible button, link, or primary action labeled Create new document.
|
||||
2. Activation opens a dedicated form view, modal, or detail panel for creating a Document.
|
||||
|
||||
Preferred first implementation:
|
||||
1. A dedicated Document create page or panel.
|
||||
2. A simple form with explicit labels.
|
||||
3. Existing Person records should be selectable through a dropdown.
|
||||
4. Text inputs are acceptable for the remaining fields in first release.
|
||||
|
||||
## 5. Create Document Form
|
||||
|
||||
The Create Document form should contain the following fields.
|
||||
|
||||
### 5.1 Required Fields
|
||||
|
||||
| UI Label | Schema Field | Input Type | Required | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Document name | name | Text input | Yes | Examples: Pioneer Days, Letter from Zenna to Omie |
|
||||
| Document type | document_type | Text input | Yes | Examples: book, letter, enlistment papers, military record, other |
|
||||
|
||||
### 5.2 Date Fields
|
||||
|
||||
| UI Label | Schema Field | Input Type | Required | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Exact date | document_date | Date input | No | Use when the exact date is known |
|
||||
| Approximate date | document_date_raw | Text input | No | Use when exact date is uncertain, approximate, or unknown |
|
||||
|
||||
Date handling rule:
|
||||
1. The form may allow both fields to be entered.
|
||||
2. If both fields are entered, `document_date` is the canonical structured date.
|
||||
3. `document_date_raw` may still be retained as the user-entered descriptive form.
|
||||
4. The UI should explain the distinction clearly.
|
||||
|
||||
Examples:
|
||||
1. Exact date: `07/13/1885`
|
||||
2. Approximate date: `c. 1885`
|
||||
3. Approximate date: `Fall 1925`
|
||||
4. Approximate date: `unknown`
|
||||
|
||||
### 5.3 Optional Metadata Fields
|
||||
|
||||
| UI Label | Schema Field | Input Type | Required | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Document location | location_created | Text input | No | Where the document was created |
|
||||
| Notes | notes | Multiline text area | No | Freeform notes about the document |
|
||||
| Archive identifier | archive_identifier | Text input | No | Free text for now; may represent inventory code, storage reference, or repository note |
|
||||
|
||||
Archive identifier guidance:
|
||||
1. First implementation should treat this as free text.
|
||||
2. Helper text may explain that this can store a repository code, box or folder reference, or storage note.
|
||||
|
||||
### 5.4 System Fields
|
||||
|
||||
| Schema Field | User Editable | Notes |
|
||||
|---|---|---|
|
||||
| created_at | No | System-generated at creation time |
|
||||
| updated_at | No | Not user-entered during creation |
|
||||
|
||||
### 5.5 Optional Related Person
|
||||
|
||||
The Create Document flow may optionally link one related person during first release.
|
||||
|
||||
| UI Label | Schema Area | Input Type | Required | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Related person | Person -> DocumentPerson | Dropdown select | No | Selects an existing Person and links as author when saved |
|
||||
|
||||
First release behavior:
|
||||
1. The user may save a Document without linking any person.
|
||||
2. If a person is linked during create, only one person is supported in first release.
|
||||
3. The selected person is linked as author.
|
||||
4. Additional people and recipient workflows are deferred to a future revision.
|
||||
|
||||
### 5.6 Related Records Not Created Directly Here
|
||||
|
||||
| Related Area | Included in Document Create | Notes |
|
||||
|---|---|---|
|
||||
| Jobs | No | Jobs are created later when transcription work begins |
|
||||
| Sources | No | Sources are added later as uploaded pages or files |
|
||||
|
||||
## 6. Related Person Workflow
|
||||
|
||||
### 6.1 User Intent
|
||||
|
||||
The user should be able to:
|
||||
1. select an existing Person to associate with the Document
|
||||
2. change the associated Person from the Document edit flow
|
||||
3. save the Document even if no person is linked
|
||||
|
||||
### 6.2 Data Model Interpretation
|
||||
|
||||
Person selection source:
|
||||
1. The UI should select from Person records.
|
||||
2. If a person is linked, the system should create a DocumentPerson record.
|
||||
3. Role handling for non-author document relationships is deferred.
|
||||
4. If the first release needs a persisted role immediately, the role can default to `author` until the relationship model is broadened.
|
||||
|
||||
This means:
|
||||
1. The user does not choose from DocumentPerson records.
|
||||
2. DocumentPerson is the relationship created after the Person is chosen or created.
|
||||
|
||||
### 6.3 Related Person UI Behavior
|
||||
|
||||
Minimum acceptable first implementation:
|
||||
1. Dropdown of existing Person records.
|
||||
2. Clear display of the selected related person before submit.
|
||||
3. Ability to change or clear the selected person in the Document edit flow.
|
||||
4. A Create new item option in the author selector that routes to Person create.
|
||||
5. A visible Create new person link near the selector.
|
||||
|
||||
If the person does not exist:
|
||||
1. The user can use Create new item from the author selector and continue from Person create.
|
||||
2. The Document create flow links existing Person records after selection.
|
||||
|
||||
## 7. Validation Rules
|
||||
|
||||
### 7.1 Required Field Validation
|
||||
|
||||
The form must reject submission if:
|
||||
1. `name` is empty
|
||||
2. `document_type` is empty
|
||||
|
||||
### 7.2 Date Validation
|
||||
|
||||
The form should allow:
|
||||
1. `document_date` only
|
||||
2. `document_date_raw` only
|
||||
3. both `document_date` and `document_date_raw`
|
||||
4. neither date field
|
||||
|
||||
If both are present:
|
||||
1. `document_date` is treated as the canonical exact date
|
||||
2. `document_date_raw` is retained as descriptive context
|
||||
|
||||
### 7.3 Related Person Validation
|
||||
|
||||
The form must not require a linked person in first release.
|
||||
|
||||
If a related person is selected or created:
|
||||
1. the selected value must resolve to a valid Person record before final save
|
||||
2. the DocumentPerson link must not be partially persisted on failure
|
||||
|
||||
## 8. Submission Behavior
|
||||
|
||||
When the user submits the form, the system should perform these logical steps:
|
||||
1. validate form inputs
|
||||
2. create the Document record
|
||||
3. create one DocumentPerson record only if an existing related person was selected
|
||||
4. persist intended records successfully before reporting success to the user
|
||||
|
||||
Expected write sequence:
|
||||
1. insert Document
|
||||
2. insert DocumentPerson link only if a person is linked
|
||||
|
||||
Recommended transactional behavior:
|
||||
1. Document and optional DocumentPerson writes should succeed or fail together
|
||||
2. Person creation is a separate workflow reached from the author selector and is not part of the same transaction
|
||||
|
||||
## 9. Expected Result After Success
|
||||
|
||||
After successful creation, the user should expect to see:
|
||||
1. confirmation that the Document was created successfully
|
||||
2. the Document name displayed in the resulting UI state
|
||||
3. the Document metadata displayed on the new Document detail page
|
||||
4. any linked person displayed in the resulting UI state
|
||||
5. a Sources section showing an empty state when no sources exist yet
|
||||
6. a Jobs section showing an empty state when no jobs exist yet
|
||||
7. a clear next step, such as adding source files
|
||||
|
||||
Recommended success route:
|
||||
1. navigate to the new Document detail page
|
||||
2. show Document summary metadata
|
||||
3. show linked people section
|
||||
4. show empty-state placeholders for Sources and Jobs
|
||||
|
||||
## 10. Expected Result After Failure
|
||||
|
||||
If submission fails, the user should expect:
|
||||
1. clear error messaging
|
||||
2. field-level validation feedback where applicable
|
||||
3. no false success message
|
||||
4. preservation of entered form values when possible
|
||||
|
||||
Examples:
|
||||
1. missing required name
|
||||
2. missing required document type
|
||||
3. failed person creation
|
||||
4. failed DocumentPerson link creation
|
||||
5. database or server error
|
||||
|
||||
## 11. Read Document Journey
|
||||
|
||||
### 11.1 User Intent
|
||||
|
||||
The user wants to open a specific Document and quickly understand:
|
||||
1. what the document is
|
||||
2. which people are linked to it
|
||||
3. whether sources exist
|
||||
4. whether jobs exist
|
||||
5. what the next action should be
|
||||
|
||||
### 11.2 Entry Points
|
||||
|
||||
A user can reach a Document detail page by:
|
||||
1. selecting a document from the Document page list
|
||||
2. being redirected after successfully creating a new document
|
||||
3. following a direct link to a known Document record
|
||||
|
||||
### 11.3 Document Detail Layout
|
||||
|
||||
The Document detail page should include:
|
||||
1. a header area with document name, document type, and key date values
|
||||
2. a metadata section with location_created, notes, and archive_identifier
|
||||
3. System metadata where created_at and updated_at are shown as read-only values
|
||||
4. a related people section
|
||||
5. a Sources section
|
||||
6. a Jobs section
|
||||
|
||||
The Document detail page should support:
|
||||
1. empty-state messaging when no related records exist
|
||||
2. clear next actions from each empty state
|
||||
3. filtered Sources and Jobs views scoped to the current document
|
||||
|
||||
### 11.4 Read Empty States
|
||||
|
||||
If no related records exist:
|
||||
1. People section says no linked people yet
|
||||
2. Sources section says no sources added yet
|
||||
3. Jobs section says no jobs created yet
|
||||
4. each section presents one clear next action
|
||||
|
||||
### 11.5 Read Success Criteria
|
||||
|
||||
A successful Read experience means:
|
||||
1. The user can identify the Document immediately
|
||||
2. The user can see whether work has started
|
||||
3. The user can navigate directly to document-scoped Jobs and Sources workflows
|
||||
|
||||
## 12. Update Document Journey
|
||||
|
||||
### 12.1 User Intent
|
||||
|
||||
The user wants to correct or enrich metadata after creation without touching jobs or source transcriptions directly.
|
||||
|
||||
### 12.2 Update Entry Point
|
||||
|
||||
From the Document detail page:
|
||||
1. The user selects Edit document
|
||||
2. UI opens edit mode or a dedicated edit view
|
||||
|
||||
### 12.3 Editable Fields
|
||||
|
||||
First release editable fields:
|
||||
1. name
|
||||
2. document_type
|
||||
3. document_date
|
||||
4. document_date_raw
|
||||
5. location_created
|
||||
6. notes
|
||||
7. archive_identifier
|
||||
|
||||
Read-only or system-managed fields:
|
||||
1. id
|
||||
2. created_at
|
||||
3. updated_at
|
||||
|
||||
### 12.4 Update Validation Rules
|
||||
|
||||
1. name remains required
|
||||
2. document_type remains required
|
||||
3. document_date and document_date_raw may both be present
|
||||
4. if both date fields are present, document_date remains canonical
|
||||
5. validation errors should be shown inline and block save
|
||||
|
||||
### 12.5 Update Save Behavior
|
||||
|
||||
On save:
|
||||
1. system validates form data
|
||||
2. system persists Document updates
|
||||
3. updated_at is refreshed by system policy
|
||||
4. UI shows a confirmation message
|
||||
5. user remains on Document detail page with refreshed values
|
||||
|
||||
### 12.6 Update Failure Behavior
|
||||
|
||||
If save fails:
|
||||
1. Show a clear error message
|
||||
2. keep user edits in form where possible
|
||||
3. do not show stale success messaging
|
||||
4. Allow retry without losing context
|
||||
|
||||
## 13. Delete Document Journey
|
||||
|
||||
### 13.1 User Intent
|
||||
|
||||
The user wants to remove a Document only when it is safe and unambiguous.
|
||||
|
||||
### 13.2 Delete Entry Point
|
||||
|
||||
From the Document detail page:
|
||||
1. The user selects Delete document
|
||||
2. UI opens a confirmation dialog explaining permanence
|
||||
|
||||
### 13.3 Delete Guardrails
|
||||
|
||||
Delete is allowed only when:
|
||||
1. the Document has no related Source records
|
||||
2. the Document has no related Job records
|
||||
|
||||
Delete is blocked when:
|
||||
1. any Source exists for the Document
|
||||
2. any Job exists for the Document
|
||||
|
||||
### 13.4 Blocked Delete UX
|
||||
|
||||
When blocked:
|
||||
1. Show an explicit reason that related Jobs or Sources exist
|
||||
2. Show which dependency types are present
|
||||
3. provide links to filtered Sources and Jobs for cleanup
|
||||
4. keep the Document unchanged
|
||||
|
||||
### 13.5 Allowed Delete UX
|
||||
|
||||
When allowed:
|
||||
1. Show final confirmation with document name
|
||||
2. perform delete
|
||||
3. show success confirmation
|
||||
4. return user to Document page list
|
||||
|
||||
### 13.6 Delete Failure Behavior
|
||||
|
||||
If delete fails due to system error:
|
||||
1. Show a clear error message
|
||||
2. keep user on Document detail page
|
||||
3. preserve ability to retry
|
||||
|
||||
## 14. Non-Goals for This Flow
|
||||
|
||||
The Document journey does not define:
|
||||
1. Source upload field-level UX
|
||||
2. Job execution internals
|
||||
3. revision editor behavior for transcriptions
|
||||
4. multi-person recipient workflows in first release
|
||||
|
||||
## 15. Relationship to Other Workflows
|
||||
|
||||
This Document workflow integrates with:
|
||||
1. Sources workflow for adding pages or files to the document
|
||||
2. Jobs workflow for transcription execution
|
||||
3. Person workflow for future expansion beyond one optional linked person
|
||||
|
||||
## 16. Relationship to Schema Mapping
|
||||
|
||||
This document is the intended UX contract.
|
||||
|
||||
The companion schema-mapping document should answer:
|
||||
1. which schema field appears on which screen
|
||||
2. whether the field is currently implemented
|
||||
3. whether the field is hidden, editable, or system-managed
|
||||
4. what the implementation gap is between intended UX and current code
|
||||
|
||||
## 17. Deferred Items
|
||||
|
||||
These topics are intentionally deferred to future revisions:
|
||||
1. multiple linked people during create and update
|
||||
2. recipient support during create and update
|
||||
3. a broader role model for non-author document relationships
|
||||
4. filtered Jobs and Sources list navigation details
|
||||
@@ -1,206 +0,0 @@
|
||||
# JobSource Schema-to-UI Mapping
|
||||
|
||||
Purpose: Map the JobSource schema to UI-facing workflows, while separating intended target behavior from current implementation.
|
||||
|
||||
Supporting entity note: JobSource does not currently have a standalone UI surface.
|
||||
|
||||
## 1. Entity Snapshot
|
||||
|
||||
- Table: job_source
|
||||
- Primary key: id (UUID)
|
||||
- Related entities: Job, Source
|
||||
- Canonical schema references:
|
||||
- src/transcription/db/models.py
|
||||
- docs/schema_v2.md
|
||||
|
||||
## 2. Mapping Rules
|
||||
|
||||
This document uses three lenses:
|
||||
1. Intended behavior: what user-facing workflows should support indirectly.
|
||||
2. Current behavior: what code supports today.
|
||||
3. Gap to target: what must change to align implementation with intended UX.
|
||||
|
||||
## 3. Field Inventory
|
||||
|
||||
| Field | DB Type | Nullable | Default/Auto Value | Intended UI Treatment | Notes |
|
||||
|---|---|---|---|---|---|
|
||||
| id | UUID | No | uuid4() | Hidden, system-managed | Primary key |
|
||||
| job_id | UUID FK | No | None | Context-managed | Selected Job context |
|
||||
| source_id | UUID FK | No | None | Context-managed | Selected Source context |
|
||||
| status | enum JobSourceStatus | No | pending | Shown in job detail source context | Per-source execution state |
|
||||
| raw_transcription | str | Yes | None | Shown read-only in review context | Machine output per source |
|
||||
| ai_metadata | JSONB/JSON | Yes | None | Hidden or advanced diagnostics | Provider metadata |
|
||||
| raw_api_response | JSONB/JSON | Yes | None | Hidden or advanced diagnostics | Low-level provider payload |
|
||||
| error_detail | str | Yes | None | Shown when status is failed | Execution failure details |
|
||||
| executed_at | datetime | No | datetime.now(UTC) | Shown read-only | Execution timestamp |
|
||||
|
||||
## 4. CREATE Mapping
|
||||
|
||||
### 4.1 Intended Create Flow
|
||||
|
||||
JobSource creation is indirect through Job and transcription workflows:
|
||||
1. Job create flow should create a JobSource row for each uploaded source page.
|
||||
2. Processing workflow may create missing JobSource rows when persisting transcription output.
|
||||
|
||||
| Field | Intended User Input | Required | Visible | Notes |
|
||||
|---|---|---|---|---|
|
||||
| job_id | None | Yes | No | Derived from active Job |
|
||||
| source_id | None | Yes | No | Derived from created/selected Source |
|
||||
| status | None | No | Indirectly | Defaults to pending at create |
|
||||
| raw_transcription | None | No | No at create | Filled after processing |
|
||||
| ai_metadata | None | No | No | Operational metadata |
|
||||
| raw_api_response | None | No | No | Operational payload |
|
||||
| error_detail | None | No | No at create | Filled on failure |
|
||||
| executed_at | None | No | No | System-generated |
|
||||
|
||||
### 4.2 Current Implementation
|
||||
|
||||
Current entry points:
|
||||
1. upload create path adds pending JobSource link in _create_upload_records().
|
||||
2. transcription update path creates or updates JobSource row during output persistence.
|
||||
|
||||
Current backend paths:
|
||||
1. src/transcription/services/store.py -> _create_upload_records()
|
||||
2. src/transcription/services/transcription.py -> update_job_transcription()
|
||||
|
||||
| Field | Current Value at Create/Update | Source | Visible to User | Evidence |
|
||||
|---|---|---|---|---|
|
||||
| id | Generated UUID | System | No | src/transcription/db/models.py |
|
||||
| job_id | Caller or workflow derived | Service/workflow | Indirectly | store.py, transcription.py |
|
||||
| source_id | Caller or workflow derived | Service/workflow | Indirectly | store.py, transcription.py |
|
||||
| status | pending at create, transcribed or failed on update | Workflow logic | Partial | transcription.py |
|
||||
| raw_transcription | Set on successful transcription update | Workflow/provider result | Yes in review context | transcription.py, jobs UI |
|
||||
| ai_metadata | Available in model; not currently filled in update path | Workflow potential | No | models.py, transcription.py |
|
||||
| raw_api_response | Available in model; not currently filled in update path | Workflow potential | No | models.py, transcription.py |
|
||||
| error_detail | Set on failed transcription update | Workflow/provider error | Partial | transcription.py |
|
||||
| executed_at | Set at row creation and refreshed on updates | System/workflow | Partial | models.py, transcription.py |
|
||||
|
||||
### 4.3 Gap to Target
|
||||
|
||||
To satisfy intended supporting behavior, implementation must add:
|
||||
1. explicit per-source status display for all linked sources in Job detail.
|
||||
2. clear surfaced error_detail for failed source executions.
|
||||
3. optional diagnostics surface for ai_metadata/raw_api_response when needed.
|
||||
4. first-class multi-source create path from Job create flow.
|
||||
|
||||
## 5. READ Mapping
|
||||
|
||||
### 5.1 Intended Read Behavior
|
||||
|
||||
Users should see JobSource data indirectly in job detail and review workflows:
|
||||
1. per-source execution status.
|
||||
2. per-source raw transcription output.
|
||||
3. per-source failure details where applicable.
|
||||
4. execution timestamp context.
|
||||
|
||||
### 5.2 Current Implementation
|
||||
|
||||
Current read behavior is partial and job-detail-centric.
|
||||
|
||||
| Field | Current Rendering | Visible to User | Notes | Evidence |
|
||||
|---|---|---|---|---|
|
||||
| status | Job-level status is visible; source-level status is limited | Partial | Source-level status not fully surfaced as a dedicated list | src/transcription/ui/pages/jobs_page.py |
|
||||
| raw_transcription | Original transcription card is visible | Yes | Primary source is shown in current detail flow | src/transcription/ui/components/transcript.py |
|
||||
| error_detail | Not prominently surfaced in current detail UI | Partial | Stored in JobSource rows during failures | src/transcription/services/transcription.py |
|
||||
| executed_at | Not first-class rendered | Partial | Available in model for future display | src/transcription/db/models.py |
|
||||
|
||||
Service read/query coverage:
|
||||
1. read_job_source() reads one row with source relation.
|
||||
2. list_job_sources() lists rows and supports job_id filtering.
|
||||
|
||||
### 5.3 Gap to Target
|
||||
|
||||
To satisfy intended read behavior, implementation must add:
|
||||
1. source-level execution table in Job detail.
|
||||
2. explicit failed-source messaging from error_detail.
|
||||
3. multi-source navigation in job review UI.
|
||||
|
||||
## 6. UPDATE Mapping
|
||||
|
||||
### 6.1 Intended Update Behavior
|
||||
|
||||
JobSource updates are workflow-managed, not directly user-edited.
|
||||
|
||||
Intended user-editable fields:
|
||||
- none in first-release behavior
|
||||
|
||||
Workflow-managed fields:
|
||||
- status
|
||||
- raw_transcription
|
||||
- error_detail
|
||||
- executed_at
|
||||
- optional diagnostics payload fields
|
||||
|
||||
### 6.2 Current Implementation
|
||||
|
||||
| Field | Updatable via UI | Updatable via Service/Workflow | Notes |
|
||||
|---|---|---|---|
|
||||
| status | No | Yes | Set by transcription update and job lifecycle handling |
|
||||
| raw_transcription | No | Yes | Persisted in update_job_transcription() |
|
||||
| error_detail | No | Yes | Persisted on transcription failure |
|
||||
| executed_at | No | Yes | Updated when existing JobSource rows are changed |
|
||||
| ai_metadata/raw_api_response | No | Potentially yes | Model supports them; active population is limited |
|
||||
|
||||
### 6.3 Gap to Target
|
||||
|
||||
Implementation should add:
|
||||
1. clearer job-detail visualization of per-source execution updates.
|
||||
2. optional operator diagnostics views for advanced troubleshooting.
|
||||
|
||||
## 7. DELETE Mapping
|
||||
|
||||
### 7.1 Intended Delete Behavior
|
||||
|
||||
JobSource deletion should be policy-driven and usually tied to Job/Source lifecycle operations.
|
||||
|
||||
Rules:
|
||||
1. direct user deletion is not required in first-release behavior.
|
||||
2. cleanup should occur through Job or Source deletion policies.
|
||||
|
||||
### 7.2 Current Implementation
|
||||
|
||||
| Action | UI Exposed | Backend Capability | Notes |
|
||||
|---|---|---|---|
|
||||
| Delete JobSource row | No | Yes | TranscriptionService.delete_job_source() exists |
|
||||
|
||||
### 7.3 Gap to Target
|
||||
|
||||
Implementation may add:
|
||||
1. maintenance tooling for cleanup operations.
|
||||
2. policy-aware cascade guidance in Job and Source delete flows.
|
||||
|
||||
## 8. Hidden and System-Managed Fields
|
||||
|
||||
| Field | Category | Why Hidden or Protected |
|
||||
|---|---|---|
|
||||
| id | System-managed | Internal identifier |
|
||||
| job_id | Context-managed | Derived from Job context |
|
||||
| source_id | Context-managed | Derived from Source context |
|
||||
| ai_metadata | Operational metadata | Advanced diagnostics payload |
|
||||
| raw_api_response | Operational metadata | Raw provider response payload |
|
||||
| executed_at | System-managed | Execution timestamp |
|
||||
|
||||
## 9. Traceability Anchors
|
||||
|
||||
Schema and models:
|
||||
- docs/schema_v2.md
|
||||
- src/transcription/db/models.py
|
||||
|
||||
Current implementation:
|
||||
- src/transcription/services/store.py
|
||||
- src/transcription/services/transcription.py
|
||||
- src/transcription/services/workflows.py
|
||||
- src/transcription/ui/pages/jobs_page.py
|
||||
- src/transcription/ui/components/transcript.py
|
||||
- tests/services/test_v2_crud.py
|
||||
|
||||
Related user-facing workflows:
|
||||
- docs/ui/entities/job/user-journey.md
|
||||
- docs/ui/entities/source/user-journey.md
|
||||
|
||||
## 10. Coverage Summary
|
||||
|
||||
- Every JobSource schema field appears in the field inventory.
|
||||
- Intended behavior is defined as supporting workflow behavior rather than standalone UI.
|
||||
- Current behavior reflects workflow/service-driven CRUD with partial job-detail visibility.
|
||||
- Gaps between intended and current behavior are explicit.
|
||||
@@ -1,154 +0,0 @@
|
||||
# Job Acceptance Criteria
|
||||
|
||||
Purpose: Define implementation-ready acceptance criteria for Job Create, Read, Update, and Delete workflows.
|
||||
|
||||
Companion documents:
|
||||
- docs/ui/entities/job/user-journey.md
|
||||
- docs/ui/entities/job/schema-mapping.md
|
||||
|
||||
## Scope
|
||||
|
||||
This checklist covers:
|
||||
1. Create flow
|
||||
2. Read flow
|
||||
3. Update flow
|
||||
4. Delete flow
|
||||
|
||||
This checklist does not cover:
|
||||
1. provider-specific transcription internals
|
||||
2. advanced workflow scheduling and queue orchestration controls
|
||||
3. multi-job bulk operations
|
||||
|
||||
## Create Acceptance Criteria
|
||||
|
||||
### CR-1 Job creation entry
|
||||
1. Given the user is on the Jobs page
|
||||
2. When the user selects Create job
|
||||
3. Then the user is taken to Job detail/create mode
|
||||
|
||||
### CR-2 Required create values
|
||||
1. document_id must be selected before submit
|
||||
2. at least one source file must be uploaded before submit
|
||||
3. each uploaded file creates a Source linked to the selected Document
|
||||
4. each created Source is linked to the new Job through JobSource
|
||||
|
||||
### CR-3 Source ordering behavior
|
||||
1. Given multi-file or folder upload
|
||||
2. When source records are created
|
||||
3. Then page ordering follows alphabetical order of original filenames
|
||||
4. Then helper text explains how filename conventions control ordering
|
||||
|
||||
### CR-4 Provider/model/prompt visibility
|
||||
1. provider, model, and prompt_name are visible in create flow when known
|
||||
2. provider, model, and prompt_name are visible in detail flow when known
|
||||
3. if values are unknown at create time, UI shows clear unknown or pending state without blocking submit
|
||||
|
||||
### CR-5 Successful create outcome
|
||||
1. Given valid inputs
|
||||
2. When the user submits create
|
||||
3. Then the Job record is created and linked to selected Document
|
||||
4. Then source and JobSource records are created for uploads
|
||||
5. Then job status is queued or processing based on execution timing
|
||||
6. Then the user is routed to Job detail mode
|
||||
|
||||
### CR-6 Create failure outcome
|
||||
1. Given create validation or persistence failure
|
||||
2. Then clear error feedback is shown
|
||||
3. Then no false success feedback is shown
|
||||
4. Then entered selections are preserved where possible
|
||||
5. Then retry path remains available
|
||||
|
||||
## Read Acceptance Criteria
|
||||
|
||||
### RD-1 Jobs list retrieval
|
||||
1. Given one or more jobs exist
|
||||
2. When the user opens the Jobs page
|
||||
3. Then all jobs are listed in a table or equivalent list surface
|
||||
|
||||
### RD-2 Jobs list fields
|
||||
1. Jobs list shows job id
|
||||
2. Jobs list shows status
|
||||
3. Jobs list shows created or updated timestamps
|
||||
4. Jobs list shows retry_count when available
|
||||
5. Jobs list provides navigation to Job detail for each row
|
||||
|
||||
### RD-3 Job detail retrieval
|
||||
1. Given a valid job id
|
||||
2. When the user opens Job detail
|
||||
3. Then job metadata for that record only is shown
|
||||
4. Then document-scoped navigation links for Sources and Jobs are shown
|
||||
|
||||
### RD-4 Detail execution context visibility
|
||||
1. provider, model, and prompt_name are displayed when known
|
||||
2. status lifecycle value is visible
|
||||
3. source-level transcription and revision context is available through Source detail navigation from Job detail
|
||||
|
||||
### RD-5 Missing and invalid id states
|
||||
1. Given an invalid job id format
|
||||
2. Then UI shows invalid job id state without crashing
|
||||
3. Given a valid but nonexistent job id
|
||||
4. Then UI shows job not found state without crashing
|
||||
|
||||
## Update Acceptance Criteria
|
||||
|
||||
### UP-1 Revision edit entry
|
||||
1. Given a job detail page
|
||||
2. When the user opens the page
|
||||
3. Then navigation links to job-scoped Sources are available
|
||||
4. Then source rows can open Source detail revision workflow
|
||||
|
||||
### UP-2 Revision validation
|
||||
1. revision save blocks empty trimmed text and shows warning feedback
|
||||
|
||||
### UP-3 Successful revision save
|
||||
1. Source detail save persists revised text and shows success feedback
|
||||
|
||||
### UP-4 Revision save failure
|
||||
1. Source detail save failure shows clear error feedback with retry path
|
||||
|
||||
### UP-5 Job lifecycle state update visibility
|
||||
1. status changes from queued to processing to terminal states are reflected in UI
|
||||
2. retry_count updates are reflected when retry logic runs
|
||||
3. users cannot directly edit lifecycle state fields in first release
|
||||
|
||||
## Delete Acceptance Criteria
|
||||
|
||||
### DL-1 Delete entry and confirmation
|
||||
1. Given a job detail context
|
||||
2. When the user opens job delete page
|
||||
3. Then a permanent-action confirmation is shown for non-processing jobs
|
||||
|
||||
### DL-2 Dependency guardrails
|
||||
1. Delete is blocked while job status is processing
|
||||
2. Related JobSource links are removed as part of allowed delete flow
|
||||
|
||||
### DL-3 Blocked delete behavior
|
||||
1. When blocked, the UI shows clear processing-state guidance
|
||||
2. The user is offered navigation back to job or jobs list
|
||||
|
||||
### DL-4 Successful delete
|
||||
1. Given an allowed delete
|
||||
2. When the user confirms delete
|
||||
3. Then the job is removed and success feedback is shown
|
||||
4. Then the user is returned to Jobs list
|
||||
|
||||
### DL-5 Delete failure
|
||||
1. Given backend failure during delete
|
||||
2. Then clear error feedback is shown
|
||||
3. Then the user remains in delete context with retry path
|
||||
|
||||
## Cross-Criteria Quality Gates
|
||||
|
||||
### QG-1 Separation of intent and implementation
|
||||
1. UX intent remains in user-journey.md
|
||||
2. Current versus target implementation mapping remains in schema-mapping.md
|
||||
|
||||
### QG-2 Traceability
|
||||
1. Each accepted behavior maps to at least one UI action or service path
|
||||
2. No acceptance criterion contradicts first-release deferred items
|
||||
|
||||
### QG-3 First-release constraints
|
||||
1. Jobs page remains list-all with explicit Create job action
|
||||
2. Job create requires Document selection and source upload
|
||||
3. provider/model/prompt_name are visible to users when known
|
||||
4. manual retry controls may remain deferred while status visibility is required
|
||||
@@ -1,233 +0,0 @@
|
||||
# Job Schema-to-UI Mapping
|
||||
|
||||
Purpose: Map the Job schema to the UI, while clearly separating intended target behavior from current implementation.
|
||||
|
||||
Companion document: user-journey.md
|
||||
Acceptance criteria: acceptance-criteria.md
|
||||
|
||||
## 1. Entity Snapshot
|
||||
|
||||
- Table: Job
|
||||
- Primary key: id (UUID)
|
||||
- Related entities: Document, JobSource, Source
|
||||
- Canonical schema references:
|
||||
- src/transcription/db/models.py
|
||||
- docs/schema_v2.md
|
||||
|
||||
## 2. Mapping Rules
|
||||
|
||||
This document uses three lenses:
|
||||
1. Intended behavior: what the UX should support.
|
||||
2. Current behavior: what the code supports today.
|
||||
3. Gap to target: what must change to align implementation with the intended UX.
|
||||
|
||||
## 3. Field Inventory
|
||||
|
||||
| Field | DB Type | Nullable | Default/Auto Value | Intended UI Treatment | Notes |
|
||||
|---|---|---|---|---|---|
|
||||
| id | UUID | No | uuid4() | Shown read-only in list and detail | Primary key |
|
||||
| document_id | UUID FK | No | None | Required create input via Document selection | Job belongs to one Document |
|
||||
| status | enum JobStatus | No | queued | Shown read-only as lifecycle state | System-managed transitions |
|
||||
| retry_count | int | No | 0 | Shown read-only | Operational counter |
|
||||
| date_created | datetime | No | datetime.now(UTC) | Shown read-only | System-managed timestamp |
|
||||
| date_updated | datetime | No | datetime.now(UTC) | Shown read-only | System-managed timestamp |
|
||||
| provider | str | Yes | None | Visible when known; editable if create-time options are available | Processing metadata |
|
||||
| model | str | Yes | None | Visible when known; editable if create-time options are available | Processing metadata |
|
||||
| prompt_name | str | Yes | None | Visible when known; editable if create-time options are available | Prompt metadata |
|
||||
|
||||
Related execution fields rendered in Job detail via relationships:
|
||||
- Job detail renders metadata and document links; source-level review/editing is reached through job-scoped Sources routes.
|
||||
|
||||
## 4. CREATE Mapping
|
||||
|
||||
### 4.1 Intended Create Flow
|
||||
|
||||
Entry point: Jobs page Create job action
|
||||
User action: open create mode, select Document, upload one or more source files or a folder, submit for transcription
|
||||
Success destination: Job detail page in detail mode
|
||||
|
||||
| Field | Intended User Input | Required | Visible | Notes |
|
||||
|---|---|---|---|---|
|
||||
| document_id | Select/search | Yes | Yes | Required create selection |
|
||||
| status | None | No | Yes (read-only) | Starts at queued and changes by workflow |
|
||||
| retry_count | None | No | Yes (read-only) | Starts at 0 |
|
||||
| date_created | None | No | Yes (read-only) | System-generated |
|
||||
| date_updated | None | No | Yes (read-only) | System-generated |
|
||||
| provider | Display or select | No | Yes | Visible when known during create and detail |
|
||||
| model | Display or select | No | Yes | Visible when known during create and detail |
|
||||
| prompt_name | Display or select | No | Yes | Visible when known during create and detail |
|
||||
|
||||
Create-related relationship rules:
|
||||
1. source file upload is required for create.
|
||||
2. each uploaded file creates a Source linked to the selected Document.
|
||||
3. each created Source must be linked to the new Job through JobSource.
|
||||
4. processing order for multi-file and folder uploads is alphabetical by original filename.
|
||||
|
||||
### 4.2 Current Implementation
|
||||
|
||||
Current entry point: Jobs page create flow
|
||||
Current user action: select Document and upload one or more files or a folder through a single upload widget
|
||||
Current backend path: job create submit -> create_job_for_document()
|
||||
|
||||
| Field | Current Value at Create | Source | Visible to User | Evidence |
|
||||
|---|---|---|---|---|
|
||||
| id | Generated UUID | System | Yes on jobs list/detail | src/transcription/ui/pages/jobs_page.py |
|
||||
| document_id | Selected existing Document id | User selection + service write | Indirectly | src/transcription/ui/pages/jobs_page.py, src/transcription/services/store.py |
|
||||
| status | queued | Service/model default | Yes | src/transcription/services/store.py, src/transcription/db/models.py |
|
||||
| retry_count | 0 | Model default | Yes | src/transcription/db/models.py, src/transcription/ui/pages/jobs_page.py |
|
||||
| date_created | current UTC timestamp | System | Yes | src/transcription/db/models.py, src/transcription/ui/pages/jobs_page.py |
|
||||
| date_updated | current UTC timestamp | System | Yes | src/transcription/db/models.py, src/transcription/ui/pages/jobs_page.py |
|
||||
| provider | None at create, set after transcription update | Workflow/service | Yes | src/transcription/services/workflows.py |
|
||||
| model | None at create, set after transcription update | Workflow/service | Yes | src/transcription/services/workflows.py |
|
||||
| prompt_name | None at create, set by workflow updates | Workflow/service | Yes | src/transcription/services/workflows.py |
|
||||
|
||||
Current create constraints:
|
||||
1. dedicated Create job action exists in the Jobs page.
|
||||
2. job create flow requires a Document selection.
|
||||
3. current upload path accepts one widget for files or folder selection.
|
||||
|
||||
### 4.3 Gap to Target
|
||||
|
||||
To satisfy intended Create flow, implementation must add:
|
||||
1. Jobs list Create job action that opens Job detail/create mode.
|
||||
2. explicit Document selection and source upload controls in create mode.
|
||||
3. multi-file and folder upload support in create mode.
|
||||
4. deterministic alphabetical page ordering and user guidance.
|
||||
5. explicit visibility of provider, model, and prompt_name in create/detail when known.
|
||||
|
||||
## 5. READ Mapping
|
||||
|
||||
### 5.1 Intended Read Behavior
|
||||
|
||||
On Job list/detail surfaces, users should be able to see:
|
||||
1. all jobs in one list.
|
||||
2. status and timeline context.
|
||||
3. selected Document context.
|
||||
4. source-level processing and transcription results.
|
||||
5. provider/model/prompt_name when known.
|
||||
|
||||
### 5.2 Current Implementation
|
||||
|
||||
Current read behavior exists in jobs list and jobs detail routes.
|
||||
|
||||
| Field | Current Rendering | Visible to User | Notes | Evidence |
|
||||
|---|---|---|---|---|
|
||||
| id | Jobs list row and detail header | Yes | Primary visible identifier | src/transcription/ui/pages/jobs_page.py |
|
||||
| status | Jobs list and detail | Yes | Chip styling for transcribed; text for others | src/transcription/ui/pages/jobs_page.py |
|
||||
| retry_count | Jobs list table | Yes | Included in row model | src/transcription/ui/components/table/jobs.py |
|
||||
| date_created | Jobs list table | Yes | Included in row model | src/transcription/ui/components/table/jobs.py |
|
||||
| date_updated | Jobs list table | Yes | Included in row model | src/transcription/ui/components/table/jobs.py |
|
||||
| document_id | Not rendered directly as labeled field | Partial | Document context exists by relationship but limited direct display | src/transcription/ui/pages/jobs_page.py |
|
||||
| provider/model/prompt_name | Rendered as labeled fields in Job detail | Yes | Shows pending fallback when unset | src/transcription/ui/pages/jobs_page.py |
|
||||
|
||||
Source-related read behavior:
|
||||
1. Job detail exposes Sources navigation for current job context.
|
||||
2. Source preview, transcription context, and revision editor are rendered in Source detail.
|
||||
3. invalid or missing job ids show explicit UI states.
|
||||
|
||||
### 5.3 Gap to Target
|
||||
|
||||
To satisfy intended Read flow, implementation must add:
|
||||
1. optional in-page source summaries in Job detail if future UX requires fewer navigation steps.
|
||||
2. richer filtering/search UX if needed.
|
||||
|
||||
## 6. UPDATE Mapping
|
||||
|
||||
### 6.1 Intended Update Behavior
|
||||
|
||||
Primary user updates in first release are source revision edits in Source detail reached from Job detail.
|
||||
|
||||
Intended editable scope (first release):
|
||||
- Source.revised_text through Source detail review
|
||||
|
||||
Intended read-only Job fields in first release:
|
||||
- id
|
||||
- document_id after create
|
||||
- status
|
||||
- retry_count
|
||||
- date_created
|
||||
- date_updated
|
||||
|
||||
Job metadata visibility policy:
|
||||
- provider, model, and prompt_name should be visible when known.
|
||||
- create-time editing of provider/model/prompt_name is optional and depends on available options.
|
||||
|
||||
### 6.2 Current Implementation
|
||||
|
||||
| Field/Area | Updatable via UI | Updatable via Service | Notes |
|
||||
|---|---|---|---|
|
||||
| Source.revised_text from Source detail | Yes | Yes | Saved via transcription service revision path from Sources page detail route |
|
||||
| status | No | Yes | Updated by workflow lifecycle services |
|
||||
| retry_count | No | Yes | Incremented by workflow retry logic |
|
||||
| provider/model/prompt_name | No | Yes | Set during transcription result finalization |
|
||||
| document_id | No | Technically via model/service update | Treated as fixed post-create in intended UX |
|
||||
|
||||
### 6.3 Gap to Target
|
||||
|
||||
Implementation now includes:
|
||||
1. create-mode handling for provider/model/prompt visibility and optional selection.
|
||||
2. detail display for provider/model/prompt and document-scoped navigation links.
|
||||
3. source revision workflow through job-scoped Sources and Source detail pages.
|
||||
4. manual controls for retry and state transitions remain deferred.
|
||||
|
||||
## 7. DELETE Mapping
|
||||
|
||||
### 7.1 Intended Delete Behavior
|
||||
|
||||
Job deletion is implemented as a dedicated delete route with processing-state guardrails.
|
||||
|
||||
Rules:
|
||||
1. deletion is allowed only when policy allows cleanup or retention handling for related JobSource records.
|
||||
2. blocked deletion must explain constraints and required cleanup path.
|
||||
3. successful deletion requires confirmation and returns user to Jobs list.
|
||||
|
||||
### 7.2 Current Implementation
|
||||
|
||||
| Action | UI Exposed | Backend Capability | Notes |
|
||||
|---|---|---|---|
|
||||
| Delete Job | Yes | Yes | Job delete page confirms permanent action and blocks when processing |
|
||||
|
||||
### 7.3 Gap to Target
|
||||
|
||||
Implementation may add in a future revision:
|
||||
1. inline delete entry in Job detail header.
|
||||
2. richer dependency messaging beyond processing-state guardrail.
|
||||
|
||||
## 8. Hidden and System-Managed Fields
|
||||
|
||||
| Field | Category | Why Hidden or Protected |
|
||||
|---|---|---|
|
||||
| status | System-managed lifecycle | Managed by worker lifecycle transitions |
|
||||
| retry_count | System-managed operational state | Reflects retry behavior, not direct user input |
|
||||
| date_created | System-managed | Audit timestamp |
|
||||
| date_updated | System-managed | Audit timestamp |
|
||||
|
||||
## 9. Traceability Anchors
|
||||
|
||||
Schema and models:
|
||||
- docs/schema_v2.md
|
||||
- src/transcription/db/models.py
|
||||
|
||||
Current implementation:
|
||||
- src/transcription/ui/pages/jobs_page.py
|
||||
- src/transcription/ui/components/table/jobs.py
|
||||
- src/transcription/ui/pages/sources_page.py
|
||||
- src/transcription/services/jobs.py
|
||||
- src/transcription/services/workflows.py
|
||||
- src/transcription/services/store.py
|
||||
- src/transcription/services/transcription.py
|
||||
|
||||
Companion UX spec:
|
||||
- docs/ui/entities/job/user-journey.md
|
||||
|
||||
Acceptance checklist:
|
||||
- docs/ui/entities/job/acceptance-criteria.md
|
||||
|
||||
## 10. Acceptance Checklist Summary
|
||||
|
||||
- Every Job schema field appears in the field inventory.
|
||||
- Intended Create behavior matches the companion user journey.
|
||||
- Current behavior reflects explicit jobs creation plus source review/editing through dedicated Sources routes.
|
||||
- Provider/model/prompt visibility intent is explicit for create and detail views.
|
||||
- Gaps between intended and current behavior are explicit.
|
||||
- Read, Update, and Delete sections distinguish target behavior from current code.
|
||||
@@ -1,291 +0,0 @@
|
||||
# Job User Journey
|
||||
|
||||
Purpose: Define how a user should interact with the UI to create and manage a Job record, including document linking, source uploads, processing status, and page-level review.
|
||||
|
||||
Scope: This document describes intended user interaction for the Job UI. It is the UX contract for the Job entity.
|
||||
|
||||
Companion schema mapping: schema-mapping.md
|
||||
Companion acceptance criteria: acceptance-criteria.md
|
||||
|
||||
## 1. Overview
|
||||
|
||||
A Job represents one transcription run for a selected Document and one or more uploaded source files.
|
||||
|
||||
Managing a Job is run-first:
|
||||
1. The user opens the Jobs page.
|
||||
2. The user selects Create job.
|
||||
3. The user lands on a Job detail/create surface.
|
||||
4. The user links a Document and uploads one or more source files.
|
||||
5. The user submits for transcription.
|
||||
6. The system creates and processes the Job.
|
||||
7. The user reviews job metadata and follows document-scoped links for Sources and Jobs.
|
||||
|
||||
## 2. User Goal
|
||||
|
||||
The user wants to:
|
||||
1. see all jobs in one place
|
||||
2. create a new transcription run intentionally
|
||||
3. attach the run to the correct Document
|
||||
4. upload source file(s) for that run
|
||||
5. submit and monitor processing state
|
||||
6. review and revise page-level outputs
|
||||
|
||||
## 3. Page Model
|
||||
|
||||
### 3.1 Jobs List Page
|
||||
|
||||
The Jobs page is the primary UI surface where users manage jobs.
|
||||
|
||||
It should support:
|
||||
1. listing all jobs
|
||||
2. searching or filtering jobs
|
||||
3. opening job detail for any row
|
||||
4. starting Create job
|
||||
5. clear empty state when no jobs exist
|
||||
|
||||
### 3.2 Job Detail/Create Page
|
||||
|
||||
The Job detail/create page is used for both creating a new Job and viewing an existing Job.
|
||||
|
||||
Create mode should include:
|
||||
1. document selection
|
||||
2. source upload controls
|
||||
3. submit for transcription action
|
||||
|
||||
Detail mode should include:
|
||||
1. job metadata and status
|
||||
2. document-scoped navigation links for the current Document
|
||||
3. provider/model/prompt visibility when known
|
||||
4. no delete action in first release
|
||||
|
||||
## 4. Entry Points
|
||||
|
||||
Primary entry points:
|
||||
1. from Jobs page, Create job
|
||||
2. from Jobs page row selection, open existing Job detail
|
||||
|
||||
Current implementation note:
|
||||
1. current code path uses explicit /jobs/new creation
|
||||
2. intended UX is explicit Create job from the Jobs page
|
||||
3. current detail view is link-oriented and routes source review/editing through dedicated Source detail
|
||||
|
||||
## 5. Create Job Flow
|
||||
|
||||
### 5.1 User Intent
|
||||
|
||||
The user wants to start a transcription run by selecting the right Document and providing source files in one guided flow.
|
||||
|
||||
### 5.2 Create Entry
|
||||
|
||||
1. The user opens the Jobs page
|
||||
2. The user selects Create job
|
||||
3. The system opens Job detail/create page in create mode
|
||||
|
||||
### 5.3 Create Inputs
|
||||
|
||||
| UI Label | Schema Area | Input Type | Required | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Document | Job.document_id | Select/search | Yes | Links the run to one Document |
|
||||
| Source files | Source upload fields | Multi-file upload or folder upload | Yes | User may select one file, many files, or a folder |
|
||||
| Processing order | Source.page_number assignment rule | System rule | Yes | If multiple files are uploaded, order is alphabetical by original filename |
|
||||
| Provider | Job.provider | Display or select | No | Visible to user when known; selectable when options are available |
|
||||
| Model | Job.model | Display or select | No | Visible to user when known; selectable when options are available |
|
||||
| Prompt | Job.prompt_name | Display or select | No | Visible to user when known; selectable when options are available |
|
||||
|
||||
### 5.4 Source Handling Rules
|
||||
|
||||
1. Each uploaded file becomes a Source linked to the selected Document
|
||||
2. Each created Source is linked to the Job through JobSource
|
||||
3. Multi-file or folder uploads are processed alphabetically by original filename
|
||||
4. upload_name stores the original filename
|
||||
5. stored filename uses UUID plus original extension in the form UUID.extension
|
||||
|
||||
Suggested helper text:
|
||||
1. Files are processed alphabetically by original filename. Use leading numbers such as 001, 002, 003 to control page order.
|
||||
|
||||
### 5.5 Validation Rules
|
||||
|
||||
Create submission must be blocked when:
|
||||
1. no Document is selected
|
||||
2. no source file is uploaded
|
||||
|
||||
Create submission should provide clear feedback when:
|
||||
1. uploaded files are invalid or unreadable
|
||||
2. persistence fails for Job, Source, or JobSource linkage
|
||||
|
||||
### 5.6 Submission Behavior
|
||||
|
||||
On submit:
|
||||
1. validate create inputs
|
||||
2. create Job record linked to selected Document
|
||||
3. create Source records for uploaded files
|
||||
4. create JobSource links for each Source in the Job
|
||||
5. queue processing for transcription
|
||||
6. route user to Job detail mode
|
||||
|
||||
Recommended transactional behavior:
|
||||
1. intended create writes should succeed or fail together
|
||||
2. The user should not receive false success when required records fail
|
||||
|
||||
### 5.7 Create Success Result
|
||||
|
||||
After successful create:
|
||||
1. job appears in Jobs list
|
||||
2. job detail shows selected Document and created source set
|
||||
3. status appears as queued or processing based on execution timing
|
||||
4. The user can monitor progress and open page-level review
|
||||
|
||||
### 5.8 Create Failure Result
|
||||
|
||||
If create fails:
|
||||
1. Show clear error message
|
||||
2. preserve entered selections where possible
|
||||
3. keep retry path available
|
||||
4. do not show false success feedback
|
||||
|
||||
## 6. Read Job Journey
|
||||
|
||||
### 6.1 User Intent
|
||||
|
||||
The user wants to quickly understand what the job is, its current status, and which source pages need review.
|
||||
|
||||
### 6.2 Jobs List Expectations
|
||||
|
||||
The Jobs list should show, at minimum:
|
||||
1. job identifier
|
||||
2. document context
|
||||
3. current status
|
||||
4. creation or update timestamp
|
||||
5. quick action to open detail
|
||||
|
||||
Optional first-release columns if available:
|
||||
1. retry count
|
||||
2. provider/model summary
|
||||
|
||||
### 6.3 Job Detail Expectations
|
||||
|
||||
The Job detail should show:
|
||||
1. job status and summary metadata
|
||||
2. selected Document context
|
||||
3. document-scoped and job-scoped navigation links
|
||||
4. source review entry through job-scoped Sources list
|
||||
|
||||
Source detail should show:
|
||||
1. source metadata and preview
|
||||
2. original transcription output per source
|
||||
3. revision editor and latest revised content
|
||||
|
||||
### 6.4 Read Empty and Missing States
|
||||
|
||||
If no jobs exist:
|
||||
1. list shows no jobs yet empty state
|
||||
2. list shows Create job action
|
||||
|
||||
If a job id is invalid or missing:
|
||||
1. Show clear not found state
|
||||
2. do not crash the page
|
||||
|
||||
If a job has no source items due to failure:
|
||||
1. Show clear warning state
|
||||
2. keep recovery guidance visible
|
||||
|
||||
## 7. Job Status Lifecycle UX
|
||||
|
||||
### 7.1 Status Values
|
||||
|
||||
The UI should map to model-backed job states:
|
||||
1. queued
|
||||
2. processing
|
||||
3. transcribed
|
||||
4. completed
|
||||
5. partial_success
|
||||
6. failed
|
||||
|
||||
### 7.2 In-Progress States
|
||||
|
||||
When status is queued or processing:
|
||||
1. Show active progress state
|
||||
2. keep detail page refresh-safe
|
||||
3. indicate that source-level results may still be arriving
|
||||
|
||||
### 7.3 Terminal States
|
||||
|
||||
When status is completed:
|
||||
1. Show completion success state
|
||||
2. direct user to revision workflow
|
||||
|
||||
When status is partial_success:
|
||||
1. Show mixed outcome state
|
||||
2. identify failed pages
|
||||
3. guide user to review available successful pages and retry strategy
|
||||
|
||||
When status is failed:
|
||||
1. Show failure state with actionable message
|
||||
2. keep navigation and retry guidance available
|
||||
|
||||
## 8. Update Job Journey
|
||||
|
||||
### 8.1 User Intent
|
||||
|
||||
The user primarily updates job-related review outcomes by editing revised transcription text per source page.
|
||||
|
||||
### 8.2 First-Release Editable Scope
|
||||
|
||||
Editable in first release:
|
||||
1. source-level revised_text through Source detail reached from job-scoped Sources navigation
|
||||
|
||||
Read-only in first release:
|
||||
1. Job.document_id after create
|
||||
2. job status values managed by processing workflow
|
||||
3. provider/model/prompt values may be system-managed, but should remain visible in UI when known
|
||||
|
||||
### 8.3 Update Save Behavior
|
||||
|
||||
On revision save:
|
||||
1. validate revised text
|
||||
2. persist revised text for selected source
|
||||
3. update revised timestamp fields by system policy
|
||||
4. show success feedback
|
||||
|
||||
On save failure:
|
||||
1. Show clear error feedback
|
||||
2. preserve entered text where possible
|
||||
3. Allow retry
|
||||
|
||||
## 9. Delete and Retention Policy
|
||||
|
||||
### 9.1 User Intent
|
||||
|
||||
The user may need to remove invalid or duplicate jobs safely.
|
||||
|
||||
### 9.2 First-Release Policy
|
||||
|
||||
Delete behavior uses explicit guardrails:
|
||||
1. deletion is blocked while status is processing
|
||||
2. blocked delete explains constraints and offers back navigation
|
||||
3. allowed delete requires explicit confirmation and then returns to Jobs list with success feedback
|
||||
|
||||
## 10. Relationship to Other Workflows
|
||||
|
||||
Job workflow integrates with:
|
||||
1. Document workflow for ownership context
|
||||
2. Source workflow for uploaded page records and ordering
|
||||
3. Revision workflow for human correction lifecycle
|
||||
4. Worker processing workflow for queued execution and status transitions
|
||||
|
||||
## 11. Relationship to Schema Mapping
|
||||
|
||||
The companion schema-mapping document should specify:
|
||||
1. field visibility per CRUD action
|
||||
2. current implementation status
|
||||
3. intended behavior
|
||||
4. gap-to-target items
|
||||
|
||||
## 12. Deferred Items
|
||||
|
||||
Deferred to future revisions:
|
||||
1. manual retry controls from job detail
|
||||
2. advanced provider/model/prompt policy controls beyond basic create-time visibility
|
||||
3. advanced bulk actions across multiple jobs
|
||||
4. live streaming progress updates beyond refresh-based updates
|
||||
5. job templates or preset configurations
|
||||
@@ -1,154 +0,0 @@
|
||||
# Job Acceptance Criteria
|
||||
|
||||
Purpose: Define implementation-ready acceptance criteria for Job Create, Read, Update, and Delete workflows.
|
||||
|
||||
Companion documents:
|
||||
- docs/ui/entities/job/user-journey.md
|
||||
- docs/ui/entities/job/schema-mapping.md
|
||||
|
||||
## Scope
|
||||
|
||||
This checklist covers:
|
||||
1. Create flow
|
||||
2. Read flow
|
||||
3. Update flow
|
||||
4. Delete flow
|
||||
|
||||
This checklist does not cover:
|
||||
1. provider-specific transcription internals
|
||||
2. advanced workflow scheduling and queue orchestration controls
|
||||
3. multi-job bulk operations
|
||||
|
||||
## Create Acceptance Criteria
|
||||
|
||||
### CR-1 Job creation entry
|
||||
1. Given the user is on the Jobs page
|
||||
2. When the user selects Create job
|
||||
3. Then the user is taken to Job detail/create mode
|
||||
|
||||
### CR-2 Required create values
|
||||
1. document_id must be selected before submit
|
||||
2. at least one source file must be uploaded before submit
|
||||
3. each uploaded file creates a Source linked to the selected Document
|
||||
4. each created Source is linked to the new Job through JobSource
|
||||
|
||||
### CR-3 Source ordering behavior
|
||||
1. Given multi-file or folder upload
|
||||
2. When source records are created
|
||||
3. Then page ordering follows alphabetical order of original filenames
|
||||
4. Then helper text explains how filename conventions control ordering
|
||||
|
||||
### CR-4 Provider/model/prompt visibility
|
||||
1. provider, model, and prompt_name are visible in create flow when known
|
||||
2. provider, model, and prompt_name are visible in detail flow when known
|
||||
3. if values are unknown at create time, UI shows clear unknown or pending state without blocking submit
|
||||
|
||||
### CR-5 Successful create outcome
|
||||
1. Given valid inputs
|
||||
2. When the user submits create
|
||||
3. Then the Job record is created and linked to selected Document
|
||||
4. Then source and JobSource records are created for uploads
|
||||
5. Then job status is queued or processing based on execution timing
|
||||
6. Then the user is routed to Job detail mode
|
||||
|
||||
### CR-6 Create failure outcome
|
||||
1. Given create validation or persistence failure
|
||||
2. Then clear error feedback is shown
|
||||
3. Then no false success feedback is shown
|
||||
4. Then entered selections are preserved where possible
|
||||
5. Then retry path remains available
|
||||
|
||||
## Read Acceptance Criteria
|
||||
|
||||
### RD-1 Jobs list retrieval
|
||||
1. Given one or more jobs exist
|
||||
2. When the user opens the Jobs page
|
||||
3. Then all jobs are listed in a table or equivalent list surface
|
||||
|
||||
### RD-2 Jobs list fields
|
||||
1. Jobs list shows job id
|
||||
2. Jobs list shows status
|
||||
3. Jobs list shows created or updated timestamps
|
||||
4. Jobs list shows retry_count when available
|
||||
5. Jobs list provides navigation to Job detail for each row
|
||||
|
||||
### RD-3 Job detail retrieval
|
||||
1. Given a valid job id
|
||||
2. When the user opens Job detail
|
||||
3. Then job metadata for that record only is shown
|
||||
4. Then document-scoped navigation links for Sources and Jobs are shown
|
||||
|
||||
### RD-4 Detail execution context visibility
|
||||
1. provider, model, and prompt_name are displayed when known
|
||||
2. status lifecycle value is visible
|
||||
3. source-level transcription and revision context is available through Source detail navigation from Job detail
|
||||
|
||||
### RD-5 Missing and invalid id states
|
||||
1. Given an invalid job id format
|
||||
2. Then UI shows invalid job id state without crashing
|
||||
3. Given a valid but nonexistent job id
|
||||
4. Then UI shows job not found state without crashing
|
||||
|
||||
## Update Acceptance Criteria
|
||||
|
||||
### UP-1 Revision edit entry
|
||||
1. Given a job detail page
|
||||
2. When the user opens the page
|
||||
3. Then navigation links to job-scoped Sources are available
|
||||
4. Then source rows can open Source detail revision workflow
|
||||
|
||||
### UP-2 Revision validation
|
||||
1. revision save blocks empty trimmed text and shows warning feedback
|
||||
|
||||
### UP-3 Successful revision save
|
||||
1. Source detail save persists revised text and shows success feedback
|
||||
|
||||
### UP-4 Revision save failure
|
||||
1. Source detail save failure shows clear error feedback with retry path
|
||||
|
||||
### UP-5 Job lifecycle state update visibility
|
||||
1. status changes from queued to processing to terminal states are reflected in UI
|
||||
2. retry_count updates are reflected when retry logic runs
|
||||
3. users cannot directly edit lifecycle state fields in first release
|
||||
|
||||
## Delete Acceptance Criteria
|
||||
|
||||
### DL-1 Delete entry and confirmation
|
||||
1. Given a job detail context
|
||||
2. When the user opens job delete page
|
||||
3. Then a permanent-action confirmation is shown for non-processing jobs
|
||||
|
||||
### DL-2 Dependency guardrails
|
||||
1. Delete is blocked while job status is processing
|
||||
2. Related JobSource links are removed as part of allowed delete flow
|
||||
|
||||
### DL-3 Blocked delete behavior
|
||||
1. When blocked, the UI shows clear processing-state guidance
|
||||
2. The user is offered navigation back to job or jobs list
|
||||
|
||||
### DL-4 Successful delete
|
||||
1. Given an allowed delete
|
||||
2. When the user confirms delete
|
||||
3. Then the job is removed and success feedback is shown
|
||||
4. Then the user is returned to Jobs list
|
||||
|
||||
### DL-5 Delete failure
|
||||
1. Given backend failure during delete
|
||||
2. Then clear error feedback is shown
|
||||
3. Then the user remains in delete context with retry path
|
||||
|
||||
## Cross-Criteria Quality Gates
|
||||
|
||||
### QG-1 Separation of intent and implementation
|
||||
1. UX intent remains in user-journey.md
|
||||
2. Current versus target implementation mapping remains in schema-mapping.md
|
||||
|
||||
### QG-2 Traceability
|
||||
1. Each accepted behavior maps to at least one UI action or service path
|
||||
2. No acceptance criterion contradicts first-release deferred items
|
||||
|
||||
### QG-3 First-release constraints
|
||||
1. Jobs page remains list-all with explicit Create job action
|
||||
2. Job create requires Document selection and source upload
|
||||
3. provider/model/prompt_name are visible to users when known
|
||||
4. manual retry controls may remain deferred while status visibility is required
|
||||
@@ -1,147 +0,0 @@
|
||||
# Person Acceptance Criteria
|
||||
|
||||
Purpose: Define implementation-ready acceptance criteria for Person Create, Read, Update, and Delete workflows.
|
||||
|
||||
Companion documents:
|
||||
- docs/ui/entities/person/user-journey.md
|
||||
- docs/ui/entities/person/schema-mapping.md
|
||||
|
||||
## Scope
|
||||
|
||||
This checklist covers:
|
||||
1. Create flow
|
||||
2. Read flow
|
||||
3. Update flow
|
||||
4. Delete flow
|
||||
|
||||
This checklist does not cover:
|
||||
1. advanced metadata_ editing UX
|
||||
2. structured-name schema migration implementation
|
||||
3. bulk merge or dedup workflow design
|
||||
|
||||
## Create Acceptance Criteria
|
||||
|
||||
### CR-1 Person creation entry
|
||||
1. Given a Person page
|
||||
2. When the user selects Create new person
|
||||
3. Then the user can open a Person create form
|
||||
|
||||
### CR-2 Required field validation
|
||||
1. full_name is required
|
||||
2. Save is blocked when full_name is empty
|
||||
3. Inline feedback is shown for required-field errors
|
||||
|
||||
### CR-3 Optional field handling
|
||||
1. Optional fields may be blank without blocking create
|
||||
2. Date raw and exact fields can coexist
|
||||
3. Exact date remains canonical when both exact and raw are provided
|
||||
4. Portrait uploads persist under uploads/portraits/person and store a relative portrait_path
|
||||
|
||||
### CR-4 Successful create outcome
|
||||
1. Given valid input
|
||||
2. When the user saves
|
||||
3. Then the Person record is created
|
||||
4. Then success feedback is shown
|
||||
5. Then the user is routed to Person detail page
|
||||
|
||||
### CR-5 Create failure outcome
|
||||
1. Given backend failure during create
|
||||
2. Then clear error feedback is shown
|
||||
3. Then entered values are retained where possible
|
||||
4. Then no false success feedback is shown
|
||||
|
||||
## Read Acceptance Criteria
|
||||
|
||||
### RD-1 Person detail retrieval
|
||||
1. Given a valid Person id
|
||||
2. When the user opens the Person detail page
|
||||
3. Then the system displays Person metadata for that record only
|
||||
|
||||
### RD-2 Metadata visibility
|
||||
1. The page shows full_name and available optional person fields
|
||||
2. created_at and updated_at are shown as system-managed, read-only values
|
||||
3. portrait_path is rendered when available, including an image preview when possible
|
||||
4. relative portrait_path values resolve through /uploads for image rendering
|
||||
|
||||
### RD-3 Linked documents section
|
||||
1. Given no linked DocumentPerson rows
|
||||
2. Then the page shows a no linked documents yet empty state
|
||||
3. Given linked documents exist
|
||||
4. Then the page shows linked document entries
|
||||
|
||||
### RD-4 Read failure state
|
||||
1. Given a nonexistent Person id
|
||||
2. Then the UI shows a clear not found state without crashing
|
||||
|
||||
## Update Acceptance Criteria
|
||||
|
||||
### UP-1 Edit entry
|
||||
1. Given a loaded Person detail page
|
||||
2. When the user selects Edit person
|
||||
3. Then editable controls are shown for allowed fields only
|
||||
|
||||
### UP-2 Editable fields
|
||||
1. Editable: full_name, display_name, maiden_name, birth/death fields, places, biography, portrait_path
|
||||
2. Not editable: id, created_at, updated_at
|
||||
3. metadata_ remains hidden in first release
|
||||
|
||||
### UP-3 Required validation
|
||||
1. full_name remains required
|
||||
2. Save is blocked with inline feedback when full_name is empty
|
||||
|
||||
### UP-4 Successful save
|
||||
1. Given valid input
|
||||
2. When the user saves
|
||||
3. Then changes persist
|
||||
4. Then success feedback is shown
|
||||
5. Then the user remains on Person detail with refreshed values
|
||||
|
||||
### UP-5 Save failure
|
||||
1. Given backend failure during save
|
||||
2. Then clear error feedback is shown
|
||||
3. Then the user-entered values remain available for retry where possible
|
||||
4. Then no false success feedback is shown
|
||||
|
||||
## Delete Acceptance Criteria
|
||||
|
||||
### DL-1 Delete entry and confirmation
|
||||
1. Given a Person detail page
|
||||
2. When the user selects Delete person
|
||||
3. Then a confirmation dialog appears with permanent-action wording
|
||||
|
||||
### DL-2 Relationship guardrails
|
||||
1. Delete is allowed only when relationship policy allows it
|
||||
2. If linked DocumentPerson rows must be removed first, delete is blocked
|
||||
|
||||
### DL-3 Blocked delete behavior
|
||||
1. When blocked
|
||||
2. Then the UI explains why deletion is blocked
|
||||
3. Then the UI identifies linked-document dependency presence
|
||||
4. Then the UI provides navigation to cleanup paths
|
||||
|
||||
### DL-4 Successful delete
|
||||
1. Given no blocking dependencies
|
||||
2. When the user confirms delete
|
||||
3. Then the Person record is removed
|
||||
4. Then success feedback is shown
|
||||
5. Then the user returns to the Person list page
|
||||
|
||||
### DL-5 Delete failure
|
||||
1. Given backend failure during delete
|
||||
2. Then clear error feedback is shown
|
||||
3. Then the user remains on Person detail with retry path
|
||||
|
||||
## Cross-Criteria Quality Gates
|
||||
|
||||
### QG-1 Separation of intent and implementation
|
||||
1. UX intent remains in user-journey.md
|
||||
2. Current versus target implementation mapping remains in schema-mapping.md
|
||||
|
||||
### QG-2 Traceability
|
||||
1. Each accepted behavior maps to at least one future UI action or service call path
|
||||
2. No acceptance criterion contradicts the deferred-item policy
|
||||
|
||||
### QG-3 First-release constraints
|
||||
1. metadata_ remains hidden in first release
|
||||
2. structured name field split remains deferred
|
||||
3. recipient and multi-person role management stays in later revisions
|
||||
@@ -1,265 +0,0 @@
|
||||
# Person Schema-to-UI Mapping
|
||||
|
||||
Purpose: Map the Person schema to the UI, while clearly separating intended target behavior from current implementation.
|
||||
|
||||
Companion document: user-journey.md
|
||||
Acceptance criteria: acceptance-criteria.md
|
||||
|
||||
## 1. Entity Snapshot
|
||||
|
||||
- Table: Person
|
||||
- Primary key: id (UUID)
|
||||
- Related entities: DocumentPerson, Document
|
||||
- Canonical schema references:
|
||||
- src/transcription/db/models.py
|
||||
- docs/schema_v2.md
|
||||
|
||||
## 2. Mapping Rules
|
||||
|
||||
This document uses three lenses:
|
||||
1. Intended behavior: what the UX should support.
|
||||
2. Current behavior: what the code supports today.
|
||||
3. Gap to target: what must change to align implementation with the intended UX.
|
||||
|
||||
## 3. Field Inventory
|
||||
|
||||
| Field | DB Type | Nullable | Default/Auto Value | Intended UI Treatment | Notes |
|
||||
|---|---|---|---|---|---|
|
||||
| id | UUID | No | uuid4() | Hidden, system-managed | Primary key |
|
||||
| full_name | str | No | None | Shown, editable on create and update | Required canonical name |
|
||||
| display_name | str | Yes | None | Shown, editable | Optional |
|
||||
| maiden_name | str | Yes | None | Shown, editable | Optional |
|
||||
| birth_date | date | Yes | None | Shown, editable | Canonical exact date when present |
|
||||
| birth_date_raw | str | Yes | None | Shown, editable | Approximate or unknown date text |
|
||||
| birth_place | str | Yes | None | Shown, editable | Optional |
|
||||
| death_date | date | Yes | None | Shown, editable | Canonical exact date when present |
|
||||
| death_date_raw | str | Yes | None | Shown, editable | Approximate or unknown date text |
|
||||
| death_place | str | Yes | None | Shown, editable | Optional |
|
||||
| biography | str | Yes | None | Shown, editable | Optional narrative |
|
||||
| portrait_path | str | Yes | None | Shown, editable | Optional path |
|
||||
| metadata_ | JSONB/JSON | Yes | None | Hidden in first release | Advanced metadata |
|
||||
| created_at | datetime | No | datetime.now(UTC) | Hidden or read-only | System-managed |
|
||||
| updated_at | datetime | No | datetime.now(UTC) | Hidden or read-only | System-managed |
|
||||
|
||||
## 4. CREATE Mapping
|
||||
|
||||
### 4.1 Intended Create Flow
|
||||
|
||||
Entry point: Person page
|
||||
User action: Create new person
|
||||
Success destination: new Person detail page
|
||||
|
||||
| Field | Intended User Input | Required | Visible | Notes |
|
||||
|---|---|---|---|---|
|
||||
| full_name | Text input | Yes | Yes | Canonical identity field |
|
||||
| display_name | Text input | No | Yes | Optional |
|
||||
| maiden_name | Text input | No | Yes | Optional |
|
||||
| birth_date | Date input | No | Yes | Structured exact date |
|
||||
| birth_date_raw | Text input | No | Yes | Approximate/uncertain date |
|
||||
| birth_place | Text input | No | Yes | Optional |
|
||||
| death_date | Date input | No | Yes | Structured exact date |
|
||||
| death_date_raw | Text input | No | Yes | Approximate/uncertain date |
|
||||
| death_place | Text input | No | Yes | Optional |
|
||||
| biography | Text area | No | Yes | Optional |
|
||||
| portrait_path | Text input | No | Yes | Optional |
|
||||
| metadata_ | None | No | No | Hidden in first release |
|
||||
| created_at | None | No | No | System-generated |
|
||||
| updated_at | None | No | No | Not user-entered |
|
||||
|
||||
Related records during intended create:
|
||||
- No DocumentPerson link is required during Person creation.
|
||||
- Document linking can be done later from Document or Person workflows.
|
||||
|
||||
### 4.2 Current Implementation
|
||||
|
||||
Current entry point: dedicated People page and Person create/edit flows
|
||||
Current user action: open Person create page, fill form fields, optionally upload portrait
|
||||
Current backend path: People page submit callbacks -> DocumentService.create_person() / update_person()
|
||||
|
||||
| Field | Current Value at Create | Source | Visible to User | Evidence |
|
||||
|---|---|---|---|---|
|
||||
| id | Generated UUID | System | No | Person model default factory in src/transcription/db/models.py |
|
||||
| full_name | Form input | User input | Yes | src/transcription/ui/pages/people_page.py |
|
||||
| display_name | Form input or None | User input | Yes | src/transcription/ui/pages/people_page.py |
|
||||
| maiden_name | Form input or None | User input | Yes | src/transcription/ui/pages/people_page.py |
|
||||
| birth_date | Form input or None | User input | Yes | src/transcription/ui/pages/people_page.py |
|
||||
| birth_date_raw | Form input or None | User input | Yes | src/transcription/ui/pages/people_page.py |
|
||||
| birth_place | Form input or None | User input | Yes | src/transcription/ui/pages/people_page.py |
|
||||
| death_date | Form input or None | User input | Yes | src/transcription/ui/pages/people_page.py |
|
||||
| death_date_raw | Form input or None | User input | Yes | src/transcription/ui/pages/people_page.py |
|
||||
| death_place | Form input or None | User input | Yes | src/transcription/ui/pages/people_page.py |
|
||||
| biography | Form input or None | User input | Yes | src/transcription/ui/pages/people_page.py |
|
||||
| portrait_path | Relative upload path or manual path | Upload helper + user input | Yes | src/transcription/ui/pages/people_page.py, src/transcription/services/store.py |
|
||||
| metadata_ | Caller-provided or None | Service/API caller | No | Person model in src/transcription/db/models.py |
|
||||
| created_at | Current UTC timestamp | System | No | Person model default in src/transcription/db/models.py |
|
||||
| updated_at | Current UTC timestamp | System | No | Person model default in src/transcription/db/models.py |
|
||||
|
||||
### 4.3 Gap to Target
|
||||
|
||||
To satisfy the intended Create flow, implementation now includes:
|
||||
1. a Person page and dedicated create form
|
||||
2. user-entered controls for Person fields
|
||||
3. create validation and success/failure UX states
|
||||
4. post-submit routing to a Person detail page
|
||||
|
||||
## 5. READ Mapping
|
||||
|
||||
### 5.1 Intended Read Behavior
|
||||
|
||||
On the Person detail page, the user should be able to see:
|
||||
1. Person identity and biographical metadata
|
||||
2. linked Documents (through DocumentPerson)
|
||||
3. empty-state behavior when no linked documents exist
|
||||
|
||||
### 5.2 Current Implementation
|
||||
|
||||
Current Person visibility is implemented in dedicated list/detail/edit/delete pages.
|
||||
|
||||
| Field | Current Rendering | Visible to User | Notes | Evidence |
|
||||
|---|---|---|---|---|
|
||||
| full_name | Rendered in header and summary | Yes | Dedicated Person page exists | `src/transcription/ui/pages/people_page.py` |
|
||||
| display_name | Rendered | Yes | Visible in detail and list contexts | src/transcription/ui/pages/people_page.py |
|
||||
| maiden_name | Rendered | Yes | Visible in detail context | src/transcription/ui/pages/people_page.py |
|
||||
| birth_date | Rendered | Yes | Visible in detail context | src/transcription/ui/pages/people_page.py |
|
||||
| birth_date_raw | Rendered | Yes | Visible in detail context | src/transcription/ui/pages/people_page.py |
|
||||
| birth_place | Rendered | Yes | Visible in detail context | src/transcription/ui/pages/people_page.py |
|
||||
| death_date | Rendered | Yes | Visible in detail context | src/transcription/ui/pages/people_page.py |
|
||||
| death_date_raw | Rendered | Yes | Visible in detail context | src/transcription/ui/pages/people_page.py |
|
||||
| death_place | Rendered | Yes | Visible in detail context | src/transcription/ui/pages/people_page.py |
|
||||
| biography | Rendered | Yes | Visible in detail context | src/transcription/ui/pages/people_page.py |
|
||||
| portrait_path | Rendered as text and image when available | Yes | Dedicated Person page exists | `src/transcription/ui/pages/people_page.py` |
|
||||
| metadata_ | Not rendered | No | Hidden advanced field | no current UI field |
|
||||
| created_at | Rendered read-only | Yes | Visible in detail context | src/transcription/ui/pages/people_page.py |
|
||||
| updated_at | Rendered read-only | Yes | Visible in detail context | src/transcription/ui/pages/people_page.py |
|
||||
|
||||
### 5.3 Gap to Target
|
||||
|
||||
To satisfy the intended Read flow, implementation now includes:
|
||||
1. metadata rendering for Person fields
|
||||
2. linked Documents section with empty states
|
||||
3. document-link navigation paths
|
||||
|
||||
## 6. UPDATE Mapping
|
||||
|
||||
### 6.1 Intended Update Behavior
|
||||
|
||||
The user should be able to edit Person metadata from the Person detail page or a dedicated edit flow.
|
||||
|
||||
Intended editable fields:
|
||||
- full_name
|
||||
- display_name
|
||||
- maiden_name
|
||||
- birth_date
|
||||
- birth_date_raw
|
||||
- birth_place
|
||||
- death_date
|
||||
- death_date_raw
|
||||
- death_place
|
||||
- biography
|
||||
- portrait_path
|
||||
|
||||
Intended system-managed fields:
|
||||
- id
|
||||
- created_at
|
||||
- updated_at
|
||||
|
||||
Hidden in first release:
|
||||
- metadata_
|
||||
|
||||
### 6.2 Current Implementation
|
||||
|
||||
| Field | Updatable via UI | Updatable via Service | Notes |
|
||||
|---|---|---|---|
|
||||
| id | No | Practically no | Primary key should be treated as immutable |
|
||||
| full_name | Yes | Yes | Editable from Person edit page via DocumentService.update_person() |
|
||||
| display_name | Yes | Yes | Editable from Person edit page via DocumentService.update_person() |
|
||||
| maiden_name | Yes | Yes | Editable from Person edit page via DocumentService.update_person() |
|
||||
| birth_date | Yes | Yes | Editable from Person edit page via DocumentService.update_person() |
|
||||
| birth_date_raw | Yes | Yes | Editable from Person edit page via DocumentService.update_person() |
|
||||
| birth_place | Yes | Yes | Editable from Person edit page via DocumentService.update_person() |
|
||||
| death_date | Yes | Yes | Editable from Person edit page via DocumentService.update_person() |
|
||||
| death_date_raw | Yes | Yes | Editable from Person edit page via DocumentService.update_person() |
|
||||
| death_place | Yes | Yes | Editable from Person edit page via DocumentService.update_person() |
|
||||
| biography | Yes | Yes | Editable from Person edit page via DocumentService.update_person() |
|
||||
| portrait_path | Yes | Yes | Editable manually and via portrait upload helper |
|
||||
| metadata_ | No | Yes | Technically updatable, hidden in first release |
|
||||
| created_at | No | Technically yes | Should remain system-managed |
|
||||
| updated_at | No | Technically yes | Should remain system-managed |
|
||||
|
||||
### 6.3 Gap to Target
|
||||
|
||||
Implementation now includes:
|
||||
1. Person edit controls in the UI
|
||||
2. validation and save behavior for Person metadata
|
||||
3. a consistent updated_at update policy for Person edits
|
||||
|
||||
## 7. DELETE Mapping
|
||||
|
||||
### 7.1 Intended Delete Behavior
|
||||
|
||||
The UI should provide a delete action for Person with guardrails.
|
||||
|
||||
Rules:
|
||||
1. Deletion can proceed when relationship policy allows no retained document links.
|
||||
2. If linked DocumentPerson records exist and policy requires cleanup first, deletion is blocked.
|
||||
3. Delete confirmation must make clear that deletion is permanent.
|
||||
|
||||
### 7.2 Current Implementation
|
||||
|
||||
| Action | UI Exposed | Backend Capability | Notes |
|
||||
|---|---|---|---|
|
||||
| Delete Person | Yes | Yes | Dedicated delete page enforces linked-document guardrails before service delete |
|
||||
|
||||
### 7.3 Gap to Target
|
||||
|
||||
Implementation includes:
|
||||
1. a Person delete control in the UI
|
||||
2. relationship-aware pre-delete checks
|
||||
3. user-facing blocked-delete messaging
|
||||
4. confirmation UX for successful delete attempts
|
||||
|
||||
## 8. Hidden and System-Managed Fields
|
||||
|
||||
| Field | Category | Why Hidden or Protected |
|
||||
|---|---|---|
|
||||
| id | System-managed | Internal identifier |
|
||||
| created_at | System-managed | Audit timestamp |
|
||||
| updated_at | System-managed | Audit timestamp |
|
||||
| metadata_ | Hidden in first release | Advanced JSON metadata not needed in initial UI |
|
||||
|
||||
## 9. Structured Name Deferred Note
|
||||
|
||||
Structured name fields are deferred to a future schema revision.
|
||||
|
||||
Current policy:
|
||||
1. full_name remains canonical and required.
|
||||
|
||||
Future revision intent:
|
||||
1. introduce first_name, middle_name, last_name, and optional suffix fields.
|
||||
2. maintain compatibility with existing full_name records during migration.
|
||||
3. define normalization and reconciliation rules when structured and canonical forms differ.
|
||||
|
||||
## 10. Traceability Anchors
|
||||
|
||||
Schema and models:
|
||||
- docs/schema_v2.md
|
||||
- src/transcription/db/models.py
|
||||
|
||||
Current implementation:
|
||||
- src/transcription/services/documents.py
|
||||
- src/transcription/ui/pages/people_page.py
|
||||
- src/transcription/services/store.py
|
||||
|
||||
Companion UX spec:
|
||||
- docs/ui/entities/person/user-journey.md
|
||||
|
||||
Acceptance checklist:
|
||||
- docs/ui/entities/person/acceptance-criteria.md
|
||||
|
||||
## 11. Acceptance Checklist Summary
|
||||
|
||||
- Every Person schema field appears in the field inventory.
|
||||
- Intended Create behavior matches the companion user journey.
|
||||
- Current Create behavior reflects dedicated UI form implementation with optional portrait upload handling.
|
||||
- Gaps between intended and current behavior are explicit.
|
||||
- Read, Update, and Delete sections distinguish target behavior from current code.
|
||||
@@ -1,292 +0,0 @@
|
||||
# Person User Journey
|
||||
|
||||
Purpose: Define how a user should interact with the UI to create and manage a Person record, including expected inputs, validation, outcomes, and links to Document relationships.
|
||||
|
||||
Scope: This document describes intended user interaction for the Person UI. It is the UX contract for the Person entity.
|
||||
|
||||
Companion schema mapping: schema-mapping.md
|
||||
Companion acceptance criteria: acceptance-criteria.md
|
||||
|
||||
## 1. Overview
|
||||
|
||||
A Person represents a historical individual who may be associated with one or more Documents.
|
||||
|
||||
Managing a Person is a profile-first workflow:
|
||||
1. The user opens the Person page.
|
||||
2. The user selects Create new person.
|
||||
3. The user enters known biographical fields.
|
||||
4. The system creates the Person record.
|
||||
5. The user can later associate the Person with one or more Documents through DocumentPerson links.
|
||||
|
||||
## 2. User Goal
|
||||
|
||||
The user wants to:
|
||||
1. create and maintain historical person records
|
||||
2. reuse the same Person across multiple Documents
|
||||
3. record both precise and approximate date values where certainty is limited
|
||||
4. link people to documents as author or recipient in future flows
|
||||
|
||||
## 3. Page Model
|
||||
|
||||
### 3.1 Person Page
|
||||
|
||||
The Person page is the general UI surface where users manage people.
|
||||
|
||||
It should support:
|
||||
1. listing or locating existing people
|
||||
2. starting the Create new person flow
|
||||
3. navigating into a specific Person after it exists
|
||||
|
||||
### 3.2 Person Detail Page
|
||||
|
||||
The Person detail page is the page for one specific Person after creation.
|
||||
|
||||
It should show:
|
||||
1. core identity fields
|
||||
2. biographical metadata
|
||||
3. portrait image when available
|
||||
4. related Documents section
|
||||
5. empty state when no linked documents exist yet
|
||||
|
||||
## 4. Entry Point
|
||||
|
||||
Entry point: Person page
|
||||
|
||||
Primary action: Create new person
|
||||
|
||||
Expected UI affordance:
|
||||
1. a visible action labeled Create new person
|
||||
2. activation opens a dedicated form view, modal, or detail panel
|
||||
|
||||
Preferred first implementation:
|
||||
1. dedicated Person create page or panel
|
||||
2. simple labeled form controls
|
||||
3. text inputs are acceptable for first release
|
||||
|
||||
## 5. Create Person Form
|
||||
|
||||
### 5.1 Required Fields
|
||||
|
||||
| UI Label | Schema Field | Input Type | Required | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Full name | full_name | Text input | Yes | Canonical identity field |
|
||||
|
||||
### 5.2 Optional Name Fields
|
||||
|
||||
| UI Label | Schema Field | Input Type | Required | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Display name | display_name | Text input | No | Friendly or abbreviated display |
|
||||
| Maiden name | maiden_name | Text input | No | Historical alternate surname |
|
||||
|
||||
### 5.3 Birth and Death Date Fields
|
||||
|
||||
| UI Label | Schema Field | Input Type | Required | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Birth date | birth_date | Date input | No | Exact known date |
|
||||
| Birth date (approximate/raw) | birth_date_raw | Text input | No | Approximate or uncertain value |
|
||||
| Death date | death_date | Date input | No | Exact known date |
|
||||
| Death date (approximate/raw) | death_date_raw | Text input | No | Approximate or uncertain value |
|
||||
|
||||
Date handling rule:
|
||||
1. exact and raw values may both be entered
|
||||
2. exact date is canonical when present
|
||||
3. raw date is retained as historical context
|
||||
|
||||
### 5.4 Optional Biographical Fields
|
||||
|
||||
| UI Label | Schema Field | Input Type | Required | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Birth place | birth_place | Text input | No | Free text |
|
||||
| Death place | death_place | Text input | No | Free text |
|
||||
| Biography | biography | Text area | No | Narrative context |
|
||||
| Portrait path | portrait_path | Text input | No | File or resource path |
|
||||
| Metadata | metadata_ | Hidden or advanced JSON editor | No | Prefer hidden in first release |
|
||||
|
||||
### 5.5 System Fields
|
||||
|
||||
| Schema Field | User Editable | Notes |
|
||||
|---|---|---|
|
||||
| id | No | System-generated |
|
||||
| created_at | No | System-generated |
|
||||
| updated_at | No | System-managed |
|
||||
|
||||
## 6. Validation Rules
|
||||
|
||||
### 6.1 Required Validation
|
||||
|
||||
1. full_name is required
|
||||
2. save is blocked when full_name is empty
|
||||
|
||||
### 6.2 Date Validation
|
||||
|
||||
1. birth_date and birth_date_raw may coexist
|
||||
2. death_date and death_date_raw may coexist
|
||||
3. exact date fields are canonical when present
|
||||
4. raw fields remain descriptive context
|
||||
|
||||
### 6.3 Integrity Validation
|
||||
|
||||
1. form accepts unknown values for optional fields
|
||||
2. missing birth or death data does not block creation
|
||||
|
||||
## 7. Submission Behavior
|
||||
|
||||
On submit:
|
||||
1. The system validates required fields
|
||||
2. The system creates the Person record
|
||||
3. The system returns the user to the Person detail page
|
||||
4. The system shows a success message
|
||||
5. If portrait upload is used, the file is stored under uploads/portraits/person and portrait_path is set to that relative file path
|
||||
|
||||
Recommended transactional behavior:
|
||||
1. Person writes are atomic
|
||||
2. no partial save state should be persisted
|
||||
|
||||
## 8. Expected Result After Success
|
||||
|
||||
After successful creation:
|
||||
1. The user sees the Person detail page for the new record
|
||||
2. full_name is visible in the header or summary
|
||||
3. empty Related Documents section is shown if no links exist
|
||||
4. The user can proceed to link this person from Document workflows
|
||||
|
||||
## 9. Expected Result After Failure
|
||||
|
||||
If creation fails:
|
||||
1. Show a clear error message
|
||||
2. Show field-level feedback for validation failures
|
||||
3. preserve entered data where possible
|
||||
4. do not show false success messaging
|
||||
|
||||
## 10. Read Person Journey
|
||||
|
||||
### 10.1 User Intent
|
||||
|
||||
The user wants to open a Person and quickly understand:
|
||||
1. identity and key biography fields
|
||||
2. whether this person is linked to any documents
|
||||
3. what next action to take
|
||||
|
||||
### 10.2 Read Surfaces
|
||||
|
||||
The Person detail page should show:
|
||||
1. full_name and display fields
|
||||
2. birth and death fields
|
||||
3. biography summary
|
||||
4. related documents list or empty state
|
||||
5. portrait preview resolved from /uploads when portrait_path is a relative path
|
||||
|
||||
### 10.3 Read Empty State
|
||||
|
||||
If no linked documents exist:
|
||||
1. Show No linked documents yet
|
||||
2. provide guidance to link from Document workflow
|
||||
|
||||
## 11. Update Person Journey
|
||||
|
||||
### 11.1 User Intent
|
||||
|
||||
The user wants to correct or enrich person metadata over time.
|
||||
|
||||
### 11.2 Editable Fields
|
||||
|
||||
Editable:
|
||||
1. full_name
|
||||
2. display_name
|
||||
3. maiden_name
|
||||
4. birth_date
|
||||
5. birth_date_raw
|
||||
6. birth_place
|
||||
7. death_date
|
||||
8. death_date_raw
|
||||
9. death_place
|
||||
10. biography
|
||||
11. portrait_path
|
||||
|
||||
System-managed:
|
||||
1. id
|
||||
2. created_at
|
||||
3. updated_at
|
||||
4. metadata_ can remain hidden in first release
|
||||
|
||||
### 11.3 Update Save Behavior
|
||||
|
||||
On save:
|
||||
1. validate required fields
|
||||
2. persist updates
|
||||
3. refresh updated_at by system policy
|
||||
4. show confirmation
|
||||
5. keep user on Person detail page
|
||||
|
||||
### 11.4 Update Failure Behavior
|
||||
|
||||
1. Show clear error feedback
|
||||
2. preserve form state where possible
|
||||
3. Allow retry
|
||||
|
||||
## 12. Delete Person Journey
|
||||
|
||||
### 12.1 User Intent
|
||||
|
||||
The user wants to remove incorrect or duplicate person records safely.
|
||||
|
||||
### 12.2 Delete Guardrails
|
||||
|
||||
Delete is allowed when:
|
||||
1. Person has no required retained relationships
|
||||
|
||||
Delete is blocked when:
|
||||
1. Person is linked to one or more Documents via DocumentPerson and unlink policy requires cleanup first
|
||||
|
||||
### 12.3 Blocked Delete UX
|
||||
|
||||
1. explain that linked Document relationships exist
|
||||
2. Show link count or list
|
||||
3. provide cleanup path
|
||||
|
||||
### 12.4 Allowed Delete UX
|
||||
|
||||
1. Show a confirmation dialog
|
||||
2. confirm permanent action
|
||||
3. delete Person
|
||||
4. return to Person list with success message
|
||||
|
||||
## 13. Relationship to Other Workflows
|
||||
|
||||
This Person workflow integrates with:
|
||||
1. Document create and update workflows through person lookup and linking
|
||||
2. DocumentPerson mapping for role assignments
|
||||
3. future recipient and multi-person enhancements
|
||||
|
||||
## 14. Relationship to Schema Mapping
|
||||
|
||||
The companion schema-mapping document should specify:
|
||||
1. field visibility per CRUD action
|
||||
2. current implementation status
|
||||
3. intended behavior
|
||||
4. gap-to-target items
|
||||
|
||||
## 15. Deferred Items
|
||||
|
||||
Deferred to future revisions:
|
||||
1. advanced metadata_ editing UI
|
||||
2. multi-person role editing in the Person UI itself
|
||||
3. richer relationship timeline views
|
||||
4. bulk merge or dedup workflows
|
||||
5. structured name fields migration (first_name, middle_name, last_name, optional suffix)
|
||||
|
||||
### 15.1 Structured Name Fields Migration Note
|
||||
|
||||
For now, `full_name` remains the canonical required name field.
|
||||
|
||||
Future revision intent:
|
||||
1. introduce structured fields such as first_name, middle_name, last_name, and optional suffix
|
||||
2. keep full_name during transition for backward compatibility and historical formatting
|
||||
3. define normalization and formatting rules for display and sorting
|
||||
4. update search and dedup workflows to use both structured and canonical forms during migration
|
||||
|
||||
Migration considerations:
|
||||
1. schema migration and backfill strategy for existing Person records
|
||||
2. validation updates for create and update forms
|
||||
3. compatibility for existing APIs and UI components that currently rely on full_name
|
||||
4. clear precedence and reconciliation rules when structured fields and full_name differ
|
||||
@@ -1,141 +0,0 @@
|
||||
# Source Acceptance Criteria
|
||||
|
||||
Purpose: Define implementation-ready acceptance criteria for Source Create, Read, Update, and Delete workflows.
|
||||
|
||||
Companion documents:
|
||||
- docs/ui/entities/source/user-journey.md
|
||||
- docs/ui/entities/source/schema-mapping.md
|
||||
|
||||
## Scope
|
||||
|
||||
This checklist covers:
|
||||
1. Create flow
|
||||
2. Read flow
|
||||
3. Update flow
|
||||
4. Delete flow
|
||||
|
||||
This checklist does not cover:
|
||||
1. advanced multi-version revision history design
|
||||
2. job orchestration state-machine behavior
|
||||
3. provider-level transcription internals
|
||||
|
||||
## Create Acceptance Criteria
|
||||
|
||||
### CR-1 Source creation entry
|
||||
1. Given the user is in job creation or job configuration flow
|
||||
2. When the user selects Add sources
|
||||
3. Then the user can upload one or more source files or a folder
|
||||
4. Then source creation is not offered as a standalone first-release document-only flow
|
||||
|
||||
### CR-2 Required create values
|
||||
1. document_id is derived from selected Document context
|
||||
2. JobSource.job_id is derived from the active Job context
|
||||
3. Each created Source is linked to the active Job through JobSource at create time
|
||||
4. page_number is assigned to preserve ordering
|
||||
5. upload_name, filename, and file_path are persisted for each created source
|
||||
|
||||
### CR-3 Ordering and filename strategy
|
||||
1. Given a multi-file or folder upload
|
||||
2. When source records are created
|
||||
3. Then page ordering follows alphabetical order of original filenames
|
||||
4. Then upload_name stores the original filename
|
||||
5. Then filename is stored using UUID plus original extension in the form UUID.extension
|
||||
|
||||
### CR-4 Successful create outcome
|
||||
1. Given valid uploads
|
||||
2. When source creation completes
|
||||
3. Then Source records are created and linked to the Document
|
||||
4. Then Source records are linked to the active Job through JobSource
|
||||
5. Then source list reflects new pages in sequence
|
||||
6. Then the user can open preview or revision workflow
|
||||
|
||||
### CR-5 Create failure outcome
|
||||
1. Given upload or persistence failure
|
||||
2. Then clear error feedback is shown
|
||||
3. Then no false success feedback is shown
|
||||
4. Then retry path remains available
|
||||
5. Then creation fails when required Document or Job linkage cannot be established
|
||||
|
||||
## Read Acceptance Criteria
|
||||
|
||||
### RD-1 Source detail retrieval
|
||||
1. Given a valid Source id in source context
|
||||
2. When the user opens source detail
|
||||
3. Then source metadata and preview are displayed for that source only
|
||||
|
||||
### RD-2 Transcription and revision visibility
|
||||
1. Original transcription context is visible read-only in Source detail
|
||||
2. Revision state is visible in Source detail
|
||||
3. If revised_text is absent, revision input opens as empty and can be edited
|
||||
|
||||
### RD-3 Missing source state
|
||||
1. Given a missing source
|
||||
2. Then UI shows clear no source available or not found messaging without crashing
|
||||
|
||||
## Update Acceptance Criteria
|
||||
|
||||
### UP-1 Revision editing entry
|
||||
1. Given a source context
|
||||
2. When the user enters revision edit flow
|
||||
3. Then revised_text input is available in Source detail
|
||||
|
||||
### UP-2 Revision validation
|
||||
1. revised_text cannot be saved as empty after trimming
|
||||
2. Warning feedback is shown for invalid empty input
|
||||
|
||||
### UP-3 Successful revision save
|
||||
1. Given valid revision text
|
||||
2. When the user saves
|
||||
3. Then revised_text persists
|
||||
4. Then date_revised is updated
|
||||
5. Then success feedback is shown
|
||||
6. Then refreshed revision content is visible
|
||||
|
||||
### UP-4 Revision save failure
|
||||
1. Given backend failure during save
|
||||
2. Then clear error feedback is shown
|
||||
3. Then the user-entered text remains available for retry where possible
|
||||
|
||||
## Delete Acceptance Criteria
|
||||
|
||||
### DL-1 Delete entry and confirmation
|
||||
1. Given a source in source context
|
||||
2. When the user selects delete source
|
||||
3. Then a permanent-action confirmation dialog appears
|
||||
|
||||
### DL-2 Dependency guardrails
|
||||
1. If policy requires cleanup of related JobSource records first, delete is blocked
|
||||
2. If policy allows dependent cleanup path, delete can proceed
|
||||
|
||||
### DL-3 Blocked delete behavior
|
||||
1. When blocked
|
||||
2. Then UI explains dependency constraints
|
||||
3. Then UI provides guidance for dependency cleanup
|
||||
|
||||
### DL-4 Successful delete
|
||||
1. Given no blocking dependencies
|
||||
2. When the user confirms deletion
|
||||
3. Then source is removed
|
||||
4. Then success feedback is shown
|
||||
5. Then the user returns to source list context
|
||||
|
||||
### DL-5 Delete failure
|
||||
1. Given backend failure during delete
|
||||
2. Then clear error feedback is shown
|
||||
3. Then the user remains in source context with retry path
|
||||
|
||||
## Cross-Criteria Quality Gates
|
||||
|
||||
### QG-1 Separation of intent and implementation
|
||||
1. UX intent remains in user-journey.md
|
||||
2. Current versus target implementation mapping remains in schema-mapping.md
|
||||
|
||||
### QG-2 Traceability
|
||||
1. Each accepted behavior maps to at least one future UI action or service path
|
||||
2. No acceptance criterion contradicts first-release deferred items
|
||||
|
||||
### QG-3 First-release constraints
|
||||
1. Source creation remains job-create-centric
|
||||
2. revised_text is the primary editable source field in first release
|
||||
3. source creation requires both Document linkage and Job linkage at create time
|
||||
4. source delete management surfaces are phased in later
|
||||
@@ -1,212 +0,0 @@
|
||||
# Source Schema-to-UI Mapping
|
||||
|
||||
Purpose: Map the Source schema to the UI, while clearly separating intended target behavior from current implementation.
|
||||
|
||||
Companion document: user-journey.md
|
||||
Acceptance criteria: acceptance-criteria.md
|
||||
|
||||
## 1. Entity Snapshot
|
||||
|
||||
- Table: Source
|
||||
- Primary key: id (UUID)
|
||||
- Related entities: Document, JobSource, Job
|
||||
- Canonical schema references:
|
||||
- src/transcription/db/models.py
|
||||
- docs/schema_v2.md
|
||||
|
||||
## 2. Mapping Rules
|
||||
|
||||
This document uses three lenses:
|
||||
1. Intended behavior: what the UX should support.
|
||||
2. Current behavior: what the code supports today.
|
||||
3. Gap to target: what must change to align implementation with the intended UX.
|
||||
|
||||
## 3. Field Inventory
|
||||
|
||||
| Field | DB Type | Nullable | Default/Auto Value | Intended UI Treatment | Notes |
|
||||
|---|---|---|---|---|---|
|
||||
| id | UUID | No | uuid4() | Hidden, system-managed | Primary key |
|
||||
| document_id | UUID FK | No | None | Hidden/context-managed | Selected Document context |
|
||||
| page_number | int | No | 1 | Shown read-only or ordered list | Sequential ordering |
|
||||
| upload_name | str | No | None | Shown read-only after upload | Original user-provided name |
|
||||
| filename | str | No | None | Shown read-only | Stored filename |
|
||||
| file_path | str | No | None | Usually hidden; preview uses path internally | Filesystem path |
|
||||
| raw_transcription | str | Yes | None | Shown indirectly or hidden | Immutable machine output context |
|
||||
| revised_text | str | Yes | None | Editable in Source detail | Human-authored correction |
|
||||
| date_uploaded | datetime | No | datetime.now(UTC) | Shown read-only | System-managed timestamp |
|
||||
| date_revised | datetime | Yes | None | Shown read-only | Set when revision is saved |
|
||||
|
||||
## 4. CREATE Mapping
|
||||
|
||||
### 4.1 Intended Create Flow
|
||||
|
||||
Entry point: Job creation or job configuration Add sources action
|
||||
User action: upload one or more source files, or a whole folder
|
||||
Success destination: source preview or revision flow in job detail context
|
||||
|
||||
| Field | Intended User Input | Required | Visible | Notes |
|
||||
|---|---|---|---|---|
|
||||
| document_id | Hidden/context | Yes | No | Comes from selected Document |
|
||||
| JobSource.job_id | Hidden/context | Yes | No | Comes from active Job; required for first release |
|
||||
| page_number | Auto or user-assisted ordering | Yes | Indirectly | Should preserve sequence |
|
||||
| upload_name | File picker name | Yes | Yes | Original display name |
|
||||
| filename | None | Yes | No or read-only | System-stored as UUID.extension |
|
||||
| file_path | None | Yes | No | Storage path |
|
||||
| raw_transcription | None | No | No | Filled by processing |
|
||||
| revised_text | None | No | No | Initially empty |
|
||||
| date_uploaded | None | No | No | System-generated |
|
||||
| date_revised | None | No | No | Null until revision |
|
||||
|
||||
### 4.2 Current Implementation
|
||||
|
||||
Current entry point: Jobs page create flow
|
||||
Current user action: upload one or more files or a folder through a single upload widget
|
||||
Current backend path: job create submit -> create_job_for_document()
|
||||
|
||||
| Field | Current Value at Create | Source | Visible to User | Evidence |
|
||||
|---|---|---|---|---|
|
||||
| id | Generated UUID | System | No | Source model default in src/transcription/db/models.py |
|
||||
| document_id | Selected existing Document id | Job create selection + service write | Indirectly | src/transcription/ui/pages/jobs_page.py, src/transcription/services/store.py |
|
||||
| page_number | Sequential assignment based on existing max and alphabetical upload order | Service | No | src/transcription/services/store.py |
|
||||
| upload_name | original filename basename | User file name transformed by service | Indirectly | src/transcription/services/store.py |
|
||||
| filename | stored generated filename | Service | Indirectly | src/transcription/services/store.py |
|
||||
| file_path | stored path | Service | Indirectly | src/transcription/services/store.py |
|
||||
| raw_transcription | None initially | System | No at create | Source model defaults |
|
||||
| revised_text | None initially | System | No at create | Source model defaults |
|
||||
| date_uploaded | current UTC timestamp | System | No | Source model default |
|
||||
| date_revised | None | System | No | Source model default |
|
||||
|
||||
### 4.3 Gap to Target
|
||||
|
||||
To satisfy intended Create flow, implementation now includes:
|
||||
1. multi-source and folder upload support in job create/configure flows
|
||||
2. deterministic page_number assignment from alphabetical original filename ordering
|
||||
3. enforced create-time Source-to-Document and Source-to-Job linkage invariants
|
||||
4. filename storage policy using UUID.extension
|
||||
|
||||
## 5. READ Mapping
|
||||
|
||||
### 5.1 Intended Read Behavior
|
||||
|
||||
On Source detail/list surfaces, users should be able to see:
|
||||
1. source page preview
|
||||
2. source metadata and ordering
|
||||
3. revision state
|
||||
4. original transcription context
|
||||
|
||||
### 5.2 Current Implementation
|
||||
|
||||
Current Source reading is centered on dedicated Sources list/detail routes with optional document/job filtering.
|
||||
|
||||
| Field | Current Rendering | Visible to User | Notes | Evidence |
|
||||
|---|---|---|---|---|
|
||||
| upload_name | Shown in Sources list and Source detail | Yes | Displayed in source context | src/transcription/ui/pages/sources_page.py |
|
||||
| filename | Shown in Sources list and Source detail | Yes | Source metadata shown in list/detail | src/transcription/ui/pages/sources_page.py |
|
||||
| file_path | Hidden from direct text rendering | No | Used internally for preview rendering | src/transcription/ui/components/document_panzoom.py |
|
||||
| page_number | Shown in Sources list and Source detail | Yes | Ordering visible in filtered/global list | src/transcription/ui/pages/sources_page.py |
|
||||
| raw_transcription | Shown read-only in Source detail | Yes | Read from latest linked JobSource context | src/transcription/ui/pages/sources_page.py |
|
||||
| revised_text | Shown and editable in Source detail | Yes | Saved through revision action | src/transcription/ui/pages/sources_page.py |
|
||||
| date_uploaded | Shown in Source detail | Yes | Read-only metadata | src/transcription/ui/pages/sources_page.py |
|
||||
| date_revised | Shown in Source detail | Yes | Read-only metadata after revision save | src/transcription/ui/pages/sources_page.py |
|
||||
|
||||
### 5.3 Gap to Target
|
||||
|
||||
To satisfy intended Read flow, implementation must add:
|
||||
1. optional list filtering controls in-page (current filtering is URL/context based)
|
||||
2. optional page-specific navigation enhancements beyond current list/detail pattern
|
||||
|
||||
## 6. UPDATE Mapping
|
||||
|
||||
### 6.1 Intended Update Behavior
|
||||
|
||||
Primary user update for Source is revised_text maintenance in Source detail.
|
||||
|
||||
Intended editable fields (first release):
|
||||
- revised_text
|
||||
|
||||
Intended read-only fields (first release):
|
||||
- document_id
|
||||
- page_number
|
||||
- upload_name
|
||||
- filename
|
||||
- file_path
|
||||
- raw_transcription
|
||||
- date_uploaded
|
||||
- date_revised
|
||||
|
||||
### 6.2 Current Implementation
|
||||
|
||||
| Field | Updatable via UI | Updatable via Service | Notes |
|
||||
|---|---|---|---|
|
||||
| revised_text | Yes | Yes | Saved via TranscriptionService.upsert_revision_for_source() from Source detail |
|
||||
| date_revised | No | Yes | Set automatically on revision save |
|
||||
| other fields | No | Technically yes in service layer | No first-class UI editing flow |
|
||||
|
||||
### 6.3 Gap to Target
|
||||
|
||||
Implementation should add in a later revision:
|
||||
1. optional future controls for page ordering and metadata corrections
|
||||
2. revision history and conflict-resolution UX beyond single revised_text updates
|
||||
|
||||
## 7. DELETE Mapping
|
||||
|
||||
### 7.1 Intended Delete Behavior
|
||||
|
||||
Source deletion is deferred in the current UI.
|
||||
|
||||
Rules:
|
||||
1. Deletion can proceed when policy allows cleanup of related JobSource records.
|
||||
2. If related execution history must be preserved first, deletion is blocked with guidance.
|
||||
|
||||
### 7.2 Current Implementation
|
||||
|
||||
| Action | UI Exposed | Backend Capability | Notes |
|
||||
|---|---|---|---|
|
||||
| Delete Source | No | Yes | TranscriptionService.delete_source() exists, no dedicated UI delete flow |
|
||||
|
||||
### 7.3 Gap to Target
|
||||
|
||||
Implementation should add in a future revision:
|
||||
1. source delete controls in source/document context UI
|
||||
2. dependency checks for JobSource links
|
||||
3. blocked-delete messaging and cleanup path guidance
|
||||
4. confirmation UX for successful delete attempts
|
||||
|
||||
## 8. Hidden and System-Managed Fields
|
||||
|
||||
| Field | Category | Why Hidden or Protected |
|
||||
|---|---|---|
|
||||
| id | System-managed | Internal identifier |
|
||||
| document_id | Context-managed | Derived from selected document context |
|
||||
| file_path | Operational/internal | Used for file storage and preview plumbing |
|
||||
| date_uploaded | System-managed | Audit timestamp |
|
||||
| date_revised | System-managed | Revision timestamp set by system |
|
||||
|
||||
## 9. Traceability Anchors
|
||||
|
||||
Schema and models:
|
||||
- docs/schema_v2.md
|
||||
- src/transcription/db/models.py
|
||||
|
||||
Current implementation:
|
||||
- src/transcription/services/store.py
|
||||
- src/transcription/services/transcription.py
|
||||
- src/transcription/ui/pages/sources_page.py
|
||||
- src/transcription/ui/pages/jobs_page.py
|
||||
- src/transcription/ui/pages/documents_page.py
|
||||
- src/transcription/ui/components/document_panzoom.py
|
||||
|
||||
Companion UX spec:
|
||||
- docs/ui/entities/source/user-journey.md
|
||||
|
||||
Acceptance checklist:
|
||||
- docs/ui/entities/source/acceptance-criteria.md
|
||||
|
||||
## 10. Acceptance Checklist Summary
|
||||
|
||||
- Every Source schema field appears in the field inventory.
|
||||
- Intended Create behavior matches the companion user journey.
|
||||
- Source create invariant requires both Document linkage and Job linkage at create time.
|
||||
- Current behavior reflects upload-centric create flow and dedicated Sources list/detail review flow.
|
||||
- Gaps between intended and current behavior are explicit.
|
||||
- Read, Update, and Delete sections distinguish target behavior from current code.
|
||||
@@ -1,229 +0,0 @@
|
||||
# Source User Journey
|
||||
|
||||
Purpose: Define how a user should interact with the UI to create and manage Source records, including page-level transcription context and revision behavior.
|
||||
|
||||
Scope: This document describes intended user interaction for the Source UI. It is the UX contract for the Source entity.
|
||||
|
||||
Companion schema mapping: schema-mapping.md
|
||||
Companion acceptance criteria: acceptance-criteria.md
|
||||
|
||||
## 1. Overview
|
||||
|
||||
A Source represents one page or file unit associated with a Document.
|
||||
|
||||
Managing Source records is page-first:
|
||||
1. The user starts from a transcription job flow.
|
||||
2. The user adds one or more source files.
|
||||
3. The system creates Source records linked to the Document and linked to the Job through JobSource.
|
||||
4. The user reviews source lists from a dedicated Sources page.
|
||||
5. The user opens Source detail to review preview, metadata, transcription text, and revision text.
|
||||
|
||||
## 2. User Goal
|
||||
|
||||
The user wants to:
|
||||
1. add page files to a Document
|
||||
2. ensure every source is attached to the transcription job context
|
||||
3. keep page order reliable
|
||||
4. review original machine output
|
||||
5. save human revisions per page
|
||||
6. navigate source pages efficiently
|
||||
|
||||
## 3. Page Model
|
||||
|
||||
### 3.1 Source List Surface
|
||||
|
||||
A Source list surface should support:
|
||||
1. listing source pages globally or filtered by selected Document or Job
|
||||
2. sorting by page_number
|
||||
3. opening the owning Document or Job context
|
||||
4. opening Source detail for a selected source
|
||||
|
||||
### 3.2 Source Detail Surface
|
||||
|
||||
Source detail supports:
|
||||
1. pan/zoom image or PDF preview
|
||||
2. read-only source metadata (page number, names, timestamps)
|
||||
3. read-only original transcription text
|
||||
4. editable revision text with save action
|
||||
|
||||
## 4. Entry Points
|
||||
|
||||
Primary entry points:
|
||||
1. from Job workflow, Add sources while creating or configuring a job
|
||||
2. from Job detail, open filtered Sources for the current Job
|
||||
3. from Document detail, open filtered Sources for the current Document
|
||||
4. from global navigation, open all Sources
|
||||
|
||||
Current implementation note:
|
||||
1. source interaction occurs in job-create flow and dedicated Sources list/detail flows
|
||||
|
||||
## 5. Create Source Flow
|
||||
|
||||
### 5.1 User Intent
|
||||
|
||||
The user wants to attach one or more files to a Document so each page can be processed and reviewed.
|
||||
|
||||
### 5.2 Create from Job Context
|
||||
|
||||
1. The user starts from a job-creation or job-configuration flow
|
||||
2. The user can upload one or more files, or upload a whole folder
|
||||
3. The system creates Source rows linked to the selected Document
|
||||
4. The system creates JobSource links for the active Job as part of this flow
|
||||
5. Source creation fails if required Document or Job linkage cannot be established
|
||||
|
||||
### 5.3 Source Create Inputs
|
||||
|
||||
| UI Label | Schema Field | Input Type | Required | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Source files | upload_name/filename/file_path | Multi-file upload or folder upload | Yes | User may select one file, many files, or a folder |
|
||||
| Processing order | page_number assignment rule | System rule | Yes | If multiple files are uploaded, processing order is alphabetical by original filename |
|
||||
| Document reference | document_id | Hidden/context | Yes | Comes from selected Document |
|
||||
| Job reference | JobSource.job_id | Hidden/context | Yes | Required for first-release source creation |
|
||||
|
||||
### 5.4 Filename Strategy
|
||||
|
||||
1. store original user filename in upload_name
|
||||
2. store persisted filename using UUID plus original extension only, in the form UUID.extension
|
||||
3. this replaces the previous UUID-upload_name.extension pattern
|
||||
|
||||
### 5.5 Ordering Guidance
|
||||
|
||||
1. multi-file or folder uploads are processed alphabetically by original filename
|
||||
2. UI should show a warning or helper note so users understand that filename conventions control order
|
||||
|
||||
Suggested helper text:
|
||||
1. Files are processed alphabetically by original filename. Use leading numbers such as 001, 002, 003 to control page order.
|
||||
|
||||
### 5.6 System-Managed Values at Create
|
||||
|
||||
| Schema Field | User Editable | Notes |
|
||||
|---|---|---|
|
||||
| id | No | System-generated |
|
||||
| date_uploaded | No | System-generated |
|
||||
| raw_transcription | No | Filled later by processing |
|
||||
| revised_text | No | Initially empty |
|
||||
| date_revised | No | Initially null |
|
||||
|
||||
### 5.7 Expected Create Result
|
||||
|
||||
After successful source create:
|
||||
1. Source is linked to the Document
|
||||
2. Source appears in page order derived from alphabetical upload filename ordering
|
||||
3. Source is linked to the Job through JobSource at create time
|
||||
4. The user can open the owning Document or Job context
|
||||
|
||||
### 5.8 Source Creation Invariant
|
||||
|
||||
For first release:
|
||||
1. every new Source must have a Document link (Source.document_id)
|
||||
2. every new Source must have a Job link through JobSource (JobSource.job_id -> JobSource.source_id)
|
||||
3. source creation is treated as part of transcription workflow, not a standalone document-only upload path
|
||||
|
||||
## 6. Read Source Journey
|
||||
|
||||
### 6.1 User Intent
|
||||
|
||||
The user wants to view each page file and understand file identity and processing context.
|
||||
|
||||
### 6.2 Read Surface Expectations
|
||||
|
||||
The UI should show:
|
||||
1. source lists for current context (all, document-filtered, or job-filtered)
|
||||
2. upload_name as the original user-provided filename
|
||||
3. filename as the stored system filename
|
||||
4. page_number and ordering context
|
||||
5. the owning Document and Job navigation context
|
||||
6. direct action to open Source detail
|
||||
|
||||
### 6.3 Read Empty and Missing States
|
||||
|
||||
If source is missing:
|
||||
1. Show clear not found or no source available messaging
|
||||
|
||||
If source metadata is partially unavailable:
|
||||
1. Show fallback labels and keep navigation available where possible
|
||||
|
||||
## 7. Update Source Journey
|
||||
|
||||
### 7.1 User Intent
|
||||
|
||||
The user primarily tracks page-level source records while preserving raw machine output in the service layer.
|
||||
|
||||
### 7.2 Intended Editable Fields
|
||||
|
||||
Editable in first release:
|
||||
1. revised_text in Source detail
|
||||
|
||||
Read-only in first release:
|
||||
1. upload_name
|
||||
2. filename
|
||||
3. file_path
|
||||
4. raw_transcription
|
||||
5. page_number
|
||||
6. date_uploaded
|
||||
7. date_revised set by system on revision save
|
||||
|
||||
### 7.3 Revision Save Behavior
|
||||
|
||||
On save:
|
||||
1. validate revision text is non-empty after trimming
|
||||
2. persist revised_text
|
||||
3. set date_revised
|
||||
4. show success feedback
|
||||
5. keep user in current source context
|
||||
|
||||
### 7.4 Revision Failure Behavior
|
||||
|
||||
If save fails:
|
||||
1. Show clear error feedback
|
||||
2. keep user input where possible
|
||||
3. Allow retry
|
||||
|
||||
## 8. Delete Source Journey
|
||||
|
||||
### 8.1 User Intent
|
||||
|
||||
The user may need to remove incorrect or duplicate source files from a Document.
|
||||
|
||||
### 8.2 Guardrails
|
||||
|
||||
Delete is allowed when:
|
||||
1. policy allows removal of related processing history
|
||||
|
||||
Delete is blocked when:
|
||||
1. policy requires preserving dependent job-source execution records until explicit cleanup
|
||||
|
||||
### 8.3 Delete UX
|
||||
|
||||
When blocked:
|
||||
1. explain dependency constraints in a future delete flow
|
||||
2. show cleanup guidance in a future delete flow
|
||||
|
||||
When allowed:
|
||||
1. confirm permanent removal in a future delete flow
|
||||
2. remove source in a future delete flow
|
||||
3. return to source list with success state in a future delete flow
|
||||
|
||||
## 9. Relationship to Other Workflows
|
||||
|
||||
Source workflow integrates with:
|
||||
1. Document workflow for ownership and page organization
|
||||
2. Job workflow for processing status and outputs
|
||||
3. revision workflow for human correction lifecycle
|
||||
|
||||
## 10. Relationship to Schema Mapping
|
||||
|
||||
The companion schema-mapping document should specify:
|
||||
1. field visibility per CRUD action
|
||||
2. current implementation status
|
||||
3. intended behavior
|
||||
4. gap-to-target items
|
||||
|
||||
## 11. Deferred Items
|
||||
|
||||
Deferred to future revisions:
|
||||
1. bulk page reordering UX
|
||||
2. multi-file upload progress and resumable upload UX
|
||||
3. revision history versions beyond a single revised_text field
|
||||
4. richer per-page status dashboards
|
||||
5. source delete UI with dependency-aware confirmation
|
||||
@@ -1,78 +0,0 @@
|
||||
# UI Entity Traceability Matrix
|
||||
|
||||
Purpose: Map acceptance criteria to concrete implementation anchors and current delivery status.
|
||||
|
||||
Updated: 2026-08-02
|
||||
|
||||
Status legend:
|
||||
- Implemented: behavior exists in current UI and service flow
|
||||
- Partial: parts exist, but user-facing behavior or guardrails are incomplete
|
||||
- Planned: documented intent with no dedicated UI implementation yet
|
||||
|
||||
## Document
|
||||
|
||||
| Criteria Group | Acceptance IDs | Status | Primary Implementation Anchors | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Read detail and metadata | RD-1, RD-2, RD-7 | Implemented | src/transcription/ui/pages/documents_page.py; src/transcription/services/documents.py; tests/ui/test_documents_page.py | Dedicated Document detail route renders metadata, read-only system timestamps, and invalid/missing-id states. |
|
||||
| Related sections and navigation | RD-3, RD-4, RD-5, RD-6 | Implemented | src/transcription/ui/pages/documents_page.py; src/transcription/services/documents.py; tests/ui/test_documents_page.py | Document detail now shows linked people plus document-scoped Sources and Jobs navigation for the current document. |
|
||||
| Update entry, validation, and author linkage | UP-1, UP-2, UP-3, UP-4, UP-5, UP-6 | Implemented | src/transcription/ui/pages/documents_page.py; src/transcription/services/documents.py; tests/ui/test_documents_page.py; tests/services/test_document_service.py | Dedicated edit page includes required-field validation messaging, date parsing rules, and author relationship selection with save path routed back to document detail. |
|
||||
| Delete controls and guardrails | DL-1, DL-2, DL-3, DL-4, DL-5 | Implemented | src/transcription/ui/pages/documents_page.py; src/transcription/services/documents.py; tests/ui/test_documents_page.py; tests/services/test_document_service.py | Dedicated delete page provides permanent-action confirmation, dependency-category blocking, and guarded backend delete behavior. |
|
||||
|
||||
## Person
|
||||
|
||||
| Criteria Group | Acceptance IDs | Status | Primary Implementation Anchors | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Create flow and validation | CR-1, CR-2, CR-3, CR-4, CR-5 | Implemented | src/transcription/ui/pages/people_page.py; src/transcription/services/documents.py; tests/ui/test_people_page.py | Dedicated Person create page with required full_name validation, optional field handling, and success routing to detail. |
|
||||
| Read detail and linked documents | RD-1, RD-2, RD-3, RD-4 | Implemented | src/transcription/ui/pages/people_page.py; src/transcription/services/documents.py; tests/ui/test_people_page.py; tests/services/test_document_service.py | Person detail route renders metadata, full-name summary, portrait preview when available, linked-document section, and invalid/missing-id states. |
|
||||
| Update behavior | UP-1, UP-2, UP-3, UP-4, UP-5 | Implemented | src/transcription/ui/pages/people_page.py; src/transcription/services/documents.py; tests/ui/test_people_page.py; tests/services/test_document_service.py | Dedicated Person edit page supports allowed fields, required full_name validation, and save path back to detail. |
|
||||
| Delete behavior and guardrails | DL-1, DL-2, DL-3, DL-4, DL-5 | Implemented | src/transcription/ui/pages/people_page.py; src/transcription/services/documents.py; tests/ui/test_people_page.py; tests/services/test_document_service.py | Dedicated delete page provides permanent-action confirmation, linked-document blocking message, and guarded backend delete behavior. |
|
||||
|
||||
## Source
|
||||
|
||||
| Criteria Group | Acceptance IDs | Status | Primary Implementation Anchors | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Create entry and required links | CR-1, CR-2, CR-4, CR-5 | Implemented | src/transcription/services/store.py; src/transcription/ui/pages/jobs_page.py; src/transcription/ui/pages/upload_page.py; tests/services/test_store.py; tests/ui/test_jobs_page.py | Source upload/create is job-create-context only (legacy upload route redirects), with required Document and JobSource linkage enforced. |
|
||||
| Ordering and filename policy | CR-3 | Implemented | src/transcription/services/store.py; tests/services/test_store.py; tests/ui/test_jobs_page.py | Multi-file/folder uploads are ordered alphabetically by original filename, helper text is visible, and stored filenames use generated unique-id plus extension. |
|
||||
| Read and navigation visibility | RD-1, RD-2, RD-3 | Implemented | src/transcription/ui/pages/sources_page.py; src/transcription/ui/pages/documents_page.py; src/transcription/ui/pages/jobs_page.py; tests/ui/test_sources_page.py; tests/ui/test_documents_page.py; tests/ui/test_jobs_page.py | Dedicated Sources list/detail routes support global, document-filtered, and job-filtered navigation plus source metadata and preview rendering. |
|
||||
| Revision update behavior | UP-1, UP-2, UP-3, UP-4 | Implemented | src/transcription/ui/pages/sources_page.py; src/transcription/services/transcription.py; tests/ui/test_sources_page.py; tests/services/test_transcription_service.py | Source detail exposes revision edit/save UX with non-empty validation, success feedback, and refreshed state after save. |
|
||||
| Delete and dependency guardrails | DL-1, DL-2, DL-3, DL-4, DL-5 | Planned | src/transcription/services/transcription.py; tests/services/test_transcription_service.py | Job-detail source delete UI was removed from the current simplified flow; backend guardrails remain for future reinstatement. |
|
||||
|
||||
## Job
|
||||
|
||||
| Criteria Group | Acceptance IDs | Status | Primary Implementation Anchors | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Create entry and required links | CR-1, CR-2, CR-5, CR-6 | Implemented | src/transcription/services/store.py; src/transcription/ui/pages/jobs_page.py; tests/ui/test_jobs_page.py; tests/services/test_store.py | Jobs list now has explicit Create entry and `/jobs/new` create flow with Document selection, combined file/folder upload widget, and submit routing to job detail. |
|
||||
| Source ordering and upload behavior | CR-3 | Implemented | src/transcription/services/store.py; src/transcription/ui/pages/jobs_page.py; tests/services/test_store.py; tests/ui/test_jobs_page.py | Multi-file and folder upload are supported through one widget, uploads are sorted alphabetically by original filename, and helper guidance is shown in create UI. |
|
||||
| Provider/model/prompt visibility | CR-4, RD-4 | Implemented | src/transcription/services/workflows.py; src/transcription/ui/pages/jobs_page.py; tests/ui/test_jobs_page.py | Provider/model/prompt fields are visible in create and detail flows when known (with pending fallback labels). |
|
||||
| Jobs list and detail read states | RD-1, RD-2, RD-3, RD-5 | Implemented | src/transcription/ui/pages/jobs_page.py; src/transcription/ui/components/table/jobs.py; tests/ui/test_jobs_page.py | Jobs list, detail route, document-scoped navigation, and invalid/missing id states are present. |
|
||||
| Revision update behavior | UP-1, UP-2, UP-3, UP-4 | Implemented | src/transcription/ui/pages/jobs_page.py; src/transcription/ui/pages/sources_page.py; src/transcription/services/transcription.py; tests/ui/test_jobs_page.py; tests/ui/test_sources_page.py; tests/services/test_transcription_service.py | Job detail routes users to job-scoped Sources where Source detail provides revision edit/save workflow. |
|
||||
| Lifecycle visibility and retry indicators | UP-5 | Implemented | src/transcription/services/jobs.py; src/transcription/services/workflows.py; src/transcription/ui/pages/jobs_page.py; tests/ui/test_jobs_page.py | Job detail now surfaces lifecycle status plus retry/update metadata while lifecycle fields remain system-managed (no direct user edit controls). |
|
||||
| Delete and dependency guardrails | DL-1, DL-2, DL-3, DL-4, DL-5 | Implemented | src/transcription/ui/pages/jobs_page.py; src/transcription/services/jobs.py; tests/ui/test_jobs_page.py; tests/services/test_job_service.py | Job delete page enforces processing-state block, confirms allowed deletes, and routes back to jobs list on success. |
|
||||
|
||||
## Quality Gate Coverage
|
||||
|
||||
| Quality Gate | Acceptance IDs | Status | Notes |
|
||||
|---|---|---|---|
|
||||
| Separation of intent vs implementation | QG-1 across entities | Implemented | user-journey.md, schema-mapping.md, and acceptance-criteria.md are maintained per entity. |
|
||||
| Traceability from criteria to implementation | QG-2 across entities | Implemented | This matrix provides criterion-to-code anchors and current status tags. |
|
||||
| First-release constraints | QG-3 across entities | Implemented | Constraints are documented and aligned with current flows: jobs-first source upload, visible provider/model/prompt context, and system-managed lifecycle fields. |
|
||||
|
||||
## Supporting Entity Coverage
|
||||
|
||||
| Supporting Entity | Documentation | Status | Notes |
|
||||
|---|---|---|---|
|
||||
| document-person | docs/ui/entities/document-person/schema-mapping.md | Completed | Supporting-entity schema mapping created; no standalone UI contract file by design. |
|
||||
| job-source | docs/ui/entities/job-source/schema-mapping.md | Completed | Supporting-entity schema mapping created; no standalone UI contract file by design. |
|
||||
|
||||
## Suggested Implementation Order
|
||||
|
||||
1. Aggregate final acceptance review across Document, Person, Source, and Job criteria.
|
||||
|
||||
## Aggregate Final Review Snapshot (2026-08-02)
|
||||
|
||||
| Entity | Acceptance IDs still not fully met | Evidence | Notes |
|
||||
|---|---|---|---|
|
||||
| Document | None | src/transcription/ui/pages/documents_page.py; tests/ui/test_documents_page.py | Document criteria are covered by dedicated detail/edit/delete pages and document-scoped related views. |
|
||||
| Person | None | src/transcription/ui/pages/people_page.py; tests/ui/test_people_page.py | Person criteria are covered by dedicated create/detail/edit/delete pages with relationship-aware delete guardrails. |
|
||||
| Source | None | src/transcription/services/store.py; src/transcription/ui/pages/sources_page.py; src/transcription/ui/pages/jobs_page.py; tests/services/test_store.py; tests/ui/test_sources_page.py; tests/services/test_transcription_service.py | Source criteria are covered by job-context create behavior, ordering/filename policy, dedicated list/detail read flow, revision flow, and delete guardrails. |
|
||||
| Job | None | src/transcription/ui/pages/jobs_page.py; src/transcription/services/jobs.py; tests/ui/test_jobs_page.py; tests/services/test_job_service.py | Job criteria are covered by create/read/revision/lifecycle visibility and delete guardrails in dedicated routes. |
|
||||
@@ -0,0 +1,150 @@
|
||||
# Documents Page Contract
|
||||
|
||||
## Purpose
|
||||
|
||||
Documents manages the archival record for each historical artifact independently of its source files and transcription jobs. A Document can be created first, linked to people in one or more roles, and used later as the parent for Sources and Jobs.
|
||||
|
||||
## Routes
|
||||
|
||||
| Route | Purpose |
|
||||
| --- | --- |
|
||||
| `/documents` | Searchable archival Document list. |
|
||||
| `/documents/new` | Create a Document. |
|
||||
| `/documents/{document_id}` | View one Document and its related records. |
|
||||
| `/documents/{document_id}/info` | View archival metadata and system logistics for one Document. |
|
||||
| `/documents/{document_id}/edit` | Edit metadata and the complete Linked People set. |
|
||||
| `/documents/{document_id}/delete` | Confirm or block deletion. |
|
||||
| `/documents/{document_id}/jobs` | Show Jobs belonging to the Document. |
|
||||
| `/documents/{document_id}/sources` | Source-image gallery for the Document. |
|
||||
| `/documents/{document_id}/print` | Preview and browser-print the persisted Document. |
|
||||
|
||||
## List Behavior
|
||||
|
||||
- The title is **Archival Documents**.
|
||||
- **Create new document** opens the create route.
|
||||
- The table defaults to Document Title order and supports search and column sorting.
|
||||
- Columns are Document Title, Author, Tags, Document Date, Type, # Sources, and Transcription Status.
|
||||
- Document Title is left-aligned; the remaining columns are centered.
|
||||
- Author lists all linked people in the `author` role.
|
||||
- # Sources reflects the count of linked Source rows for each Document.
|
||||
- Transcription Status reflects the most recent Job status for that Document; documents with no Jobs show a blank marker.
|
||||
- Date display prefers exact date, then approximate date, then `Unknown`.
|
||||
- Selecting a row opens Document Detail.
|
||||
- Row navigation includes list context so Document Detail provides **Back to Documents**.
|
||||
- No records displays `No documents found in repository.`
|
||||
|
||||
## Create and Edit Behavior
|
||||
|
||||
Required:
|
||||
|
||||
- Document name.
|
||||
- Document type selected from the Document Type registry.
|
||||
|
||||
Optional:
|
||||
|
||||
- Exact date.
|
||||
- Approximate date.
|
||||
- Document location.
|
||||
- Archive identifier.
|
||||
- Notes.
|
||||
- Tags.
|
||||
- Linked People, with exactly one Person Role per linked Person.
|
||||
|
||||
Rules:
|
||||
|
||||
- Exact date must parse as `YYYY-MM-DD`; browser presentation may follow locale.
|
||||
- The exact-date input is labeled **Document date**.
|
||||
- Existing people appear with disambiguating labels.
|
||||
- Tag assignment supports selecting existing tags and adding new labels inline.
|
||||
- **Create new person** opens Person creation.
|
||||
- `person_id` may preselect that Person in the author role on Document creation.
|
||||
- An invalid requested Person produces a warning rather than a broken form.
|
||||
- `return_to=jobs_new` returns a successful create to Job creation with the new Document selected.
|
||||
- Edit includes active and inactive Document Types so historical values remain maintainable.
|
||||
- One Linked People table contains Select, Person, and Role columns.
|
||||
- Add and Edit use an inline Person/Role editor; Save, Cancel, and Delete change staged UI state only.
|
||||
- A Person may appear once per Document regardless of role.
|
||||
- Existing inactive-role links remain visible; only active roles may be newly assigned.
|
||||
- Document fields and the complete staged link set commit atomically on the main save.
|
||||
- Save success returns to Document Detail.
|
||||
|
||||
## Detail Behavior
|
||||
|
||||
- The heading shows name, type, and internal ID.
|
||||
- The header includes a contextual back action: **Back to Documents** by default, **Back to Person** when opened from Person Detail, and **Back to Job** when opened from Job Detail.
|
||||
- The detail workspace shows a Source-style pan/zoom media viewer with **Previous Page** / **Next Page** navigation for document source pages.
|
||||
- The center column is **Editable Revision** for the active source page.
|
||||
- Related People are grouped by role and link to Person Detail.
|
||||
- **Source Pages & Transcriptions** shows source/job counts and actions for source-image gallery, document jobs, and adding a Job.
|
||||
- **Edit Document**, **Print**, **Document Details**, **View Source Detail**, and **Delete** are available from the header.
|
||||
- Invalid IDs and missing Documents produce explicit states without rendering a partial page.
|
||||
|
||||
## Document Source Images Behavior
|
||||
|
||||
- `/documents/{document_id}/sources` shows the current Document's source pages in a thumbnail gallery.
|
||||
- Each card shows the page number, stored filename, and an **Open Source Detail** action.
|
||||
- The page includes a **Back to Document** action.
|
||||
- No source pages displays an explicit empty state.
|
||||
|
||||
## Document Info Behavior
|
||||
|
||||
- `/documents/{document_id}/info` contains **Archival Metadata** and **System Logistics**.
|
||||
- It includes a **Back to Document** action.
|
||||
- Archival metadata includes authors, document type, tags, document date, location (linked when present), archive identifier, and notes.
|
||||
|
||||
## Print Behavior
|
||||
|
||||
- Print opens a dedicated preview for persisted Document data.
|
||||
- **Facsimile** places each Source image beside its current transcription and starts every Source on a new printed sheet.
|
||||
- **Text only** omits images, joins single line breaks inside paragraphs, and preserves blank-line paragraph boundaries.
|
||||
- Non-null revised text takes precedence over raw transcription, including an intentionally empty revision.
|
||||
- Archival metadata resolves Author through the hidden built-in semantic identity, not its mutable label.
|
||||
- Archival metadata includes the Document Type label.
|
||||
- Metadata tables use a narrow non-wrapping label column and wider wrapping data columns rather than stretching across the page.
|
||||
- Job metadata uses one oldest-to-newest column per Job and ends with Status.
|
||||
- Stored text is escaped and Source media uses record-validated application URLs rather than local file paths.
|
||||
- Printing uses the browser print dialog; server-generated PDFs are not provided.
|
||||
|
||||
## Document Jobs Behavior
|
||||
|
||||
- The page lists the Document's Jobs newest first with status and Job ID.
|
||||
- **Open Job** navigates to Job Detail.
|
||||
- **Create Job** opens Job creation with the Document selected.
|
||||
- No jobs displays an explicit empty state.
|
||||
|
||||
## Delete Behavior
|
||||
|
||||
- Deletion is blocked while any Source or Job belongs to the Document.
|
||||
- The blocked state names the dependency categories and provides navigation back and to Jobs.
|
||||
- An unlinked Document requires an explicit permanent-delete action.
|
||||
- Success returns to the Documents list.
|
||||
|
||||
## Acceptance Checklist
|
||||
|
||||
- List columns, alignment, search, sorting, date fallback, and row navigation match this contract.
|
||||
- Create/edit enforce name, registered type, and valid exact-date input.
|
||||
- Linked People staging enforces one role and one row per Person.
|
||||
- Document and Linked People writes never partially commit.
|
||||
- Person-first Document creation preselects the requested Person as author.
|
||||
- Detail links people, Sources, and Jobs to the correct records.
|
||||
- Delete never removes a Document with Source or Job dependencies.
|
||||
- Both print formats preserve the frozen content, ordering, text-precedence, and safety contracts.
|
||||
- Service failures use the shared error presenter and never report false success.
|
||||
|
||||
## Implementation Anchors
|
||||
|
||||
- `src/transcription/ui/pages/documents_page.py`
|
||||
- `src/transcription/ui/components/table/documents.py`
|
||||
- `src/transcription/services/documents.py`
|
||||
- `src/transcription/services/people.py`
|
||||
- `src/transcription/services/workflows.py`
|
||||
- `src/transcription/ui/components/linked_people.py`
|
||||
- `src/transcription/ui/pages/print_preview_page.py`
|
||||
- `src/transcription/api/print_api.py`
|
||||
- `tests/ui/test_documents_page.py`
|
||||
- `tests/services/test_document_service.py`
|
||||
|
||||
## Known Limitations and Deferred Work
|
||||
|
||||
- Source page ordering remains read-only.
|
||||
- Printing other entities, batch printing, and server-side export formats are deferred.
|
||||
@@ -0,0 +1,64 @@
|
||||
# Home Page Contract
|
||||
|
||||
## Purpose
|
||||
|
||||
Home provides a user-maintained landing page for the local archive. It combines a database-backed image gallery with Markdown text and lets the operator edit both without changing application source or prompt assets.
|
||||
|
||||
## Routes
|
||||
|
||||
| Route | Browser path | Purpose |
|
||||
| --- | --- | --- |
|
||||
| `/homepage` | `/ui/homepage` | View homepage gallery and Markdown. |
|
||||
| `/homepage/edit` | `/ui/homepage/edit` | Upload images, manage image metadata, and edit Markdown. |
|
||||
|
||||
The application root and `/ui` redirect to `/ui/homepage`.
|
||||
|
||||
## View Behavior
|
||||
|
||||
- The visible page heading is **Home**; the browser tab title is **VibeScribe Home**.
|
||||
- The featured homepage image (`photo.is_primary`) is shown first; remaining images are shown in random order.
|
||||
- The current image appears in the shared dark-room viewer with its description.
|
||||
- Saved Markdown is rendered in the **Home Text** card.
|
||||
- Missing text displays `No homepage text saved yet.`
|
||||
- Missing image displays the viewer's empty state.
|
||||
- **Edit Home Page** opens the edit route.
|
||||
- The same Home Text content is also editable from **Settings → Home Page Text**.
|
||||
|
||||
## Edit Behavior
|
||||
|
||||
- The image upload accepts JPEG, PNG, GIF, WebP, BMP, and TIFF files and supports multi-file uploads.
|
||||
- A successful upload immediately stores files in the shared `photo` table/media layout and displays a positive notification.
|
||||
- The editor supports per-image description edits, setting a featured image, and deleting the current image.
|
||||
- The Markdown textarea is initialized from the currently stored homepage text.
|
||||
- **Save** writes the textarea content, displays `Homepage saved`, and returns to Home.
|
||||
- **Cancel** returns to Home without saving textarea changes. An image already uploaded during the edit session remains stored.
|
||||
|
||||
## Storage Contract
|
||||
|
||||
- Homepage markdown text is mutable application data at `UPLOAD_DIR/homepage.md`.
|
||||
- Homepage images are stored as `photo` rows (`person_id = NULL`) with files under `UPLOAD_DIR/photos/`.
|
||||
- Uploaded images are renamed to `{photo_id}{suffix}`.
|
||||
- Homepage images are database records; markdown remains file-backed.
|
||||
|
||||
## Acceptance Checklist
|
||||
|
||||
- `/`, `/ui`, and the application brand reach Home.
|
||||
- Home renders with or without stored Markdown and image content.
|
||||
- Edit loads existing Markdown.
|
||||
- Supported image upload stores one or more images and makes the first image featured when no featured image exists yet.
|
||||
- Save persists Markdown and returns to Home.
|
||||
- Cancel does not save changed Markdown.
|
||||
|
||||
## Implementation Anchors
|
||||
|
||||
- `src/transcription/ui/pages/home_page.py`
|
||||
- `src/transcription/ui/homepage_store.py`
|
||||
- `src/transcription/ui/components/app_shell.py`
|
||||
- `tests/ui/test_upload_page.py`
|
||||
- `tests/ui/test_navigation_and_mounts.py`
|
||||
- `tests/ui/test_pages_registration.py`
|
||||
|
||||
## Known Limitations
|
||||
|
||||
- Homepage markdown storage location is `UPLOAD_DIR/homepage.md` and must remain writable in the active runtime environment.
|
||||
- Uploading an image is immediate and is not rolled back by Cancel.
|
||||
@@ -0,0 +1,99 @@
|
||||
# Jobs Page Contract
|
||||
|
||||
## Purpose
|
||||
|
||||
Jobs manages transcription processing runs. A Job belongs to one Document, links one or more Source pages, records processing provenance, and exposes lifecycle actions without making lifecycle fields directly editable.
|
||||
|
||||
## Routes
|
||||
|
||||
| Route | Purpose |
|
||||
| --- | --- |
|
||||
| `/jobs` | Searchable processing Job list. |
|
||||
| `/jobs/new` | Create and queue a Job. |
|
||||
| `/jobs/{job_id}` | View status, execution logistics, and related records. |
|
||||
| `/jobs/{job_id}/cancel` | Confirm cancellation. |
|
||||
| `/jobs/{job_id}/resubmit` | Confirm resubmission of failed Sources. |
|
||||
| `/jobs/{job_id}/delete` | Confirm or block deletion. |
|
||||
|
||||
## List Behavior
|
||||
|
||||
- The title is **Transcription Pipeline Jobs**.
|
||||
- **Create job** opens Job creation and **Refresh** reloads the table.
|
||||
- Columns are Job ID, Status, Document Name, # Sources, Retries, and Updated.
|
||||
- Updated is the primary date/sort field.
|
||||
- Search covers Job ID, document name, and status.
|
||||
- Status is displayed as a semantic status chip.
|
||||
- Selecting a row opens Job Detail.
|
||||
- Global Job-list row navigation includes list context so Job Detail provides **Back to Jobs**.
|
||||
- No records displays `No job records found in repository.`
|
||||
|
||||
## Create Behavior
|
||||
|
||||
- A Target Document and at least one source file are required.
|
||||
- `document_id` may preselect a Target Document.
|
||||
- If no Documents exist, the page explains the prerequisite and links to Document creation with a return path.
|
||||
- Provider and Model are selectable when creating a new Job.
|
||||
- Upload accepts JPEG, PNG, TIFF, and PDF files and supports multiple/folder selection.
|
||||
- The visible upload queue is sorted alphabetically by original filename.
|
||||
- Files can be removed individually or cleared before submission.
|
||||
- Helper text explains numeric filename prefixes for page ordering.
|
||||
- Submission creates the Job, Source records, and JobSource links, notifies the worker, and opens Job Detail.
|
||||
- When opened with `source_id`, creation becomes a retranscription flow: Source and Document are locked, Provider is
|
||||
read-only, Model is restricted to `PROVIDER_MODELS`, no upload is accepted, and one existing Source is linked.
|
||||
|
||||
## Detail and Lifecycle Behavior
|
||||
|
||||
- The heading shows Job ID and a status badge.
|
||||
- Job Detail includes a contextual back action: **Back to Jobs** by default and **Back to Document** when opened from a Document-filtered Job list.
|
||||
- Execution Logistics shows provider, model, prompt, retry count, and last update.
|
||||
- Document Links show a clickable Document Name (with Job context), Sources count, and a single **View Sources** action using job filtering.
|
||||
- Queued and processing Jobs show an auto-refresh notice and reload every four seconds.
|
||||
- Polling stops when the Job becomes terminal or a refresh fails.
|
||||
- Queued and processing Jobs expose **Cancel**.
|
||||
- Jobs other than `transcribed` expose **Resubmit** under the current UI rule. The service blocks resubmission while processing is active or when no failed Sources exist.
|
||||
- All Jobs expose **Delete Job**, subject to explicit evidence-deletion guardrails.
|
||||
- Invalid and missing IDs produce explicit states.
|
||||
|
||||
## Cancel Behavior
|
||||
|
||||
- The confirmation explains that processing stops and remaining pending Sources become cancelled.
|
||||
- The service decides whether the current state permits cancellation.
|
||||
- Success updates the Job, notifies the worker, and returns to Job Detail.
|
||||
|
||||
## Resubmit Behavior
|
||||
|
||||
- The page shows current status and failed Source count.
|
||||
- The page explains that resubmission queues failed linked Sources while preserving immutable prior attempt evidence.
|
||||
- The service blocks submission while processing is active or when no failed Sources exist.
|
||||
- `JobSource` remains the latest compatibility projection, while every provider call appends an `ExecutionAttempt`.
|
||||
- The selected `Source.raw_transcription` projection remains available while a retry is pending or fails.
|
||||
- Success reports the number of resubmitted Sources and returns to Job Detail.
|
||||
|
||||
## Delete Behavior
|
||||
|
||||
- Deletion is blocked while status is `processing`.
|
||||
- Allowed deletion explicitly warns that related `JobSource` projections,
|
||||
immutable execution attempts, captured transport responses, and attempt-owned
|
||||
artifacts are permanently removed.
|
||||
- Source records and source files remain available for separate deletion.
|
||||
- Success returns to the Jobs list.
|
||||
|
||||
## Acceptance Checklist
|
||||
|
||||
- Job creation cannot proceed without a valid Document and at least one Source.
|
||||
- Upload ordering and removal controls match the displayed queue.
|
||||
- Detail shows current status and provenance summary with correct related links.
|
||||
- Active Jobs refresh without overlapping permanent polling after terminal state.
|
||||
- Cancel, resubmit, and delete honor service guardrails and show actionable failures.
|
||||
- Lifecycle fields cannot be edited directly.
|
||||
|
||||
## Implementation Anchors
|
||||
|
||||
- `src/transcription/ui/pages/jobs_page.py`
|
||||
- `src/transcription/ui/components/table/jobs.py`
|
||||
- `src/transcription/services/jobs.py`
|
||||
- `src/transcription/services/store.py`
|
||||
- `src/transcription/services/workflows.py`
|
||||
- `tests/ui/test_jobs_page.py`
|
||||
- `tests/services/test_job_service.py`
|
||||
- `tests/services/test_store.py`
|
||||
@@ -0,0 +1,106 @@
|
||||
# People Page Contract
|
||||
|
||||
## Purpose
|
||||
|
||||
People manages reusable historical-person records. A Person may appear in many Documents under different relationship roles and may optionally carry one or more photos plus a FamilySearch identifier.
|
||||
|
||||
## Routes
|
||||
|
||||
| Route | Purpose |
|
||||
| --- | --- |
|
||||
| `/people` | Searchable People list. |
|
||||
| `/people/new` | Create a Person. |
|
||||
| `/people/{person_id}` | View one Person and linked Documents. |
|
||||
| `/people/{person_id}/photos` | Manage Person photos. |
|
||||
| `/people/{person_id}/edit` | Edit the Person. |
|
||||
| `/people/{person_id}/delete` | Confirm permanent deletion. |
|
||||
|
||||
## List Behavior
|
||||
|
||||
- The title is **Archival Entities: People**.
|
||||
- **Create new person** opens the create route.
|
||||
- The table defaults to Name order (`Last Name, First & Middle`) and supports search and column sorting.
|
||||
- Columns are Last Name, First & Middle; Tags; FamilySearch ID; Birth Date; Death Date; and # Documents.
|
||||
- Name and Tags are left-aligned; FamilySearch ID, date columns, and # Documents are centered.
|
||||
- # Documents reflects how many linked Documents each Person is connected to.
|
||||
- Birth and death values independently prefer exact date, then approximate date, then `Unknown`.
|
||||
- Selecting a row opens Person Detail.
|
||||
- Row navigation includes list context so Person Detail provides **Back to People**.
|
||||
- No records displays `No person records found in repository.`
|
||||
|
||||
## Create and Edit Behavior
|
||||
|
||||
Required:
|
||||
|
||||
- Last name.
|
||||
- First & middle names.
|
||||
|
||||
Optional:
|
||||
|
||||
- Exact and approximate birth/death dates.
|
||||
- Birth/death places.
|
||||
- Biography.
|
||||
- FamilySearch ID.
|
||||
- Tags.
|
||||
|
||||
Rules:
|
||||
|
||||
- Missing last name or first/middle names blocks save with a warning.
|
||||
- Exact date inputs are native browser date inputs.
|
||||
- FamilySearch IDs are normalized and validated by `PeopleService`.
|
||||
- Tags use the shared Tag registry and support inline add/select behavior.
|
||||
- Photos are managed from Person Detail via `/people/{person_id}/photos` (not in create/edit form fields).
|
||||
- Metadata JSON remains hidden.
|
||||
- Save success returns to Person Detail.
|
||||
|
||||
## Detail Behavior
|
||||
|
||||
- The header provides **New Document**, **Edit Person**, **Edit Photo(s)**, and **Delete**.
|
||||
- The header includes a contextual back action: **Back to People** by default, and **Back to Document** when opened from Document Detail.
|
||||
- **New Document** opens Document creation with this Person requested for author preselection.
|
||||
- Person Detail shows a single-photo viewer with **Previous/Next** navigation; the page-level **Edit Photo(s)** header action opens photo management.
|
||||
- Photo management (upload, description edit, set-primary, delete) is intentionally moved to `/people/{person_id}/photos`.
|
||||
- Biographical Record shows split names, computed full name, tags, compact birth/death dates, and places.
|
||||
- Birth and death place values are clickable links to Google Maps when present.
|
||||
- FamilySearch ID is shown as a metadata value and is clickable to the FamilySearch person details route when present.
|
||||
- Biography has an explicit empty value.
|
||||
- Linked Documents render as a table with **Document Name**, **Document Date**, **Role**, and **Number of Pages**; selecting a row opens Document Detail.
|
||||
- No links shows both an empty state and guidance to link from a Document workflow.
|
||||
- System Logistics shows created and updated timestamps.
|
||||
|
||||
## Delete Behavior
|
||||
|
||||
- The page warns when linked Document relationships exist.
|
||||
- Delete is blocked when related Photos exist.
|
||||
- Confirmed deletion removes the Person and its relationship links; it does not delete Documents.
|
||||
- Success returns to the People list.
|
||||
- Missing or already-deleted records return to a safe list state.
|
||||
|
||||
## Photo Gallery Behavior (`/people/{person_id}/photos`)
|
||||
|
||||
- Upload is triggered from a header-level **Upload Photo(s)** control beside **Back to Person**.
|
||||
- The gallery renders all photos in a responsive grid (3-4 tiles wide on larger screens).
|
||||
- Description text is shown as an overlay at the bottom of each image for quick context.
|
||||
- The editor provides **Save Description**, **Set Primary** (when applicable), and **Delete Photo** actions.
|
||||
|
||||
## Acceptance Checklist
|
||||
|
||||
- List fields, alignment, date fallback, search, sorting, and navigation match this contract.
|
||||
- Last name and first/middle names are enforced on create and edit.
|
||||
- FamilySearch ID validation and link generation use the fixed supported identifier format.
|
||||
- Photo upload and rendering remain constrained to supported media paths.
|
||||
- New Document carries the Person context.
|
||||
- Linked Documents show the correct role and target.
|
||||
- Delete wording distinguishes removal of relationship links from deletion of Documents.
|
||||
|
||||
## Implementation Anchors
|
||||
|
||||
- `src/transcription/ui/pages/people_page.py`
|
||||
- `src/transcription/ui/components/table/people.py`
|
||||
- `src/transcription/services/people.py`
|
||||
- `tests/ui/test_people_page.py`
|
||||
- `tests/services/test_v2_crud.py`
|
||||
|
||||
## Deferred Work
|
||||
|
||||
- Structured name fields, merge/deduplication, advanced metadata editing, and Person-side relationship editing are not current behavior.
|
||||
@@ -0,0 +1,55 @@
|
||||
# Settings Page Contract
|
||||
|
||||
## Purpose
|
||||
|
||||
Settings manages installation-local registries, safe runtime .env settings, and editable text assets from one route.
|
||||
|
||||
## Route
|
||||
|
||||
| Route | Purpose |
|
||||
| --- | --- |
|
||||
| `/settings` | Manage Runtime Settings, Document Types, Person Roles, Tags, Prompts, Home Page Text, and Maintenance runs. |
|
||||
|
||||
## Behavior
|
||||
|
||||
- The page title is **Settings**.
|
||||
- Configuration surfaces are grouped as tabs:
|
||||
- **Document Types**
|
||||
- **Person Roles**
|
||||
- **Tags**
|
||||
- **Prompts**
|
||||
- **Home Page Text**
|
||||
- **Maintenance**
|
||||
- **Runtime Settings**
|
||||
- Runtime Settings exposes an allowlisted set of non-secret fields synchronized with `Settings` model fields except excluded secret/unsafe fields.
|
||||
- Runtime Settings is rendered as a compact two-column editor (**Setting**, **Value**) in a centered, narrower responsive container.
|
||||
- Runtime Settings persists changes to the resolved runtime env file, validates by constructing a `Settings` instance, and reports validation failures through the shared UI error presenter.
|
||||
- `Settings` resolves its env file in this order: explicit `_env_file`, `ENV_FILE`, then the repository-root `.env.production`.
|
||||
- Runtime Settings resolves its write target in this order: explicit function override (tests/tools), `RUNTIME_SETTINGS_ENV_FILE` environment variable (deployment override), `ENV_FILE`, then the repository-root `.env.production`.
|
||||
- Runtime Settings changes require application restart to take effect.
|
||||
- Runtime Settings renders a host-side restart command (`docker compose -f docker-compose.production.yml up -d --force-recreate app worker`) so operators can apply saved values without granting Docker control to the app container.
|
||||
- Runtime Settings includes an explicit "Other settings not shown here" markdown table listing:
|
||||
- secrets (`OPENROUTER_API_KEY`, `DATABASE__PASSWORD`)
|
||||
- high-risk database connection settings (`DATABASE__DRIVER`, `DATABASE__PATH`, `DATABASE__HOST`, `DATABASE__PORT`, `DATABASE__DATABASE`, `DATABASE__USER`)
|
||||
and deployment/helper keys (`CLOUDFLARE_TUNNEL_TOKEN`, `BACKUP_DIR`, `BACKUP_RETENTION_DAYS`, `RUNTIME_SETTINGS_ENV_FILE`, `ENV_FILE`, `COMPOSE_FILE`) plus legacy/deprecated keys (`POSTGRES_*`, `DATABASE_BACKUP_DIR`, `APP_DATA_BACKUP_DIR`, `UPLOADS_BACKUP_DIR`, `SYNOLOGY_BACKUP_DIR`), and directs edits for those keys to the resolved runtime env file path.
|
||||
- Document Types, Person Roles, and Tags support Add/Edit/Delete with existing guardrails.
|
||||
- Prompts exposes only `transcribe_document.md` for editing and restore-from-backup.
|
||||
- Home Page Text edits the same Markdown content rendered on `/homepage`.
|
||||
- Maintenance provides queue-backed **Run Backup** and **Run Storage Reconciliation** actions.
|
||||
- Maintenance also provides GEDCOM upload and **Run GEDCOM Import** actions, using the same queue-backed `MaintenanceRun` history/log flow.
|
||||
- Maintenance run history shows job type, status, started/finished timestamps, duration, summary, and log view/download actions.
|
||||
- Maintenance actions enqueue work and signal the worker; the page itself does not execute shell commands directly.
|
||||
|
||||
## Acceptance Checklist
|
||||
|
||||
- `/ui/settings` renders all seven tabs.
|
||||
- Registry and prompt workflows keep existing validation and error handling.
|
||||
- Runtime Settings excludes secret fields and rejects invalid values.
|
||||
- Saving Home Page Text persists content for the homepage view.
|
||||
|
||||
## Implementation Anchors
|
||||
|
||||
- `src/transcription/ui/pages/settings_page.py`
|
||||
- `src/transcription/ui/runtime_settings_store.py`
|
||||
- `src/transcription/ui/homepage_store.py`
|
||||
- `tests/ui/test_pages_registration.py`
|
||||
@@ -0,0 +1,99 @@
|
||||
# Sources Page Contract
|
||||
|
||||
## Purpose
|
||||
|
||||
Sources manages individual archived page/file records. It provides source-media viewing, current processing context, provider evidence inspection, previous/next page navigation, and human revision without allowing machine output to be edited.
|
||||
|
||||
## Routes
|
||||
|
||||
| Route | Purpose |
|
||||
| --- | --- |
|
||||
| `/sources` | Document-filtered or Job-filtered Source list; global route redirects to Documents. |
|
||||
| `/sources/{source_id}` | View media, transcription, revision, metadata, and evidence. |
|
||||
| `/sources/{source_id}/delete` | Confirm or block deletion. |
|
||||
|
||||
The list accepts optional `document_id` and `job_id` query parameters. Document context takes precedence if both parse successfully.
|
||||
|
||||
## List Behavior
|
||||
|
||||
- The global `/sources` route redirects to `/documents`.
|
||||
- Filtered list titles are **Sources for Document** and **Sources for Job**.
|
||||
- Filtered context provides **Back to Document** or **Back to Job**.
|
||||
- Rows are ordered by page number and then upload name.
|
||||
- Columns are Upload Title, Page Number, Document Name, Status, and Error Detail.
|
||||
- Document Name, Upload Title, and Error Detail are left-aligned; Status is centered.
|
||||
- Status labels are presented in uppercase for consistency with Jobs.
|
||||
- Stored Filename is intentionally absent from the list.
|
||||
- Selecting a row opens Source Detail.
|
||||
- No records displays `No source asset records found in repository.`
|
||||
|
||||
## Detail Behavior
|
||||
|
||||
- The heading shows page number, upload name, and Source ID.
|
||||
- **Back to Document** returns to Document Detail for the active source page.
|
||||
- **Retranscribe Source** opens Create Processing Job with this Source and its Document locked.
|
||||
- **Delete Source** opens the guarded delete route.
|
||||
- Previous and Next navigate only among Sources belonging to the same Document in page order; unavailable boundary actions are disabled.
|
||||
- The media viewer resolves the stored Source path through the configured upload root.
|
||||
- The top layout is adaptive:
|
||||
- Standard pages use three columns with a wider Editable Revision column than the image column.
|
||||
- Wide+narrow landscape images switch to a stacked left layout (image above Editable Revision) with metadata on the right.
|
||||
- Editable Revision is seeded from an existing revision or the preferred machine transcription.
|
||||
- Source Metadata shows upload name, stored filename, page number, Document Name, Document ID, and stored path. Source ID appears in the page-header subtitle.
|
||||
- SourceJob Metadata shows latest status (uppercase display), Job ID, execution time, provider, model, prompt, and failure detail.
|
||||
- Revision Logistics shows revised state, last-revised time, and upload time.
|
||||
- Candidate Machine Transcriptions appears below the image/revision area, remains compact until expanded, then compares it with the preferred
|
||||
machine result and requires confirmation before **Use this transcription**.
|
||||
- Candidate promotion does not alter a human revision. Empty states distinguish no machine result from no candidates.
|
||||
- An orientation-normalized artifact appears in evidence only when recognized metadata required a physical rotation.
|
||||
|
||||
## Provider Evidence
|
||||
|
||||
- Provider Evidence is associated with the latest JobSource execution.
|
||||
- New attempts display separate expandable Request Manifest, Transport Response, OpenRouter SDK Response Snapshot,
|
||||
Normalized Metadata, Software Context, and Derived Artifacts sections.
|
||||
- Historical `raw_api_response` values are labeled as OpenRouter SDK response snapshots.
|
||||
- Missing evidence has an explicit empty state.
|
||||
- Historical executions explicitly state that exact transport evidence was not captured.
|
||||
- Quality warning artifacts remain attached to their machine attempt and are not recomputed during page rendering.
|
||||
- **Export Evidence** downloads a versioned package containing source identity, attempts, artifacts, relationships,
|
||||
schema versions, and integrity digests without source binaries, credentials, or machine-local source paths.
|
||||
|
||||
## Revision Behavior
|
||||
|
||||
- Machine transcription is never edited directly.
|
||||
- A revision must contain non-whitespace text.
|
||||
- Save persists revised text and updates the saved timestamp without leaving the page.
|
||||
- Reset restores the in-memory revision from page load or the most recent successful save. When no revision exists, it restores the machine transcription; it does not re-read the database.
|
||||
- A failed latest execution displays guidance that a human revision can preserve corrected text.
|
||||
|
||||
## Delete Behavior
|
||||
|
||||
- Deletion is allowed only when the Source has no JobSource links.
|
||||
- A linked Source shows cleanup guidance and navigation to Jobs.
|
||||
- An unlinked Source requires explicit permanent deletion.
|
||||
- Success returns to the Sources list.
|
||||
|
||||
## Acceptance Checklist
|
||||
|
||||
- Global, Document-filtered, and Job-filtered lists show the correct context and return action.
|
||||
- List columns and alignments match this contract and omit Stored Filename.
|
||||
- Previous/next navigation never crosses Document boundaries.
|
||||
- Detail keeps machine output read-only and human revision separately editable.
|
||||
- Retranscription, candidate comparison, warnings, and explicit promotion preserve every prior attempt.
|
||||
- Empty, failed, and missing-evidence states remain explicit.
|
||||
- JSON evidence is readable without being mislabeled as native transport evidence.
|
||||
- Delete cannot remove a Source with processing-history links.
|
||||
|
||||
## Implementation Anchors
|
||||
|
||||
- `src/transcription/ui/pages/sources_page.py`
|
||||
- `src/transcription/ui/components/table/sources.py`
|
||||
- `src/transcription/services/sources.py`
|
||||
- `tests/ui/test_sources_page.py`
|
||||
- `tests/services/test_transcription_service.py`
|
||||
- `tests/services/test_v2_crud.py`
|
||||
|
||||
## Planned Changes
|
||||
|
||||
- Source page reordering remains deferred unless a demonstrated workflow need emerges.
|
||||
@@ -1,308 +0,0 @@
|
||||
# System Architecture (Version 1)
|
||||
|
||||
This document describes the production architecture of the personal historical-document transcription system. The system is intentionally optimized for single-user operation, low operational overhead, and clean internal boundaries that support future growth without rewrites.
|
||||
|
||||
## Architecture Objectives
|
||||
|
||||
The production architecture is designed to:
|
||||
|
||||
- preserve verbatim family-history source material as searchable text
|
||||
- keep operational complexity low for a personal deployment
|
||||
- support asynchronous transcription without requiring distributed infrastructure
|
||||
- maintain clear module boundaries so extensions can be added incrementally
|
||||
|
||||
## Production Scope And Scale
|
||||
|
||||
The deployed system targets personal use and a corpus of several thousand documents processed over time. The architecture favors simple, composable building blocks over distributed orchestration.
|
||||
|
||||
Current scope includes:
|
||||
|
||||
- content source upload and metadata capture
|
||||
- asynchronous transcription jobs
|
||||
- prompt-library driven transcription behavior, with one Markdown file per prompt
|
||||
- original transcription review and optional revision review
|
||||
- full-text search over accepted transcripts
|
||||
- export of transcript data
|
||||
|
||||
## Deployment Topology
|
||||
|
||||
The production deployment uses [Docker Compose](https://docs.docker.com/compose/) and treats containerized databases as extremely lightweight operational dependencies.
|
||||
|
||||
Running [PostgreSQL](https://www.postgresql.org/docs/) in its own container is considered simple by default for this system.
|
||||
|
||||
Running [MongoDB](https://www.mongodb.com/docs/) in its own container is also considered simple when document-centric storage is enabled.
|
||||
|
||||
Container count is not a hard architectural limit; a three-container deployment (app, PostgreSQL, MongoDB) is an acceptable baseline.
|
||||
|
||||
### Baseline Topology (Two Containers)
|
||||
|
||||
- one application container
|
||||
- one PostgreSQL container
|
||||
- embedded background worker execution inside the app process
|
||||
|
||||
### Expanded Topology (Three Containers)
|
||||
|
||||
- application container
|
||||
- PostgreSQL container
|
||||
- MongoDB container
|
||||
|
||||
No additional queue, scheduler, or search-engine containers are required in the baseline production setup.
|
||||
|
||||
## Runtime Architecture
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
User[Browser User] --> App[FastAPI + NiceGUI Service]
|
||||
App --> Worker[In-process Background Worker]
|
||||
App --> PG[(PostgreSQL)]
|
||||
App --> MG[(MongoDB Document Store)]
|
||||
Worker --> AI[Transcription Provider]
|
||||
Worker --> PG
|
||||
Worker --> MG
|
||||
```
|
||||
|
||||
## Runtime Ownership And Startup Policy
|
||||
|
||||
The current implementation now uses explicit lifespan-owned runtime resources.
|
||||
|
||||
- application lifespan initializes and disposes database runtime resources
|
||||
- worker lifecycle is owned by application lifespan startup/shutdown
|
||||
- worker receives lifespan-owned database engine dependency explicitly
|
||||
- schema bootstrap policy is environment-aware and explicit:
|
||||
- development/test default to bootstrap enabled
|
||||
- production defaults to bootstrap disabled
|
||||
- explicit override is available via configuration
|
||||
|
||||
This aligns implementation toward REQ-7 and REQ-10 while preserving personal-scale operational simplicity.
|
||||
|
||||
## Layered Module Structure
|
||||
|
||||
### Interface Layer
|
||||
|
||||
Responsibility:
|
||||
|
||||
- HTTP API and UI routes
|
||||
- request/response validation
|
||||
- status and result presentation
|
||||
|
||||
Out of scope:
|
||||
|
||||
- business-rule enforcement
|
||||
- data-access implementation
|
||||
|
||||
### Application Layer
|
||||
|
||||
Responsibility:
|
||||
|
||||
- upload and job orchestration
|
||||
- state transitions and retry policy
|
||||
- coordination across domain and infrastructure ports
|
||||
|
||||
Out of scope:
|
||||
|
||||
- provider-specific protocol details
|
||||
- ORM or storage-specific logic
|
||||
|
||||
### Domain Layer
|
||||
|
||||
Responsibility:
|
||||
|
||||
- verbatim transcription policy
|
||||
- revision and provenance invariants
|
||||
- confidence and annotation semantics
|
||||
|
||||
Out of scope:
|
||||
|
||||
- web framework concerns
|
||||
- database and network I/O
|
||||
|
||||
### Infrastructure Layer
|
||||
|
||||
Responsibility:
|
||||
|
||||
- persistence adapters (PostgreSQL and MongoDB)
|
||||
- transcription-provider adapter
|
||||
|
||||
Out of scope:
|
||||
|
||||
- business policy decisions
|
||||
|
||||
## Processing Workflow
|
||||
|
||||
Production transcription flow:
|
||||
|
||||
1. A user uploads one or more content sources through the UI or API.
|
||||
2. The application validates payloads and creates document, source, and job records.
|
||||
3. The in-process worker de-queues the job and calls the transcription provider.
|
||||
4. The application persists original transcription output on the job, plus confidence metadata and provenance events.
|
||||
5. Job status transitions from queued to processing to transcribed or failed.
|
||||
6. The UI and API expose status, optional revision to original transcription, and searchable transcription text.
|
||||
|
||||
## Data Model Ownership
|
||||
|
||||
System-of-record entities:
|
||||
|
||||
- documents and content sources
|
||||
- transcription jobs, original transcription, and status events
|
||||
- transcript revisions
|
||||
- provenance metadata
|
||||
|
||||
### Original Transcription And Revision Ownership
|
||||
|
||||
- each processing job stores the original immutable provider output (`text`)
|
||||
- provider metadata (`provider`, `model`, `prompt_name`) and failure detail (`error_detail`) are job-owned processing artifacts
|
||||
- revisions are optional user-authored edits linked to a content source
|
||||
- a revision can be created from original `job.text`
|
||||
- many jobs will have zero revisions; revisions are additive and never overwrite original provider output
|
||||
- a document groups one or more content sources (images, PDFs, and future source types)
|
||||
|
||||
Storage strategy:
|
||||
|
||||
- PostgreSQL for relational system-of-record entities
|
||||
- MongoDB for document-oriented payloads and large transcription artifacts
|
||||
- versioned prompt artifacts stored as individual Markdown files for human editing and refinement
|
||||
- in-memory execution state treated as ephemeral
|
||||
|
||||
## Transcription Prompt Asset Policy
|
||||
|
||||
The production system treats transcription prompts as maintainable content assets.
|
||||
|
||||
- each transcription prompt is stored in its own Markdown file
|
||||
- prompt files are designed for direct human editing and iterative refinement
|
||||
- prompt updates are independent and do not require bundling unrelated prompt changes
|
||||
- prompt file identity and revision history are tracked through normal repository version control
|
||||
|
||||
## Simplicity Guardrails
|
||||
|
||||
The production system enforces these constraints to prevent accidental over-engineering:
|
||||
|
||||
- PostgreSQL in a container is treated as a lightweight default dependency
|
||||
- MongoDB in a container is treated as a lightweight optional dependency
|
||||
- three containers (app, PostgreSQL, MongoDB) is an acceptable simple deployment
|
||||
- no dedicated queue or search cluster is introduced without measured need
|
||||
- external infrastructure is added only behind existing ports/adapters
|
||||
|
||||
## Extension Path
|
||||
|
||||
The architecture supports additive growth without changing domain contracts.
|
||||
|
||||
### Stage 1: Foundation (Current)
|
||||
|
||||
- upload, transcription, review, search, export
|
||||
- in-process worker execution
|
||||
- single provider adapter
|
||||
- app plus PostgreSQL deployment
|
||||
|
||||
### Stage 2: Throughput Hardening
|
||||
|
||||
- optional MongoDB document-store enablement
|
||||
- optional external worker/queue process
|
||||
- stronger retry and dead-letter handling
|
||||
|
||||
### Stage 3: Intelligence Features
|
||||
|
||||
- entity extraction and cross-document linking
|
||||
- timeline and narrative assembly
|
||||
- optional multi-provider routing
|
||||
|
||||
Each stage preserves existing module boundaries and keeps migration risk low.
|
||||
|
||||
## Test Strategy
|
||||
|
||||
The test strategy is aligned to personal-scale operation with fast, deterministic feedback.
|
||||
|
||||
### Unit Tests
|
||||
|
||||
- domain transcription rules and annotation behavior
|
||||
- revision-history invariants
|
||||
- job state-transition logic
|
||||
|
||||
### Integration Tests
|
||||
|
||||
- repository behavior and transaction boundaries
|
||||
- persistence-adapter and provider adapter contract mapping
|
||||
- upload-to-persistence roundtrip
|
||||
|
||||
### End-to-End Tests
|
||||
|
||||
- happy path: upload, transcribe, review, search, export
|
||||
- failure path: provider error, retry, surfaced failed status
|
||||
|
||||
### CI Execution Model
|
||||
|
||||
- fast suite on each push
|
||||
- optional slower provider-sandbox checks on scheduled runs
|
||||
|
||||
## Risks And Controls
|
||||
|
||||
### Runtime Responsiveness
|
||||
|
||||
Risk:
|
||||
|
||||
- long jobs can reduce responsiveness in a single-process deployment
|
||||
|
||||
Control:
|
||||
|
||||
- bounded concurrency and visible job status in the UI
|
||||
|
||||
### Database Concurrency Limits
|
||||
|
||||
Risk:
|
||||
|
||||
- contention can appear under sustained concurrent writes in personal-scale infrastructure
|
||||
|
||||
Control:
|
||||
|
||||
- tuned connection pooling and phased use of MongoDB for document-heavy workloads
|
||||
|
||||
### Provider Output Variance
|
||||
|
||||
Risk:
|
||||
|
||||
- transcription quality varies by content source type, handwriting legibility, and source quality
|
||||
|
||||
Control:
|
||||
|
||||
- first-class human review and immutable revision history
|
||||
|
||||
---
|
||||
|
||||
## Technology References
|
||||
|
||||
- [FastAPI documentation](https://fastapi.tiangolo.com/)
|
||||
- [NiceGUI documentation](https://nicegui.io/documentation)
|
||||
- [Docker Compose documentation](https://docs.docker.com/compose/)
|
||||
- [PostgreSQL documentation](https://www.postgresql.org/docs/)
|
||||
- [MongoDB documentation](https://www.mongodb.com/docs/)
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v1.md)
|
||||
- [System Design Intent](intent.md)
|
||||
- [Transcription Methodology](transcription_methodology.md)
|
||||
- System Architecture (this document)
|
||||
- [System Requirements](requirements_v1.md)
|
||||
- [Data model](schema_v1.md)
|
||||
- [Error Handling Policy](error_handling_v1.md)
|
||||
- [Implementation Plan](implementation_plan_v1.md)
|
||||
|
||||
## Glossary
|
||||
|
||||
- Adapter: A component that translates between internal interfaces and external systems such as databases or AI services.
|
||||
- Background job: Work executed outside the request/response path so the UI remains responsive.
|
||||
- Boundary: A strict separation between modules with different responsibilities.
|
||||
- CI (Continuous Integration): Automated test execution for code changes.
|
||||
- Contract test: A test that verifies an adapter follows expected input/output behavior at a boundary.
|
||||
- Domain layer: The module that contains core business rules and invariants.
|
||||
- End-to-end test: A test that validates a full user flow across the running system.
|
||||
- Full-text search: Text indexing and querying optimized for natural-language search.
|
||||
- In-process worker: A background executor that runs within the same application process.
|
||||
- Integration test: A test that verifies interactions between real modules and infrastructure components.
|
||||
- MongoDB: A document-oriented database used for flexible, high-variance data structures.
|
||||
- Modular monolith: A single deployable application with strongly separated internal modules.
|
||||
- Port/Interface: A stable contract used by application/domain code to call infrastructure implementations.
|
||||
- Prompt artifact: A single Markdown file that defines one transcription prompt and can be revised independently.
|
||||
- Provenance: Metadata that records where generated data came from and how it was produced.
|
||||
- Revision history: Optional versioned record of user-authored transcription edits over time.
|
||||
- System of record: The authoritative persistent store for canonical data.
|
||||
- Vertical slice: A minimal end-to-end feature path spanning UI/API, application logic, and persistence.
|
||||
@@ -1,288 +0,0 @@
|
||||
# Error Handling Policy
|
||||
|
||||
This document defines the canonical error-handling policy for the document transcription system. It is the single source of truth for how errors are classified, surfaced to users, logged for diagnosis, and handled across UI, API, service, worker, and provider boundaries.
|
||||
|
||||
## Error Handling Objectives
|
||||
|
||||
The production error-handling model is designed to:
|
||||
|
||||
- make failures visible to the user in clear, actionable language
|
||||
- preserve enough diagnostic detail for fast troubleshooting
|
||||
- keep module behavior consistent across all boundaries
|
||||
- distinguish expected domain failures from unexpected defects
|
||||
- support safe retries for transient failures without hiding persistent faults
|
||||
|
||||
## Scope And Authority
|
||||
|
||||
This page governs error-handling behavior for:
|
||||
|
||||
- UI interactions (NiceGUI pages)
|
||||
- API endpoints (FastAPI routes)
|
||||
- application services and orchestration logic
|
||||
- in-process background worker execution
|
||||
- external provider adapters and persistence adapters
|
||||
|
||||
If implementation behavior conflicts with this document, this document is authoritative and implementation should be updated.
|
||||
|
||||
## Core Principles
|
||||
|
||||
- **Clarity first:** user-facing messages should explain what failed in plain language.
|
||||
- **Actionability required:** each surfaced error should include a suggested next step.
|
||||
- **Safety by default:** internal details are logged; sensitive details are not exposed by default in UI/API.
|
||||
- **Consistency across boundaries:** category and structure should remain stable from source to surface.
|
||||
- **Fail explicitly:** silent failure is prohibited.
|
||||
- **Traceability:** every non-trivial error should be traceable with an error reference ID.
|
||||
|
||||
## Error Taxonomy
|
||||
|
||||
The system uses stable, implementation-independent categories:
|
||||
|
||||
| Category | Definition | Typical Source | Retriable |
|
||||
| --- | --- | --- | --- |
|
||||
| `validation_error` | Payload or parameter shape/content is invalid | UI/API input validation, service guards | no |
|
||||
| `user_input_error` | User-provided artifact is unacceptable though structurally valid | unsupported file type, empty file, oversized upload | sometimes |
|
||||
| `not_found_error` | Requested resource does not exist | missing job/document/source/revision | no |
|
||||
| `conflict_error` | Requested operation violates current state constraints | invalid state transition | no |
|
||||
| `external_provider_error` | External AI/provider call fails | upstream HTTP/API/provider failures | sometimes |
|
||||
| `infrastructure_transient_error` | Temporary environment issue | network timeout, DB connection reset | yes |
|
||||
| `infrastructure_persistent_error` | Non-transient environment issue | missing permissions, misconfiguration | no |
|
||||
| `internal_unexpected_error` | Unhandled defect or unknown failure | uncaught exceptions, logic errors | unknown (default no) |
|
||||
|
||||
### Classification Rules
|
||||
|
||||
- Classification occurs as close as possible to the origin boundary.
|
||||
- Provider-specific exceptions must be normalized into taxonomy categories before crossing service boundaries.
|
||||
- Unknown exceptions are classified as `internal_unexpected_error` and logged with traceback.
|
||||
- Category names are stable contracts and must not be changed casually.
|
||||
|
||||
## User-Facing Error Experience Contract
|
||||
|
||||
When an error is shown in the GUI, it must include:
|
||||
|
||||
1. **Title** (short context, e.g., “Upload failed”)
|
||||
2. **Message** (plain-language explanation)
|
||||
3. **Suggested action** (explicit next step)
|
||||
4. **Error reference ID** (for support/debug traceability)
|
||||
5. **Technical details** (optional/collapsible for advanced users)
|
||||
|
||||
### UI Message Rules
|
||||
|
||||
- Do not expose raw stack traces by default.
|
||||
- Do not expose secrets, credentials, connection strings, or filesystem internals unless explicitly in debug tooling.
|
||||
- Prefer domain language over implementation language.
|
||||
- Use persistent visibility for important failures (dialog/card), not only transient toasts.
|
||||
|
||||
### Suggested Action Requirements
|
||||
|
||||
Every user-visible error must include a suggested course of action, such as:
|
||||
|
||||
- retry the operation
|
||||
- check file type/size constraints
|
||||
- refresh the jobs page
|
||||
- verify environment configuration
|
||||
- contact operator with error ID and timestamp
|
||||
|
||||
## API Error Response Contract
|
||||
|
||||
API errors should return a structured envelope with stable fields:
|
||||
|
||||
- `error_id`: short unique reference ID
|
||||
- `category`: taxonomy category
|
||||
- `message`: safe human-readable summary
|
||||
- `suggestion`: recommended next step
|
||||
- `details`: optional, only when safe and appropriate
|
||||
- `timestamp`: UTC ISO-8601
|
||||
|
||||
HTTP status mapping guidance:
|
||||
|
||||
- `validation_error`, `user_input_error` -> `400`
|
||||
- `not_found_error` -> `404`
|
||||
- `conflict_error` -> `409`
|
||||
- `external_provider_error` -> `502` or `503` (depending on failure mode)
|
||||
- `infrastructure_transient_error` -> `503`
|
||||
- `infrastructure_persistent_error` -> `500`
|
||||
- `internal_unexpected_error` -> `500`
|
||||
|
||||
## Logging And Observability Contract
|
||||
|
||||
All logged errors must include, where available:
|
||||
|
||||
- `error_id`
|
||||
- `category`
|
||||
- `operation` (e.g., `upload.submit`, `worker.process_job`, `jobs.refresh`)
|
||||
- `exception_type`
|
||||
- `job_id`, `document_id`, `source_id` (when relevant)
|
||||
- UTC timestamp
|
||||
|
||||
Rules:
|
||||
|
||||
- Use structured logging fields where practical.
|
||||
- Use full traceback for unexpected errors (`internal_unexpected_error`).
|
||||
- Log at boundary handoff points to preserve causal trail.
|
||||
- Avoid duplicate noisy logging for the same exception at every layer.
|
||||
|
||||
## Recovery And Retry Policy
|
||||
|
||||
### Retriable Conditions
|
||||
|
||||
Retriable failures include:
|
||||
|
||||
- transient network/provider timeouts
|
||||
- intermittent provider unavailability
|
||||
- temporary DB/network interruptions
|
||||
|
||||
### Non-Retriable Conditions
|
||||
|
||||
Non-retriable failures include:
|
||||
|
||||
- invalid file formats
|
||||
- missing required data
|
||||
- permission/configuration failures
|
||||
- deterministic domain conflicts
|
||||
|
||||
### Worker Behavior
|
||||
|
||||
- The worker must classify and persist failure details consistently.
|
||||
- Retries should be bounded by configured limits.
|
||||
- Exhausted retries must end in explicit failed status with recorded reason.
|
||||
- No infinite retry loops are allowed.
|
||||
|
||||
## Boundary-Specific Responsibilities
|
||||
|
||||
### UI Layer
|
||||
|
||||
Responsibility:
|
||||
|
||||
- display user-safe error summaries and suggested actions
|
||||
- show persistent error visibility for critical failures
|
||||
- include error reference IDs in visible output
|
||||
|
||||
Out of scope:
|
||||
|
||||
- low-level exception parsing
|
||||
- provider-specific protocol interpretation
|
||||
|
||||
### API Layer
|
||||
|
||||
Responsibility:
|
||||
|
||||
- map application exceptions into stable error envelopes and HTTP statuses
|
||||
- preserve category and error_id continuity
|
||||
|
||||
Out of scope:
|
||||
|
||||
- domain-specific remediation logic
|
||||
|
||||
### Service Layer
|
||||
|
||||
Responsibility:
|
||||
|
||||
- classify domain and infrastructure exceptions
|
||||
- convert adapter-specific failures into taxonomy categories
|
||||
- return deterministic error types to callers
|
||||
|
||||
Out of scope:
|
||||
|
||||
- presentation formatting for UI
|
||||
|
||||
### Worker Layer
|
||||
|
||||
Responsibility:
|
||||
|
||||
- execute retry policy for retriable failures
|
||||
- persist terminal failure details for jobs
|
||||
- emit operational logs with category and identifiers
|
||||
|
||||
Out of scope:
|
||||
|
||||
- direct UI messaging
|
||||
|
||||
### Provider Adapter Layer
|
||||
|
||||
Responsibility:
|
||||
|
||||
- normalize provider SDK/HTTP failures into domain-neutral exceptions
|
||||
- preserve raw provider context for logs (safely)
|
||||
|
||||
Out of scope:
|
||||
|
||||
- choosing user-facing wording
|
||||
|
||||
## Error Lifecycle Workflow
|
||||
|
||||
Standard lifecycle:
|
||||
|
||||
1. Failure occurs at a boundary or operation.
|
||||
2. Exception is classified into taxonomy category.
|
||||
3. `error_id` is created (or propagated).
|
||||
4. Error is logged with required structured fields.
|
||||
5. User/API receives safe message + suggested action.
|
||||
6. Persistent job/resource state is updated when applicable.
|
||||
7. Tests verify contract behavior for the pathway.
|
||||
|
||||
## Test Strategy For Error Handling
|
||||
|
||||
### Unit Tests
|
||||
|
||||
- category classification behavior
|
||||
- retry eligibility decisions
|
||||
- exception-to-message mapping safety
|
||||
|
||||
### Integration Tests
|
||||
|
||||
- UI pathways show clear message + suggested action for known failures
|
||||
- API returns structured error envelope with expected status/category
|
||||
- worker persists failed status and failure detail as required
|
||||
|
||||
### Regression Tests
|
||||
|
||||
- each previously observed production issue should have a guarding test
|
||||
- contract tests must cover adapter error normalization behavior
|
||||
|
||||
## Known Failure Patterns And Prescribed Responses
|
||||
|
||||
| Pattern | Category | User Message | Suggested Action |
|
||||
| --- | --- | --- | --- |
|
||||
| Upload payload cannot be parsed by UI handler | `internal_unexpected_error` (until narrowed) | Upload failed due to unexpected processing error | Retry once; if repeated, report error ID and check runtime version compatibility |
|
||||
| Unsupported extension | `user_input_error` | File type is not supported | Upload JPG, PNG, TIFF, or PDF |
|
||||
| Empty file upload | `validation_error` | Uploaded file is empty | Choose a valid non-empty file and retry |
|
||||
| Provider timeout | `external_provider_error` or `infrastructure_transient_error` | Transcription provider timed out | Retry from jobs page; if repeated, check provider status |
|
||||
| Job lookup missing | `not_found_error` | Requested job was not found | Refresh jobs list and open a valid job |
|
||||
|
||||
## Governance And Update Process
|
||||
|
||||
This document is a living policy artifact.
|
||||
|
||||
Update this document when:
|
||||
|
||||
- new error categories are introduced
|
||||
- handling behavior changes at any boundary
|
||||
- a production incident reveals missing guidance
|
||||
- API/UI error contracts change
|
||||
|
||||
Change requirements:
|
||||
|
||||
- update this document and associated tests in the same change set
|
||||
- preserve taxonomy stability; if changed, document migration impact
|
||||
- record noteworthy policy changes in project release notes or changelog
|
||||
|
||||
---
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v1.md)
|
||||
- [System Design Intent](intent.md)
|
||||
- [Transcription Methodology](transcription_methodology.md)
|
||||
- [System Architecture](architecture_v1.md)
|
||||
- [System Requirements](requirements_v1.md)
|
||||
- [Data model](schema_v1.md)
|
||||
- Error Handling Policy (this document)
|
||||
- [Implementation Plan](implementation_plan_v1.md)
|
||||
|
||||
## Glossary
|
||||
|
||||
- Error category: Stable classification used to drive handling, messaging, and status mapping.
|
||||
- Error envelope: Structured API payload describing a failure.
|
||||
- Error reference ID: Short identifier used to correlate user-visible failure with logs.
|
||||
- Retriable error: Failure likely to succeed on a later attempt without code changes.
|
||||
- Terminal failure: Failure state after retries are exhausted or retry is not allowed.
|
||||
@@ -1,203 +0,0 @@
|
||||
# Version 1 Implementation Plan
|
||||
|
||||
This plan defines the path from current implementation to **Version 1 complete**, aligned to the updated domain model:
|
||||
|
||||
- `Document` groups one or more content `Source` records
|
||||
- `Job` owns original immutable provider output (`text`) and processing metadata
|
||||
- `Revision` stores optional user-authored edits linked to a `Source`
|
||||
|
||||
The objective is to complete V1 scope with production readiness while keeping non-V1 enhancements out of active delivery.
|
||||
|
||||
---
|
||||
|
||||
## V1 Completion Definition
|
||||
|
||||
V1 is complete when all of the following are true:
|
||||
|
||||
1. **Functional complete**
|
||||
- Upload, queue, processing, status display, and transcription result inspection work end-to-end.
|
||||
- Optional revision workflow is implemented (create/view/update single revision).
|
||||
2. **Data-model complete**
|
||||
- Runtime behavior, persistence, and tests all align to `Document` / `Source` / `Job` / `Revision`.
|
||||
3. **Operational complete**
|
||||
- Error handling, logs, and runbooks support reliable operation.
|
||||
4. **Documentation complete**
|
||||
- Architecture, requirements, schema, error handling, and index are consistent and current.
|
||||
|
||||
---
|
||||
|
||||
## Phase 1 — Data Contract Stabilization (Schema-First)
|
||||
|
||||
**Goal:** Lock a single canonical contract before further feature work.
|
||||
|
||||
### Tasks
|
||||
1. Confirm and document invariants:
|
||||
- `Job.text` is original immutable transcription output.
|
||||
- `Revision` is optional and user-authored.
|
||||
- Revisions are derived from the original `Job.text`.
|
||||
2. Verify relationship cardinality assumptions:
|
||||
- `Document` -> many `Source`
|
||||
- `Document` -> many `Job`
|
||||
- `Source` -> one `Job`
|
||||
- `Source` -> one `Revision`
|
||||
3. Ensure field naming consistency (`date_created`, `date_updated`, `date_uploaded`) across code and docs.
|
||||
4. Freeze V1 status lifecycle to current implementation (`queued`, `processing`, `transcribed`, `failed`).
|
||||
|
||||
### Deliverables
|
||||
- Updated `schema_v1.md` and `requirements.md` traceability alignment.
|
||||
- Explicit V1 data invariants section in architecture docs.
|
||||
|
||||
### Exit Criteria
|
||||
- No conflicting definitions of ownership/cardinality/status remain in docs.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2 — Service Layer Refactor To New Model
|
||||
|
||||
**Goal:** Remove all obsolete `Transcript` assumptions from service/workflow code.
|
||||
|
||||
### Tasks
|
||||
1. Refactor `services/transcription.py`:
|
||||
- Replace transcript CRUD assumptions with job-output + revision operations.
|
||||
2. Refactor `services/jobs.py`:
|
||||
- Replace old timestamp/relationship accessors with current model fields.
|
||||
3. Refactor `services/documents.py` and `services/store.py`:
|
||||
- Ensure upload creates and links `Document`, `Source`, and `Job` correctly.
|
||||
4. Refactor `services/workflows.py`:
|
||||
- Persist original provider output to `Job`.
|
||||
- Persist failure detail to `Job.error_detail`.
|
||||
- Use `Revision` only for user-authored edits.
|
||||
|
||||
### Deliverables
|
||||
- Service layer fully aligned with new schema.
|
||||
|
||||
### Exit Criteria
|
||||
- No service module imports or persists `Transcript` model artifacts.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3 — UI Contract Alignment
|
||||
|
||||
**Goal:** Align pages/components to source/job/revision semantics.
|
||||
|
||||
### Tasks
|
||||
1. Update job detail and related UI components:
|
||||
- Display original immutable transcription from `Job.text`.
|
||||
- Display optional revision sourced from `Source.revision` (0 or 1).
|
||||
2. Align date fields with new schema naming.
|
||||
3. Preserve clear user messaging when no revisions exist.
|
||||
|
||||
### Deliverables
|
||||
- Updated jobs page and detail components.
|
||||
|
||||
### Exit Criteria
|
||||
- UI behavior and labels match documentation and domain model.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4 — Database Bootstrap, Migration, and Safety
|
||||
|
||||
**Goal:** Make schema transition safe in dev/test and repeatable for deployment.
|
||||
|
||||
### Tasks
|
||||
1. Update bootstrap compatibility logic in `db/operations.py`:
|
||||
- Remove obsolete transcript-table assumptions.
|
||||
- Add forward-compatible patches for current tables only.
|
||||
2. Define migration/backfill approach for existing local data.
|
||||
3. Document rollback and recovery steps.
|
||||
4. Rehearse migration path against representative data.
|
||||
|
||||
### Deliverables
|
||||
- Migration/upgrade runbook.
|
||||
- Validated bootstrap behavior for dev/test.
|
||||
|
||||
### Exit Criteria
|
||||
- Migration path is documented and tested with no unresolved data-loss risk.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5 — Test Suite Realignment
|
||||
|
||||
**Goal:** Restore full confidence after the schema redesign.
|
||||
|
||||
### Tasks
|
||||
1. Rewrite model tests for:
|
||||
- `Document`, `Source`, `Job`, `Revision` relationships and invariants.
|
||||
2. Rewrite service/integration tests:
|
||||
- Worker success/failure paths using `Job.text` / `Job.error_detail`.
|
||||
- Optional single-revision creation/update behavior.
|
||||
3. Update UI tests for new job-detail/revision rendering behavior.
|
||||
4. Re-enable strict CI quality gates (lint, type, tests).
|
||||
|
||||
### Deliverables
|
||||
- Updated test matrix and passing CI.
|
||||
|
||||
### Exit Criteria
|
||||
- Critical user flows and failure paths are covered and green.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6 — Reliability, Operations, and Release Readiness
|
||||
|
||||
**Goal:** Ensure V1 is operable and launch-safe.
|
||||
|
||||
### Tasks
|
||||
1. Verify error taxonomy behavior across UI/API/service/worker.
|
||||
2. Confirm structured logging includes relevant identifiers (`job_id`, `document_id`, `source_id` when applicable).
|
||||
3. Validate retry behavior and terminal failure handling.
|
||||
4. Finalize release checklist, deployment steps, and rollback procedure.
|
||||
5. Execute final acceptance run against requirements traceability.
|
||||
|
||||
### Deliverables
|
||||
- V1 release checklist and acceptance evidence.
|
||||
- `runbook_v1.md` for incident response and operator workflows.
|
||||
- `release_checklist_v1.md` for release sign-off.
|
||||
|
||||
### Exit Criteria
|
||||
- Stakeholder sign-off and launch readiness achieved.
|
||||
|
||||
---
|
||||
|
||||
## Requirement Traceability Focus
|
||||
|
||||
The plan must keep clear evidence against these requirement groups:
|
||||
|
||||
- **Core flow:** REQ-0 to REQ-6
|
||||
- **Runtime and operations constraints:** REQ-7 to REQ-12
|
||||
- **Revision workflow:** REQ-13
|
||||
|
||||
A lightweight traceability table should be maintained with:
|
||||
|
||||
- requirement ID
|
||||
- implementation status (`not started` / `in progress` / `done`)
|
||||
- validation evidence (test name, screenshot, or runbook step)
|
||||
|
||||
---
|
||||
|
||||
## Suggested Execution Rhythm
|
||||
|
||||
- **Weekly:** requirement status and risk review
|
||||
- **Per PR:** contract checks (model names, field names, lifecycle values)
|
||||
- **Milestone checks:** end of Phases 2, 4, and 6
|
||||
|
||||
---
|
||||
|
||||
## Scope Discipline Rule (V1 Focus)
|
||||
|
||||
- Only work required to satisfy V1 requirements enters this plan.
|
||||
- Nice-to-have enhancements are captured in a separate backlog document.
|
||||
- Schema or contract changes after Phase 1 require explicit approval and traceability impact review.
|
||||
|
||||
---
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v1.md)
|
||||
- [System Design Intent](intent.md)
|
||||
- [Transcription Methodology](transcription_methodology.md)
|
||||
- [System Architecture](architecture_v1.md)
|
||||
- [System Requirements](requirements_v1.md)
|
||||
- [Data model](schema_v1.md)
|
||||
- [Error Handling Policy](error_handling_v1.md)
|
||||
- Implementation Plan (this document)
|
||||
|
||||
@@ -1,58 +0,0 @@
|
||||
## Document Transcription System Overview
|
||||
|
||||
This project is a production application for transcribing and preserving historical family documents. It is intentionally designed for personal-scale use, with a simplicity-first architecture that is easy to operate and easy to extend.
|
||||
|
||||
## Start Here
|
||||
|
||||
Read [architecture_v1.md](architecture_v1.md) first.
|
||||
|
||||
The architecture page is the primary technical reference and defines:
|
||||
|
||||
- deployed topology and infrastructure limits
|
||||
- module boundaries and dependency flow
|
||||
- processing life cycle and data ownership
|
||||
- test strategy, risk controls, and extension path
|
||||
|
||||
## What The Application Does
|
||||
|
||||
At a high level, users upload images or PDFs as content sources for handwritten, typed, or typeset documents, run asynchronous transcription jobs, review optional revisions, and search across accepted text.
|
||||
|
||||
### Core capabilities:
|
||||
|
||||
- document grouping with one or more content sources and metadata capture
|
||||
- asynchronous transcription with visible job status
|
||||
- immutable original transcription persisted with each job (plus provider/model/prompt metadata)
|
||||
- transcription prompt management with one Markdown file per prompt for human refinement over time
|
||||
- optional revisions for user-authored edits of original immutable transcription text
|
||||
- full-text search over accepted transcripts
|
||||
- export of transcript data
|
||||
|
||||
## Production Operating Model
|
||||
|
||||
The system runs with minimal operational overhead:
|
||||
|
||||
- PostgreSQL in a dedicated Docker container is considered extremely lightweight and simple for this system
|
||||
- MongoDB in a dedicated Docker container is also considered extremely lightweight and simple for document-centric persistence
|
||||
- a three-container deployment (app, PostgreSQL, MongoDB) is a simple and acceptable baseline
|
||||
- no required queue or search-engine containers in the baseline setup
|
||||
|
||||
This operating model keeps deployment and maintenance simple while preserving clean boundaries for future scale.
|
||||
|
||||
---
|
||||
|
||||
## Documentation Map
|
||||
|
||||
- System Overview (this document)
|
||||
- [System Design Intent](intent.md)
|
||||
- [Transcription Methodology](transcription_methodology.md)
|
||||
- [System Architecture](architecture_v1.md)
|
||||
- [System Requirements](requirements_v1.md)
|
||||
- [Data model](schema_v1.md)
|
||||
- [Error Handling Policy](error_handling_v1.md)
|
||||
- [Implementation Plan](implementation_plan_v1.md)
|
||||
|
||||
## Glossary
|
||||
|
||||
- Document-oriented persistence: Storing data as flexible records instead of fixed relational rows.
|
||||
- Prompt artifact: A single Markdown file that defines one transcription prompt and is edited independently.
|
||||
- System of record: The authoritative persistent store for canonical data.
|
||||
@@ -1,45 +0,0 @@
|
||||
# V1 Release Readiness Checklist
|
||||
|
||||
Use this checklist before declaring V1 operationally complete.
|
||||
|
||||
## A) Functional Readiness
|
||||
|
||||
- [ ] Upload flow works for supported file types.
|
||||
- [ ] Worker transitions jobs through `queued -> processing -> transcribed|failed`.
|
||||
- [ ] Job detail displays immutable original transcription from `Job.text`.
|
||||
- [ ] Revision workflow supports create/update/view/delete for optional single revision.
|
||||
|
||||
## B) Reliability and Error Handling
|
||||
|
||||
- [ ] Error categories surface with actionable messages in UI/API pathways.
|
||||
- [ ] Failed jobs persist `error_detail` and terminal state.
|
||||
- [ ] Stale processing recovery verified on restart.
|
||||
- [ ] Retry/timeout behavior validated against configured limits.
|
||||
|
||||
## C) Operational Readiness
|
||||
|
||||
- [ ] `runbook_v1.md` reviewed and current.
|
||||
- [ ] `migration_v1.md` reviewed and current.
|
||||
- [ ] Backup and rollback procedures tested at least once.
|
||||
- [ ] Incident escalation packet template is known to operators.
|
||||
|
||||
## D) Quality Gates
|
||||
|
||||
- [ ] Lint/type checks pass.
|
||||
- [ ] `pytest -m "not external" -q` passes.
|
||||
- [ ] Targeted external/provider checks executed (if credentials available).
|
||||
- [ ] Release evidence recorded in `release_evidence_v1.md`.
|
||||
|
||||
## E) Traceability and Documentation
|
||||
|
||||
- [ ] `requirements_v1.md` aligns with implemented V1 behavior.
|
||||
- [ ] `architecture_v1.md`, `schema_v1.md`, and `error_handling_v1.md` are consistent.
|
||||
- [ ] `traceability_v1.md` is updated with current implementation and test evidence.
|
||||
- [ ] `implementation_plan_v1.md` phase status updated with evidence references.
|
||||
- [ ] REQ traceability evidence links recorded (tests/runbook/checks).
|
||||
|
||||
## Release Sign-Off
|
||||
|
||||
- [ ] Technical sign-off complete.
|
||||
- [ ] Operational sign-off complete.
|
||||
- [ ] V1 completion date recorded.
|
||||
@@ -1,39 +0,0 @@
|
||||
# V1 Release Evidence Log
|
||||
|
||||
## Step 5 Quality Gates (2026-07-29)
|
||||
|
||||
### Lint
|
||||
|
||||
- Command: `python -m ruff check .`
|
||||
- Result: ✅ pass
|
||||
- Notes: initial findings were auto-fixed (`ruff --fix`) plus small manual line-wrap/annotation adjustments.
|
||||
|
||||
### Tests (primary gate)
|
||||
|
||||
- Command: `python -m pytest -m "not external" -q`
|
||||
- Result: ✅ pass (`[100%]`)
|
||||
|
||||
### Tests (external smoke)
|
||||
|
||||
- Command: `python -m pytest -m external -q`
|
||||
- Result: ✅ pass (`[100%]`)
|
||||
|
||||
### Type Check
|
||||
|
||||
- Command: `python -m ty check src tests`
|
||||
- Result: ⚠️ not passing
|
||||
- Summary: existing SQLModel/SQLAlchemy typing incompatibilities and test double typing mismatches remain.
|
||||
|
||||
Key current blocker families:
|
||||
|
||||
1. SQLModel relationship/query attribute typing (`selectinload`, `order_by`, `.any()`)
|
||||
2. SQLAlchemy join clause typing in `services/transcription.py`
|
||||
3. Test fake client type mismatch for `OpenRouterTranscriptionProvider(client=...)`
|
||||
4. `Settings(**defaults)` typed-dict strictness in `tests/test_config.py`
|
||||
|
||||
## Current Gate Status
|
||||
|
||||
- Lint: pass
|
||||
- Non-external tests: pass
|
||||
- External smoke tests: pass
|
||||
- Type check: **blocked** (requires dedicated typing cleanup pass)
|
||||
@@ -1,98 +0,0 @@
|
||||
## Document Transcription System Requirements
|
||||
|
||||
This page captures a SysML v1.6-style requirements baseline for the production system described in [index_v1.md](index_v1.md). The model is represented as concise tables and traceability lists that preserve SysML-style IDs and relationship semantics.
|
||||
|
||||
## Scope
|
||||
|
||||
- System of interest: the single Python application service (NiceGUI + FastAPI) with PostgreSQL as the relational system of record and optional MongoDB for document-oriented persistence.
|
||||
- Operational context: local-first execution with Docker Compose and an intentionally lightweight production trajectory.
|
||||
- Primary concern: end-to-end transcription job lifecycle from upload through completion or failure.
|
||||
|
||||
## Requirements Model (Concise Text Form)
|
||||
|
||||
### Requirements
|
||||
|
||||
| ID | Category | Requirement | Risk | Verify Method |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| REQ-0 | System | Provide end-to-end document transcription with persistent, inspectable lifecycle state. | medium | demonstration |
|
||||
| REQ-1 | Functional | Allow users to upload one or more images or PDFs as sources from the web UI. | low | test |
|
||||
| REQ-2 | Functional | Run each upload through asynchronous processing that returns an original transcription or explicit failure. | high | test |
|
||||
| REQ-3 | Functional | Persist and expose job states: queued, processing, transcribed, failed. | high | inspection |
|
||||
| REQ-4 | Functional | Persist transcription output, processing history, and failure details. | medium | test |
|
||||
| REQ-5 | Interface | Expose API and UI views for status inspection and completed transcription reading. | medium | demonstration |
|
||||
| REQ-6 | Performance | Trigger background processing on upload to preserve UI responsiveness. | medium | analysis |
|
||||
| REQ-7 | Design Constraint | Keep lifespan-owned runtime resources: SQLAlchemy engine, async session factory, worker resources, provider clients. | medium | inspection |
|
||||
| REQ-8 | Design Constraint | Initialize configuration and logging once at startup through centralized mechanisms. | low | inspection |
|
||||
| REQ-9 | Design Constraint | Use Docker Compose baseline of app plus PostgreSQL; allow optional MongoDB container when enabled. | medium | demonstration |
|
||||
| REQ-10 | Design Constraint | Keep schema bootstrap explicit and opt-in; normal startup does not mutate production schema. | high | inspection |
|
||||
| REQ-11 | Design Constraint | Use service-backed persistence for core document and job data. | medium | inspection |
|
||||
| REQ-12 | Design Constraint | Store transcription prompts as individual Markdown artifacts for iterative refinement. | medium | inspection |
|
||||
| REQ-13 | Functional | Allow users to create one optional revision of transcription text derived from the original job transcription. | low | test |
|
||||
|
||||
### Requirement Relationships
|
||||
|
||||
- Contains: REQ-0 contains REQ-1 through REQ-13.
|
||||
- Derives: REQ-2 -> REQ-3, REQ-3 -> REQ-4.
|
||||
- Traces: REQ-5 -> REQ-3.
|
||||
- Refines: REQ-6 -> REQ-2.
|
||||
|
||||
### Architecture Elements
|
||||
|
||||
| Element | Type | Doc Reference |
|
||||
| --- | --- | --- |
|
||||
| UI | NiceGUI pages | src/transcription/ui/pages |
|
||||
| API | FastAPI routes | src/transcription/api/routes.py |
|
||||
| GRAPH | Async processing workflow | src/transcription/services, src/transcription/ai |
|
||||
| DBREL | PostgreSQL + SQLModel relational persistence | src/transcription/db |
|
||||
| DBDOC | MongoDB document persistence | src/transcription/db, src/transcription/services |
|
||||
| OPS | Docker Compose runtime | docker-compose.yml |
|
||||
| PROMPTS | Transcription prompt artifact library (Markdown files) | .github/prompts, docs |
|
||||
| TESTS | Pytest verification suite | tests |
|
||||
|
||||
### Satisfaction Mapping
|
||||
|
||||
- UI satisfies REQ-1, REQ-5, REQ-13.
|
||||
- API satisfies REQ-5.
|
||||
- GRAPH satisfies REQ-2, REQ-6.
|
||||
- DBREL satisfies REQ-3, REQ-10, REQ-13.
|
||||
- DBDOC satisfies REQ-4, REQ-11.
|
||||
- OPS satisfies REQ-9.
|
||||
- PROMPTS satisfies REQ-12.
|
||||
|
||||
### Verification Mapping
|
||||
|
||||
- TESTS verifies REQ-1, REQ-2, REQ-3, REQ-4, REQ-5, REQ-10, REQ-11, REQ-12, REQ-13.
|
||||
|
||||
## Requirement Notes
|
||||
|
||||
- Requirement IDs (`REQ-*`) are stable references for planning, implementation, and test traceability.
|
||||
- The model uses compact tables and traceability lists for renderer compatibility while preserving SysML-style requirement IDs and relationship semantics.
|
||||
- Requirement categories (functional, interface, performance, and design constraints) are preserved as explicit REQ entries and relationship labels to keep change impact visible.
|
||||
- PostgreSQL containerization and optional MongoDB containerization are both treated as extremely lightweight and simple operational choices in this architecture.
|
||||
|
||||
## Verification Intent
|
||||
|
||||
- Demonstration: validate end-to-end behavior via running system flows and operator-visible outcomes.
|
||||
- Inspection: verify architecture and startup/runtime policies in code and configuration.
|
||||
- Analysis: evaluate asynchronous execution behavior and design sufficiency.
|
||||
- Test: automate behavioral checks through pytest suites and service-level tests.
|
||||
|
||||
---
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v1.md)
|
||||
- [System Design Intent](intent.md)
|
||||
- [Transcription Methodology](transcription_methodology.md)
|
||||
- [System Architecture](architecture_v1.md)
|
||||
- System Requirements (this document)
|
||||
- [Data model](schema_v1.md)
|
||||
- [Error Handling Policy](error_handling_v1.md)
|
||||
- [Implementation Plan](implementation_plan_v1.md)
|
||||
|
||||
## Glossary
|
||||
|
||||
- Document-oriented persistence: A storage approach that uses flexible document structures for variable data shapes.
|
||||
- Prompt artifact: A single Markdown file that defines one transcription prompt and is revised independently.
|
||||
- SysML: Systems Modeling Language used to express structured requirements and traceability.
|
||||
- System of record: The authoritative persistent store for canonical business data.
|
||||
@@ -1,129 +0,0 @@
|
||||
# V1 Operations Runbook
|
||||
|
||||
This runbook provides day-2 operational procedures for the V1 baseline.
|
||||
|
||||
## Scope
|
||||
|
||||
Applies to:
|
||||
|
||||
- local/hosted V1 runtime
|
||||
- SQLite-backed persistence
|
||||
- in-process worker lifecycle
|
||||
- OpenRouter provider integration
|
||||
|
||||
## Preconditions
|
||||
|
||||
- `.env` contains `OPENROUTER_API_KEY`
|
||||
- app starts successfully
|
||||
- `uploads/` and `prompts/` are writable
|
||||
- health endpoint responds at `/healthz`
|
||||
|
||||
## Standard Startup Procedure
|
||||
|
||||
1. Start the app using the project-standard command.
|
||||
2. Open `/healthz` and verify `{"status":"ok"}`.
|
||||
3. Open `/ui/upload` and submit a small valid file.
|
||||
4. Confirm job transitions from `queued` -> `processing` -> `transcribed` (or `failed` with detail).
|
||||
|
||||
## Standard Shutdown Procedure
|
||||
|
||||
1. Stop the application process.
|
||||
2. Ensure no active process still holds the SQLite file.
|
||||
3. If maintenance is planned, copy the DB file before edits:
|
||||
- `transcription.db` (or configured `DATABASE_URL` file path)
|
||||
|
||||
## Incident: Jobs Stuck In `processing`
|
||||
|
||||
### Symptoms
|
||||
|
||||
- Jobs remain `processing` for longer than provider timeout
|
||||
- New uploads queue but do not complete
|
||||
- provider usage increases but no terminal job state is visible
|
||||
|
||||
### Checks
|
||||
|
||||
1. Confirm app process is still running.
|
||||
2. Confirm worker loop is active (startup logs include worker lifespan start).
|
||||
3. Inspect recent app logs for:
|
||||
- `worker.process_job`
|
||||
- `error_id`
|
||||
- `category`
|
||||
- `job_id` / `document_id` / `source_id`
|
||||
4. Verify provider credentials and provider status.
|
||||
|
||||
### Recovery
|
||||
|
||||
1. Restart the app to trigger stale-processing recovery.
|
||||
2. On startup, app re-queues stale processing jobs based on timeout policy.
|
||||
3. Re-check jobs page and confirm terminal state progression.
|
||||
4. If persistent, capture logs + error IDs and move to deep investigation.
|
||||
|
||||
## Incident: Provider Authentication Failures
|
||||
|
||||
### Symptoms
|
||||
|
||||
- failures categorized as provider/auth
|
||||
- jobs fail quickly with authentication guidance
|
||||
|
||||
### Recovery
|
||||
|
||||
1. Validate `OPENROUTER_API_KEY` value.
|
||||
2. Restart app after updating env.
|
||||
3. Re-run a small transcription to confirm recovery.
|
||||
|
||||
## Incident: Upload Failures
|
||||
|
||||
### Symptoms
|
||||
|
||||
- UI reports upload errors
|
||||
- unsupported extension or empty payload
|
||||
|
||||
### Recovery
|
||||
|
||||
1. Validate file extension (`.jpg`, `.jpeg`, `.png`, `.tif`, `.tiff`, `.pdf`).
|
||||
2. Validate file is not empty.
|
||||
3. Validate upload directory permissions.
|
||||
4. Retry upload.
|
||||
|
||||
## Incident: Database File/Permission Issues
|
||||
|
||||
### Symptoms
|
||||
|
||||
- persistence errors during upload/job update
|
||||
- startup failures around schema/runtime
|
||||
|
||||
### Recovery
|
||||
|
||||
1. Confirm the configured DB file path exists and is writable.
|
||||
2. Confirm parent directory permissions.
|
||||
3. Restore from last known backup copy if corruption is suspected.
|
||||
4. Restart app and run smoke test.
|
||||
|
||||
## Logging Requirements (Operational)
|
||||
|
||||
Operational triage should always capture:
|
||||
|
||||
- `error_id`
|
||||
- category
|
||||
- operation name
|
||||
- `job_id`, `document_id`, `source_id` when applicable
|
||||
- UTC timestamp
|
||||
|
||||
## Escalation Packet (When opening an issue)
|
||||
|
||||
Include:
|
||||
|
||||
- exact timestamp window
|
||||
- one failing `job_id`
|
||||
- relevant `error_id` values
|
||||
- latest 100 lines of app logs
|
||||
- environment summary (`DATABASE_URL` type, app version/commit)
|
||||
|
||||
## Post-Incident Validation
|
||||
|
||||
After mitigation, verify:
|
||||
|
||||
1. Upload works.
|
||||
2. One job reaches `transcribed`.
|
||||
3. One induced failure reaches `failed` with error detail.
|
||||
4. Jobs page and detail page render correctly.
|
||||
@@ -1,98 +0,0 @@
|
||||
## Database Schema (V1 Baseline)
|
||||
|
||||
This document describes the current relational schema for the transcription system.
|
||||
|
||||
All primary and foreign keys in the domain models are UUID-based in V1.
|
||||
|
||||
---
|
||||
|
||||
## Schema Diagram
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
DOCUMENT {
|
||||
UUID id PK
|
||||
TEXT name
|
||||
}
|
||||
|
||||
JOB {
|
||||
UUID id PK
|
||||
UUID document_id FK
|
||||
TEXT status
|
||||
INTEGER retry_count
|
||||
DATETIME date_created
|
||||
DATETIME date_updated
|
||||
TEXT provider
|
||||
TEXT model
|
||||
TEXT prompt_name
|
||||
TEXT text
|
||||
TEXT error_detail
|
||||
}
|
||||
|
||||
SOURCE {
|
||||
UUID id PK
|
||||
UUID document_id FK
|
||||
UUID job_id FK
|
||||
TEXT upload_name
|
||||
TEXT filename
|
||||
TEXT file_path
|
||||
DATETIME date_uploaded
|
||||
}
|
||||
|
||||
REVISION {
|
||||
UUID id PK
|
||||
UUID source_id "FK, UK"
|
||||
INTEGER revision
|
||||
TEXT text
|
||||
DATETIME date_created
|
||||
}
|
||||
|
||||
DOCUMENT ||--o{ SOURCE : has_many
|
||||
DOCUMENT ||--o{ JOB : has_many
|
||||
JOB ||--o{ SOURCE : referenced_by
|
||||
SOURCE ||--o| REVISION : has_optional_one
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Table Relationships and Constraints
|
||||
|
||||
- A `Document` can have zero or more `Source` records.
|
||||
- A `Document` can have zero or more `Job` records.
|
||||
- A `Source` belongs to exactly one `Document` and one `Job`.
|
||||
- A `Source` may have one optional `Revision`.
|
||||
- Optional `0..1` revision cardinality is enforced by uniqueness on `revision.source_id`.
|
||||
|
||||
### Invariants
|
||||
|
||||
- `Job.text` stores immutable original provider transcription output.
|
||||
- `Revision` rows are optional user-authored edits derived from original transcription.
|
||||
- Revisions do not overwrite original `Job.text`.
|
||||
- Job status lifecycle values are: `queued`, `processing`, `transcribed`, `failed`.
|
||||
|
||||
### Timestamp Fields
|
||||
|
||||
- `Job.date_created`
|
||||
- `Job.date_updated`
|
||||
- `Source.date_uploaded`
|
||||
- `Revision.date_created`
|
||||
|
||||
---
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v1.md)
|
||||
- [System Design Intent](intent.md)
|
||||
- [Transcription Methodology](transcription_methodology.md)
|
||||
- [System Architecture](architecture_v1.md)
|
||||
- [System Requirements](requirements_v1.md)
|
||||
- Data model (this document)
|
||||
- [Error Handling Policy](error_handling_v1.md)
|
||||
- [Implementation Plan](implementation_plan_v1.md)
|
||||
|
||||
## Glossary
|
||||
|
||||
- **Document**: logical grouping for one or more transcribed sources.
|
||||
- **Source**: uploaded file content (image/PDF) linked to a job.
|
||||
- **Job**: processing record that stores lifecycle status and original output.
|
||||
- **Revision**: optional single user-authored edited text linked to a source.
|
||||
@@ -1,40 +0,0 @@
|
||||
# V1 Traceability Matrix
|
||||
|
||||
This matrix provides implementation and validation evidence for V1 requirements (`REQ-0` through `REQ-13`).
|
||||
|
||||
Status values:
|
||||
|
||||
- `done`: implemented and evidence recorded
|
||||
- `in progress`: partially implemented or evidence incomplete
|
||||
- `not started`: no implementation/evidence yet
|
||||
|
||||
## Requirement Evidence Table
|
||||
|
||||
| Requirement | Status | Implementation Evidence | Validation Evidence |
|
||||
| --- | --- | --- | --- |
|
||||
| REQ-0 | done | End-to-end upload + worker pipeline in `src/transcription/services/store.py`, `src/transcription/worker.py`, `src/transcription/services/workflows.py` | `tests/integration/test_pipeline_flow.py` |
|
||||
| REQ-1 | done | Upload UI/page flow in `src/transcription/ui/pages/upload_page.py`, `src/transcription/ui/components/upload.py` | `tests/ui/test_upload_page.py`, `tests/integration/test_pipeline_flow.py` |
|
||||
| REQ-2 | done | Async worker execution and provider call orchestration in `src/transcription/worker.py`, `src/transcription/services/workflows.py` | `tests/integration/test_pipeline_flow.py`, `tests/services/test_workflows_reliability.py` |
|
||||
| REQ-3 | done | Job lifecycle state model + transitions in `src/transcription/models.py`, `src/transcription/services/jobs.py`, `src/transcription/services/workflows.py` | `tests/services/test_job_service.py`, `tests/ui/test_jobs_page.py` |
|
||||
| REQ-4 | done | Persistence of original output and failure detail in `src/transcription/services/transcription.py`, `src/transcription/services/workflows.py` | `tests/integration/test_pipeline_flow.py`, `tests/services/test_workflows_reliability.py` |
|
||||
| REQ-5 | done | Status/result inspection via UI pages and API health route in `src/transcription/ui/pages/jobs_page.py`, `src/transcription/api/health.py` | `tests/ui/test_jobs_page.py`, `tests/ui/test_pages_registration.py`, `tests/api/test_health.py` |
|
||||
| REQ-6 | done | Background processing trigger/worker notifier and non-blocking workflow in `src/transcription/ui/components/upload.py`, `src/transcription/worker.py` | `tests/test_app.py`, `tests/services/test_workflows_reliability.py` |
|
||||
| REQ-7 | done | Lifespan-owned runtime resources in `src/transcription/app.py`, `src/transcription/db/runtime.py` | `tests/test_app.py`, `tests/test_db.py` |
|
||||
| REQ-8 | done | Centralized settings/logging initialization in `src/transcription/config.py`, `src/transcription/app.py` | `tests/test_config.py`, `tests/test_app.py` |
|
||||
| REQ-9 | done | Containerized runtime baseline in `docker-compose.yml`, `Dockerfile` | `release_checklist_v1.md` (Ops checklist), manual demonstration step |
|
||||
| REQ-10 | done | Explicit schema bootstrap policy + runtime controls in `src/transcription/config.py`, `src/transcription/app.py`, `src/transcription/db/operations.py` | `tests/test_db.py`, `tests/test_config.py` |
|
||||
| REQ-11 | done | Service/workflow persistence boundaries in `src/transcription/services/*.py`, `src/transcription/services/workflows.py` | `tests/services/test_job_service.py`, `tests/services/test_transcription_service.py` |
|
||||
| REQ-12 | done | Prompt artifacts in `prompts/` and loading/validation in `src/transcription/services/transcription.py` | `tests/test_prompts.py` |
|
||||
| REQ-13 | done | Optional single revision create/update/view/delete in `src/transcription/services/transcription.py`, `src/transcription/ui/pages/jobs_page.py` | `tests/services/test_transcription_service.py`, `tests/ui/test_jobs_page.py` |
|
||||
|
||||
## Operational Evidence (Step 3 Artifacts)
|
||||
|
||||
- Runbook: `runbook_v1.md`
|
||||
- Migration/backfill/rollback guidance: `migration_v1.md`
|
||||
- Release readiness checklist: `release_checklist_v1.md`
|
||||
|
||||
## Verification Cadence
|
||||
|
||||
- Per change: maintain `tests/test_traceability.py` mappings for touched requirements.
|
||||
- Per milestone: update this table status and evidence links.
|
||||
- Pre-release: confirm all rows are `done` and non-external suite is green.
|
||||
@@ -1,35 +0,0 @@
|
||||
# AI Coding Assistant Project Briefing & Context
|
||||
|
||||
## Project Mission
|
||||
This application is a family history archival and transcription platform. Its primary goal is to accept scanned document images (letters, postcards, logbooks, diaries), execute OCR and structured transcription via AI vision models (OpenAI GPT-4o, Anthropic Claude 3.5 Sonnet), and manage historical metadata (authors, recipients, dates, and locations).
|
||||
|
||||
---
|
||||
|
||||
## Technical Stack & Architecture
|
||||
* **Database:** PostgreSQL 13+ with native `UUID` (`gen_random_uuid()`) and `JSONB` columns.
|
||||
* **Backend Runtime / Concurrency:** Python utilizing `asyncio` for concurrent HTTP API calls to AI providers, with strict rate-limiting via `asyncio.Semaphore`.
|
||||
* **Validation & Types:** Python with **Pydantic** model definitions. Incoming AI responses must be parsed and validated with Pydantic models *before* database insertion.
|
||||
* **ORM / Database Access:** SQLModel and SQLAlchemy, using parameterized statements and PostgreSQL-native types.
|
||||
|
||||
---
|
||||
|
||||
## Core System Directives for AI Code Generation
|
||||
|
||||
### 1. Data Immutability vs. Human Corrections
|
||||
* `job_source.raw_transcription` and `source.raw_transcription` represent original, point-in-time machine outputs and are **immutable**.
|
||||
* Human corrections occur on `source.revised_text`.
|
||||
* When fetching text for the UI, always display `COALESCE(source.revised_text, source.raw_transcription)`.
|
||||
|
||||
### 2. Async Execution & Batching Rules
|
||||
* A `job` represents an overarching execution run for a folder/group of images belonging to a single `document`.
|
||||
* Images are submitted to AI APIs **one at a time in rapid succession** using `asyncio` worker pools.
|
||||
* Each single-image API call populates a row in `job_source` with its own `status`, `raw_transcription`, `ai_metadata`, and `raw_api_response`.
|
||||
* If 9 of 10 pages succeed and 1 fails, `job_source.status` for the failed image becomes `'failed'`, while `job.status` becomes `'partial_success'`. Do not mark the entire batch as failed if partial results exist.
|
||||
|
||||
### 3. Entity Relationships
|
||||
* **Authors/Recipients:** A `document` can have multiple authors and recipients. Do NOT put direct `author_id` foreign keys on `document`. Query authors/recipients via `document_person` where `role = 'author'` or `role = 'recipient'`.
|
||||
* **Page Ordering:** Multi-page documents must always be queried using `ORDER BY page_number ASC`.
|
||||
|
||||
### 4. Database Mutations
|
||||
* Always use parameterized SQL queries (`$1`, `$2`) to prevent SQL injection.
|
||||
* Store datetimes using UTC ISO 8601 strings or native PostgreSQL `TIMESTAMPTZ`.
|
||||
@@ -1,416 +0,0 @@
|
||||
# SQLModel Table Models
|
||||
|
||||
These models implement the canonical [Version 2 database schema](../schema_v2.md). Each schema entity is represented by exactly one `SQLModel` table class. Because `SQLModel` is built on Pydantic and SQLAlchemy, these classes provide application validation and PostgreSQL mappings without parallel row and create models.
|
||||
|
||||
Database-generated UUIDs and timestamps are `None` until PostgreSQL supplies their values during insert. The database columns remain non-nullable. `Person.metadata_` maps to the `metadata` column because `metadata` is reserved by SQLAlchemy's declarative API.
|
||||
|
||||
```python
|
||||
from datetime import date
|
||||
from datetime import datetime
|
||||
from enum import StrEnum
|
||||
from uuid import UUID
|
||||
|
||||
from pydantic import JsonValue
|
||||
from sqlalchemy import Column
|
||||
from sqlalchemy import Date
|
||||
from sqlalchemy import DateTime
|
||||
from sqlalchemy import ForeignKey
|
||||
from sqlalchemy import Index
|
||||
from sqlalchemy import Integer
|
||||
from sqlalchemy import String
|
||||
from sqlalchemy import Text
|
||||
from sqlalchemy import UniqueConstraint
|
||||
from sqlalchemy import text
|
||||
from sqlalchemy.dialects.postgresql import JSONB
|
||||
from sqlalchemy.dialects.postgresql import UUID as PostgreSQLUUID
|
||||
from sqlmodel import Field
|
||||
from sqlmodel import Relationship
|
||||
from sqlmodel import SQLModel
|
||||
|
||||
|
||||
class PersonRole(StrEnum):
|
||||
AUTHOR = "author"
|
||||
RECIPIENT = "recipient"
|
||||
|
||||
|
||||
class JobStatus(StrEnum):
|
||||
QUEUED = "queued"
|
||||
PROCESSING = "processing"
|
||||
COMPLETED = "completed"
|
||||
PARTIAL_SUCCESS = "partial_success"
|
||||
FAILED = "failed"
|
||||
|
||||
|
||||
class JobSourceStatus(StrEnum):
|
||||
PENDING = "pending"
|
||||
TRANSCRIBED = "transcribed"
|
||||
FAILED = "failed"
|
||||
|
||||
|
||||
class Person(SQLModel, table=True):
|
||||
__tablename__ = "person"
|
||||
__table_args__ = (Index("idx_person_full_name", "full_name"),)
|
||||
|
||||
id: UUID | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
PostgreSQLUUID(as_uuid=True),
|
||||
primary_key=True,
|
||||
server_default=text("gen_random_uuid()"),
|
||||
),
|
||||
)
|
||||
full_name: str = Field(sa_column=Column(Text, nullable=False))
|
||||
display_name: str | None = Field(default=None, sa_column=Column(Text))
|
||||
maiden_name: str | None = Field(default=None, sa_column=Column(Text))
|
||||
birth_date: date | None = Field(default=None, sa_column=Column(Date))
|
||||
birth_date_raw: str | None = Field(default=None, sa_column=Column(Text))
|
||||
birth_place: str | None = Field(default=None, sa_column=Column(Text))
|
||||
death_date: date | None = Field(default=None, sa_column=Column(Date))
|
||||
death_date_raw: str | None = Field(default=None, sa_column=Column(Text))
|
||||
death_place: str | None = Field(default=None, sa_column=Column(Text))
|
||||
biography: str | None = Field(default=None, sa_column=Column(Text))
|
||||
portrait_path: str | None = Field(default=None, sa_column=Column(Text))
|
||||
metadata_: JsonValue | None = Field(
|
||||
default_factory=dict,
|
||||
sa_column=Column(
|
||||
"metadata",
|
||||
JSONB,
|
||||
server_default=text("'{}'::jsonb"),
|
||||
),
|
||||
)
|
||||
created_at: datetime | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
DateTime(timezone=True),
|
||||
nullable=False,
|
||||
server_default=text("now()"),
|
||||
),
|
||||
)
|
||||
updated_at: datetime | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
DateTime(timezone=True),
|
||||
nullable=False,
|
||||
server_default=text("now()"),
|
||||
),
|
||||
)
|
||||
|
||||
document_people: list["DocumentPerson"] = Relationship(
|
||||
back_populates="person",
|
||||
sa_relationship_kwargs={"lazy": "raise", "passive_deletes": True},
|
||||
)
|
||||
|
||||
|
||||
class Document(SQLModel, table=True):
|
||||
__tablename__ = "document"
|
||||
__table_args__ = (Index("idx_document_date", "document_date"),)
|
||||
|
||||
id: UUID | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
PostgreSQLUUID(as_uuid=True),
|
||||
primary_key=True,
|
||||
server_default=text("gen_random_uuid()"),
|
||||
),
|
||||
)
|
||||
name: str = Field(sa_column=Column(Text, nullable=False))
|
||||
document_type: str | None = Field(default=None, sa_column=Column(Text))
|
||||
document_date: date | None = Field(default=None, sa_column=Column(Date))
|
||||
document_date_raw: str | None = Field(default=None, sa_column=Column(Text))
|
||||
location_created: str | None = Field(default=None, sa_column=Column(Text))
|
||||
notes: str | None = Field(default=None, sa_column=Column(Text))
|
||||
archive_identifier: str | None = Field(default=None, sa_column=Column(Text))
|
||||
created_at: datetime | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
DateTime(timezone=True),
|
||||
nullable=False,
|
||||
server_default=text("now()"),
|
||||
),
|
||||
)
|
||||
updated_at: datetime | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
DateTime(timezone=True),
|
||||
nullable=False,
|
||||
server_default=text("now()"),
|
||||
),
|
||||
)
|
||||
|
||||
document_people: list["DocumentPerson"] = Relationship(
|
||||
back_populates="document",
|
||||
sa_relationship_kwargs={"lazy": "raise", "passive_deletes": True},
|
||||
)
|
||||
jobs: list["Job"] = Relationship(
|
||||
back_populates="document",
|
||||
sa_relationship_kwargs={"lazy": "raise", "passive_deletes": True},
|
||||
)
|
||||
sources: list["Source"] = Relationship(
|
||||
back_populates="document",
|
||||
sa_relationship_kwargs={"lazy": "raise", "passive_deletes": True},
|
||||
)
|
||||
|
||||
|
||||
class DocumentPerson(SQLModel, table=True):
|
||||
__tablename__ = "document_person"
|
||||
__table_args__ = (
|
||||
UniqueConstraint(
|
||||
"document_id",
|
||||
"person_id",
|
||||
"role",
|
||||
name="unique_document_person_role",
|
||||
),
|
||||
Index("idx_document_person_doc", "document_id"),
|
||||
Index("idx_document_person_per", "person_id"),
|
||||
)
|
||||
|
||||
id: UUID | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
PostgreSQLUUID(as_uuid=True),
|
||||
primary_key=True,
|
||||
server_default=text("gen_random_uuid()"),
|
||||
),
|
||||
)
|
||||
document_id: UUID = Field(
|
||||
sa_column=Column(
|
||||
PostgreSQLUUID(as_uuid=True),
|
||||
ForeignKey("document.id", ondelete="CASCADE"),
|
||||
nullable=False,
|
||||
),
|
||||
)
|
||||
person_id: UUID = Field(
|
||||
sa_column=Column(
|
||||
PostgreSQLUUID(as_uuid=True),
|
||||
ForeignKey("person.id", ondelete="CASCADE"),
|
||||
nullable=False,
|
||||
),
|
||||
)
|
||||
role: PersonRole = Field(sa_column=Column(String(20), nullable=False))
|
||||
created_at: datetime | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
DateTime(timezone=True),
|
||||
nullable=False,
|
||||
server_default=text("now()"),
|
||||
),
|
||||
)
|
||||
|
||||
document: Document | None = Relationship(
|
||||
back_populates="document_people",
|
||||
sa_relationship_kwargs={"lazy": "raise"},
|
||||
)
|
||||
person: Person | None = Relationship(
|
||||
back_populates="document_people",
|
||||
sa_relationship_kwargs={"lazy": "raise"},
|
||||
)
|
||||
|
||||
|
||||
class Job(SQLModel, table=True):
|
||||
__tablename__ = "job"
|
||||
__table_args__ = (Index("idx_job_document", "document_id"),)
|
||||
|
||||
id: UUID | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
PostgreSQLUUID(as_uuid=True),
|
||||
primary_key=True,
|
||||
server_default=text("gen_random_uuid()"),
|
||||
),
|
||||
)
|
||||
document_id: UUID = Field(
|
||||
sa_column=Column(
|
||||
PostgreSQLUUID(as_uuid=True),
|
||||
ForeignKey("document.id", ondelete="CASCADE"),
|
||||
nullable=False,
|
||||
),
|
||||
)
|
||||
status: JobStatus = Field(
|
||||
default=JobStatus.QUEUED,
|
||||
sa_column=Column(
|
||||
String(50),
|
||||
nullable=False,
|
||||
server_default=text("'queued'"),
|
||||
),
|
||||
)
|
||||
retry_count: int = Field(
|
||||
default=0,
|
||||
sa_column=Column(
|
||||
Integer,
|
||||
nullable=False,
|
||||
server_default=text("0"),
|
||||
),
|
||||
)
|
||||
provider: str = Field(sa_column=Column(Text, nullable=False))
|
||||
model: str = Field(sa_column=Column(Text, nullable=False))
|
||||
prompt_name: str | None = Field(default=None, sa_column=Column(Text))
|
||||
date_created: datetime | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
DateTime(timezone=True),
|
||||
nullable=False,
|
||||
server_default=text("now()"),
|
||||
),
|
||||
)
|
||||
date_updated: datetime | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
DateTime(timezone=True),
|
||||
nullable=False,
|
||||
server_default=text("now()"),
|
||||
),
|
||||
)
|
||||
|
||||
document: Document | None = Relationship(
|
||||
back_populates="jobs",
|
||||
sa_relationship_kwargs={"lazy": "raise"},
|
||||
)
|
||||
job_sources: list["JobSource"] = Relationship(
|
||||
back_populates="job",
|
||||
sa_relationship_kwargs={"lazy": "raise", "passive_deletes": True},
|
||||
)
|
||||
|
||||
|
||||
class Source(SQLModel, table=True):
|
||||
__tablename__ = "source"
|
||||
__table_args__ = (
|
||||
Index("idx_source_document", "document_id"),
|
||||
Index("idx_source_page_order", "document_id", "page_number"),
|
||||
)
|
||||
|
||||
id: UUID | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
PostgreSQLUUID(as_uuid=True),
|
||||
primary_key=True,
|
||||
server_default=text("gen_random_uuid()"),
|
||||
),
|
||||
)
|
||||
document_id: UUID = Field(
|
||||
sa_column=Column(
|
||||
PostgreSQLUUID(as_uuid=True),
|
||||
ForeignKey("document.id", ondelete="CASCADE"),
|
||||
nullable=False,
|
||||
),
|
||||
)
|
||||
page_number: int = Field(
|
||||
default=1,
|
||||
sa_column=Column(
|
||||
Integer,
|
||||
nullable=False,
|
||||
server_default=text("1"),
|
||||
),
|
||||
)
|
||||
upload_name: str = Field(sa_column=Column(Text, nullable=False))
|
||||
filename: str = Field(sa_column=Column(Text, nullable=False))
|
||||
file_path: str = Field(sa_column=Column(Text, nullable=False))
|
||||
raw_transcription: str | None = Field(default=None, sa_column=Column(Text))
|
||||
revised_text: str | None = Field(default=None, sa_column=Column(Text))
|
||||
date_uploaded: datetime | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
DateTime(timezone=True),
|
||||
nullable=False,
|
||||
server_default=text("now()"),
|
||||
),
|
||||
)
|
||||
date_revised: datetime | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(DateTime(timezone=True)),
|
||||
)
|
||||
|
||||
document: Document | None = Relationship(
|
||||
back_populates="sources",
|
||||
sa_relationship_kwargs={"lazy": "raise"},
|
||||
)
|
||||
job_sources: list["JobSource"] = Relationship(
|
||||
back_populates="source",
|
||||
sa_relationship_kwargs={"lazy": "raise", "passive_deletes": True},
|
||||
)
|
||||
|
||||
|
||||
class JobSource(SQLModel, table=True):
|
||||
__tablename__ = "job_source"
|
||||
__table_args__ = (
|
||||
UniqueConstraint("job_id", "source_id", name="unique_job_source"),
|
||||
Index("idx_job_source_job", "job_id"),
|
||||
Index("idx_job_source_source", "source_id"),
|
||||
Index(
|
||||
"idx_job_source_ai_metadata",
|
||||
"ai_metadata",
|
||||
postgresql_using="gin",
|
||||
),
|
||||
)
|
||||
|
||||
id: UUID | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
PostgreSQLUUID(as_uuid=True),
|
||||
primary_key=True,
|
||||
server_default=text("gen_random_uuid()"),
|
||||
),
|
||||
)
|
||||
job_id: UUID = Field(
|
||||
sa_column=Column(
|
||||
PostgreSQLUUID(as_uuid=True),
|
||||
ForeignKey("job.id", ondelete="CASCADE"),
|
||||
nullable=False,
|
||||
),
|
||||
)
|
||||
source_id: UUID = Field(
|
||||
sa_column=Column(
|
||||
PostgreSQLUUID(as_uuid=True),
|
||||
ForeignKey("source.id", ondelete="CASCADE"),
|
||||
nullable=False,
|
||||
),
|
||||
)
|
||||
status: JobSourceStatus = Field(
|
||||
default=JobSourceStatus.PENDING,
|
||||
sa_column=Column(
|
||||
String(50),
|
||||
nullable=False,
|
||||
server_default=text("'pending'"),
|
||||
),
|
||||
)
|
||||
raw_transcription: str | None = Field(default=None, sa_column=Column(Text))
|
||||
ai_metadata: JsonValue | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(JSONB),
|
||||
)
|
||||
raw_api_response: JsonValue | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(JSONB),
|
||||
)
|
||||
error_detail: str | None = Field(default=None, sa_column=Column(Text))
|
||||
executed_at: datetime | None = Field(
|
||||
default=None,
|
||||
sa_column=Column(
|
||||
DateTime(timezone=True),
|
||||
nullable=False,
|
||||
server_default=text("now()"),
|
||||
),
|
||||
)
|
||||
|
||||
job: Job | None = Relationship(
|
||||
back_populates="job_sources",
|
||||
sa_relationship_kwargs={"lazy": "raise"},
|
||||
)
|
||||
source: Source | None = Relationship(
|
||||
back_populates="job_sources",
|
||||
sa_relationship_kwargs={"lazy": "raise"},
|
||||
)
|
||||
```
|
||||
|
||||
The enum annotations validate application values while the mapped columns retain the `VARCHAR` types specified by the DDL. PostgreSQL owns generated UUIDs and timestamps through `server_default`; call `session.refresh(instance)` after a flush or commit when those generated values are needed immediately.
|
||||
|
||||
`ai_metadata`, `raw_api_response`, and `metadata_` accept any JSON value supported by `JSONB`. Validate provider-specific payload structure before assigning it to these fields, while preserving the complete raw response in `raw_api_response`.
|
||||
|
||||
Relationships use `lazy="raise"` to prevent implicit database I/O in async code. Queries must explicitly load relationships they need, for example with `selectinload()`.
|
||||
|
||||
The schema's behavioral invariants are enforced outside the table shape where appropriate:
|
||||
|
||||
- `PersonRole`, `JobStatus`, and `JobSourceStatus` define the exact values listed by the schema.
|
||||
- `unique_document_person_role` enforces role uniqueness for `(document_id, person_id, role)`.
|
||||
- Services order document sources by `Source.document_id` and `Source.page_number`.
|
||||
- Services derive aggregate `Job.status` from related `JobSource.status` values.
|
||||
- Services preserve `JobSource.raw_transcription` and `JobSource.raw_api_response` as point-in-time outputs while updating the active text on `Source`.
|
||||
@@ -1,136 +0,0 @@
|
||||
# System Architecture (Version 2)
|
||||
|
||||
This document describes the V2 production architecture of the personal historical-document transcription system.
|
||||
|
||||
## Architecture Objectives
|
||||
|
||||
* Preserve source material as immutable transcribed text alongside page-level spatial AI metadata.
|
||||
* Support batching multi-image and folder uploads cleanly into sequential pages (`page_number`).
|
||||
* Leverage asynchronous worker pools (`asyncio`) for parallel single-image API execution bounded by rate limiters (`asyncio.Semaphore`).
|
||||
* Migrate persistence to PostgreSQL using native `UUID`, `TIMESTAMPTZ`, and `JSONB` document storage.
|
||||
* Standardize all data validation, API parsing, and database models on **Pydantic V2**.
|
||||
* Support rich historical attribution (multi-author and multi-recipient relationships).
|
||||
|
||||
## Runtime Topology
|
||||
|
||||
The V2 runtime operates as an asynchronous Python application:
|
||||
|
||||
* FastAPI + NiceGUI web application process.
|
||||
* In-process `asyncio` background task orchestrator for parallel API execution.
|
||||
* Relational persistence via PostgreSQL (using `asyncpg` or `psycopg3`).
|
||||
* Pydantic V2 validation layer wrapping API payloads and PostgreSQL `JSONB` schemas.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
U[Browser User] --> A[FastAPI + NiceGUI App]
|
||||
A --> W[Asyncio Worker Engine]
|
||||
A --> DB[(PostgreSQL Database)]
|
||||
W --> P[Vision Provider APIs\nOpenAI / Claude]
|
||||
W --> DB
|
||||
```
|
||||
|
||||
## Lifecycle Ownership
|
||||
|
||||
Application lifespan owns runtime setup/teardown:
|
||||
|
||||
* Initialize environment logging and Pydantic configuration.
|
||||
* Manage asynchronous PostgreSQL connection pools (`asyncpg` / `psycopg3`).
|
||||
* Execute database migrations and index initialization.
|
||||
* Recover stale processing jobs on startup.
|
||||
* Manage graceful shutdown of active `asyncio` worker pools.
|
||||
|
||||
## Layered Module Structure
|
||||
|
||||
### Interface Layer
|
||||
|
||||
* `src/transcription/ui/**` (NiceGUI pages, multi-page renderers, person cards)
|
||||
* `src/transcription/api/**` (FastAPI routes and JSON error handlers)
|
||||
|
||||
### Application & Async Worker Layer
|
||||
|
||||
* `src/transcription/services/workflows.py`
|
||||
* `src/transcription/worker.py`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
* Batch orchestration and status transitions (`queued` -> `processing` -> `completed` | `partial_success` | `failed`).
|
||||
* Parallel single-image API execution using `asyncio.gather` bounded by `asyncio.Semaphore`.
|
||||
* Pydantic schema parsing (`PageAIMetadata`) and validation prior to database storage.
|
||||
|
||||
### Domain & Service Layer
|
||||
|
||||
* `src/transcription/db/models.py` (SQLModel/Pydantic V2 schema definitions for the current implementation)
|
||||
* `src/transcription/services/*.py` (Transactional operations for `Document`, `Person`, `Source`, `Job`, and `JobSource`)
|
||||
|
||||
### Infrastructure Layer
|
||||
|
||||
* `src/transcription/db/**` (PostgreSQL connection pooling and raw parameterized SQL execution)
|
||||
* `src/transcription/providers/**` (OpenAI & Anthropic Vision SDK adapters)
|
||||
|
||||
## Processing Workflow
|
||||
|
||||
1. User uploads a folder or batch of images for a `Document`.
|
||||
2. System creates `Document`, `Job(status='queued')`, and ordered `Source` pages (`page_number = 1..N`).
|
||||
3. Worker claims job, sets `Job.status = 'processing'`, and spawns parallel `asyncio` tasks bounded by semaphore.
|
||||
4. Each task calls Vision API for a **single** `Source` image.
|
||||
5. On task completion:
|
||||
* Writes a `JobSource` record containing `status='transcribed'`, `raw_transcription`, `ai_metadata` (bounding boxes/confidence), and `raw_api_response`.
|
||||
* Caches active text to `Source.raw_transcription`.
|
||||
|
||||
|
||||
6. On page failure:
|
||||
* Writes `JobSource` record with `status='failed'` and `error_detail`.
|
||||
|
||||
|
||||
7. Once all page tasks resolve:
|
||||
* Marks `Job.status` as `completed` (100% success), `partial_success` (at least 1 success, 1 failure), or `failed` (all failed).
|
||||
|
||||
|
||||
|
||||
## Domain Ownership & Invariants
|
||||
|
||||
* **Immutable AI Outputs:** `source.raw_transcription` and `job_source.raw_transcription` store original, point-in-time machine output and are immutable.
|
||||
* **Inlined Revisions:** Human corrections occur on `source.revised_text`. UI renders `COALESCE(revised_text, raw_transcription)`.
|
||||
* **Sequential Integrity:** Multi-page documents are strictly ordered by `source.page_number ASC`.
|
||||
* **Page Execution Isolation:** A failure on one page image does not invalidate successful transcriptions on sister pages in the same batch job.
|
||||
|
||||
## Data Model Summary
|
||||
|
||||
* `Document` has many `Source` pages, many `Job` runs, and many `Person` records via `DocumentPerson` junction (`author` or `recipient`).
|
||||
* `Source` belongs to one `Document` and can be processed across many `JobSource` executions.
|
||||
* `Job` has many `JobSource` execution records.
|
||||
|
||||
## Test Strategy
|
||||
|
||||
* Unit tests for Pydantic V2 schemas, custom validators, and JSONB serialization.
|
||||
* Integration tests for async PostgreSQL connection handling and parameterized queries.
|
||||
* Async workflow tests using mock AI providers to verify `partial_success` and retry logic.
|
||||
* UI integration tests for multi-page rendering and person management.
|
||||
|
||||
---
|
||||
|
||||
## Technology References
|
||||
|
||||
- [FastAPI documentation](https://fastapi.tiangolo.com/)
|
||||
- [NiceGUI documentation](https://nicegui.io/documentation)
|
||||
- [PostgreSQL documentation](https://www.postgresql.org/docs/)
|
||||
- [Python asyncio](https://docs.python.org/3/library/asyncio.html#module-asyncio)
|
||||
- [Pydantic Validation](https://pydantic.dev/docs/validation/latest/get-started/)
|
||||
- [Pydantic AI](https://pydantic.dev/docs/ai/overview/)
|
||||
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v2.md)
|
||||
- [System Design Intent](invariant/intent.md)
|
||||
- [Transcription Methodology](invariant/transcription_methodology.md)
|
||||
- System Architecture (this document)
|
||||
- [System Requirements](requirements_v2.md)
|
||||
- [Data model](schema_v2.md)
|
||||
- [Error Handling Policy](error_handling_v2.md)
|
||||
- [Implementation Plan](implementation_plan_v2.md)
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
@@ -1,88 +0,0 @@
|
||||
# Error Handling Policy (Version 2)
|
||||
|
||||
This document defines the canonical error-handling policy for the V2 document transcription system.
|
||||
|
||||
## Error Handling Objectives
|
||||
|
||||
* Make failures visible in clear, actionable language at both the document and individual page levels.
|
||||
* Support **isolated failure handling** in multi-image batches so single page errors do not crash an entire batch job.
|
||||
* Preserve diagnostic detail (Pydantic validation errors, raw provider responses) in PostgreSQL `JSONB` for fast troubleshooting.
|
||||
* Ensure consistent error envelope structure across API, UI, and async worker boundaries.
|
||||
|
||||
## Scope And Authority
|
||||
|
||||
Governs error behavior across NiceGUI pages, FastAPI routes, service orchestration, `asyncio` background tasks, PostgreSQL interactions, and AI provider adapters.
|
||||
|
||||
## Error Taxonomy
|
||||
|
||||
| Category | Definition | Retriable |
|
||||
| --- | --- | --- |
|
||||
| `validation_error` | Pydantic payload or parameter schema validation failure | no |
|
||||
| `user_input_error` | Unacceptable user file (unsupported image type, corrupt file) | no |
|
||||
| `not_found_error` | Requested resource (`Document`, `Source`, `Person`, `Job`) missing | no |
|
||||
| `conflict_error` | Operation violates state constraints (e.g., duplicate `document_person` role) | no |
|
||||
| `external_provider_error` | AI Provider API failure (rate limit, vision execution error) | yes |
|
||||
| `infrastructure_transient_error` | Temporary DB connection reset or HTTP timeout | yes |
|
||||
| `infrastructure_persistent_error` | Database down, missing API credentials, misconfiguration | no |
|
||||
| `internal_unexpected_error` | Uncaught Python exception or logic defect | no |
|
||||
|
||||
## Async Batch & Page-Level Error Behavior
|
||||
|
||||
In multi-image `asyncio` batch processing:
|
||||
|
||||
1. **Page Isolation:** Exceptions caught during individual page calls are caught within the `asyncio` task wrapper.
|
||||
2. **Page Record Logging:** Page failure detail is written directly to `job_source.error_detail` and `job_source.status = 'failed'`.
|
||||
3. **Batch Aggregate State:**
|
||||
* If **all** page tasks succeed -> `job.status = 'completed'`.
|
||||
* If **some** page tasks fail -> `job.status = 'partial_success'`.
|
||||
* If **all** page tasks fail -> `job.status = 'failed'`.
|
||||
|
||||
|
||||
4. **Retry Strategy:** The UI exposes a "Retry Failed Pages" option for `partial_success` jobs, which spawns a new targeted `Job` containing *only* the `Source` IDs marked as `failed`.
|
||||
|
||||
## API Error Response Contract
|
||||
|
||||
API error responses return a structured JSON envelope:
|
||||
```json
|
||||
{
|
||||
"error_id": "err_uuid_12345",
|
||||
"category": "validation_error",
|
||||
"message": "The uploaded payload failed schema validation.",
|
||||
"suggestion": "Check file format and metadata fields, then try again.",
|
||||
"details": {
|
||||
"pydantic_errors": [...]
|
||||
},
|
||||
"timestamp": "2026-07-31T07:55:00Z"
|
||||
}
|
||||
```
|
||||
|
||||
HTTP Status Mappings:
|
||||
|
||||
* `validation_error`, `user_input_error` -> `400`
|
||||
* `not_found_error` -> `404`
|
||||
* `conflict_error` -> `409`
|
||||
* `external_provider_error` -> `502` / `503`
|
||||
* `infrastructure_transient_error` -> `503`
|
||||
* `infrastructure_persistent_error`, `internal_unexpected_error` -> `500`
|
||||
|
||||
---
|
||||
|
||||
## Technology References
|
||||
|
||||
- [FastAPI documentation](https://fastapi.tiangolo.com/)
|
||||
- [NiceGUI documentation](https://nicegui.io/documentation)
|
||||
- [PostgreSQL documentation](https://www.postgresql.org/docs/)
|
||||
- [Python asyncio](https://docs.python.org/3/library/asyncio.html#module-asyncio)
|
||||
- [Pydantic Validation](https://pydantic.dev/docs/validation/latest/get-started/)
|
||||
- [Pydantic AI](https://pydantic.dev/docs/ai/overview/)
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v2.md)
|
||||
- [System Design Intent](invariant/intent.md)
|
||||
- [Transcription Methodology](invariant/transcription_methodology.md)
|
||||
- [System Architecture](architecture_v2.md)
|
||||
- [System Requirements](requirements_v2.md)
|
||||
- [Data model](schema_v2.md)
|
||||
- Error Handling Policy (this document)
|
||||
- [Implementation Plan](implementation_plan_v2.md)
|
||||
@@ -1,64 +0,0 @@
|
||||
# implementation_plan_v2
|
||||
|
||||
## Goal
|
||||
|
||||
Replace the current V1 SQLModel schema with the approved V2 schema and make sure every database operation works through the existing async SQLAlchemy/SQLModel session layer.
|
||||
|
||||
Use a fresh database. There will be no migrations, data conversion, legacy compatibility shims, or parallel V1/V2 code paths.
|
||||
|
||||
## Current Project Impact
|
||||
|
||||
- `src/transcription/db/models.py` still defines the V1 `Document`, `Source`, `Job`, and `Revision` tables.
|
||||
- The V2 target adds `Person`, `DocumentPerson`, and `JobSource`, moves revisions onto `Source`, and removes the direct `Source.job_id` relationship.
|
||||
- The engine, session factory, transaction handling, and PostgreSQL async support already exist and do not need to be rewritten.
|
||||
- Async CRUD currently lives in `DocumentService`, `JobService`, `TranscriptionService`, and the upload record helper. Their queries and eager-loading options depend on V1 relationships.
|
||||
- Existing tests cover only part of the schema and CRUD surface.
|
||||
|
||||
## Implementation
|
||||
|
||||
### 1. Update the schema
|
||||
|
||||
- Replace the models in `src/transcription/db/models.py` with the approved V2 tables, enums, relationships, foreign keys, constraints, and indexes.
|
||||
- Remove `Revision`, `Source.job_id`, and the transcription fields that no longer belong on `Job`.
|
||||
- Keep `create_all()` as the schema bootstrap for a fresh database.
|
||||
- Delete `_ensure_sqlite_compat_columns()` and all schema patching from `src/transcription/db/operations.py`.
|
||||
- Keep the Python models and `docs/schema_v2.md` consistent.
|
||||
|
||||
### 2. Align the async CRUD methods
|
||||
|
||||
- Keep the existing `ServiceBase` session and transaction pattern.
|
||||
- Update document CRUD to load and manage its ordered `Source` rows and `DocumentPerson` links.
|
||||
- Update job CRUD and queue queries to use `JobSource` instead of `Source.job_id`.
|
||||
- Add the missing async CRUD operations for `Person`, `Source`, `DocumentPerson`, and `JobSource` using the existing service style. Do not add another repository abstraction.
|
||||
- Replace revision CRUD with direct updates to `Source.revised_text` and `Source.date_revised`.
|
||||
- Remove the temporary transcript compatibility aliases instead of redirecting them.
|
||||
- Update only direct database call sites that construct or query these records; UI and worker feature changes are not part of this work.
|
||||
|
||||
### 3. Verify the schema and CRUD
|
||||
|
||||
- Update the schema bootstrap test to expect `person`, `document`, `document_person`, `source`, `job`, and `job_source`, with no `revision` table.
|
||||
- Add async create, read, update, delete, list, and filtered-query tests for each entity that exposes those operations.
|
||||
- Test relationship loading, page ordering, uniqueness constraints, delete behavior, status values, and `JobSource` JSON fields.
|
||||
- Test both service-owned sessions and caller-provided sessions so flush/commit behavior remains correct.
|
||||
- Run the focused database and service tests, then the full suite with `uv run pytest`.
|
||||
|
||||
### 4. Update the UI for the V2 schema
|
||||
|
||||
- Review the UI components and views that display document, job, person, and source data so they reference the V2 schema instead of V1 relationships.
|
||||
- Update upload, detail, and listing screens to show the new person and source associations, revised-source fields, and the revised status values.
|
||||
- Keep the UI behavior aligned with the updated service layer and ensure the existing UI tests continue to pass with the V2 data model.
|
||||
- Consider the guidance in `docs/ui_style_guide.md` when making UI changes so the updated views remain consistent with the project’s visual and interaction conventions.
|
||||
|
||||
## Done When
|
||||
|
||||
- A fresh database is created directly from the V2 SQLModel metadata.
|
||||
- All async CRUD methods pass against the V2 relationships and fields.
|
||||
- No code references `Revision`, `Source.job_id`, removed `Job` transcription fields, or compatibility aliases.
|
||||
- The focused tests and full test suite pass.
|
||||
|
||||
## Out of Scope
|
||||
|
||||
- Database migrations or preservation of V1 data
|
||||
- Legacy compatibility code
|
||||
- Database engine or session-layer rewrites
|
||||
- UI redesign, batch orchestration, worker concurrency, deployment, and operational runbooks
|
||||
@@ -1,47 +0,0 @@
|
||||
# Document Transcription System Overview (Version 2)
|
||||
|
||||
This project is a personal-scale application for transcribing, indexing, and preserving historical family documents, letters, postcards, and journals.
|
||||
|
||||
## Start Here
|
||||
|
||||
Read [architecture_v2.md](architecture_v2.md) first for technical overview and system design.
|
||||
|
||||
## Core V2 Capabilities
|
||||
|
||||
* **Folder & Multi-Image Ingestion:** Upload whole folders or image batches that map sequentially (`page_number`) under a single `Document`.
|
||||
* **Parallel Async AI Vision Engine:** Concurrently process single-page image transcriptions using Python `asyncio` bounded by rate limiters.
|
||||
* **Robust PostgreSQL Storage:** Relational storage for entities with native `UUID`, `TIMESTAMPTZ`, and `JSONB` for deep AI spatial metadata and raw envelopes.
|
||||
* **Pydantic V2 Validation:** End-to-end type safety, DB row mapping, and JSONB payload validation.
|
||||
* **Historical Person Management:** Track authors and recipients across documents with rich biographical entities (`Person`).
|
||||
* **Page-Level Execution Auditing & Revisions:** Store immutable point-in-time machine output per run while enabling inline human corrections (`revised_text`).
|
||||
* **Partial Failure Recovery:** Bounded batch execution that isolates single-page API errors (`partial_success`) for simple retries.
|
||||
|
||||
## Technical Stack
|
||||
|
||||
* **Application Web Framework:** FastAPI + NiceGUI
|
||||
* **Persistence Engine:** PostgreSQL 18+
|
||||
* **Data Validation & Schemas:** Pydantic V2
|
||||
* **Concurrency & Workers:** Python `asyncio` worker pool with `asyncio.Semaphore`
|
||||
* **Vision Providers:** OpenAI (GPT-4o) and Anthropic (Claude 3.5 Sonnet) via native SDKs
|
||||
|
||||
---
|
||||
|
||||
## Technology References
|
||||
|
||||
- [FastAPI documentation](https://fastapi.tiangolo.com/)
|
||||
- [NiceGUI documentation](https://nicegui.io/documentation)
|
||||
- [PostgreSQL documentation](https://www.postgresql.org/docs/)
|
||||
- [Python asyncio](https://docs.python.org/3/library/asyncio.html#module-asyncio)
|
||||
- [Pydantic Validation](https://pydantic.dev/docs/validation/latest/get-started/)
|
||||
- [Pydantic AI](https://pydantic.dev/docs/ai/overview/)
|
||||
|
||||
## Documentation Index
|
||||
|
||||
- System Overview (this document)
|
||||
- [System Design Intent](invariant/intent.md)
|
||||
- [Transcription Methodology](invariant/transcription_methodology.md)
|
||||
- [System Architecture](architecture_v2.md)
|
||||
- [System Requirements](requirements_v2.md)
|
||||
- [Data model](schema_v2.md)
|
||||
- [Error Handling Policy](error_handling_v2.md)
|
||||
- [Implementation Plan](implementation_plan_v2.md)
|
||||
@@ -1,42 +0,0 @@
|
||||
# Document Transcription System Requirements (Version 2)
|
||||
|
||||
This document captures the **Version 2 baseline requirements** for the production implementation.
|
||||
|
||||
## Requirements Model
|
||||
|
||||
| ID | Category | Requirement | Verify Method |
|
||||
| --- | --- | --- | --- |
|
||||
| REQ-0 | System | Provide end-to-end multi-page document transcription with persistent, inspectable async job states. | demonstration |
|
||||
| REQ-1 | Functional | Allow users to upload folders or multi-image batches as sequential `Source` pages under a `Document`. | test |
|
||||
| REQ-2 | Functional | Process multi-page jobs asynchronously using an `asyncio` worker pool bounded by rate limits. | test |
|
||||
| REQ-3 | Functional | Persist page-level execution outputs (`raw_transcription`, `ai_metadata`, `raw_api_response`) on `JobSource`. | test |
|
||||
| REQ-4 | Functional | Support job states (`queued`, `processing`, `completed`, `partial_success`, `failed`) and page states (`pending`, `transcribed`, `failed`). | inspection |
|
||||
| REQ-5 | Functional | Allow users to manage historical `Person` records and link multiple authors/recipients to a `Document` via `DocumentPerson`. | test |
|
||||
| REQ-6 | Functional | Maintain immutable original machine output on `Source.raw_transcription` while permitting inline human edits on `Source.revised_text`. | test |
|
||||
| REQ-7 | Data Constraint | Store all persistent domain data in PostgreSQL using native `UUID`, `TIMESTAMPTZ`, and `JSONB` columns. | inspection |
|
||||
| REQ-8 | Data Constraint | Validate all API requests, database rows, and JSONB structures using Pydantic V2 schemas. | test |
|
||||
| REQ-9 | Interface | Render multi-page transcriptions sequentially by `page_number` in the web UI with author/recipient metadata. | demonstration |
|
||||
| REQ-10 | Operations | Allow operators to retry only failed pages for jobs in a `partial_success` state. | test |
|
||||
|
||||
## Element Satisfaction Mapping
|
||||
|
||||
* **UI (NiceGUI):** Satisfies REQ-1, REQ-5, REQ-6, REQ-9, REQ-10.
|
||||
* **API (FastAPI):** Satisfies REQ-1, REQ-4, REQ-5, REQ-8.
|
||||
* **WORKER (asyncio):** Satisfies REQ-2, REQ-3, REQ-4, REQ-10.
|
||||
* **PERSISTENCE (PostgreSQL):** Satisfies REQ-3, REQ-6, REQ-7.
|
||||
* **MODELS (Pydantic V2):** Satisfies REQ-8.
|
||||
|
||||
---
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v2.md)
|
||||
- [System Design Intent](invariant/intent.md)
|
||||
- [Transcription Methodology](invariant/transcription_methodology.md)
|
||||
- [System Architecture](architecture_v2.md)
|
||||
- System Requirements (this document)
|
||||
- [Data model](schema_v2.md)
|
||||
- [Error Handling Policy](error_handling_v2.md)
|
||||
- [Implementation Plan](implementation_plan_v2.md)
|
||||
|
||||
|
||||
@@ -1,137 +0,0 @@
|
||||
# Database Schema (Version 2)
|
||||
|
||||
This document describes the PostgreSQL relational schema for the transcription platform. It incorporates multi-image batch orchestration via `asyncio`, page-level execution tracking, many-to-many author/recipient attribution, and JSONB document storage for AI vision outputs.
|
||||
|
||||
All primary and foreign keys are PostgreSQL native UUIDs (`gen_random_uuid()`).
|
||||
|
||||
## Entity Relationship Diagram
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
PERSON {
|
||||
UUID id PK
|
||||
TEXT full_name
|
||||
TEXT display_name
|
||||
TEXT maiden_name
|
||||
DATE birth_date
|
||||
TEXT birth_date_raw
|
||||
TEXT birth_place
|
||||
DATE death_date
|
||||
TEXT death_date_raw
|
||||
TEXT death_place
|
||||
TEXT biography
|
||||
TEXT portrait_path
|
||||
JSONB metadata
|
||||
TIMESTAMPTZ created_at
|
||||
TIMESTAMPTZ updated_at
|
||||
}
|
||||
|
||||
DOCUMENT {
|
||||
UUID id PK
|
||||
TEXT name
|
||||
TEXT document_type
|
||||
DATE document_date
|
||||
TEXT document_date_raw
|
||||
TEXT location_created
|
||||
TEXT notes
|
||||
TEXT archive_identifier
|
||||
TIMESTAMPTZ created_at
|
||||
TIMESTAMPTZ updated_at
|
||||
}
|
||||
|
||||
DOCUMENT_PERSON {
|
||||
UUID id PK
|
||||
UUID document_id FK
|
||||
UUID person_id FK
|
||||
VARCHAR role "author | recipient"
|
||||
TIMESTAMPTZ created_at
|
||||
}
|
||||
|
||||
JOB {
|
||||
UUID id PK
|
||||
UUID document_id FK
|
||||
VARCHAR status "queued | processing | transcribed | completed | partial_success | failed"
|
||||
INTEGER retry_count
|
||||
TEXT provider
|
||||
TEXT model
|
||||
TEXT prompt_name
|
||||
TIMESTAMPTZ date_created
|
||||
TIMESTAMPTZ date_updated
|
||||
}
|
||||
|
||||
SOURCE {
|
||||
UUID id PK
|
||||
UUID document_id FK
|
||||
INTEGER page_number
|
||||
TEXT upload_name
|
||||
TEXT filename
|
||||
TEXT file_path
|
||||
TEXT raw_transcription
|
||||
TEXT revised_text
|
||||
TIMESTAMPTZ date_uploaded
|
||||
TIMESTAMPTZ date_revised
|
||||
}
|
||||
|
||||
JOB_SOURCE {
|
||||
UUID id PK
|
||||
UUID job_id FK
|
||||
UUID source_id FK
|
||||
VARCHAR status "pending | transcribed | failed"
|
||||
TEXT raw_transcription
|
||||
JSONB ai_metadata
|
||||
JSONB raw_api_response
|
||||
TEXT error_detail
|
||||
TIMESTAMPTZ executed_at
|
||||
}
|
||||
|
||||
DOCUMENT ||--o{ DOCUMENT_PERSON : "has_people"
|
||||
PERSON ||--o{ DOCUMENT_PERSON : "participates_in"
|
||||
DOCUMENT ||--o{ JOB : "has_jobs"
|
||||
DOCUMENT ||--o{ SOURCE : "contains_pages"
|
||||
JOB ||--o{ JOB_SOURCE : "executes"
|
||||
SOURCE ||--o{ JOB_SOURCE : "processed_in"
|
||||
```
|
||||
|
||||
## Domain Invariants & Rules
|
||||
|
||||
### Page-Level Execution & AI Outputs
|
||||
|
||||
* Execution Granularity: Every single image execution by an AI model produces a dedicated record in job_source.
|
||||
* Source vs Execution Status: `source` does not carry a `status` column. Per-source execution state is tracked in `job_source.status` (`pending`, `transcribed`, `failed`).
|
||||
* Point-in-Time Auditability: job_source.raw_api_response stores the unparsed REST response envelope for that specific image page call. job_source.ai_metadata stores spatial bounding boxes, token usage, and layout details for that specific image page call.
|
||||
* Active Output Caching: Upon successful completion of an image call, source.raw_transcription is updated with the latest output string from job_source.raw_transcription for fast UI rendering.
|
||||
|
||||
### Page Ordering & Revisions
|
||||
|
||||
* Sequential Integrity: source.page_number dictates page ordering within a document. Reads assembling full documents must query ORDER BY source.document_id, source.page_number ASC.
|
||||
* Inlined Human Corrections: User edits occur at the page level inside source.revised_text. source.raw_transcription remains immutable. If source.revised_text is non-null, application frontends must render source.revised_text.
|
||||
|
||||
### Async Job Lifecycle & Failure Isolation
|
||||
|
||||
* Batch Orchestrator: A job represents an overarching execution run across one or more source images belonging to a document.
|
||||
* Isolated Failures: API requests run concurrently (e.g., using asyncio). A failure on page 3 does not invalidate successful transcriptions on page 1 or 2.
|
||||
* Job States:
|
||||
- queued: Created, awaiting worker execution.
|
||||
- processing: Concurrent HTTP tasks actively running.
|
||||
- completed: 100% of linked job_source tasks succeeded (transcribed).
|
||||
- partial_success: At least one job_source succeeded and at least one failed.
|
||||
- failed: All linked job_source tasks failed or a job-level runtime error occurred.
|
||||
|
||||
### Attribution & Person Roles
|
||||
|
||||
* Multi-Person Roles: Documents support zero, one, or many authors and recipients linked via document_person.
|
||||
* Role Uniqueness: (document_id, person_id, role) must be unique to prevent duplicate role tagging.
|
||||
|
||||
---
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v2.md)
|
||||
- [System Design Intent](invariant/intent.md)
|
||||
- [Transcription Methodology](invariant/transcription_methodology.md)
|
||||
- [System Architecture](architecture_v2.md)
|
||||
- [System Requirements](requirements_v2.md)
|
||||
- Data model (this document)
|
||||
- [Error Handling Policy](error_handling_v2.md)
|
||||
- [Implementation Plan](implementation_plan_v2.md)
|
||||
|
||||
@@ -1,142 +0,0 @@
|
||||
# System Architecture (Version 3)
|
||||
|
||||
This document describes the V3 production architecture of the personal historical-document transcription system.
|
||||
|
||||
## Architecture Objectives
|
||||
|
||||
* Preserve source material as immutable transcribed text alongside page-level spatial AI metadata and complete provider API envelopes.
|
||||
* Support batching multi-image and folder uploads cleanly into sequential pages (`page_number`).
|
||||
* Capture complete input prompt provenance (`system_prompt`, `user_prompt`, `prompt_hash`) and execution parameters (`temperature`, `top_p`) at submission time on `Job`.
|
||||
* Leverage asynchronous worker pools (`asyncio`) for parallel single-image API execution bounded by rate limiters (`asyncio.Semaphore`).
|
||||
* Maintain relational data-model portability across the supported backends by using SQLModel/SQLAlchemy and compatibility types so the same domain schema works in SQLite for local development/testing and PostgreSQL in production.
|
||||
* Keep operator tooling and local maintenance workflows OS-independent by using Python or other cross-platform interfaces for canonical project automation.
|
||||
* Verify image asset integrity via SHA-256 file hashing (`file_hash`) while storing binary assets on the local filesystem.
|
||||
* Standardize all data validation, API parsing, and database models on **Pydantic V2** and **SQLModel**.
|
||||
* Support rich historical attribution (multi-author and multi-recipient relationships via `DocumentPerson`).
|
||||
|
||||
## Runtime Topology
|
||||
|
||||
The V3 runtime operates as an asynchronous Python application:
|
||||
|
||||
* FastAPI + NiceGUI web application process.
|
||||
* In-process `asyncio` background task orchestrator for parallel API execution.
|
||||
* Relational persistence via SQLModel / SQLAlchemy, using SQLite for local development/testing and PostgreSQL as the production persistence target.
|
||||
* Pydantic V2 validation layer wrapping API payloads, prompt configurations, and JSON metadata schemas.
|
||||
* Cross-platform operator workflows implemented in Python so core local operations run consistently on Windows, Linux, and macOS.
|
||||
|
||||
^^^mermaid
|
||||
flowchart LR
|
||||
U[Browser User] --> A[FastAPI + NiceGUI App]
|
||||
A --> W[Asyncio Worker Engine]
|
||||
A --> DB[(Relational DB\nSQLite / PostgreSQL)]
|
||||
W --> P[Vision Provider APIs\nOpenAI / Claude / OpenRouter]
|
||||
W --> DB
|
||||
^^^
|
||||
|
||||
## Lifecycle Ownership
|
||||
|
||||
Application lifespan owns runtime setup/teardown:
|
||||
|
||||
* Initialize environment logging, directory paths, and Pydantic configuration.
|
||||
* Manage asynchronous database engine connection pools (`aiosqlite` or `asyncpg`).
|
||||
* Execute database bootstrap (`SQLModel.metadata.create_all()`) or migrations.
|
||||
* Recover stale or interrupted processing jobs on startup.
|
||||
* Manage graceful shutdown of active `asyncio` worker pools.
|
||||
|
||||
## Layered Module Structure
|
||||
|
||||
### Interface Layer
|
||||
|
||||
* `src/transcription/ui/**` (NiceGUI pages, multi-page renderers, person cards)
|
||||
* `src/transcription/api/**` (FastAPI routes and JSON error handlers)
|
||||
|
||||
### Application & Async Worker Layer
|
||||
|
||||
* `src/transcription/services/workflows.py`
|
||||
* `src/transcription/worker.py`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
* Batch orchestration and status transitions (`queued` -> `processing` -> `transcribed` | `partial_success` | `failed`).
|
||||
* Parallel single-image API execution using `asyncio.gather` bounded by `asyncio.Semaphore`.
|
||||
* Resolve prompt configuration at submission time and persist frozen snapshot fields on `Job`.
|
||||
* Pydantic schema parsing and validation prior to database storage.
|
||||
|
||||
### Domain & Service Layer
|
||||
|
||||
* `src/transcription/db/models.py` (SQLModel schema definitions for Document, Source, Job, JobSource, Person, DocumentPerson)
|
||||
* `src/transcription/services/*.py` (Transactional operations for `Document`, `Person`, `Source`, `Job`, and `JobSource`)
|
||||
|
||||
### Infrastructure Layer
|
||||
|
||||
* `src/transcription/db/**` (Async database session factory, engine creation, and JSON dialect abstractions)
|
||||
* `src/transcription/providers/**` (OpenAI, Anthropic, and OpenRouter Vision SDK adapters)
|
||||
|
||||
## Processing Workflow
|
||||
|
||||
1. User uploads a folder or batch of images for a `Document`.
|
||||
2. System hashes each image file (SHA-256), writes image files to filesystem storage, and creates `Document`, `Job(status='queued')`, and ordered `Source` pages (`page_number = 1..N`).
|
||||
3. Worker claims job, sets `Job.status = 'processing'`, and spawns parallel `asyncio` tasks bounded by semaphore.
|
||||
4. Each task reads the frozen prompt snapshot from `Job` and calls Vision API for a **single** `Source` image.
|
||||
5. On task completion:
|
||||
* Writes a `JobSource` record containing `status='transcribed'`, `raw_transcription`, operational `ai_metadata`, and complete unedited `raw_api_response`.
|
||||
* Caches active output text to `Source.raw_transcription`.
|
||||
|
||||
|
||||
6. On page failure:
|
||||
* Writes `JobSource` record with `status='failed'` and `error_detail`.
|
||||
|
||||
|
||||
7. Once all page tasks resolve:
|
||||
* Marks `Job.status` as `transcribed` (100% success), `partial_success` (at least 1 success, 1 failure), or `failed` (all failed).
|
||||
|
||||
|
||||
|
||||
## Domain Ownership & Invariants
|
||||
|
||||
* **Immutable AI Outputs:** `source.raw_transcription` and `job_source.raw_transcription` store original, point-in-time machine output and are immutable.
|
||||
* **Complete Input & Output Provenance:** Every `job` stores the exact frozen input configuration sent to the model, and every `job_source` stores per-page output evidence including the complete REST response envelope returned.
|
||||
* **Inlined Revisions:** Human corrections occur on `source.revised_text`. UI renders `COALESCE(revised_text, raw_transcription)`.
|
||||
* **Sequential Integrity:** Multi-page documents are strictly ordered by `source.page_number ASC`.
|
||||
* **Page Execution Isolation:** A failure on one page image does not invalidate successful transcriptions on sister pages in the same batch job.
|
||||
|
||||
## Data Model Summary
|
||||
|
||||
* `Document` has many `Source` pages, many `Job` runs, and many `Person` records via `DocumentPerson` junction (`author` or `recipient`).
|
||||
* `Source` belongs to one `Document` and can be processed across many `JobSource` executions.
|
||||
* `Job` has many `JobSource` execution records.
|
||||
* `JobSource` holds page-level execution status, output text, and raw response JSON.
|
||||
|
||||
## Test Strategy
|
||||
|
||||
* Unit tests for SQLModel/Pydantic V2 models, JSON cross-dialect serialization, and file hashing functions.
|
||||
* Integration tests for async database connection handling, session management, and queries.
|
||||
* Async workflow tests using mock AI providers to verify `partial_success`, page-level failure isolation, and retry logic.
|
||||
* UI integration tests for multi-page rendering and person attribution management.
|
||||
|
||||
---
|
||||
|
||||
## Technology References
|
||||
|
||||
* [FastAPI documentation](https://fastapi.tiangolo.com/)
|
||||
* [NiceGUI documentation](https://nicegui.io/documentation)
|
||||
* [SQLModel documentation](https://sqlmodel.tiangolo.com/)
|
||||
* [SQLAlchemy Async I/O documentation](https://docs.sqlalchemy.org/en/20/orm/extensions/asyncio.html)
|
||||
* [Python asyncio](https://www.google.com/search?q=https://docs.python.org/3/library/asyncio.html%23module-asyncio)
|
||||
* [Pydantic Validation](https://pydantic.dev/docs/validation/latest/get-started/)
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v3.md)
|
||||
- [System Design Intent](invariant/intent.md)
|
||||
- [Transcription Methodology](invariant/transcription_methodology.md)
|
||||
- System Architecture (this document)
|
||||
- [System Requirements](requirements_v3.md)
|
||||
- [Data model](schema_v3.md)
|
||||
- [Error Handling Policy](error_handling_v3.md)
|
||||
- [Implementation Plan](implementation_plan_v3.md)
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
@@ -1,87 +0,0 @@
|
||||
# Error Handling Policy (Version 3)
|
||||
|
||||
This document defines the canonical error-handling policy for the v3 document transcription system.
|
||||
|
||||
## Error Handling Objectives
|
||||
|
||||
* Make failures visible in clear, actionable language at both the document and individual page levels.
|
||||
* Support **isolated failure handling** in multi-image batches so single page errors do not crash an entire batch job.
|
||||
* Preserve diagnostic detail (Pydantic validation errors, raw provider REST envelopes, exact input prompts) in generic database JSON structures for fast troubleshooting.
|
||||
* Ensure consistent error envelope structure across API, UI, and async worker boundaries.
|
||||
|
||||
## Scope And Authority
|
||||
|
||||
Governs error behavior across NiceGUI pages, FastAPI routes, service orchestration, `asyncio` background tasks, database interactions, and AI provider adapters.
|
||||
|
||||
## Error Taxonomy
|
||||
|
||||
| Category | Definition | Retriable |
|
||||
| --- | --- | --- |
|
||||
| `validation_error` | Pydantic payload or parameter schema validation failure | no |
|
||||
| `user_input_error` | Unacceptable user file (unsupported image type, corrupt file) | no |
|
||||
| `not_found_error` | Requested resource (`Document`, `Source`, `Person`, `Job`) missing | no |
|
||||
| `conflict_error` | Operation violates state constraints (e.g., duplicate `document_person` role) | no |
|
||||
| `external_provider_error` | AI Provider API failure (rate limit, vision execution error) | yes |
|
||||
| `infrastructure_transient_error` | Temporary DB connection reset or HTTP timeout | yes |
|
||||
| `infrastructure_persistent_error` | Database down, missing API credentials, misconfiguration | no |
|
||||
| `internal_unexpected_error` | Uncaught Python exception or logic defect | no |
|
||||
|
||||
## Async Batch & Page-Level Error Behavior
|
||||
|
||||
In multi-image `asyncio` batch processing:
|
||||
|
||||
1. **Page Isolation:** Exceptions caught during individual page calls are trapped within the `asyncio` task wrapper.
|
||||
2. **Page Record Logging:** Page failure details, along with the prompt inputs and hyperparameters attempted, are written directly to `job_source.error_detail` and `job_source.status = 'failed'`.
|
||||
3. **Batch Aggregate State:**
|
||||
* If **all** page tasks succeed -> `job.status = 'completed'`.
|
||||
* If **some** page tasks fail -> `job.status = 'partial_success'`.
|
||||
* If **all** page tasks fail -> `job.status = 'failed'`.
|
||||
|
||||
|
||||
4. **Retry Strategy:** The UI exposes a "Retry Failed Pages" option for `partial_success` jobs, which spawns a new targeted `Job` containing *only* the `Source` IDs marked as `failed`.
|
||||
|
||||
## API Error Response Contract
|
||||
|
||||
API error responses return a structured JSON envelope:
|
||||
^^^json
|
||||
{
|
||||
"error_id": "err_uuid_12345",
|
||||
"category": "validation_error",
|
||||
"message": "The uploaded payload failed schema validation.",
|
||||
"suggestion": "Check file format and metadata fields, then try again.",
|
||||
"details": {
|
||||
"pydantic_errors": [...]
|
||||
},
|
||||
"timestamp": "2026-08-08T15:00:00Z"
|
||||
}
|
||||
^^^
|
||||
|
||||
HTTP Status Mappings:
|
||||
|
||||
* `validation_error`, `user_input_error` -> `400`
|
||||
* `not_found_error` -> `404`
|
||||
* `conflict_error` -> `409`
|
||||
* `external_provider_error` -> `502` / `503`
|
||||
* `infrastructure_transient_error` -> `503`
|
||||
* `infrastructure_persistent_error`, `internal_unexpected_error` -> `500`
|
||||
|
||||
---
|
||||
|
||||
## Technology References
|
||||
|
||||
* [FastAPI documentation](https://fastapi.tiangolo.com/)
|
||||
* [NiceGUI documentation](https://nicegui.io/documentation)
|
||||
* [SQLModel documentation](https://sqlmodel.tiangolo.com/)
|
||||
* [Python asyncio](https://www.google.com/search?q=https://docs.python.org/3/library/asyncio.html%23module-asyncio)
|
||||
* [Pydantic Validation](https://pydantic.dev/docs/validation/latest/get-started/)
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v3.md)
|
||||
- [System Design Intent](invariant/intent.md)
|
||||
- [Transcription Methodology](invariant/transcription_methodology.md)
|
||||
- [System Architecture](architecture_v3.md)
|
||||
- [System Requirements](requirements_v3.md)
|
||||
- [Data model](schema_v3.md)
|
||||
- Error Handling Policy (this document)
|
||||
- [Implementation Plan](implementation_plan_v3.md)
|
||||
@@ -1,81 +0,0 @@
|
||||
# Implementation Plan (Version 3)
|
||||
|
||||
## Goal
|
||||
|
||||
Replace the current v2 SQLModel schema with the approved v3 schema and make sure every database operation works through the existing async SQLAlchemy/SQLModel session layer.
|
||||
|
||||
Use a fresh database. There will be no migrations, data conversion, legacy compatibility shims, or parallel v2/v3 code paths.
|
||||
|
||||
## Current Project Impact
|
||||
|
||||
* `src/transcription/db/models.py` defines the SQLModel tables. It must be updated to match the approved v3 schema (`Document`, `Person`, `DocumentPerson`, `Source`, `Job`, `JobSource`).
|
||||
* The v3 target adds frozen submission-time prompt snapshot fields (`prompt_name`, `prompt_hash`, `system_prompt`, `user_prompt`, `temperature`, `top_p`) to `Job` and full output payloads (`raw_api_response`, `ai_metadata`) to `JobSource`.
|
||||
* The v3 target adds image asset verification fields (`file_hash`, `file_size_bytes`) to `Source`.
|
||||
* Database operations must utilize `JSONBCompat` and the existing SQLModel/SQLAlchemy abstractions to preserve the same logical schema and JSON behavior across the supported backends, while keeping PostgreSQL as the intended production database.
|
||||
* Async CRUD lives in `DocumentService`, `JobService`, `TranscriptionService`, and upload helpers. Their queries and relationship loading must be updated for v3 fields.
|
||||
* Canonical operator tooling must remain OS-independent; safety workflows such as destructive-test backup and restore should run through Python or other cross-platform entry points rather than platform-specific shells.
|
||||
|
||||
## Implementation
|
||||
|
||||
### 1. Update the Schema and Domain Models
|
||||
|
||||
* Replace the models in `src/transcription/db/models.py` with the approved v3 tables, enums, relationships, foreign keys, constraints, and indexes.
|
||||
* Ensure all JSON fields use `JSONBCompat` for dialect portability across SQLite and PostgreSQL.
|
||||
* Keep `SQLModel.metadata.create_all()` as the schema bootstrap for fresh databases.
|
||||
* Delete `_ensure_sqlite_compat_columns()` and all legacy schema patching from `src/transcription/db/operations.py`.
|
||||
* Keep the Python models and `docs/schema_v3.md` perfectly synchronized.
|
||||
|
||||
### 2. Update Data Services and Async Worker Layer
|
||||
|
||||
* Update job creation and worker orchestration so prompt configuration is resolved at submission and frozen onto `Job` (`prompt_name`, `prompt_hash`, `system_prompt`, `user_prompt`, `temperature`, `top_p`) before execution starts.
|
||||
* Update `TranscriptionService` and provider adapters to store the complete unedited API REST response dictionary into `job_source.raw_api_response` alongside operational metrics in `job_source.ai_metadata`.
|
||||
* Update upload handlers to calculate and store file metadata (`file_hash` via SHA-256, `file_size_bytes`) on `Source` records during file ingestion.
|
||||
* Remove legacy single-source compatibility flows so worker paths persist per-page outcomes only through `JobSource` updates.
|
||||
|
||||
### 3. Update Integration Tests and Mock AI Providers
|
||||
|
||||
* Update mock provider fixtures in test suites to return realistic complete API response envelopes.
|
||||
* Verify test coverage for `JSONBCompat` field writes and reads under SQLite in-memory test databases.
|
||||
* Add assertions in async workflow tests to verify frozen prompt snapshot fields on `Job`, plus per-page failure isolation and output evidence on `JobSource`.
|
||||
|
||||
### 4. Update the UI for the v3 Schema
|
||||
|
||||
* Review the UI components and views displaying document, job, person, and source data so they reference v3 schema properties instead of v2 relationships.
|
||||
* Ensure the UI correctly renders `COALESCE(revised_text, raw_transcription)` for page viewing and inline editing.
|
||||
* Ensure resubmit actions only queue failed pages and preserve frozen prompt snapshot behavior on the existing `Job`.
|
||||
* Consider the guidance in `docs/ui_style_guide.md` when making UI changes so updated views remain consistent with the project’s visual conventions.
|
||||
|
||||
### 5. Keep Operational Tooling Portable
|
||||
|
||||
* Implement destructive-test backup and restore workflows in Python so the canonical path runs on Windows, Linux, and macOS.
|
||||
* Avoid making core developer or recovery procedures depend on PowerShell-only or shell-specific semantics.
|
||||
* Keep operational documentation aligned with the cross-platform command path used by the repository.
|
||||
|
||||
## Done When
|
||||
|
||||
* A fresh database is created directly from the v3 SQLModel metadata.
|
||||
* Frozen prompt input provenance is captured on `Job` for each submission, and full per-page output evidence is captured on `JobSource` for every AI execution task.
|
||||
* The focused tests and full test suite pass on both SQLite and PostgreSQL backends.
|
||||
* Canonical operator workflows required for development and destructive-test recovery run without a Windows-only shell dependency.
|
||||
|
||||
## Out of Scope
|
||||
|
||||
* Database migrations or preservation of v2 data
|
||||
* Legacy compatibility code
|
||||
* UI redesign, batch orchestration, worker concurrency, deployment, and operational runbooks
|
||||
|
||||
---
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v3.md)
|
||||
- [System Design Intent](invariant/intent.md)
|
||||
- [Transcription Methodology](invariant/transcription_methodology.md)
|
||||
- [System Architecture](architecture_v3.md)
|
||||
- [System Requirements](requirements_v3.md)
|
||||
- [Data model](schema_v3.md)
|
||||
- [Error Handling Policy](error_handling_v3.md)
|
||||
- Implementation Plan (this document)
|
||||
|
||||
|
||||
|
||||
@@ -1,49 +0,0 @@
|
||||
# Document Transcription System Overview (Version 3)
|
||||
|
||||
This project is a personal-scale application for transcribing, indexing, and preserving historical family documents, letters, postcards, and journals.
|
||||
|
||||
## Start Here
|
||||
|
||||
Read [architecture_v3.md](https://www.google.com/search?q=architecture_v3.md) first for technical overview and system design.
|
||||
|
||||
## Core V3 Capabilities
|
||||
|
||||
* **Folder & Multi-Image Ingestion:** Upload one or more images that map sequentially (`page_number`) under a single `Document`.
|
||||
* **Parallel Async AI Vision Engine:** Concurrently process single-page image transcriptions using Python `asyncio` bounded by rate limiters.
|
||||
* **Portable Relational Storage:** SQLModel and SQLAlchemy preserve a portable relational model across the supported backends, with SQLite for local development/testing and PostgreSQL as the production database target.
|
||||
* **Cross-Platform Operations:** Canonical developer and recovery workflows run through Python-based, OS-independent tooling rather than platform-specific shell scripts.
|
||||
* **Complete Auditability & Provenance:** Capture frozen submission-time input prompts (`system_prompt`, `user_prompt`) and hyperparameters (`temperature`, `top_p`) on `Job`, plus per-page operational metrics (`ai_metadata`) and full provider response envelopes (`raw_api_response`) on `JobSource`.
|
||||
* **Asset Integrity Tracking:** Calculate and store cryptographic hashes (SHA-256) and file sizes on `Source` image records while preserving clean filesystem storage.
|
||||
* **Pydantic V2 Validation:** End-to-end type safety, DB row mapping, and JSON payload validation.
|
||||
* **Historical Person Management:** Track authors and recipients across documents with rich biographical entities (`Person`).
|
||||
* **Page-Level Execution Auditing & Revisions:** Store immutable point-in-time machine output per run while enabling inline human corrections (`revised_text`).
|
||||
* **Partial Failure Recovery:** Bounded batch execution that isolates single-page API errors (`partial_success`) for simple retries.
|
||||
|
||||
## Technical Stack
|
||||
|
||||
* **Application Web Framework:** FastAPI + NiceGUI
|
||||
* **Persistence Engine:** SQLModel / SQLAlchemy (SQLite for development/testing, PostgreSQL for production)
|
||||
* **Data Validation & Schemas:** Pydantic V2
|
||||
* **Concurrency & Workers:** Python `asyncio` worker pool with `asyncio.Semaphore`
|
||||
* **Vision Providers:** OpenAI, Anthropic, and OpenRouter Vision models via native SDK adapters
|
||||
|
||||
---
|
||||
|
||||
## Technology References
|
||||
|
||||
* [FastAPI documentation](https://fastapi.tiangolo.com/)
|
||||
* [NiceGUI documentation](https://nicegui.io/documentation)
|
||||
* [SQLModel documentation](https://sqlmodel.tiangolo.com/)
|
||||
* [Python asyncio](https://www.google.com/search?q=https://docs.python.org/3/library/asyncio.html%23module-asyncio)
|
||||
* [Pydantic Validation](https://pydantic.dev/docs/validation/latest/get-started/)
|
||||
|
||||
## Documentation Index
|
||||
|
||||
- System Overview (this document)
|
||||
- [System Design Intent](invariant/intent.md)
|
||||
- [Transcription Methodology](invariant/transcription_methodology.md)
|
||||
- [System Architecture](architecture_v3.md)
|
||||
- [System Requirements](requirements_v3.md)
|
||||
- [Data model](schema_v3.md)
|
||||
- [Error Handling Policy](error_handling_v3.md)
|
||||
- [Implementation Plan](implementation_plan_v3.md)
|
||||
@@ -1,45 +0,0 @@
|
||||
# Document Transcription System Requirements (Version 3)
|
||||
|
||||
This document captures the **Version 3 baseline requirements** for the production implementation.
|
||||
|
||||
## Requirements Model
|
||||
|
||||
| ID | Category | Requirement | Verify Method |
|
||||
| --- | --- | --- | --- |
|
||||
| REQ-0 | System | Provide end-to-end multi-page document transcription with persistent, inspectable async job states. | demonstration |
|
||||
| REQ-1 | Functional | Allow users to upload multi-image batches as sequential `Source` pages under a `Document`. | test |
|
||||
| REQ-2 | Functional | Process multi-page jobs asynchronously using an `asyncio` worker pool bounded by rate limits. | test |
|
||||
| REQ-3 | Functional | Persist frozen submission-time execution parameters and full input prompts (`system_prompt`, `user_prompt`, `prompt_name`, `prompt_hash`, `temperature`, `top_p`) on `Job`, and persist page-level output responses (`raw_transcription`, `ai_metadata`, `raw_api_response`) on `JobSource`. | test |
|
||||
| REQ-4 | Functional | Support job states (`queued`, `processing`, `completed`, `partial_success`, `failed`) and page states (`pending`, `transcribed`, `failed`). | inspection |
|
||||
| REQ-5 | Functional | Allow users to manage historical `Person` records and link multiple authors/recipients to a `Document` via `DocumentPerson`. | test |
|
||||
| REQ-6 | Functional | Maintain immutable original machine output on `Source.raw_transcription` while permitting inline human edits on `Source.revised_text`. | test |
|
||||
| REQ-7 | Data Constraint | Use SQLModel/SQLAlchemy to preserve a portable relational domain model and compatible data shape across the supported backends, with SQLite for local development/testing and PostgreSQL as the production system of record. | inspection |
|
||||
| REQ-8 | Data Constraint | Validate all API requests, database rows, and JSON structures using Pydantic V2 schemas and SQLModel. | test |
|
||||
| REQ-9 | Interface | Render multi-page transcriptions sequentially by `page_number` in the web UI with author/recipient metadata. | demonstration |
|
||||
| REQ-10 | Operations | Allow operators to resubmit only failed pages for queued reprocessing while preserving the frozen prompt snapshot on the existing `Job`. | test |
|
||||
| REQ-11 | Data Constraint | Calculate and store cryptographic file hashes (SHA-256) and file sizes for uploaded source images to track asset integrity. | test |
|
||||
| REQ-12 | Operations Constraint | Keep core development, testing, restore, and recovery workflows OS-independent across Windows, Linux, and macOS; do not require a platform-specific shell for canonical project processes. | inspection |
|
||||
|
||||
## Element Satisfaction Mapping
|
||||
|
||||
* **UI (NiceGUI):** Satisfies REQ-1, REQ-5, REQ-6, REQ-9, REQ-10.
|
||||
* **API (FastAPI):** Satisfies REQ-1, REQ-4, REQ-5, REQ-8.
|
||||
* **WORKER (asyncio):** Satisfies REQ-2, REQ-3, REQ-4, REQ-10.
|
||||
* **PERSISTENCE (SQLModel/SQLAlchemy):** Satisfies REQ-3, REQ-6, REQ-7, REQ-11.
|
||||
* **MODELS (Pydantic V2 / SQLModel):** Satisfies REQ-8.
|
||||
* **OPERATIONS TOOLING (Python / OS-neutral automation):** Satisfies REQ-12.
|
||||
|
||||
---
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v3.md)
|
||||
- [System Design Intent](invariant/intent.md)
|
||||
- [Transcription Methodology](invariant/transcription_methodology.md)
|
||||
- [System Architecture](architecture_v3.md)
|
||||
- System Requirements (this document)
|
||||
- [Data model](schema_v3.md)
|
||||
- [Error Handling Policy](error_handling_v3.md)
|
||||
- [Implementation Plan](implementation_plan_v3.md)
|
||||
|
||||
|
||||
@@ -1,137 +0,0 @@
|
||||
# Database Schema (Version 3)
|
||||
|
||||
This document describes the relational schema for the transcription platform. It incorporates multi-image batch orchestration, page-level execution tracking, many-to-many author/recipient attribution, submission-time prompt snapshot capture, and raw API payload evidence for archival auditing.
|
||||
|
||||
The schema uses generic JSON columns compatible with SQLite in local development and PostgreSQL native JSONB/UUID types in production.
|
||||
|
||||
## Entity Relationship Diagram
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
PERSON {
|
||||
UUID id PK
|
||||
TEXT full_name
|
||||
TEXT display_name
|
||||
TEXT maiden_name
|
||||
DATE birth_date
|
||||
TEXT birth_date_raw
|
||||
TEXT birth_place
|
||||
DATE death_date
|
||||
TEXT death_date_raw
|
||||
TEXT death_place
|
||||
TEXT biography
|
||||
TEXT portrait_path
|
||||
JSONB metadata
|
||||
TIMESTAMPTZ created_at
|
||||
TIMESTAMPTZ updated_at
|
||||
}
|
||||
|
||||
DOCUMENT {
|
||||
UUID id PK
|
||||
TEXT name
|
||||
TEXT document_type
|
||||
DATE document_date
|
||||
TEXT document_date_raw
|
||||
TEXT location_created
|
||||
TEXT notes
|
||||
TEXT archive_identifier
|
||||
TIMESTAMPTZ created_at
|
||||
TIMESTAMPTZ updated_at
|
||||
}
|
||||
|
||||
DOCUMENT_PERSON {
|
||||
UUID id PK
|
||||
UUID document_id FK
|
||||
UUID person_id FK
|
||||
VARCHAR role "author | recipient"
|
||||
TIMESTAMPTZ created_at
|
||||
}
|
||||
|
||||
JOB {
|
||||
UUID id PK
|
||||
UUID document_id FK
|
||||
VARCHAR status "queued | processing | transcribed | completed | partial_success | failed"
|
||||
INTEGER retry_count
|
||||
TEXT provider
|
||||
TEXT model
|
||||
TEXT prompt_name
|
||||
TEXT prompt_hash
|
||||
TEXT system_prompt
|
||||
TEXT user_prompt
|
||||
FLOAT temperature
|
||||
FLOAT top_p
|
||||
TIMESTAMPTZ date_created
|
||||
TIMESTAMPTZ date_updated
|
||||
}
|
||||
|
||||
SOURCE {
|
||||
UUID id PK
|
||||
UUID document_id FK
|
||||
INTEGER page_number
|
||||
TEXT upload_name
|
||||
TEXT filename
|
||||
TEXT file_path
|
||||
TEXT file_hash
|
||||
BIGINT file_size_bytes
|
||||
TEXT raw_transcription
|
||||
TEXT revised_text
|
||||
TIMESTAMPTZ date_uploaded
|
||||
TIMESTAMPTZ date_revised
|
||||
}
|
||||
|
||||
JOB_SOURCE {
|
||||
UUID id PK
|
||||
UUID job_id FK
|
||||
UUID source_id FK
|
||||
VARCHAR status "pending | transcribed | failed"
|
||||
TEXT raw_transcription
|
||||
JSONB ai_metadata
|
||||
JSONB raw_api_response
|
||||
TEXT error_detail
|
||||
TIMESTAMPTZ executed_at
|
||||
}
|
||||
|
||||
DOCUMENT ||--o{ DOCUMENT_PERSON : "has_people"
|
||||
PERSON ||--o{ DOCUMENT_PERSON : "participates_in"
|
||||
DOCUMENT ||--o{ JOB : "has_jobs"
|
||||
DOCUMENT ||--o{ SOURCE : "contains_pages"
|
||||
JOB ||--o{ JOB_SOURCE : "executes"
|
||||
SOURCE ||--o{ JOB_SOURCE : "processed_in"
|
||||
```
|
||||
|
||||
## Domain Invariants & Provenance Rules
|
||||
|
||||
### Page-Level Execution & AI Outputs
|
||||
|
||||
* **Execution Granularity:** Every single image execution attempt by an AI model produces a dedicated record in `job_source`.
|
||||
* **Submission Snapshot Provenance:** Every `job` captures the frozen prompt identifier details (`prompt_name`, `prompt_hash`), full prompt text strings (`system_prompt`, `user_prompt`), and hyperparameters (`temperature`, `top_p`) at submission time.
|
||||
* **Point-in-Time Output Auditability:** `job_source.raw_api_response` stores the complete, unedited provider REST response envelope for that specific image page call. `job_source.ai_metadata` stores spatial bounding boxes, normalized token usage, latency, and cost details for fast querying.
|
||||
* **Active Output Caching:** Upon successful completion of an image call, `source.raw_transcription` is updated with the latest output string from `job_source.raw_transcription` for fast UI rendering.
|
||||
|
||||
### Image Storage & Integrity
|
||||
|
||||
* **Filesystem Storage:** Binary images are stored on disk in the local file system. The `source` table holds the relative `file_path`.
|
||||
* **File Integrity Tracking:** `source` captures `file_hash` (SHA-256) and `file_size_bytes` at upload time to guarantee document file integrity and duplicate checking over long-term preservation.
|
||||
|
||||
### Page Ordering & Revisions
|
||||
|
||||
* **Sequential Integrity:** `source.page_number` dictates page ordering within a document. Reads assembling full documents must query `ORDER BY source.document_id, source.page_number ASC`.
|
||||
* **Inlined Human Corrections:** User edits occur at the page level inside `source.revised_text`. `source.raw_transcription` remains immutable. If `source.revised_text` is non-null, application frontends must render `source.revised_text`.
|
||||
|
||||
### Async Job Lifecycle & Failure Isolation
|
||||
|
||||
* **Batch Orchestrator:** A job represents an overarching execution run across one or more source images belonging to a document.
|
||||
* **Isolated Failures:** API requests run concurrently (e.g., using `asyncio`). A failure on page 3 does not invalidate successful transcriptions on page 1 or 2.
|
||||
* **Job States:**
|
||||
* `queued`: Created, awaiting worker execution.
|
||||
* `processing`: Concurrent HTTP tasks actively running.
|
||||
* `completed`: 100% of linked `job_source` tasks succeeded (`transcribed`).
|
||||
* `partial_success`: At least one `job_source` succeeded and at least one failed.
|
||||
* `failed`: All linked `job_source` tasks failed or a job-level runtime error occurred.
|
||||
|
||||
|
||||
|
||||
### Attribution & Person Roles
|
||||
|
||||
* **Multi-Person Roles:** Documents support zero, one, or many authors and recipients linked via `document_person`.
|
||||
* **Role Uniqueness:** `(document_id, person_id, role)` must be unique to prevent duplicate role tagging.
|
||||
@@ -1,144 +0,0 @@
|
||||
# System Architecture (Version 4)
|
||||
|
||||
This document describes the production architecture of the document transcription system.
|
||||
|
||||
## Architecture Objectives
|
||||
|
||||
- Preserve original source material and immutable machine transcription output.
|
||||
- Support batching one or more images into ordered multi-page documents.
|
||||
- Capture complete submission-time prompt provenance and per-page provider response evidence.
|
||||
- Execute page transcription concurrently with bounded `asyncio` workers.
|
||||
- Maintain relational portability across SQLite and PostgreSQL.
|
||||
- Keep operator workflows cross-platform and Python-driven.
|
||||
- Support many-to-many document-person relationships with extensible roles.
|
||||
- Support registry-driven document type classification.
|
||||
|
||||
## Runtime Topology
|
||||
|
||||
The runtime operates as an asynchronous Python application:
|
||||
|
||||
- FastAPI + NiceGUI web application process.
|
||||
- In-process `asyncio` worker engine for transcription execution.
|
||||
- Relational persistence via SQLModel / SQLAlchemy.
|
||||
- Pydantic V2 validation across API payloads, prompt configuration, and structured metadata.
|
||||
|
||||
^^^mermaid
|
||||
flowchart LR
|
||||
U[Browser User] --> A[FastAPI + NiceGUI App]
|
||||
A --> W[Asyncio Worker Engine]
|
||||
A --> DB[(Relational DB)]
|
||||
W --> P[Vision Provider APIs]
|
||||
W --> DB
|
||||
^^^
|
||||
|
||||
## Lifecycle Ownership
|
||||
|
||||
Application lifespan owns runtime setup and teardown:
|
||||
|
||||
- Initialize logging, settings, directories, and prompt configuration.
|
||||
- Manage asynchronous database engine connection pools.
|
||||
- Execute database bootstrap or migrations.
|
||||
- Recover stale or interrupted jobs on startup.
|
||||
- Manage graceful shutdown of active background tasks.
|
||||
|
||||
## Layered Module Structure
|
||||
|
||||
### Interface Layer
|
||||
|
||||
- `src/transcription/ui/**`
|
||||
- `src/transcription/api/**`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- Render document, source, person, job, and classification views.
|
||||
- Accept user input for uploads, editing, linking, and revisions.
|
||||
- Present structured validation and conflict feedback.
|
||||
|
||||
### Application and Async Worker Layer
|
||||
|
||||
- `src/transcription/services/workflows.py`
|
||||
- `src/transcription/worker.py`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- Orchestrate uploads, job creation, and status transitions.
|
||||
- Execute per-page provider calls through bounded concurrency.
|
||||
- Persist page-level outcomes and update aggregate job state.
|
||||
|
||||
### Domain and Service Layer
|
||||
|
||||
- `src/transcription/db/models.py`
|
||||
- `src/transcription/services/*.py`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- Manage transactional operations for documents, people, types, links, sources, jobs, and job sources.
|
||||
- Apply deterministic conflict handling for relationship-role writes.
|
||||
- Use set-based synchronization for many-to-many relationship updates.
|
||||
- Resolve and validate registry-backed document types.
|
||||
|
||||
### Infrastructure Layer
|
||||
|
||||
- `src/transcription/db/**`
|
||||
- `src/transcription/providers/**`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- Provide async database sessions and engine configuration.
|
||||
- Provide provider adapters for vision model execution.
|
||||
|
||||
## Core Workflows
|
||||
|
||||
### 1. Multi-Page Transcription
|
||||
|
||||
1. User uploads one or more images for a `Document`.
|
||||
2. System stores files, hashes them, creates ordered `Source` rows, and creates a `Job`.
|
||||
3. Worker claims the job, marks it `processing`, and executes page calls concurrently.
|
||||
4. Each page writes a `JobSource` result with raw output, metadata, and full provider response evidence.
|
||||
5. Aggregate status becomes `completed`, `partial_success`, or `failed`.
|
||||
|
||||
### 2. Document-Person Relationship Management
|
||||
|
||||
1. User opens a document or person edit flow.
|
||||
2. UI loads existing links grouped by role.
|
||||
3. User adds or removes people within one or more roles.
|
||||
4. Service computes add/remove deltas rather than replacing all links blindly.
|
||||
5. Conflict checks enforce uniqueness and deterministic write semantics before persistence commits.
|
||||
|
||||
### 3. Document Type Management
|
||||
|
||||
1. User selects a registry-backed document type for a document.
|
||||
2. Service resolves the stable type code or id.
|
||||
3. Persistence stores the `document_type_id` reference.
|
||||
4. Inactive types remain valid for historical rows but are excluded from default selectors.
|
||||
|
||||
## Domain Invariants
|
||||
|
||||
- `Source.raw_transcription` stores immutable machine output.
|
||||
- Human corrections occur only in `Source.revised_text`.
|
||||
- Prompt and parameter provenance is frozen on `Job` at submission time.
|
||||
- Provider output evidence is stored on `JobSource` for each page execution.
|
||||
- `DocumentPerson` links are unique for `(document_id, person_id, role_id)`.
|
||||
- Relationship mutations are deterministic and set-based.
|
||||
- `DocumentType.code` is stable; `DocumentType.label` may evolve.
|
||||
|
||||
## Data Model Summary
|
||||
|
||||
- `Document` has one `DocumentType`, many `Source` pages, many `Job` runs, and many `Person` records through `DocumentPerson`.
|
||||
- `Source` belongs to one `Document` and may participate in many `JobSource` executions.
|
||||
- `Job` has many `JobSource` rows.
|
||||
- `PersonRole` defines available relationship roles.
|
||||
|
||||
## Test Strategy
|
||||
|
||||
- Unit tests for models, validation, hashing, and registry resolution.
|
||||
- Service tests for CRUD, set-based sync, uniqueness conflicts, and deterministic relationship writes.
|
||||
- Async workflow tests for page isolation, partial failure handling, and stored evidence.
|
||||
- UI integration tests for multi-page rendering, role grouping, and document type selection.
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v4.md)
|
||||
- [System Requirements](requirements_v4.md)
|
||||
- [Data Model](schema_v4.md)
|
||||
- [Error Handling Policy](error_handling_v4.md)
|
||||
@@ -1,110 +0,0 @@
|
||||
# Error Handling Policy (Version 4)
|
||||
|
||||
This document defines the canonical error-handling policy for the document transcription system.
|
||||
|
||||
## Error Handling Objectives
|
||||
|
||||
- Make failures visible in clear, actionable language at both the document and page levels.
|
||||
- Support isolated failure handling in multi-page jobs so one failing page does not invalidate successful pages.
|
||||
- Preserve diagnostic detail for validation failures, provider failures, and policy conflicts.
|
||||
- Ensure consistent error envelope structure across API, UI, service, and worker boundaries.
|
||||
|
||||
## Scope and Authority
|
||||
|
||||
This policy governs error behavior across:
|
||||
|
||||
- NiceGUI pages
|
||||
- FastAPI routes
|
||||
- Service-layer orchestration
|
||||
- `asyncio` worker tasks
|
||||
- Database interactions
|
||||
- Provider adapters
|
||||
|
||||
## Error Taxonomy
|
||||
|
||||
| Category | Definition | Retriable |
|
||||
| --- | --- | --- |
|
||||
| `validation_error` | Payload, parameter, or schema validation failure | no |
|
||||
| `user_input_error` | Unacceptable file, invalid selection, or malformed request from the operator | no |
|
||||
| `not_found_error` | Requested `Document`, `Source`, `Person`, `Job`, role, or type does not exist | no |
|
||||
| `conflict_error` | Operation violates uniqueness or relationship-write policy | no |
|
||||
| `external_provider_error` | Provider API failure, rate limit, or execution problem | yes |
|
||||
| `infrastructure_transient_error` | Temporary DB, file-system, or network instability | yes |
|
||||
| `infrastructure_persistent_error` | Persistent configuration, credential, or database availability failure | no |
|
||||
| `internal_unexpected_error` | Uncaught exception or logic defect | no |
|
||||
|
||||
## Async Batch and Page-Level Error Behavior
|
||||
|
||||
In multi-page `asyncio` processing:
|
||||
|
||||
1. Exceptions from individual page calls are trapped within the page task wrapper.
|
||||
2. Failed page detail is written to `JobSource.error_detail` and the page state becomes `failed`.
|
||||
3. Aggregate job status is derived from page outcomes:
|
||||
- all pages succeed -> `completed`
|
||||
- some succeed and some fail -> `partial_success`
|
||||
- all fail -> `failed`
|
||||
4. Successful pages remain valid even when sister pages fail.
|
||||
|
||||
## Relationship and Classification Conflict Behavior
|
||||
|
||||
When relationship or document-type writes fail policy checks:
|
||||
|
||||
1. Reject the full write operation.
|
||||
2. Return structured conflict detail including target identifiers and the violated rule.
|
||||
3. Preserve existing persisted relationships unchanged.
|
||||
|
||||
## API Error Response Contract
|
||||
|
||||
API error responses return a structured envelope:
|
||||
|
||||
^^^json
|
||||
{
|
||||
"error_id": "err_uuid_12345",
|
||||
"category": "conflict_error",
|
||||
"message": "Relationship write conflicts with existing links.",
|
||||
"suggestion": "Adjust the requested relationship links and retry.",
|
||||
"details": {
|
||||
"document_id": "...",
|
||||
"person_id": "...",
|
||||
"attempted_role": "recipient",
|
||||
"operation": "add_link",
|
||||
"conflict_reason": "duplicate document-person-role link"
|
||||
},
|
||||
"timestamp": "2026-08-10T15:00:00Z"
|
||||
}
|
||||
^^^
|
||||
|
||||
HTTP status mappings:
|
||||
|
||||
- `validation_error`, `user_input_error` -> `400`
|
||||
- `not_found_error` -> `404`
|
||||
- `conflict_error` -> `409`
|
||||
- `external_provider_error` -> `502` or `503`
|
||||
- `infrastructure_transient_error` -> `503`
|
||||
- `infrastructure_persistent_error`, `internal_unexpected_error` -> `500`
|
||||
|
||||
## UI Error Presentation Rules
|
||||
|
||||
- Display concise failure summaries with the next action the operator can take.
|
||||
- Keep form state in context when feasible.
|
||||
- Distinguish validation issues, conflict issues, provider failures, and infrastructure failures.
|
||||
- For bulk relationship updates, identify the specific role or person that caused a conflict.
|
||||
|
||||
## Logging and Audit Expectations
|
||||
|
||||
- Log worker failures with correlation IDs and provider context.
|
||||
- Log relationship and classification conflicts with machine-readable detail.
|
||||
- Log persisted provider errors and page-level execution failures.
|
||||
|
||||
## Retry Guidance
|
||||
|
||||
- Do not auto-retry validation or conflict failures.
|
||||
- Permit user-driven retry after the input or selection changes.
|
||||
- Allow bounded retry for transient provider or infrastructure failures when the operation is idempotent.
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v4.md)
|
||||
- [System Requirements](requirements_v4.md)
|
||||
- [Data Model](schema_v4.md)
|
||||
- [System Architecture](architecture_v4.md)
|
||||
@@ -1,99 +0,0 @@
|
||||
# Implementation Plan (Version 4)
|
||||
|
||||
## Goal
|
||||
|
||||
Implement the Version 4 project definition from the current repository state while preserving existing data by default.
|
||||
|
||||
## Migration Policy
|
||||
|
||||
- Database changes are non-destructive by default.
|
||||
- Exception: the legacy `document_type` text field may be replaced by a `document_type_id` reference without migrating existing text values.
|
||||
- Exception: `document_person` links may be recreated manually.
|
||||
|
||||
## Current Project Impact
|
||||
|
||||
- `src/transcription/db/models.py` requires full schema alignment with the V4 core documents.
|
||||
- `src/transcription/services/documents.py` requires set-based document-person sync and document-type resolution.
|
||||
- API modules require additive role-aware relationship behavior and document-type selection behavior.
|
||||
- UI pages require grouped role displays, multi-role editing, and registry-backed document-type selection.
|
||||
- Existing tests require updates for role enforcement, document-type selection, and regression safety.
|
||||
|
||||
## Implementation Phases
|
||||
|
||||
### 1. Finalize the Transition Documents
|
||||
|
||||
- Confirm the reset scope.
|
||||
- Confirm the database exception policy.
|
||||
- Keep core V4 documents as the only authoritative product definition.
|
||||
|
||||
### 2. Align the Persistence Layer
|
||||
|
||||
- Update SQLModel definitions to match the final V4 schema.
|
||||
- Add `person_role` and `document_type` support.
|
||||
- Replace legacy document-type storage with `document_type_id`.
|
||||
- Apply the accepted manual exception strategy for `document_type` and `document_person` data.
|
||||
- Preserve all other data structures non-destructively.
|
||||
|
||||
### 3. Update Services and Write Semantics
|
||||
|
||||
- Implement set-based synchronization for document-person updates.
|
||||
- Implement deterministic uniqueness and relationship-write conflict checks.
|
||||
- Remove suggestion-related service behavior.
|
||||
- Add document-type resolution and validation by stable code or id.
|
||||
|
||||
### 4. Update API Contracts
|
||||
|
||||
- Keep API evolution additive.
|
||||
- Add role-aware relationship retrieval and write behavior.
|
||||
- Add document-type catalog retrieval and code-based selection for document writes.
|
||||
- Remove suggestion-related API surfaces from the V4 target state.
|
||||
|
||||
### 5. Update UI Workflows
|
||||
|
||||
- Replace single-person link editing with grouped multi-role editing.
|
||||
- Render grouped role links on document and person detail views.
|
||||
- Replace free-text document type entry with registry-backed selection.
|
||||
- Preserve clear validation and conflict messaging.
|
||||
|
||||
### 6. Verification and Hardening
|
||||
|
||||
- Add or update service tests for many-per-role behavior, uniqueness conflict handling, and set-based sync correctness.
|
||||
- Add API tests for relationship behavior and document-type selection.
|
||||
- Add UI tests or walkthrough coverage for grouped roles and type selection.
|
||||
- Add regression coverage for delete and cleanup semantics.
|
||||
- Enforce backup-first test execution for AI-run unit tests: backup `./data` before tests, then always prompt for restore after successful tests.
|
||||
- Keep restore confirmation-gated by default so code and test outcomes can be reviewed before data is reverted.
|
||||
|
||||
## Done When
|
||||
|
||||
- Core V4 documents and code paths agree on the final project definition.
|
||||
- Relationship-role writes are deterministic and non-destructive.
|
||||
- Relationship-write conflict rules are enforced consistently.
|
||||
- Document type selection is registry-backed.
|
||||
- The accepted manual exceptions for `document_type` and `document_person` are completed.
|
||||
- The focused test coverage passes.
|
||||
|
||||
## Out of Scope
|
||||
|
||||
- Suggested/asserted relationship state.
|
||||
- Suggestion review or extraction workflows.
|
||||
- Global person entity-resolution engine.
|
||||
- Automated semantic document-type classification.
|
||||
|
||||
## Delivery Order Recommendation
|
||||
|
||||
1. Freeze scope boundary and implementation plan.
|
||||
2. Freeze core V4 documents.
|
||||
3. Align persistence models.
|
||||
4. Align services and API behavior.
|
||||
5. Align UI behavior.
|
||||
6. Run focused verification and regression checks.
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [V4 Scope Boundary](scope_boundary_v4.md)
|
||||
- [System Overview](index_v4.md)
|
||||
- [System Requirements](requirements_v4.md)
|
||||
- [Data Model](schema_v4.md)
|
||||
- [System Architecture](architecture_v4.md)
|
||||
- [Error Handling Policy](error_handling_v4.md)
|
||||
@@ -1,40 +0,0 @@
|
||||
# Document Transcription System Overview (Version 4)
|
||||
|
||||
This project is a personal-scale application for transcribing, organizing, and preserving historical documents, images, and related people records.
|
||||
|
||||
## Start Here
|
||||
|
||||
Read [architecture_v4.md](architecture_v4.md) first for the technical overview and system design.
|
||||
|
||||
## Core Capabilities
|
||||
|
||||
- Folder and multi-image ingestion into sequential `Source` pages under a single `Document`.
|
||||
- Parallel asynchronous AI vision transcription using Python `asyncio` bounded by rate limits.
|
||||
- Portable relational storage using SQLModel and SQLAlchemy across SQLite and PostgreSQL.
|
||||
- Complete prompt and response provenance for every transcription job and page execution.
|
||||
- File-integrity tracking through SHA-256 hashing and stored file sizes.
|
||||
- Historical `Person` management with many-to-many document links and extensible relationship roles.
|
||||
- Registry-driven `DocumentType` classification with stable codes and controlled selection.
|
||||
- Inline human revision of transcribed pages while preserving immutable machine output.
|
||||
- Partial-failure recovery for multi-page jobs.
|
||||
- Cross-platform operational workflows driven by Python-based tooling.
|
||||
|
||||
## Technical Stack
|
||||
|
||||
- Application Web Framework: FastAPI + NiceGUI
|
||||
- Persistence Engine: SQLModel / SQLAlchemy
|
||||
- Data Validation and Schemas: Pydantic V2
|
||||
- Concurrency and Workers: Python `asyncio`
|
||||
- Vision Providers: OpenAI, Anthropic, and OpenRouter adapters
|
||||
|
||||
## Core Documentation Index
|
||||
|
||||
- [System Architecture](architecture_v4.md)
|
||||
- [System Requirements](requirements_v4.md)
|
||||
- [Data Model](schema_v4.md)
|
||||
- [Error Handling Policy](error_handling_v4.md)
|
||||
|
||||
## Transition Documents
|
||||
|
||||
- [Scope Boundary](scope_boundary_v4.md)
|
||||
- [Implementation Plan](implementation_plan_v4.md)
|
||||
@@ -1,51 +0,0 @@
|
||||
# Document Transcription System Requirements (Version 4)
|
||||
|
||||
This document defines the baseline requirements for the document transcription system.
|
||||
|
||||
## Requirements Model
|
||||
|
||||
| ID | Category | Requirement | Verify Method |
|
||||
| --- | --- | --- | --- |
|
||||
| REQ-0 | System | Provide end-to-end multi-page document transcription with persistent, inspectable async job states. | demonstration |
|
||||
| REQ-1 | Functional | Allow users to upload one or more images as ordered `Source` pages under a `Document`. | test |
|
||||
| REQ-2 | Functional | Process page transcription asynchronously using an `asyncio` worker pool bounded by rate limits. | test |
|
||||
| REQ-3 | Functional | Persist submission-time prompt configuration and full page-level provider response evidence for every job execution. | test |
|
||||
| REQ-4 | Functional | Support job states `queued`, `processing`, `completed`, `partial_success`, and `failed`, plus page states `pending`, `transcribed`, and `failed`. | inspection |
|
||||
| REQ-5 | Functional | Allow users to manage historical `Person` records and link multiple people per role to a `Document`. | test |
|
||||
| REQ-6 | Functional | Support an extensible role taxonomy for document-person relationships. | inspection |
|
||||
| REQ-7 | Policy Constraint | Enforce deterministic relationship-role writes with uniqueness on `(document_id, person_id, role_id)` and explicit conflict responses for invalid duplicate link attempts. | test |
|
||||
| REQ-8 | Functional | Use set-based synchronization for document-person mutations so updates add and remove only the intended links. | test |
|
||||
| REQ-9 | Functional | Maintain immutable machine output on `Source.raw_transcription` while permitting inline human edits on `Source.revised_text`. | test |
|
||||
| REQ-10 | Functional | Support a registry-driven `DocumentType` taxonomy with stable codes, mutable labels, and active/inactive lifecycle control. | test |
|
||||
| REQ-11 | Data Constraint | Store `Document` type as a controlled reference to `DocumentType`. | test |
|
||||
| REQ-12 | Interface | Render multi-page transcriptions sequentially by `page_number` with document, people, and document-type metadata. | demonstration |
|
||||
| REQ-13 | Interface | Document create/edit UI must support selecting multiple people per role and selecting an active document type from the registry. | demonstration |
|
||||
| REQ-14 | API Constraint | Expose additive, role-aware retrieval and write behavior for document-person links and code-based selection for document types. | test |
|
||||
| REQ-15 | Data Constraint | Calculate and store cryptographic file hashes (SHA-256) and file sizes for uploaded source images. | test |
|
||||
| REQ-16 | Data Constraint | Preserve a portable relational model across supported backends using SQLModel, SQLAlchemy, SQLite, and PostgreSQL. | inspection |
|
||||
| REQ-17 | Reliability | Ensure delete and update flows for documents, people, and relationship links remain deterministic and safe. | test |
|
||||
| REQ-18 | Operations Constraint | Keep canonical development, testing, restore, and recovery workflows OS-independent; for AI-run unit tests, require a pre-test backup of `./data` and an always-shown post-success confirmation prompt before any restore action. | inspection |
|
||||
| REQ-19 | Quality | Provide automated coverage for async transcription workflows, relationship-role enforcement, document-type selection, and regression behavior. | test |
|
||||
|
||||
## Clarifying Constraints
|
||||
|
||||
1. `DocumentType.code` and `PersonRole.code` are stable machine identifiers.
|
||||
2. `DocumentType.label` and `PersonRole.label` may evolve without changing canonical identity.
|
||||
3. Relationship-write policy and conflict handling must be consistent across UI, API, services, and persistence.
|
||||
4. Many-per-role behavior is required for document-person links.
|
||||
5. Relationship conflicts must fail deterministically without partial mutation.
|
||||
|
||||
## Element Satisfaction Mapping
|
||||
|
||||
- UI (NiceGUI): Satisfies REQ-0, REQ-1, REQ-5, REQ-9, REQ-12, REQ-13.
|
||||
- API (FastAPI): Satisfies REQ-1, REQ-4, REQ-5, REQ-7, REQ-8, REQ-14.
|
||||
- Worker (`asyncio`): Satisfies REQ-2, REQ-3, REQ-4.
|
||||
- Persistence (SQLModel / SQLAlchemy): Satisfies REQ-3, REQ-9, REQ-10, REQ-11, REQ-15, REQ-16, REQ-17.
|
||||
- Test Suite: Verifies all test-marked requirements and satisfies REQ-19.
|
||||
|
||||
## Related Local References
|
||||
|
||||
- [System Overview](index_v4.md)
|
||||
- [System Architecture](architecture_v4.md)
|
||||
- [Data Model](schema_v4.md)
|
||||
- [Error Handling Policy](error_handling_v4.md)
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user