100 Commits
Author SHA1 Message Date
Jim Lancaster be3d0b7ee0 Some tweaks to accommodate relocating development to loca
Quality Gate / gate (push) Successful in 1m13s
2026-09-10 11:33:52 -05:00
Jim Lancaster 7eca9fe7dc V6.2 Add GEDCOM data
Quality Gate / gate (push) Successful in 1m27s
2026-09-03 05:34:13 -05:00
Jim Lancaster e54c2d9f26 Update the roadmap_plan.md
Quality Gate / gate (push) Successful in 2m15s
2026-09-02 19:03:35 -05:00
Jim Lancaster 065acad125 Scrub references to older versions
Quality Gate / gate (push) Successful in 2m39s
2026-09-02 17:02:48 -05:00
Jim Lancaster eeb1888aa3 Update instructions - epilog (Code review, remediation now complete)
Quality Gate / gate (push) Successful in 2m35s
2026-09-02 16:48:00 -05:00
Jim Lancaster 321c454a4f Implementat Rec - PHase 5 complete
Quality Gate / gate (push) Successful in 2m37s
2026-09-02 16:36:06 -05:00
Jim Lancaster 6f4accf275 Implement rec - Phase 4 complete
Quality Gate / gate (push) Successful in 2m38s
2026-09-02 16:23:09 -05:00
Jim Lancaster 27ec81ca5b Implement Rec - Phase 3 complete
Quality Gate / gate (push) Successful in 2m37s
2026-09-02 16:15:33 -05:00
Jim Lancaster 4410d23f5c Implement Rec - Phase 2
Quality Gate / gate (push) Successful in 2m36s
2026-09-02 16:05:18 -05:00
Jim Lancaster d5e798825d Implement Rec - Phase 1
Quality Gate / gate (push) Successful in 2m31s
2026-09-02 15:59:56 -05:00
Jim Lancaster 02dca888d2 Code-review complete
Quality Gate / gate (push) Failing after 14s
2026-09-02 15:47:53 -05:00
Jim Lancaster a896a11d2e Update instructions - final
Quality Gate / gate (push) Successful in 2m32s
2026-09-02 14:54:37 -05:00
Jim Lancaster 0b48c80d87 Update instructions - Part 2, phase 3
Quality Gate / gate (push) Successful in 2m34s
2026-09-02 14:42:20 -05:00
Jim Lancaster 70f8d6182e Update instructions - part 2, phase 2
Quality Gate / gate (push) Successful in 2m33s
2026-09-02 14:36:45 -05:00
Jim Lancaster fc4288ff88 Update instructions - part 2
Quality Gate / gate (push) Successful in 2m34s
2026-09-02 14:31:40 -05:00
Jim Lancaster e5ef4d4422 Update instructions, agents, skills - part 1
Quality Gate / gate (push) Successful in 2m41s
2026-09-02 14:05:41 -05:00
Jim Lancaster 88cef169c4 V6.1 update roadmap_plan.md
Quality Gate / gate (push) Successful in 2m35s
2026-09-02 13:42:15 -05:00
Jim Lancaster 15a4814e23 V6.1 runtime settings ->.env.production fix
Quality Gate / gate (push) Successful in 2m29s
2026-09-02 12:31:04 -05:00
Jim Lancaster 16391463d6 V6.1 yet more backup/.env cleanup
Quality Gate / gate (push) Successful in 2m33s
2026-09-02 11:43:07 -05:00
Jim Lancaster 96bb80d91f V6.1 backups: once more into the breach.
Quality Gate / gate (push) Successful in 2m29s
2026-09-02 10:59:28 -05:00
Jim Lancaster 494f378e48 V6.1 continue beating on the backup
Quality Gate / gate (push) Failing after 2m34s
2026-09-02 10:04:38 -05:00
Jim Lancaster 75dc946123 V6.1 continue refining backup: switch to Synology Drive Client for remote backup
Quality Gate / gate (push) Failing after 1m10s
2026-09-01 18:47:03 -05:00
Jim Lancaster 929f8d6de9 V6.1 Fix backup script
Quality Gate / gate (push) Failing after 2m30s
2026-09-01 16:08:25 -05:00
Jim Lancaster 72909c1fd5 V6.1 Another dockerfile fix
Quality Gate / gate (push) Failing after 1m1s
2026-09-01 15:59:40 -05:00
Jim Lancaster a5ec9f40a8 V6.1 dockerfile fix.
Quality Gate / gate (push) Failing after 1m0s
2026-09-01 15:36:34 -05:00
Jim Lancaster 0d5206fe97 V6.1 update database for maintenance runs
Quality Gate / gate (push) Failing after 2m29s
2026-09-01 13:51:38 -05:00
Jim Lancaster 32b6b5f29a V6.1 refinements - add link to source image gallery on Document detail page. Fix Maintenance tab actions (again).
Quality Gate / gate (push) Failing after 2m30s
2026-09-01 13:32:46 -05:00
Jim Lancaster 9b6bb9ae66 V6.1 more fixes
Quality Gate / gate (push) Failing after 2m35s
2026-09-01 12:36:59 -05:00
Jim Lancaster 4c877dd6a2 V6.1 fixes
Quality Gate / gate (push) Failing after 2m59s
2026-09-01 11:54:34 -05:00
Jim Lancaster 9990583345 V6.1 UI refinements, add Maintenance jobs to Settings
Quality Gate / gate (push) Failing after 2m57s
2026-08-31 11:31:25 -05:00
Jim Lancaster daa1642933 Cloudflare-related tweak
Quality Gate / gate (push) Failing after 50s
2026-08-27 12:23:08 -05:00
Jim Lancaster d198cc0c68 Separate CLoudflare from transcribe
Quality Gate / gate (push) Failing after 49s
2026-08-27 10:19:34 -05:00
Jim LancasterandCopilot App 34c7b16675 Refine Synology backup workflow
Quality Gate / gate (push) Failing after 47s
Co-authored-by: Copilot App <[email protected]>
2026-08-26 17:32:20 -05:00
Jim Lancaster ca29bc8b74 Continued work on backup script, methodology
Quality Gate / gate (push) Failing after 48s
2026-08-26 14:26:15 -05:00
Jim Lancaster 89ed83239e Update backup script to include everything needed to do a bare metal restore of the new V6 server
Quality Gate / gate (push) Failing after 49s
2026-08-26 12:21:40 -05:00
Jim Lancaster 2c6ef46f5f V6 complete
Quality Gate / gate (push) Failing after 50s
2026-08-26 12:04:43 -05:00
Jim Lancaster c2ed98c16d V6 Post migration cleanup
Quality Gate / gate (push) Failing after 50s
2026-08-26 11:08:10 -05:00
Jim Lancaster c85cc6be20 V6 Phase 4 - Postgres backup to Synology
Quality Gate / gate (push) Failing after 49s
2026-08-25 13:09:14 -05:00
Jim Lancaster 234476ba6d V6.0 Phase 3 - Add cloudflare tunnel. Remote access to app via the tunnel now working.
Quality Gate / gate (push) Failing after 49s
2026-08-25 13:05:47 -05:00
Jim Lancaster faa30fd27b V6 Phase 2 complete
Quality Gate / gate (push) Failing after 50s
2026-08-25 11:44:59 -05:00
Jim Lancaster 867cc9eb78 V6 Phase 1 complete
Quality Gate / gate (push) Failing after 49s
2026-08-25 10:38:45 -05:00
Jim Lancaster 0e43094b80 Refine Settings page
Quality Gate / gate (push) Failing after 49s
2026-08-24 17:16:40 -05:00
Jim Lancaster c05c054b94 Refine Settings page
Quality Gate / gate (push) Failing after 47s
2026-08-24 17:08:09 -05:00
Jim Lancaster c0112c2714 Add ability to edit safe .env.example options to Settings
Quality Gate / gate (push) Failing after 47s
2026-08-24 16:14:49 -05:00
Jim Lancaster 96af7b2dc0 Minor UI refinement
Quality Gate / gate (push) Failing after 49s
2026-08-24 12:58:34 -05:00
Jim Lancaster 9a7970c533 UI refinement: Back buttons
Quality Gate / gate (push) Failing after 49s
2026-08-24 12:44:02 -05:00
Jim LancasterandCopilot App f15c9834e4 Remove completed remediation handoff and mark the review closed
Quality Gate / gate (push) Failing after 47s
The handoff brief was a work order for phases 2-5. That work is done, so the
document now describes a future that already happened and would misdirect
anyone who found it.

The review report itself had the same problem in weaker form: its findings read
as open. Adds a status banner marking it closed and retained for reasoning only.

The banner also records that two of its recommendations were wrong on contact.
The HIGH-03 fix as written would have stripped root-cause data from
ExecutionAttempt provenance, and the HIGH-01 fix had to preserve per-page
durability the report never mentioned. Leaving that unstated invites someone to
'restore' the report's version later.

Co-authored-by: Copilot App <[email protected]>
2026-08-23 19:25:39 -05:00
Jim LancasterandCopilot App 2c59cbd2c7 Add AGENTS.md to make repo context auto-discoverable
Nothing in this repo was auto-read by an agent at the start of a session.
The .github/instructions files only attach once a matching file is edited,
which is too late to steer strategy, and the canonical docs set is not
discoverable without already knowing to look for it.

AGENTS.md routes rather than duplicates: it states the authority order
(docs/* first, docs/reviews/** explicitly non-canonical), the uv-only
command set, the test-enforced boundaries, and the change protocol.

The Traps section records failure modes this codebase has actually produced
rather than generic advice: the message/detail split that leaked paths in one
direction and degraded provenance in the other, the two competing atomicity
invariants in workflows.py where the obvious simplification breaks multi-page
durability, and the habit of changing a shared symbol without enumerating its
consumers.

Verified: cited test paths exist, full suite passes, ruff/ty clean.

Co-authored-by: Copilot App <[email protected]>
2026-08-23 19:23:26 -05:00
Jim LancasterandCopilot App 5ff66a8c40 Align python-reviewer agent with the python-code-reviewer skill
The agent file had drifted from the skill it delegates to, in two ways that
would corrupt a review run.

Report target: the agent said write reports to ./docs, but the skill targets
./docs/reviews/<date>-code-review.md and explicitly marks docs/reviews/** as
non-canonical. Following the agent would place a dated, opinionated review
inside the canonical authority set that findings are supposed to resolve
against.

Verification commands: the agent said run 'ruff check', 'pytest', and 'ty'.
None are on PATH in this uv project, so an agent following its own instruction
gets command-not-found and is pushed toward guessing instead of verifying.

The agent now defers to the skill for all specifics rather than restating them,
which is what let the two copies drift apart. Also carries forward the
consumer-tracing rule and states the read-only scope explicitly.

Co-authored-by: Copilot App <[email protected]>
2026-08-23 19:17:33 -05:00
Jim LancasterandCopilot App 626b5d4b10 Harden python-code-reviewer skill with lessons from executing its own review
Three gaps surfaced by implementing the 2026-08-23 review's recommendations.

1. Recommendations were never verified the way claims were. The report's fix for
   the error path leak would have stripped root-cause data from evidence records,
   because the review traced one consumer of AppError.message and missed that
   format_error_detail writes it to ExecutionAttempt.error_detail. Adds workflow
   step 9 (validate recommendations against consumers), a Blast Radius field on
   findings, and the worked example so the failure mode is concrete.

2. Fixes that sit between competing invariants were not flagged. The atomicity
   recommendation did not note that per-page durability and terminal-status
   atomicity pull in opposite directions, so the obvious simplification silently
   breaks multi-page durability. Recommendations must now name both invariants,
   the test guarding each, and the over-correction to avoid.

3. Severity could not express reachability. Two findings were latent behind a
   default setting and a single-instance deployment, which is a sequencing
   constraint: they must be fixed before the change that makes them live. Adds an
   explicit Reachability field with Live / Latent / Theoretical.

Verified: meta contract guards and traceability tests pass.

Co-authored-by: Copilot App <[email protected]>
2026-08-23 19:14:35 -05:00
Jim LancasterandCopilot App 12f125761a test: invert UI boundary guard to allowlist
Quality Gate / gate (push) Failing after 47s
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:57:13 -05:00
Jim LancasterandCopilot App 803237371e fix: log omitted OpenRouter request manifests
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:55:33 -05:00
Jim LancasterandCopilot App d4ae97c1b1 chore: remove uncertain orphaned definitions
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:53:54 -05:00
Jim LancasterandCopilot App c52d41ec33 test: extend orphan sweep to public class methods
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:52:06 -05:00
Jim LancasterandCopilot App 6c3eac0a44 test: prune low-signal assertions and tighten guards
Quality Gate / gate (push) Failing after 47s
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:48:16 -05:00
Jim LancasterandCopilot App e5410708e4 refactor: extract blocking and sequence retry helpers
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:46:19 -05:00
Jim LancasterandCopilot App efbae26f16 refactor: route workflow sessions through unit of work
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:43:07 -05:00
Jim LancasterandCopilot App 3873810022 perf: move blocking homepage and hash I/O off loop
Quality Gate / gate (push) Failing after 47s
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:41:01 -05:00
Jim LancasterandCopilot App 2093eb6fb3 fix: retry execution attempt number conflicts
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:35:59 -05:00
Jim LancasterandCopilot App f9261a1af3 fix: derive worker shutdown wait from timeout budget
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:32:40 -05:00
Jim LancasterandCopilot App 86cdb4035c fix: gate retries by error category with backoff
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:30:54 -05:00
Jim LancasterandCopilot App f193b2800b feat: run stale-job recovery periodically
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:28:21 -05:00
Jim LancasterandCopilot App 736d0c06f4 chore: enable logging format lint rule
Quality Gate / gate (push) Failing after 48s
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:19:05 -05:00
Jim LancasterandCopilot App 26f9c83f54 chore: make ty check blocking with targeted suppressions
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:17:44 -05:00
Jim LancasterandCopilot App a2bb1acd6b chore: enforce ruff format in pre-commit
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:14:52 -05:00
Jim LancasterandCopilot App 2a56365847 style: apply ruff formatting sweep
Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:13:38 -05:00
Jim LancasterandCopilot App 4aaa9bd581 Add remediation handoff brief for review phases 2-5
Quality Gate / gate (push) Failing after 47s
Records the implementation plan derived from the 2026-08-23 review: per-task
acceptance criteria, the verification baseline, and environment constraints.

Lives in docs/reviews/ so it is discoverable from the repo rather than from
session state, and is indexed from docs/reviews/README.md. Non-canonical, like
everything under docs/reviews/**.

Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:04:55 -05:00
Jim LancasterandCopilot App de18c2e9da Fix workflow commit atomicity, error path leak, and UI error boundary
Quality Gate / gate (push) Failing after 48s
Phase 1 of docs/reviews/2026-08-23-code-review.md.

HIGH-01: process_queued_job committed page evidence and the terminal job
status in separate transactions, so a crash between them left a transcript
persisted against a job stuck in PROCESSING that the worker never reclaims.
The final page's write is now deferred into _finalize_batch_outcome so it
shares the terminal transaction. Intermediate pages remain individually
durable, and the terminal commit is shielded against cancellation the same
way per-page writes already were.

HIGH-04: added tests/integration/test_pipeline_atomicity.py covering both
Transaction B and Transaction C. Confirmed failing against the previous
implementation before the fix.

HIGH-03: classify_unexpected_error interpolated the raw exception into
AppError.message, which the UI renders and the API serializes, leaking the
database path from OperationalError. message is now generic. Because message
also feeds format_error_detail, which writes evidence records, the root cause
is preserved on a new internal-only AppError.detail field rather than
discarded.

HIGH-02: replaced 8 hand-rolled ui.notify error calls in home_page and
people_page with error_presenter.show_error, restoring the correlation
error_id, canonical category, and suggestion. Added an AST guard to
test_ui_boundaries.py so pages cannot hand-roll error notifications again.

Docs updated per documentation-sync: the message/detail split in
docs/error_handling.md and the multi-page atomicity rule in
services.instructions.md.

Verification: ruff clean, 381 tests passing, ty unchanged at 10 known
SQLAlchemy descriptor false positives.

Co-authored-by: Copilot App <[email protected]>
2026-08-23 18:02:04 -05:00
Jim LancasterandCopilot App 8d3c60fce1 Add 2026-08-23 architecture and code review report
Quality Gate / gate (push) Failing after 46s
Full review per .github/skills/python-code-reviewer/skill.md, with escalations to the evidence-provenance-auditor and test-effectiveness-auditor skills.

Verification: ruff clean, 377 tests passing, 10 advisory ty diagnostics.
Outcome: 0 critical, 4 high, 5 medium, 11 low.

Report only; no source changes.

Co-authored-by: Copilot App <[email protected]>
2026-08-23 17:13:10 -05:00
Jim LancasterandCopilot App 8a30231adf Triage high-value ty diagnostics
Quality Gate / gate (push) Failing after 47s
Apply the highest-value typing fixes from the ty baseline pass:
- align migration row typing with SQLAlchemy RowMapping sequences
- accept refreshable callback return type in homepage gallery
- guard nullable media URL before ui.image in people photos
- guard nullable source MIME type before startswith checks
- fix tests/test_db collect() return annotation to match 4-tuple

This clears all actionable ty findings from that set and leaves only
known SQLModel/SQLAlchemy descriptor false positives.

Co-authored-by: Copilot App <[email protected]>
2026-08-23 16:54:34 -05:00
Jim LancasterandCopilot App 67feeb28af Repair the pre-commit quality gate and clear the ruff backlog
Quality Gate / gate (push) Failing after 47s
The pre-commit hooks declared `language: system` with bare `ruff`/`ty`
entries, but both are uv-managed dev dependencies and are not on PATH, so every
commit failed with `Executable 'ruff' not found`. Route both through
`uv run`; keep ruff blocking and make ty advisory (verbose) until its 18
whole-project diagnostics are cleared.

With the gate working, clear `ruff check .` to zero:

- 18 auto-fixes (import sorting, blank lines, `max()` simplification,
  `with` merging, unused imports).
- Real defects: `SourceNavigation` annotated but never imported in
  sources_page; two naive `datetime.now()` calls in migration.py now use
  `datetime.now(UTC)`.
- Dead parameters removed: `source_has_photo_table` (computed, passed, never
  read), `_serialize_value(key=...)`, and unused `request` on two NiceGUI
  page handlers where the framework injects it optionally.
- Mechanical line-length wrapping and one `startswith` tuple collapse.
- `# noqa: PLR0915` / `# noqa: PLR1702` on five long UI/migration
  functions, following the convention already used in jobs_page and
  settings_page, rather than refactoring during stabilization.

Full suite green (377 tests, `-m "not external"`).

Co-authored-by: Copilot App <[email protected]>
2026-08-23 16:47:11 -05:00
Jim LancasterandCopilot App c6ed3126e0 Enforce the four unenforced reviewer checks with guard tests
The reviewer skill recorded four deterministic checks as unenforced or partial. Add tests so they fail the build instead of relying on a reviewer noticing.

tests/test_model_contract_guards.py:
- Status vocabulary: flags string literals compared against or assigned to status/purpose attributes, plus a narrower sweep that requires every status-valued literal in the package to be a known non-status use.
- Relationship loading: every Relationship must declare lazy='raise' except documented exceptions, and the exception set must match the Relationship Loading Contract in docs/schema.md.
- Schema fidelity: the Field-Accurate Table Contracts tables must match db/models.py on table coverage, field names, and declaration order, and the Authoritative Enumerations section must match the enum members.

tests/test_orphan_sweep.py:
- Locks the set of unreferenced public definitions. Route handlers registered by decorator are exempt, string entrypoint references count, and tests/ and tools/ count as consumers. KNOWN_ORPHANS records the four current orphans with rationale; a new one fails the build.

Each guard was mutation-tested: reverting the fix below, dropping a documented field, widening a lazy strategy, and adding a stranded function each fail their respective test.

Also fix the one violation the status guard found: sources_page.py compared attempt.status.value to the literal 'transcribed' instead of JobSourceStatus.TRANSCRIBED, which would survive an enum rename.

Co-authored-by: Copilot App <[email protected]>
2026-08-23 16:35:48 -05:00
Jim LancasterandCopilot App 5566f48fc0 Align python-code-reviewer skill with repo ground truth
Quality Gate / gate (push) Failing after 11s
Update the reviewer skill so its procedure matches how this repo actually works:

- Route review reports to docs/reviews/ and mark them non-canonical, resolving the conflict where reports landed in the same docs/ tree they resolve findings against.
- Pin verification commands to uv (uv run ruff check / ty check / pytest -m 'not external').
- Record the pytest contract: strict markers, strict asyncio mode, and the never-awaited-coroutine warning promoted to an error.
- Convert the deterministic checks to a table with an Enforced by column; three checks are unenforced and one only partial, which are now findings by construction.
- Add a consequence-based severity rubric and a Direction column for bidirectional drift.
- Escalate test-suite concerns to test-effectiveness-auditor.

Also fix tests/test_db.py, which was missing 'from sqlalchemy import text' while using it in 14 places. Three tests were failing with NameError. Wrapped the pre-existing long lines in the same file so it lints clean.

Document the deliberate nicegui==3.13.0 pin in pyproject.toml, a new runbook dependency upgrade policy, and the reviewer skill, so the pin is not flagged as a defect or widened as incidental cleanup.

Co-authored-by: Copilot App <[email protected]>
2026-08-23 16:27:44 -05:00
Jim Lancaster ed6998d8da V5.1 Update tests and documentation
Quality Gate / gate (push) Failing after 11s
2026-08-23 16:01:53 -05:00
Jim Lancaster ebf659b26c V5.1 UI refinements
Quality Gate / gate (push) Failing after 11s
2026-08-23 15:28:41 -05:00
Jim Lancaster ae3483ec2e V5.1 Modify Person table: split full name into first & last, added tags support
Quality Gate / gate (push) Failing after 11s
2026-08-23 12:23:57 -05:00
Jim Lancaster 141ee1fa85 V5.0 Minor change to UI
Quality Gate / gate (push) Failing after 11s
2026-08-23 10:49:40 -05:00
Jim Lancaster 0f30d902b9 V5.0 fixes and revisions
Quality Gate / gate (push) Failing after 11s
2026-08-23 10:32:20 -05:00
Jim Lancaster 86b8e83ff4 v5.0 Introduce centralized homepage & portrait photo management
Quality Gate / gate (push) Failing after 11s
2026-08-23 09:11:36 -05:00
Jim Lancaster efe7785392 Fix document drift caused by adding tags 2026-08-23 06:46:43 -05:00
Jim Lancaster 94db493756 Revise Source detail page to improve line-wrap issues.
Quality Gate / gate (push) Failing after 11s
2026-08-22 18:56:22 -05:00
Jim Lancaster 0d554c0648 V4.11 Added tags + lots of little changes to the UI
Quality Gate / gate (push) Failing after 12s
2026-08-22 18:32:52 -05:00
Jim Lancaster 63c21d4a14 v4.10 revision to remove "legacy compatibility" code
Quality Gate / gate (push) Failing after 11s
2026-08-22 11:21:18 -05:00
Jim Lancaster cf49c3c127 V4.10
Quality Gate / gate (push) Failing after 11s
2026-08-22 10:19:30 -05:00
Jim Lancaster bf2f3ac09c Remove references to "v4" throughout the code and documentation
Quality Gate / gate (push) Failing after 11s
2026-08-20 16:35:20 -05:00
Jim Lancaster eaeb0bc806 claude-sonnet-5 review: Phase 5 (final) implemented by gpt-5.3-codex
Quality Gate / gate (push) Failing after 37s
2026-08-20 16:14:44 -05:00
Jim Lancaster 8b08478c9d claude-sonnet-5 review: Phase 4 (by gpt-5.3-codex)
Quality Gate / gate (push) Failing after 11s
2026-08-20 16:04:37 -05:00
Jim Lancaster 450d33d507 claude-sonnet-5 review Phase 3 (by gpt-5.3-codex)
Quality Gate / gate (push) Failing after 11s
2026-08-20 15:36:17 -05:00
Jim Lancaster 796216087c claude-sonnet-5 review: Phase 2 by gpt-5.3-codex
Quality Gate / gate (push) Failing after 11s
2026-08-20 15:22:23 -05:00
Jim Lancaster afd1dba4d4 Phase 1 - minor fix to Sources page
Quality Gate / gate (push) Failing after 11s
2026-08-20 15:14:34 -05:00
Jim Lancaster 7daa0b9808 claude-sonnet-5 review: Phase 1 implemented by gpt-5.3-codex
Quality Gate / gate (push) Failing after 12s
2026-08-20 15:05:17 -05:00
Jim Lancaster 7c4300f9c2 gpt-5.3 codex review: Phase 7 and the addition of the new test-effectiveness-auditor skill.
Quality Gate / gate (push) Failing after 12s
2026-08-20 11:50:10 -05:00
Jim Lancaster 443a1e29c8 gpt-5.3-codesx review: Phase 5 Release Readiness & Contract Enforcement
Quality Gate / gate (push) Failing after 10s
2026-08-20 08:42:51 -05:00
Jim Lancaster cdd846fe29 gpt-5.3 codex review: Phase 4
Quality Gate / gate (push) Failing after 11s
2026-08-19 21:16:36 -05:00
Jim Lancaster 30fcef3892 gpt-5.3-codex review Phase 3
Quality Gate / gate (push) Successful in 34s
2026-08-19 20:50:21 -05:00
Jim Lancaster de8cdb6e1a Phase 1 of Phase 1 results (I'm losing track of the phases) - Update the schema doc
Quality Gate / gate (push) Successful in 35s
2026-08-19 18:29:20 -05:00
Jim Lancaster b6a5a89a84 gpt-5.3-codex review phase 2 - update instructions & skills
Quality Gate / gate (push) Successful in 34s
2026-08-19 18:22:06 -05:00
Jim Lancaster c261fbb3bd gpt-5.3-codex review phase 1 (revised)
Quality Gate / gate (push) Successful in 35s
2026-08-19 15:28:51 -05:00
Jim Lancaster 5404224079 gpt-5.3-codex review phase 1 - Flatten the documentation
Quality Gate / gate (push) Successful in 33s
2026-08-19 14:54:24 -05:00
Jim Lancaster 2c26177d0c Prep for GPT-5.3-codex architecture & code review.
Quality Gate / gate (push) Successful in 35s
2026-08-19 14:25:42 -05:00
208 changed files with 16841 additions and 7374 deletions
-56
View File
@@ -1,56 +0,0 @@
# --- NiceGUI Server ---
# HOST=`0.0.0.0` (default)
# PORT=8000 (default)
# LOG_LEVEL: [`critical`, `error`, `warning`, `info` (default), `debug`, `trace`]
# RELOAD=false (default)
# --- AI provider ---
# PROVIDER=[`openrouter`(default), `google_genai`]
PROVIDER=openrouter
# OPENROUTER_API_KEY - Required when `PROVIDER=openrouter`
OPENROUTER_API_KEY=your-api-key-goes-here
# GEMINI_API_KEY - Required when `PROVIDER=google_genai`
# PROVIDER_MODEL= specify model. If left blank OpenRouter will supply default.
PROVIDER_MODEL=google/gemini-2.5-flash
# Optional JSON allowlist for model selection. The default above is always first.
# PROVIDER_MODELS=["google/gemini-2.5-flash","google/gemini-2.5-pro","anthropic/claude-sonnet-4"]
# OPENROUTER_HTTP_REFERER=https://example.com
# OPENROUTER_APP_TITLE="Google: Gemini 2.5 Flash (openrouter)"
# --- runtime environment ---
# ENVIRONMENT: [`development`(default), `test`, `production`]
# --- persistence ---
# Use nested settings with double underscore because env_nested_delimiter="__".
# SQLite example:
# DATABASE__DRIVER=sqlite
# DATABASE__PATH=app.db
#
# SQLite with custom relative path:
# DATABASE__DRIVER=sqlite
DATABASE__PATH=./data/transcription.db
#
# Postgres example:
# DATABASE__DRIVER=postgres
# DATABASE__HOST=localhost
# DATABASE__PORT=5432
# DATABASE__DATABASE=transcription
# DATABASE__USER=postgres
# DATABASE__PASSWORD=change-me
#
# Optional persistence flags:
# BOOTSTRAP_SCHEMA_ON_STARTUP=false
# SQLITE_CHECK_SAME_THREAD=false
# --- filesystem paths ---
UPLOAD_DIR="./data"
PROMPT_DIR="./prompts"
# --- worker reliability ---
WORKER_MAX_RETRIES=0
WORKER_RETRY_BACKOFF_SECONDS=0
# WORKER_PROVIDER_TIMEOUT_SECONDS=180
WORKER_PROVIDER_TIMEOUT_SECONDS=180
WORKER_MIN_TRANSCRIPTION_CHARS=0
WORKER_MIN_TRANSCRIPTION_LINES=0
WORKER_FAIL_ON_FINISH_REASON_LENGTH=false
+67
View File
@@ -0,0 +1,67 @@
# Production environment example for docker-compose.production.yml
# --- NiceGUI Server ---
HOST=0.0.0.0
PORT=8000
LOG_LEVEL=info
RELOAD=false
ENVIRONMENT=production
# TRANSCRIPTION_COMMIT=
RUN_EMBEDDED_WORKER=false
LOG_DIR=/app/data/logs
LOG_FILE_NAME=transcription.log
LOG_FILE_MAX_BYTES=10485760
LOG_FILE_BACKUP_COUNT=5
# --- AI provider ---
PROVIDER=openrouter
OPENROUTER_API_KEY=replace-with-real-key
PROVIDER_MODEL=google/gemini-2.5-flash
# PROVIDER_MODELS=["google/gemini-2.5-flash","anthropic/claude-sonnet-4"]
# OPENROUTER_HTTP_REFERER=
# OPENROUTER_APP_TITLE=
DEFAULT_PROMPT_NAME=transcribe_document.md
# TRANSCRIPTION_TEMPERATURE=
# TRANSCRIPTION_TOP_P=
# --- persistence ---
# Common database settings:
DATABASE__DRIVER=postgres
DATABASE__DATABASE=transcription
DATABASE__USER=transcription
DATABASE__PASSWORD=replace-with-strong-password
BOOTSTRAP_SCHEMA_ON_STARTUP=false
# SQLite-specific settings:
# DATABASE__PATH=./data/transcription.db
# SQLITE_CHECK_SAME_THREAD=false
# Postgres-specific settings:
DATABASE__HOST=postgres
DATABASE__PORT=5432
# --- filesystem paths ---
UPLOAD_DIR=/app/uploads
PROMPT_DIR=/app/prompts
# --- backup workflow helpers (not Runtime Settings model fields) ---
BACKUP_DIR=/backup
BACKUP_RETENTION_DAYS=14
# --- worker reliability ---
WORKER_MAX_RETRIES=0
WORKER_PROVIDER_TIMEOUT_SECONDS=30.0
WORKER_STALE_JOB_SECONDS=90.0
WORKER_RETRY_BACKOFF_SECONDS=1.0
WORKER_SHUTDOWN_GRACE_SECONDS=5.0
WORKER_POLL_INTERVAL_SECONDS=1.0
WORKER_MIN_TRANSCRIPTION_CHARS=0
WORKER_MIN_TRANSCRIPTION_LINES=0
WORKER_FAIL_ON_FINISH_REASON_LENGTH=false
# --- cloudflare tunnel ---
# Required for token-based tunnel startup.
CLOUDFLARE_TUNNEL_TOKEN=replace-with-cloudflare-tunnel-token
# --- deployment wiring helpers ---
# Runtime Settings writes target this file path inside the app container.
RUNTIME_SETTINGS_ENV_FILE=/app/.env.production
+67
View File
@@ -0,0 +1,67 @@
# Production environment
# --- NiceGUI Server ---
HOST=0.0.0.0
PORT=8000
LOG_LEVEL=info
RELOAD=false
ENVIRONMENT=production
# TRANSCRIPTION_COMMIT=
RUN_EMBEDDED_WORKER=false
LOG_DIR=/app/data/logs
LOG_FILE_NAME=transcription.log
LOG_FILE_MAX_BYTES=10485760
LOG_FILE_BACKUP_COUNT=5
# --- AI provider ---
PROVIDER=openrouter
OPENROUTER_API_KEY=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
PROVIDER_MODEL=google/gemini-2.5-flash
PROVIDER_MODELS=["google/gemini-2.5-pro","google/gemini-2.5-flash","anthropic/claude-opus-5","anthropic/claude-sonnet-4","openai/gpt-5.6","openai/gpt-4o"]
# OPENROUTER_HTTP_REFERER=
# OPENROUTER_APP_TITLE=
DEFAULT_PROMPT_NAME=transcribe_document.md
# TRANSCRIPTION_TEMPERATURE=
# TRANSCRIPTION_TOP_P=
# --- persistence ---
# Common database settings:
DATABASE__DRIVER=postgres
DATABASE__DATABASE=transcription
DATABASE__USER=transcription
DATABASE__PASSWORD=<password>
BOOTSTRAP_SCHEMA_ON_STARTUP=false
# SQLite-specific settings:
# DATABASE__PATH=./data/transcription.db
# SQLITE_CHECK_SAME_THREAD=false
# Postgres-specific settings:
DATABASE__HOST=postgres
DATABASE__PORT=5432
# --- filesystem paths ---
UPLOAD_DIR=/app/uploads
PROMPT_DIR=/app/prompts
# --- backup workflow helpers (not Runtime Settings model fields) ---
BACKUP_DIR=/backup
BACKUP_RETENTION_DAYS=14
# --- worker reliability ---
WORKER_MAX_RETRIES=0
WORKER_PROVIDER_TIMEOUT_SECONDS=30.0
WORKER_STALE_JOB_SECONDS=30.0
WORKER_RETRY_BACKOFF_SECONDS=1.0
WORKER_SHUTDOWN_GRACE_SECONDS=5.0
WORKER_POLL_INTERVAL_SECONDS=1.0
WORKER_MIN_TRANSCRIPTION_CHARS=0
WORKER_MIN_TRANSCRIPTION_LINES=0
WORKER_FAIL_ON_FINISH_REASON_LENGTH=false
# --- cloudflare tunnel ---
# Required for token-based tunnel startup.
CLOUDFLARE_TUNNEL_TOKEN=CLOUDFLARE_TUNNEL_TOKEN=eyJhIjoiYTRhNjM0NzNhNzBiZjhhYmY3OWUyNjE4ZTcyNjgwZmMiLCJ0IjoiZWY1MjFkNWItYzY1ZS00ZGFmLTlmYTMtMzQyOGYzMGUyMDY4IiwicyI6IlpHWm1NVGRsTVRVdFpEYzNaaTAwWkRJeUxXRmhPRFV0TmpKallXRmhPRFJrWXpSaSJ9
# --- deployment wiring helpers ---
# Runtime Settings writes target this file path inside the app container.
RUNTIME_SETTINGS_ENV_FILE=/app/.env.production
+67
View File
@@ -0,0 +1,67 @@
# Production environment
# --- NiceGUI Server ---
HOST=0.0.0.0
PORT=8000
LOG_LEVEL=info
RELOAD=false
ENVIRONMENT=production
# TRANSCRIPTION_COMMIT=
RUN_EMBEDDED_WORKER=false
LOG_DIR=/app/data/logs
LOG_FILE_NAME=transcription.log
LOG_FILE_MAX_BYTES=10485760
LOG_FILE_BACKUP_COUNT=5
# --- AI provider ---
PROVIDER=openrouter
OPENROUTER_API_KEY=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
PROVIDER_MODEL=google/gemini-2.5-flash
PROVIDER_MODELS=["google/gemini-2.5-pro","google/gemini-2.5-flash","anthropic/claude-opus-5","anthropic/claude-sonnet-4","openai/gpt-5.6","openai/gpt-4o"]
# OPENROUTER_HTTP_REFERER=
# OPENROUTER_APP_TITLE=
DEFAULT_PROMPT_NAME=transcribe_document.md
# TRANSCRIPTION_TEMPERATURE=
# TRANSCRIPTION_TOP_P=
# --- persistence ---
# Common database settings:
DATABASE__DRIVER=sqlite
# DATABASE__DATABASE=transcription
# DATABASE__USER=transcription-local
# DATABASE__PASSWORD=My!3sons
# BOOTSTRAP_SCHEMA_ON_STARTUP=false
# SQLite-specific settings:
DATABASE__PATH=./data-local/transcription-local.db
SQLITE_CHECK_SAME_THREAD=false
# Postgres-specific settings:
# DATABASE__HOST=postgres
# DATABASE__PORT=5432
# --- filesystem paths ---
UPLOAD_DIR=./data-local
PROMPT_DIR=/data/prompts
# --- backup workflow helpers (not Runtime Settings model fields) ---
BACKUP_DIR=/backup
BACKUP_RETENTION_DAYS=14
# --- worker reliability ---
WORKER_MAX_RETRIES=0
WORKER_PROVIDER_TIMEOUT_SECONDS=30.0
WORKER_STALE_JOB_SECONDS=30.0
WORKER_RETRY_BACKOFF_SECONDS=1.0
WORKER_SHUTDOWN_GRACE_SECONDS=5.0
WORKER_POLL_INTERVAL_SECONDS=1.0
WORKER_MIN_TRANSCRIPTION_CHARS=0
WORKER_MIN_TRANSCRIPTION_LINES=0
WORKER_FAIL_ON_FINISH_REASON_LENGTH=false
# --- cloudflare tunnel ---
# Required for token-based tunnel startup.
CLOUDFLARE_TUNNEL_TOKEN=CLOUDFLARE_TUNNEL_TOKEN=eyJhIjoiYTRhNjM0NzNhNzBiZjhhYmY3OWUyNjE4ZTcyNjgwZmMiLCJ0IjoiZWY1MjFkNWItYzY1ZS00ZGFmLTlmYTMtMzQyOGYzMGUyMDY4IiwicyI6IlpHWm1NVGRsTVRVdFpEYzNaaTAwWkRJeUxXRmhPRFV0TmpKallXRmhPRFJrWXpSaSJ9
# --- deployment wiring helpers ---
# Runtime Settings writes target this file path inside the app container.
RUNTIME_SETTINGS_ENV_FILE=/app/.env.production
+8 -9
View File
@@ -1,12 +1,6 @@
--- ---
name: Python Architect Reviewer name: Python Architect Reviewer
description: Evidence-based senior architect reviewer for FastAPI, NiceGUI, and SQLModel codebases. description: Evidence-based senior architect reviewer for FastAPI, NiceGUI, and SQLModel codebases.
tools:
- read_file
- list_dir
- file_search
- grep_search
- run_in_terminal
skills: skills:
- python-code-reviewer - python-code-reviewer
--- ---
@@ -15,10 +9,15 @@ skills:
You are a Senior Python Architect performing an evidence-based, read-only code review. You are a Senior Python Architect performing an evidence-based, read-only code review.
> No `tools:` allowlist is declared here on purpose. Tool identifiers differ between the runtimes
> this agent is invoked from, so a hard-coded list silently under-tools the agent in one of them.
> Read-only discipline is enforced by the **Read-Only Scope** rule below, not by the frontmatter.
## Operating Principles ## Operating Principles
- **Stack Context:** Python 3.12+, FastAPI, NiceGUI, SQLModel, SQLAlchemy (SQLite/PostgreSQL), Pydantic V2, asyncio workers, and OpenRouter adapters. - **Stack Context:** Python 3.12+, FastAPI, NiceGUI, SQLModel, SQLAlchemy (SQLite/PostgreSQL), Pydantic V2, asyncio workers, and OpenRouter adapters.
- **Evidence-Based:** Always inspect real files. Every finding must reference concrete file paths and line numbers (e.g., `app/services/worker.py:45-78`). Do not speculate. - **Evidence-Based:** Always inspect real files. Every finding must reference concrete file paths and line numbers (e.g., `app/services/worker.py:45-78`). Do not speculate.
- **Tool Verification:** Run linters and tests via the terminal (`ruff check`, `pytest`, `ty`) to verify issues before reporting. - **Tool Verification:** This is a `uv` project; the toolchain is not on `PATH`. Verify with `uv run ruff check .`, `uv run ty check`, and `uv run pytest -q -m "not external"`, and record the exact commands and outcomes. Never report a lint, type, or test claim you did not run.
- **Skill Execution:** Adhere strictly to the review dimensions, duplication analysis, and report scaffolding defined in the `python-code-reviewer` skill. - **Verify Recommendations, Not Just Findings:** Before recommending a change to a shared symbol, enumerate its consumers and confirm the fix is safe for each. See the skill's consumer-tracing step and `Blast Radius` field.
- **Report Target:** Output all complete review reports as Markdown files written to `./docs`. - **Skill Is Canonical:** The `python-code-reviewer` skill defines the review workflow, deterministic checks, severity and reachability rubrics, report location, and report template. Follow it exactly. Where this file and the skill disagree, the skill wins — do not restate its specifics here.
- **Read-Only Scope:** Do not modify source, tests, docs, instructions, or configuration. The review report is the only artifact you produce.
@@ -0,0 +1,41 @@
---
description: Require documentation updates whenever code changes alter contracts, behavior, or scope.
applyTo: 'src/transcription/**/*.py'
---
# Documentation Sync Requirements
Keep docs in sync in the same change whenever implementation alters a documented contract, behavior, or roadmap decision.
Documentation targets below always refer to the **current** baseline. `docs/index.md` states which
baseline that is; resolve any version-specific document from there. Never cite a superseded version
tree by name in this file or in the docs you update — retired revision trees are not authority, and
`tests/test_meta_contract_guards.py` fails active contract files that route authority through them.
## Update documentation when any of these change
1. **Schema/Data contract**
- Models, fields, enums, constraints, indexes, relationships, loading semantics.
- **Required doc update:** `docs/schema.md`.
2. **Configuration contract**
- `Settings` keys, defaults, required/optional environment values.
- **Required doc update:** `.env.production.example` and any directly related setup docs.
3. **User-visible UI behavior**
- Page flow, routes, button/action behavior, labels, status wording, empty/error states.
- **Required doc update:** relevant `docs/ui/pages/*.md` docs and feature docs when applicable.
4. **Error handling semantics**
- Error categories, retry behavior, envelope structure, translation boundaries.
- **Required doc update:** `docs/error_handling.md` and `docs/invariant/error_handling.md`.
5. **Roadmap/scope decisions**
- Version targets, sequencing, deferrals, and accepted alternatives.
- **Required doc update:** `docs/roadmap_plan.md`, plus any backlog or feature document for the
current baseline. Locate it through `docs/index.md` rather than assuming a version-named path.
## Working rule
If none of the categories above changed, documentation edits are optional.
If any category changed, update docs in the same PR/change set rather than deferring.
@@ -0,0 +1,121 @@
---
description: Cross-cutting error handling rules for services, API, and UI.
applyTo: 'src/transcription/**/*.py'
---
# Error Handling (Cross-cutting)
Primary references:
- `docs/error_handling.md`
- `docs/invariant/error_handling.md`
- `docs/requirements.md`
## Taxonomy and Categories
Use category-driven semantics aligned to canonical policy:
- `validation`
- `not_found`
- `conflict`
- `external`
- `timeout`
- `internal`
Do not invent ad hoc categories in user/API-facing envelopes unless canonical docs are updated.
Runtime/internal categories may be more specific for diagnostics and persistence, but they must map
deterministically to the canonical envelope categories through the centralized mapper in
`transcription.errors.canonical_error_category`.
Current internal categories:
- `validation_error`
- `user_input_error`
- `not_found_error`
- `conflict_error`
- `external_provider_error`
- `external_timeout_error`
- `processing_error`
- `infrastructure_transient_error`
- `infrastructure_persistent_error`
- `internal_unexpected_error`
Required internal -> canonical mapping:
- `validation_error`, `user_input_error` -> `validation`
- `not_found_error` -> `not_found`
- `conflict_error` -> `conflict`
- `external_provider_error` -> `external`
- `external_timeout_error`, `infrastructure_transient_error` -> `timeout`
- `processing_error`, `infrastructure_persistent_error`, `internal_unexpected_error` -> `internal`
## Translation Boundaries
- **Provider/adapters:** raise provider/domain exceptions; do not emit UI text.
- **Services:** map raw exceptions into internal categories and preserve causal chain (`raise ... from ...`).
- **UI/API:** map internal category -> canonical envelope category and emit user-safe, actionable messages.
## Retry Rules
- No auto-retry for `validation`, `not_found`, `conflict`.
- `external`/`timeout` may be retried when operation semantics are safe.
- Preserve each retry as new evidence where applicable (no history rewrite).
## Job/Page Failure Semantics
- Page-level (`JobSource`): `pending`, `transcribed`, `failed`, `cancelled`.
- Job terminals: `transcribed`, `partial_success`, `failed`.
- Cancellation must keep job-level and page-level semantics explicit and consistent.
- Do not emit legacy terminal state language such as `completed` in active user/API lifecycle contracts.
## User-Safe Messaging
- Never leak stack traces, credentials, auth headers, or local filesystem paths in user-facing output.
- Include actionable remediation guidance aligned to category.
- Keep envelope structure consistent across API endpoints.
### `AppError.message` vs `AppError.detail`
`AppError` carries two texts with different audiences, and they must not be collapsed. Getting this
wrong has already caused a real defect in this repository, in both directions.
| Attribute | Audience | Reaches | Rule |
| :--- | :--- | :--- | :--- |
| `message` | User and API clients | `ErrorEnvelope.message`, UI notifications | Stays generic. Never embed exception text, provider payloads, or filesystem paths. |
| `detail` | Internal only | Logs, and `format_error_detail` -> `ExecutionAttempt.error_detail` and `MaintenanceRun.error_detail` | Carries the root cause. Never rendered to users or serialized into an envelope without a sanitizing projection. |
- Putting root-cause data in `message` leaks infrastructure detail to users.
- Omitting it from `detail` silently degrades the provenance record this system exists to preserve —
a failed attempt whose `error_detail` says nothing is an attempt that cannot be diagnosed later.
- Any render boundary that displays persisted `error_detail` must apply the same no-local-path rule
as `message`: sanitize machine-local absolute paths before the text becomes user-visible.
- When you raise from a caught exception, populate **both**: a generic `message` and a `detail`
carrying `type(exc).__name__` and the exception text, with `raise ... from exc`.
- `detail` is optional (`None`). A read path that assumes it is populated must handle its absence.
- Before changing either attribute, or any helper that formats them, enumerate every consumer —
evidence writes, maintenance runs, logging, API envelopes, and UI presentation all read these
fields, and tests assert on the persisted text.
Canonical definitions live in `src/transcription/errors.py`; see also `docs/error_handling.md`.
## Logging and Diagnostics
- Log operation identifiers and error IDs where available.
- Preserve category + cause-chain context.
- Distinguish no-response timeout/network failures from returned provider error responses.
## Guardrails
- No broad catch-and-swallow patterns.
- No success-shaped fallback values after exceptions.
- Category mapping must remain deterministic and testable.
## Contract Sync Rule
If taxonomy, retries, or envelope semantics change:
1. Update canonical docs (`docs/error_handling.md`, and invariant docs if needed).
2. Update tests in the same change.
3. Update related instruction/skill references.
4. If change affects persisted status/category fields, update `docs/schema.md` when applicable.
@@ -0,0 +1,122 @@
---
description: Provider adapter rules for evidence capture, secret safety, and client lifecycle.
applyTo: 'src/transcription/providers/**/*.py'
---
# Provider Adapters
Primary references:
- `docs/invariant/ai_evidence_and_provenance.md` (canonical; provider adapters own provider-boundary evidence capture)
- `docs/architecture.md`
- `docs/schema.md`
The provider layer is where an external API becomes application data. It is also the only place
that can capture what actually crossed the wire — once a response reaches a service, the evidence
it did not preserve is gone permanently. Treat capture correctness as the primary job of this
layer and text extraction as secondary.
## Layer Boundary
- Adapters may depend on `transcription.config`, `transcription.providers.*`, the HTTP client, and
the provider SDK. They must not import `services`, `db`, `ui`, or `api`.
- Provider specifics — headers, model slugs, payload shapes, SDK types, error classes — stop here.
Callers receive `TranscriptionResult` and `ProviderError` subclasses only.
- Adapters raise provider/domain exceptions. They must not emit user-facing text, notifications,
or remediation wording; that translation belongs to services and UI. See
[error-handling instructions](./error-handling.instructions.md).
- Adapters do not persist. They return evidence; services decide what is written and when.
Enforced by `tests/test_provider_boundaries.py`.
## Contract Surface
- Every adapter satisfies the `TranscriptionProvider` protocol in `base.py`. Failed-call evidence is
returned through the caller-owned `ProviderCallEvidence` sink passed to `transcribe()`, so
evidence stays scoped to one invocation instead of living on mutable adapter instance state.
- `TranscriptionResult`, `RequestManifest`, and `TransportEvidence` are `extra="forbid"` and frozen.
Add a field to the contract rather than smuggling data through an untyped dict.
- Evidence contracts in `evidence.py` are versioned (`schema_name` + `schema_version`). A change to
the meaning or shape of a captured field requires a version bump, not a silent redefinition —
stored evidence must keep its original meaning.
## Transport Evidence
The rules below implement `docs/invariant/ai_evidence_and_provenance.md` §3.4-3.5. That document
wins if this file drifts from it.
- Capture the response body **at the HTTP boundary, before SDK parsing**, so fields the SDK does
not model are not lost. `_CapturingAsyncClient` exists for this; do not replace it with a
post-parse `model_dump()` and call the result transport evidence.
- Reset per-call capture state at the start of every call. Without it, a connection failure can
attach the *previous* call's response as evidence for this one. Guarded by
`tests/test_evidence_provenance.py::test_openrouter_does_not_reuse_prior_response_on_connection_failure`.
- Keep transport capture scoped to the call, not the adapter instance. Concurrent `transcribe()`
calls on one adapter must not be able to overwrite each other's response evidence.
- Handle the streamed-body case (`httpx.ResponseNotRead`) rather than assuming `response.content`
is always available.
- When no response arrives — timeout, DNS, connection reset — emit
`TransportEvidence(response_received=False)`. Absence of a response is itself evidence and must
be explicit, never an empty body or a missing record.
- Preserve safe response evidence for **unsuccessful** calls too, whenever a response was received.
- Never relabel an SDK snapshot or normalized metadata as transport evidence, and never backfill
it into an execution that predates capture.
## Secret Safety
- Persist response headers only through `filter_safe_response_headers` and the
`SAFE_RESPONSE_HEADERS` allowlist. Allowlist, never denylist: capture-then-redact is prohibited,
because an unknown header is unsafe by default.
- Adding a header to the allowlist is a deliberate evidence decision. Confirm it carries no
credential, cookie, or session material, and state why it is needed for correlation, content
interpretation, rate-limit diagnosis, or audit.
- API keys, `Authorization`, and cookies must never appear in a manifest, evidence record, log
line, or exception message.
- The request manifest references source content by identity (digest, size, media type, page).
Do not duplicate base64 source bytes into it — `_replace_embedded_media` exists for this.
## Execution Specification
The manifest must let a reader reconstruct what was asked, per invariant §3.3:
- Provider, requested model, full effective prompt text, and prompt digest.
- Every explicitly supplied parameter, and — separately — which optional parameters were
**omitted**. Omission is not the same as a null value or an assumed provider default; the
`optional_parameter_states` distinction between `omitted`, `null`, and `value` is deliberate.
- Timeout budget, retry policy, source reference, and `SoftwareContext` versions.
- Manifest digests use `canonical_json_bytes`. Do not hash a plain `json.dumps()`; key order and
separators must stay deterministic or digests become uncomparable.
## Client Lifecycle and Async Safety
- Reuse one pooled `AsyncClient` per adapter instance; do not construct a client per request.
- Accept an injected client so tests can drive the adapter without network access.
- Derive timeouts from `Settings` (`worker_provider_timeout_seconds`) rather than hard-coding, and
keep the client timeout aligned with the configured budget so the SDK cannot expire first and
hide the real failure.
- Implement `aclose()` and release pooled resources. An adapter that creates a client owns closing
it; one given a client must not close a caller-owned resource it did not create.
- Never block the event loop. Offload CPU-bound work (hashing large payloads, image encoding) with
`asyncio.to_thread`.
- Propagate `asyncio.CancelledError` untouched — do not convert cancellation into a provider error.
## Failure Handling
- Raise `ProviderAuthError` for authentication, `ProviderResponseError` for malformed or unusable
responses, and `ProviderError` otherwise.
- Always attach `request_manifest`, `transport_evidence`, and an accurate `failure_phase` to raised
errors. `failure_phase` must distinguish a received-but-failed response from a call that never
reached the provider.
- Validate responses with Pydantic rather than indexing into raw dicts.
- Invalid *optional* metadata (for example unparsable token counts) must not discard an otherwise
valid transcript. Degrade the metadata, not the result.
## Contract Sync Rule
If capture behavior, evidence schema, or the header allowlist changes:
1. Update `docs/invariant/ai_evidence_and_provenance.md` only if the durable preservation contract
itself is changing — that revision is deliberate and reviewed, not incidental.
2. Update `docs/schema.md` when persisted evidence fields change.
3. Update or add tests in the same change (`tests/providers/`, `tests/test_evidence_provenance.py`).
4. Bump the affected evidence `schema_version` when a field's meaning changes.
+105 -34
View File
@@ -17,9 +17,19 @@ applyTo: 'src/transcription/services/*.py'
- **A service module must not import another service module.** This is enforced by - **A service module must not import another service module.** This is enforced by
[test_service_boundaries](../../tests/test_service_boundaries.py). Shared types go in a [test_service_boundaries](../../tests/test_service_boundaries.py). Shared types go in a
neutral module that defines no service class (see [errors](../../src/transcription/services/errors.py)). neutral module that defines no service class (see [errors](../../src/transcription/services/errors.py)).
- Not every module in this package is a service. Helper modules that define no `*Service` - Not every module in this package is a service. Modules fall into three kinds:
class (`base`, `errors`, `normalization`, `prompts`, `quality`, `media_storage`, - **Aggregate services** own models and define a `*Service` class: `documents.py`, `sources.py`,
`source_media`) are free-function modules and are exempt from the service rules below. `jobs.py`, `people.py`, `photos.py`, `maintenance.py`, and `evidence.py` (read/projection only,
owns nothing).
- **Orchestration modules** define no service class and compose writes across aggregates:
`store.py`, `workflows.py`. They are the sanctioned place to create or delete rows owned by more
than one service — see [Service Composition](#service-composition).
- **Shared infrastructure and free-function helpers** are exempt from the service rules below:
`base.py` (`ServiceBase`), `registry.py` (`RegistryService`, a generic base for lookup tables —
not an aggregate owner itself), `unit_of_work.py`, `errors.py`, `normalization.py`, `prompts.py`,
`quality.py`, `media_storage.py`, `source_media.py`. `__init__.py` exposes `ServiceBundle`.
- Cross-cutting error behavior must follow
[error-handling instructions](./error-handling.instructions.md).
## Model Ownership ## Model Ownership
@@ -28,11 +38,38 @@ is the only service that may **create or delete** its rows.
| Model | Owner | | Model | Owner |
| --- | --- | | --- | --- |
| `Document`, `DocumentType` | `DocumentService` | | `Document`, `DocumentType`, `DocumentTag` | `DocumentService` |
| `Source`, `JobSource` | `SourceService` | | `Source`, `JobSource` | `SourceService` |
| `Job` | `JobService` | | `Job` | `JobService` |
| `Person`, `PersonRole`, `DocumentPerson` | `PeopleService` | | `Person`, `PersonRole`, `DocumentPerson`, `PersonTag` | `PeopleService` |
| `ExecutionAttempt` | `EvidenceService` | | `GenealogyPerson`, `GenealogyFamily`, `GenealogyFamilyChild`, `GenealogyCitation` | `MaintenanceService` |
| `Photo` | `PhotosService` |
| `MaintenanceRun` | `MaintenanceService` |
| `ExecutionAttempt` | `SourceService` |
| `Tag` | shared — see below |
Keep this table complete: every table in `src/transcription/db/models.py` appears exactly once,
except `Tag`. When you add a model, add its owner here in the same change.
### `Tag` is deliberately shared
`Tag` is one table reached through two `RegistryService[Tag]` facades that differ only in the
reference model they count usage through: `TagRegistry` (`documents.py`, via `DocumentTag`) and
`PersonTagRegistry` (`people.py`, via `PersonTag`). Both create and delete `Tag` rows through the
generic registry. This is the single sanctioned exception to one-owner-per-model — do not "fix" it by
assigning `Tag` to one service, because the other facade would then be creating rows it does not own.
Any change to `Tag` semantics, labels, or normalization must be validated against **both** facades
and the junction table each one counts.
`DocumentTag` and `PersonTag` follow the junction rule below: each is created and deleted only by the
service on its own side.
### Registries
`DocumentTypeRegistry`, `TagRegistry`, `PersonRoleRegistry`, and `PersonTagRegistry` are
`RegistryService` subclasses, not independent services. A registry belongs to the aggregate service
whose module declares it and shares that service's ownership. Registry CRUD uses `<operation>_entry`
naming (see [CRUD Methods](#crud-methods)).
### Junction tables ### Junction tables
@@ -48,18 +85,22 @@ lifecycle owner. The service on the other side may read through the junction (vi
Two consequences follow, and both are deliberate: Two consequences follow, and both are deliberate:
- **Cascade deletion is not a violation.** A service deleting the aggregate root it owns - **Cascade deletion is not a violation.** A service deleting the aggregate root it owns
may delete junction rows referencing that root, because they cannot outlive it may delete rows referencing that root which cannot outlive it
(`JobService.delete_job_with_guardrails`). (`JobService.delete_job_with_guardrails` deletes the job's `job_source` rows).
- **Evidence deletion is an explicit workflow, not a runtime path.** `JobService.delete_job_and_evidence`
deletes `ExecutionAttempt` rows owned by `SourceService`. That is sanctioned because it is the
named retention workflow that `delete_job_with_guardrails` refuses to perform implicitly — that
method *blocks* deletion when attempts exist. Append-only means runtime code never rewrites or
removes history to represent a new outcome; it does not forbid a deliberate, operator-invoked
retention operation. Do not add a second path that deletes attempts.
- **Ownership governs creation and deletion, not every state transition.** `job_source` is - **Ownership governs creation and deletion, not every state transition.** `job_source` is
both a link and the transcription work queue. `JobService.cancel_job` and both a link and the transcription work queue. `JobService.cancel_job` and
`resubmit_failed_sources` transition `job_source.status` across a whole job, because that `resubmit_failed_sources` transition `job_source.status` across a whole job, because that
transition is a Job lifecycle event, not a per-page outcome. They create and delete transition is a Job lifecycle event, not a per-page outcome. They create and delete
nothing. nothing.
`EvidenceService.promote_machine_attempt` writes two fields on `Source` `EvidenceService` is read-focused and projection-focused. It may coordinate selection
(`preferred_execution_attempt_id`, `raw_transcription`). This is allowed on the same flows, but append-only attempt creation remains in `SourceService` write paths.
principle: selecting which attempt a Source presents is an evidence decision that happens
to land on `Source`. It is scoped to those two projection fields.
If a new operation cannot be expressed within one owner, it belongs in an orchestration If a new operation cannot be expressed within one owner, it belongs in an orchestration
module, not in a cross-service import. module, not in a cross-service import.
@@ -71,14 +112,18 @@ module, not in a cross-service import.
which defines no service class and is therefore importable by any of them. which defines no service class and is therefore importable by any of them.
- Use a context manager for large `try/except` blocks, like `handle_transcription_errors` in - Use a context manager for large `try/except` blocks, like `handle_transcription_errors` in
[sources](../../src/transcription/services/sources.py). [sources](../../src/transcription/services/sources.py).
- Category mapping, retry behavior, and translation boundaries are defined in
[error-handling instructions](./error-handling.instructions.md).
- Service-edge exception translation must be deterministic: map to canonical categories and preserve clear provider->service->API/UI boundaries.
## Checklist ## Checklist
- [ ] Uses `ServiceBase` for common logic - [ ] Uses `ServiceBase` for common logic
- [ ] Session kwarg for `AsyncSession` to pass a session object into each method - [ ] Session kwarg for `AsyncSession` to pass a session object into each method
- [ ] Services use `self._session_scope` in their methods to pass the session through - [ ] Services use `self._session_scope` in their methods to pass the session through
- Multiple operations on the same object(s) require sharing a session between all the methods used - Multiple operations on the same object(s) require sharing a session between all the methods used
- [ ] Every model the module touches is either owned by it or reached read-only - [ ] Every model the module touches is either owned by it or reached read-only
- [ ] Evidence writes preserve append-only semantics
## CRUD Methods ## CRUD Methods
@@ -86,7 +131,8 @@ module, not in a cross-service import.
- Where a service exposes create/read/update/delete for its root model, define them at the - Where a service exposes create/read/update/delete for its root model, define them at the
top of the class in that order, before derived reads and workflow helpers. top of the class in that order, before derived reads and workflow helpers.
- Not every aggregate needs all four. `ExecutionAttempt` is append-only evidence written by - Not every aggregate needs all four. `ExecutionAttempt` is append-only evidence written by
`workflows.py`, so `EvidenceService` deliberately exposes reads and no create or delete. `SourceService` workflow-facing methods, so `EvidenceService` deliberately exposes reads and
no create or delete.
Do not add unused CRUD methods to satisfy symmetry. Do not add unused CRUD methods to satisfy symmetry.
- `RegistryService` is generic across small lookup models and uses `<operation>_entry` - `RegistryService` is generic across small lookup models and uses `<operation>_entry`
naming instead. naming instead.
@@ -99,18 +145,9 @@ When a service method accepts an optional `session` kwarg, write methods must us
- If `session` is provided: the method must **not** commit; it should `flush()` so IDs and FK values are available to the caller's transaction. - If `session` is provided: the method must **not** commit; it should `flush()` so IDs and FK values are available to the caller's transaction.
- Use `refresh()` on returned ORM objects when the caller needs DB-populated values (defaults, triggers, merged state). - Use `refresh()` on returned ORM objects when the caller needs DB-populated values (defaults, triggers, merged state).
Recommended helper behavior:
- Inputs: active session object, original `session` arg (or a boolean ownership flag), and an optional list of objects to refresh.
- Logic: `commit` when service-owned session, `flush` when caller-owned session, then refresh requested objects.
This keeps orchestration functions atomic: they can pass one shared session across multiple services and commit exactly once at the workflow boundary.
## Workflow Transaction Boundaries ## Workflow Transaction Boundaries
For multi-step job lifecycles (for example queued transcription jobs), orchestration functions must use explicit transaction phases. For multi-step job lifecycles, orchestration functions must use explicit transaction phases.
Required boundary model:
- **Transaction A (claim):** transition `JobStatus.QUEUED -> JobStatus.PROCESSING` and commit immediately. - **Transaction A (claim):** transition `JobStatus.QUEUED -> JobStatus.PROCESSING` and commit immediately.
- Perform provider/network work **outside** database transactions. - Perform provider/network work **outside** database transactions.
@@ -124,21 +161,55 @@ Atomicity rules:
- Terminal state (`TRANSCRIBED` or `FAILED`) and transcript row changes must succeed or roll back together. - Terminal state (`TRANSCRIBED` or `FAILED`) and transcript row changes must succeed or roll back together.
- Retry persistence (`QUEUED` + retry increment + error detail) must succeed or roll back together. - Retry persistence (`QUEUED` + retry increment + error detail) must succeed or roll back together.
Separation of concerns: ### Multi-page batches
- Worker modules should stay lightweight and delegate lifecycle transitions to service/workflow orchestration functions. These two requirements are in tension for multi-page jobs: each page should be durable as
- In `workflows.py`, `process_queued_job` should own one complete attempt lifecycle: `QUEUED -> PROCESSING -> TRANSCRIBED|FAILED`. soon as its provider call returns, but the last page must commit together with the terminal
- In `workflows.py`, `advance_job` should coordinate broader status progression around attempts (for example retry scheduling from `FAILED -> QUEUED`). status. `process_queued_job` resolves it by committing every page except the last one
- Services should expose session-aware write helpers (flush on caller-owned session) so orchestration controls commit boundaries. individually, then deferring the final page's write into `_finalize_batch_outcome` so it
- Backoff/sleep behavior must run outside transactional scopes. shares the terminal transaction.
Both paths are shielded against cancellation, so the final page is no less durable than the
pages before it. Enforced by `tests/integration/test_pipeline_atomicity.py`; per-page
durability is separately enforced by
`tests/services/test_workflows_reliability.py::TestWorkflowReliability::test_transcribed_page_is_committed_before_next_provider_call_finishes`.
### Stale-reclaim safety
- `WORKER_STALE_JOB_SECONDS` must remain **greater than** `WORKER_PROVIDER_TIMEOUT_SECONDS`; stale
recovery must not be able to fire before one provider call can legitimately finish.
- Long-running multi-page orchestration must refresh job liveness explicitly between intermediate
page commits. Do not rely on incidental row updates or provider metadata writes to keep
`Job.date_updated` fresh.
- Enforced by `tests/test_config.py` and
`tests/services/test_workflows_reliability.py::TestWorkflowReliability::test_intermediate_page_commit_advances_job_liveness_timestamp`.
## Contract Alignment
- Treat `docs/` as the active architecture and requirements baseline.
- Legacy revision trees are out of scope for active implementation decisions and must not be referenced as authoritative service guidance.
- Treat `src/transcription/db/models.py` as runtime schema ground truth and `docs/schema.md` as the field-accurate contract mirror.
- `Job.status` success path is `TRANSCRIBED`.
- `JobSource.status` is queue/projection state only (`PENDING`, `TRANSCRIBED`, `FAILED`, `CANCELLED`).
- Source ingest may normalize media before persistence; persisted bytes/hash are canonical for processing and provenance.
- `ExecutionAttempt` is append-only evidence history; do not mutate historical attempt rows in runtime code.
- `Source.raw_transcription` is a projection, not authoritative history.
- Service/UI read paths that touch relationships must be eager-loaded for `lazy="raise"` compatibility.
- If model fields, enums, constraints, indexes, or relationship-loading semantics change, update `docs/schema.md` in the same change.
- If `Settings` fields or defaults change in `src/transcription/config.py`, update `.env.production.example` in the same change so keys/defaults remain synchronized and no stale settings remain documented.
## Schema Drift and Legacy Compatibility Policy
- Prefer schema migration over startup reconciliation or runtime compatibility paths in service writes.
- Do not add legacy read/write compatibility code in service workflows by default.
- If drift is discovered and a migration decision is ambiguous (for example, one-way destructive DDL, uncertain data retention impact, or unknown deployment sequence), pause and ask the user to choose migration vs compatibility before coding.
- If a temporary compatibility path is explicitly approved, document an expiration/removal plan in the same change.
# Service Composition # Service Composition
A service method may read across models it does not own, using eager loads from its own A service method may read across models it does not own, using eager loads from its own
aggregate root. What it may not do is import another service. aggregate root. What it may not do is import another service.
Operations that must **write** models owned by more than one service — uploading a picture, Operations that must **write** models owned by more than one service are composed in an orchestration module
for example — are composed in an orchestration module
([store](../../src/transcription/services/store.py), ([store](../../src/transcription/services/store.py),
[workflows](../../src/transcription/services/workflows.py)). Orchestration modules define no [workflows](../../src/transcription/services/workflows.py)).
service class, may import any service, and own the commit boundary.
+119
View File
@@ -0,0 +1,119 @@
---
description: Authoring rules for the test suite, including markers, async discipline, and guard-test design.
applyTo: 'tests/**/*.py'
---
# Tests
Primary references:
- `AGENTS.md` (Change Protocol — failing test first)
- `docs/index.md` and `docs/invariant/*`
- `.github/skills/test-effectiveness-auditor/skill.md` (periodic audit of this suite)
The suite is not only regression protection here — it is where several architectural rules are
*defined*. `tests/test_service_boundaries.py`, `tests/test_ui_boundaries.py`,
`tests/test_provider_boundaries.py`, `tests/test_model_contract_guards.py`, and
`tests/test_meta_contract_guards.py` are the enforcement layer named in the `AGENTS.md` authority
order. A weak test in this repository does not merely fail to catch a bug; it can silently repeal a
documented invariant.
The baseline is green. `uv run pytest -q -m "not external"` must report zero failures and zero
errors, and there is no tolerated set of known-failing tests.
## Write the Failing Test First
For any behavioral fix, write the test before the fix and confirm it fails *for the reason you
expect*. A test that passes against the broken code proves nothing, and several defects in this
repository were subtle enough that a test written afterward would have done exactly that. If the
new test passes immediately, you have not reproduced the defect yet.
## Runner Configuration
Configured in `pyproject.toml`; do not work around these:
- `--strict-markers` — an unregistered marker is an error. Register new markers in
`[tool.pytest.ini_options] markers` with a description rather than inventing one at the call site.
- `asyncio_mode = "strict"` — every async test needs an explicit `@pytest.mark.asyncio`, and async
fixtures use `@pytest_asyncio.fixture`. There is no implicit promotion.
- `filterwarnings = ["error:coroutine .* was never awaited:RuntimeWarning"]` — an un-awaited
coroutine is an error, not a warning. This usually means a mock replaced an async callable with a
sync one, or an `await` was dropped. Fix the call; never silence the warning.
## Markers and Layout
- `unit` — pure logic, no framework or database.
- `integration` — touches framework, database, or multi-component contracts.
- `external` — calls live services; slow and credential-dependent.
`external` tests must also carry their own `skipif` so the suite stays green without credentials
(see `tests/services/test_transcription_external.py`). Local and documented runs use
`-m "not external"`; CI intentionally runs unfiltered, which is equivalent because those tests skip
themselves. Never let an unmarked test reach the network.
Place tests by the layer under test: `tests/services/`, `tests/ui/`, `tests/api/`,
`tests/providers/`, `tests/integration/`, with cross-cutting guards at the top level.
## Fixtures and Isolation
- Prefer the shared fixtures in `tests/conftest.py` (`default_settings`, `async_session`,
`default_session_factory`, and the per-aggregate service fixtures) over building settings or
engines by hand.
- `Settings` is isolated suite-wide by the session-scoped autouse fixture in `conftest.py`, because
`env_file` resolves against the working directory. Tests that need env-file loading pass
`_env_file=` explicitly; tests asserting declared defaults need nothing. Do not reintroduce
reliance on a developer's local env file. Guarded by `tests/test_config_isolation.py`.
- Database fixtures refuse to run against anything but the per-test path, and that refusal is
deliberate. Never relax it to point a destructive fixture at a real database.
- Tests must not leave artifacts outside `tmp_path`.
## Assertion Strength
Assert on the domain effect, not on the fact that code ran.
- Prefer persisted state, status transitions, error categories, and evidence records over
"no exception raised", "not None", or a bare status code.
- **Read committed state through a separate session.** Asserting against the same session that
performed the write can pass on unflushed in-memory state and prove nothing about durability.
This is how the atomicity guarantees in `tests/services/test_workflows_reliability.py` and
`tests/integration/test_pipeline_atomicity.py` are made real.
- Critical paths need negative-path coverage — timeouts, provider failures, validation errors,
cancellation. Happy-path-only coverage of a critical module is a gap, not a suite.
- Avoid count-threshold assertions as a proxy for correctness. A test asserting "at least N items
were discovered" passes indefinitely while the thing it was meant to protect rots; assert on a
specific known member instead.
## Guard Tests
Structural guards carry extra obligations, because they are cited as proof that a rule holds.
- **Guard the guard.** Every scanning guard needs a companion assertion that the scan actually found
something, following the existing `test_*_are_discovered` pattern. A guard that silently scans an
empty set passes forever.
- **Scope must match the claim.** A guard's name and docstring must describe only what it actually
verifies. A test covering one function while appearing to enforce a repo-wide rule is worse than
no test, because it stops anyone from writing the real one.
- **Prove non-vacuity by injected fault.** Temporarily introduce the violation, confirm the guard
fails with a comprehensible message, then revert. Do this whenever you add or materially change a
guard. Revert with an explicit edit if the file has uncommitted changes — `git checkout --` will
discard them.
- **Prefer structural analysis to substring matching.** AST inspection of imports and definitions is
resistant to false negatives; a bare-name search across the repository is not, since an unrelated
mention anywhere makes dead code look reachable.
- Failure messages should name the offending file, symbol, and the remedy. These fire for people who
did not write the guard.
- Any new file under `.github/**` must be added to `ACTIVE_CONTRACT_FILES` in
`tests/test_meta_contract_guards.py`, or the completeness guard fails by design.
## Redundancy
Duplicate coverage across layers costs runtime and dilutes signal. Pick the canonical layer for a
behavior — unit for logic, integration for wiring — and let the other layer assert only what is
unique to it. Retire tests superseded by a stronger guard instead of accumulating both, and record
deliberate retentions with a rationale rather than leaving them unexplained.
## Contract Sync Rule
When a test encodes or relaxes a documented rule, update the corresponding instruction file or
`docs/*` page in the same change. When a guard test is the enforcement for a rule stated in
`AGENTS.md` or an instruction file, cite the test by name there so the link survives refactoring.
+28 -3
View File
@@ -11,6 +11,9 @@ Keep dependencies flowing in this direction:
Pages may depend on application services and framework-provided dependencies. Components may depend on smaller components and shared presentation helpers. Services and domain modules must never depend on the UI. Pages may depend on application services and framework-provided dependencies. Components may depend on smaller components and shared presentation helpers. Services and domain modules must never depend on the UI.
Cross-cutting error behavior must follow
[error-handling instructions](./error-handling.instructions.md).
## Package Root ## Package Root
- Keep `ui/__init__.py` as the UI composition root: register global assets, register pages, and mount NiceGUI on FastAPI. - Keep `ui/__init__.py` as the UI composition root: register global assets, register pages, and mount NiceGUI on FastAPI.
@@ -37,17 +40,39 @@ Pages may depend on application services and framework-provided dependencies. Co
- Keep app-wide navigation and layout primitives in `components/app_shell.py`. - Keep app-wide navigation and layout primitives in `components/app_shell.py`.
- Keep generic table/event adaptation in `components/table/common.py`; feature-specific columns, row read models, and formatting belong in the feature table module. - Keep generic table/event adaptation in `components/table/common.py`; feature-specific columns, row read models, and formatting belong in the feature table module.
- Keep exception normalization and user-facing error display in `components/error_presenter.py`; preserve `AppError` details and operation identifiers at page/component boundaries. - Keep exception normalization and user-facing error display in `components/error_presenter.py`; preserve `AppError` details and operation identifiers at page/component boundaries.
- Use `components/media_urls.py` for media URL generation; do not hand-build upload/static paths in page code.
## CSS Assets ## CSS Assets
- Keep all application CSS in `ui/static/theme.css`; do not add page- or component-specific stylesheets or embed style blocks in Python components. - Keep all application CSS in `ui/static/theme.css`; do not add page- or component-specific stylesheets or embed style blocks in Python components.
- Load `theme.css` once from the composition root with `ui.add_css(..., shared=True)`. - Load `theme.css` once from the composition root with `ui.add_css(..., shared=True)`.
- Read stylesheet text through `importlib.resources.files(...)` so loading works from installed packages and is independent of the working directory. - Read stylesheet text through `importlib.resources.files(...)` so loading works from installed packages and is independent of the working directory.
- Centralize CSS reading in one typed helper cached by resource path with `functools.cache` or an equivalent unbounded `lru_cache`. Cache the immutable stylesheet text to avoid repeated resource I/O; keep NiceGUI registration at the composition root. - Centralize CSS reading in one typed helper cached by resource path.
- Do not encode application behavior in CSS or other static assets. - Do not encode application behavior in CSS or other static assets.
## State and Side Effects ## State and Side Effects
- Limit component state to ephemeral interaction state such as loading flags, form values, dialogs, and expansion state. - Limit component state to ephemeral interaction state such as loading flags, form values, dialogs, and expansion state.
- Application and worker state must be resolved at the page or application boundary and passed through narrow interfaces such as callbacks or notifier protocols. - Application and worker state must be resolved at the page or application boundary and passed through narrow interfaces.
- Keep filesystem, network, provider, and worker orchestration behind application services or dedicated adapters. UI code may trigger those operations but must not implement them. - Keep filesystem, network, provider, and worker orchestration behind application services or dedicated adapters.
## Media Route Safety Rules
Two patterns are approved:
1. **Record-validated API routes** for print/export contexts.
2. **Controlled upload URL resolver** (`components/media_urls.py`) for general UI media.
Prohibited patterns:
- Direct `file://` links or exposing local filesystem paths.
- Manual URL construction from raw `Path` values in pages/components.
- User-facing payloads containing local absolute paths.
## Contract Alignment
- Treat `docs/` as the active baseline.
- Resolve lifecycle and status semantics against `src/transcription/db/models.py` and `docs/schema.md`; do not introduce alternate status labels or implied legacy states in UI behavior.
- Use status vocabulary exactly as modeled (`queued`, `processing`, `transcribed`, `partial_success`, `failed`; and `pending`, `transcribed`, `failed`, `cancelled`).
- Print/export media flows must use record-validated routes; direct local filesystem paths are prohibited.
- If lifecycle wording/behavior changes, update corresponding `docs/ui/pages/*.md` contracts in the same change.
@@ -9,7 +9,7 @@ agent: Python Architect Reviewer
Execute a comprehensive, evidence-based code review of the target codebase. Execute a comprehensive, evidence-based code review of the target codebase.
## Target Scope ## Target Scope
- **Review Target:** ${{input:target_path:./}} - **Review Target:** the repository root, unless the invoker names a narrower path; review that path instead.
- **Source Root:** `src/` - **Source Root:** `src/`
- **Docs Root:** `docs/` - **Docs Root:** `docs/`
- **Focus Areas:** FastAPI endpoints, NiceGUI components, SQLModel persistence, asyncio workers, Pydantic V2 models, and OpenRouter provider adapters. - **Focus Areas:** FastAPI endpoints, NiceGUI components, SQLModel persistence, asyncio workers, Pydantic V2 models, and OpenRouter provider adapters.
@@ -17,7 +17,7 @@ Execute a comprehensive, evidence-based code review of the target codebase.
## Execution Rules ## Execution Rules
1. Map repository layout, dependency manifests, and configuration files from the project root before inspecting modules. 1. Map repository layout, dependency manifests, and configuration files from the project root before inspecting modules.
2. Read real code modules under `src/` (or the specified target path); cite exact file paths and line ranges for every finding. 2. Read real code modules under `src/` (or the specified target path); cite exact file paths and line ranges for every finding.
3. Validate issues using terminal tools (`ruff check`, `pytest`, `ty`) where appropriate. 3. Validate issues by running `uv run ruff check .`, `uv run ty check`, and `uv run pytest -q -m "not external"`. Nothing is on `PATH` in this `uv` project, so bare `ruff`/`ty`/`pytest` will fail.
4. Check for duplication, divergent implementations, and extractable helpers. 4. Check for duplication, divergent implementations, and extractable helpers.
5. Format the entire review following the standardized 6-section template defined in the `python-code-reviewer` skill. 5. Format the entire review following the standardized 10-section template defined in the `python-code-reviewer` skill.
6. Write the final report as a Markdown file to `./docs/code-review-${{current_date}}.md`. 6. Write the final report to `./docs/reviews/<YYYY-MM-DD>-code-review.md`, using today's date. This path is defined by the skill; do not write the report anywhere else.
@@ -0,0 +1,80 @@
---
name: evidence-provenance-auditor
description: Deterministic reviewer for transcription evidence/provenance guarantees. Use when changes touch execution attempts, source storage, retries, transport evidence, artifact provenance, or evidence exports.
---
# Evidence & Provenance Auditor
Perform focused, deterministic audits of evidence integrity and provenance behavior.
## When to Use
- Reviewing changes in:
- `src/transcription/services/sources.py`
- `src/transcription/services/store.py`
- `src/transcription/services/workflows.py`
- `src/transcription/services/evidence.py`
- `src/transcription/db/models.py`
- Auditing evidence exports/imports or evidence-display behavior.
- Verifying no drift from canonical provenance invariants.
## Normative References (must be used)
1. `docs/invariant/ai_evidence_and_provenance.md`
2. `docs/schema.md`
3. `docs/requirements.md`
4. `docs/error_handling.md`
## Deterministic Pass/Fail Checks
### A. Append-only history
- Every provider call results in a new `ExecutionAttempt`.
- Runtime paths do not mutate historical attempts to represent new outcomes.
- Retry behavior appends attempts rather than rewriting prior rows.
### B. Projection vs authority separation
- `Source.raw_transcription` and preferred pointers are mutable projection surfaces.
- Attempt rows remain authoritative historical evidence.
- Candidate promotion updates projection pointers without rewriting history.
### C. Transport evidence semantics
- Transport evidence is correctly labeled as application-boundary capture.
- SDK snapshots/normalized metadata are not mislabeled as native upstream payload.
- No-response timeout/network states are explicit.
### D. Canonical source identity
- Canonical stored bytes/hash/size are internally consistent.
- If ingest normalization is applied, code/docs consistently represent resulting canonical identity.
- Post-ingest derivatives do not overwrite canonical source bytes.
### E. Secret safety
- No credentials/auth headers/cookies/unrestricted headers persisted.
- Header persistence uses explicit allowlist semantics.
### F. Route/path safety
- Print/export source access is record-validated.
- UI/media path construction does not expose local filesystem paths.
### G. Schema/docs alignment
- Evidence-related model fields and semantics align with canonical docs.
- Evidence model changes require same-change doc updates.
### H. Canonical authority boundaries
- Active guidance resolves against `docs/*` and current instruction files.
## Review Workflow
1. Read normative references first.
2. Inspect model + service + workflow write paths.
3. Inspect evidence read/display/export paths.
4. Report high-confidence findings with concrete path/line evidence.
5. Classify each finding by invariant family (A-H).
## Output Format
Use this structure:
- Verdict by invariant family (A-H)
- Findings with `Location`, `Observed Behavior`, `Risk`, `Recommended Fix`
- Drift table (`Doc claim` vs `Code reality` vs `Action`)
- Regression guards needed
+146 -16
View File
@@ -11,7 +11,7 @@ Perform thorough, evidence-based code reviews for Python projects. Every finding
- Performing an architectural or code quality review of a Python codebase. - Performing an architectural or code quality review of a Python codebase.
- Auditing applications using FastAPI, NiceGUI, SQLModel/SQLAlchemy, Pydantic V2, or asyncio workers. - Auditing applications using FastAPI, NiceGUI, SQLModel/SQLAlchemy, Pydantic V2, or asyncio workers.
- Generating structured Markdown review reports in `./docs`. - Generating structured Markdown review reports in `./docs/reviews`.
## Technical Stack Scope ## Technical Stack Scope
@@ -21,15 +21,63 @@ Perform thorough, evidence-based code reviews for Python projects. Every finding
- **Validation & Settings:** Pydantic V2 and pydantic-settings - **Validation & Settings:** Pydantic V2 and pydantic-settings
- **Concurrency:** Python asyncio workers - **Concurrency:** Python asyncio workers
- **Vision/LLM Integration:** OpenRouter / provider adapters - **Vision/LLM Integration:** OpenRouter / provider adapters
- **Image & Print Pipeline:** Pillow-backed media handling and print/export rendering
- **Quality & Testing:** pytest, pytest-asyncio, Ruff, and ty - **Quality & Testing:** pytest, pytest-asyncio, Ruff, and ty
NiceGUI is pinned to an exact version (`nicegui==3.13.0` in `pyproject.toml`); API guidance
must be correct for that release rather than for the latest published version. The exact pin is
a deliberate release-stability decision recorded in `docs/production-runbook.md` ("Dependency
upgrade policy") — do not report it as a defect or recommend widening it.
## Review Workflow ## Review Workflow
1. **Map the Repository First:** Inspect entry points, package layout, configurations, dependency manifests, and any project-specific rule files (`AGENTS.md`, `CLAUDE.md`, `.github/instructions/`). Project-specific conventions override generic advice. 1. **Map the Repository First:** Inspect entry points, package layout, configurations, dependency manifests, and any project-specific rule files (`AGENTS.md`, `.github/instructions/`, `.github/skills/`). Project-specific conventions override generic advice.
2. **Read Representative Modules:** Sample across all layers (routes/pages, UI components, services, workers, persistence, provider adapters, settings, tests) before drawing conclusions. 2. **Establish Canonical Authority First:** Read architecture/contracts (`docs/*`, `docs/invariant/*`, UI docs) and active instructions/skills before evaluating source behavior.
3. **Verify Claims:** Run or reference project tooling (`ruff check`, `ty`, `pytest`) rather than guessing. 3. **Read Representative Modules:** Sample across all layers (routes/pages, UI components, services, workers, persistence, provider adapters, settings, tests) before drawing conclusions.
4. **Prioritize Hot Paths:** Focus deeply on request handling, database sessions, background workers, and external API calls. 4. **Run Drift Analysis:** Compare documented intended behavior versus repository ground truth; identify both implementation drift and undocumented-but-repeatable conventions that should be formalized.
5. **Enforce Read-Only Safety:** Do not modify code unless explicitly instructed. 5. **Run Dead-Code/Orphan Sweep:** Identify candidate orphan modules/functions/classes with zero inbound references, then verify expected exceptions (entrypoints, framework/plugin registration, dynamic imports/reflection, CLI hooks, test-only utilities) before marking as orphaned.
6. **Assess Boundary and Coupling Health:** Evaluate UI/service/persistence/provider dependency flow, identify circular dependencies, leaky abstractions, and transaction ownership ambiguity.
7. **Assess Invariant Placement:** For each hard rule, decide whether it belongs in docs (rationale), instructions (active steering), skills (periodic audit procedure), or deterministic tests (enforcement).
8. **Verify Claims:** This is a `uv` project (`uv.lock`, root `ruff.toml`). Run `uv run ruff check .`, `uv run ty check`, and `uv run pytest -q -m "not external"` rather than guessing, and record the exact commands and their outcomes in the report.
9. **Validate Recommendations Against Consumers:** A recommendation is a claim about the future and must be verified like any other. Before recommending a change to a shared symbol — a model field, an exception attribute, a helper's return value, a function signature — enumerate **every** consumer of that symbol (`grep` the whole repo, including tests) and confirm the fix is safe for each one. Record the consumers in the finding's **Blast Radius**. A fix that is correct for the path that produced the finding can silently break a second consumer, and evidence/provenance and logging paths are the usual casualties because they read the same fields the UI does.
10. **Prioritize Hot Paths:** Focus deeply on request handling, database sessions, background workers, and external API calls.
11. **Enforce Read-Only Safety:** Do not modify code unless explicitly instructed.
12. **Escalate Provenance Audits:** For evidence/provenance-heavy changes, apply invariant checks from `.github/skills/evidence-provenance-auditor/skill.md` and include pass/fail outcomes in the report.
13. **Escalate Test-Suite Audits:** When findings touch test coverage, redundancy, or assertion strength, apply `.github/skills/test-effectiveness-auditor/skill.md` and include its outcomes alongside the provenance results.
### Worked example: why step 9 exists
The 2026-08-23 review recommended fixing a filesystem-path leak in
`classify_unexpected_error` by making `AppError.message` generic and logging the exception
detail instead. The analysis of the leak was correct, and the fix was implemented as written.
It was wrong. `AppError.message` had a second consumer the review never traced:
`format_error_detail`, which writes `ExecutionAttempt.error_detail` — a **provenance record**.
The recommended fix closed a privacy leak by silently stripping root-cause data from the
evidence history this system exists to preserve. It was caught only because an unrelated
integration test asserted on the persisted error text.
The correct fix separated the audiences — a user-safe `message` and an internal-only `detail`
that still reaches evidence and logs. One `grep` for consumers of `.message` during the review
would have found this. Treat any recommendation that changes a widely-read field as unverified
until its consumers are enumerated.
## Repo-Specific Deterministic Checks (Transcription)
When reviewing this repository, always include explicit pass/fail checks for the following.
Where **Enforced by** reads *unenforced*, recommending a deterministic test is itself a finding.
| # | Check | Enforced by |
| :-- | :--- | :--- |
| 1 | **Service boundary rule:** no service-to-service imports | `tests/test_service_boundaries.py` |
| 2 | **UI boundary rule:** pages/components do not perform persistence access | `tests/test_ui_boundaries.py` |
| 3 | **Status vocabulary conformance:** `JobStatus`/`JobSourceStatus`/`JobPurpose` usage matches current enums in `src/transcription/db/models.py`; no stringly-typed status literals | `tests/test_model_contract_guards.py` |
| 4 | **Evidence ownership conformance:** append-only attempt history is preserved and projection writes are not mistaken for history mutation (`src/transcription/services/sources.py`, `src/transcription/services/evidence.py`) | `tests/test_evidence_provenance.py::test_attempts_are_append_only_and_exported_with_integrity` |
| 5 | **Canonical authority:** findings must resolve against `docs/*` first | `tests/test_meta_contract_guards.py::test_canonical_authority_references_are_present` |
| 6 | **Schema contract fidelity:** when model/persistence behavior changes, `docs/schema.md` remains field-accurate with `src/transcription/db/models.py` | `tests/test_model_contract_guards.py` (field names, ordering, enum members, table coverage), `tests/test_meta_contract_guards.py` (presence and references) |
| 7 | **Media boundary conformance:** print/export media is record-validated and UI media URL generation uses controlled resolver paths | `tests/test_media_path_safety.py`, `tests/ui/test_media_urls.py` |
| 8 | **Eager-loading conformance:** service/UI read paths satisfy `lazy="raise"` expectations | `tests/test_model_contract_guards.py` (declaration-side; documented `noload` exceptions must match `docs/schema.md`) |
| 9 | **Cross-cutting error conformance:** service/API/UI translation and retry behavior align with `.github/instructions/error-handling.instructions.md` | `tests/test_errors.py`, `tests/api/test_error_responses.py`, `tests/ui/test_error_presenter.py` |
| 10 | **Orphaned/dead-code conformance:** include a deterministic orphan sweep and report confirmed orphans removed/retained with rationale | `tests/test_orphan_sweep.py` (`KNOWN_ORPHANS` records each retained orphan and its rationale) |
## Core Review Areas ## Core Review Areas
@@ -70,13 +118,61 @@ Perform thorough, evidence-based code reviews for Python projects. Every finding
### 8. Testing & Quality Tooling ### 8. Testing & Quality Tooling
- **Test Isolation:** Verify tests do not rely on live external services, real clocks, or shared global state. - **Test Isolation:** Verify tests do not rely on live external services, real clocks, or shared global state.
- **Async Test Setup:** Check `pytest-asyncio` configuration and fixture lifecycle. - **Async Test Setup:** Check `pytest-asyncio` configuration and fixture lifecycle.
- **Project Test Contract (`pyproject.toml`):** `--strict-markers` is enabled, so every marker must be declared; `asyncio_mode = "strict"` requires explicit `@pytest.mark.asyncio`; declared markers are `unit`, `integration`, and `external`, and `external` must be excluded from default verification runs. `filterwarnings` promotes `coroutine ... was never awaited` to an **error** — treat any unawaited coroutine as a hard failure and a Critical/High finding, never a warning.
- **Suite Signal Quality:** For low-value, redundant, or tautological tests, escalate to `.github/skills/test-effectiveness-auditor/skill.md` and fold its outcomes into the report.
### 9. Duplication & Consolidation ### 9. Duplication & Consolidation
- Identify repeated code blocks, candidate helper abstractions, divergent patterns for identical operations, and duplicated domain constants. - Identify repeated code blocks, candidate helper abstractions, divergent patterns for identical operations, and duplicated domain constants.
### 10. Orphaned/Dead Code Audit
- Find candidate orphan modules/functions/classes with no inbound references.
- Validate each candidate against dynamic wiring exceptions (entrypoints, plugin registration, reflection/dynamic imports, CLI hooks, test utilities).
- Report outcomes as: removed orphan, retained-with-justification, or uncertain-follow-up.
### 11. Architecture & Governance
- **Architectural Drift:** Compare intended architecture rules against implementation behavior and cite concrete drift points.
- **Systemic Health:** Evaluate domain cohesion, dependency direction, lifecycle consistency, and operational reliability seams.
- **Invariant Routing:** Recommend the correct enforcement layer per rule (docs vs instructions vs skills vs tests).
- **Meta-Tooling Alignment:** Recommend updates for instruction files and skills when repository patterns or contracts evolve.
## Severity Rubric
Severity reflects concrete consequence, never style preference or effort to fix.
- **Critical:** Data or evidence loss/corruption; provenance or append-only history violated; secret leakage; silent wrong output presented as authoritative.
- **High:** Architectural boundary violated (service/UI/persistence/provider); runtime failure or unhandled exception on a hot path (request handling, DB sessions, worker loop, external API calls); documented invariant contradicted by implementation.
- **Medium:** Correctness risk under load or edge conditions (N+1, missing eager load, leaked task, missing timeout); drift between docs and code with no immediate runtime impact.
- **Low:** Maintainability, typing completeness, duplication, naming, or dead code with no behavioral risk.
### Reachability
Severity states how bad the consequence is; **Reachability** states whether it can happen today.
They are independent, and a finding is not complete without both. Record one of:
- **Live:** reachable in the current configuration and deployment.
- **Latent:** the defective code is present but unreachable because of a current setting, single-
instance deployment, or absent caller. **State the exact condition that unblocks it.**
- **Theoretical:** requires a combination the project has explicitly ruled out.
Latent findings carry a scheduling constraint that severity alone cannot express: a latent defect
must usually be fixed *before* the change that makes it live, not after. Say so explicitly in the
finding and reflect the ordering in the §9 action plan — for example, "fix the retry-category gate
before raising `worker_max_retries` above 0," or "handle this `IntegrityError` before deploying a
second worker replica." Do not downgrade severity merely because a finding is latent.
### Conflicting invariants
When a fix sits between two invariants that pull in opposite directions, say so in the
**Recommendation** and name both, along with the test that guards each. Flag explicitly what the
over-correction would be, because the simplest-looking fix usually satisfies one invariant by
silently destroying the other. A recommendation that resolves one side without naming the other is
incomplete and will be implemented incorrectly.
## Output Report Structure & Template ## Output Report Structure & Template
Generate Markdown reports in `./docs` following this exact template structure: Generate Markdown reports at `./docs/reviews/<YYYY-MM-DD>-code-review.md` following this exact
template structure. Reports are dated, non-canonical artifacts: `docs/reviews/**` is explicitly
**not** part of the canonical authority set that the canonical-authority check resolves against.
```markdown ```markdown
# Architecture & Code Review Report # Architecture & Code Review Report
@@ -91,13 +187,24 @@ Generate Markdown reports in `./docs` following this exact template structure:
--- ---
## 2. Findings by Severity ## 2. Executive Architecture Assessment
- High-level verdict on domain cohesion, boundary clarity, and architecture fitness.
- Top 3-5 systemic risks or bottlenecks.
---
## 3. Findings by Severity
### Critical Severity ### Critical Severity
#### [CRIT-01] Title #### [CRIT-01] Title
- **Location:** `path/to/file.py:lines` - **Location:** `path/to/file.py:lines`
- **Reachability:** Live / Latent (state the exact condition that unblocks it) / Theoretical
- **Problem & Consequence:** Concrete consequence, not a style opinion. - **Problem & Consequence:** Concrete consequence, not a style opinion.
- **Recommendation:** Fix with before/after sketch. - **Blast Radius:** Every consumer of the symbols the recommendation changes, each confirmed
safe. Write `None — change is local` only after actually searching. If the fix touches a
shared field or helper, list the call sites (including tests and evidence/logging paths).
- **Recommendation:** Fix with before/after sketch. If two invariants conflict here, name both,
name the test guarding each, and state what the over-correction would be.
- **Effort:** S / M / L - **Effort:** S / M / L
### High Severity ### High Severity
@@ -114,7 +221,23 @@ Generate Markdown reports in `./docs` following this exact template structure:
--- ---
## 3. Stack-Specific Analysis ## 4. Architectural Drift & Gap Analysis
`Direction` is `doc->code` (implementation must change to match documented intent) or
`code->doc` (an undocumented but repeatable convention that should be formalized).
| Area / Component | Direction | Documented / Intended Rule | Actual Implementation State | Severity | Recommended Resolution |
| :--- | :--- | :--- | :--- | :--- | :--- |
---
## 5. Invariant Inventory & Routing Recommendations
| Invariant / Constraint | Current Location | Recommended Target Layer | Rationale |
| :--- | :--- | :--- | :--- |
---
## 6. Stack-Specific Analysis
- Python 3.12+ Best Practices - Python 3.12+ Best Practices
- FastAPI - FastAPI
- NiceGUI - NiceGUI
@@ -126,7 +249,7 @@ Generate Markdown reports in `./docs` following this exact template structure:
--- ---
## 4. Duplication & Consolidation Report ## 7. Duplication & Consolidation Report
| Pattern / Duplication | Locations | Proposed Canonical Home | Estimated Lines Removed | | Pattern / Duplication | Locations | Proposed Canonical Home | Estimated Lines Removed |
| :--- | :--- | :--- | :--- | | :--- | :--- | :--- | :--- |
@@ -135,12 +258,19 @@ Generate Markdown reports in `./docs` following this exact template structure:
--- ---
## 5. Prioritized Action Plan ## 8. Meta-Tooling & Instruction Update Recommendations
1. **Phase 1: Quick Wins (PR 1-2)** - Required updates to docs, instructions, skills, or tests to keep enforcement current.
2. **Phase 2: Reliability & Concurrency (PR 3-4)**
3. **Phase 3: Consolidation & Refactoring (PR 5-6)**
--- ---
## 6. Preserved Strengths ## 9. Prioritized Dependency-Ordered Action Plan
1. **Phase 1: Blocking fixes**
2. **Phase 2: Enforcement hardening**
3. **Phase 3: Reliability & concurrency**
4. **Phase 4: Consolidation & refactoring**
5. **Phase 5: Non-blocking governance/documentation depth**
---
## 10. Preserved Strengths
- Existing patterns worth maintaining. - Existing patterns worth maintaining.
@@ -0,0 +1,98 @@
---
name: test-effectiveness-auditor
description: Periodic reviewer for test-suite signal quality. Detects low-value or redundant tests, validates contract coverage, and recommends pruning or strengthening actions.
---
# Test Effectiveness Auditor
Run a deterministic audit of test usefulness. Focus on whether tests catch real regressions, not whether they merely execute code.
## When to Use
- Monthly/quarterly test-health review.
- Pre-release hardening when test count grows quickly.
- After major AI-assisted test generation.
- When suite runtime is increasing without clear quality gains.
## Primary Objectives
1. Identify tests that are weak, redundant, or non-diagnostic.
2. Confirm critical contracts are guarded by meaningful assertions.
3. Produce a prune/strengthen backlog with explicit risk and effort.
## Normative References (Transcription Repo)
1. `docs/*`
2. `docs/invariant/*`
3. `.github/instructions/*.instructions.md`
4. `tests/test_meta_contract_guards.py`
5. Contract-specific guards (`tests/test_service_boundaries.py`, `tests/test_ui_boundaries.py`, worker/evidence/media/error suites)
## Deterministic Audit Checks
### A. Contract Traceability
- Each high-risk contract maps to at least one focused regression test file.
- Missing mapping is a gap.
### B. Assertion Strength
- Flag tests that only assert status code, non-null, or “no exception” without validating state transitions or persisted outcomes.
- Prefer assertions on domain effects: DB rows, status changes, error categories, evidence writes, or emitted payload shape.
### C. Failure-Path Coverage
- Critical paths must include negative-path tests (timeouts, provider errors, validation failures, cancellation paths, retries).
- Happy-path-only coverage on critical modules is a gap.
### D. Redundancy and Noise
- Detect near-duplicate tests asserting the same behavior at multiple layers with no extra signal.
- Recommend canonical location (unit/integration) and prune overlaps.
### E. Mutation/Change Sensitivity
- Prefer mutation testing for high-risk modules when practical.
- If not run, identify tests likely to survive meaningful code mutations (low sensitivity).
### F. Drift Guards
- Verify config/doc/instruction contracts have deterministic guards and are current.
- Ensure settings/docs synchronization checks remain active.
## Evidence Standards
- Every finding must include concrete file paths and line ranges.
- No speculative claims.
- Distinguish clearly between:
- **Confirmed ineffective tests**
- **Likely weak tests (needs mutation/probe confirmation)**
## Output Format
Produce a Markdown report at `docs/reviews/<YYYY-MM-DD>-test-effectiveness.md`, using today's date.
Like code review reports, it is a dated, non-canonical artifact: `docs/reviews/**` is not part of the
canonical authority set.
```markdown
# Test Effectiveness Audit Report
## 1. Executive Verdict
- Effective / Effective with Conditions / Needs Remediation
- Top risks to confidence
## 2. Contract Coverage Matrix
| Contract | Guarding Tests | Signal Quality | Gap | Action |
| :--- | :--- | :--- | :--- | :--- |
## 3. Weak/Redundant Test Findings
| Finding ID | Location | Why Low-Signal | Risk | Recommendation |
| :--- | :--- | :--- | :--- | :--- |
## 4. Prune/Strengthen Backlog
| Task ID | Goal | Files | Acceptance Criteria | Validation |
| :--- | :--- | :--- | :--- | :--- |
## 5. Confidence Recommendation
- Go / Go with Conditions / No-Go for release confidence
```
## Decision Rules
- Do not recommend deleting a test unless equivalent or stronger coverage is identified.
- Prefer strengthening assertions before adding more tests.
- Prioritize deterministic contract guards over broad snapshot-style tests.
+9 -4
View File
@@ -1,7 +1,7 @@
name: Quality Gate name: Quality Gate
# V4.7 Phase 6 / review log [40]. Before this, ruff, ty and pytest were enforced # Repository quality gate. Before this workflow existed, ruff, ty, and pytest were
# only by .pre-commit-config.yaml, and only for developers who had actually run # enforced only by .pre-commit-config.yaml for developers who had run
# `pre-commit install`. # `pre-commit install`.
on: on:
@@ -27,13 +27,13 @@ jobs:
- name: Write placeholder configuration - name: Write placeholder configuration
# Settings requires openrouter_api_key and 115 tests cannot construct # Settings requires openrouter_api_key and 115 tests cannot construct
# Settings without it. This is written to a .env file rather than exported # Settings without it. This is written to .env.production rather than exported
# as an environment variable on purpose: the external tests guard on # as an environment variable on purpose: the external tests guard on
# os.getenv("OPENROUTER_API_KEY"), which reads the process environment and # os.getenv("OPENROUTER_API_KEY"), which reads the process environment and
# not the file, so writing the file reproduces the local result exactly - # not the file, so writing the file reproduces the local result exactly -
# the 4 external tests skip instead of running against a fake key and # the 4 external tests skip instead of running against a fake key and
# failing. Exporting it instead produces 3 failures. # failing. Exporting it instead produces 3 failures.
run: echo "OPENROUTER_API_KEY=ci-placeholder-not-a-real-key" > .env run: echo "OPENROUTER_API_KEY=ci-placeholder-not-a-real-key" > .env.production
- name: Lint and type check - name: Lint and type check
# Runs the hooks defined in .pre-commit-config.yaml instead of repeating # Runs the hooks defined in .pre-commit-config.yaml instead of repeating
@@ -42,4 +42,9 @@ jobs:
run: uv run pre-commit run --all-files --show-diff-on-failure run: uv run pre-commit run --all-files --show-diff-on-failure
- name: Tests - name: Tests
# Deliberately unfiltered, unlike the "-m 'not external'" form the guidance files
# use for local runs. Tests marked "external" skip themselves when live-service
# credentials are absent, so CI gets the same effective set plus a real run of any
# external test whose credentials are configured. Not drift -- do not "fix" this to
# match the local command without also giving those tests a way to run.
run: uv run pytest run: uv run pytest
+16 -4
View File
@@ -11,13 +11,25 @@ wheels/
# Environment secrets # Environment secrets
.env .env
.env.production
# SQLite database # SQLite database
*.db *.db
# Document images # All data including db, backups, document images, photos, and logs:
uploads/*
data/* data/*
data/backups/*
data/documents/*
data/logs/*
data/photos/*
data-local/*
# Local destructive-test backups
.test-backups/ # Migration tests
data-migration-test/*
.migration-bundle/*
# Cloudflare tunnel local runtime files
deploy/cloudflared/config.yml
deploy/cloudflared/config.yaml
deploy/cloudflared/credentials.json
+17 -6
View File
@@ -1,18 +1,29 @@
# Quality gate for V4.6 [HIGH-06]. Both hooks are blocking: a regression in # Quality gate: `ruff check`, `ruff format --check`, and `ty check`
# `ruff check` or `ty check` fails the commit. # are blocking once known `ty` false positives are suppressed inline with rationale.
#
# Both tools are uv-managed dev dependencies and are not on PATH, so each entry must
# go through `uv run`.
repos: repos:
- repo: local - repo: local
hooks: hooks:
- id: ruff - id: ruff
name: ruff check name: ruff check
entry: ruff check entry: uv run ruff check
language: system language: system
types_or: [python, pyi] types_or: [python, pyi]
require_serial: true require_serial: true
- id: ty - id: ruff-format
name: ty check name: ruff format check
entry: ty check entry: uv run ruff format --check .
language: system language: system
types_or: [python, pyi] types_or: [python, pyi]
pass_filenames: false pass_filenames: false
require_serial: true require_serial: true
- id: ty
name: ty check
entry: uv run ty check
language: system
types_or: [python, pyi]
pass_filenames: false
require_serial: true
verbose: true
+117
View File
@@ -0,0 +1,117 @@
# AGENTS.md
Orientation for AI agents working in this repository. This file is a **router**, not a spec:
it points at canonical authority and flags the traps that are expensive to discover by trial.
Where this file and `docs/*` disagree, `docs/*` wins.
## What This Is
A document transcription system that preserves durable archival records (Documents, Sources,
People) and executes page transcription asynchronously through vision/LLM providers. Its
defining constraint is **evidence**: every machine attempt is recorded append-only with
request/response provenance. Features that would lose, mutate, or obscure that history are
wrong regardless of how convenient they are.
Stack: Python 3.12+ · FastAPI + NiceGUI · SQLModel/SQLAlchemy (SQLite-first, PostgreSQL-
compatible) · Pydantic V2 · asyncio worker · OpenRouter adapter.
## Commands
This is a `uv` project. **Nothing is on `PATH`**`ruff`, `ty`, and `pytest` all require
`uv run`. Bare invocations fail with command-not-found.
```bash
uv run ruff check . # lint (blocking in pre-commit)
uv run ruff format --check . # format (blocking in pre-commit)
uv run ty check # types (blocking in pre-commit)
uv run pytest -q -m "not external" # default verification run
```
`external` marks tests that hit live services; always exclude it unless explicitly asked.
All four commands are expected to pass clean — there is no tolerated baseline of failures.
If `ty` reports something, fix it or suppress it inline *with a rationale comment*; a bare
`ignore` will not survive review.
## Authority Order
Resolve every question in this order, and stop at the first that answers it:
1. **`docs/*`** — canonical. Start at [`docs/index.md`](docs/index.md), which defines the
reading order. `docs/invariant/*` holds cross-version rules that outlive any release.
2. **`.github/instructions/*.md`** — active steering, auto-attached when you edit matching
paths. Covers services, UI, providers, tests, error handling, and documentation sync.
3. **`.github/skills/*`** — periodic audit procedures (code review, provenance, test
effectiveness).
4. **`tests/`** — deterministic enforcement. A guard test is the ground truth for whatever
rule it encodes.
`.github/agents/` and `.github/prompts/` hold named workflows that are loaded only when
invoked explicitly, so they never override the order above. They are how a review or audit
is *started*, not a source of rules.
`docs/reviews/**` is **not** canonical. Those are dated, opinionated snapshots that were
accurate when written and may since have been fixed, superseded, or found wrong.
## Layout
| Path | Role |
| :--- | :--- |
| `src/transcription/ui/**`, `api/**` | Interface. No direct persistence access. |
| `src/transcription/services/**` | Domain logic and transaction ownership. |
| `src/transcription/db/**` | Models and persistence. |
| `src/transcription/providers/**` | Provider adapters; provider details stop here. |
| `src/transcription/worker.py` | Asyncio worker loop. |
| `tests/` | Includes boundary/contract guards, not just behavior tests. |
## Enforced Boundaries
These are not conventions — a test fails if you break them:
- **No service-to-service imports** (`test_service_boundaries.py`). Compose in the caller.
- **No persistence access from pages/components** (`test_ui_boundaries.py`, allowlist-based).
- **No hand-rolled error notifications in UI** — use the shared error presenter.
- **No stringly-typed status literals** — use the enums (`test_model_contract_guards.py`).
- **Attempt history is append-only** (`test_evidence_provenance.py`).
- **`docs/schema.md` stays field-accurate** with `db/models.py`.
- **Orphans are tracked, not tolerated** — `test_orphan_sweep.py` records each retained
orphan with rationale in `KNOWN_ORPHANS`.
## Traps
Non-obvious things that have already caused real bugs here:
- **`AppError.message` vs `AppError.detail`.** `message` is user/API-facing and must stay
generic — never put exception text or filesystem paths in it. `detail` is internal-only and
is what reaches logs and `ExecutionAttempt.error_detail`. Putting root-cause data in
`message` leaks; removing it from `detail` silently degrades provenance. See
`docs/error_handling.md`.
- **Two competing atomicity invariants in `services/workflows.py`.** Intermediate pages must
commit individually (durability across a long multi-page job); the *final* page must commit
atomically with the terminal job status. Collapsing the batch into one transaction satisfies
the second and destroys the first. Both are guarded — `test_workflows_reliability.py` and
`tests/integration/test_pipeline_atomicity.py`.
- **Shared symbols have more consumers than the obvious one.** Before changing a model field,
exception attribute, or helper return value, grep for every consumer including tests.
Evidence and logging paths frequently read the same fields the UI does.
- **`Tag` is owned by two services, on purpose.** Every other model has exactly one owning
service, so the ownership rule reads as absolute — it isn't. `Tag` is a single table reached
through two `RegistryService[Tag]` facades, `TagRegistry` (documents) and `PersonTagRegistry`
(people), which count usage through `DocumentTag` and `PersonTag` respectively. Changing tag
semantics through one facade silently changes the other. Consolidating them under one service
is not a cleanup; it makes the other side a cross-aggregate writer.
- **Import style:** ruff `isort` runs with `force-single-line = true`. One import per line.
- **Latent defects have ordering constraints.** Some code is unreachable only because of a
current setting or single-instance deployment. Fix it *before* the change that unblocks it,
not after.
## Change Protocol
- **Write the failing test first** for behavioral fixes, and confirm it actually fails for the
reason you think. Several bugs here were subtle enough that a test written afterward would
have passed against the broken code.
- **Update docs in the same change** when you alter a contract, behavior, or scope — see
`.github/instructions/documentation-sync.instructions.md`.
- **Do not commit unless asked.** Making a requested change is not consent to commit it.
- **Do not push or open PRs on your own initiative.**
- **Scope discipline:** fix what was asked plus what your change genuinely breaks. Pre-existing
unrelated issues are a separate conversation.
+8
View File
@@ -13,6 +13,8 @@ RUN uv sync --frozen --no-dev --no-install-project
COPY src ./src COPY src ./src
COPY prompts ./prompts COPY prompts ./prompts
COPY tools ./tools
COPY deploy ./deploy
RUN uv sync --frozen --no-dev RUN uv sync --frozen --no-dev
@@ -27,12 +29,18 @@ ENV PYTHONDONTWRITEBYTECODE=1 \
WORKDIR /app WORKDIR /app
RUN apt-get update \
&& apt-get install -y --no-install-recommends postgresql-client \
&& rm -rf /var/lib/apt/lists/*
RUN groupadd --system --gid 1001 appgroup \ RUN groupadd --system --gid 1001 appgroup \
&& useradd --system --uid 1001 --gid appgroup --create-home appuser && useradd --system --uid 1001 --gid appgroup --create-home appuser
COPY --from=builder /app/.venv /app/.venv COPY --from=builder /app/.venv /app/.venv
COPY --from=builder /app/src /app/src COPY --from=builder /app/src /app/src
COPY --from=builder /app/prompts /app/prompts COPY --from=builder /app/prompts /app/prompts
COPY --from=builder /app/tools /app/tools
COPY --from=builder /app/deploy /app/deploy
RUN mkdir -p /app/uploads /app/data \ RUN mkdir -p /app/uploads /app/data \
&& chown -R appuser:appgroup /app && chown -R appuser:appgroup /app
+43 -6
View File
@@ -22,13 +22,13 @@ uv sync
### 2) Configure environment ### 2) Configure environment
Create a `.env` file in the project root with the required OpenRouter API key: Create a `.env.production` file in the project root with the required OpenRouter API key:
```env ```env
OPENROUTER_API_KEY=your_openrouter_api_key OPENROUTER_API_KEY=your_openrouter_api_key
``` ```
Settings are read from CLI arguments first, then environment variables, then `.env`, then the defaults below. Settings are read from CLI arguments first, then environment variables, then `.env.production`, then the defaults below.
### Configuration Source Precedence ### Configuration Source Precedence
@@ -37,13 +37,13 @@ When the same setting is provided in multiple places, the value is chosen in thi
1. CLI arguments (for example `--port 8000`) 1. CLI arguments (for example `--port 8000`)
2. Settings constructor arguments (used mainly in tests) 2. Settings constructor arguments (used mainly in tests)
3. Environment variables 3. Environment variables
4. `.env` file values 4. `.env.production` file values
5. Model defaults in code 5. Model defaults in code
Practical examples: Practical examples:
- `--port 8000` overrides both `PORT=8000` in the shell and `PORT=7000` in `.env`. - `--port 8000` overrides both `PORT=8000` in the shell and `PORT=7000` in `.env.production`.
- `DATABASE__PATH=prod.db` in the shell overrides `DATABASE__PATH=dev.db` in `.env`. - `DATABASE__PATH=prod.db` in the shell overrides `DATABASE__PATH=dev.db` in `.env.production`.
#### Server and runtime #### Server and runtime
@@ -54,6 +54,7 @@ Practical examples:
| `LOG_LEVEL` | `info` | Uvicorn and application log level. | | `LOG_LEVEL` | `info` | Uvicorn and application log level. |
| `RELOAD` | `false` | Restart the development server when source files change. | | `RELOAD` | `false` | Restart the development server when source files change. |
| `ENVIRONMENT` | `development` | Runtime environment: `development`, `test`, or `production`. | | `ENVIRONMENT` | `development` | Runtime environment: `development`, `test`, or `production`. |
| `RUN_EMBEDDED_WORKER` | `true` | Run worker loop inside web app process. Set `false` when using a dedicated worker service. |
#### Provider #### Provider
@@ -92,7 +93,7 @@ DATABASE__USER=postgres
DATABASE__PASSWORD=change-me DATABASE__PASSWORD=change-me
``` ```
This uses Pydantic nested settings (`env_nested_delimiter='__'`) and avoids JSON blobs in `.env`. A top-level `DATABASE={...}` JSON value is still supported as a fallback, and nested keys such as `DATABASE__PATH` take precedence over conflicting JSON keys. This uses Pydantic nested settings (`env_nested_delimiter='__'`) and avoids JSON blobs in env files. A top-level `DATABASE={...}` JSON value is still supported as a fallback, and nested keys such as `DATABASE__PATH` take precedence over conflicting JSON keys.
`BOOTSTRAP_SCHEMA_ON_STARTUP` creates missing tables when the app starts. When unset, it is enabled in `development` and `test`, and disabled in `production`; set it explicitly to override that policy. `SQLITE_CHECK_SAME_THREAD` defaults to `false`. `BOOTSTRAP_SCHEMA_ON_STARTUP` creates missing tables when the app starts. When unset, it is enabled in `development` and `test`, and disabled in `production`; set it explicitly to override that policy. `SQLITE_CHECK_SAME_THREAD` defaults to `false`.
@@ -122,6 +123,33 @@ This starts the development server with SQLite, creates missing tables, and enab
Replace `localhost` with the server's hostname or IP address when connecting from another machine. Replace `localhost` with the server's hostname or IP address when connecting from another machine.
## Production stack (Phase 1)
Use the production compose profile for split app/worker deployment with PostgreSQL and Cloudflare Tunnel:
```bash
copy .env.production.example .env.production
docker compose -f docker-compose.production.yml up -d --build
```
Services:
- `app`: FastAPI + NiceGUI runtime (`RUN_EMBEDDED_WORKER=false`)
- `worker`: standalone queue processor (`python -m transcription.worker_service`)
- `postgres`: primary datastore
- `cloudflared`: tunnel client using mounted ingress config + `CLOUDFLARE_TUNNEL_TOKEN`
Operational defaults in the production compose file:
- worker healthcheck is disabled (the worker process has no HTTP `/healthz` endpoint)
- cloudflared is pinned to HTTP/2 with explicit DNS resolvers (`1.1.1.1`, `1.0.0.1`) for restricted LXC/container DNS environments
Cloudflare setup files:
1. `copy deploy\cloudflared\config.yml.example deploy\cloudflared\config.yml`
2. set `CLOUDFLARE_TUNNEL_TOKEN` in `.env.production`
3. update ingress hostnames in `deploy\cloudflared\config.yml`
## How to navigate the GUI ## How to navigate the GUI
- **Upload page** (`/ui`) - **Upload page** (`/ui`)
@@ -147,6 +175,15 @@ not a path. Each job snapshots the validated prompt text, SHA-256 hash, and samp
The canonical MVP prompt is: The canonical MVP prompt is:
- `prompts/transcribe_document.md` - `prompts/transcribe_document.md`
## Database migration workflow
Schema upgrades use an explicit export/import rebuild flow (no runtime legacy write compatibility).
See `docs/data_migration.md` for commands and cutover steps.
## Backup and restore workflow
Production backup/restore (PostgreSQL + uploads + deployment config) steps are documented in `docs/backup_restore.md`.
## Destructive test procedure (with data backup) ## Destructive test procedure (with data backup)
AI execution policy: before the first unit-test run in a test/fix cycle, create one backup of `./data`. Reuse that same backup for every subsequent test run in the cycle. After tests succeed, always pause and ask whether to restore now. AI execution policy: before the first unit-test run in a test/fix cycle, create one backup of `./data`. Reuse that same backup for every subsequent test run in the cycle. After tests succeed, always pause and ask whether to restore now.
+69
View File
@@ -0,0 +1,69 @@
#!/usr/bin/env sh
set -eu
BACKUP_DIR="${BACKUP_DIR:-./data/backups}"
RETENTION_DAYS="${BACKUP_RETENTION_DAYS:-14}"
UPLOAD_DIR="${UPLOAD_DIR:-/app/uploads}"
PROMPT_DIR="${PROMPT_DIR:-/app/prompts}"
DATABASE_DRIVER="${DATABASE__DRIVER:-postgres}"
DATABASE_HOST="${DATABASE__HOST:-postgres}"
DATABASE_PORT="${DATABASE__PORT:-5432}"
DATABASE_NAME="${DATABASE__DATABASE:-}"
DATABASE_USER="${DATABASE__USER:-}"
DATABASE_PASSWORD="${DATABASE__PASSWORD:-}"
timestamp="$(date -u +%Y%m%d-%H%M%S)"
postgres_file="postgres-${timestamp}.dump"
manifest_file="backup-${timestamp}.manifest"
mkdir -p "${BACKUP_DIR}"
if [ "${DATABASE_DRIVER}" != "postgres" ]; then
echo "create_postgres_backup.sh requires DATABASE__DRIVER=postgres." >&2
exit 1
fi
if [ -z "${DATABASE_NAME}" ] || [ -z "${DATABASE_USER}" ] || [ -z "${DATABASE_PASSWORD}" ]; then
echo "DATABASE__DATABASE, DATABASE__USER, and DATABASE__PASSWORD must be set." >&2
exit 1
fi
if ! command -v pg_dump >/dev/null 2>&1; then
echo "pg_dump is not installed in this environment." >&2
exit 1
fi
PGPASSWORD="${DATABASE_PASSWORD}" pg_dump \
-h "${DATABASE_HOST}" \
-p "${DATABASE_PORT}" \
-U "${DATABASE_USER}" \
-d "${DATABASE_NAME}" \
-Fc \
> "${BACKUP_DIR}/${postgres_file}"
cat > "${BACKUP_DIR}/${manifest_file}" <<EOF
created_at_utc=${timestamp}
postgres_dump=${postgres_file}
backup_dir=${BACKUP_DIR}
uploads_backup_dir=${BACKUP_DIR}/uploads
prompts_backup_dir=${BACKUP_DIR}/prompts
EOF
find "${BACKUP_DIR}" -type f \( \
-name 'postgres-*.dump' -o \
-name 'backup-*.manifest' \
\) -mtime +"${RETENTION_DAYS}" -delete
uploads_backup_dir="${BACKUP_DIR}/uploads"
mkdir -p "${uploads_backup_dir}"
if [ -d "${UPLOAD_DIR}" ]; then
cp -an "${UPLOAD_DIR}/." "${uploads_backup_dir}/"
fi
prompts_backup_dir="${BACKUP_DIR}/prompts"
mkdir -p "${prompts_backup_dir}"
if [ -d "${PROMPT_DIR}" ]; then
cp -a "${PROMPT_DIR}/." "${prompts_backup_dir}/"
fi
echo "Created backup set:"
echo " ${BACKUP_DIR}/${postgres_file}"
echo " ${BACKUP_DIR}/${manifest_file}"
echo " ${uploads_backup_dir}/ (incremental uploads mirror)"
echo " ${prompts_backup_dir}/ (prompts mirror)"
@@ -0,0 +1,36 @@
#!/usr/bin/env sh
set -eu
# Example only. Copy to a local script and replace placeholder values.
# Do NOT commit secrets.
SHARE="//nas-host-or-ip/share-name"
MOUNT_POINT="/mnt/nas-backups"
CREDENTIALS_FILE="/etc/samba/credentials/nas-share-credentials"
USERNAME="replace-with-nas-user"
PASSWORD="replace-with-nas-password"
if [ "${USERNAME}" = "replace-with-nas-user" ] || [ "${PASSWORD}" = "replace-with-nas-password" ]; then
echo "Edit USERNAME and PASSWORD placeholders before running this script."
exit 1
fi
mkdir -p "${MOUNT_POINT}"
mkdir -p "$(dirname "${CREDENTIALS_FILE}")"
cat > "${CREDENTIALS_FILE}" <<'EOF'
username=__USERNAME__
password=__PASSWORD__
EOF
sed -i "s|__USERNAME__|${USERNAME}|g" "${CREDENTIALS_FILE}"
sed -i "s|__PASSWORD__|${PASSWORD}|g" "${CREDENTIALS_FILE}"
chmod 600 "${CREDENTIALS_FILE}"
mount -t cifs "${SHARE}" "${MOUNT_POINT}" \
-o "credentials=${CREDENTIALS_FILE},vers=3.0,iocharset=utf8,uid=0,gid=0,file_mode=0600,dir_mode=0700"
echo ""
echo "Mounted ${SHARE} at ${MOUNT_POINT}"
echo ""
echo "To persist across reboot, add this line to /etc/fstab:"
echo "${SHARE} ${MOUNT_POINT} cifs credentials=${CREDENTIALS_FILE},vers=3.0,iocharset=utf8,uid=0,gid=0,file_mode=0600,dir_mode=0700,_netdev,nofail,x-systemd.automount 0 0"
+86
View File
@@ -0,0 +1,86 @@
#!/usr/bin/env sh
set -eu
if [ "$#" -lt 1 ]; then
echo "Usage: $0 <path-to-postgres-dump>"
exit 1
fi
dump_file="$1"
COMPOSE_FILE="${COMPOSE_FILE:-docker-compose.production.yml}"
ENV_FILE="${ENV_FILE:-.env.production}"
SYNOLOGY_BACKUP_DIR="${SYNOLOGY_BACKUP_DIR:-}"
if [ ! -f "${dump_file}" ]; then
echo "Backup file not found: ${dump_file}"
exit 1
fi
backup_dir="$(dirname "${dump_file}")"
backup_name="$(basename "${dump_file}")"
timestamp="$(printf '%s' "${backup_name}" | sed -n 's/^postgres-\([0-9]\{8\}-[0-9]\{6\}\)\.dump$/\1/p')"
uploads_file=""
config_file=""
# If SYNOLOGY_BACKUP_DIR wasn't exported in the shell, read it from ENV_FILE.
if [ -z "${SYNOLOGY_BACKUP_DIR}" ] && [ -f "${ENV_FILE}" ]; then
SYNOLOGY_BACKUP_DIR="$(
sed -n 's/^SYNOLOGY_BACKUP_DIR=//p' "${ENV_FILE}" | tail -n 1
)"
fi
if [ -n "${timestamp}" ]; then
# Legacy local full-archive naming.
candidate_uploads="${backup_dir}/uploads-${timestamp}.tar.gz"
candidate_config="${backup_dir}/config-${timestamp}.tar.gz"
if [ -f "${candidate_uploads}" ]; then
uploads_file="${candidate_uploads}"
fi
if [ -f "${candidate_config}" ]; then
config_file="${candidate_config}"
fi
fi
docker compose --env-file "${ENV_FILE}" -f "${COMPOSE_FILE}" exec -T postgres sh -lc \
"PGPASSWORD=\"\$POSTGRES_PASSWORD\" psql -U \"\$POSTGRES_USER\" -d postgres -c \"DROP DATABASE IF EXISTS \\\"\$POSTGRES_DB\\\";\""
docker compose --env-file "${ENV_FILE}" -f "${COMPOSE_FILE}" exec -T postgres sh -lc \
"PGPASSWORD=\"\$POSTGRES_PASSWORD\" psql -U \"\$POSTGRES_USER\" -d postgres -c \"CREATE DATABASE \\\"\$POSTGRES_DB\\\";\""
cat "${dump_file}" | docker compose --env-file "${ENV_FILE}" -f "${COMPOSE_FILE}" exec -T postgres sh -lc \
"PGPASSWORD=\"\$POSTGRES_PASSWORD\" pg_restore -U \"\$POSTGRES_USER\" -d \"\$POSTGRES_DB\" --clean --if-exists --no-owner --no-privileges"
if [ -n "${SYNOLOGY_BACKUP_DIR}" ] && [ -d "${SYNOLOGY_BACKUP_DIR}/uploads" ]; then
docker compose --env-file "${ENV_FILE}" -f "${COMPOSE_FILE}" run --rm --no-deps \
-v "${SYNOLOGY_BACKUP_DIR}:/backup" \
--entrypoint sh app -lc \
"mkdir -p /app/uploads/documents /app/uploads/photos && \
find /app/uploads/documents -mindepth 1 -delete && \
find /app/uploads/photos -mindepth 1 -delete && \
if [ -d /backup/uploads/documents ]; then cp -a /backup/uploads/documents/. /app/uploads/documents/; fi && \
if [ -d /backup/uploads/photos ]; then cp -a /backup/uploads/photos/. /app/uploads/photos/; fi && \
if [ -f /backup/uploads/homepage.md ]; then cp /backup/uploads/homepage.md /app/uploads/homepage.md; else rm -f /app/uploads/homepage.md; fi"
elif [ -n "${uploads_file}" ]; then
# Legacy local full-archive restore.
cat "${uploads_file}" | docker compose --env-file "${ENV_FILE}" -f "${COMPOSE_FILE}" run --rm --no-deps --entrypoint sh app -lc \
"mkdir -p /app/uploads && find /app/uploads -mindepth 1 -delete && tar -xzf - -C /app/uploads"
fi
if [ -n "${SYNOLOGY_BACKUP_DIR}" ] && [ -n "${timestamp}" ] && [ -f "${SYNOLOGY_BACKUP_DIR}/config-${timestamp}.tar.gz" ]; then
tar -xzf "${SYNOLOGY_BACKUP_DIR}/config-${timestamp}.tar.gz" -C .
elif [ -n "${config_file}" ]; then
# Legacy local full-archive restore.
tar -xzf "${config_file}" -C .
fi
echo "Restore complete from: ${dump_file}"
if [ -n "${SYNOLOGY_BACKUP_DIR}" ] && [ -d "${SYNOLOGY_BACKUP_DIR}/uploads" ]; then
echo "Restored uploads mirror from: ${SYNOLOGY_BACKUP_DIR}/uploads"
elif [ -n "${uploads_file}" ]; then
echo "Restored uploads archive: ${uploads_file}"
fi
if [ -n "${SYNOLOGY_BACKUP_DIR}" ] && [ -n "${timestamp}" ] && [ -f "${SYNOLOGY_BACKUP_DIR}/config-${timestamp}.tar.gz" ]; then
echo "Restored config archive: ${SYNOLOGY_BACKUP_DIR}/config-${timestamp}.tar.gz"
elif [ -n "${config_file}" ]; then
echo "Restored config archive: ${config_file}"
fi
+14
View File
@@ -0,0 +1,14 @@
ingress:
# Primary transcription app endpoint.
- hostname: transcription.example.com
service: http://app:8000
# Optional: generic remote access endpoints for other internal services.
# Replace hostnames and targets for your LAN.
- hostname: homeassistant.example.com
service: http://192.168.1.50:8123
- hostname: pihole.example.com
service: http://192.168.1.60:80
# Required catch-all.
- service: http_status:404
+70
View File
@@ -0,0 +1,70 @@
services:
app:
build:
context: .
dockerfile: Dockerfile
image: transcription:prod
env_file:
- .env.production
environment:
RUN_EMBEDDED_WORKER: "false"
RUNTIME_SETTINGS_ENV_FILE: "/app/.env.production"
depends_on:
postgres:
condition: service_healthy
ports:
- "8000:8000"
volumes:
- app_uploads:/app/uploads
- app_data:/app/data
- ./backup:/backup
- ./.env.production:/app/.env.production
- ./prompts:/app/prompts:ro
restart: unless-stopped
healthcheck:
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/healthz')"]
interval: 30s
timeout: 5s
retries: 5
start_period: 20s
worker:
image: transcription:prod
env_file:
- .env.production
command: ["python", "-m", "transcription.worker_service"]
depends_on:
postgres:
condition: service_healthy
volumes:
- app_uploads:/app/uploads
- app_data:/app/data
- ./backup:/backup
- ./.env.production:/app/.env.production:ro
- ./prompts:/app/prompts:ro
healthcheck:
disable: true
restart: unless-stopped
postgres:
image: postgres:16-alpine
env_file:
- .env.production
environment:
POSTGRES_DB: ${DATABASE__DATABASE}
POSTGRES_USER: ${DATABASE__USER}
POSTGRES_PASSWORD: ${DATABASE__PASSWORD}
volumes:
- postgres_data:/var/lib/postgresql/data
restart: unless-stopped
healthcheck:
test: ["CMD-SHELL", "pg_isready -U $$POSTGRES_USER -d $$POSTGRES_DB"]
interval: 10s
timeout: 5s
retries: 10
start_period: 10s
volumes:
postgres_data:
app_uploads:
app_data:
+1 -1
View File
@@ -5,7 +5,7 @@ services:
dockerfile: Dockerfile dockerfile: Dockerfile
container_name: transcription-app container_name: transcription-app
env_file: env_file:
- .env - .env.production
environment: environment:
# Database configuration uses nested settings names (env_nested_delimiter="__"). # Database configuration uses nested settings names (env_nested_delimiter="__").
# DATABASE_URL is NOT read by the application and must not be used here. # DATABASE_URL is NOT read by the application and must not be used here.
+230
View File
@@ -0,0 +1,230 @@
# System Architecture (Current Baseline: V6.1)
This document defines the current V6.1 architecture baseline.
## Architecture Objectives
- Preserve durable archival records for Documents, Sources, People, and processing runs.
- Execute page transcription asynchronously with bounded worker behavior.
- Preserve append-only machine-attempt evidence with request/response provenance.
- Keep UI, API, service, persistence, and provider boundaries explicit and testable.
## Technical Stack
- **Runtime:** Python 3.12+
- **Web application:** FastAPI + NiceGUI
- **Persistence:** SQLModel / SQLAlchemy — PostgreSQL in production, SQLite for local development and tests
- **Validation and settings:** Pydantic V2 + pydantic-settings
- **Concurrency:** asyncio worker loop
- **Provider integration:** OpenRouter adapter behind provider interface
- **Deployment:** Docker Compose (app, worker, PostgreSQL, Cloudflare Tunnel)
- **Quality and tests:** Ruff, ty, pytest, pytest-asyncio
## Runtime Topology
```mermaid
flowchart LR
U[Browser User] --> A[FastAPI + NiceGUI App]
A --> W[Asyncio Worker]
A --> DB[(PostgreSQL / SQLite)]
W --> P[Provider Adapter]
W --> DB
```
The worker loop drains two queues in the same pass: queued transcription Jobs and queued
`MaintenanceRun` records. When neither has work, it idles.
### Production deployment
Production runs as a Docker Compose stack with the app and worker as separate services, so the app
process runs with `RUN_EMBEDDED_WORKER=false` and the worker process owns queue draining. Local
development runs a single process with the worker embedded.
```mermaid
flowchart LR
I[Internet] --> CF[cloudflared tunnel + Access]
CF --> APP[app service]
APP --> PG[(postgres service)]
WK[worker service] --> PG
WK --> PROV[OpenRouter]
```
Deployment, rollback, and recovery procedures are in [Production Runbook](production-runbook.md);
backup configuration and restore are in [Backup and Restore](backup_restore.md).
## Layered Boundaries
### Interface Layer
- `src/transcription/ui/**`
- `src/transcription/api/**`
Responsibilities:
- Route registration, page orchestration, presentation adapters.
- Structured user messaging through shared error presenter.
- No direct persistence access from pages/components.
### Service and Orchestration Layer
Aggregate services:
- `src/transcription/services/documents.py`
- `src/transcription/services/people.py`
- `src/transcription/services/jobs.py`
- `src/transcription/services/sources.py`
- `src/transcription/services/photos.py`
- `src/transcription/services/maintenance.py`
- `src/transcription/services/evidence.py` (read/projection only)
Orchestration modules:
- `src/transcription/services/store.py`
- `src/transcription/services/workflows.py`
Responsibilities:
- Aggregate ownership and invariants.
- Transaction-aware write helpers.
- Cross-service workflows in orchestration modules (`store.py`, `workflows.py`).
- Lookup-table CRUD through the generic `RegistryService` base (`registry.py`), which is not an
aggregate owner itself.
Module classification and per-model ownership are defined in
[services instructions](../.github/instructions/services.instructions.md).
### Persistence Layer
- `src/transcription/db/**`
Responsibilities:
- SQLModel definitions, async session/engine runtime, registry bootstrap.
- Loader helpers that enforce explicit eager loading with `lazy="raise"` relationships.
### Provider Layer
- `src/transcription/providers/**`
Responsibilities:
- Provider API encapsulation.
- Request manifest and transport evidence capture.
- Normalized transcription result contract.
## Core Domain Model
- `Document` owns archival metadata and links to `Source`, `Job`, and `DocumentPerson`.
- `Source` is a document page/file record with selected machine projection and human revision.
- `Job` is an aggregate processing run with status and frozen prompt/runtime settings.
- `JobSource` is queue/membership state for one `(job, source)` pair.
- `ExecutionAttempt` is append-only evidence for each provider call.
- `Photo` is person imagery owned by `PhotosService`.
- `MaintenanceRun` is one queued or executed operational maintenance run.
- `GenealogyPerson`, `GenealogyFamily`, `GenealogyFamilyChild`, and `GenealogyCitation` store
imported GEDCOM genealogy data and citation provenance.
- `DocumentType` and `PersonRole` are UUID-backed registries with optional protected `semantic_key`.
- `Tag` is a shared registry reached through both document and person tagging, linked by
`DocumentTag` and `PersonTag`.
## Processing and Evidence Workflow
1. User creates/updates Document metadata and linked People atomically through workflow orchestration.
2. User creates a Job by uploading one or more Source files or by retranscribing an existing Source.
3. Source files are validated and stored; orientation normalization may be applied at ingest, and stored bytes become the canonical processing bytes.
4. Worker claims queued Job, transitions to `processing`, and processes pending pages in deterministic order.
5. Each provider call writes one immutable `ExecutionAttempt` with:
- request manifest + hash
- transport evidence (when response exists)
- SDK snapshot and normalized metadata
- outcome, timing, and error details when applicable
6. `JobSource` status is updated as queue/projection state; `Source.raw_transcription` is set on first successful attempt and can be explicitly re-pointed by candidate promotion.
7. Job terminal status resolves to `transcribed`, `partial_success`, or `failed`.
## Status Semantics
- **Job statuses:** `queued`, `processing`, `transcribed`, `partial_success`, `failed`
- Operational success path resolves to `transcribed`.
- **JobSource statuses:** `pending`, `transcribed`, `failed`, `cancelled`
- **MaintenanceRun statuses:** `queued`, `processing`, `succeeded`, `failed`
- Maintenance uses `succeeded` rather than `transcribed`; the transcription vocabulary does not
apply to operational runs.
- **Maintenance job types:** `backup`, `storage_reconciliation`, `gedcom_import`
## Maintenance Execution
Operational maintenance is queue-backed rather than run inline from the UI, so it survives request
lifetime and is recorded:
1. Settings enqueues a `MaintenanceRun` with `status=queued` and a `triggered_by` marker.
2. The worker claims the oldest queued run with a conditional update, moving it to `processing`.
3. `backup` runs the deploy backup script; `storage_reconciliation` compares stored media against
`Document`/`Source` records; `gedcom_import` parses the latest uploaded `.ged` file and upserts
genealogy records.
4. The run finalizes to `succeeded` or `failed` with summary, timing, log path, and `error_detail`.
`MaintenanceRun` records operational history and is not evidence in the `ExecutionAttempt` sense;
append-only guarantees apply to transcription attempts.
## Security and Path Handling Boundaries
- Print media delivery uses record-validated API route:
- `src/transcription/api/print_api.py`
- General UI media links resolve through:
- `src/transcription/ui/components/media_urls.py`
- Local filesystem paths must never be accepted from user input as trusted media routes.
## Concurrency and Reliability Principles
- Worker loop reuses service bundle/provider resources for pooled calls.
- Provider-call timeout is explicit and bounded.
- Non-retriable worker-loop faults are surfaced and stop loop spin.
- Per-page outcomes are durably persisted before processing next page.
## Design Decisions and Rationale
### Why `transcribed` is the success terminal state
- The worker and job orchestration resolve successful completion to `JobStatus.TRANSCRIBED`, with mixed and failure outcomes represented by `partial_success` and `failed`.
- This keeps terminal status vocabulary aligned with what the pipeline actually produces: transcribed page content and evidence, not a generic completion marker.
### Why evidence history is append-only while page text is a projection
- `ExecutionAttempt` stores immutable per-call evidence and preserves full attempt history across retries.
- `Source.raw_transcription` is intentionally a mutable projection so UI and exports can show a selected current machine text without mutating historical evidence.
- This split keeps auditability and UX both first-class: history is durable, presentation is editable.
### Why orchestration modules own cross-service workflows
- Service modules do not import each other; aggregate ownership remains local to each service.
- Multi-aggregate writes are coordinated in orchestration modules (`store.py`, `workflows.py`) so transaction boundaries are explicit and testable.
- This avoids circular dependencies and keeps cross-cutting workflow logic centralized.
### Why explicit eager loading is required
- ORM relationships are configured with `lazy="raise"` in key paths, so code must request needed relationships up front.
- This prevents hidden query behavior in UI/service code and makes read shape deterministic and reviewable.
### Why canonical source bytes may be ingest-normalized
- Ingest normalization can correct orientation before persistence so provider calls, evidence hashes, and rendered processing source are consistent.
- The canonical stored bytes, digest, and size become the durable processing identity for that source.
### Why media access uses controlled routes/helpers
- Print/export media uses record-validated API endpoints to avoid direct filesystem path exposure.
- General UI media URLs are generated through shared resolver helpers to keep path handling consistent and centralized.
## Scope Boundary
Current architecture rules live in `docs/*`.
## Related References
- [System Requirements](requirements.md)
- [Data Model](schema.md)
- [Error Handling Policy](error_handling.md)
- [Production Runbook](production-runbook.md)
- [Backup and Restore](backup_restore.md)
- [Error Handling invariant](./invariant/error_handling.md)
- [AI evidence invariant](./invariant/ai_evidence_and_provenance.md)
-547
View File
@@ -1,547 +0,0 @@
# Architecture & Code Review Report
> ## Status: Historical Snapshot - Superseded
>
> **This document describes the codebase as it stood on 2026-08-17, before the V4.6 remediation release.** It is retained because it is the canonical registry of the finding IDs (`CRIT-01`, `HIGH-06`, `MED-14`, and so on) cited throughout the V4.6, V4.7, and V4.8 planning documents. Six documents reference it; deleting it would orphan all 32 finding IDs.
>
> **Do not read it as a description of current state.** Every file path, line number, and metric below is pre-V4.6 and most are now wrong. The baseline figures in particular are stale: `ruff check` is clean, `ty check` reports 0 diagnostics, and the suite is at 292 passed / 4 skipped.
>
> **Disposition of all 32 findings:**
>
> | Status | Findings |
> | :--- | :--- |
> | Addressed in V4.6 | All 32 were dispositioned - fixed, consciously accepted, or explicitly deferred. See [V4.6 scope boundary](ver4.6/scope_boundary_v4_6.md) and [V4.6 implementation plan](ver4.6/implementation_plan_v4_6.md). |
> | Carried into V4.7 | `MED-14` (SourceService decomposition) and `HIGH-06` (CI enforcement of the quality gate). See [V4.7 scope boundary](ver4.7/scope_boundary_v4_7.md). |
>
> **Where later thinking supersedes this report:** the [V4.6 review log](ver4.6/review_log_v4_6.md) records the decisions, deviations, and revisions made during implementation. Where this report and a committed planning document disagree, **the planning document wins**.
>
> Two recommendations here were later revised on evidence. `MED-14` proposed extracting a `services/artifacts.py`; V4.7 cancels that in favour of deleting the `ProcessingArtifact` subsystem outright. The report also treats `job_source` and `execution_attempt` as complementary; measurement showed their evidence columns are fully duplicated.
**Repository Target:** `C:\GitHub\transcription\`
**Target Stack:** Python 3.12+ | FastAPI | NiceGUI | SQLModel/SQLAlchemy | Pydantic V2 | asyncio | OpenRouter
**Review Date:** 2026-08-17
**Baseline verified:** `pytest` → 264 passed, 4 skipped. `ruff check` → 6 errors. `ty check` → 197 diagnostics.
---
## 1. Executive Summary
- **The codebase is disciplined and unusually well-structured for its size (~9k src LOC).** Error taxonomy (`errors.py`), evidence capture (`providers/evidence.py`), transaction-ownership helpers (`ServiceBase._finalize`), and CSS/asset discipline in the UI are genuinely strong and should be preserved.
- **Top risk is the job-claim path.** `JobService.read_next_queued_job` (`services/jobs.py:170-187`) selects *every* `QUEUED` job with three levels of eager loading, has no `LIMIT`, no `FOR UPDATE SKIP LOCKED`, and no compare-and-swap on the status transition. The code even carries a comment acknowledging the race (`services/workflows.py:193-194`) without fixing it.
- **Second risk is model-level eager loading.** Nearly every `Relationship` in `db/models.py` sets `lazy="selectin"` on *both* sides of bidirectional links (`Document.jobs``Job.document`, `Job.job_sources``JobSource.job`, `JobSource.source``Source.job_sources`). Reading one `Job` cascades into loading effectively the whole related graph, and it makes every explicit `selectinload()` in the services redundant.
- **No index exists on `Job.status` or `Job.date_created`**, yet the worker polls `WHERE status='queued' ORDER BY date_created` once per second. Every poll is a full table scan.
- **`worker_provider_timeout_seconds` is hard-capped at `le=20.0`** (`config.py:110`). Vision transcription of a full document page routinely exceeds 20s; this cap makes systematic timeouts unconfigurable-away.
- **The provider HTTP client is destroyed and rebuilt for every single job** (`worker.py:157-174`), defeating connection pooling and TLS session reuse on the hottest path.
- **`app_state.py` is dead code containing a guaranteed `TypeError`** (verified at runtime): `resolve_session_factory` calls `get_session_factory()` with no arguments. `@functools.cache` erases the signature, so `ty` cannot see it.
- **`ty` is configured as a dev dependency but is not usable as a gate.** 197 diagnostics, ~160 of which are SQLModel relationship false positives already suppressed with `# pyright: ignore` comments that `ty` does not honor.
- **Schema evolution is hand-rolled** in `db/operations.py` with raw `ALTER TABLE`/`CREATE INDEX IF NOT EXISTS` and a SQLite-shaped `CHAR(32)` UUID column. There is no Alembic. Postgres portability is claimed but not actually exercised.
- **Meaningful duplication exists in the UI layer** (~500 lines): media-URL resolution, `_parse_uuid`, settings resolution, delete-confirmation scaffolds, and hand-rolled tables are each reimplemented 3-5 times.
- **Meaningful duplication also exists in the service layer** (~400 lines): `DocumentType` and `PersonRole` registry CRUD are structurally identical, 38 "not-found" raises are hand-written, and three media-storage flows are reimplemented.
---
## 1a. Post-Review Addendum
The findings below were established during the V4.6 scoping discussion that followed the original review. They restate severity in light of the project's confirmed operating context and add findings discovered during that discussion. **The original finding IDs are stable and remain the canonical reference for the V4.6 documents.**
### Confirmed Operating Context
| Question | Answer |
| :--- | :--- |
| Database | **SQLite only.** PostgreSQL is the intended destination but is deferred beyond V4.6. `JSONBCompat` is retained. |
| Topology | **Single user, single process** today. Multi-user server is the stated direction. |
| Schema evolution | **Re-level from current metadata.** No Alembic. The app is pre-production and the schema is still moving. |
| Existing data | Rebuilt from scratch during implementation; migrated from backup as the final step. |
| Release character | **Pure remediation.** No new features. |
| Scope band | Critical through Low, inclusive. |
### Severity Re-Grades
| ID | Original | Re-graded | Rationale |
| :--- | :--- | :--- | :--- |
| CRIT-01 | Critical | **High** | With one process and one worker there is no live duplicate-processing race. The missing `.limit(1)` and the eager-load cost remain genuine defects; the atomic claim becomes forward-compatibility work for the multi-user direction rather than an active-incident fix. |
| CRIT-02 | Critical | **Critical** (unchanged) | Read amplification is independent of both topology and dialect. It costs on every read today. |
| HIGH-05 | High | **High** (reframed) | The remedy is **not** Alembic. Because the schema is pre-production and the data is disposable, the correct fix is to delete `upgrade_schema` and the three `_upgrade_*` functions outright and re-level the schema from current SQLModel metadata. This automatically resolves the `CHAR(32)` defect. |
| MED-01 | Medium | **Medium** (low urgency) | Single-user operation means event-loop stalls are self-inflicted only. Remains in scope. |
### Items Added During Scoping
These are recorded as [HIGH-08], [MED-10] through [MED-14], and [LOW-08] below.
---
## 2. Findings by Severity
### Critical Severity
#### [CRIT-01] Queued-job claim has no row lock, no CAS, and no LIMIT — duplicate processing and full-queue load
> **Re-graded to High.** See [§1a](#severity-re-grades). The single-process deployment removes the live duplicate-processing race; the missing `.limit(1)` and the eager-load cost are still real, and the atomic claim is retained as forward-compatibility work.
- **Location:** `src/transcription/services/jobs.py:170-187`; claim logic at `src/transcription/services/workflows.py:188-196`; divergent duplicate at `src/transcription/db/operations.py:143-151`
- **Problem & Consequence:** `read_next_queued_job` issues `SELECT ... WHERE status = 'queued' ORDER BY date_created, id` with `selectinload(Job.document)` and `selectinload(Job.job_sources).selectinload(JobSource.source)` — and **no `.limit(1)`**. It materializes the entire queue plus its document/job_source/source graph on every worker tick just to call `.first()`. With a backlog of N jobs this is O(N) rows and several extra SELECT round-trips per second.
Worse, the claim is a read-then-write with no atomicity: `process_queued_job` reads status `QUEUED`, then separately calls `mark_job_status(job.id, PROCESSING)`. Two workers (or an app replica plus the in-process worker) can both read the same row as `QUEUED` and both transcribe it — double provider spend and duplicate `ExecutionAttempt` evidence rows. The comment at `workflows.py:193-194` explicitly names this hazard ("otherwise other workers may see the job as still QUEUED") but the committed fix only narrows the window rather than closing it.
Note also that `db/operations.py:get_next_queued_job` is a second, *different* implementation of the same concept that *does* have `.limit(1)`. The worse implementation is the live one.
- **Recommendation:** Replace the read-then-write with a single atomic claim, and delete the duplicate.
```python
# Before (services/jobs.py) — no limit, no lock
query = select(Job).options(...).where(Job.status == JobStatus.QUEUED).order_by(Job.date_created, Job.id)
return (await _session.exec(query)).first()
# After — atomic claim, one row, dialect-aware
async def claim_next_queued_job(self, *, session=None) -> Job | None:
async with self._session_scope(session) as s:
stmt = (
select(Job)
.where(Job.status == JobStatus.QUEUED)
.order_by(Job.date_created, Job.id)
.limit(1)
)
if s.bind.dialect.name == "postgresql":
stmt = stmt.with_for_update(skip_locked=True)
job = (await s.exec(stmt)).first()
if job is None:
return None
job.status = JobStatus.PROCESSING
job.date_updated = datetime.now(UTC)
await self._finalize(session=s, caller_session=session, refresh=(job,))
return job
```
Load the eager relationships in a *second* query after the claim succeeds, so the hot poll stays a single narrow row. On SQLite, wrap the claim in `BEGIN IMMEDIATE` or accept single-worker-only and document it.
- **Effort:** M
#### [CRIT-02] Bidirectional `lazy="selectin"` on every relationship causes cascading read amplification
- **Location:** `src/transcription/db/models.py:71-73, 89-91, 108-115, 160-168, 211-212, 269-276, 332-333`
- **Problem & Consequence:** Every `Relationship` in the domain model sets `sa_relationship_kwargs={"lazy": "selectin"}`, including both sides of each pair. Fetching a single `Job` triggers: `Job` → `Job.document` → `Document.jobs` (all jobs for that document) → `Document.sources` → `Document.document_people` → `DocumentPerson.person` / `.role_ref` → each `Job.job_sources` → `JobSource.source` → `Source.job_sources` → … SQLAlchemy's identity map prevents infinite recursion but does **not** prevent the extra SELECT round trips per level.
Two concrete consequences: (a) the per-second worker poll is far more expensive than it appears from reading `jobs.py`; (b) the dozens of explicit `selectinload(...)` options in `documents.py`, `jobs.py`, `sources.py`, and `people.py` are dead weight — the relationship default already does it — and they are the source of ~160 of the 197 `ty` diagnostics.
- **Recommendation:** Flip the model default to `lazy="raise"` (or `"noload"`, as already correctly done for `Source.processing_artifacts` at `models.py:279` and `JobSource.execution_attempts` at `models.py:336`) and rely on the per-query `selectinload()` that services already declare. `lazy="raise"` converts silent N+1 into a loud test failure and would prove which eager loads are actually needed.
```python
# models.py
jobs: list["Job"] = Relationship(back_populates="document", sa_relationship_kwargs={"lazy": "raise"})
```
Roll out per-model with the existing test suite as the safety net; the suite already covers the read paths.
- **Effort:** M
### High Severity
#### [HIGH-01] `app_state.py` is unreferenced dead code containing a guaranteed `TypeError`
- **Location:** `src/transcription/app_state.py:29-34`; the called function at `src/transcription/db/session.py:20-26`
- **Problem & Consequence:** `resolve_session_factory` falls back to `get_session_factory()` with no arguments, but the signature is `get_session_factory(database_url: str)`. Verified at runtime:
```
TypeError: get_session_factory() missing 1 required positional argument: 'database_url'
```
`@functools.cache` wraps the function in a `_lru_cache_wrapper`, which erases the signature — so `ty check src\transcription\app_state.py` reports "All checks passed". The whole module has **zero importers** anywhere in `src`, `tests`, or `tools`, so the bug is currently latent; anyone wiring this helper up hits an immediate crash on the fallback path.
- **Recommendation:** Delete `app_state.py`. Its three live behaviors already exist elsewhere (`db/session.py:resolve_session_factory`, `db/runtime.py:get_database_runtime`, `worker.py:resolve_worker_notifier`). If retained instead, fix the fallback to `resolve_session_factory()` from `db.session`, and add a typed non-cached wrapper around cached functions so type checkers keep the signature.
- **Effort:** S
#### [HIGH-02] Provider HTTP client is rebuilt and torn down once per job
- **Location:** `src/transcription/worker.py:148-174` (`finally: await services.sources.aclose()`), driven by the tight inner loop at `src/transcription/worker.py:134-142`; client construction at `src/transcription/providers/openrouter.py:197-201`
- **Problem & Consequence:** `process_next_queued_job` constructs a fresh `ServiceBundle` per call and unconditionally closes the provider in `finally`. Since `workflows.py:243` accesses `services.sources.provider`, a new `httpx.AsyncClient` + `OpenRouter` SDK client is created and destroyed for **every job**. This throws away the connection pool and forces a full TLS handshake per job — added latency on the single most latency-sensitive path, plus churn of file descriptors during backlog drain.
- **Recommendation:** Hoist the `ServiceBundle` to worker-loop scope (or reuse `app.state.services`, which the lifespan already builds at `app.py:45-50`) and close the provider once at loop shutdown.
```python
# worker.py — before
async def process_next_queued_job(...):
services = ServiceBundle(...)
try: ...
finally: await services.sources.aclose()
# after: build once in run_worker_loop / lifespan, pass in, close in the lifespan finally
async def run_worker_loop(*, services: ServiceBundle, ...):
try:
while True: ... await process_next_queued_job(services=services, ...)
finally:
await services.sources.aclose()
```
- **Effort:** M
#### [HIGH-03] Provider timeout is capped at 20 seconds by configuration
- **Location:** `src/transcription/config.py:110`
- **Problem & Consequence:** `worker_provider_timeout_seconds: float = Field(default=20.0, gt=0.0, le=20.0)`. The `le=20.0` bound makes 20s both the default *and* the maximum. `workflows.py:238-248` wraps the provider call in `asyncio.wait_for(..., timeout=that_value)`. Multi-modal transcription of a full-page historical document commonly exceeds 20s; operators cannot raise the ceiling without editing source. Every such job fails with `failure_phase="local_timeout"`, and with `worker_max_retries` defaulting to `0` (`config.py:108`) it fails permanently on the first attempt.
Compounding this, `httpx.AsyncClient(follow_redirects=True)` at `openrouter.py:198` sets no explicit `timeout`, so it inherits httpx's 5-second default for connect/read/write/pool unless the OpenRouter SDK overrides it.
- **Recommendation:** Remove the `le=20.0` cap (keep `gt=0.0`), raise the default to something realistic (120s), and set an explicit `httpx.Timeout` derived from the same setting so the transport and the `wait_for` agree.
```python
worker_provider_timeout_seconds: float = Field(default=120.0, gt=0.0)
# openrouter.py
httpx.AsyncClient(follow_redirects=True, timeout=httpx.Timeout(settings.worker_provider_timeout_seconds))
```
- **Effort:** S
#### [HIGH-04] No index on the columns the worker polls every second
- **Location:** `src/transcription/db/models.py:171-212` (`Job.status`, `Job.date_created`, `Job.document_id` all lack `index=True`); also `Source.document_id:252`, `JobSource.job_id:313`, `JobSource.source_id:314`
- **Problem & Consequence:** The worker executes `WHERE status = 'queued' ORDER BY date_created` once per second (`worker.py:130`, `jobs.py:183-185`). Without a composite index this is a full scan plus sort on every tick, and it grows linearly with total job history — not with queue depth. The `JobSource` foreign keys are joined on every job read; PostgreSQL does not auto-index FKs.
- **Recommendation:** Add a composite index for the poll and plain indexes on the hot FKs.
```python
class Job(SQLModel, table=True):
__table_args__ = (Index("ix_job_status_date_created", "status", "date_created"),)
document_id: UUID = Field(foreign_key="document.id", index=True)
```
Note these must also be added to the hand-rolled upgrade path in `db/operations.py` (see [HIGH-05]).
- **Effort:** S
#### [HIGH-05] Hand-rolled schema migrations with SQLite-shaped DDL block the claimed Postgres support
- **Location:** `src/transcription/db/operations.py:25-109`
- **Problem & Consequence:** Schema evolution is a chain of `_upgrade_*` functions issuing raw `ALTER TABLE` / `CREATE INDEX IF NOT EXISTS` against whatever database is present, executed inside `create_all()`. Specific defects:
- `operations.py:77` adds `preferred_execution_attempt_id CHAR(32)` — but the model declares it a `UUID` FK to `execution_attempt.id` (`models.py:260-264`). On PostgreSQL this creates a `char(32)` column that will not compare or join against a native `uuid` column, and the declared foreign key is never created at all.
- Every upgrade is unversioned and re-inspected on each startup; there is no down path, no history table, and no way to tell whether a production database is current.
- `asyncpg` and `psycopg2-binary` are both dependencies (`pyproject.toml:17,21`) and `JSONBCompat` (`models.py:27-35`) carefully supports JSONB, so Postgres is clearly an intended target — but no test exercises it. All 264 tests run on SQLite.
- **Recommendation:** **Re-level the schema from current metadata; do not adopt Alembic.** The application is pre-production, the schema is still evolving, and the existing data is disposable and backed up. Delete `upgrade_schema` and the three `_upgrade_*` functions (`operations.py:25-109`) together with their tests (`tests/test_db.py:109-172`), drop the database, and let `create_all()` generate the schema from SQLModel metadata. This removes the `CHAR(32)` defect at the root rather than patching it, because SQLModel emits the correct column type per dialect automatically (verified: it emits native `UUID` and `JSONB` under the PostgreSQL dialect). `Settings.should_bootstrap_schema` (`config.py:140-145`) already gates the bootstrap path correctly. Reintroduce a migration tool only when the schema stabilizes and real data must survive upgrades.
- **Sequencing:** This must land in the *same* pass as [HIGH-04] (missing indexes), [CRIT-02] (`lazy` flip), and [HIGH-08] (`use_alter`), because all four regenerate the same schema.
- **Effort:** M
#### [HIGH-06] `ty` is a configured dev tool but produces 197 diagnostics and cannot gate CI
- **Location:** `pyproject.toml:38`; suppression comments throughout, e.g. `src/transcription/services/jobs.py:67-68,103-104,123-124,156,180-181`
- **Problem & Consequence:** The project pins `ty` as its type checker, but the codebase suppresses SQLModel relationship typing with `# pyright: ignore[reportArgumentType]` — a *pyright* directive that `ty` does not honor. Result: `ty check` emits 197 diagnostics (160 `invalid-argument-type`, 18 `unresolved-attribute`, 13 `not-subscriptable`), so nobody can run it as a gate, and genuine errors hide in the noise. Two real bugs are buried in there:
- `tests/ui/test_sources_page.py:25` — `Source(...)` constructed without the required `document_id`.
- `tools/run_destructive_tests.py:76,80` — `fcntl` is imported and used, but `fcntl` does not exist on Windows, which is this project's development platform.
- **Recommendation:** Pick one checker and commit to it. If `ty`: replace `# pyright: ignore[...]` with `# ty: ignore[...]`, or better, eliminate the root cause by adopting [CRIT-02]'s `lazy="raise"` change plus typed column accessors, which removes most `selectinload` diagnostics outright. Then wire `ty check` into pre-commit (`pre-commit` is already a dev dependency at `pyproject.toml:35`).
- **Effort:** M
#### [HIGH-07] UI pages own persistence and ORM-loader concerns (violates `ui.instructions.md`)
- **Location:** `src/transcription/ui/pages/jobs_page.py:17,185-192`; `src/transcription/ui/pages/sources_page.py:13,439`; `src/transcription/ui/components/document_panzoom.py:12,61,65`
- **Problem & Consequence:** `ui.instructions.md` states pages must not import sessions or manage transactions, and components must not resolve app state. Three violations:
- `jobs_page.py` imports `transcription.db.session.session_scope` and manages the session lifecycle itself around `create_job_for_document`, while every sibling call site goes through a service.
- `sources_page.py:439` imports `sqlalchemy.inspect` and reads `inspect(attempt).unloaded` to decide rendering — the presentation layer is now coupled to the loader strategy, and will silently misbehave if a service changes its deferred columns.
- `document_panzoom.py` calls `get_settings()` inside a component and re-implements upload-path resolution.
- **Effort:** M
#### [HIGH-08] Circular foreign-key cycle makes `create_all` fail on PostgreSQL
- **Location:** `src/transcription/db/models.py:260-264` (`Source.preferred_execution_attempt_id`), with the cycle running `source` → `job_source` → `execution_attempt` → `source`
- **Problem & Consequence:** Verified by compiling the SQLModel metadata against the PostgreSQL dialect, which emits:
> `SAWarning: Cannot correctly sort tables; there are unresolvable cycles between tables "execution_attempt, job_source, source", which is usually caused by mutually dependent foreign key constraints.`
The resulting sort order places `execution_attempt` **before** `source`, but `execution_attempt.source_id` is a foreign key to `source.id`. On PostgreSQL, where foreign keys are enforced inline at `CREATE TABLE` time, this is a hard `create_all()` failure. SQLite does not enforce the ordering, so the defect is completely invisible on the current test suite and will surface only at the moment of the Postgres cutover.
- **Recommendation:** Mark the nullable leg of the cycle with `use_alter=True` so SQLAlchemy emits it as a deferred `ALTER TABLE ... ADD CONSTRAINT` after all tables exist. Verified to silence the warning and produce a correct ordering.
```python
# models.py — Source
preferred_execution_attempt_id: UUID | None = Field(
default=None,
sa_column=Column(
GUID(),
ForeignKey("execution_attempt.id", use_alter=True, name="fk_source_preferred_attempt"),
nullable=True,
),
)
```
This is cheap, harmless on SQLite, and should land with the schema re-level ([HIGH-05]) so the Postgres path is unblocked whenever it is taken.
- **Effort:** S
### Medium Severity
#### [MED-01] Blocking filesystem and CPU work on the async event loop
- **Location:** `src/transcription/services/store.py:363`; `src/transcription/services/people.py:623`; `src/transcription/services/sources.py:908-923,941,1297,1354`; `src/transcription/services/normalization.py:52-101`; `src/transcription/services/prompts.py:138,162,173-176`; `src/transcription/ui/homepage_store.py:22,28,41`; `src/transcription/ui/pages/home_page.py:87,92`; `src/transcription/ui/pages/people_page.py:495`
- **Problem & Consequence:** All media persistence and artifact I/O is synchronous, called from `async def` paths. `_write_external_artifact` (`sources.py:908`) additionally calls `os.fsync()`, which can block for tens of milliseconds. `normalize_orientation` (`normalization.py:52`) runs full Pillow decode/transpose/re-encode at `quality=95, subsampling=0` inline — that is CPU-bound work measured in hundreds of milliseconds for a scanned page. Every one of these stalls the single event loop shared by the FastAPI API, all NiceGUI clients, and the worker.
- **Recommendation:** Route blocking work through `asyncio.to_thread` at the service boundary (one wrapper per operation, not per call site). For NiceGUI handlers, `nicegui.run.io_bound` / `run.cpu_bound` are the idiomatic equivalents. `normalize_orientation` is the highest-value single conversion.
- **Effort:** M
#### [MED-02] Dead configuration surface: three settings are defined and tested but never read
- **Location:** `src/transcription/config.py:99` (`sqlite_check_same_thread`), `:109` (`worker_retry_backoff_seconds`)
- **Problem & Consequence:** `sqlite_check_same_thread` is never read — `engine.py:43` hardcodes `{"check_same_thread": False}`. `worker_retry_backoff_seconds` is never read either; `tests/test_config.py:152` asserts its default, which gives false confidence that backoff exists. `services.instructions.md:59` mandates a retry path with backoff, and `workflows.py:159-169` implements the `FAILED → QUEUED` transition, but nothing ever sleeps between attempts. Additionally, `advance_job` is invoked exactly once per `process_next_queued_job` call, so a job that transitions `FAILED → QUEUED` is only retried on a later poll — a documented behavior that reads as accidental.
- **Recommendation:** Either wire `worker_retry_backoff_seconds` into the retry scheduler (a `next_attempt_at` column filtered in the claim query is the correct shape — sleeping in the worker loop would stall all other jobs) or delete both settings and their tests. Honor `sqlite_check_same_thread` in `engine.py:43` or remove it.
- **Effort:** S
#### [MED-03] Runtime `inspect.signature` and `getattr` duck-typing at the provider boundary
- **Location:** `src/transcription/services/sources.py:1237-1238,1242-1244`; `src/transcription/services/sources.py:152-154`; `src/transcription/services/workflows.py:297-302`
- **Problem & Consequence:** The `TranscriptionProvider` Protocol (`providers/base.py:102-117`) already declares `requested_model` as a parameter, yet `sources.py:1237` re-checks for it at runtime via `inspect.signature(adapter.transcribe).parameters` on **every transcription call**, then builds an untyped `dict` of kwargs. Similarly, `aclose` and `current_request_manifest` / `current_transport_evidence` are accessed via `getattr(..., None)` even though they are part of the de-facto contract. This defeats static checking on the most important interface in the system, adds per-call reflection overhead, and means a provider that silently drops `requested_model` fails only at runtime.
- **Recommendation:** Extend the Protocol to declare `aclose()`, `current_request_manifest`, and `current_transport_evidence`; then call `adapter.transcribe(...)` with real keyword arguments and drop the `inspect` import.
```python
class TranscriptionProvider(Protocol):
current_request_manifest: RequestManifest | None
current_transport_evidence: TransportEvidence | None
async def transcribe(self, *, prompt_text: str, ..., requested_model: str | None = None) -> TranscriptionResult: ...
async def aclose(self) -> None: ...
```
- **Effort:** S
#### [MED-04] `@cache` on `get_settings(**kwargs)` and on engine/session factories creates cross-test and cross-tenant coupling
- **Location:** `src/transcription/config.py:148-151`; `src/transcription/db/engine.py:39-55`; `src/transcription/db/session.py:20-26`
- **Problem & Consequence:** `get_settings(**kwargs: Any)` is `@cache`-decorated with arbitrary keyword arguments — any unhashable value raises `TypeError`, and the cache key is the kwargs tuple, so `get_settings()` and `get_settings(environment="test")` return different singletons. More seriously, `dispose_engine(database_url)` (`engine.py:50-55`) calls `get_engine.cache_clear()`, which evicts **all** cached engines, not just the one being disposed; a multi-database process would silently lose its other engines' pools. The same pattern applies to `dispose_session_factory` (`session.py:48-50`).
- **Recommendation:** Replace the caches with an explicit registry keyed by URL that supports targeted eviction. `db/runtime.py` already models lifespan-owned resources correctly — extend that pattern rather than layering `functools.cache` beneath it. Separately, drop `**kwargs` from `get_settings` and keep it a true zero-argument singleton.
- **Effort:** M
#### [MED-05] Dead compatibility aliases and a three-way import path for one function
- **Location:** `src/transcription/services/store.py:35,382-383`; `src/transcription/services/transcription.py:12,36`; imports at `store.py:26`, `workflows.py:42`, `sources.py:1263`
- **Problem & Consequence:** `build_prompt_execution` is defined in `sources.py:1263` and imported through three different paths: `store.py` uses `from .transcription import build_prompt_execution`, `workflows.py` uses `from .sources import ...`, and `tests/test_prompts.py:12` uses a third. `transcription.py` (41 lines) exists solely as a re-export shim. Alongside it, `UploadError = SourceStorageError` (`store.py:35`), `create_upload_job = create_document_job` (`store.py:382`), and `store_file = store_source_file` (`store.py:383`) are aliases with zero remaining callers.
- **Recommendation:** Delete the three aliases and the `transcription.py` shim; standardize all imports on `services.sources`.
- **Effort:** S
#### [MED-06] `ServiceBundle` default factories construct four services against global settings
- **Location:** `src/transcription/services/__init__.py:15-22`; consumed at `src/transcription/worker.py:157-158`
- **Problem & Consequence:** `ServiceBundle` declares `field(default_factory=DocumentService)` for all four services. Instantiating `ServiceBundle()` therefore calls `get_settings()` and `resolve_session_factory()` four times, binding to process-global state. `worker.py:157` takes exactly this path whenever `session_factory is None`. This is the "global singleton instead of injected dependency" pattern the FastAPI DI system exists to avoid, and it makes the worker's database target implicit.
- **Recommendation:** Remove the default factories and require explicit construction, plus a single `ServiceBundle.from_session_factory(factory, settings)` classmethod — which also removes the four-way duplication of the same construction block at `app.py:45-50` and `worker.py:160-165`.
- **Effort:** S
#### [MED-07] Unused `asyncio.Queue` allocated in every service instance
- **Location:** `src/transcription/services/base.py:20,26,30`
- **Problem & Consequence:** `ServiceBase.__init__` does `self.queue = queue or asyncio.Queue()`. No code anywhere reads `self.queue`. The annotation is the unparameterized `asyncio.Queue`. Constructing an `asyncio.Queue` also binds to the running event loop policy, so building a `ServiceBundle` outside a loop is a latent hazard, and per [MED-06] this happens four times per bundle.
- **Recommendation:** Delete the `queue` attribute and constructor parameter.
- **Effort:** S
#### [MED-08] Exception swallowed to `None` in an ORM model property
- **Location:** `src/transcription/db/models.py:220-233`
- **Problem & Consequence:** `Job.filename` reaches into `job_source.__dict__` to dodge lazy loading, then catches `DetachedInstanceError` *and* bare `Exception` (`models.py:227`), returning the string `"unknown"`. Any genuine error — a corrupted row, a mapper misconfiguration — is silently rendered as "unknown" in the UI with no log line. The workaround exists only because of the eager-loading design in [CRIT-02].
- **Recommendation:** Remove the property from the model and compute the display value in the feature table read model (`ui/components/table/jobs.py`), which is where `ui.instructions.md` says presentation formatting belongs. If it stays, drop the bare `except Exception` and log the `DetachedInstanceError` case.
- **Effort:** S
#### [MED-09] Large inline SVG asset embedded in a Python module
- **Location:** `src/transcription/ui/theme.py:36-40` (single 23,317-character line)
- **Problem & Consequence:** `VIBESCRIBE_LOGO_SVG` is a 23KB string literal inside a Python source file. It trips `ruff`'s `line-too-long`, makes the module unreadable and undiffable, and contradicts `ui.instructions.md`'s rule that static assets live under `ui/static/` and be read via `importlib.resources`. The project already has exactly the right helper for this — `ui/resources.py:10-19`'s cached `importlib.resources` reader.
- **Effort:** S
#### [MED-10] `DATABASE_URL` is silently ignored by `Settings`
- **Location:** `src/transcription/config.py` (`Settings`, nested `database` config); `docker-compose.yml:10`
- **Problem & Consequence:** `docker-compose.yml:10` sets `DATABASE_URL`, plainly intending to point the application at a different database. `Settings` reads its database configuration from a *nested* `database` model with `env_nested_delimiter="__"` and `extra="ignore"`, so `DATABASE_URL` matches nothing and is discarded without warning. Verified at runtime: with `DATABASE_URL=postgresql://...` exported, `get_settings().database` still resolves to `driver='sqlite' path='./data/transcription.db'`. An operator following the committed compose file gets SQLite while believing they configured PostgreSQL — silent, and the failure mode is data written to the wrong place.
- **Recommendation:** Pick one contract and make the other loud. Either add an explicit `DATABASE_URL` field that parses a full URL into the nested settings, or delete `DATABASE_URL` from `docker-compose.yml` and document `DATABASE__DRIVER` / `DATABASE__PATH`. Given [HIGH-05] defers PostgreSQL, the correct V4.6 action is to remove the misleading compose variable and document the real nested names. Escalates to Critical the moment PostgreSQL is enabled.
- **Effort:** S
#### [MED-11] `DocumentType` and `PersonRole` registry CRUD is duplicated wholesale
- **Location:** `src/transcription/services/documents.py:49-61,64-72,350-500`; `src/transcription/services/people.py:49-79,214-378`
- **Problem & Consequence:** The two models are structurally identical (`id, semantic_key, label, normalized_label, is_active, created_at, updated_at`) and carry identical operation sets, guards, and error mappings:
| Operation | `DocumentType` | `PersonRole` |
| :--- | :--- | :--- |
| label normalizer + casefold key | `documents.py:49-61` | `people.py:49-75` |
| summary dataclass | `documents.py:64-72` | `people.py:79` |
| list / list summaries with counts | `documents.py:350-388` | `people.py:214-249` |
| create, `IntegrityError` → conflict | `documents.py:390-413` | `people.py:251-274` |
| read, not-found raise | `documents.py:415-430` | `people.py:276-291` |
| update, `IntegrityError` → conflict | `documents.py:432-461` | `people.py:293-322` |
| delete, built-in guard + referenced guard | `documents.py:463-491` | `people.py:324-352` |
| `is_*_referenced` | `documents.py:493-500` | `people.py:354-378` |
The duplication extends to the wording of the user-facing suggestion strings ("Deactivate the type instead" / "Deactivate the role instead"). Any fix to one — a normalization bug, a missing guard, an error-category correction — has to be remembered twice.
- **Recommendation:** Introduce a generic `RegistryService[ModelT]` base that owns the eight operations, the label normalization, and the `IntegrityError` mapping. Each concrete registry declares its model, its error class, its reference query, and its noun for message templating. Collapses roughly 200 lines and makes a third registry nearly free.
- **Effort:** M
#### [MED-12] 38 hand-written "not found" raises; the helper that solves it exists and is used once
- **Location:** `src/transcription/services/people.py` (15 sites), `sources.py` (14), `documents.py` (9); helper at `documents.py:123-132`
- **Problem & Consequence:** The pattern `entity = await session.get(Model, id)` / `if entity is None: raise <Error>(f"... {id} not found", category=ErrorCategory.NOT_FOUND, suggestion=...)` is written out longhand 38 times across the service layer, roughly 150 lines. `DocumentService._get_document_or_raise` (`documents.py:123-132`) already implements exactly this — but it is called from only one site (`documents.py:533`), while the identical block is still hand-written at `documents.py:174`, `210`, and `291` in the same file. The abstraction was created and then not adopted, which is the worst of both outcomes: the maintenance burden of a helper plus the drift risk of copies.
- **Recommendation:** Promote the helper to `ServiceBase` and adopt it everywhere.
```python
# services/base.py
async def _get_or_raise[T](
self, session: AsyncSession, model: type[T], entity_id: UUID, *,
error: type[AppError], noun: str, suggestion: str,
) -> T: ...
```
- **Effort:** M
#### [MED-13] Three parallel media-storage implementations
- **Location:** `src/transcription/services/store.py:319-379`; `src/transcription/services/people.py:596-631`; `src/transcription/ui/homepage_store.py:31-44`
- **Problem & Consequence:** `store_source_file`, `store_person_portrait`, and the homepage image writer each independently perform: empty-content check → extension allowlist check → `mkdir(parents=True, exist_ok=True)` → `write_bytes` → wrap `OSError` in a domain error → log. They differ in which of those steps they actually do, so the guarantees are inconsistent — only one of the three hashes its content. All three also block the event loop ([MED-01]).
- **Recommendation:** Consolidate into `services/media_storage.py` per §4, wrapping the write in `asyncio.to_thread`. Resolves this finding and [MED-01] together.
- **Effort:** M
#### [MED-14] `SourceService` owns four domain models, violating the project's own service rule
- **Location:** `src/transcription/services/sources.py` (1254 lines)
- **Problem & Consequence:** `.github/instructions/services.instructions.md:12` states "1 service class per data model." `SourceService` owns `Source`, `JobSource`, `ExecutionAttempt`, and `ProcessingArtifact`:
| Responsibility | Lines |
| :--- | :--- |
| Source CRUD, navigation, listing | 157-345 |
| JobSource association CRUD | 347-510 |
| Evidence write (`update_job_source_transcription`) | 511-670 |
| Attempt promotion and listing | 672-725 |
| Artifact storage (JSON, binary, external, verify) | 727-992 |
| Evidence export | 994-1091 |
| Revisions | 1093-1141 |
The clearest symptom is `update_job_source_transcription` — 160 lines, 17 keyword parameters, mutating five models in one call. The same instruction file (line 13) says an operation spanning more than one service "needs to have a separate orchestration function"; this method *is* that orchestration function, living inside a service. The size also made `sources.py` an import hub: `documents.py:24` and `store.py:24-25` both import from it, and `documents.py:24` importing `source_mime_type` violates the "services are completely independent" rule at line 13.
- **Recommendation:** Extract `ExecutionAttempt` and `ProcessingArtifact` into their own services and relocate `update_job_source_transcription` to `workflows.py` as orchestration. Keep `Source` and `JobSource` together — they are written in the same transaction on every path, and separating them would add ceremony without benefit. Move `source_mime_type` to a shared module so `documents.py` no longer imports a sibling service.
- **Deferred to V4.7.** This touches the transcription write path and is too large to absorb alongside the V4.6 schema re-level.
- **Effort:** L
### Low Severity
#### [LOW-01] `ruff check` fails on 6 issues, 5 auto-fixable
- **Location:** `src/transcription/ui/theme.py:38,40`; `tests/test_app.py:23`; `tests/ui/test_upload_page.py:44`; plus 2 others
- **Recommendation:** Run `ruff check --fix`; the only non-trivial one is the SVG line, addressed by [MED-09].
- **Effort:** S
#### [LOW-02] Stale path reference in project instructions
- **Location:** `.github/instructions/services.instructions.md:10`
- **Problem:** Points to `src/transcription/models.py`; the actual location is `src/transcription/db/models.py`.
- **Effort:** S
#### [LOW-03] `list_jobs` accepts and discards a parameter
- **Location:** `src/transcription/services/jobs.py:113-120` (`_ = load_docs`)
- **Problem:** A dead parameter kept alive only to satisfy `ARG` linting. Callers may believe it changes behavior.
- **Recommendation:** Remove the parameter and update callers.
- **Effort:** S
#### [LOW-04] `resolve_worker_notifier` returns unvalidated `getattr` results
- **Location:** `src/transcription/worker.py:54-61`
- **Problem:** Any non-`None` attribute is returned as a `WorkerNotifier` without checking it has `notify`. Compare `app_state.py:15-18`, which correctly uses `isinstance`.
- **Effort:** S
#### [LOW-05] Untyped handler parameters and loosely-typed dict returns in UI
- **Location:** `ui/pages/jobs_page.py:491`; `ui/pages/people_page.py:492`; `ui/pages/home_page.py:85`; `_render_document_form_fields` / `_render_person_form_fields` returning `dict[str, Any]`
- **Recommendation:** Annotate with `nicegui.events.UploadEventArguments`; replace the form-field dicts with frozen dataclasses.
- **Effort:** S
#### [LOW-06] Auto-refresh timer deactivated but never cancelled; magic interval
- **Location:** `ui/pages/jobs_page.py:245,251,253`
- **Problem:** `ui.timer(4.0, refresh_job)` is toggled via `.active = False` rather than `.cancel()`; `4.0` is an unnamed literal. Client-scoped, so impact is bounded.
- **Effort:** S
#### [LOW-07] `people_page.py:504` catches `Exception` and discards it entirely
- **Location:** `ui/pages/people_page.py:504`
- **Effort:** S
#### [LOW-08] Four avoidable query inefficiencies in `sources.py`
- **Location:** `src/transcription/services/sources.py:233-244, 338-343, 961, 1012-1013`
- **Problem & Consequence:**
- `list_sources_detail:338-343` filters by `job_id` **in Python**, after loading every `Source` row and its eager graph, instead of joining `JobSource` in SQL. Cost grows with the whole table rather than with the result set.
- `read_source_navigation:233-244` fetches the complete ordered id list for a document to identify two neighbours. Two `LIMIT 1` queries (`page_number < n ORDER BY page_number DESC`, and the mirror) return the same answer at constant cost.
- `list_processing_artifacts:961` has no `limit` parameter while its sibling `list_processing_artifact_summaries:980` does, and it loads `inline_payload` blobs that the caller frequently does not need.
- `build_evidence_export:1012-1013` re-reads and re-hashes every external artifact file synchronously on the event loop before serializing. Integrity verification is correct to perform, but it belongs in `asyncio.to_thread` ([MED-01]).
- **Recommendation:** Push the `job_id` filter into SQL, replace the navigation scan with two bounded queries, add a `limit` to `list_processing_artifacts`, and move artifact hashing off the loop.
- **Effort:** S
---
## 3. Stack-Specific Analysis
### Python 3.12+ Best Practices
Modern syntax is used consistently and correctly: `type` statements (`db/session.py:17,45,73,102`), `X | None` unions, `StrEnum`, `match` statements (`db/engine.py:16-31`, `db/session.py:84-92`, `workflows.py:153-171`), frozen `dataclass(slots=True)`, and `pathlib` throughout — no `os.path` anywhere. Gaps: unparameterized `asyncio.Queue` ([MED-07]), untyped `prompt_execution` parameters (`store.py:192,247`), the untyped kwargs dict at `sources.py:1229-1238` ([MED-03]), and the swallowed exception at `models.py:227` ([MED-08]). Broad `except Exception` appears frequently but is almost always accompanied by `# noqa: BLE001` and immediate normalization through `classify_unexpected_error` — that is a defensible boundary pattern, not a defect.
### FastAPI
`create_app` (`app.py:87-111`) is a clean factory using the modern `lifespan` context manager, not the deprecated `@app.on_event`. Routers are domain-organized with prefixes and tags. `response_model` is declared on every route. Dependency injection is used correctly in `api/v4_documents.py:111-128`, with the useful touch that `get_document_service` prefers lifespan-owned state and falls back gracefully. Two gaps: service methods called from `async def` endpoints perform synchronous file I/O ([MED-01]), and `_recover_stale_processing_jobs` (`app.py:73-84`) constructs a throwaway `JobService` rather than using the bundle built five lines earlier.
### NiceGUI
The strongest layer of the codebase in terms of convention adherence. CSS discipline is exemplary: a single `ui.add_css(read_css("theme.css"), shared=True)` at the composition root (`ui/__init__.py:28`), read through `importlib.resources` with a `@cache`-backed loader and path validation (`ui/resources.py:10-19`), and zero inline `.style()` calls or `<style>` blocks in components. **No cross-client state leakage was found** — per-request state lives in page-function closures, and the only module-level globals are idempotent registration flags (`theme.py:15`, `_register_panzoom_assets`'s `lru_cache`). `error_presenter.show_error/summarize_error` is applied uniformly and preserves `AppError` id/category/suggestion. The table architecture (generic `build_table` + per-feature row read models) matches the documented split. Defects are the boundary violations in [HIGH-07], the blocking I/O in [MED-01], and the duplication catalogued in §4.
### SQLModel & SQLAlchemy
Session lifecycle is the clear high point: `ServiceBase._finalize` (`services/base.py:41-60`) implements a genuinely well-reasoned commit-vs-flush ownership protocol that lets orchestration functions commit exactly once at the workflow boundary, and `session_scope` / `transaction_scope` (`db/session.py:53-105`) express the two modes cleanly, with `transaction_scope` correctly rejecting a supplied session that has no active transaction. `JSONBCompat` (`models.py:27-35`) is the right cross-dialect abstraction, and `BigInteger` for `file_size_bytes` and `StaticPool` for in-memory SQLite show real attention to portability. Against that, the eager-loading defaults ([CRIT-02]), missing indexes ([HIGH-04]), unlocked job claim ([CRIT-01]), and hand-rolled DDL ([HIGH-05]) are the four issues that most need attention. Note also that `models.py` sets `updated_at` / `date_updated` via `default_factory` only — there is no `onupdate`, so these columns are stale unless a service sets them by hand (`jobs.py:166` does; most other update paths do not).
### Pydantic V2 & Settings
Fully V2-native. No `@validator`, no `class Config`, no `.dict()`, no `parse_obj` anywhere. `model_config = ConfigDict(...)` is used consistently, usually with `extra="forbid", frozen=True` — a good default that catches provider payload drift. `Settings` is a single `BaseSettings` source of truth with `env_nested_delimiter`, a discriminated `DatabaseSettings` union, `SecretStr` for credentials, and constrained `Annotated` types (`NonEmptyStr`, `Probability`, `Temperature`). There are **no scattered `os.getenv` calls** in `src`. The one wart is `object.__setattr__` in `normalize_provider_models` (`config.py:130,137`) to mutate a frozen model — functional but fragile; `model_copy(update=...)` or a computed property would express it more safely. Issues: [MED-02] dead settings, [HIGH-03] the timeout cap, [MED-04] the `@cache` signature.
### Asyncio Workers
`worker_consumer_lifespan` (`worker.py:64-94`) is well-built: it holds a strong reference to the task, sets the stop event, wakes the loop, waits with a bounded timeout, and escalates to `cancel()` + `suppress(CancelledError)` on timeout — a correct graceful-shutdown sequence. The `WorkerNotifier` Protocol with `Event`/`Noop` implementations is a clean seam. `_persist_page_outcome_durably` (`workflows.py:486-499`) uses `asyncio.shield` with correct cancellation re-raise so a shutdown mid-job cannot lose provider evidence — a genuinely subtle piece of code done right. Remaining concerns: the single-worker assumption is unenforced ([CRIT-01]), there is no backpressure or concurrency limit (jobs are processed strictly serially, so a large backlog drains slowly while the provider sits idle), and blocking I/O inside the loop ([MED-01]) stalls both the worker and all HTTP/UI clients on the same loop.
### OpenRouter / Adapter Boundary
Encapsulation is good — no OpenRouter-specific header, model name, or payload shape appears in `services`, `api`, or `ui`. `providers/__init__.py`'s factory keeps the concrete adapter behind `get_transcription_provider`. Responses are validated through real Pydantic schemas (`OpenRouterResponse`, `ResponseChoice`, `ResponseUsage`), and `_CapturingAsyncClient` / `_CapturingAsyncByteStream` (`openrouter.py:45-91`) is a thoughtful mechanism for retaining raw transport bytes for evidence without disturbing SDK parsing. The failures are lifecycle and typing: per-job client churn ([HIGH-02]), no explicit httpx timeout ([HIGH-03]), and runtime `inspect`/`getattr` duck-typing instead of an honest Protocol ([MED-03]).
### Testing & Quality Tooling
264 tests pass with 4 skipped; markers (`unit`/`integration`/`external`) are declared and `--strict-markers` is on; `filterwarnings` escalates never-awaited coroutines to errors — a good async-specific guard. Coverage is broad across services, providers, API, and UI pages. The material gaps: (a) `asyncio_mode = "strict"` is set but no `asyncio_default_fixture_loop_scope` is configured, which pytest-asyncio warns about and which will change behavior on upgrade; (b) **every test runs on SQLite**, so the Postgres support that `JSONBCompat`, `asyncpg`, and `psycopg2-binary` all exist to provide is entirely unverified — [HIGH-05]'s `CHAR(32)` bug is exactly the class of defect this would catch; (c) no test asserts concurrent job-claim safety, which is why [CRIT-01] survives; (d) `ty` cannot gate ([HIGH-06]); (e) `tools/run_destructive_tests.py:76,80` uses `fcntl`, unavailable on this project's Windows development platform.
---
## 4. Duplication & Consolidation Report
| Pattern / Duplication | Locations | Proposed Canonical Home | Est. Lines Removed |
| :--- | :--- | :--- | :--- |
| Delete-confirmation page scaffold (blocked-deps card + confirm/cancel row) | `ui/pages/documents_page.py:354-433`; `jobs_page.py:358-428`; `sources_page.py:204-277`; `people_page.py:~270-310` | `ui/components/confirm_delete.py` | ~120 |
| Upload/media URL resolution (`_resolve_*_src`, `_to_absolute_upload_url`) | `ui/pages/sources_page.py:711-770`; `people_page.py:525-579`; `ui/components/document_panzoom.py:59-75` | `ui/components/media_urls.py` (pure, takes `upload_dir` + `base_url`) | ~110 |
| Invalid-id / not-found guard (parse → red label → return) | `documents_page.py:172-184,219-231,275-287,359-371`; `jobs_page.py:212-221,310-319,363-372`; `sources_page.py:111-137,207-222` | `ui/components/guards.py:load_or_render_error(...)` | ~90 |
| Hand-rolled `ui.table` instead of `build_table` | `ui/pages/settings_page.py:65-113` + person-roles table; `ui/components/linked_people.py:66-75`; `print_preview_page.py:125-161` | `ui/components/table/common.py:build_table` (add selection / no-search options) | ~70 |
| File-picker upload wiring | `people_page.py:491-522`; `jobs_page.py:449-503`; `home_page.py:38-40,85-89` | `ui/components/upload_panel.py` | ~50 |
| `_parse_uuid` | `documents_page.py:591`; `jobs_page.py:558`; `sources_page.py:780`; `people_page.py:589`; `linked_people.py:175` | `ui/components/formatters.py` | ~35 |
| `_resolve_runtime_settings(request)` | `jobs_page.py:567`; `sources_page.py:773`; `people_page.py:582` | Shared page-helper module | ~18 |
| `_parse_iso_date` | `documents_page.py:600`; `people_page.py:598` | `ui/components/formatters.py` | ~14 |
| `ServiceBundle` construction block (4 identical service instantiations) | `app.py:45-50`; `worker.py:160-165`; `services/__init__.py:19-22` | `ServiceBundle.from_session_factory(...)` classmethod | ~20 |
| "Next queued job" query, two divergent implementations | `services/jobs.py:170-187` (no `LIMIT`); `db/operations.py:143-151` (has `LIMIT`) | `JobService.claim_next_queued_job` (per [CRIT-01]); delete the `operations.py` copy | ~12 |
| `build_prompt_execution` re-export shim + legacy aliases | `services/transcription.py` (whole module); `services/store.py:35,382,383` | `services/sources.py` (single import path) | ~45 |
| `store_source_file` / `store_person_portrait` / `store_homepage_image` — three near-identical validate-hash-write-bytes flows | `services/store.py:319-379`; `services/people.py:596-631`; `ui/homepage_store.py:31-44` | `services/media_storage.py` (one async, `to_thread`-wrapped writer) | ~60 |
| Registry CRUD (list / summaries / create / read / update / delete / referenced) for `DocumentType` and `PersonRole` | `services/documents.py:350-500`; `services/people.py:214-378` | `services/registry.py:RegistryService[ModelT]` ([MED-11]) | ~200 |
| Label normalization + casefold key + summary dataclass | `services/documents.py:49-72`; `services/people.py:49-79` | `services/registry.py` (base) | ~35 |
| `get(...)` → `if None: raise ...NOT_FOUND` guard, written longhand 38 times | `services/people.py` (15), `sources.py` (14), `documents.py` (9) | `ServiceBase._get_or_raise` ([MED-12]) | ~150 |
### Proposed Canonical Abstractions
```python
# src/transcription/services/media_storage.py
async def store_media(
*, filename: str, content: bytes, root: Path, relative_directory: Path | None = None,
filename_stem: str | None = None, validate: Callable[[str, bytes], None] | None = None,
) -> StoredMedia: ... # StoredMedia = frozen dataclass(path, sha256, byte_size, media_type)
# wraps the blocking write in asyncio.to_thread — resolves [MED-01]
# src/transcription/services/__init__.py
@classmethod
def from_session_factory(cls, factory: SessionFactory, settings: Settings | None = None) -> ServiceBundle: ...
# src/transcription/services/jobs.py
async def claim_next_queued_job(self, *, session: AsyncSession | None = None) -> Job | None: ...
# atomic QUEUED -> PROCESSING with LIMIT 1 + FOR UPDATE SKIP LOCKED
# src/transcription/services/registry.py
class RegistryService[ModelT: RegistryModel](ServiceBase):
"""Shared CRUD for semantic-key registries (DocumentType, PersonRole)."""
model: type[ModelT]
error: type[AppError]
noun: str
async def list_all(self, *, active_only: bool = True, session=None) -> Sequence[ModelT]: ...
async def list_summaries(self, *, session=None) -> Sequence[RegistrySummary]: ...
async def create(self, *, label: str, is_active: bool = True, session=None) -> ModelT: ...
async def read(self, entity_id: UUID, *, session=None) -> ModelT: ...
async def update(self, entity_id: UUID, *, label: str, is_active: bool, session=None) -> ModelT: ...
async def delete(self, entity_id: UUID, *, session=None) -> None: ...
async def is_referenced(self, entity_id: UUID, *, session=None) -> bool: ...
def _reference_query(self, entity: ModelT) -> Select[tuple[UUID]]: ... # subclass hook
# src/transcription/services/base.py
async def _get_or_raise[T](
self, session: AsyncSession, model: type[T], entity_id: UUID, *,
error: type[AppError], noun: str, suggestion: str,
) -> T: ... # absorbs 38 hand-written not-found blocks — resolves [MED-12]
# src/transcription/ui/components/media_urls.py
def build_upload_url(*, file_path: Path, upload_dir: Path, base_url: str) -> str | None: ...
# src/transcription/ui/components/confirm_delete.py
def render_confirm_delete(
*, title: str, blockers: Sequence[str], on_confirm: Callable[[], Awaitable[None]],
on_cancel: Callable[[], None],
) -> None: ...
# src/transcription/ui/components/guards.py
def parse_uuid_or_render_error(raw: str, *, entity: str) -> UUID | None: ...
```
---
## 5. Prioritized Action Plan
> **Superseded for V4.6.** The three phases below are the original review's sequencing. The V4.6 release restructures this into seven phases against the confirmed operating context in [§1a](#1a-post-review-addendum); see [`ver4.6/implementation_plan_v4_6.md`](ver4.6/implementation_plan_v4_6.md). The material differences are: Alembic is replaced by a schema re-level; the schema-affecting items are merged into a single pass; the service-layer consolidation ([MED-11], [MED-12], [MED-13]) is added; and the `SourceService` split ([MED-14]) is deferred to V4.7.
### Phase 1: Quick Wins (PR 1-2)
1. Delete `src/transcription/app_state.py` — dead module with a live `TypeError` ([HIGH-01]).
2. Remove `le=20.0` from `worker_provider_timeout_seconds`, raise the default, and pass an explicit `httpx.Timeout` to the OpenRouter client ([HIGH-03]).
3. Add `Index("ix_job_status_date_created", "status", "date_created")` and `index=True` on the hot foreign keys ([HIGH-04]).
4. Add `.limit(1)` to `read_next_queued_job` — a one-line change that removes the full-queue load ahead of the full [CRIT-01] fix.
5. Delete `services/transcription.py`, the three `store.py` aliases, and `ServiceBase.queue`; standardize `build_prompt_execution` imports ([MED-05], [MED-07]).
6. Move the 23KB SVG to `ui/static/` and run `ruff check --fix` ([MED-09], [LOW-01]).
7. Resolve or delete `sqlite_check_same_thread` and `worker_retry_backoff_seconds` ([MED-02]).
### Phase 2: Reliability & Concurrency (PR 3-4)
8. Implement `claim_next_queued_job` with `LIMIT 1` + `FOR UPDATE SKIP LOCKED`, delete the `db/operations.py` duplicate, and add a concurrency test that runs two claimers against one queued job ([CRIT-01]).
9. Hoist `ServiceBundle` and the provider client to worker-loop scope so the HTTP connection pool survives across jobs ([HIGH-02], [MED-06]).
10. Wrap blocking media/artifact I/O and Pillow normalization in `asyncio.to_thread` behind a single `services/media_storage.py` ([MED-01]).
11. ~~Adopt Alembic~~ — **superseded**: re-level the schema from current metadata and delete the `_upgrade_*` chain ([HIGH-05]), landing together with [HIGH-04], [HIGH-08], and [CRIT-02] in one pass.
12. Extend `TranscriptionProvider` Protocol to cover `aclose` and the evidence attributes; delete the `inspect.signature` reflection ([MED-03]).
### Phase 3: Consolidation & Refactoring (PR 5-6)
13. Flip relationship defaults to `lazy="raise"` model by model, letting the existing suite prove which explicit `selectinload()` calls are load-bearing ([CRIT-02]). This also removes most of the `# pyright: ignore` comments.
14. Standardize on `ty`, convert remaining suppressions to `# ty: ignore[...]`, and wire `ty check` into the existing pre-commit setup ([HIGH-06]).
15. Fix the three UI boundary violations: session ownership in `jobs_page`, `sqlalchemy.inspect` in `sources_page`, `get_settings()` in `document_panzoom` ([HIGH-07]).
16. Extract the UI duplication per §4, highest value first: `confirm_delete` → `media_urls` → `guards` → `formatters` (~500 lines removed).
17. Replace `functools.cache` on engine/session factories with an explicit URL-keyed registry supporting targeted eviction ([MED-04]).
---
## 6. Preserved Strengths
- **`ServiceBase._finalize` (`services/base.py:41-60`)** — the commit-vs-flush ownership protocol is the single best idea in the codebase. It lets orchestration functions compose multiple services into one atomic transaction without any service knowing about the others, and it is documented in `services.instructions.md`. Keep it and keep enforcing it.
- **Error taxonomy (`errors.py`)** — `AppError` carrying `category`, `suggestion`, `retriable`, and a short shareable `error_id`, with `classify_unexpected_error` normalizing at every boundary and `format_error_detail` producing a stable persisted string. It is applied consistently from services through API handlers to `ui/components/error_presenter.py`.
- **Evidence capture pipeline** — `_CapturingAsyncClient` / `_CapturingAsyncByteStream` (`openrouter.py:45-91`) plus `ExecutionAttempt` / `ProcessingArtifact` with content-addressed digests and integrity verification (`sources.py:925-959`) is a serious, well-executed provenance design that is rare to see done properly.
- **`asyncio.shield` around page-outcome persistence (`workflows.py:486-499`)** — correctly written, including the `await task` before re-raising `CancelledError`, so provider results survive shutdown mid-job.
- **Worker lifespan shutdown (`worker.py:64-94`)** — strong task reference, stop event, wake, bounded wait, then cancel-and-suppress. Textbook correct.
- **Pydantic V2 discipline** — zero V1 residue, `extra="forbid"` + `frozen=True` as the house default, constrained `Annotated` types, `SecretStr` for credentials, discriminated union for database config, and no `os.getenv` anywhere in `src`.
- **UI CSS and asset discipline** — one `add_css` at the composition root, `importlib.resources` with a `@cache`d reader and path validation, semantic `ui-*` classes, no inline styles. This is exactly what `ui.instructions.md` prescribes, followed without exception.
- **No cross-client state leakage in NiceGUI** — per-request state lives in page-function closures; the only module globals are idempotent registration flags. This is the most common NiceGUI defect and this codebase avoids it entirely.
- **Cross-dialect care** — `JSONBCompat`, `BigInteger` for byte sizes, `native_enum=False` with `values_callable` for stable enum storage, `StaticPool` for in-memory SQLite. The intent is right; it just needs Postgres CI to make it real.
- **The instruction files themselves** — `.github/instructions/services.instructions.md` and `ui.instructions.md` are specific, enforceable, and largely followed. Most findings in this report are deviations from rules the project already wrote down, which is a much healthier position than having no rules at all.
+60
View File
@@ -0,0 +1,60 @@
# Backup and Restore (V6.1)
This guide defines operational backup/restore for clean-slate recovery of the Docker runtime using a host-visible backup folder.
## 1. Backup artifacts
- Backup target root: `BACKUP_DIR` (recommended production value: `/backup`)
- Database artifact per run:
- `postgres-YYYYMMDD-HHMMSS.dump` (PostgreSQL custom dump via `pg_dump -Fc`)
- `backup-YYYYMMDD-HHMMSS.manifest` (run manifest)
- Media/config mirrors under `BACKUP_DIR`:
- `uploads/**` (incremental copy: new files only)
- `prompts/**` (prompt directory mirror)
Retention:
- `BACKUP_RETENTION_DAYS` applies to `postgres-*.dump` and `backup-*.manifest` files.
## 2. Creating backups
Run from repository root:
```bash
sh deploy/backup/create_postgres_backup.sh
```
Environment variables used by the backup script:
- `BACKUP_DIR` (default `./data/backups`)
- `BACKUP_RETENTION_DAYS` (default `14`)
- `UPLOAD_DIR` (default `/app/uploads`)
- `PROMPT_DIR` (default `/app/prompts`)
- `DATABASE__DRIVER` (must be `postgres`)
- `DATABASE__HOST` (default `postgres`)
- `DATABASE__PORT` (default `5432`)
- `DATABASE__DATABASE` (required)
- `DATABASE__USER` (required)
- `DATABASE__PASSWORD` (required)
Recommended production setup:
- Mount a host-visible folder into `/backup` for both `app` and `worker`.
- Set `BACKUP_DIR=/backup` in `.env.production`.
- Use host-level tooling (for example Synology Drive Client on the host) to replicate that folder externally.
## 3. Restoring from backup
Restore requires downtime for app + worker writes.
1. Stop app and worker:
- `docker compose --env-file .env.production -f docker-compose.production.yml stop app worker`
2. Restore database:
- `sh deploy/backup/restore_postgres_backup.sh /backup/postgres-YYYYMMDD-HHMMSS.dump`
3. Start app and worker:
- `docker compose --env-file .env.production -f docker-compose.production.yml start app worker`
4. Validate `/healthz` and run one smoke workflow.
Notes:
- `restore_postgres_backup.sh` still supports legacy archive restore paths for older backup sets.
+66
View File
@@ -0,0 +1,66 @@
# Cloudflare Tunnel and Access Setup
This guide defines the repository-supported setup for exposing app and selected LAN services through Cloudflare Tunnel with Cloudflare Access protection.
## 1. Files used by this deployment
1. `deploy/cloudflared/config.yml` (local copy from `config.yml.example`)
2. `.env.production` (`CLOUDFLARE_TUNNEL_TOKEN`)
3. `docker-compose.production.yml` (`cloudflared` service reads token + mounts config)
Do not commit `config.yml` or `.env.production`.
## 2. Configure cloudflared
1. Copy `deploy/cloudflared/config.yml.example` to `deploy/cloudflared/config.yml`.
2. Update hostname -> service mappings in `ingress`.
3. Keep the final catch-all ingress `http_status:404`.
4. Set `CLOUDFLARE_TUNNEL_TOKEN` in `.env.production`.
Example app route:
- `transcription.example.com` -> `http://app:8000`
Optional generic remote-access routes:
- `homeassistant.example.com` -> `http://<home-assistant-lan-ip>:8123`
- `pihole.example.com` -> `http://<pihole-lan-ip>:80`
## 3. Cloudflare Access policy baseline
Create one Access app policy per exposed hostname:
1. Include: your allowed identities/groups only.
2. Exclude: none by default.
3. Require: identity provider login (and MFA if available).
Recommended baseline:
- App endpoint (`transcription.*`): your admin identity set.
- Other internal endpoints (`homeassistant.*`, `pihole.*`, etc.): explicit least-privilege groups.
## 4. Startup
Start production stack:
```bash
docker compose --env-file .env.production -f docker-compose.production.yml up -d --build
```
Validate tunnel container:
```bash
docker compose --env-file .env.production -f docker-compose.production.yml logs cloudflared
```
LXC/proxied-network note:
- The `cloudflared` service is pinned to `--protocol http2` with explicit DNS resolvers (`1.1.1.1`, `1.0.0.1`) in `docker-compose.production.yml`.
- This avoids environments where Docker's embedded resolver (`127.0.0.11`) cannot resolve `region*.v2.argotunnel.com`, which causes connector precheck failure and tunnel shutdown.
- If tunnel status is still down, verify host/container egress for DNS and TCP 443 to `api.cloudflare.com` and `*.argotunnel.com`.
## 5. Security notes
- Keep `postgres` and other internal-only services off public hostnames unless required.
- Use distinct hostnames per service; avoid path-based multiplexing for unrelated admin surfaces.
- Rotate `CLOUDFLARE_TUNNEL_TOKEN` and Access policy memberships on a regular schedule.
+94
View File
@@ -0,0 +1,94 @@
# Database Rebuild Migration Workflow
This project uses an explicit **export/import rebuild workflow** for schema migration.
Policy:
- Do not add runtime legacy-compatibility write paths.
- Rebuild a fresh target database from current models.
- Export current data/media, then import into the fresh target.
## Commands
### 1) Export current DB + uploads into a bundle
```bash
uv run python tools/export_import_migration.py export --bundle-dir .migration-bundle
```
Optional source overrides:
- `--source-db <path-or-sqlalchemy-url>`
- `--source-upload-dir <path>`
### 2) Import bundle into a fresh target (SQLite or PostgreSQL)
```bash
uv run python tools/export_import_migration.py import --bundle-dir .migration-bundle --target-db .\data\transcription-new.db --target-upload-dir .\data-new
```
PostgreSQL target example:
```bash
uv run python tools/export_import_migration.py import --bundle-dir .migration-bundle --target-db postgresql://transcription:change-me@localhost:5432/transcription --target-upload-dir .\data-new
```
### 3) Verify migration parity and integrity
```bash
uv run python tools/export_import_migration.py verify --source-db .\data\transcription.db --target-db postgresql://transcription:change-me@localhost:5432/transcription
```
The verify command checks:
- row-count parity across migration tables
- orphan-reference checks for `source`, `job`, `job_source`, and `execution_attempt`
- duplicate `(job_id, source_id, attempt_number)` in `execution_attempt`
Exit code:
- `0` when counts and integrity checks pass
- `1` when mismatches or integrity violations are detected
### 4) One-shot export+import
```bash
uv run python tools/export_import_migration.py migrate --bundle-dir .migration-bundle --target-db .\data\transcription-new.db --target-upload-dir .\data-new
```
## What gets migrated
- Tables (in dependency order): `document_type`, `person_role`, `tag`, `document`, `person`, `photo`, `document_person`, `document_tag`, `person_tag`, `job`, `source`, `job_source`, `execution_attempt`.
- Media tree under `UPLOAD_DIR`.
The bundle contains:
- `database.json` (row export)
- `uploads/` (copied media files)
Path normalization during export/import:
- `source.file_path` is normalized to `documents/...` (upload-root-relative POSIX).
- `photo.path` is normalized to `photos/...` (upload-root-relative POSIX).
Legacy V4.x portrait/homepage backfill in the export step:
- If the source DB has no `photo` table, the exporter synthesizes `photo` rows from legacy `person.portrait_path` values and from legacy homepage image files under `UPLOAD_DIR/homepage`.
- Legacy portrait and homepage image files are copied into the unified `UPLOAD_DIR/photos/{photo_id}{suffix}` layout in the migration bundle.
- Legacy homepage markdown is relocated from `UPLOAD_DIR/homepage/homepage.md` to `UPLOAD_DIR/homepage.md`.
- Legacy `person.full_name` values are split into `given_names` + `last_name` for V5.1 schema compatibility.
## Cutover (SQLite -> PostgreSQL)
After importing to a fresh target:
1. Stop app and worker services to freeze writes.
2. Export a migration bundle from the last SQLite state.
3. Import bundle to PostgreSQL target.
4. Run `verify` against source and target before switching runtime.
5. Switch runtime config to PostgreSQL (`DATABASE__DRIVER=postgres` and related `DATABASE__*` values).
6. Start app and worker services.
7. Run smoke checks (`/healthz`, create/upload/process one job).
## Rollback
If verify or smoke checks fail:
1. Stop app and worker services.
2. Revert runtime config to SQLite.
3. Start app and worker against pre-cutover SQLite database.
4. Preserve failed migration bundle and logs for analysis.
+137
View File
@@ -0,0 +1,137 @@
# Error Handling Policy (Current Baseline: V6.1)
This policy defines the active V6.1 error taxonomy, translation boundaries, and retry semantics.
## Error Categories
| Category | Meaning | Typical Origin | User Treatment |
| :--- | :--- | :--- | :--- |
| `validation` | Input payload/selection is invalid | UI form parsing, service validators | Inline correction guidance |
| `not_found` | Target record is missing | ID lookup in service layer | Non-blocking warning or redirect |
| `conflict` | State prevents requested action | lifecycle transitions, duplicate semantic keys | Explain required precondition |
| `external` | Provider/network dependency failure | OpenRouter/provider adapter | Retry path and evidence retained |
| `timeout` | Provider call exceeded configured bound | worker/provider client timeout | Retry path and bounded messaging |
| `internal` | Unexpected local failure | unhandled service/runtime faults | Safe generic message + diagnostics capture |
## Runtime Taxonomy and Canonical Mapping
Runtime code uses a richer internal taxonomy for diagnostics and persisted evidence, then maps that
taxonomy to the six canonical categories at the API/UI envelope boundary.
### Internal runtime categories
- `validation_error`
- `user_input_error`
- `not_found_error`
- `conflict_error`
- `external_provider_error`
- `external_timeout_error`
- `processing_error`
- `infrastructure_transient_error`
- `infrastructure_persistent_error`
- `internal_unexpected_error`
### Internal -> Canonical mapping
| Internal category | Canonical envelope category |
| :--- | :--- |
| `validation_error` | `validation` |
| `user_input_error` | `validation` |
| `not_found_error` | `not_found` |
| `conflict_error` | `conflict` |
| `external_provider_error` | `external` |
| `external_timeout_error` | `timeout` |
| `infrastructure_transient_error` | `timeout` |
| `processing_error` | `internal` |
| `infrastructure_persistent_error` | `internal` |
| `internal_unexpected_error` | `internal` |
`ExecutionAttempt.error_category` stores the internal category value so diagnostics remain specific.
## Translation Boundaries
- **Provider layer:** raise provider-scoped exceptions with provider context; do not emit UI text.
- **Service layer:** map raw exceptions into internal categories and preserve causal chain.
- **UI/API layer:** convert internal categories to canonical categories using the centralized mapping.
## Decision Context
### Why taxonomy is category-based (not exception-class-based)
- Categories encode operator-facing recovery semantics (fix input, retry later, investigate internal failure) independent of low-level exception type.
- This keeps retry and messaging behavior consistent even when provider/client libraries change.
### Why page-level failure is isolated
- Multi-page archival documents often contain a mix of readable and degraded pages.
- Isolating failures to page scope preserves successful results and avoids all-or-nothing loss when one page fails.
- Aggregate job status then communicates overall outcome (`transcribed`, `partial_success`, `failed`) without hiding page detail.
### Why retries append evidence instead of mutating rows
- Retry operations are new observations, not corrections of history.
- Appending attempts preserves forensic traceability, timing history, and provider variability analysis.
- Projection updates remain explicit user/workflow decisions, separate from immutable evidence.
## Job and Page Failure Semantics
### Page-Level (`JobSource`)
- `pending` -> `transcribed` when attempt succeeds.
- `pending` -> `failed` when attempt fails terminally.
- `pending` -> `cancelled` on job cancellation before processing.
### Job-Level (`Job`)
- `transcribed` when all pages transcribe successfully.
- `partial_success` when mixed success/failure outcomes exist.
- `failed` when no page transcribes successfully.
## Retry and Retranscription Rules
1. Failed/cancelled pages may be re-queued through retranscription workflows.
2. Retry attempts must append new `ExecutionAttempt` rows; prior evidence remains immutable.
3. Selecting a better candidate must update projection pointers, not mutate historical attempt rows.
## Logging and Diagnostics Rules
1. Persist sufficient attempt error metadata (`error_category`, `error_message`, transport evidence) for post-hoc analysis.
2. Avoid leaking stack traces or local paths into user-facing message envelopes.
3. Preserve causal exception chains for internal diagnostics.
### Message vs detail split
Rules 1 and 2 pull in opposite directions: evidence records need the root cause, and
user-facing envelopes must not carry it. `AppError` therefore separates the two audiences:
| Field | Audience | Carries root cause | Surfaces |
| --- | --- | --- | --- |
| `message` | User-facing and API-facing | No | `show_error`, `build_error_envelope` |
| `detail` | Internal only | Yes | `format_error_detail` (evidence), logs, sanitized UI projection only |
`classify_unexpected_error` builds a generic `message` and puts the exception type and
text on `detail`. Anything rendered to a user or serialized into an API envelope must
read `message`; anything persisted as provenance or logged may read `detail`. When a UI
surface needs to show persisted `error_detail`, it must route through a sanitizing
projection that preserves the category, suggestion, and error reference while reducing
machine-local absolute paths to basenames only.
Enforced by `tests/test_errors.py::test_unexpected_error_does_not_leak_filesystem_paths`
and `tests/test_error_message_safety.py`.
## Operator Recovery Guidance
- **validation/conflict:** correct input or state and retry manually.
- **external/timeout:** allow bounded retries and keep prior attempt evidence visible.
- **internal:** stop automatic retries, surface a safe message, and inspect diagnostics with correlation context.
## UI Messaging Contract
- User-visible errors must be actionable, bounded, and category-consistent.
- Multi-page jobs must show partial outcomes instead of collapsing into a single opaque failure.
- Recovery actions (`retry`, `retranscribe`, `edit input`) must be offered where available.
## Cross-Reference
- [Error Handling invariant](./invariant/error_handling.md)
- [System Requirements](requirements.md)
- [Data Model](schema.md)
+36
View File
@@ -0,0 +1,36 @@
# Document Transcription System Overview (Current Baseline: V6.1)
This directory is the single source of truth for current V6.1 behavior and architecture.
## Canonical Reading Order
1. [System Architecture](architecture.md) for runtime topology, boundaries, and lifecycle ownership.
2. [System Requirements](requirements.md) for verifiable current-state requirements.
3. [Data Model](schema.md) for entities, constraints, and evidence persistence rules.
4. [Error Handling Policy](error_handling.md) for category, translation, and retry behavior.
## Cross-Version Invariants
- [Historical Document Transcription Design Intent](./invariant/intent.md)
- [Transcription Methodology](./invariant/transcription_methodology.md)
- [Error Handling](./invariant/error_handling.md)
- [Digital Evidence and AI Processing Provenance](./invariant/ai_evidence_and_provenance.md)
- [UI Style Guide](./invariant/ui_style_guide.md)
## Deployment and Operations
- [Production Runbook](production-runbook.md) for deploy, rollback, and recovery.
- [Backup and Restore](backup_restore.md) for backup configuration and restore procedure.
- [Data Migration](data_migration.md) for the SQLite to PostgreSQL migration path.
- [Cloudflare Tunnel and Access](cloudflare_tunnel_access.md) for remote exposure and access control.
## Baseline Statement
The current V6.1 baseline includes the architectural cleanup, person-schema redesign,
containerized PostgreSQL deployment, and the navigation, Document Detail, and worker-backed
maintenance refinements reflected across this canonical document set.
Use this `docs/*` canonical set for active design and implementation decisions.
Every canonical document above states this same baseline; `tests/test_meta_contract_guards.py`
fails if one of them falls behind. Forward-looking work is tracked in
[`roadmap_plan.md`](roadmap_plan.md) and is not part of the baseline.
+13 -9
View File
@@ -10,7 +10,7 @@ The application exists to preserve historical source material and produce useful
The application distinguishes five kinds of information: The application distinguishes five kinds of information:
1. **Source evidence**: the original uploaded media and the facts needed to identify and verify it. 1. **Source evidence**: the canonical stored media used for processing and the facts needed to identify and verify it.
2. **Execution specification**: the frozen instructions, parameters, source identity, and software context for one processing attempt. 2. **Execution specification**: the frozen instructions, parameters, source identity, and software context for one processing attempt.
3. **Transport evidence**: the response received at the application/provider boundary, including safe protocol metadata. 3. **Transport evidence**: the response received at the application/provider boundary, including safe protocol metadata.
4. **Normalized data**: selected fields extracted for search, display, accounting, and workflow behavior. 4. **Normalized data**: selected fields extracted for search, display, accounting, and workflow behavior.
@@ -20,12 +20,12 @@ Normalized data and derived artifacts never replace source or transport evidence
## 3. Core Invariants ## 3. Core Invariants
### 3.1 Original Source Preservation ### 3.1 Canonical Source Preservation
1. The original uploaded bytes are the primary evidence and must be preserved without transformation. 1. Each source must have one canonical stored byte stream used for processing and provenance.
2. Each source must have a cryptographic content digest, byte size, and stable identity. 2. Canonical storage may apply deterministic ingest normalization before persistence.
3. Processing may use transformed derivatives, but those derivatives must not overwrite the original. 3. Canonical stored bytes must have a cryptographic content digest, byte size, and stable identity.
4. A derivative used for processing must record its relationship to the original, its transformation, and its own digest. 4. Post-ingest processing derivatives must not overwrite canonical stored bytes.
5. Moving or renaming a stored file must not change its evidence identity. 5. Moving or renaming a stored file must not change its evidence identity.
### 3.2 Append-Only Processing History ### 3.2 Append-Only Processing History
@@ -45,7 +45,7 @@ Each execution must preserve enough information to understand what the applicati
3. Prompt asset name and content digest when a prompt asset is used. 3. Prompt asset name and content digest when a prompt asset is used.
4. Every explicitly supplied generation or processing parameter. 4. Every explicitly supplied generation or processing parameter.
5. Whether an optional parameter was explicitly set or omitted. 5. Whether an optional parameter was explicitly set or omitted.
6. Source and derivative digests, media type, dimensions or page geometry when known, and page identity. 6. Canonical source digest (and derivative digests when used), media type, dimensions or page geometry when known, and page identity.
7. A secret-safe representation of the request structure. 7. A secret-safe representation of the request structure.
8. Application, provider-adapter, and client-library versions sufficient to interpret the execution. 8. Application, provider-adapter, and client-library versions sufficient to interpret the execution.
@@ -108,7 +108,7 @@ Provenance supports explanation, comparison, and best-effort reproduction; it do
Identical requests may produce different results because of model updates, provider routing, nondeterministic computation, undocumented defaults, safety systems, or retired endpoints. The application must preserve whether a parameter was omitted rather than pretending to know the provider default used at that time. Identical requests may produce different results because of model updates, provider routing, nondeterministic computation, undocumented defaults, safety systems, or retired endpoints. The application must preserve whether a parameter was omitted rather than pretending to know the provider default used at that time.
Likewise, preserving a general vision-model response does not create OCR coordinates that were never returned. Future coordinate extraction remains possible because the original source evidence is preserved and can be processed again by a suitable system. Likewise, preserving a general vision-model response does not create OCR coordinates that were never returned. Future coordinate extraction remains possible because canonical source evidence is preserved and can be processed again by a suitable system.
## 5. Model Evaluation Policy ## 5. Model Evaluation Policy
@@ -123,11 +123,15 @@ Evaluation should:
5. Preserve the exact model, endpoint or route, parameters, prompt, source digest, and scoring method for every comparison. 5. Preserve the exact model, endpoint or route, parameters, prompt, source digest, and scoring method for every comparison.
6. Treat model rankings as corpus- and version-specific, not permanent declarations of a universal “best” model. 6. Treat model rankings as corpus- and version-specific, not permanent declarations of a universal “best” model.
The deterministic scorer for these comparisons lives in `src/transcription/benchmarking.py`; it is
retained as evaluation-policy infrastructure even though application runtime paths do not call it
directly.
Benchmark material containing family records remains private application data unless explicitly approved for publication. Benchmark material containing family records remains private application data unless explicitly approved for publication.
## 6. Ownership and Change Policy ## 6. Ownership and Change Policy
1. Versioned architecture, schema, scope, and implementation documents define how a release satisfies this invariant. 1. Canonical V6.1 architecture, schema, requirements, and error-policy documents define how current behavior satisfies this invariant.
2. Provider adapters own the capture of provider-boundary evidence. 2. Provider adapters own the capture of provider-boundary evidence.
3. Services own validation, persistence, retention, and export behavior. 3. Services own validation, persistence, retention, and export behavior.
4. UI pages may inspect evidence through service contracts but do not define evidence semantics. 4. UI pages may inspect evidence through service contracts but do not define evidence semantics.
+7 -7
View File
@@ -4,14 +4,14 @@
This guide defines non-negotiable UI styling rules for the transcription application. This guide defines non-negotiable UI styling rules for the transcription application.
The design system is token-first and class-driven: The design system is token-first and class-driven:
1. Theme tokens are defined in [src/transcription/ui/static/theme.css](src/transcription/ui/static/theme.css). 1. Theme tokens are defined in [src/transcription/ui/static/theme.css](../../src/transcription/ui/static/theme.css).
2. Python UI code composes semantic classes instead of inline color values. 2. Python UI code composes semantic classes instead of inline color values.
3. Pages and components should share a single visual language across Documents, Jobs, People, and Sources flows. 3. Pages and components should share a single visual language across Documents, Jobs, People, and Sources flows.
## 2. Source of Truth ## 2. Source of Truth
Use these files as the style authority: Use these files as the style authority:
1. [src/transcription/ui/static/theme.css](src/transcription/ui/static/theme.css) for color tokens, semantic utility classes, table styles, and viewer surfaces. 1. [src/transcription/ui/static/theme.css](../../src/transcription/ui/static/theme.css) for color tokens, semantic utility classes, table styles, and viewer surfaces.
2. [src/transcription/ui/theme.py](src/transcription/ui/theme.py) for runtime NiceGUI theme bridge and shared UI helpers. 2. [src/transcription/ui/theme.py](../../src/transcription/ui/theme.py) for runtime NiceGUI theme bridge and shared UI helpers.
If this document conflicts with implementation, update this document to match the code immediately after intentional style changes. If this document conflicts with implementation, update this document to match the code immediately after intentional style changes.
@@ -25,7 +25,7 @@ If this document conflicts with implementation, update this document to match th
## 4. Token System ## 4. Token System
### 4.1 Palette Tokens ### 4.1 Palette Tokens
Base palette variables live under :root in [src/transcription/ui/static/theme.css](src/transcription/ui/static/theme.css): Base palette variables live under :root in [src/transcription/ui/static/theme.css](../../src/transcription/ui/static/theme.css):
1. --palette-carbon-black: #1c2321 1. --palette-carbon-black: #1c2321
2. --palette-cool-steel: #7d98a1 2. --palette-cool-steel: #7d98a1
3. --palette-blue-slate: #5e6572 3. --palette-blue-slate: #5e6572
@@ -81,7 +81,7 @@ Semantic tokens currently include:
2. ui-table-header 2. ui-table-header
3. ui-table-body 3. ui-table-body
Use existing class combinations from [src/transcription/ui/components](src/transcription/ui/components) and [src/transcription/ui/pages](src/transcription/ui/pages) as reference implementations. Use existing class combinations from [src/transcription/ui/components](../../src/transcription/ui/components) and [src/transcription/ui/pages](../../src/transcription/ui/pages) as reference implementations.
## 6. Legacy Class Policy ## 6. Legacy Class Policy
Legacy `vibe-` presentation classes are prohibited. Use `ui-` semantic classes from `theme.css`. Legacy `vibe-` presentation classes are prohibited. Use `ui-` semantic classes from `theme.css`.
@@ -96,7 +96,7 @@ Legacy `vibe-` presentation classes are prohibited. Use `ui-` semantic classes f
## 8. Implementation Rules For Contributors ## 8. Implementation Rules For Contributors
1. Prefer composing existing semantic classes before creating new ones. 1. Prefer composing existing semantic classes before creating new ones.
2. If a new class is required, add it to [src/transcription/ui/static/theme.css](src/transcription/ui/static/theme.css) with a semantic name, then reuse it. 2. If a new class is required, add it to [src/transcription/ui/static/theme.css](../../src/transcription/ui/static/theme.css) with a semantic name, then reuse it.
3. Keep behavior ownership in Python and appearance ownership in CSS. 3. Keep behavior ownership in Python and appearance ownership in CSS.
4. Update UI tests that assert exact text or labels when intentional copy changes are made. 4. Update UI tests that assert exact text or labels when intentional copy changes are made.
5. Avoid introducing class churn unrelated to the feature being changed. 5. Avoid introducing class churn unrelated to the feature being changed.
@@ -104,7 +104,7 @@ Legacy `vibe-` presentation classes are prohibited. Use `ui-` semantic classes f
## 9. Verification Checklist ## 9. Verification Checklist
Before merging UI changes, verify: Before merging UI changes, verify:
1. No new inline hex colors were introduced in UI pages/components. 1. No new inline hex colors were introduced in UI pages/components.
2. New styles are token-backed and added to [src/transcription/ui/static/theme.css](src/transcription/ui/static/theme.css). 2. New styles are token-backed and added to [src/transcription/ui/static/theme.css](../../src/transcription/ui/static/theme.css).
3. Primary buttons, links, cards, and tables still render with consistent semantics. 3. Primary buttons, links, cards, and tables still render with consistent semantics.
4. Keyboard focus ring visibility is preserved. 4. Keyboard focus ring visibility is preserved.
5. Relevant UI and integration tests pass. 5. Relevant UI and integration tests pass.
+184
View File
@@ -0,0 +1,184 @@
# Production Runbook
This runbook is the operational checklist for releasing and monitoring the transcription system.
## 1. Pre-release gate checklist
1. Run the full suite: `uv run pytest`
2. Confirm contract guardrails are green:
- `uv run pytest tests/test_meta_contract_guards.py`
3. Confirm health endpoint includes worker liveness payload (`/healthz` returns `worker.state`).
4. Confirm required runtime settings are present in deployment environment:
- `OPENROUTER_API_KEY`
- `DATABASE__*`
- filesystem paths for data/logs/backups.
- `CLOUDFLARE_TUNNEL_TOKEN`
5. Confirm schema contract alignment is current:
- `src/transcription/db/models.py`
- `docs/schema.md`
## 2. Release execution steps
1. Deploy artifact/config to target environment.
- Production stack: `docker compose -f docker-compose.production.yml up -d --build`
- `Settings` loads from explicit `_env_file`, then `ENV_FILE`, then the repository-root `.env.production`; it does not resolve relative to the process working directory.
- For Runtime Settings writes in production, mount `.env.production` into the app container and set `RUNTIME_SETTINGS_ENV_FILE=/app/.env.production`.
- If deployment uses a non-default env-file location, set both `ENV_FILE` and `RUNTIME_SETTINGS_ENV_FILE` to that absolute path so startup reads and Settings-page writes stay aligned.
- For SQLite -> PostgreSQL cutover, run `uv run python tools/export_import_migration.py verify --source-db <sqlite-path-or-url> --target-db <postgres-url>` before switching runtime.
2. Validate service startup:
- `/healthz` responds `200`
- if `RUN_EMBEDDED_WORKER=true`, `worker.state` is `running`
- if `RUN_EMBEDDED_WORKER=false`, validate `worker` container is running in Compose
(worker healthcheck is intentionally disabled because it does not expose `/healthz`)
- validate `cloudflared` logs show active tunnel routes and no ingress errors
3. Execute one smoke workflow:
- create a document/job with at least one source
- verify terminal job outcome updates
- verify execution evidence row appended
4. Verify log flow:
- stdout aggregation receives events
- file logs are written under `./data/logs`
5. Create a fresh PostgreSQL backup after successful deployment:
- `sh deploy/backup/create_postgres_backup.sh` (creates DB dump plus uploads/prompts backups under `BACKUP_DIR`)
## 3. Rollback triggers and actions
### Trigger conditions
1. `/healthz` reports `worker.state=failed`
2. Repeated provider timeout/error spikes beyond normal baseline
3. Evidence write failures or DB persistence failures
### Actions
1. Roll back app artifact and config to previous release.
2. Restart service and re-check `/healthz`.
3. Re-run smoke workflow and confirm worker returns to `running`.
4. Preserve incident evidence:
- `./data/logs`
- relevant DB rows (`job`, `job_source`, `execution_attempt`)
5. If persistence regression is confirmed, restore the latest valid DB dump:
- `sh deploy/backup/restore_postgres_backup.sh <dump-file>`
- Synology media mirror and paired config snapshot (same timestamp) are restored automatically when present.
## 4. Post-release monitoring checklist
## First 24 hours
1. Monitor `/healthz` periodically for `worker.state`.
2. Track job terminal distribution (`transcribed`, `partial_success`, `failed`).
3. Sample timeout/error categories for abnormal increase.
4. Spot-check new `execution_attempt` records for append-only growth and timing metadata.
## First 72 hours
1. Re-check error/timeout trend versus 24h baseline.
2. Verify no recurring worker-failed states.
3. Verify storage growth and rotation behavior under `./data/logs`.
4. Confirm incident response notes are captured for any production anomalies.
## 5. Operator playbook for common incidents
### Worker failed
1. Check `/healthz` payload (`error_id`, `error_category`).
2. Locate matching error in logs.
3. If non-transient defect persists, roll back.
### Provider timeout spike
1. Confirm provider reachability and rate limits.
2. Review timeout frequency and impacted job volume.
3. If sustained, execute rollback criteria and notify stakeholders.
### Partial-success increase
1. Inspect affected `job_source` and `execution_attempt` records.
2. Confirm failures are category-aligned (`external`/`timeout`/`internal`).
3. Triage whether issue is source quality, provider, or runtime regression.
### Cloudflare ingress/access failure
1. Check `cloudflared` container logs for ingress parse, DNS, or auth failures.
2. Confirm `deploy/cloudflared/config.yml` hostname mappings are correct.
3. Confirm `CLOUDFLARE_TUNNEL_TOKEN` in `.env.production` matches the tunnel configured in Cloudflare.
4. Confirm Cloudflare Access app policy includes the intended identity/group for that hostname.
5. If logs show SRV/DNS failures via `127.0.0.11`, use the compose-defined resolver override
(`dns: 1.1.1.1, 1.0.0.1`) and ensure outbound TCP 443 is allowed.
### Backup or restore failure
1. Verify `postgres` container is healthy and accepting connections.
2. Confirm dump file exists and is non-zero size.
3. Re-run backup/restore scripts with explicit `ENV_FILE` and `COMPOSE_FILE` if using non-default paths.
4. If direct Synology copy fails, keep local backup and resolve mount/network before next backup cycle.
- For LXC setups, use `deploy/backup/mount_synology_cifs.example.sh` as the persistent mount template.
## 6. Dependency upgrade policy
Dependencies are declared in `pyproject.toml` and resolved through the committed
`uv.lock`. The lockfile guarantees reproducible installs; the version specifiers
control what a deliberate `uv lock --upgrade` is allowed to move.
### NiceGUI is pinned exactly (`nicegui==3.13.0`)
1. **Rationale.** NiceGUI bundles Quasar and Vue. Minor releases change component
props, slots, and styling, which surfaces as visual and interaction regressions
rather than import or type errors. The UI suite under `tests/ui/` asserts
structure and behavior, not rendered appearance, so a NiceGUI bump can pass the
full test suite and still degrade the interface.
2. **Scope of risk.** All NiceGUI usage is confined to `src/transcription/ui/` and
uses only the public `nicegui.ui` and `nicegui.events` surfaces. The coupling is
shallow, so the pin is about release stability, not about unpicking deep
framework entanglement.
3. **Current stance.** Hold the exact pin through release stabilization. Do not
widen it as incidental cleanup, and do not let automated dependency updates move
it. This includes forgoing patch releases, which is the accepted cost.
4. **Revisiting.** Treat a NiceGUI upgrade as scheduled work with its own change
window: bump the pin deliberately, run `uv run pytest -m "not external"`, then
manually verify each page contract in `docs/ui/pages/` before accepting.
### All other dependencies
Declared with `>=` floors and moved by explicit `uv lock --upgrade`. Verify with
`uv run ruff check .`, `uv run ty check`, and `uv run pytest -q -m "not external"`
before committing a changed lockfile.
## 7. Type-check suppression policy
`uv run ty check` is a blocking pre-commit gate. Suppressions are allowed only for
proven SQLAlchemy descriptor false positives where runtime behavior is correct and
the checker cannot represent the descriptor protocol at that call site.
Every suppression must be:
1. **Targeted** to a single rule (for example `# ty: ignore[unresolved-attribute]`).
2. **Inline** on the expression it suppresses (not file-wide).
3. Followed by a **one-line rationale** stating it is a SQLAlchemy descriptor false positive.
Do not use broad or rationale-free suppressions. If a diagnostic is not a known
false positive, fix the code instead of suppressing it.
## 8. Worker shutdown budget
Worker shutdown waits for at most:
`WORKER_PROVIDER_TIMEOUT_SECONDS + WORKER_SHUTDOWN_GRACE_SECONDS`
`WORKER_PROVIDER_TIMEOUT_SECONDS` covers an in-flight provider call, and
`WORKER_SHUTDOWN_GRACE_SECONDS` is extra time for the loop to persist outcomes
and exit cleanly after the call returns.
Set the container or service termination grace period **above this total**
budget. If termination grace is shorter, the process may be killed before
terminal status and evidence writes are finalized.
## 9. Horizontal scaling precondition
Multiple worker replicas can race on execution-attempt numbering for the same
`(job_id, source_id)` pair. The runtime now retries boundedly on unique-key
conflicts (`uq_execution_attempt_number`) and surfaces a conflict-domain error
if retries are exhausted.
Do not deploy additional worker replicas unless this conflict-retry path and its
tests are present and green in the target build.
+118
View File
@@ -0,0 +1,118 @@
# System Requirements (Current Baseline: V6.1)
These requirements define the active V6.1 contract and align to current implementation.
Requirement IDs encode the baseline that introduced them (`REQ-4-*` from V4, `REQ-6-*` from V6) and
are stable. Never renumber an existing ID; retire it explicitly instead.
## Functional Requirements
### Domain and Record Management
- **REQ-4-001 Document Registry:** The system must create and update `Document` records with title, type, date metadata, optional location, optional archive identifier, and optional notes.
- **REQ-4-002 Source Registry:** The system must create and update `Source` records linked to exactly one `Document`.
- **REQ-4-003 People Registry:** The system must create and update `Person` records, support many-to-many links to `Document` with role, and support many-to-many Person tagging via the shared Tag registry.
- **REQ-4-004 Registry Semantics:** Document types and person roles must support optional immutable semantic keys and hard-delete only when unreferenced.
### Job and Workflow Behavior
- **REQ-4-010 Job Creation:** The system must create `Job` records from uploaded sources and from retranscription of existing sources.
- **REQ-4-011 Prompt Snapshotting:** Job creation must persist effective prompt and runtime settings as immutable per-job snapshots.
- **REQ-4-012 Queue Membership:** Each `(job, source)` pair must be represented by one `JobSource` row.
- **REQ-4-013 Job Status Lifecycle:** `Job.status` must use one of `queued`, `processing`, `transcribed`, `partial_success`, `failed`.
- **REQ-4-014 JobSource Status Lifecycle:** `JobSource.status` must use one of `pending`, `transcribed`, `failed`, `cancelled`.
- **REQ-4-015 Terminal Job Resolution:** Job terminal status must derive from page outcomes as `transcribed`, `partial_success`, or `failed`.
- **REQ-4-016 Cancellation Semantics:** Job cancellation must set remaining `pending` page entries to `cancelled`.
### Transcription and Evidence
- **REQ-4-020 Attempt Evidence:** Each provider call must emit one append-only `ExecutionAttempt` record.
- **REQ-4-021 Attempt Payload:** `ExecutionAttempt` must retain request manifest/hash, outcome, timing, model/provider fields, and error details when present.
- **REQ-4-022 Transport Evidence:** Provider response evidence must be attached to the attempt when a response is available.
- **REQ-4-023 Source Projection Rule:** `Source.raw_transcription` is a projection chosen from attempt outcomes and can be repointed by explicit promotion.
- **REQ-4-024 Candidate Visibility:** UI must expose candidate attempts with metadata needed for comparative review and selection.
### Media and Access
- **REQ-4-030 Ingest Canonicalization:** Stored source bytes may be normalized at ingest (for example orientation correction); stored bytes are the canonical processing source.
- **REQ-4-031 Path Safety:** Client-facing media URLs must be generated from controlled application paths only.
- **REQ-4-032 Print Media Validation:** Print/export source media must be served through record-validated API routes.
### Error and UX Contracts
- **REQ-4-040 Error Envelope:** Service/API errors must map to structured, user-safe error categories and messages.
- **REQ-4-041 Partial Failure Visibility:** Mixed page outcomes must be visible at job and page level.
- **REQ-4-042 Retry Support:** Failed and cancelled pages must support targeted retranscription without requiring full document recreation.
### Person Imagery
- **REQ-6-001 Photo Records:** The system must store reusable `Photo` records that are either owned by a `Person` or unowned for homepage gallery use.
- **REQ-6-002 Primary Photo:** At most one photo per owning `Person` may be marked `is_primary`; setting a new primary must clear the previous one.
- **REQ-6-003 Primary Reassignment:** Deleting an owner's primary photo must promote a remaining photo of that owner rather than leaving the owner without a primary.
### Operational Maintenance
- **REQ-6-010 Queued Maintenance Runs:** Settings-initiated maintenance must persist a `MaintenanceRun` and execute in the worker, not inline in the request that started it.
- **REQ-6-011 Maintenance Run Types:** `MaintenanceRun.job_type` must use one of `backup`, `storage_reconciliation`, `gedcom_import`.
- **REQ-6-012 Maintenance Status Lifecycle:** `MaintenanceRun.status` must use one of `queued`, `processing`, `succeeded`, `failed`.
- **REQ-6-013 Single Claim:** A queued run must be claimed by at most one worker, using a conditional status update rather than read-then-write.
- **REQ-6-014 Run History:** Completed runs must retain status, timing, summary, log reference, and error detail, and expose the log for viewing and download.
- **REQ-6-015 GEDCOM Upload Import:** Settings must support manual `.ged` upload and queue-backed import into genealogy tables.
- **REQ-6-016 GEDCOM Idempotent Upsert:** GEDCOM import must upsert `GenealogyPerson` and `GenealogyFamily` by FamilySearch IDs and avoid duplicate imported citations on re-run.
### Deployment and Runtime Configuration
- **REQ-6-020 Production Persistence:** Production must run against PostgreSQL; SQLite remains supported for local development and tests.
- **REQ-6-021 Split Worker Deployment:** Production must support running the worker as its own process with the app started at `RUN_EMBEDDED_WORKER=false`.
- **REQ-6-022 Runtime Settings Persistence:** Runtime settings edits must persist to the mounted production environment file and survive container restart.
- **REQ-6-023 Configuration Contract Sync:** `.env.production.example` must stay synchronized with `Settings` keys and production-safe defaults.
- **REQ-6-024 Health Reporting:** The deployed stack must report app and worker health through `/healthz`.
- **REQ-6-025 Backup and Restore:** Database and media/config backups must be produced on a host-visible path with a tested restore procedure.
## Non-Functional Requirements
- **REQ-4-100 Boundary Integrity:** UI pages/components must not access persistence directly and must call service APIs.
- **REQ-4-101 Service Ownership:** Aggregate writes must occur in owning service/workflow modules, not in UI handlers.
- **REQ-4-102 Deterministic Loading:** ORM relationship reads in service/UI code must use explicit eager loading compatible with `lazy="raise"`.
- **REQ-4-103 Async Safety:** Long-running provider calls must not block UI event handlers directly.
- **REQ-4-104 Evidence Durability:** Attempt evidence must survive process restart once the transaction commits.
- **REQ-4-105 Test Guardrails:** Architecture boundary tests must remain in place for services and UI boundaries.
## Requirement Interpretation Notes
### Status and lifecycle semantics
- `REQ-4-013` and `REQ-4-015` intentionally bind success to `transcribed`, not a generic `completed`, so docs, tests, and runtime transitions stay consistent.
- `REQ-4-016` and `REQ-4-042` distinguish cancellation from failure at page level (`cancelled` vs `failed`) while still allowing targeted retranscription.
### Evidence semantics
- `REQ-4-020` through `REQ-4-024` separate authoritative history (`ExecutionAttempt`) from operational projection (`Source.raw_transcription`).
- This supports immutable provenance while allowing explicit candidate promotion for operator workflows.
### Boundary and loading semantics
- `REQ-4-100` and `REQ-4-101` codify aggregate/service ownership and keep UI out of persistence concerns.
- `REQ-4-102` exists to enforce deterministic query shape under `lazy="raise"` and avoid hidden data access in rendering callbacks.
## Verification Anchors
- Service boundary enforcement: `tests/test_service_boundaries.py`
- UI boundary enforcement: `tests/test_ui_boundaries.py`
- Job lifecycle reliability and terminal status behavior: `tests/services/test_workflows_reliability.py`
- Evidence append-only and projection behavior: `tests/services/test_store.py`, `tests/services/test_transcription_service.py`
- Person photo ownership and primary selection: `tests/services/test_photo_service.py`
- Maintenance run lifecycle and worker execution: `tests/services/test_maintenance_service.py`
- Runtime settings persistence: `tests/ui/test_runtime_settings_store.py`, `tests/services/test_settings_services.py`
- Configuration contract synchronization: `tests/test_meta_contract_guards.py`
- Deployment health reporting: `tests/api/test_health.py`
## Traceability Notes
- Source of truth for status enums:
- `src/transcription/db/models.py`
- Source of truth for workflow transitions:
- `src/transcription/services/workflows.py`
- `src/transcription/services/jobs.py`
- Source of truth for attempt evidence writes:
- `src/transcription/services/sources.py`
+488
View File
@@ -0,0 +1,488 @@
# Architecture & Code Review Report
**Repository Target:** `transcription/`
**Target Stack:** Python 3.12+ | FastAPI | NiceGUI | SQLModel/SQLAlchemy | Pydantic V2 | asyncio | OpenRouter
**Review date:** 2026-08-23
**Governing procedure:** `.github/skills/python-code-reviewer/skill.md`
**Escalations applied:** `.github/skills/evidence-provenance-auditor/skill.md`, `.github/skills/test-effectiveness-auditor/skill.md`
**Scope:** 77 Python modules / ~13k LOC under `src/transcription`, 57 test files (377 collected non-external tests), 23 documents under `docs/`, 9 active rule files.
> **Status: closed.** Every finding below was remediated in the phases following this
> review. This document is retained as a record of the reasoning, **not** as a list of
> open work, and it is not canonical authority.
>
> Two recommendations were wrong on contact and were corrected during implementation:
> the HIGH-03 fix as written would have stripped root-cause data from `ExecutionAttempt`
> provenance, and the HIGH-01 fix needed to preserve per-page durability that the report
> did not mention. Where this text and the current code or guard tests disagree, the code
> and tests are correct.
### Verification commands and outcomes
| Command | Outcome |
| :--- | :--- |
| `uv run ruff check .` | **Pass**`All checks passed!` |
| `uv run pytest -q -m "not external"` | **Pass** — 377 passed |
| `uv run ty check` | **10 diagnostics** — all SQLModel/SQLAlchemy column-descriptor false positives (`services/photos.py` ×8, `tests/test_storage_reconciliation.py` ×2). Advisory only; no suppression strategy exists. |
---
## 1. Executive Summary
- **Overall health is good.** The codebase has genuine architectural discipline: layered `ui → services → db`, a single Pydantic-V2 settings source, an atomic compare-and-swap job claim, append-only evidence history, and eleven deterministic guard tests that enforce structural rules rather than describing them.
- **No Critical findings.** The highest-risk category for this domain — secret leakage into stored provenance — was explicitly audited and **passes**: request headers are never persisted, response headers use an allowlist, and the API key is `SecretStr` end-to-end.
- **The top risk is a transaction-atomicity violation on the worker hot path.** Page evidence and terminal job status commit in two separate transactions (`workflows.py:549-598`), directly contradicting `services.instructions.md`. A crash between them leaves a transcript persisted against a job stuck in `PROCESSING`.
- **That violation is invisible to the test suite.** The test-effectiveness audit confirms no test can fail on a split commit — the pipeline tests assert the happy-path end state, which passes either way. The invariant is documented and steered but *not enforced*.
- **Stale-job recovery is startup-only** (`app.py:79`), with a 30-second staleness threshold. A job orphaned shortly before a fast restart is not recovered and remains `PROCESSING` indefinitely, because the worker only claims `QUEUED` rows.
- **The mandated error-presentation boundary is bypassed at 8 sites.** `home_page.py` and `people_page.py` hand-roll `ui.notify(str(exc), ...)`, discarding the `error_id`, category, and suggestion that `error_presenter.show_error` provides. `people_page.py` imports the correct helpers and still bypasses them.
- **User-facing output can leak filesystem paths.** `classify_unexpected_error` (`errors.py:94`) interpolates the raw exception into a message rendered in the UI; a SQLAlchemy `OperationalError` embeds the database file path. This contradicts an explicit rule in `error-handling.instructions.md`.
- **The retry gate ignores error category** (`workflows.py:185`), so non-retriable faults would be requeued. Currently latent because `worker_max_retries` defaults to `0`.
- **Highest-leverage work is enforcement, not refactoring.** Two atomicity tests, a `ty` suppression strategy that lets the pre-commit hook become blocking, and `ruff format --check` in the gate would convert three documented-but-unenforced invariants into deterministic ones.
---
## 2. Executive Architecture Assessment
**Verdict: architecturally sound with a concentrated reliability gap in the worker's commit boundary.**
Domain cohesion is strong. The `services/` layer owns transactions and business rules, `ui/` owns presentation, `db/` owns schema, and `providers/` isolates the OpenRouter adapter behind a `TranscriptionProvider` protocol. Dependency direction is correct and — unusually — *mechanically enforced*: `test_service_boundaries.py` AST-scans for service-to-service imports and `test_ui_boundaries.py` scans pages/components for persistence access. Provider details do not leak upward; `workflows.py` imports only the abstract `providers` types, never `openrouter`.
The evidence/provenance model is the strongest part of the system. `ExecutionAttempt` is genuinely append-only, retries append rather than rewrite, projection writes onto `JobSource` are clearly distinguished from history mutation, and all 14 provenance-auditor invariant checks pass.
**Top systemic risks:**
1. **Split commit boundary on the worker path (High).** Evidence durability and job terminal status are two transactions. This is the one place where the architecture's own written contract is contradicted by the implementation, on the hottest path in the system.
2. **Recovery is a startup-only, time-thresholded sweep (Medium).** There is no runtime reconciliation, so the self-healing property depends on restart cadence rather than on a bounded interval.
3. **Enforcement coverage has known holes (Medium).** Atomicity, error-presenter usage, and formatting are all documented rules with no deterministic test. The repo's own strength — routing invariants into tests — has not been applied to these three.
4. **Leaky transaction ownership (Medium).** `workflows.py` reaches into `services.jobs._session_scope()` and `services.sources._session_scope()` — private members of two different services — to open transactions. Session ownership is ambiguous exactly where it most needs to be explicit.
5. **A 10-diagnostic type-checker baseline with no suppression policy (Low).** The signal is currently ignorable, which means a real regression would blend into the noise.
---
## 3. Findings by Severity
### Critical Severity
**None identified.**
The secret-leakage check — the only plausible Critical for this system — passes explicitly. `OpenRouterProvider` stores an allowlisted subset of *response* headers only (`providers/evidence.py:130-134`, `SAFE_RESPONSE_HEADERS`); request headers containing `Authorization` are never captured into `TransportEvidence`; and the key is held as `SecretStr` from `config.py` through to the client. Append-only evidence history is likewise intact and test-enforced.
---
### High Severity
#### [HIGH-01] Page evidence and terminal job status commit in separate transactions
- **Location:** `src/transcription/services/workflows.py:549-565` (`_finalize_batch_outcome`), `src/transcription/services/workflows.py:584-598` (`_persist_page_outcome`)
- **Problem & Consequence:** `.github/instructions/services.instructions.md` states: *"Never commit transcript updates separately from the paired terminal/retry job status change."* The implementation does exactly that. `_persist_page_outcome` opens its own scope and commits page evidence (line 592-594); `_finalize_batch_outcome` later opens a *second* scope and commits the terminal `JobStatus` (line 558-560). For a single-page job these are two transactions with a window between them. A process crash, container eviction, or unhandled error in that window persists the transcript while the job remains `PROCESSING`. Because the worker only claims `QUEUED` rows, that job is not reprocessed; it is recoverable only by the startup sweep, and only if it has aged past the staleness threshold (see MED-01). The user sees a job that never completes despite the transcription having succeeded and been billed.
This is a deliberate design tension, not an oversight: `_persist_page_outcome_durably` (line 568-581) wraps the page write in `asyncio.shield` precisely so per-page evidence survives cancellation mid-batch. That goal is correct for *multi*-page jobs. The defect is that the single-page and final-page cases inherit the split unnecessarily.
- **Recommendation:** Keep per-page durability for intermediate pages, but commit the final page outcome and the terminal status in one transaction.
```python
# Before — two scopes, two commits
await _persist_page_outcome_durably(job=job, services=services, page=page, session=None)
...
await _finalize_batch_outcome(job=job, services=services, status=status, session=None)
# After — final page and terminal status share one transaction
async with services.jobs.session_scope() as tx:
for page in intermediate_pages:
await _persist_page_outcome_durably(job=job, services=services, page=page, session=None)
await _write_page_outcome(job=job, services=services, page=final_page, session=tx)
await services.jobs.mark_job_status(job.id, status, session=tx)
await tx.commit()
```
Pair this with the atomicity test in HIGH-04 so the boundary cannot silently regress.
- **Effort:** M
---
#### [HIGH-02] Mandated error-presentation boundary bypassed at 8 sites
- **Location:** `src/transcription/ui/pages/home_page.py:212,220,228,255`; `src/transcription/ui/pages/people_page.py:265,321,330,339`
- **Problem & Consequence:** `.github/instructions/ui.instructions.md:42` requires all user-facing error display to route through `components/error_presenter.py`. Seven of nine pages comply. These two hand-roll `ui.notify(str(exc), type="negative")`. The consequence is not cosmetic: `show_error` (`error_presenter.py:52-67`) surfaces the correlation `error_id`, the canonical error category, and the actionable `suggestion` field. Bypassing it means a user hitting a failure on the home or people page gets a bare exception string with **no error reference to report**, making these two pages unsupportable in production — precisely the pages most likely to be a user's entry point.
`people_page.py` already imports `run_ui_action` and `show_error` at lines 28-29 and uses them elsewhere in the same module, so the bypass is inconsistency rather than missing infrastructure.
- **Recommendation:** Replace each site with the canonical helper. The unused `summarize_error` helper in `error_presenter.py` (currently a retained orphan — see LOW-07) is the natural fit where a compact string is genuinely needed.
```python
# Before
except AppError as exc:
ui.notify(str(exc), type="negative")
# After
except AppError as exc:
show_error(exc)
```
Then close the hole permanently by extending `tests/test_ui_boundaries.py` with an AST check that no module under `PAGES_DIR` calls `ui.notify(...)` with `type="negative"`.
- **Effort:** S
---
#### [HIGH-03] Unexpected-error path leaks filesystem paths into user-facing output
- **Location:** `src/transcription/errors.py:91-98` (line 94), rendered via `src/transcription/ui/components/error_presenter.py:52-67`
- **Problem & Consequence:** `classify_unexpected_error` builds `f"Unexpected error during {operation}: {exc}"` and stores it as `AppError.message`. `show_error` renders `error.message` directly to the user. Any exception whose `str()` contains infrastructure detail is therefore displayed verbatim — a SQLAlchemy `OperationalError` embeds the absolute SQLite database path, and an `OSError` from the media layer embeds the storage root. `.github/instructions/error-handling.instructions.md:74` states: *"Never leak … local filesystem paths in user-facing output."* This is the generic catch-all path, so it applies to every unanticipated failure across the application.
- **Recommendation:** Split the diagnostic detail from the user-facing message. Log the full exception with the `error_id` as the correlation key; show the user a stable message plus that id.
```python
# Before
return AppError(
f"Unexpected error during {operation}: {exc}",
category=ErrorCategory.INTERNAL_UNEXPECTED,
...
)
# After
error = AppError(
f"Unexpected error during {operation}.",
category=ErrorCategory.INTERNAL_UNEXPECTED,
suggestion="Retry once. If it persists, report the error reference id.",
retriable=False,
)
logger.exception("error_id=%s operation=%s", error.error_id, operation)
return error
```
Add a case to `tests/ui/test_error_presenter.py` asserting that a raised `OperationalError` carrying a path does not surface that path in the rendered message.
- **Effort:** S
---
#### [HIGH-04] Transaction-atomicity invariants have no enforcing test
- **Location:** Contract at `.github/instructions/services.instructions.md` §"Workflow Transaction Boundaries"; gap confirmed across `tests/integration/test_pipeline_flow.py:66-160` and `tests/services/test_job_service.py:41-59`
- **Problem & Consequence:** The test-effectiveness audit establishes that **neither** Transaction B (transcript + `TRANSCRIBED`) nor Transaction C (retry: `error_detail` + `retry_count` + `QUEUED`) is enforced. The existing pipeline test asserts the final state after a successful run — which passes identically whether the writes shared one commit or used two. To fail on a split-commit regression a test must inject a fault *between* the writes; no such test exists.
The consequence is that HIGH-01 shipped undetected and any future refactor of `advance_job` can reintroduce it just as silently. This is a *governance* failure rather than a code defect: the repo's stated model is that hard rules belong in deterministic tests, and this rule is the most consequential one that never made the transition.
- **Recommendation:** Add `tests/integration/test_pipeline_atomicity.py` with two tests that patch the session to raise after `flush()` but before `commit()`, then assert that *neither* side of the pair is visible in a fresh session. These tests should **fail against the current implementation** and pass once HIGH-01 is fixed — write them first.
- **Effort:** M
---
### Medium Severity
#### [MED-01] Stale-job recovery runs only at startup, behind a 30-second threshold
- **Location:** `src/transcription/app.py:71-81` (`_recover_stale_processing_jobs`), sole caller at `app.py:79` inside `_lifespan`
- **Problem & Consequence:** `requeue_stale_processing_jobs` has exactly one call site, in the lifespan startup handler. There is no runtime re-check. The staleness predicate is `updated_at < now - worker_provider_timeout_seconds` (default **30.0s**, `config.py:116`). A job orphaned less than 30 seconds before a fast container restart therefore fails the predicate at the only moment recovery is attempted, and stays `PROCESSING` forever — the worker claims only `QUEUED` rows. It self-heals only on some *later, unrelated* restart. In a frequently-redeployed environment, restarts are exactly when orphans are created, so the recovery window is systematically misaligned with the failure it exists to handle.
- **Recommendation:** Move the sweep onto a periodic task in the worker loop (e.g. every `max(30, provider_timeout * 2)` seconds) in addition to the startup call, and derive the threshold from a dedicated `worker_stale_job_seconds` setting rather than reusing the provider timeout, so the two can be tuned independently.
- **Effort:** M
---
#### [MED-02] Retry gate ignores `error_category`, so non-retriable failures would be requeued
- **Location:** `src/transcription/services/workflows.py:184-194`
- **Problem & Consequence:** The `JobStatus.FAILED` branch gates solely on `job.retry_count < settings.worker_max_retries`. It does not consult `error_category` or the `AppError.retriable` flag. `.github/instructions/error-handling.instructions.md` classifies `validation`, `not_found`, and `conflict` as non-retriable; under this gate a malformed source or a missing record would be retried to exhaustion, consuming provider quota on calls that cannot succeed and delaying the terminal failure the user needs to see. There is also no backoff — retries requeue immediately.
Currently **latent**: `worker_max_retries` defaults to `0` (`config.py:113`) and is commented out in `.env`, so the branch always falls through to the max-retries log. It becomes live the moment anyone enables retries.
- **Recommendation:** Gate on retriability *and* count, and add exponential backoff before requeue.
```python
case JobStatus.FAILED:
if job.error_category in NON_RETRIABLE_CATEGORIES:
logger.error("Job %s failed non-retriably (%s).", job.id, job.error_category)
return
if job.retry_count < settings.worker_max_retries:
...
```
Cover with a test that a `validation`-category failure is not requeued even when `worker_max_retries > 0`.
- **Effort:** S
---
#### [MED-03] `IntegrityError` on the attempt-number flush is uncaught, risking evidence loss
- **Location:** `src/transcription/services/sources.py:540-546` (attempt-number computation), `sources.py:587` (unguarded `flush()`)
- **Problem & Consequence:** `attempt_number` is derived read-then-write as `MAX(attempt_number) + 1`, and `uq_execution_attempt_number` enforces uniqueness (`db/models.py:507`, documented at `docs/schema.md:273`). The sibling `JobSource` insert *does* catch `IntegrityError` (`sources.py:531-534`), but the `ExecutionAttempt` flush at line 587 does not. Two concurrent attempt writes for the same job source would raise an unhandled `IntegrityError` and lose an evidence row — the one class of data this system exists to preserve. Not currently reachable: the worker is single-instance and processes sources sequentially. It becomes reachable the moment a second worker replica is deployed.
- **Recommendation:** Mirror the `JobSource` handling — catch `IntegrityError`, recompute `MAX(attempt_number) + 1`, and retry the insert a bounded number of times, raising a domain error on exhaustion. Note this constraint as a horizontal-scaling precondition in `docs/production-runbook.md`.
- **Effort:** M
---
#### [MED-04] Shutdown timeout is shorter than the provider timeout
- **Location:** `src/transcription/worker.py:146` (`asyncio.wait_for(worker_task, timeout=2.0)`); provider timeout at `config.py:116` (default 30.0s)
- **Problem & Consequence:** Graceful shutdown waits 2 seconds for the worker task, but the stop event is only checked *between* jobs and an in-flight provider call may run for up to 30 seconds. Any shutdown during a provider call therefore cancels mid-flight. Combined with HIGH-01's split commit, a cancellation that lands between the evidence commit and the status commit produces exactly the stuck-`PROCESSING` state described there — so this finding materially raises HIGH-01's probability rather than being independent of it.
- **Recommendation:** Derive the shutdown budget from the provider timeout (`worker_provider_timeout_seconds + small_grace`) instead of hardcoding `2.0`, and ensure the container's termination grace period exceeds it. Document both in `docs/production-runbook.md`.
- **Effort:** S
---
#### [MED-05] `workflows.py` reaches into two services' private `_session_scope`
- **Location:** `src/transcription/services/workflows.py:558` (`services.jobs._session_scope()`), `workflows.py:592` (`services.sources._session_scope()`)
- **Problem & Consequence:** The orchestration module opens transactions by calling a private member on two different service objects. This is the concrete mechanism behind HIGH-01: because transaction ownership is expressed through a private back-door rather than a declared boundary, nothing in the design makes it obvious that two scopes are being opened for one logical unit of work. It also couples `workflows.py` to a service implementation detail that `test_service_boundaries.py` cannot see (it checks imports, not attribute access).
- **Recommendation:** Promote a single explicit transaction entry point — a `session_scope()` on `ServiceBundle`, or a module-level `unit_of_work(services)` helper — and make `workflows.py` use only that. Extend `test_service_boundaries.py` with an AST check forbidding `_session_scope` attribute access outside the owning service module.
- **Effort:** M
---
### Low Severity
#### [LOW-01] `hashlib.sha256` over full file bytes runs on the event loop
- **Location:** `src/transcription/services/store.py:401`
- **Problem & Consequence:** Digest computation is CPU-bound and synchronous inside an `async def`. For large uploads this blocks the loop, stalling both the NiceGUI UI and the worker. Every sibling I/O path in the codebase correctly uses `asyncio.to_thread` (`media_storage.py:43`, `normalization.py:117`, `photos.py:176`, `sources.py:740,753`), so this is an isolated deviation.
- **Recommendation:** `digest = await asyncio.to_thread(lambda: hashlib.sha256(file_bytes).hexdigest())`.
- **Effort:** S
#### [LOW-02] `homepage_store.py` performs synchronous file I/O from async callers
- **Location:** `src/transcription/ui/homepage_store.py:25,32`; called from `src/transcription/ui/pages/home_page.py:170`
- **Problem & Consequence:** Same class as LOW-01 — reads/writes the homepage JSON directly rather than via `asyncio.to_thread`. Impact is small (a tiny file), but it is a second deviation from an otherwise universal convention.
- **Recommendation:** Wrap both calls in `asyncio.to_thread`.
- **Effort:** S
#### [LOW-03] Worker poll interval is hardcoded outside `Settings`
- **Location:** `src/transcription/app.py:62` (`poll_interval_seconds=1.0`)
- **Problem & Consequence:** The single operational knob controlling worker latency-vs-load cannot be tuned without a code change, contradicting the otherwise-clean rule that all configuration lives in `config.py` (zero `os.getenv` calls exist outside it).
- **Recommendation:** Add `worker_poll_interval_seconds: float = 1.0` to `Settings` and read it at the call site.
- **Effort:** S
#### [LOW-04] `_build_request_manifest` returns `None` silently, producing incomplete evidence
- **Location:** `src/transcription/providers/openrouter.py:347`
- **Problem & Consequence:** When `source_reference is None` the manifest is skipped with no log line. The attempt is still recorded but its provenance is quietly incomplete, and there is no signal that it happened — the failure mode is undetectable after the fact.
- **Recommendation:** Log at `warning` with the job/source identifiers before returning `None`, so incomplete provenance is at least attributable.
- **Effort:** S
#### [LOW-05] Ten `ty` diagnostics with no suppression strategy
- **Location:** `src/transcription/services/photos.py` (8), `tests/test_storage_reconciliation.py` (2)
- **Problem & Consequence:** All ten are SQLModel/SQLAlchemy false positives — column descriptors are typed as their Python value type (`UUID`, `datetime`, `bool`), so `.is_()`, `.asc()`, `func.count()`, and `group_by()` appear invalid. Because there is no suppression policy, the pre-commit hook must run `ty` in advisory mode, which means a *genuine* new type error would print alongside the known ten and block nothing.
- **Recommendation:** Add targeted `# ty: ignore[...]` comments with a one-line rationale at each of the ten sites, then flip the pre-commit hook to blocking. This converts a permanently-ignored signal into a real gate.
- **Effort:** M
#### [LOW-06] `ruff format` is not enforced; 35 files have drifted
- **Location:** `.pre-commit-config.yaml`, `ruff.toml`
- **Problem & Consequence:** `ruff check` is blocking but `ruff format --check` is absent from the gate, so formatting drift accumulates silently and inflates unrelated diffs whenever anyone does run the formatter.
- **Recommendation:** Run `uv run ruff format .` once as a single isolated commit, then add `ruff format --check` to the pre-commit gate.
- **Effort:** S
#### [LOW-07] Four retained orphans, all recorded as "uncertain — follow-up"
- **Location:** `tests/test_orphan_sweep.py:33-52` (`KNOWN_ORPHANS`): `BenchmarkManifest`, `dispose_all_engines`, `refresh_engine`, `summarize_error`
- **Problem & Consequence:** Every entry carries the weakest possible justification. `summarize_error` is the notable one: it is an unused helper in `error_presenter.py` *while two pages hand-roll error display* (HIGH-02) — the orphan and the boundary violation are the same problem viewed from two directions. `dispose_all_engines` / `refresh_engine` are plausibly test-support utilities and should be classified as such rather than left uncertain.
- **Recommendation:** Resolve each to a definite outcome — `summarize_error` becomes used by the HIGH-02 fix; classify the engine helpers as test-support or delete them; decide on `BenchmarkManifest`.
- **Effort:** S
#### [LOW-08] Orphan sweep only scans module-level public definitions
- **Location:** `tests/test_orphan_sweep.py`
- **Problem & Consequence:** Methods and private functions are out of scope, so dead code inside classes — the most common kind in a service-oriented codebase — is structurally invisible to the sweep.
- **Recommendation:** Extend the AST walk to public methods on service classes, seeding `KNOWN_ORPHANS` with the current result set to keep the change non-breaking.
- **Effort:** M
#### [LOW-09] f-string interpolation in logging calls
- **Location:** `src/transcription/services/workflows.py:193` and similar sites
- **Problem & Consequence:** `logger.error(f"Job {job.id} has failed...")` formats eagerly regardless of level and prevents structured-logging backends from grouping by template. Ruff's `flake8-logging-format` (`G`) rules are not enabled, so this is unenforced.
- **Recommendation:** Use `logger.error("Job %s has failed and reached max retries.", job.id)` and enable ruff rule set `G`.
- **Effort:** S
#### [LOW-10] Low-signal and always-true assertions in the test suite
- **Location:** `tests/test_traceability.py:54-57`; `tests/integration/test_pipeline_flow.py:135-140,446-452`; `tests/test_orphan_sweep.py:119`; `tests/services/test_workflows_reliability.py:105,178,241,317,375`
- **Problem & Consequence:** Per the test-effectiveness audit: `test_traceability.py:54-57` asserts properties of dict literals defined in the same file (can only fail if the test itself is edited); `assert processed is True` in the pipeline tests is unfalsifiable because `read_job` raises rather than returning `None`; the `>= 200` orphan threshold is a historical snapshot that tolerates ±40 drift; and the `assert result is not None` guards are shadowed by the attribute assertions that follow. Together these overstate effective coverage.
- **Recommendation:** Apply the prune/strengthen backlog in §6 (Testing).
- **Effort:** S
#### [LOW-11] Wall-clock timing dependencies risk CI flakiness
- **Location:** `tests/services/test_workflows_reliability.py:157-196` (real `time.sleep(0.40)`, upper bound `< 540ms` with only 10% slack); `test_workflows_reliability.py:341` (`asyncio.wait_for(..., timeout=2)`)
- **Problem & Consequence:** On a loaded CI runner, a 200ms asyncio task plus 400ms blocking setup can exceed the 540ms bound, producing false failures that erode trust in the suite.
- **Recommendation:** Widen the slack factor to `0.8` or replace the blocking sleep with a controlled clock mock.
- **Effort:** S
---
## 4. Architectural Drift & Gap Analysis
`Direction` is `doc->code` (implementation must change to match documented intent) or `code->doc` (an undocumented but repeatable convention that should be formalized).
| Area / Component | Direction | Documented / Intended Rule | Actual Implementation State | Severity | Recommended Resolution |
| :--- | :--- | :--- | :--- | :--- | :--- |
| Worker commit boundary | `doc->code` | `services.instructions.md`: never commit transcript updates separately from the paired terminal status change | `workflows.py:549-598` commits page evidence and terminal status in two separate sessions | High | Fix per HIGH-01; enforce per HIGH-04 |
| UI error presentation | `doc->code` | `ui.instructions.md:42`: all user-facing error display routes through `error_presenter.py` | 8 hand-rolled `ui.notify` sites in `home_page.py` and `people_page.py` | High | Fix per HIGH-02; add AST guard to `test_ui_boundaries.py` |
| Unexpected-error messaging | `doc->code` | `error-handling.instructions.md:74`: never leak local filesystem paths in user-facing output | `errors.py:94` interpolates raw `exc` into the rendered message | High | Fix per HIGH-03 |
| Retry policy | `doc->code` | `error-handling.instructions.md`: validation / not_found / conflict are non-retriable | `workflows.py:185` gates on retry count only | Medium | Fix per MED-02 |
| Stale-job recovery | `code->doc` | Not documented as startup-only or time-thresholded | Single startup call site; 30s threshold reuses the provider timeout | Medium | Fix per MED-01, then document the recovery contract in `docs/production-runbook.md` |
| Transaction ownership | `code->doc` | `services.instructions.md` assigns transaction ownership to services | `workflows.py` opens scopes via two services' private `_session_scope` | Medium | Fix per MED-05; document the single unit-of-work entry point |
| Blocking-I/O convention | `code->doc` | Not stated as a rule; followed at 5 of 7 sites | `store.py:401` and `homepage_store.py:25,32` deviate | Low | Fix per LOW-01/LOW-02, then state the `asyncio.to_thread` rule in `services.instructions.md` |
| Configuration centralization | `code->doc` | Zero `os.getenv` outside `config.py` — a real, held convention | Held everywhere except the hardcoded `poll_interval_seconds` at `app.py:62` | Low | Fix per LOW-03, then formalize the rule and add a deterministic guard |
| Type-check baseline | `code->doc` | No documented policy for `ty` diagnostics | 10 tolerated false positives; hook is advisory-only | Low | Adopt the suppression strategy in LOW-05 and document it |
| Formatting | `code->doc` | `ruff.toml` configures the formatter | `ruff format --check` absent from the gate; 35 files drifted | Low | Fix per LOW-06 |
| Dependency pin | — | `docs/production-runbook.md` "Dependency upgrade policy" records the exact `nicegui==3.13.0` pin as a deliberate stability decision | Matches | — | **No action** — correctly documented, not a defect |
---
## 5. Invariant Inventory & Routing Recommendations
| Invariant / Constraint | Current Location | Recommended Target Layer | Rationale |
| :--- | :--- | :--- | :--- |
| Transcript + terminal status commit atomically | Instructions only | **Deterministic test** (`tests/integration/test_pipeline_atomicity.py`) | Highest-consequence rule in the system with zero enforcement; steering alone already failed to prevent HIGH-01 |
| Retry writes commit atomically | Instructions only | **Deterministic test** (same file) | Same class; a partial retry commit corrupts `retry_count` accounting |
| All UI errors route through `error_presenter` | Instructions (`ui.instructions.md:42`) | **Deterministic test** (extend `test_ui_boundaries.py`) | Mechanically checkable via AST; 8 live violations prove instructions are insufficient here |
| No filesystem paths in user-facing output | Instructions (`error-handling.instructions.md:74`) | **Deterministic test** (extend `tests/ui/test_error_presenter.py`) | Checkable by asserting a path-bearing exception does not surface its path |
| Non-retriable categories are never requeued | Instructions | **Deterministic test** (`tests/services/test_workflows_reliability.py`) | Latent today; a test freezes the correct behavior before retries are enabled |
| Blocking I/O runs via `asyncio.to_thread` | Convention only (5/7 sites) | **Instructions** (`services.instructions.md`) | Judgment-dependent (thresholds vary by payload size); steering fits better than a hard test |
| Transaction opened through one owned entry point | Convention, violated | **Instructions + test** | Document the entry point; AST-guard against `_session_scope` access outside its owning module |
| Append-only `ExecutionAttempt` history | Docs + 3 tests | **Keep as-is** | Correctly routed and genuinely mutation-sensitive; the model to imitate |
| Service/UI boundary rules | Instructions + 2 AST tests | **Keep as-is** | Working exactly as intended |
| Status vocabulary conformance | `docs/schema.md` + contract guards | **Keep as-is** | Enum drift would fail the suite |
| No secrets in stored evidence | Docs + provenance skill + allowlist in code | **Keep as-is** | Allowlist is the right mechanism — fails closed by construction |
| `ty` diagnostic suppression policy | Nonexistent | **Docs + blocking hook** | Needs a written rationale per suppression before the gate can be trusted |
| NiceGUI exact pin | `docs/production-runbook.md` | **Keep as-is** | Deliberate, documented, correctly excluded from review findings |
---
## 6. Stack-Specific Analysis
### Python 3.12+ Best Practices
Modern syntax is used consistently: `X | None` unions throughout, builtin generics, no `typing.List`/`Optional` legacy forms, `pathlib` over `os.path`. Type-annotation coverage is high, with no bare `Any` on public service signatures. Broad `except Exception` appears where it belongs — the per-page handler at `workflows.py:352` deliberately isolates one page's failure from the batch, which is correct. `# noqa: PLR0915` / `PLR1702` are used sparingly and consistently. Minor gaps: f-strings in logging (LOW-09), and two blocking-I/O deviations (LOW-01/LOW-02).
### FastAPI
Lifespan is handled correctly via an `asynccontextmanager` `_lifespan` (`app.py:36-68`) rather than deprecated `@app.on_event`. Routers are domain-organized with typed path/query parameters and `response_model` declarations. Error handling is centralized through `register_error_handlers`, and the full internal→canonical category mapping is round-trip tested at the HTTP layer (`tests/api/test_error_responses.py:59-95`). `print_api.py:42-49` performs correct `relative_to`-based path containment for media serving. No blocking calls found in `async def` route handlers.
### NiceGUI
Separation of concerns is good — pages delegate to services and `test_ui_boundaries.py` mechanically prevents persistence access from pages and components. Client state is client-scoped; no cross-session global-state leaks found. API usage is correct for the pinned 3.13.0 release. The two defects are the error-presenter bypass (HIGH-02) and synchronous file I/O in `homepage_store.py` (LOW-02).
### SQLModel & SQLAlchemy
The strongest layer. `lazy="raise"` is declared on relationships and correctly paired with `expire_on_commit=False`, which together make N+1 access a loud failure rather than a silent performance cost — no N+1 patterns found. The job claim is a genuine atomic compare-and-swap (`jobs.py:212-222`: conditional `UPDATE ... WHERE status = QUEUED ... RETURNING`), which is the correct primitive and correctly implemented. Hot-path indexes are declared and test-verified (`test_db.py:131`). Cross-dialect portability is handled for SQLite and PostgreSQL. Weaknesses are transaction *ownership* (MED-05, HIGH-01) rather than query construction, plus the uncaught `IntegrityError` at MED-03.
### Pydantic V2 & Settings
Fully migrated — no `@validator`, no `Config` class, no `.dict()` or `parse_obj` anywhere. `model_config = ConfigDict(...)` and `@field_validator` are used correctly. `config.py` is a clean single source of truth: **zero** `os.getenv` calls exist outside it, `.env` is untracked and gitignored, and the API key is `SecretStr` end-to-end. The only deviation is the hardcoded poll interval (LOW-03).
### Asyncio Workers
Task lifecycle is handled properly: task references are retained (no GC risk), `CancelledError` is re-raised rather than swallowed, the provider call happens outside any DB transaction, timeouts resolve to terminal states, and there is no tight polling spin. `_persist_page_outcome_durably`'s use of `asyncio.shield` (`workflows.py:568-581`) is a thoughtful durability mechanism. The defects are the split commit boundary (HIGH-01), the shutdown-vs-provider timeout mismatch (MED-04), and startup-only recovery (MED-01).
### OpenRouter / Adapter Boundary
Encapsulation is clean — `workflows.py` imports only abstract types from `providers`, never `openrouter` directly, so provider specifics do not leak into business logic. The `AsyncClient` is shared with configured timeouts and is properly closed: `worker.py:248,271` → `services.aclose()` → `sources.aclose()` (`sources.py:129-133`) → provider `aclose()` (`openrouter.py:86-87,233-235`). Responses are Pydantic-validated. **All 14 evidence-provenance-auditor invariant checks pass**, including the critical one: the API key is never persisted, request headers are never stored, and `TransportEvidence` captures response headers through an explicit allowlist (`evidence.py:130-134`). Only LOW-04 applies here.
### Testing & Quality Tooling
377 tests pass with `-m "not external"`. The project test contract is honored: `--strict-markers` with all three markers (`unit`, `integration`, `external`) declared, `asyncio_mode = "strict"` with **every** `async def test_` correctly decorated across all 17 async test files, `external` properly excluded from default runs, and **no unawaited-coroutine warnings** — the `filterwarnings` error promotion is clean.
Contract coverage is genuinely strong for structural rules. Confirmed *mutation-sensitive* enforcement exists for: append-only evidence history (3 independent tests, including full before/after field-tuple snapshots), stuck-in-`PROCESSING` prevention, the complete 10-category error mapping, and both boundary rules.
The critical gap is transaction atomicity (HIGH-04) — the audit verdict is **"Effective with Conditions / Go with Conditions"**, blocking on the two missing atomicity tests. Secondary items are the low-signal assertions (LOW-10) and wall-clock flakiness (LOW-11).
**Prune/strengthen backlog:**
| Priority | Task | Location |
| :--- | :--- | :--- |
| High | Add Transaction B atomicity test (fault injected between transcript and status writes) | new `tests/integration/test_pipeline_atomicity.py` |
| High | Add Transaction C atomicity test (retry: `error_detail` + `retry_count` + `QUEUED`) | same file |
| Medium | Delete tautological assertions on same-file dict literals | `tests/test_traceability.py:54-57` |
| Medium | Remove unfalsifiable `assert processed is True` | `tests/integration/test_pipeline_flow.py:135-140,446-452` |
| Medium | Replace `>= 200` snapshot threshold with set-membership assertion | `tests/test_orphan_sweep.py:119` |
| Medium | Assert mapped test files contain ≥1 test, not merely that they exist | `tests/test_traceability.py:59-60` |
| Low | Widen timing slack or mock the clock | `tests/services/test_workflows_reliability.py:157-196` |
| Low | Drop `assert result is not None` guards shadowed by following assertions | `tests/services/test_workflows_reliability.py:105,178,241,317,375` |
---
## 7. Duplication & Consolidation Report
| Pattern / Duplication | Locations | Proposed Canonical Home | Estimated Lines Removed |
| :--- | :--- | :--- | :--- |
| Hand-rolled `ui.notify(str(exc), type="negative")` | `home_page.py:212,220,228,255`; `people_page.py:265,321,330,339` | `ui/components/error_presenter.py::show_error` (already exists) | ~16 |
| Optional-session `if session is None: async with _session_scope()` preamble | `workflows.py:557-561`, `workflows.py:591-595`, and sibling service write paths | `services/base.py::unit_of_work(services, session)` context manager | ~30 |
| Synchronous I/O not wrapped in `asyncio.to_thread` | `store.py:401`, `homepage_store.py:25,32` | `services/base.py::run_blocking` helper | ~6 |
| Read-then-increment `MAX(n) + 1` with uniqueness retry | `sources.py:540-546` (uncaught) vs `sources.py:531-534` (caught) | `services/base.py::insert_with_sequence_retry` | ~20 |
### Proposed Canonical Abstractions
```python
# src/transcription/services/base.py
@asynccontextmanager
async def unit_of_work(
services: ServiceBundle,
session: AsyncSession | None = None,
) -> AsyncIterator[AsyncSession]:
"""Single transaction entry point. Yields a session and commits once on clean exit.
Replaces the `if session is None: async with X._session_scope()` preamble and the
private-member access at workflows.py:558,592. Makes the two-commit split of
HIGH-01 structurally hard to reintroduce.
"""
async def run_blocking[T](fn: Callable[[], T]) -> T:
"""Run a CPU- or disk-bound callable off the event loop."""
return await asyncio.to_thread(fn)
async def insert_with_sequence_retry(
session: AsyncSession,
*,
build: Callable[[int], SQLModel],
next_value: Callable[[], Awaitable[int]],
attempts: int = 3,
) -> SQLModel:
"""Insert a row carrying a derived sequence number, retrying on IntegrityError."""
```
---
## 8. Meta-Tooling & Instruction Update Recommendations
1. **Add `tests/integration/test_pipeline_atomicity.py`** (HIGH-04). The single highest-value enforcement change. Write it before fixing HIGH-01 so it demonstrably fails first.
2. **Extend `tests/test_ui_boundaries.py`** with an AST check forbidding `ui.notify(..., type="negative")` in `PAGES_DIR`, routing all error display through `error_presenter`. Converts `ui.instructions.md:42` from steering into enforcement.
3. **Extend `tests/ui/test_error_presenter.py`** with a case asserting that a path-bearing exception does not surface its path, enforcing `error-handling.instructions.md:74`.
4. **Adopt a `ty` suppression policy** — targeted `# ty: ignore[...]` with rationale at the 10 known sites, documented in `docs/` — then **flip the pre-commit `ty` hook from advisory to blocking**. Until this happens the type checker provides no gate.
5. **Add `ruff format --check` to the pre-commit gate**, preceded by one isolated formatting commit across the 35 drifted files.
6. **Enable ruff rule set `G`** (`flake8-logging-format`) to catch f-string logging (LOW-09).
7. **Extend `tests/test_orphan_sweep.py`** to public methods on service classes, seeding `KNOWN_ORPHANS` with current results (LOW-08). Then resolve all four existing "uncertain" entries to definite outcomes.
8. **Extend `tests/test_service_boundaries.py`** with an AST check forbidding `_session_scope` attribute access outside its owning service module (MED-05). Also address the noted classification gap: the test excludes orchestration modules by hardcoded stem name (`store`, `workflows`, `__init__`), so a new orchestration module under a different name would be misclassified as a service.
9. **Update `.github/instructions/services.instructions.md`** to state the `asyncio.to_thread` rule for blocking I/O and to name the single `unit_of_work` transaction entry point.
10. **Update `docs/production-runbook.md`** with the stale-job recovery contract (interval, threshold, and its relationship to the container termination grace period), and note single-worker as a current precondition until MED-03 is fixed.
11. **Note for `test_ui_boundaries.py`:** the forbidden-import lists are fixed string sets, so a future persistence helper under a new name would escape the check. Consider inverting to an allowlist of permitted imports for pages.
---
## 9. Prioritized Dependency-Ordered Action Plan
**Phase 1: Blocking fixes**
1. Write the two atomicity tests (HIGH-04) and confirm they **fail** against current `main`.
2. Fix the split commit boundary (HIGH-01) and confirm the tests now pass.
3. Fix the filesystem-path leak in `classify_unexpected_error` (HIGH-03).
4. Replace the 8 hand-rolled error notifications with `show_error` (HIGH-02).
**Phase 2: Enforcement hardening**
5. Add the `ui.notify` AST guard and the path-leak presenter test, locking in items 3-4.
6. Adopt the `ty` suppression policy and make the pre-commit hook blocking (LOW-05).
7. Run `ruff format .` as an isolated commit, then add `ruff format --check` to the gate (LOW-06).
8. Enable ruff rule set `G` and fix the resulting logging call sites (LOW-09).
**Phase 3: Reliability & concurrency**
9. Move stale-job recovery to a periodic worker task with a dedicated setting (MED-01).
10. Gate retries on `error_category` and add backoff (MED-02) — do this before ever raising `worker_max_retries` above 0.
11. Derive the shutdown budget from the provider timeout (MED-04).
12. Handle `IntegrityError` on the attempt-number flush (MED-03) — a hard precondition for running more than one worker replica.
13. Move `sha256` and homepage-store I/O off the event loop (LOW-01, LOW-02); move the poll interval into `Settings` (LOW-03).
**Phase 4: Consolidation & refactoring**
14. Introduce `unit_of_work` and migrate `workflows.py` off private `_session_scope` access (MED-05); add the corresponding boundary guard.
15. Extract `run_blocking` and `insert_with_sequence_retry` (§7).
16. Prune the low-signal assertions and reduce timing flakiness (LOW-10, LOW-11).
**Phase 5: Non-blocking governance/documentation depth**
17. Extend the orphan sweep to methods and resolve the four uncertain orphans (LOW-07, LOW-08).
18. Update `services.instructions.md` and `docs/production-runbook.md` per §8 items 9-10.
19. Log incomplete request manifests (LOW-04).
20. Consider inverting the UI boundary check to an allowlist.
---
## 10. Preserved Strengths
- **Evidence and provenance integrity is exemplary.** All 14 provenance-auditor invariants pass. `ExecutionAttempt` history is genuinely append-only, retries append rather than rewrite, and projection writes are cleanly distinguished from history mutation. Three independent tests — including full before/after field-tuple snapshots — make any mutation regression fail loudly.
- **Secret hygiene is correct by construction.** The response-header **allowlist** (`evidence.py:130-134`) fails closed: a newly-introduced sensitive header is excluded by default rather than requiring someone to remember to block it. Request headers are never captured, and `SecretStr` is used end-to-end.
- **Atomic job claiming.** `jobs.py:212-222` uses a conditional `UPDATE ... WHERE status = QUEUED ... RETURNING` — a true compare-and-swap that makes double-claiming impossible under concurrency, rather than the common read-then-write race.
- **`lazy="raise"` paired with `expire_on_commit=False`.** This combination turns accidental lazy loads into immediate errors instead of silent N+1 queries, and it is the reason no N+1 patterns exist in the codebase. Keep it.
- **Architectural rules are mechanically enforced, not merely documented.** AST-based boundary tests for service-to-service imports and UI persistence access are the right pattern; this review's main recommendation is simply to apply that same pattern to three more rules.
- **Configuration discipline.** Zero `os.getenv` calls outside `config.py`, `.env` untracked and gitignored, clean Pydantic V2 throughout with no V1 residue.
- **Path containment on media serving.** `print_api.py:42-49` uses proper `relative_to` validation rather than string prefix matching.
- **Async worker fundamentals.** Task references retained, `CancelledError` re-raised, provider calls outside DB transactions, timeouts resolving to terminal states, no tight polling loop. `asyncio.shield` in `_persist_page_outcome_durably` is a genuinely thoughtful durability mechanism — the fix in HIGH-01 should preserve it for intermediate pages.
- **Test contract rigor.** `--strict-markers`, `asyncio_mode = "strict"` honored across all 17 async test files with no missing decorators, and coroutine-never-awaited promoted to a hard error with a clean run.
+896
View File
@@ -0,0 +1,896 @@
# Architecture & Code Review Report
**Repository Target:** `C:\Github\transcription\`
**Target Stack:** Python 3.12+ | FastAPI | NiceGUI | SQLModel/SQLAlchemy | Pydantic V2 | asyncio | OpenRouter
**Review Date:** 2026-09-02
**Canonical Baseline:** V6.1 (`docs/index.md`)
---
## 0. Verification Commands and Outcomes
All four commands were executed in this checkout before any finding was written. This report
records the exact outcomes rather than assuming them.
| Command | Outcome |
| :--- | :--- |
| `uv run pytest -q -m "not external"` | **410 passed, 0 failed, 0 errors** (exit 0) |
| `uv run ruff check .` | **All checks passed!** |
| `uv run ruff format --check .` | **191 files already formatted** |
| `uv run ty check` | **All checks passed!** |
The stated green baseline is real. No finding below is a test failure; every finding is a
behavior, contract, or guard-coverage defect that the passing suite does not detect.
---
## 1. Executive Summary
- **The system's core evidence guarantees hold.** `ExecutionAttempt` is genuinely append-only,
attempt numbering is allocated with bounded conflict retry, transport evidence is captured at
the HTTP boundary before SDK parsing, and header persistence uses a true allowlist. Provenance
invariant families AE and G pass.
- **Both competing atomicity invariants in `services/workflows.py` are real and both guards
genuinely enforce them.** I injected-fault-verified the tests rather than trusting the
docstrings: `test_pipeline_atomicity.py` fails on a split final-page commit, and
`test_workflows_reliability.py:318` reads intermediate attempts through a *separate session*,
so it would fail if intermediate pages stopped committing individually.
- **The most significant defect is a privacy leak that a prior review believed it had closed.**
The 2026-08-23 review moved root-cause text out of `AppError.message` into `AppError.detail`
to keep filesystem paths away from users. That text now reaches users anyway, because the UI
renders `ExecutionAttempt.error_detail` verbatim (HIGH-01). The leak was relocated, not closed.
- **A second, independent path leak exists in five explicit `raise` sites** that the existing
guard never covered — it tests only `classify_unexpected_error` (HIGH-02).
- **Provenance invariant family F (path safety) fails**, and it fails *inconsistently within one
file*: `sources_page.py:443` carefully sanitizes a stored path through
`public_media_path_label`, then `sources_page.py:484` dumps raw `error_detail` forty lines later.
- **The orphan sweep does not do what its docstring claims.** It matches definitions by bare name,
so an entirely dead *module* passes whenever its function names collide with live ones.
`ui/pages/tags_page.py` is the proof: 93 lines never imported by anything (MED-01/LOW-01).
- **On the three flagged open items:** the V4/V6.1 doc drift is confirmed (MED-02); the
`.env.production` coupling is real but currently correct and loud-failing, so Medium not High
(MED-03); and the Tags roadmap is **right** — the route is genuinely not registered, so the
module is dead code rather than a live retired route.
- **Two latent concurrency defects carry ordering constraints** and must be fixed *before* the
changes that would make them live (MED-04, MED-05), not after.
- **Guidance-file accuracy:** the recently revised `.github/instructions/*` files were verified
against code rather than trusted. They are accurate as written; the code is what diverges from
them. The one exception is that `error-handling.instructions.md` states a `detail` rule the UI
layer has never followed, which makes it an unenforced claim rather than a wrong one.
---
## 2. Executive Architecture Assessment
**Verdict: architecturally sound, with a concentrated failure in the *last mile* of error
presentation.**
Domain cohesion and dependency direction are good and, unusually, mechanically enforced.
`test_service_boundaries.py` and `test_ui_boundaries.py` AST-scan for violations using
*allowlists* rather than blocklists, which is the correct choice — a newly added persistence
helper cannot slip through under an unlisted name. `workflows.py` imports only the abstract
`providers` types and never `openrouter`, so provider details genuinely stop at the adapter.
Transaction ownership is explicit and well-reasoned: `ServiceBase._finalize` commits for
service-owned sessions and flushes for caller-owned ones, which is what lets orchestration
modules compose multi-aggregate writes without services importing each other.
The evidence layer is the strongest part of the system and shows real care. The distinction
between transport response, SDK-parsed response, and normalized metadata is maintained in code,
not just in prose — `_CapturingAsyncClient` exists specifically to retain the exact wire body
before the SDK can discard unknown fields, and `TransportEvidence(response_received=False)`
explicitly represents "no response was received" rather than conflating it with an empty one.
The weakness is at the boundary where internal diagnostic text becomes pixels. Every layer
*below* the UI respects the message/detail split; the UI layer reads the internal field directly
and renders it. The architecture defines the contract correctly and then has no enforcement at
the one layer that violates it.
**Top systemic risks:**
1. **Internal diagnostic text reaches users through the evidence display path** (HIGH-01). The
rule is documented in three places and enforced in none of them at the UI boundary.
2. **Path-safety discipline is applied per-call-site rather than structurally** (HIGH-02, HIGH-01).
It is correct wherever someone remembered; there is no guard that makes forgetting fail.
3. **Guard coverage is narrower than guard docstrings claim.** Two guards
(`test_orphan_sweep.py`, `test_errors.py`) assert something meaningfully weaker than the
invariant they are named for, which converts them into a false sense of enforcement.
4. **Worker safety currently rests on single-process sequential execution, not on configuration**
(MED-04, MED-05). Nothing is wrong today; two plausible future changes each make something wrong.
---
## 3. Findings by Severity
### Critical Severity
*None.* No evidence loss, append-only violation, secret leakage, or silent-wrong-output defect
was found. The candidates in this class (provider evidence mis-attribution, stale-job double
processing) are latent and are reported at High/Medium with their unblocking conditions.
---
### High Severity
#### [HIGH-01] Internal-only `error_detail` is rendered directly to users, reopening the leak the 2026-08-23 fix was meant to close
- **Location:**
- Write side: `src/transcription/errors.py:99-139` (`classify_unexpected_error``detail`, `format_error_detail` → persisted text)
- Persist: `src/transcription/services/workflows.py:720` (`error_detail=format_error_detail(page.error)`)
- **Render (Source Detail):** `src/transcription/ui/pages/sources_page.py:481-484`
- **Render (Sources list):** `src/transcription/ui/pages/sources_page.py:121``src/transcription/ui/components/table/sources.py:39,90-95` ("Error Detail" column)
- **Render (Maintenance):** `src/transcription/ui/pages/settings_page.py:562`, written by `src/transcription/services/maintenance.py:206`
- Contract violated: `docs/error_handling.md:107-114`; `.github/instructions/error-handling.instructions.md:86`; `docs/invariant/error_handling.md:59`; `docs/invariant/ai_evidence_and_provenance.md:103`
- **Reachability:** **Live.** Concrete path, no configuration required: a page fails with any
non-`AppError` exception → `workflows.py:388` calls `classify_unexpected_error(exc)`
`errors.py:118` sets `detail=f"{type(exc).__name__}: {exc}"``format_error_detail`
(`errors.py:135-139`) emits `... | detail=OSError: [Errno 13] Permission denied: '/app/uploads/documents/<uuid>/page-1.jpg' | ...`
→ persisted to `ExecutionAttempt.error_detail` → rendered verbatim at
`sources_page.py:484` and in the `/sources` table column. A SQLAlchemy `OperationalError`
carries the database path by the same route.
- **Problem & Consequence:** `docs/error_handling.md:110` states `detail` is *"Internal only"* and
that its only surfaces are `format_error_detail` (evidence) and logs;
`error-handling.instructions.md:86` says *"Never rendered to users or serialized into an
envelope."* The UI reads it anyway. The consequence is not hypothetical drift — it is the
precise defect the previous review's fix existed to prevent. That fix made `message` generic and
moved the root cause to `detail` on the stated grounds that `detail` never reaches users. That
premise was never true: `error_detail` had a UI consumer the whole time. The result is that the
filesystem-path leak was relocated from the notification banner to the Source Detail card and
the Sources table, while the test suite records the leak as fixed
(`tests/test_errors.py:56-78`).
The inconsistency is visible inside a single file: `sources_page.py:443` deliberately routes a
stored path through `public_media_path_label` (`ui/components/media_urls.py:58-72`), which
correctly degrades an absolute path to its bare filename — and then `sources_page.py:484`
renders unsanitized text that may contain an absolute path.
- **Blast Radius:** Enumerated by grepping every reader of `.detail` and `error_detail`:
- `errors.py:137``format_error_detail`, the only reader of `AppError.detail`. **Must keep the root cause.**
- `services/workflows.py:720` — the only writer of `ExecutionAttempt.error_detail`.
- `services/maintenance.py:206` — the only writer of `MaintenanceRun.error_detail`.
- `services/evidence.py:195``build_evidence_export` emits `error_detail`. Export is an
operator-initiated evidence artifact; per invariant 3.7.1 it **must** retain it.
- `db/models.py:508-522``Source.latest_error_detail` projection, consumed only by `sources_page.py:121`.
- `ui/pages/sources_page.py:481-484`, `ui/components/table/sources.py`, `ui/pages/settings_page.py:562` — the three render sites.
- Tests asserting on persisted text: `tests/test_v42_evidence.py:284`,
`tests/services/test_workflows_reliability.py` (timeout detail),
`tests/services/test_maintenance_service.py`. A fix that changes *what is stored* breaks these;
a fix that changes *what is displayed* does not.
- **Recommendation — two invariants conflict here; both must be named.**
**Invariant 1 (evidence):** `ExecutionAttempt.error_detail` must retain the root cause.
`docs/requirements.md:30` (REQ-4-021) and `docs/invariant/ai_evidence_and_provenance.md:33`
require it; guarded by `tests/test_v42_evidence.py::test_attempts_are_append_only_and_exported_with_integrity`
and `tests/services/test_workflows_reliability.py`.
**Invariant 2 (privacy):** user-facing surfaces must not expose local filesystem details.
`docs/invariant/error_handling.md:59`; guarded (partially) by
`tests/test_errors.py::test_unexpected_error_does_not_leak_filesystem_paths`.
**The over-correction to avoid is stripping root-cause text out of `detail` or
`format_error_detail` to make the UI safe.** That is exactly the mistake documented in the
reviewer skill's worked example, and it would silently destroy the provenance record this
system exists to preserve while making every guard still pass.
Fix at the **render** boundary, not the write boundary. Add a presentation-layer projection and
route all three UI sites through it, leaving the persisted evidence untouched:
```python
# src/transcription/ui/components/error_presenter.py (new)
def display_failure_detail(error_detail: str | None) -> str | None:
"""Render persisted failure detail without machine-local paths.
`ExecutionAttempt.error_detail` is provenance and keeps the full root cause
(docs/error_handling.md). This projection is the only thing a page may show.
"""
```
It should preserve the `[category]`, `suggestion=`, and `error_id=` segments (which are what
make the display actionable) and reduce any absolute path inside `detail=` to its basename,
mirroring `public_media_path_label`. The operator keeps diagnosability — required by
`docs/ui/pages/sources.md:43` and `docs/requirements.md:59` (REQ-6-014) — without the container
filesystem layout being published to the browser.
Then decide and record which resolution was chosen: either the UI shows the sanitized
projection (recommended), or `docs/error_handling.md:107-114` and
`error-handling.instructions.md:86` are revised to state that operator-facing evidence displays
may render `error_detail` **and** that the guarantee moves to "no machine-local detail ever
enters `detail`" — which would be a much harder guarantee to keep. Do not leave the current
state, where the docs claim one thing and three pages do another.
- **Effort:** M
---
#### [HIGH-02] Absolute filesystem paths are embedded in user-facing `AppError.message` at five explicit raise sites
- **Location:**
- `src/transcription/services/sources.py:856` — `f"Prompt file not found: {prompt_path}"`
- `src/transcription/services/sources.py:864` — `f"Prompt file is empty: {prompt_path}"`
- `src/transcription/services/sources.py:914` — `f"Source file not found: {path}"`
- `src/transcription/services/prompts.py:99` — `f"Prompt directory is unavailable: {root}"`
- `src/transcription/services/prompts.py:186-191` — `_filesystem_error` builds `f"{message}: {exc}"`
- Contract violated: `.github/instructions/error-handling.instructions.md:74,85`; `docs/invariant/error_handling.md:59`
- **Reachability:** **Live**, on an ordinary user path. `sources.py:845` resolves
`prompt_root = runtime_settings.prompt_dir.resolve()`, so `prompt_path` is absolute
(`/app/prompts/transcribe_document.md` in the container). `load_prompt_text` is invoked by
`build_prompt_execution` (`sources.py:829-831`), which runs on **every document upload** via
`services/store.py:94` and `store.py:162`. The resulting `PromptLoadError` is an `AppError`
subclass, so it flows through `run_ui_action` → `show_error`
(`ui/components/error_presenter.py:51-66`), which renders `error.message` into both a
`ui.notify` banner and a card label, and through `build_error_envelope` (`errors.py:88-96`)
into API responses.
- **Problem & Consequence:** `error-handling.instructions.md:85` requires `message` to *"Stay
generic. Never embed exception text, provider payloads, or filesystem paths."* These five sites
embed exactly that. `prompts.py:186-191` violates the rule in **both** directions at once: it
puts `{exc}` — an `OSError` whose `str()` includes the offending filename — into `message`, and
it sets **no `detail=`**, so the internal field that is supposed to carry the root cause is
empty while the user-facing field carries all of it.
This is not a new regression; it is coverage that the existing guard never had.
`tests/test_errors.py:56-78` verifies only that `classify_unexpected_error` — the *catch-all*
path — does not leak. Every deliberate `raise SomeError(f"... {path}")` in the codebase is
outside its scope, so the suite reports the invariant as enforced while five live sites violate it.
- **Blast Radius:** Verified by grepping all consumers of these exception types.
`PromptLoadError`/`PromptStoreError`/`TranscriptionError` messages are consumed by:
`ui/components/error_presenter.py:55,63` (render), `errors.py:92` (API envelope),
`errors.py:135` (`format_error_detail` → evidence). Because the recommended change *adds* a
`detail` and *shortens* `message`, `format_error_detail` output still contains the path — so
evidence value is preserved, not reduced. Tests asserting on these messages:
`tests/test_prompts.py`, `tests/services/test_prompt_store.py`,
`tests/services/test_transcription_service.py`. These assert on message prefixes
(`"Prompt file not found"`), not on the interpolated path, and were checked to survive the change —
but re-run them, since `prompts.py:186` currently produces a message whose suffix some
assertion could depend on.
- **Recommendation:** Apply the pattern `errors.py:113-119` already establishes — generic
`message`, root cause on `detail`, `raise ... from exc`. Use `path.name` when a filename is
genuinely useful to the user.
```python
# sources.py:855 — before
raise PromptLoadError(f"Prompt file not found: {prompt_path}", ...)
# after
raise PromptLoadError(
f"Prompt file not found: {prompt_path.name}",
category=ErrorCategory.INFRA_PERSISTENT,
suggestion="Verify PROMPT_DIR and prompt file configuration, then retry.",
detail=f"Prompt file missing at {prompt_path}",
)
# prompts.py:186 — before
return PromptStoreError(f"{message}: {exc}", category=..., suggestion=...)
# after
return PromptStoreError(
message,
category=ErrorCategory.INFRA_PERSISTENT,
suggestion="Check prompt directory permissions and available disk space, then retry.",
detail=f"{type(exc).__name__}: {exc}",
)
```
Then widen the guard so this class cannot recur — see MED-07. Note the dependency: HIGH-02 and
HIGH-01 must be fixed **together**, because moving the path from `message` to `detail` while the
UI still renders `error_detail` relocates the leak instead of closing it. That is the same
mistake that produced HIGH-01.
- **Effort:** S (fix) / M (with the guard)
---
### Medium Severity
#### [MED-01] The orphan sweep matches by bare name and therefore cannot detect a dead module
- **Location:** `tests/test_orphan_sweep.py:101-163` (`_public_definitions`, `_orphans`)
- **Reachability:** **Live** — the guard is running now and reporting a clean sweep that is not clean.
- **Problem & Consequence:** `_public_definitions()` keys definitions by bare name
(`definitions[node.name]`, line 111) and `_orphans()` marks a definition referenced if that
bare name appears **anywhere** in `src/`, `tests/`, or `tools/` (lines 157-162). Two different
modules that define the same public name are therefore indistinguishable, and neither can ever
be reported as an orphan.
`src/transcription/ui/pages/tags_page.py` demonstrates the consequence. Its only public
definition is `register_page` (line 18). Seven live page modules define a function of the same
name and `ui/__init__.py:37-43` calls all seven — so `register_page` is heavily referenced and
`tags_page.register_page` is scored as reachable. In fact **nothing imports `tags_page` at all**
(verified: the only repo-wide references to the module are the file itself and
`tests/ui/test_tags_page.py`, which merely asserts the route 404s). 93 lines of code, including a
lazy-load-unsafe relationship traversal at `tags_page.py:71-74`, sit outside the sweep's reach.
The sweep also never asks whether a *module* is imported, only whether its definitions' names
appear somewhere — so this is a structural gap, not a one-off miss.
- **Blast Radius:** `tests/test_orphan_sweep.py` only; `KNOWN_ORPHANS` entries are keyed by the
same bare/dotted names and would need re-keying if qualification is added. Expect the stricter
sweep to surface additional true orphans on first run — triage them into `KNOWN_ORPHANS` with
rationales rather than weakening the check.
- **Recommendation:** Qualify definitions by module (`f"{module_path}:{name}"`) and add a separate,
cheap module-reachability pass: a module under `src/transcription/` is reachable if any other
module imports it, or it is a declared entrypoint (`app.py`, `__main__.py`, `worker_service.py`).
Report unreachable modules as orphans in their own right. Also fix
`test_public_definitions_are_discovered` (line 169), whose `>= 420` snapshot threshold is a
weak assertion that drifts upward silently — the 2026-08-23 review already flagged the same
pattern at the then-current `>= 200` and it was raised rather than replaced.
- **Effort:** M
---
#### [MED-02] Canonical invariant document declares a V4 baseline while the canonical baseline is V6.1
- **Location:** `docs/invariant/ai_evidence_and_provenance.md:130`
- **Reachability:** **Live** (documentation), no runtime impact.
- **Problem & Consequence:** Section 6.1 reads *"Canonical V4 architecture, schema, requirements,
and error-policy documents define how current behavior satisfies this invariant."*
`docs/index.md:1,29-32` establishes V6.1 as the baseline and states that every canonical document
asserts the same baseline. This is the **ownership clause of the invariant that governs the
entire evidence model** — the clause that tells a reader which documents are authoritative — and
it points at a superseded generation. A reader following it lands on stale authority precisely
when resolving an evidence question, which is the highest-stakes case.
A baseline-currency guard **does** exist —
`tests/test_meta_contract_guards.py::test_canonical_docs_declare_one_consistent_baseline`
(lines 89-112) — and `docs/invariant/ai_evidence_and_provenance.md` is **not** in
`BASELINE_SCAN_EXCLUSIONS` (lines 56-64), so the file is scanned. The claim escapes for two
independent reasons, either of which alone would be sufficient:
1. `_CURRENT_VERSION_CLAIM` (line 67) matches only the words `current` or `active` before a
version. This line says "**Canonical** V4", a third phrasing the pattern does not know.
2. Both patterns require `V(\d+\.\d+)` — a mandatory minor version. The bare token `V4` cannot
match either regex under any phrasing.
The guard is therefore not absent but *phrase-shaped*: it enforces currency only for the two
sentence forms someone thought of, against version strings that carry a minor. That is a weaker
property than its docstring implies ("Every canonical doc that names the current baseline must
name the same one").
- **Blast Radius:** Documentation only; no code reads this string. Widening the guard's patterns
will re-scan all canonical docs — expect it to surface further stale mentions on first run
(`docs/architecture.md`, `docs/schema.md`, `docs/requirements.md`, and `docs/error_handling.md`
each contain 2-3 version tokens), which should be triaged rather than excluded.
- **Recommendation:** Two parts, and the second matters more than the first.
1. Change "Canonical V4" to "Canonical V6.1" at line 130.
2. Fix the guard's shape rather than adding a third phrase to the list. Accept an optional minor
(`V(\d+)(?:\.(\d+))?`) and invert the matching: flag **every** `V<n>` token in a scanned
canonical doc that is not the declared baseline, rather than only those preceded by an
approved adjective. Phrase-list matching fails open — each new phrasing silently reopens the
hole — whereas token matching fails closed and forces an explicit exclusion.
See §8.1 for the alternative the maintainer is considering: dropping version labels from
canonical docs entirely, which removes the failure mode instead of guarding it.
- **Effort:** S
---
#### [MED-03] Settings resolve `.env.production` relative to the process working directory, and the isolation fix exists only in the test harness
- **Location:** `src/transcription/config.py:66-75` (`env_file=".env.production"`);
workaround at `tests/conftest.py:27-50`; guarded by `tests/test_config_isolation.py`;
depended on by `.github/workflows/quality-gate.yml` and `docker-compose.production.yml`
- **Reachability:** **Live but currently correct.** I verified the production path rather than
assuming it: `Dockerfile` sets `WORKDIR /app` in the runtime stage, and
`docker-compose.production.yml` mounts `./.env.production` to `/app/.env.production` for both the
`app` and `worker` services, so the relative path resolves correctly today.
- **Problem & Consequence:** Correct configuration loading depends on an **implicit, undocumented
contract between `config.py` and the process working directory.** Nothing in `config.py` states
it, and nothing tests it. The failure mode is not silent — `openrouter_api_key` is required with
no default, so a wrong cwd produces a `ValidationError` at startup rather than a partially
configured process — which is why this is Medium rather than High.
The more telling symptom is what the coupling forced on the test harness. `conftest.py:45-50`
cannot escape it by passing an argument; it must **mutate the Pydantic class-level
`model_config` dict at runtime** and restore it in a `finally`. That is a global, order-sensitive
side effect adopted because the module offers no seam. It also silently repairs a second
consumer: `ui/runtime_settings_store.py:402` reads the same `Settings.model_config["env_file"]`
to decide where the Settings page writes. Two subsystems are coupled through a mutable class
attribute.
- **Blast Radius:** Every `Settings` construction. Consumers of `model_config["env_file"]`:
`ui/runtime_settings_store.py:402` (write-target resolution, contract documented at
`docs/ui/pages/settings.md:27`) and `tests/conftest.py:45-50`. A change must preserve the
documented three-step resolution order — explicit override, `RUNTIME_SETTINGS_ENV_FILE`, then the
configured default — or `docs/ui/pages/settings.md:27` becomes wrong.
- **Recommendation:** Introduce one explicit resolution function that both `Settings` construction
and `runtime_settings_store` call, honoring an `ENV_FILE` environment variable and falling back
to a path anchored to a known root rather than to `os.getcwd()`. Tests then pass a path instead of
mutating class state, and `tests/test_config_isolation.py` can assert against the seam rather
than against the monkeypatch. If instead the cwd contract is accepted as deliberate, document it
in `config.py` and in `docs/production-runbook.md` and add a guard asserting `WORKDIR`/cwd
alignment — an implicit contract with a container image is exactly the kind of rule the invariant
routing table exists to place.
- **Effort:** M
---
#### [MED-04] Stale-job reclaim threshold is not derived from maximum job duration; safety currently comes from single-process sequencing
- **Location:** `src/transcription/config.py:116-117`
(`worker_provider_timeout_seconds=30.0`, `worker_stale_job_seconds=30.0`);
sweep at `src/transcription/worker.py:222-228`; reclaim at
`src/transcription/services/jobs.py:242-268`
- **Reachability:** **Latent.** Unblocked by *either* of: (a) running more than one worker replica
(adding `deploy.replicas > 1` to the `worker` service in `docker-compose.production.yml`), or
(b) setting `RUN_EMBEDDED_WORKER=true` on the `app` service while the standalone `worker`
container is also running. It is safe today only because
`docker-compose.production.yml` sets `RUN_EMBEDDED_WORKER: "false"` on `app` and defines exactly
one `worker`, and because within a single loop `run_worker_loop` awaits
`process_next_queued_job` to completion before returning to the stale sweep — so the sweep can
never observe a job that this same process is actively working.
- **Problem & Consequence:** The stale threshold (30s) **equals** the per-page provider timeout
(30s), leaving zero margin even for a single-page job. A multi-page document is legitimately
`PROCESSING` for up to N × 30s. `Job.date_updated` carries an `onupdate`
(`db/models.py:372-375`), but between the initial claim and the terminal write the only touch
is `sources.py:522-523` reassigning `job.provider`/`job.model` to values they usually already
hold, which SQLAlchemy resolves to no net change and therefore no `UPDATE`. I did not empirically
confirm the no-`UPDATE` behavior, so treat that specific step as unverified — but the finding does
not depend on it, because even a per-page refresh leaves only a 30s margin against a 30s timeout.
With a second concurrent worker, the sweep would requeue a job that is mid-provider-call. Both
workers then process the same job, producing duplicate `ExecutionAttempt` rows for the same
logical work and racing terminal status writes. Append-only history would be *preserved* but no
longer *faithful*: the evidence would show attempts that do not correspond to distinct
application decisions.
This is worth flagging because `jobs.py:191-197` explicitly implements and documents
`SKIP LOCKED` row locking "so concurrent workers never contend for the same job." The claim path
is built for multi-worker operation; the reclaim path is not. A reader who trusts the claim
docstring would reasonably scale the worker.
- **Blast Radius:** `requeue_stale_processing_jobs` has one production caller (`worker.py:226`) and
tests in `tests/test_worker.py` and `tests/services/test_job_service.py`. Changing the *default*
affects `tests/test_config.py` declared-defaults assertions — check those before editing the default.
- **Recommendation:** **Fix before adding a second worker replica, not after.** Two parts:
(1) Make the threshold a function of the real bound rather than a coincidental peer of the
page timeout — at minimum default `worker_stale_job_seconds` to a multiple of
`worker_provider_timeout_seconds` with headroom, and add a model validator rejecting a stale
threshold at or below the provider timeout.
(2) Preferably make reclaim heartbeat-based: have `_persist_page_outcome` bump `Job.date_updated`
explicitly so liveness reflects progress rather than elapsed time since claim.
Add a guard asserting a multi-page job in flight is not reclaimed by a concurrently-invoked sweep.
- **Effort:** M
---
#### [MED-05] Provider evidence capture is per-instance mutable state, making the adapter non-reentrant by contract
- **Location:** `src/transcription/providers/openrouter.py:197-199, 264-267, 274-275, 297-298, 397-412`;
`_CapturingAsyncClient.last_response`/`last_body` at `openrouter.py:66-94`;
contract at `src/transcription/providers/base.py:110-118`
(`current_request_manifest`, `current_transport_evidence`)
- **Reachability:** **Latent.** Unblocked by any concurrent `transcribe()` on a single adapter
instance — most plausibly by processing a job's pages in parallel (`workflows.py:274` is
currently a sequential `for` loop) or by any second consumer sharing one
`SourceService.provider`. Verified safe today: `workflows.py:272` resolves one provider for the
loop and awaits each page; the worker's `ServiceBundle` (`worker.py:206`) is distinct from
`app.state.services` (`app.py:43`), so the UI cannot share the worker's adapter instance, and
the UI only enqueues jobs (`ui/pages/jobs_page.py:186-208`).
- **Problem & Consequence:** The `TranscriptionProvider` protocol defines evidence retrieval as
"the most recent call" state read *after* the fact. `workflows.py:369-370` relies on this on the
timeout path, reading `provider.current_request_manifest` / `current_transport_evidence` when no
result object exists. Under concurrency, page B's response overwrites
`_CapturingAsyncClient.last_response` before page A's timeout handler reads it, and page A's
`ExecutionAttempt` is written with page B's transport evidence.
The consequence is **evidence mis-attribution** — a provenance-integrity failure, which this
project's own rubric treats as its most serious class. It would also be near-undetectable after
the fact: the attempt row would be well-formed, internally consistent, and wrong. The
application-level design that makes this safe (sequential pages) is not expressed in the
provider contract, so the constraint lives only in `workflows.py`'s loop structure.
- **Blast Radius:** Changing the protocol touches `providers/base.py:102-136`,
`providers/openrouter.py:221-231`, the two read sites at `workflows.py:369-370`, and the fakes in
`tests/providers/test_openrouter.py`, `tests/services/test_workflows_reliability.py`, and
`tests/test_provider_boundaries.py`, all of which implement or assert the current property-based
contract.
- **Recommendation:** **Fix before introducing any intra-job page concurrency.** The durable fix is
to stop returning evidence through instance state: attach `request_manifest` and
`transport_evidence` to the raised exception on every failure path — which `ProviderError`
already supports (`providers/base.py:18-29`) and which the timeout path cannot currently use
because `asyncio.wait_for` raises `TimeoutError` from outside the adapter. A narrower option is
to have `transcribe()` accept a caller-owned capture sink so evidence is scoped to the call
rather than to the adapter. As an immediate, near-zero-cost step, document the non-reentrancy on
the protocol in `providers/base.py` so the constraint is visible where it is depended upon.
- **Effort:** M
---
#### [MED-06] Provider error bodies reach user-facing text while three provider failure paths persist no `detail`
- **Location:** `src/transcription/services/sources.py:923-947` (`handle_transcription_errors`);
message construction at `src/transcription/providers/openrouter.py:414-431`
(`_transport_error_message`)
- **Reachability:** **Live** for the message half (any provider failure during a UI-initiated
transcription surfaces through `show_error`).
- **Problem & Consequence:** Two mirrored halves of the same rule are broken in one function.
- `sources.py:943` builds `f"Provider transcription failed: {exc}"`, and `exc` is a
`ProviderError` whose message may embed up to 500 characters of the provider's error body
(`openrouter.py:430`). That is a provider payload in `message`, which
`error-handling.instructions.md:85` explicitly forbids.
- None of the three handlers (lines 929, 935, 942) passes `detail=`. Per
`error-handling.instructions.md:89-92`, omitting it degrades the provenance record.
I checked whether the provenance half is actually harmful before reporting it, and it is
**substantially mitigated**: `workflows.py:391` calls `_find_provider_error`, which walks
`__cause__`/`__context__` (`workflows.py:806-813`) to recover the original `ProviderError` and
persists its `transport_evidence` — status code, safe headers, and the exact response body — onto
the attempt. So the root cause is preserved in transport evidence even though `error_detail` is
thin. This is why the finding is Medium rather than High. The residual cost is that the
human-readable failure summary is uninformative for the two paths (`ProviderAuthError`,
`ProviderResponseError`) whose messages are entirely generic.
- **Blast Radius:** `handle_transcription_errors` is used on the transcription path in
`sources.py`; `TranscriptionError.message` is consumed by `error_presenter.show_error`,
`build_error_envelope`, and `format_error_detail`. Assertions on these messages live in
`tests/services/test_transcription_service.py` and `tests/providers/test_openrouter.py`.
- **Recommendation:** Move the interpolated provider text from `message` to `detail` on all three
handlers, keeping the generic message the other two already use:
```python
except ProviderError as exc:
raise TranscriptionError(
"Provider transcription failed",
category=ErrorCategory.EXTERNAL_PROVIDER,
suggestion="Retry the transcription from jobs. If repeated, check provider availability.",
retriable=True,
detail=f"{type(exc).__name__}: {exc}",
) from exc
```
Apply the same `detail=` addition to the `ProviderAuthError` and `ProviderResponseError`
handlers. Note the interaction with HIGH-01: until the render boundary is sanitized, moving text
into `detail` still reaches users through the `error_detail` display. Sequence accordingly.
- **Effort:** S
---
#### [MED-07] No deterministic guard covers the `message`/`detail` split at explicit raise sites
- **Location:** `tests/test_errors.py:56-78`; rule at
`.github/instructions/error-handling.instructions.md:78-98`; canonical statement at
`docs/error_handling.md:102-115`
- **Reachability:** **Live** — this coverage gap is what allowed HIGH-02 and MED-06 to exist in a
fully green suite.
- **Problem & Consequence:** `docs/error_handling.md:115` names
`tests/test_errors.py::test_unexpected_error_does_not_leak_filesystem_paths` as the enforcement
for the message/detail split. That test exercises exactly one function,
`classify_unexpected_error`. Every direct `raise SomeAppError(...)` in `src/` — roughly 50 sites
by grep — is unenforced. The documentation therefore overstates the enforcement, which is worse
than having no guard: a contributor reading `error_handling.md:115` reasonably concludes the rule
is mechanically protected.
Per the reviewer skill, where a check is unenforced, recommending the deterministic test is
itself a finding.
- **Blast Radius:** Tests only.
- **Recommendation:** Add an AST guard, `tests/test_error_message_safety.py`, that scans `src/`
for `raise <AppError subclass>(...)` and fails when the first positional argument is an f-string
containing a formatted value whose name matches a path-like or exception-like identifier
(`path`, `_path`, `root`, `dir`, `exc`, `err`, `e`). Model it on the existing AST guards, which
are the established pattern here (`test_ui_boundaries.py`, `test_service_boundaries.py`,
`test_orphan_sweep.py`). Pair it with a second guard asserting that no UI module reads
`error_detail` without routing through the sanitizing projection from HIGH-01 — that one closes
the render side, which is where the real leak is.
- **Effort:** M
---
### Low Severity
#### [LOW-01] `ui/pages/tags_page.py` is dead code; the V6.1 roadmap is correct
- **Location:** `src/transcription/ui/pages/tags_page.py` (93 lines);
registration list at `src/transcription/ui/__init__.py:37-43`
- **Reachability:** **Not reachable.** This resolves the flagged open item: the route is genuinely
**not** registered. `register_pages` calls seven page registrars and `tags_page` is not among
them; nothing anywhere imports the module. `docs/roadmap_plan.md:47` ("Retire the Tags page") is
accurate, and `tests/ui/test_tags_page.py` correctly asserts `/ui/tags` returns 404 — though it
passes trivially, since an unimported module cannot register anything.
- **Problem & Consequence:** No runtime risk; purely stranded code. It is worth noting that if it
*were* ever re-registered, `tags_page.py:71-74` traverses `document.document_tags` and
`link.tag_ref` inside a page render, and those relationships are configured `lazy="raise"`
(`docs/architecture.md:200-203`) — so re-enabling this module without adding eager loads to
`list_documents` would raise on first render.
- **Recommendation:** Delete `src/transcription/ui/pages/tags_page.py`. Retain
`tests/ui/test_tags_page.py` as the retirement guard. Fixing MED-01 first would make this
finding reproducible by the suite rather than by manual inspection.
- **Effort:** S
#### [LOW-02] `benchmarking.py` ships in the runtime package but is referenced only by tests
- **Location:** `src/transcription/benchmarking.py` (69 lines); sole consumers
`tests/test_v42_evidence.py:15-16` (`EditorialAssessment`, `score_transcription`)
- **Reachability:** Live as importable API; never invoked by application code.
- **Problem & Consequence:** No defect. It supports the model-evaluation policy in
`docs/invariant/ai_evidence_and_provenance.md:113-126`, which is legitimate, but it currently has
no production caller and no tooling entrypoint, so it is indistinguishable from drift.
- **Recommendation:** Either move it under `tools/` alongside the other operator utilities, or add
a `KNOWN_ORPHANS`-style rationale recording that it is retained as the evaluation-policy
implementation. Do not silently keep it unlabeled.
- **Effort:** S
#### [LOW-03] Two overlapping prompt error types split across modules
- **Location:** `src/transcription/services/errors.py:15-16` (`PromptLoadError`) and
`src/transcription/services/prompts.py:19` (`PromptStoreError`)
- **Reachability:** Live; no misbehavior observed.
- **Problem & Consequence:** `services/errors.py:1-8` documents itself as the neutral home for
exceptions raised by more than one service, precisely so a caller's `except` clause does not
change when an operation moves. `PromptStoreError` is defined outside that module and covers an
overlapping domain (prompt file access), so a caller wanting to handle "any prompt failure" must
import from two modules and know which is which. `sources.py:855` raises `PromptLoadError` for a
missing prompt file while `prompts.py:132` raises `PromptStoreError` for the same condition
reached through the Settings page.
- **Recommendation:** Move `PromptStoreError` into `services/errors.py` next to `PromptLoadError`,
or make one a subclass of the other so a single `except` covers prompt failures. Low urgency; do
it opportunistically when HIGH-02 touches both files anyway.
- **Effort:** S
---
## 4. Architectural Drift & Gap Analysis
| Area / Component | Direction | Documented / Intended Rule | Actual Implementation State | Severity | Recommended Resolution |
| :--- | :--- | :--- | :--- | :--- | :--- |
| Error presentation | `doc->code` | `docs/error_handling.md:110` — `detail` is internal only, surfaced by `format_error_detail` and logs | `sources_page.py:484`, `table/sources.py:90`, `settings_page.py:562` render `error_detail` verbatim to users | High | Sanitizing render projection (HIGH-01); do **not** strip `detail` |
| User-facing messages | `doc->code` | `invariant/error_handling.md:59` — no local filesystem detail in user-facing messages | 5 live sites interpolate absolute paths into `AppError.message` | High | Generic `message`, path on `detail` (HIGH-02) |
| Evidence invariant ownership | `doc->doc` | `docs/index.md:1` — baseline is V6.1 | `invariant/ai_evidence_and_provenance.md:130` names "Canonical V4"; the currency guard scans the file but its regexes match neither the phrasing nor a minor-less `V4` | Medium | Update text; make the guard token-based, or drop version labels entirely (MED-02, §8.1) |
| Enforcement claim | `doc->code` | `docs/error_handling.md:115` — split "Enforced by `tests/test_errors.py::…`" | That test covers only `classify_unexpected_error`; explicit raises unguarded | Medium | Add AST guard (MED-07) |
| Orphan sweep | `doc->code` | `test_orphan_sweep.py:1-13` — sweep is "deterministic" and "conservative" | Bare-name matching; cannot see a dead module (`tags_page.py`) | Medium | Qualify by module + module-reachability pass (MED-01) |
| Worker scaling | `code->doc` | `jobs.py:191-197` — `SKIP LOCKED` so "concurrent workers never contend" | Claim path is multi-worker-safe; stale-reclaim path is not | Medium | Derive stale threshold from job duration; document single-worker constraint until fixed (MED-04) |
| Provider adapter contract | `code->doc` | `providers/base.py:110-118` — evidence read as "most recent call" state | Contract is silently non-reentrant; safety lives in `workflows.py`'s sequential loop | Medium | Scope evidence to the call; document non-reentrancy (MED-05) |
| Settings env file | `code->doc` | `config.py:66-75` — `env_file=".env.production"` | Correctness depends on an undocumented cwd contract with `Dockerfile` `WORKDIR /app` | Medium | Explicit resolver seam, or document + guard the contract (MED-03) |
| Tags page | *(no drift)* | `roadmap_plan.md:47` — Tags page retired | Route genuinely unregistered; module is stranded code | Low | Delete the module (LOW-01) |
---
## 5. Invariant Inventory & Routing Recommendations
| Invariant / Constraint | Current Location | Recommended Target Layer | Rationale |
| :--- | :--- | :--- | :--- |
| `detail`/`error_detail` never rendered to users | docs + instructions | **Deterministic test** + sanitizing projection | Stated in three documents and violated in three files; prose has demonstrably failed to hold it |
| `message` carries no paths or exception text | instructions; partial test | **Deterministic test** (AST, all raise sites) | Existing guard covers one function; the gap produced HIGH-02 |
| `ExecutionAttempt.error_detail` retains root cause | docs + `test_v42_evidence.py` | **Keep in tests** — already correct | Counterweight to the above; must be named in any fix so it is not over-corrected |
| Intermediate pages commit individually | `workflows.py` docstring + `test_workflows_reliability.py:318` | **Keep in tests** — verified genuine | Cross-session read makes it a real durability assertion |
| Final page atomic with terminal status | `services.instructions.md` + `test_pipeline_atomicity.py` | **Keep in tests** — verified genuine | Fault injection makes a split commit fail |
| Canonical baseline version consistency | `docs/index.md` + `test_meta_contract_guards.py:89` | **Repair existing test, or remove the labels** | Guard exists but matches by approved phrase and requires a minor version, so it fails open on new phrasings (MED-02) |
| Module-level reachability / dead modules | `test_orphan_sweep.py` (ineffective) | **Deterministic test** (repair existing) | Guard exists but cannot detect the case (MED-01) |
| Stale threshold > max job duration | *(unenforced)* | **Config validator + test** | Currently a coincidence of two equal defaults (MED-04) |
| Provider adapter non-reentrancy | *(unenforced, implicit)* | **Instructions** + protocol docstring | A design constraint callers must know before adding concurrency (MED-05) |
| Env-file resolution independent of cwd | `tests/conftest.py` monkeypatch | **Code seam** + `docs/production-runbook.md` | A test-only fix for a production coupling is misrouted enforcement (MED-03) |
---
## 6. Stack-Specific Analysis
**Python 3.12+.** Modern and consistent. PEP 695 generics are used correctly and non-trivially
(`RegistryService[ModelT: RegistryEntry]` in `services/registry.py:58`, `UiActionOutcome[T]`,
`_get_or_raise[ModelT]`), `type` statements appear in `db/session.py:15,50`, and `X | None` is
used throughout. `structural Protocol` bounds (`RegistryEntry`, `WorkerNotifier`,
`TranscriptionProvider`) are used to avoid type suppressions rather than to decorate. `ty` passes
clean with no suppressions found. The two `# noqa` uses (`workflows.py:228` `PLR0915`,
`workflows.py:383` `BLE001`) are both justified in context — the broad catch is a deliberate
per-page containment boundary that immediately classifies and re-records.
**FastAPI.** Lifespan is handled via `@asynccontextmanager` (`app.py:36`), not the deprecated
`@app.on_event`. Session factories are injected through `Depends` (`SessionFactoryDep`,
`db/session.py:50`) rather than reached as globals from routes. `api/errors.py` centralizes
envelope translation. One residual: `get_settings` is `@cache`d and read as a module-level
fallback in ~10 modules; this is acceptable given the documented restart-to-apply contract
(`docs/ui/pages/settings.md:28`) but means the cache is process-lifetime and unclearable.
**NiceGUI (pinned `3.13.0`).** The pin is a recorded release-stability decision and is not
reported as a defect. Boundaries are enforced structurally: `test_ui_boundaries.py` uses an
import **allowlist**, which is the right polarity. Blocking work is dispatched off the event loop
via `run_blocking` (`settings_page.py:818,822`). The one boundary that is *not* enforced is
presentation of internal fields (HIGH-01) — pages are prevented from touching persistence but not
from rendering internal-only text.
**SQLModel / SQLAlchemy.** Strong. `lazy="raise"` on relationships forces explicit eager loading;
read paths declare `selectinload` chains with comments explaining *why* each is needed
(`sources.py:309-316` is a good example). `expire_on_commit=False` (`db/session.py:28`) is set
deliberately, which is what makes post-commit attribute access in `evidence.py:159-202` safe.
`claim_next_queued_job` (`jobs.py:186-241`) branches correctly on dialect — `SKIP LOCKED` on
PostgreSQL, conditional `UPDATE ... RETURNING` on SQLite — rather than assuming one engine.
Attempt-number allocation uses `begin_nested()` with bounded retry (`sources.py:598-616`), the
right pattern for a monotonic per-parent sequence. No N+1 patterns were found in the read paths
sampled.
**Pydantic V2 & Settings.** Fully V2; no `@validator`, `class Config`, `.dict()`, or `parse_obj`
anywhere. Evidence contracts use `ConfigDict(extra="forbid", frozen=True)` (`providers/evidence.py:47`),
which is exactly right for persisted provenance — an unexpected field fails loudly rather than
being silently dropped. `SecretStr` guards the API key. The discriminated
`SqliteSettings | PostgresSettings` union is clean. `normalize_provider_models` correctly runs
`mode="before"` so the derived tuple is produced by construction rather than by mutating a frozen
model — a subtlety that is easy to get wrong. Sole issue: the cwd-coupled `env_file` (MED-03).
**Asyncio Workers.** Notably careful. `asyncio.shield` wraps both the per-page commit and the
terminal commit (`workflows.py:601-614`, `650-663`), with the `except CancelledError: await task;
raise` pattern that actually completes the shielded work rather than merely deferring cancellation —
a detail most implementations get wrong. `handle_worker_exceptions` (`worker.py:157-182`)
distinguishes retriable from non-retriable faults and stops the loop rather than spinning.
`_advance_job_with_containment` (`worker.py`/`workflows.py:507-540`) guarantees a claimed job
cannot strand in `PROCESSING`. `worker_consumer_lifespan` has a bounded shutdown with escalation to
`cancel()`. Gaps are MED-04 and MED-05, both latent and both with stated unblocking conditions.
**OpenRouter / Adapter Boundary.** Encapsulation holds: `test_provider_boundaries.py` enforces it,
and `workflows.py` imports only `providers` abstractions. `_CapturingAsyncClient` is a
well-judged design — it captures the exact transport body before SDK parsing without altering what
the SDK consumes, including the streamed case. Timeout construction (`openrouter.py:200-206`)
correctly overrides httpx's 5s per-phase default that would otherwise silently cap the configured
budget. `SAFE_RESPONSE_HEADERS` (`providers/evidence.py:29-41`) was reviewed field-by-field:
all nine entries are non-secret correlation, content, or rate-limit headers, and
`filter_safe_response_headers` is a true allowlist filter with no redaction-after-capture — this
satisfies invariant 3.8.2 exactly. `_replace_embedded_media` correctly substitutes a source
reference for base64 payloads, satisfying 3.8.3. The one structural weakness is MED-05.
**Testing & Quality Tooling.** 410 tests, all green, with genuinely strong contract guards
(boundaries, model contract, media path safety, evidence append-only, atomicity). Marker strictness
and `asyncio_mode = "strict"` are configured, and no unawaited-coroutine warnings appeared. Two
guards, however, assert meaningfully less than their names and docstrings claim
(`test_orphan_sweep.py` — MED-01; `test_errors.py` path-leak coverage — MED-07), and the
`>= 420` snapshot threshold at `test_orphan_sweep.py:169` repeats a weak-assertion pattern the
2026-08-23 review already flagged at `>= 200`; it was raised rather than replaced with
set-membership.
---
## 7. Duplication & Consolidation Report
| Pattern / Duplication | Locations | Proposed Canonical Home | Estimated Lines Removed |
| :--- | :--- | :--- | :--- |
| `f"{type(exc).__name__}: {exc}"` detail construction | `errors.py:118`, `maintenance.py:71,95,135,206`, `runtime_settings_store.py:388,476,554` | `errors.py::exception_detail(exc)` | ~8 (consistency > line count) |
| Filesystem `AppError` construction from `OSError` | `prompts.py:186-191`, `runtime_settings_store.py:384-389,472-477,550-555` | `errors.py::filesystem_error(message, exc, *, suggestion)` | ~20 |
| Overlapping prompt error types | `services/errors.py:15`, `services/prompts.py:19` | `services/errors.py` (LOW-03) | ~5 |
| Duplicated `provider_duration_ms` / `processing_duration_ms` max-clamp arithmetic | `workflows.py:341-345, 361-368, 397-404` | `workflows.py::_page_durations(started_at, finished_at, monotonic_started_at)` | ~20 |
| `_utc_now_naive` defined per module | `workflows.py:51`, `jobs.py:29`, `db/models.py`, `sources.py` | Single helper in `db/models.py`, imported | ~12 |
### Proposed Canonical Abstractions
```python
# src/transcription/errors.py
def exception_detail(exc: BaseException) -> str:
"""Internal-only root-cause text for AppError.detail. Never user-facing."""
def filesystem_error[E: AppError](error_type: type[E], message: str, exc: OSError, *, suggestion: str) -> E:
"""Build a filesystem AppError with a generic message and the path on detail."""
# src/transcription/ui/components/error_presenter.py
def display_failure_detail(error_detail: str | None) -> str | None:
"""Sanitize persisted failure detail for UI rendering (HIGH-01)."""
```
---
## 8. Meta-Tooling & Instruction Update Recommendations
1. **`docs/invariant/ai_evidence_and_provenance.md:130`** — resolve the V4 label. Two viable
routes, and the maintainer has proposed the second:
- **(a) Repair the guard.** Fix the text to V6.1 and make
`test_canonical_docs_declare_one_consistent_baseline` token-based rather than phrase-based
(MED-02). Keeps version labels as navigational anchors.
- **(b) Remove version labels from canonical docs.** While the project has a single principal
user and no released versions to support, "canonical" and "current" are the same thing, so the
label carries no information a reader can act on — it only creates a second thing to keep in
sync. Retain the baseline declaration in `docs/index.md` alone as the release marker, keep
version language in `docs/roadmap_plan.md` and the migration/deployment docs (already
excluded from the scan for exactly this reason), and replace in-body references with
unversioned phrasing ("the canonical architecture, schema, requirements, and error-policy
documents"). The guard then inverts: assert that no canonical doc outside the exclusion set
contains a version token at all, which is a stricter and much cheaper property to hold than
agreement between many labels. Requirement IDs (`REQ-4-021`, `REQ-6-014`) are stable
identifiers, not currency claims, and should be left alone.
2. **`docs/error_handling.md:107-115`** — either add the sanitizing-projection rule for UI display
of `error_detail`, or revise the `detail` "Surfaces" row to admit operator-facing evidence
displays. Update the "Enforced by" line once MED-07's guard lands, since it currently overstates
coverage.
3. **`.github/instructions/error-handling.instructions.md`** — add an explicit clause under
"User-Safe Messaging" stating that *persisted* `error_detail` is subject to the same no-paths
rule at any render boundary. The current table (line 86) states the rule for `AppError.detail`
and stops there, so the persisted-then-rendered path falls between the lines.
4. **`.github/instructions/providers.instructions.md`** — record the adapter non-reentrancy
constraint (MED-05); it is currently an undocumented precondition of `workflows.py`.
5. **`.github/instructions/services.instructions.md`** — the two competing atomicity invariants are
well described and both guards verified; no change needed. Worth adding the stale-reclaim
threshold constraint (MED-04) alongside them, since it is a third worker-lifecycle rule with no
documented home.
6. **`tests/test_orphan_sweep.py`** — repair per MED-01 and replace the `>= 420` threshold with
set-membership assertions.
7. **New `tests/test_error_message_safety.py`** — AST guard per MED-07, covering both the raise
sites and the UI render sites.
8. **`docs/production-runbook.md`** — document the cwd/`WORKDIR` contract for `.env.production`
resolution if MED-03 is resolved by documentation rather than by a code seam.
---
## 9. Prioritized Dependency-Ordered Action Plan
**Phase 1 — Blocking fixes (privacy; ordered, HIGH-01 first)**
1. **HIGH-01** — add `display_failure_detail` and route `sources_page.py:484`,
`table/sources.py:90-95`, and `settings_page.py:562` through it. Do this **first**: it closes
the render boundary, so the Phase-1.2 fix cannot relocate a leak again.
2. **HIGH-02** — move paths from `message` to `detail` at the five sites, including the
`prompts.py:186` double violation.
3. **MED-06** — move provider payload text to `detail`; add `detail=` to all three
`handle_transcription_errors` handlers.
**Phase 2 — Enforcement hardening (make Phase 1 permanent)**
4. **MED-07** — AST guard for raise-site `message` safety **and** for UI reads of `error_detail`.
5. **MED-01** — qualify orphan definitions by module; add module-reachability; replace the
snapshot threshold.
6. **MED-02** — fix the V4/V6.1 text and extend the meta-contract guard to baseline-version currency.
**Phase 3 — Reliability & concurrency (latent; each must precede its unblocking change)**
7. **MED-04** — derive `worker_stale_job_seconds` from `worker_provider_timeout_seconds` with a
rejecting validator, ideally plus a progress heartbeat. **Must land before any second worker
replica.**
8. **MED-05** — scope provider evidence to the call rather than the instance. **Must land before
any intra-job page concurrency.** Document non-reentrancy immediately as an interim step.
**Phase 4 — Consolidation & refactoring**
9. **MED-03** — explicit env-file resolution seam shared by `Settings` and `runtime_settings_store`;
remove the `model_config` monkeypatch from `conftest.py`.
10. **LOW-01** — delete `tags_page.py` (after MED-01, so the suite reproduces the finding).
11. **LOW-03** and the §7 consolidations — fold in opportunistically while Phase 1 touches these files.
**Phase 5 — Non-blocking governance/documentation depth**
12. **LOW-02** — relocate or annotate `benchmarking.py`.
13. Instruction/doc updates §8.3–§8.5, §8.8.
---
## 10. Preserved Strengths
- **Append-only evidence is real, not aspirational.** Every provider call produces a distinct
`ExecutionAttempt`; no runtime path mutates a historical row. Projection writes onto
`Source.raw_transcription` are clearly separated from history, and `promote_machine_attempt`
(`evidence.py:117-146`) repoints the projection without rewriting evidence — with a docstring
that explains exactly why that one write lives in a read-oriented service.
- **Transport-layer terminology is honored in code.** `_CapturingAsyncClient` exists specifically so
the stored body is the application-boundary capture rather than an SDK-parsed object, and
`TransportEvidence(response_received=False)` explicitly represents "no response" instead of
conflating it with an empty one. This is invariant 3.4/3.5 implemented rather than asserted.
- **Header allowlisting is done the hard, correct way** — filter-before-store with an explicit
frozenset, never capture-then-redact (`providers/evidence.py:29-41,130-134`).
- **Boundaries are enforced by allowlist, not blocklist.** `test_ui_boundaries.py:20-25` states the
reasoning explicitly; it means a newly added persistence helper cannot slip through under an
unlisted name.
- **The two competing atomicity invariants are both correctly implemented and both genuinely
guarded**, with the tests structured so that the naive over-correction fails.
- **Cancellation safety in the worker is unusually well handled** — `asyncio.shield` plus
`await task` on `CancelledError` actually completes the commit rather than merely deferring
cancellation.
- **Comments explain rationale, not mechanics.** `workflows.py:269-271`, `openrouter.py:200-202`,
`config.py:114-115`, and `jobs.py:191-197` each record *why* a non-obvious choice was made,
several citing the review log entry that motivated it. This is what made verifying the atomicity
and timeout invariants tractable in this review.
- **Documentation-to-code traceability is strong overall.** Page contracts, schema field tables,
and requirement IDs are maintained and guarded; the drift found in this review is narrow and
specific rather than systemic.
---
## Appendix A — Repo-Specific Deterministic Checks
| # | Check | Result | Evidence |
| :-- | :--- | :--- | :--- |
| 1 | Service boundary rule: no service-to-service imports | **Pass** | `tests/test_service_boundaries.py` green; AST scan, allowlist-based; `workflows.py` composes via `ServiceBundle` |
| 2 | UI boundary rule: no persistence access from pages/components | **Pass (structurally)** | `tests/test_ui_boundaries.py` green. Caveat: it guards *data access*, not presentation of internal-only fields — see HIGH-01 |
| 3 | Status vocabulary conformance; no stringly-typed literals | **Pass** | `tests/test_model_contract_guards.py` green; enum members verified against `db/models.py` |
| 4 | Evidence ownership: append-only history, projections not history mutation | **Pass** | `test_v42_evidence.py::test_attempts_are_append_only_and_exported_with_integrity` verified non-vacuous (asserts both retained attempts and export integrity at lines 281-290) |
| 5 | Canonical authority: findings resolve against `docs/*` first | **Pass with defect** | `test_canonical_authority_references_are_present` green. The companion baseline-currency guard (`test_canonical_docs_declare_one_consistent_baseline`) scans the offending file but fails open on its phrasing and on minor-less version tokens — MED-02 |
| 6 | Schema contract fidelity: `docs/schema.md` field-accurate | **Pass** | `test_model_contract_guards.py` + `test_meta_contract_guards.py` green |
| 7 | Media boundary: record-validated media, controlled URL resolver | **Pass** | `test_media_path_safety.py`, `tests/ui/test_media_urls.py` green; `public_media_path_label` verified path-safe |
| 8 | Eager-loading conformance vs `lazy="raise"` | **Pass** | Declaration-side guard green; sampled read paths declare explicit `selectinload` chains. Note: dead `tags_page.py:71-74` would violate it if re-registered (LOW-01) |
| 9 | Cross-cutting error conformance | **FAIL** | Guards green but coverage is narrower than documented: HIGH-01, HIGH-02, MED-06, MED-07 |
| 10 | Orphan/dead-code conformance | **FAIL** | Guard green but structurally unable to detect a dead module: MED-01, proven by LOW-01 |
## Appendix B — Evidence & Provenance Auditor Families
| Family | Subject | Result | Evidence |
| :--- | :--- | :--- | :--- |
| A | Attempt history append-only | **Pass** | No update/delete path to `ExecutionAttempt`; insert-only with `begin_nested` + bounded sequence retry (`sources.py:556-616`) |
| B | Attempt numbering monotonic per source | **Pass** | `insert_with_sequence_retry`; uniqueness constraint plus retry on conflict |
| C | Transport evidence captured at the transport boundary | **Pass** | `_CapturingAsyncClient` retains the exact wire body pre-SDK-parse (`openrouter.py:66-94`) |
| D | Absent response distinguished from empty response | **Pass** | `TransportEvidence.response_received` is explicit, not inferred |
| E | Response header persistence is allowlist-based | **Pass** | `SAFE_RESPONSE_HEADERS` (`providers/evidence.py:29-41`) — all nine entries verified non-secret; filter-before-store |
| F | No machine-local detail on user-facing surfaces | **FAIL** | `error_detail` rendered verbatim at three UI sites (HIGH-01); paths in `message` at five sites (HIGH-02) |
| G | Request manifest excludes embedded media payloads | **Pass** | `_replace_embedded_media` (`openrouter.py:377-395`) substitutes a source reference for base64 data |
| H | Evidence attribution is correct under concurrency | **Pass today / at risk** | Correct in the current sequential single-worker deployment; the contract itself is non-reentrant (MED-05) and reclaim has no margin (MED-04) |
+21
View File
@@ -0,0 +1,21 @@
# Review Reports
Dated architecture and code review reports generated by
`.github/skills/python-code-reviewer/skill.md`.
**These files are not canonical authority.** Everything in `docs/reviews/**` is a
point-in-time observation, not a contract. Canonical intent lives in `docs/index.md`,
`docs/architecture.md`, `docs/requirements.md`, `docs/schema.md`,
`docs/error_handling.md`, and `docs/invariant/**`. When a report and a canonical
document disagree, the canonical document wins until it is deliberately updated.
Naming: `<YYYY-MM-DD>-code-review.md` for review reports, and
`<YYYY-MM-DD>-remediation-handoff.md` for the implementation plan derived from one.
## Current
- [`2026-08-23-code-review.md`](./2026-08-23-code-review.md) — full review. 0 critical,
4 high, 5 medium, 11 low. **All findings remediated.** Retained as a record of the
reasoning, not as a list of open work. Note that a few of its recommendations were
wrong on contact and were corrected during implementation; the code and the guard
tests are authoritative over the report text.
+210
View File
@@ -0,0 +1,210 @@
# Roadmap Plan (Starting at V6.0)
This roadmap starts at **V6.0** and tracks forward-looking work only.
## V6.0 - Hosting Migration
Objective: move from local-only operation to secure, stable remote hosting.
Status: **Completed**
Detailed plan: [`v6_0_hosting_migration_plan.md`](v6_0_hosting_migration_plan.md)
### Scope
1. Containerize app runtime for production deployment.
2. Run PostgreSQL in Docker and migrate from SQLite.
3. Add Cloudflare Tunnel exposure with Access protection.
4. Add operational safeguards (health checks, restart policies, backups).
### Deliverables
- Production-ready `docker-compose` deployment for app + database + tunnel.
- Environment-based configuration for DB, uploads, prompts, and logging.
- Verified data migration path into PostgreSQL.
- Runbook updates for deploy, rollback, and backup/restore.
### Exit Criteria
- `/healthz` reports healthy app and worker in deployed environment.
- One end-to-end document -> source -> job workflow succeeds remotely.
- Backup and restore procedure is tested.
### Accomplished
1. Delivered production Docker deployment with split `app`/`worker`, `postgres`, and `cloudflared`.
2. Landed SQLite -> PostgreSQL migration tooling and runbook coverage.
3. Added production health/reliability wiring and operational runbooks for deploy/rollback/recovery.
4. Established host-visible backup workflow and restore path for PostgreSQL plus media/config assets.
## V6.1 - Testing and Refinement
Objective: improve navigation and operational workflows after user feedback.
Status: **Completed**
### Scope
1. Make Document Detail the primary source-page workspace:
- Use Source-style pan/zoom + previous/next page controls.
- Move editable revision controls into Document Detail.
- Move archival/system metadata to dedicated Document Info route.
2. Simplify top navigation:
- Remove top-level Tags and Sources entries.
- Retire the Tags page and the global Source Asset Records entry flow.
3. Improve list/detail clarity:
- Add Document transcription status to Archival Documents list.
- Add Document Date in People Detail -> Linked Documents table.
4. Add worker-backed Settings maintenance runs:
- Add `maintenance_run` persistence (`id`, `job_type`, `status`, `started_at`, `finished_at`, `triggered_by`, `summary`, `log_path`, `error_detail`).
- Add Run Backup and Run Storage Reconciliation actions that enqueue runs and execute in the worker.
- Add run history with status, duration, summary, and log view/download.
- Defer daily/weekly scheduling controls to V6.2.
### Accomplished
1. Refactored Document Detail into the primary source-page workspace (pan/zoom viewer, previous/next page navigation, editable revision flow) and moved archival/system metadata to Document Info.
2. Simplified top navigation by removing Tags/Sources entries and retiring the Tags page/global Source Asset Records flow.
3. Improved data clarity with document transcription status in Archival Documents and Document Date in People Detail linked documents.
4. Implemented queue-backed maintenance operations (`maintenance_run` model/service/worker/UI) with run history and log view/download.
5. Hardened runtime settings operations in production:
- runtime settings writes target mounted `.env.production`,
- fallback write path for single-file bind mounts,
- explicit hidden/deployment-key disclosure in Settings UI.
6. Simplified backup configuration and behavior:
- standardized on `BACKUP_DIR` + `BACKUP_RETENTION_DAYS`,
- backup script uses `DATABASE__*` persistence keys,
- compose maps Postgres container init values from `DATABASE__*`,
- env contract drift tests now guard `.env.production.example`.
## V6.2 - GEDCOM Data Layer
Objective: introduce a genealogical data layer sourced from GEDCOM exports, bridged to
existing `Person` records via FamilySearch ID, without disrupting document-focused Person
workflows.
### Scope
1. Manual `.ged` file upload only. No FamilySearch credentials are stored or used by the
app; the user runs the third-party `getmyancestors` tool themselves and uploads the
resulting export.
2. Four new tables: `genealogy_person`, `genealogy_family`, `genealogy_family_child`, and
`genealogy_citation` (raw GEDCOM `SOUR` citations, reusable in a later version to record
when a transcribed document itself becomes citation evidence for FamilySearch).
3. Upsert-based import keyed on FamilySearch ID (`fs_id`) so repeat imports update existing
records in place without breaking existing `Person.family_search_id` links or duplicating
surrogate keys.
4. Reuse the existing V6.1 worker-backed `maintenance_run` pattern for import runs (run
history, status, summary, log view/download) rather than new infrastructure.
### Deliverables
- GEDCOM parser/importer producing the four genealogy tables.
- `MaintenanceJobType` entry for GEDCOM import with upsert semantics and a run summary
(records added/updated).
- Settings UI entry to upload a `.ged` file, trigger an import run, and view history.
### Exit Criteria
- Importing the same `.ged` file twice does not duplicate or orphan data.
- Existing `Person.family_search_id` values continue to resolve to the correct
`genealogy_person` row after import.
- Import run history is visible with status, duration, and summary, consistent with other
maintenance runs.
## V6.3 - Reporting and Genealogy-Enriched Features
Objective: improve research value with person-centric outputs, grounded in both archival
documents and the V6.2 genealogical data layer.
This version is broken into five sequential sub-versions because of real dependency
ordering: entity linking must exist before GEDCOM data can be targeted per-person; the
Facts/Events mechanism must exist before timelines or reconciliation have anything
meaningful to consume.
### V6.3.1 - Manual Entity Linking
- Search/browse UI over `genealogy_person` to find and link a candidate match to an
application `Person`, setting `family_search_id`. Linking is reversible (unlink).
- Once linked, GEDCOM vitals display alongside the `Person` record without requiring any
schema change to `Person`.
### V6.3.2 - Person Facts and Events
- New fact/event table capturing: person, fact type (birth/death/event/free-form), date
(+raw), place, free-text description, and a link to the source document as evidence.
- Manual tagging UI while reviewing a transcribed document: select a passage, choose the
person and fact type, record the date/description.
- One-time migration of existing `Person.birth_date`/`birth_date_raw`/`birth_place`/
`death_date`/`death_date_raw`/`death_place` values into fact/event rows (tagged as
legacy/no-document-evidence where no source document is known), followed by retiring those
six columns from `Person`. Birth/death become Facts/Events like any other locally-known
fact, for both linked and unlinked people. `Person` permanently keeps `last_name`,
`given_names`, `biography`, `family_search_id`, `metadata_`, tags, photos, and document
associations.
### V6.3.3 - Person Timelines
- Timeline query merging GEDCOM milestones (birth, marriage, children's births, death) for
linked persons with locally recorded Facts/Events.
- Timeline UI on Person Detail with clear ordering/filters; entries link back to their
originating document or GEDCOM record.
### V6.3.4 - Reconciliation
- Compares Facts/Events (the real, document-evidenced local signal) against corresponding
`genealogy_person` fields for linked persons.
- Persisted reconciliation record: person, field, local value with evidence-document link,
GEDCOM value, and status (open / submitted / dismissed).
- Re-evaluated automatically as part of each GEDCOM import maintenance run: opens new
discrepancies, auto-resolves ones where GEDCOM now matches, leaves others unchanged.
- Reconciliation review UI functions as a manual to-do list for updating FamilySearch; the
app does not write back to FamilySearch itself.
### V6.3.5 - AI-Assisted Biography Generation
- Prompted narrative generation grounded in GEDCOM facts, Facts/Events, and relevant
document snippets as structured input, using existing evidence-safe prompting patterns.
- Output cites back to source documents and FamilySearch records.
- Saved/printable report presentation for review; reports do not modify archival source
data.
### Exit Criteria (applies across V6.3.1-V6.3.5)
- Entity links are reversible and do not alter document associations.
- Timelines are reproducible from persisted records.
- Reconciliation items always carry a link to the document evidence justifying the local
value, and re-running GEDCOM import correctly opens, resolves, or leaves items unchanged.
- Narrative generation is traceable to source records and prompts.
- Reports can be reviewed without modifying archival source data.
## V6.4 - Access Control and Multi-User Readiness
[ *More thoughts on user accounts:*
* *Create a generic "view only" user that does not have the rights to alter any of the data*
* *Limit user accounts access to data by Tag. I have distant family members that I would want to share the transcribed data with, but they would only be interested in a subset of it. For example my Cochran cousins would have no interest in Lancaster documents, so limit the Cochra Clan cousins to view-only access to documents tagged "cochran clan"* ]
Objective: prepare for managed collaboration beyond single-user operation.
### Scope
1. Introduce application-level authentication.
2. Add role-based authorization (admin/editor/contributor/viewer).
3. Add audit visibility for user-attributed write actions.
### Deliverables
- User identity model and login/session flow.
- Route/page/service authorization enforcement.
- Audit metadata for sensitive create/update/delete workflows.
### Exit Criteria
- Unauthorized operations are blocked consistently across UI/API.
- Role policies are enforced by deterministic tests.
- User-attributed changes are visible for audit/review.
## Deferred / Future Ideas (not committed scope)
Captured for later consideration, not yet scheduled to a version:
* AI-assisted entity disambiguation (kinship co-occurrence, chronological plausibility
filtering) when linking document mentions to people.
* Kinship-aware `@mention` tagging while transcribing.
* Relationship-calculator badges (e.g., "3rd Great-Grandmother") in the document viewer.
* Interactive migration/geography mapping from GEDCOM and document place mentions.
* AI-suggested document discovery by date/location overlap with known persons.
* Ability to search within a document to find potential people to add to the People table.
## Planning Notes
- Keep architecture, schema, and UI contracts synchronized in `docs/` as each version lands.
- Prefer explicit schema migration over runtime compatibility write paths.
- Preserve evidence/provenance guarantees when adding new AI-powered features.
- GEDCOM/FamilySearch data is external, collaborative, and mutable; treat it as a managed
cache bridged via `fs_id`, never as a replacement for archival evidence recorded from
transcribed documents.
+391
View File
@@ -0,0 +1,391 @@
# Data Model and Persistence Schema (Current Baseline: V6.1)
This document is the field-accurate V6.1 schema contract aligned to `src/transcription/db/models.py`.
## Source of Truth Anchors
- `src/transcription/db/models.py` (status and purpose enums, including maintenance lifecycle enums)
- `src/transcription/db/models.py:80-120` (`DocumentType`, `PersonRole`)
- `src/transcription/db/models.py:122-172` (`Tag`, `Document`)
- `src/transcription/db/models.py` (`Person`, `GenealogyPerson`, `GenealogyFamily`, `GenealogyFamilyChild`, `GenealogyCitation`)
- `src/transcription/db/models.py` (`Photo`, `DocumentPerson`, `DocumentTag`)
- `src/transcription/db/models.py:285-347` (`Job`)
- `src/transcription/db/models.py` (`MaintenanceRun`)
- `src/transcription/db/models.py:350-462` (`Source`, `JobSource`)
- `src/transcription/db/models.py:465-522` (`ExecutionAttempt`)
## Entity Relationship Overview
```mermaid
erDiagram
DocumentType ||--o{ Document : classifies
Document ||--o{ Job : has
Document ||--o{ Source : has
Document ||--o{ DocumentPerson : links
Document ||--o{ DocumentTag : tagged
Person ||--o{ DocumentPerson : links
Person ||--o{ PersonTag : tagged
Person ||--o{ Photo : owns
GenealogyPerson ||--o{ GenealogyFamily : husband
GenealogyPerson ||--o{ GenealogyFamily : wife
GenealogyPerson ||--o{ GenealogyFamilyChild : child
GenealogyFamily ||--o{ GenealogyFamilyChild : includes
GenealogyPerson ||--o{ GenealogyCitation : cited
GenealogyFamily ||--o{ GenealogyCitation : cited
Document ||--o{ GenealogyCitation : evidence
PersonRole ||--o{ DocumentPerson : labels
Tag ||--o{ DocumentTag : labels
Tag ||--o{ PersonTag : labels
Job ||--o{ JobSource : includes
Source ||--o{ JobSource : participates
JobSource ||--o{ ExecutionAttempt : attempts
MaintenanceRun {
uuid id PK
}
```
## Authoritative Enumerations
### JobStatus
- `queued`
- `processing`
- `transcribed`
- `partial_success`
- `failed`
### JobSourceStatus
- `pending`
- `transcribed`
- `failed`
- `cancelled`
### JobPurpose
- `transcription`
- `retranscription`
### MaintenanceJobType
- `backup`
- `storage_reconciliation`
- `gedcom_import`
### MaintenanceRunStatus
- `queued`
- `processing`
- `succeeded`
- `failed`
## Field-Accurate Table Contracts
### `DocumentType`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `semantic_key` | `str \| None` | nullable unique, indexed |
| `label` | `str` | required |
| `normalized_label` | `str` | unique, indexed |
| `is_active` | `bool` | default `True` |
| `created_at` | `datetime` | default now |
| `updated_at` | `datetime` | default now, onupdate |
### `PersonRole`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `semantic_key` | `str \| None` | nullable unique, indexed |
| `label` | `str` | required |
| `normalized_label` | `str` | unique, indexed |
| `is_active` | `bool` | default `True` |
| `created_at` | `datetime` | default now |
| `updated_at` | `datetime` | default now, onupdate |
### `Tag`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `semantic_key` | `str \| None` | nullable unique, indexed |
| `label` | `str` | required |
| `normalized_label` | `str` | unique, indexed |
| `is_active` | `bool` | default `True` |
| `created_at` | `datetime` | default now |
| `updated_at` | `datetime` | default now, onupdate |
### `Document`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `name` | `str` | required |
| `document_type_id` | `UUID \| None` | FK -> `document_type.id`, indexed |
| `document_date` | `date \| None` | optional |
| `document_date_raw` | `str \| None` | optional |
| `location_created` | `str \| None` | optional |
| `notes` | `str \| None` | optional |
| `archive_identifier` | `str \| None` | optional |
| `created_at` | `datetime` | default now |
| `updated_at` | `datetime` | default now, onupdate |
### `Person`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `last_name` | `str` | required |
| `given_names` | `str` | required |
| `birth_date` | `date \| None` | optional |
| `birth_date_raw` | `str \| None` | optional |
| `birth_place` | `str \| None` | optional |
| `death_date` | `date \| None` | optional |
| `death_date_raw` | `str \| None` | optional |
| `death_place` | `str \| None` | optional |
| `biography` | `str \| None` | optional |
| `family_search_id` | `str \| None` | nullable unique |
| `metadata_` | `dict[str, JsonValue] \| None` | stored as DB column `metadata` (`JSONBCompat`) |
| `created_at` | `datetime` | default now |
| `updated_at` | `datetime` | default now, onupdate |
### `GenealogyPerson`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `fs_id` | `str` | unique, indexed FamilySearch identifier |
| `full_name` | `str` | required |
| `birth_date` | `date \| None` | optional |
| `birth_date_raw` | `str \| None` | optional |
| `birth_place` | `str \| None` | optional |
| `death_date` | `date \| None` | optional |
| `death_date_raw` | `str \| None` | optional |
| `death_place` | `str \| None` | optional |
| `created_at` | `datetime` | default now |
| `updated_at` | `datetime` | default now, onupdate |
### `GenealogyFamily`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `fs_family_id` | `str` | unique, indexed FamilySearch family identifier |
| `husband_id` | `UUID \| None` | nullable FK -> `genealogy_person.id`, indexed |
| `wife_id` | `UUID \| None` | nullable FK -> `genealogy_person.id`, indexed |
| `marriage_date` | `date \| None` | optional |
| `marriage_date_raw` | `str \| None` | optional |
| `marriage_place` | `str \| None` | optional |
| `created_at` | `datetime` | default now |
| `updated_at` | `datetime` | default now, onupdate |
### `GenealogyFamilyChild`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `family_id` | `UUID` | FK -> `genealogy_family.id`, indexed |
| `child_id` | `UUID` | FK -> `genealogy_person.id`, indexed |
| `relationship_type` | `str \| None` | optional |
| `created_at` | `datetime` | default now |
Constraint:
- `UniqueConstraint(family_id, child_id)` named `uq_genealogy_family_child`
### `GenealogyCitation`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `genealogy_person_id` | `UUID \| None` | nullable FK -> `genealogy_person.id`, indexed |
| `genealogy_family_id` | `UUID \| None` | nullable FK -> `genealogy_family.id`, indexed |
| `fact_type` | `GenealogyCitationFactType` | enum: `birth`, `death`, `marriage`, `other` |
| `raw_citation_text` | `str` | required raw GEDCOM citation text |
| `source_kind` | `GenealogyCitationSourceKind` | enum: `familysearch_imported`, `transcription_evidence` |
| `document_id` | `UUID \| None` | nullable FK -> `document.id`, indexed |
| `created_at` | `datetime` | default now |
### `Photo`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `person_id` | `UUID \| None` | nullable FK -> `person.id`, indexed (`NULL` = homepage photo) |
| `path` | `str` | required upload-root-relative POSIX path (`photos/...`) |
| `description` | `str \| None` | optional |
| `is_primary` | `bool` | default `False`; owner-level "featured/primary" marker |
| `created_at` | `datetime` | default now |
| `updated_at` | `datetime` | default now, onupdate |
### `DocumentPerson`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `document_id` | `UUID` | FK -> `document.id`, indexed |
| `person_id` | `UUID` | FK -> `person.id`, indexed |
| `role_id` | `UUID` | FK -> `person_role.id`, indexed |
| `created_at` | `datetime` | default now |
| `updated_at` | `datetime` | default now, onupdate |
Constraint:
- `UniqueConstraint(document_id, person_id)` named `uq_document_person`
### `DocumentTag`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `document_id` | `UUID` | FK -> `document.id`, indexed |
| `tag_id` | `UUID` | FK -> `tag.id`, indexed |
| `created_at` | `datetime` | default now |
| `updated_at` | `datetime` | default now, onupdate |
Constraint:
- `UniqueConstraint(document_id, tag_id)` named `uq_document_tag`
### `PersonTag`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `person_id` | `UUID` | FK -> `person.id`, indexed |
| `tag_id` | `UUID` | FK -> `tag.id`, indexed |
| `created_at` | `datetime` | default now |
| `updated_at` | `datetime` | default now, onupdate |
Constraint:
- `UniqueConstraint(person_id, tag_id)` named `uq_person_tag`
### `Job`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `document_id` | `UUID` | FK -> `document.id`, indexed |
| `status` | `JobStatus` | non-null enum (stored as enum values) |
| `retry_count` | `int` | default `0`, `ge=0` |
| `purpose` | `JobPurpose` | non-null enum, default `transcription` |
| `date_created` | `datetime` | default now |
| `date_updated` | `datetime` | default now, onupdate |
| `provider` | `str \| None` | optional |
| `model` | `str \| None` | optional |
| `prompt_name` | `str \| None` | optional |
| `prompt_hash` | `str \| None` | optional |
| `system_prompt` | `str \| None` | optional |
| `user_prompt` | `str \| None` | optional |
| `temperature` | `float \| None` | optional |
| `top_p` | `float \| None` | optional |
Index:
- `Index("ix_job_status_date_created", "status", "date_created")`
### `MaintenanceRun`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `job_type` | `MaintenanceJobType` | non-null enum |
| `status` | `MaintenanceRunStatus` | non-null enum, default `queued` |
| `started_at` | `datetime \| None` | optional |
| `finished_at` | `datetime \| None` | optional |
| `triggered_by` | `str \| None` | optional |
| `summary` | `str \| None` | optional |
| `log_path` | `str \| None` | optional, log-root-relative POSIX path |
| `error_detail` | `str \| None` | optional internal failure detail |
| `created_at` | `datetime` | default now |
| `updated_at` | `datetime` | default now, onupdate |
### `Source`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `document_id` | `UUID` | FK -> `document.id`, indexed |
| `page_number` | `int` | default `1`, `ge=1` |
| `upload_name` | `str` | required |
| `filename` | `str` | required |
| `file_path` | `str` | required upload-root-relative POSIX path (`documents/...`) |
| `file_hash` | `str` | required |
| `file_size_bytes` | `int` | `BigInteger`, non-null |
| `raw_transcription` | `str \| None` | projection field |
| `preferred_execution_attempt_id` | `UUID \| None` | nullable FK -> `execution_attempt.id`, indexed (`use_alter`) |
| `revised_text` | `str \| None` | optional human revision |
| `date_uploaded` | `datetime` | default now |
| `date_revised` | `datetime \| None` | optional |
### `JobSource`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `job_id` | `UUID` | FK -> `job.id`, indexed |
| `source_id` | `UUID` | FK -> `source.id`, indexed |
| `status` | `JobSourceStatus` | non-null enum, default `pending` |
Constraint:
- `UniqueConstraint(job_id, source_id)` named `uq_job_source_job_source`
Runtime reconciliation:
- Startup database operations remove retired V4.6 `job_source` evidence columns (`raw_transcription`, `ai_metadata`, `raw_api_response`, `error_detail`, `executed_at`) when present so persisted schema matches this contract.
### `ExecutionAttempt`
| Field | Type | Notes |
| :--- | :--- | :--- |
| `id` | `UUID` | PK |
| `job_source_id` | `UUID` | FK -> `job_source.id`, indexed |
| `job_id` | `UUID` | FK -> `job.id`, indexed |
| `source_id` | `UUID` | FK -> `source.id`, indexed |
| `attempt_number` | `int` | `ge=1` |
| `status` | `JobSourceStatus` | non-null enum, value-stable with `JobSource.status` |
| `provider` | `str` | required |
| `model` | `str \| None` | optional |
| `request_manifest` | `dict[str, JsonValue] \| None` | JSONBCompat |
| `request_manifest_sha256` | `str \| None` | optional |
| `request_manifest_schema_version` | `str \| None` | optional |
| `response_received` | `bool` | default `False` |
| `transport_status_code` | `int \| None` | optional |
| `transport_body` | `bytes \| None` | LargeBinary |
| `transport_content_type` | `str \| None` | optional |
| `transport_content_encoding` | `str \| None` | optional |
| `transport_safe_headers` | `dict[str, JsonValue] \| None` | JSONBCompat |
| `router_request_id` | `str \| None` | optional |
| `router_generation_id` | `str \| None` | optional |
| `sdk_response_snapshot` | `dict[str, JsonValue] \| None` | JSONBCompat |
| `normalized_metadata` | `dict[str, JsonValue] \| None` | JSONBCompat; may include app-namespaced `processing_timing` (`provider_call_duration_ms`, `processing_duration_ms`) |
| `software_context` | `dict[str, JsonValue] \| None` | JSONBCompat |
| `raw_transcription` | `str \| None` | optional |
| `error_category` | `str \| None` | optional |
| `error_detail` | `str \| None` | optional |
| `failure_phase` | `str \| None` | optional |
| `started_at` | `datetime` | required |
| `finished_at` | `datetime` | required |
| `duration_ms` | `int` | `ge=0` |
| `created_at` | `datetime` | default now |
Constraint:
- `UniqueConstraint(job_id, source_id, attempt_number)` named `uq_execution_attempt_number`
## Relationship Loading Contract
- Most ORM relationships are configured with `lazy="raise"`.
- `JobSource.execution_attempts` is intentionally `lazy="noload"` with ordered attempts.
- Service/UI read paths must explicitly eager-load required relationships before access.
## Persistence Invariants (Ground Truth)
1. `ExecutionAttempt` is append-only runtime evidence.
2. `JobSource.status` represents queue/projection execution state and is not a full evidence container.
3. `Source.raw_transcription` is a mutable projection and not authoritative attempt history.
4. `Job` terminal status derives from page outcomes (`JobSource` state), not from a separate summary table.
5. `DocumentType.semantic_key` and `PersonRole.semantic_key` are nullable-unique semantic identifiers.
## Cross-Reference
- [System Architecture](architecture.md)
- [System Requirements](requirements.md)
- [Error Handling Policy](error_handling.md)
- [AI Evidence and Provenance Invariant](./invariant/ai_evidence_and_provenance.md)
+4 -4
View File
@@ -13,6 +13,7 @@ These documents are written for maintainers and AI contributors. They are behavi
- [People](pages/people.md) - [People](pages/people.md)
- [Jobs](pages/jobs.md) - [Jobs](pages/jobs.md)
- [Sources](pages/sources.md) - [Sources](pages/sources.md)
- [Settings](pages/settings.md)
NiceGUI registers the routes shown in each contract without the `/ui` prefix. The application mounts NiceGUI under `/ui`, so `/documents` in page code is served to a browser as `/ui/documents`. NiceGUI registers the routes shown in each contract without the `/ui` prefix. The application mounts NiceGUI under `/ui`, so `/documents` in page code is served to a browser as `/ui/documents`.
@@ -25,9 +26,8 @@ When documents disagree, use this order:
3. UI dependency and ownership boundaries: [UI contributor instructions](../../.github/instructions/ui.instructions.md). 3. UI dependency and ownership boundaries: [UI contributor instructions](../../.github/instructions/ui.instructions.md).
4. Durable failure behavior: [Error Handling invariant](../invariant/error_handling.md). 4. Durable failure behavior: [Error Handling invariant](../invariant/error_handling.md).
5. Durable AI evidence behavior: [Digital Evidence and AI Processing Provenance](../invariant/ai_evidence_and_provenance.md). 5. Durable AI evidence behavior: [Digital Evidence and AI Processing Provenance](../invariant/ai_evidence_and_provenance.md).
6. Data definitions and relationships: current models plus the [V4 schema](../ver4/schema_v4.md). 6. Data definitions and relationships: current models plus the [schema contract](../schema.md).
7. Planned behavior changes: the applicable V4.x scope and implementation documents. 7. Implementation truth: current code and tests.
8. Implementation truth: current code and tests.
If code intentionally changes accepted page behavior, update the corresponding page contract in the same change. If code accidentally differs, correct the implementation rather than rewriting intent to match a defect. If code intentionally changes accepted page behavior, update the corresponding page contract in the same change. If code accidentally differs, correct the implementation rather than rewriting intent to match a defect.
@@ -57,4 +57,4 @@ Each page contract contains:
## Current Baseline ## Current Baseline
These contracts describe the completed V4 through V4.5 behavior. These contracts describe the current V6.1 baseline.
+29 -9
View File
@@ -11,10 +11,11 @@ Documents manages the archival record for each historical artifact independently
| `/documents` | Searchable archival Document list. | | `/documents` | Searchable archival Document list. |
| `/documents/new` | Create a Document. | | `/documents/new` | Create a Document. |
| `/documents/{document_id}` | View one Document and its related records. | | `/documents/{document_id}` | View one Document and its related records. |
| `/documents/{document_id}/info` | View archival metadata and system logistics for one Document. |
| `/documents/{document_id}/edit` | Edit metadata and the complete Linked People set. | | `/documents/{document_id}/edit` | Edit metadata and the complete Linked People set. |
| `/documents/{document_id}/delete` | Confirm or block deletion. | | `/documents/{document_id}/delete` | Confirm or block deletion. |
| `/documents/{document_id}/jobs` | Show Jobs belonging to the Document. | | `/documents/{document_id}/jobs` | Show Jobs belonging to the Document. |
| `/documents/{document_id}/sources` | Redirect to the Document-filtered Sources list. | | `/documents/{document_id}/sources` | Source-image gallery for the Document. |
| `/documents/{document_id}/print` | Preview and browser-print the persisted Document. | | `/documents/{document_id}/print` | Preview and browser-print the persisted Document. |
## List Behavior ## List Behavior
@@ -22,11 +23,14 @@ Documents manages the archival record for each historical artifact independently
- The title is **Archival Documents**. - The title is **Archival Documents**.
- **Create new document** opens the create route. - **Create new document** opens the create route.
- The table defaults to Document Title order and supports search and column sorting. - The table defaults to Document Title order and supports search and column sorting.
- Columns are Document Title, Type, Author, Document Date, and Archive Ref. - Columns are Document Title, Author, Tags, Document Date, Type, # Sources, and Transcription Status.
- Document Title is left-aligned; the remaining columns are centered. - Document Title is left-aligned; the remaining columns are centered.
- Author lists all linked people in the `author` role. - Author lists all linked people in the `author` role.
- # Sources reflects the count of linked Source rows for each Document.
- Transcription Status reflects the most recent Job status for that Document; documents with no Jobs show a blank marker.
- Date display prefers exact date, then approximate date, then `Unknown`. - Date display prefers exact date, then approximate date, then `Unknown`.
- Selecting a row opens Document Detail. - Selecting a row opens Document Detail.
- Row navigation includes list context so Document Detail provides **Back to Documents**.
- No records displays `No documents found in repository.` - No records displays `No documents found in repository.`
## Create and Edit Behavior ## Create and Edit Behavior
@@ -43,12 +47,15 @@ Optional:
- Document location. - Document location.
- Archive identifier. - Archive identifier.
- Notes. - Notes.
- Tags.
- Linked People, with exactly one Person Role per linked Person. - Linked People, with exactly one Person Role per linked Person.
Rules: Rules:
- Exact date must parse as `YYYY-MM-DD`; browser presentation may follow locale. - Exact date must parse as `YYYY-MM-DD`; browser presentation may follow locale.
- The exact-date input is labeled **Document date**.
- Existing people appear with disambiguating labels. - Existing people appear with disambiguating labels.
- Tag assignment supports selecting existing tags and adding new labels inline.
- **Create new person** opens Person creation. - **Create new person** opens Person creation.
- `person_id` may preselect that Person in the author role on Document creation. - `person_id` may preselect that Person in the author role on Document creation.
- An invalid requested Person produces a warning rather than a broken form. - An invalid requested Person produces a warning rather than a broken form.
@@ -64,14 +71,27 @@ Rules:
## Detail Behavior ## Detail Behavior
- The heading shows name, type, and internal ID. - The heading shows name, type, and internal ID.
- The first Source, when present, appears in the dark-room viewer. - The header includes a contextual back action: **Back to Documents** by default, **Back to Person** when opened from Person Detail, and **Back to Job** when opened from Job Detail.
- Archival Metadata shows authors, Document Type, compact Document date, location, and archive identifier. Notes appear in a separate archival-notes block within the same card. - The detail workspace shows a Source-style pan/zoom media viewer with **Previous Page** / **Next Page** navigation for document source pages.
- System Logistics shows created and updated timestamps. - The center column is **Editable Revision** for the active source page.
- Related People are grouped by role and link to Person Detail. - Related People are grouped by role and link to Person Detail.
- **Sources & Pipeline Jobs** shows counts and actions for filtered Sources, Document Jobs, and adding a Job. - **Source Pages & Transcriptions** shows source/job counts and actions for source-image gallery, document jobs, and adding a Job.
- **Edit Document**, **Print**, and **Delete** are available from the header. - **Edit Document**, **Print**, **Document Details**, **View Source Detail**, and **Delete** are available from the header.
- Invalid IDs and missing Documents produce explicit states without rendering a partial page. - Invalid IDs and missing Documents produce explicit states without rendering a partial page.
## Document Source Images Behavior
- `/documents/{document_id}/sources` shows the current Document's source pages in a thumbnail gallery.
- Each card shows the page number, stored filename, and an **Open Source Detail** action.
- The page includes a **Back to Document** action.
- No source pages displays an explicit empty state.
## Document Info Behavior
- `/documents/{document_id}/info` contains **Archival Metadata** and **System Logistics**.
- It includes a **Back to Document** action.
- Archival metadata includes authors, document type, tags, document date, location (linked when present), archive identifier, and notes.
## Print Behavior ## Print Behavior
- Print opens a dedicated preview for persisted Document data. - Print opens a dedicated preview for persisted Document data.
@@ -120,11 +140,11 @@ Rules:
- `src/transcription/services/workflows.py` - `src/transcription/services/workflows.py`
- `src/transcription/ui/components/linked_people.py` - `src/transcription/ui/components/linked_people.py`
- `src/transcription/ui/pages/print_preview_page.py` - `src/transcription/ui/pages/print_preview_page.py`
- `src/transcription/api/v4_print.py` - `src/transcription/api/print_api.py`
- `tests/ui/test_documents_page.py` - `tests/ui/test_documents_page.py`
- `tests/services/test_document_service.py` - `tests/services/test_document_service.py`
## Known Limitations and Deferred Work ## Known Limitations and Deferred Work
- Source page ordering remains read-only in V4.4. - Source page ordering remains read-only.
- Printing other entities, batch printing, and server-side export formats are deferred. - Printing other entities, batch printing, and server-side export formats are deferred.
+15 -14
View File
@@ -2,48 +2,50 @@
## Purpose ## Purpose
Home provides a user-maintained landing page for the local archive. It combines one current image with Markdown text and lets the operator edit both without changing application source or prompt assets. Home provides a user-maintained landing page for the local archive. It combines a database-backed image gallery with Markdown text and lets the operator edit both without changing application source or prompt assets.
## Routes ## Routes
| Route | Browser path | Purpose | | Route | Browser path | Purpose |
| --- | --- | --- | | --- | --- | --- |
| `/homepage` | `/ui/homepage` | View current homepage image and Markdown. | | `/homepage` | `/ui/homepage` | View homepage gallery and Markdown. |
| `/homepage/edit` | `/ui/homepage/edit` | Upload an image and edit Markdown. | | `/homepage/edit` | `/ui/homepage/edit` | Upload images, manage image metadata, and edit Markdown. |
The application root and `/ui` redirect to `/ui/homepage`. The application root and `/ui` redirect to `/ui/homepage`.
## View Behavior ## View Behavior
- The visible page heading is **Home**; the browser tab title is **VibeScribe Home**. - The visible page heading is **Home**; the browser tab title is **VibeScribe Home**.
- The latest homepage image appears in the shared dark-room viewer. - The featured homepage image (`photo.is_primary`) is shown first; remaining images are shown in random order.
- The current image appears in the shared dark-room viewer with its description.
- Saved Markdown is rendered in the **Home Text** card. - Saved Markdown is rendered in the **Home Text** card.
- Missing text displays `No homepage text saved yet.` - Missing text displays `No homepage text saved yet.`
- Missing image displays the viewer's empty state. - Missing image displays the viewer's empty state.
- **Edit Home Page** opens the edit route. - **Edit Home Page** opens the edit route.
- The same Home Text content is also editable from **Settings → Home Page Text**.
## Edit Behavior ## Edit Behavior
- The image upload accepts JPEG, PNG, GIF, WebP, BMP, and TIFF files. - The image upload accepts JPEG, PNG, GIF, WebP, BMP, and TIFF files and supports multi-file uploads.
- A successful upload immediately stores the file, updates the preview to that image, and displays a positive notification. - A successful upload immediately stores files in the shared `photo` table/media layout and displays a positive notification.
- The editor supports per-image description edits, setting a featured image, and deleting the current image.
- The Markdown textarea is initialized from the currently stored homepage text. - The Markdown textarea is initialized from the currently stored homepage text.
- **Save** writes the textarea content, displays `Homepage saved`, and returns to Home. - **Save** writes the textarea content, displays `Homepage saved`, and returns to Home.
- **Cancel** returns to Home without saving textarea changes. An image already uploaded during the edit session remains stored. - **Cancel** returns to Home without saving textarea changes. An image already uploaded during the edit session remains stored.
## Storage Contract ## Storage Contract
- Homepage content is mutable application data under `data/homepage`. - Homepage markdown text is mutable application data at `UPLOAD_DIR/homepage.md`.
- Markdown is stored in `homepage.md`. - Homepage images are stored as `photo` rows (`person_id = NULL`) with files under `UPLOAD_DIR/photos/`.
- Uploaded images keep a sanitized basename. - Uploaded images are renamed to `{photo_id}{suffix}`.
- The view selects the supported image with the most recent modification time. - Homepage images are database records; markdown remains file-backed.
- Homepage files are not transcription prompts and are not database records.
## Acceptance Checklist ## Acceptance Checklist
- `/`, `/ui`, and the application brand reach Home. - `/`, `/ui`, and the application brand reach Home.
- Home renders with or without stored Markdown and image content. - Home renders with or without stored Markdown and image content.
- Edit loads existing Markdown. - Edit loads existing Markdown.
- A supported image upload updates the preview and becomes the latest homepage image. - Supported image upload stores one or more images and makes the first image featured when no featured image exists yet.
- Save persists Markdown and returns to Home. - Save persists Markdown and returns to Home.
- Cancel does not save changed Markdown. - Cancel does not save changed Markdown.
@@ -58,6 +60,5 @@ The application root and `/ui` redirect to `/ui/homepage`.
## Known Limitations ## Known Limitations
- Homepage storage is fixed under the repository/application `data` directory rather than a configured application-data root. - Homepage markdown storage location is `UPLOAD_DIR/homepage.md` and must remain writable in the active runtime environment.
- Uploading an image is immediate and is not rolled back by Cancel. - Uploading an image is immediate and is not rolled back by Cancel.
- The editor does not currently delete or select among previously uploaded images.
+8 -5
View File
@@ -19,10 +19,12 @@ Jobs manages transcription processing runs. A Job belongs to one Document, links
- The title is **Transcription Pipeline Jobs**. - The title is **Transcription Pipeline Jobs**.
- **Create job** opens Job creation and **Refresh** reloads the table. - **Create job** opens Job creation and **Refresh** reloads the table.
- Columns are Job ID, Status, Source Filename, Retries, Created, and Updated. - Columns are Job ID, Status, Document Name, # Sources, Retries, and Updated.
- Search covers Job ID, filename, and status. - Updated is the primary date/sort field.
- Search covers Job ID, document name, and status.
- Status is displayed as a semantic status chip. - Status is displayed as a semantic status chip.
- Selecting a row opens Job Detail. - Selecting a row opens Job Detail.
- Global Job-list row navigation includes list context so Job Detail provides **Back to Jobs**.
- No records displays `No job records found in repository.` - No records displays `No job records found in repository.`
## Create Behavior ## Create Behavior
@@ -30,7 +32,7 @@ Jobs manages transcription processing runs. A Job belongs to one Document, links
- A Target Document and at least one source file are required. - A Target Document and at least one source file are required.
- `document_id` may preselect a Target Document. - `document_id` may preselect a Target Document.
- If no Documents exist, the page explains the prerequisite and links to Document creation with a return path. - If no Documents exist, the page explains the prerequisite and links to Document creation with a return path.
- Provider and Model are optional request overrides. - Provider and Model are selectable when creating a new Job.
- Upload accepts JPEG, PNG, TIFF, and PDF files and supports multiple/folder selection. - Upload accepts JPEG, PNG, TIFF, and PDF files and supports multiple/folder selection.
- The visible upload queue is sorted alphabetically by original filename. - The visible upload queue is sorted alphabetically by original filename.
- Files can be removed individually or cleared before submission. - Files can be removed individually or cleared before submission.
@@ -42,8 +44,9 @@ Jobs manages transcription processing runs. A Job belongs to one Document, links
## Detail and Lifecycle Behavior ## Detail and Lifecycle Behavior
- The heading shows Job ID and a status badge. - The heading shows Job ID and a status badge.
- Job Detail includes a contextual back action: **Back to Jobs** by default and **Back to Document** when opened from a Document-filtered Job list.
- Execution Logistics shows provider, model, prompt, retry count, and last update. - Execution Logistics shows provider, model, prompt, retry count, and last update.
- Document Links open the parent Document and Job-filtered Sources. - Document Links show a clickable Document Name (with Job context), Sources count, and a single **View Sources** action using job filtering.
- Queued and processing Jobs show an auto-refresh notice and reload every four seconds. - Queued and processing Jobs show an auto-refresh notice and reload every four seconds.
- Polling stops when the Job becomes terminal or a refresh fails. - Polling stops when the Job becomes terminal or a refresh fails.
- Queued and processing Jobs expose **Cancel**. - Queued and processing Jobs expose **Cancel**.
@@ -53,7 +56,7 @@ Jobs manages transcription processing runs. A Job belongs to one Document, links
## Cancel Behavior ## Cancel Behavior
- The confirmation explains that processing stops and remaining non-transcribed Sources become failed. - The confirmation explains that processing stops and remaining pending Sources become cancelled.
- The service decides whether the current state permits cancellation. - The service decides whether the current state permits cancellation.
- Success updates the Job, notifies the worker, and returns to Job Detail. - Success updates the Job, notifies the worker, and returns to Job Detail.
+31 -15
View File
@@ -2,7 +2,7 @@
## Purpose ## Purpose
People manages reusable historical-person records. A Person may appear in many Documents under different relationship roles and may optionally carry a portrait and FamilySearch identifier. People manages reusable historical-person records. A Person may appear in many Documents under different relationship roles and may optionally carry one or more photos plus a FamilySearch identifier.
## Routes ## Routes
@@ -11,6 +11,7 @@ People manages reusable historical-person records. A Person may appear in many D
| `/people` | Searchable People list. | | `/people` | Searchable People list. |
| `/people/new` | Create a Person. | | `/people/new` | Create a Person. |
| `/people/{person_id}` | View one Person and linked Documents. | | `/people/{person_id}` | View one Person and linked Documents. |
| `/people/{person_id}/photos` | Manage Person photos. |
| `/people/{person_id}/edit` | Edit the Person. | | `/people/{person_id}/edit` | Edit the Person. |
| `/people/{person_id}/delete` | Confirm permanent deletion. | | `/people/{person_id}/delete` | Confirm permanent deletion. |
@@ -18,61 +19,76 @@ People manages reusable historical-person records. A Person may appear in many D
- The title is **Archival Entities: People**. - The title is **Archival Entities: People**.
- **Create new person** opens the create route. - **Create new person** opens the create route.
- The table defaults to Full Name order and supports search and column sorting. - The table defaults to Name order (`Last Name, First & Middle`) and supports search and column sorting.
- Columns are Full Name, Display Name, Maiden Name, Birth Date, and Death Date. - Columns are Last Name, First & Middle; Tags; FamilySearch ID; Birth Date; Death Date; and # Documents.
- Full Name is left-aligned; Display Name, Maiden Name, and date columns are centered. - Name and Tags are left-aligned; FamilySearch ID, date columns, and # Documents are centered.
- # Documents reflects how many linked Documents each Person is connected to.
- Birth and death values independently prefer exact date, then approximate date, then `Unknown`. - Birth and death values independently prefer exact date, then approximate date, then `Unknown`.
- Selecting a row opens Person Detail. - Selecting a row opens Person Detail.
- Row navigation includes list context so Person Detail provides **Back to People**.
- No records displays `No person records found in repository.` - No records displays `No person records found in repository.`
## Create and Edit Behavior ## Create and Edit Behavior
Required: Required:
- Full name. - Last name.
- First & middle names.
Optional: Optional:
- Display name and maiden name.
- Exact and approximate birth/death dates. - Exact and approximate birth/death dates.
- Birth/death places. - Birth/death places.
- Biography. - Biography.
- Portrait path or uploaded portrait.
- FamilySearch ID. - FamilySearch ID.
- Tags.
Rules: Rules:
- Missing Full name blocks save with a warning. - Missing last name or first/middle names blocks save with a warning.
- Exact date inputs are native browser date inputs. - Exact date inputs are native browser date inputs.
- FamilySearch IDs are normalized and validated by `PeopleService`. - FamilySearch IDs are normalized and validated by `PeopleService`.
- Portrait uploads are stored under the configured upload root in a Person-specific directory and update Portrait path. - Tags use the shared Tag registry and support inline add/select behavior.
- Photos are managed from Person Detail via `/people/{person_id}/photos` (not in create/edit form fields).
- Metadata JSON remains hidden. - Metadata JSON remains hidden.
- Save success returns to Person Detail. - Save success returns to Person Detail.
## Detail Behavior ## Detail Behavior
- The header provides **New Document**, **Edit Person**, and **Delete**. - The header provides **New Document**, **Edit Person**, **Edit Photo(s)**, and **Delete**.
- The header includes a contextual back action: **Back to People** by default, and **Back to Document** when opened from Document Detail.
- **New Document** opens Document creation with this Person requested for author preselection. - **New Document** opens Document creation with this Person requested for author preselection.
- The portrait viewer resolves supported relative upload paths and absolute HTTP/data URLs. - Person Detail shows a single-photo viewer with **Previous/Next** navigation; the page-level **Edit Photo(s)** header action opens photo management.
- Biographical Record shows names, compact birth/death dates, places, and an **Open in FamilySearch** link when an ID exists. - Photo management (upload, description edit, set-primary, delete) is intentionally moved to `/people/{person_id}/photos`.
- Biographical Record shows split names, computed full name, tags, compact birth/death dates, and places.
- Birth and death place values are clickable links to Google Maps when present.
- FamilySearch ID is shown as a metadata value and is clickable to the FamilySearch person details route when present.
- Biography has an explicit empty value. - Biography has an explicit empty value.
- Linked Documents show Document name, relationship role, and an action to open Document Detail. - Linked Documents render as a table with **Document Name**, **Document Date**, **Role**, and **Number of Pages**; selecting a row opens Document Detail.
- No links shows both an empty state and guidance to link from a Document workflow. - No links shows both an empty state and guidance to link from a Document workflow.
- System Logistics shows created and updated timestamps. - System Logistics shows created and updated timestamps.
## Delete Behavior ## Delete Behavior
- The page warns when linked Document relationships exist. - The page warns when linked Document relationships exist.
- Delete is blocked when related Photos exist.
- Confirmed deletion removes the Person and its relationship links; it does not delete Documents. - Confirmed deletion removes the Person and its relationship links; it does not delete Documents.
- Success returns to the People list. - Success returns to the People list.
- Missing or already-deleted records return to a safe list state. - Missing or already-deleted records return to a safe list state.
## Photo Gallery Behavior (`/people/{person_id}/photos`)
- Upload is triggered from a header-level **Upload Photo(s)** control beside **Back to Person**.
- The gallery renders all photos in a responsive grid (3-4 tiles wide on larger screens).
- Description text is shown as an overlay at the bottom of each image for quick context.
- The editor provides **Save Description**, **Set Primary** (when applicable), and **Delete Photo** actions.
## Acceptance Checklist ## Acceptance Checklist
- List fields, alignment, date fallback, search, sorting, and navigation match this contract. - List fields, alignment, date fallback, search, sorting, and navigation match this contract.
- Full name is enforced on create and edit. - Last name and first/middle names are enforced on create and edit.
- FamilySearch ID validation and link generation use the fixed supported identifier format. - FamilySearch ID validation and link generation use the fixed supported identifier format.
- Portrait upload and rendering remain constrained to supported media paths. - Photo upload and rendering remain constrained to supported media paths.
- New Document carries the Person context. - New Document carries the Person context.
- Linked Documents show the correct role and target. - Linked Documents show the correct role and target.
- Delete wording distinguishes removal of relationship links from deletion of Documents. - Delete wording distinguishes removal of relationship links from deletion of Documents.
+55
View File
@@ -0,0 +1,55 @@
# Settings Page Contract
## Purpose
Settings manages installation-local registries, safe runtime .env settings, and editable text assets from one route.
## Route
| Route | Purpose |
| --- | --- |
| `/settings` | Manage Runtime Settings, Document Types, Person Roles, Tags, Prompts, Home Page Text, and Maintenance runs. |
## Behavior
- The page title is **Settings**.
- Configuration surfaces are grouped as tabs:
- **Document Types**
- **Person Roles**
- **Tags**
- **Prompts**
- **Home Page Text**
- **Maintenance**
- **Runtime Settings**
- Runtime Settings exposes an allowlisted set of non-secret fields synchronized with `Settings` model fields except excluded secret/unsafe fields.
- Runtime Settings is rendered as a compact two-column editor (**Setting**, **Value**) in a centered, narrower responsive container.
- Runtime Settings persists changes to the resolved runtime env file, validates by constructing a `Settings` instance, and reports validation failures through the shared UI error presenter.
- `Settings` resolves its env file in this order: explicit `_env_file`, `ENV_FILE`, then the repository-root `.env.production`.
- Runtime Settings resolves its write target in this order: explicit function override (tests/tools), `RUNTIME_SETTINGS_ENV_FILE` environment variable (deployment override), `ENV_FILE`, then the repository-root `.env.production`.
- Runtime Settings changes require application restart to take effect.
- Runtime Settings renders a host-side restart command (`docker compose -f docker-compose.production.yml up -d --force-recreate app worker`) so operators can apply saved values without granting Docker control to the app container.
- Runtime Settings includes an explicit "Other settings not shown here" markdown table listing:
- secrets (`OPENROUTER_API_KEY`, `DATABASE__PASSWORD`)
- high-risk database connection settings (`DATABASE__DRIVER`, `DATABASE__PATH`, `DATABASE__HOST`, `DATABASE__PORT`, `DATABASE__DATABASE`, `DATABASE__USER`)
and deployment/helper keys (`CLOUDFLARE_TUNNEL_TOKEN`, `BACKUP_DIR`, `BACKUP_RETENTION_DAYS`, `RUNTIME_SETTINGS_ENV_FILE`, `ENV_FILE`, `COMPOSE_FILE`) plus legacy/deprecated keys (`POSTGRES_*`, `DATABASE_BACKUP_DIR`, `APP_DATA_BACKUP_DIR`, `UPLOADS_BACKUP_DIR`, `SYNOLOGY_BACKUP_DIR`), and directs edits for those keys to the resolved runtime env file path.
- Document Types, Person Roles, and Tags support Add/Edit/Delete with existing guardrails.
- Prompts exposes only `transcribe_document.md` for editing and restore-from-backup.
- Home Page Text edits the same Markdown content rendered on `/homepage`.
- Maintenance provides queue-backed **Run Backup** and **Run Storage Reconciliation** actions.
- Maintenance also provides GEDCOM upload and **Run GEDCOM Import** actions, using the same queue-backed `MaintenanceRun` history/log flow.
- Maintenance run history shows job type, status, started/finished timestamps, duration, summary, and log view/download actions.
- Maintenance actions enqueue work and signal the worker; the page itself does not execute shell commands directly.
## Acceptance Checklist
- `/ui/settings` renders all seven tabs.
- Registry and prompt workflows keep existing validation and error handling.
- Runtime Settings excludes secret fields and rejects invalid values.
- Saving Home Page Text persists content for the homepage view.
## Implementation Anchors
- `src/transcription/ui/pages/settings_page.py`
- `src/transcription/ui/runtime_settings_store.py`
- `src/transcription/ui/homepage_store.py`
- `tests/ui/test_pages_registration.py`
+13 -11
View File
@@ -8,7 +8,7 @@ Sources manages individual archived page/file records. It provides source-media
| Route | Purpose | | Route | Purpose |
| --- | --- | | --- | --- |
| `/sources` | Global or filtered Source list. | | `/sources` | Document-filtered or Job-filtered Source list; global route redirects to Documents. |
| `/sources/{source_id}` | View media, transcription, revision, metadata, and evidence. | | `/sources/{source_id}` | View media, transcription, revision, metadata, and evidence. |
| `/sources/{source_id}/delete` | Confirm or block deletion. | | `/sources/{source_id}/delete` | Confirm or block deletion. |
@@ -16,12 +16,13 @@ The list accepts optional `document_id` and `job_id` query parameters. Document
## List Behavior ## List Behavior
- The title is **Source Asset Records**, **Sources for Document**, or **Sources for Job** according to context. - The global `/sources` route redirects to `/documents`.
- Global context provides **Create Job**. - Filtered list titles are **Sources for Document** and **Sources for Job**.
- Filtered context provides **Back to Document** or **Back to Job**. - Filtered context provides **Back to Document** or **Back to Job**.
- Rows are ordered by page number and then upload name. - Rows are ordered by page number and then upload name.
- Columns are Document Name, Page Number, Upload Title, Status, and Error Detail. - Columns are Upload Title, Page Number, Document Name, Status, and Error Detail.
- Document Name, Upload Title, and Error Detail are left-aligned; Status is centered. - Document Name, Upload Title, and Error Detail are left-aligned; Status is centered.
- Status labels are presented in uppercase for consistency with Jobs.
- Stored Filename is intentionally absent from the list. - Stored Filename is intentionally absent from the list.
- Selecting a row opens Source Detail. - Selecting a row opens Source Detail.
- No records displays `No source asset records found in repository.` - No records displays `No source asset records found in repository.`
@@ -29,18 +30,19 @@ The list accepts optional `document_id` and `job_id` query parameters. Document
## Detail Behavior ## Detail Behavior
- The heading shows page number, upload name, and Source ID. - The heading shows page number, upload name, and Source ID.
- **Back to Sources** returns to the global list. - **Back to Document** returns to Document Detail for the active source page.
- **Retranscribe Source** opens Create Processing Job with this Source and its Document locked. - **Retranscribe Source** opens Create Processing Job with this Source and its Document locked.
- **Delete Source** opens the guarded delete route. - **Delete Source** opens the guarded delete route.
- Previous and Next navigate only among Sources belonging to the same Document in page order; unavailable boundary actions are disabled. - Previous and Next navigate only among Sources belonging to the same Document in page order; unavailable boundary actions are disabled.
- The media viewer resolves the stored Source path through the configured upload root. - The media viewer resolves the stored Source path through the configured upload root.
- Transcription Text is read-only and displays the preferred machine projection, with a legacy latest-JobSource - The top layout is adaptive:
fallback only when no Source projection exists. - Standard pages use three columns with a wider Editable Revision column than the image column.
- Editable Revision is seeded from an existing revision or the machine transcription. - Wide+narrow landscape images switch to a stacked left layout (image above Editable Revision) with metadata on the right.
- Editable Revision is seeded from an existing revision or the preferred machine transcription.
- Source Metadata shows upload name, stored filename, page number, Document Name, Document ID, and stored path. Source ID appears in the page-header subtitle. - Source Metadata shows upload name, stored filename, page number, Document Name, Document ID, and stored path. Source ID appears in the page-header subtitle.
- SourceJob Metadata shows latest status, Job ID, execution time, provider, model, prompt, and failure detail. - SourceJob Metadata shows latest status (uppercase display), Job ID, execution time, provider, model, prompt, and failure detail.
- Revision Logistics shows revised state, last-revised time, and upload time. - Revision Logistics shows revised state, last-revised time, and upload time.
- Candidate Machine Transcriptions remains compact until a candidate is expanded, then compares it with the preferred - Candidate Machine Transcriptions appears below the image/revision area, remains compact until expanded, then compares it with the preferred
machine result and requires confirmation before **Use this transcription**. machine result and requires confirmation before **Use this transcription**.
- Candidate promotion does not alter a human revision. Empty states distinguish no machine result from no candidates. - Candidate promotion does not alter a human revision. Empty states distinguish no machine result from no candidates.
- An orientation-normalized artifact appears in evidence only when recognized metadata required a physical rotation. - An orientation-normalized artifact appears in evidence only when recognized metadata required a physical rotation.
@@ -94,4 +96,4 @@ The list accepts optional `document_id` and `job_id` query parameters. Document
## Planned Changes ## Planned Changes
- Source page reordering is deferred beyond V4.3 and may be reconsidered if a demonstrated workflow need emerges. - Source page reordering remains deferred unless a demonstrated workflow need emerges.
-146
View File
@@ -1,146 +0,0 @@
# Implementation Plan (Version 4.1)
## Goal
Deliver the V4.1 usability revision as a small, behavior-safe increment over the V4 baseline.
## Implementation Principles
- Keep presentation formatting in UI components and route orchestration in pages.
- Keep persistence and cross-record queries behind service boundaries.
- Reuse shared table and date-label helpers instead of duplicating fallback logic.
- Make the FamilySearch schema change additive and nullable.
- Add focused tests for changed behavior before broad regression verification.
## Current Project Impact
| Area | Expected impact |
| --- | --- |
| Persistence | Add nullable `Person.family_search_id`; provide the repository's supported schema-upgrade path for existing databases. |
| People service | Normalize and validate FamilySearch IDs at the domain/service boundary if model validation does not fully cover writes. |
| Documents UI | Add table data, improve relationship labels/links, compact date display, and combine processing navigation. |
| People UI | Add table date fields, Person-first Document creation, compact date display, and FamilySearch controls. |
| Sources service/UI | Query adjacent document Sources and add bounded navigation; revise list columns and wrapping. |
| Jobs UI | Refresh the active detail read model on a timer until terminal status. |
| Shared UI | Add reusable constrained/wrapped table presentation and compact date formatting where appropriate. |
| Tests | Update model/service and UI coverage for all affected workflows. |
## Implementation Phases
### 1. Add Shared Presentation Rules
- Review `ui/components/table/common.py` and packaged theme CSS for the narrowest reusable table-width solution.
- Add reusable styles or column slots for constrained, wrapping, left-aligned text.
- Add a shared formatter for exact/approximate/unknown dates if it can be reused without coupling components to persistence.
- Preserve sorting and search behavior for rendered display values.
### 2. Update Archival List Tables
- Extend the Document table read model with author names and the compact document date.
- Build author display from eagerly loaded document-person links using the `author` role.
- Apply title/type alignment and constrained title wrapping.
- Extend the Person table read model with compact birth and death date values.
- Apply Display Name and Maiden Name alignment.
- Remove Stored Filename from the Source table read model only if no other list behavior consumes it; always remove its rendered column.
- Constrain and left-align the requested Source columns.
- Add or update UI component tests for serialized rows, columns, and fallback formatting.
### 3. Improve Document Relationship Workflows
- Introduce one person-label formatter that combines preferred Display Name, Full Name context, and known birth year without implying uniqueness.
- Use Person UUIDs as selector values.
- Apply the formatter to every relationship role selector.
- Change Related People rows into actions that navigate to `/people/{person_id}`.
- Replace separate exact/approximate rows in view mode with one conditional Document Date row.
- Combine Pipeline Jobs and Sources into one related-processing card beneath Related People.
- Preserve existing job/source counts and navigation actions.
### 4. Add the Person-First Document Workflow
- Add a New Document action on Person Detail.
- Pass the Person UUID through a narrowly defined query parameter to `/documents/new`.
- Validate the requested UUID against the loaded people list.
- Preselect that person in the intended default relationship role. Use `author` unless a different role is explicitly encoded later.
- Ignore invalid or unavailable preselection values with the application's normal visible error/notification behavior.
- Confirm ordinary `/documents/new` behavior remains unchanged.
### 5. Add FamilySearch Person References
- Add nullable, unique `family_search_id` to the `Person` model and schema.
- Implement a non-destructive upgrade for existing SQLite and PostgreSQL databases using the repository's established schema-management approach.
- Normalize values by trimming and uppercasing.
- Validate the `XXXX-XXX` alphanumeric identifier shape and return a clear validation error for malformed input.
- Report duplicate identifiers as a deterministic conflict rather than a generic persistence failure.
- Add the field to Person create/edit forms and preserve it during updates.
- Add a URL builder that safely inserts only a validated identifier into the fixed FamilySearch details URL.
- Render a FamilySearch action on Person Detail only when an identifier is present.
- Add persistence, normalization, validation, form, and link-generation tests.
### 6. Add Source Page Navigation
- Add a Sources service query that returns previous/current/next context for a Source within its Document.
- Define ordering by `page_number`, with a stable secondary key such as Source UUID for defensive determinism.
- Keep navigation bounded to the current `document_id`.
- Render previous and next actions adjacent to the source viewer or detail header.
- Disable or omit unavailable boundary actions.
- Test first, middle, last, single-page, and cross-document cases.
### 7. Add Job Detail Auto-Refresh
- Make Job Detail content refreshable without rebuilding unrelated global navigation.
- Start a NiceGUI timer only for queued or processing jobs.
- On each tick, re-read the Job through `JobService` and refresh the detail content.
- Use a 4-second default interval.
- Stop or deactivate the timer when status becomes completed, partial success, failed, or cancelled, according to the model's actual terminal states.
- Prevent overlapping refresh callbacks.
- Retain existing error presentation if a refresh read fails.
- Add UI tests for timer creation, refresh, and terminal-state stopping.
### 8. Simplify View-Mode Date Rows
- On Document Detail, show exact date, else approximate date, else one not-set value.
- On Person Detail, apply the same independent rule to birth and death.
- Do not hide either input in create/edit mode.
- Test each exact, approximate, and absent state.
### 9. Verification and Documentation Alignment
- Run the focused model/service/UI tests covering changed surfaces.
- Run the existing regression suite appropriate to persistence and UI changes.
- Confirm SQLite and PostgreSQL model compatibility at the schema-definition level.
- Update V4.1 documentation if implementation reveals a necessary boundary change; do not silently expand scope.
## Recommended Delivery Order
1. Shared formatters and table presentation.
2. Additive Person schema change and FamilySearch validation.
3. Document and Person list/detail changes.
4. Person-first Document workflow.
5. Source navigation.
6. Job polling.
7. Focused and regression verification.
## Done When
- Every V4.1 acceptance criterion is demonstrated or covered by a focused test.
- Existing Person rows remain valid after the nullable schema addition.
- Duplicate FamilySearch references cannot be assigned to multiple local Person records.
- FamilySearch links are generated only from normalized, validated IDs.
- Auto-refresh performs no polling after a terminal job state.
- Adjacent Source navigation never crosses Document boundaries.
- The existing V4 workflows remain operational.
## Out of Scope
- Page reordering.
- Settings management.
- External genealogy API integration.
- Raw `.env` editing.
- Theme editing.
## Related Local References
- [V4.1 Scope Boundary](scope_boundary_v4_1.md)
- [V4 Implementation Plan](../ver4/implementation_plan_v4.md)
- [V4 Requirements](../ver4/requirements_v4.md)
- [V4 Error Handling Policy](../ver4/error_handling_v4.md)
-142
View File
@@ -1,142 +0,0 @@
# V4.1 Scope Boundary
This document defines the scope of the first incremental revision to Version 4. V4 remains the product and architecture baseline; V4.1 adds focused usability improvements and one additive Person field.
## Purpose
- Improve common archival record workflows without redesigning the application.
- Resolve table overflow, ambiguous person selection, and unnecessary navigation.
- Add a manually maintained FamilySearch person reference without introducing external API integration.
## In Scope
### 1. Archival Documents List
- Keep every table column within the available page width.
- Limit and wrap long Document Title values.
- Left-align Document Title and Type.
- Add Author and Document Date columns.
- Display all people linked through the `author` role in the Author column.
- Display exact document date when present, otherwise approximate date when present, otherwise `Unknown`.
### 2. Document Detail and Editing
- Use an unambiguous label in person selectors. The label should prefer Display Name, retain Full Name for context, and include the birth year when known.
- Do not require Display Name to be unique.
- Link each Related People entry to its Person Detail page.
- In view mode, display only the populated exact or approximate document date row. Display a single unknown/not-set state when neither exists.
- Keep both exact and approximate inputs available in create/edit mode.
- Move source navigation out from beneath the media viewer.
- Present Pipeline Jobs and Sources together in one related-processing card with counts and actions.
### 3. People List
- Left-align Display Name and Maiden Name.
- Display exact birth date when present, otherwise approximate birth date when present, otherwise `Unknown`.
- Add a Death Date column with the same fallback rule.
### 4. Person Detail and Editing
- Add a New Document action that opens Document creation with the current person preselected.
- Preserve the existing Document-first workflow.
- In view mode, display only the populated exact or approximate row for each of birth and death date. Display a single unknown/not-set state when neither value exists.
- Keep both exact and approximate inputs available in create/edit mode.
### 5. FamilySearch Reference
- Add a nullable, unique `family_search_id` field to `Person`.
- Allow the field to be entered and changed in Person create/edit flows.
- Trim whitespace, normalize the identifier to uppercase, and validate it against the supported
`XXXX-XXX` alphanumeric shape before persistence.
- When an identifier exists, show a FamilySearch action on Person Detail linking to:
`https://www.familysearch.org/tree/person/details/{family_search_id}`
- Construct the URL in application code; do not persist the full URL.
### 6. Source List and Detail
- Keep every Source Asset Records table column within the available page width.
- Limit and wrap long Document Name, Upload Title, and Error Detail values.
- Left-align Document Name, Upload Title, and Error Detail.
- Remove Stored Filename only from the Source Asset Records table. Continue storing it and showing it on Source Detail.
- On Source Detail, add previous and next navigation for Sources belonging to the same Document, ordered by `page_number`.
- Disable or omit the previous/next action at the first/last page.
### 7. Job Detail
- Automatically refresh Job Detail while the job is in a non-terminal state.
- Use a modest interval in the 3-5 second range.
- Stop polling when the job reaches a terminal state or the page is no longer active.
- Preserve manual navigation and existing job actions.
### 8. Homepage Storage Decision
- Continue treating homepage markdown and images as mutable application data, not prompt artifacts or packaged source assets.
- Keep homepage content separate from `prompts`.
- Defer relocation to a configurable application-data root unless the existing location prevents normal installed or deployed operation.
## Out of Scope
- Source page renumbering or reordering.
- A Settings page.
- Editing `.env` or secrets through the UI.
- Runtime theme editing.
- FamilySearch authentication, API calls, search, import, synchronization, or conflict resolution.
- Ancestry references or other genealogy providers.
- Google Maps links from place fields.
- Enforcing unique Display Name values.
- Changes to transcription execution or provider behavior.
- Changes to the V4 API solely to expose the V4.1 presentation enhancements.
## Locked Design Decisions
### A. Person Selector Identity
- Selection values remain internal Person UUIDs.
- Display labels provide disambiguating context but are not identity keys.
- Duplicate Full Name and Display Name values remain valid.
### B. Date Presentation
- Exact dates take precedence over approximate/raw dates for compact list and view presentation.
- Create/edit forms retain both fields so either representation can be maintained.
- V4.1 does not introduce a new mutual-exclusion database constraint.
### C. FamilySearch Storage
- Store only the FamilySearch person identifier.
- Treat a FamilySearch person identifier as unique across local Person records.
- Use one dedicated nullable Person field while FamilySearch is the only supported external genealogy reference.
- Reconsider a generic external-reference model only when a second provider or multiple references per person are required.
### D. Stored Filename
- Stored Filename remains part of the Source model and Source Detail diagnostics.
- Only the list-table column is removed.
## Data and Compatibility Policy
- The `family_search_id` addition must be nullable and non-destructive for existing Person rows.
- Existing records, routes, relationships, jobs, Sources, prompt provenance, and uploaded media remain valid.
- UI changes must preserve both Document-first and Person-first workflows.
- V4.1 must remain portable across SQLite and PostgreSQL.
## Acceptance Criteria
1. Document, Person, and Source tables fit their page containers at supported desktop widths without losing requested columns.
2. Long table text wraps or is constrained without forcing important columns outside the table container.
3. Duplicate-named people can be distinguished in every document relationship selector.
4. Related People entries navigate to the correct Person Detail page.
5. Compact date displays consistently use exact, then approximate, then unknown fallback behavior.
6. Starting from Person Detail can create a Document with that person preselected without breaking normal Document creation.
7. A valid FamilySearch ID is persisted and produces the correct Person Detail hyperlink; absent IDs produce no action.
8. Source previous/next actions remain within the same Document and follow `page_number`.
9. Active Job Detail pages update without manual refresh and stop polling after terminal status.
10. Stored Filename is absent from the Source list table but remains available on Source Detail.
11. Focused automated tests pass and unaffected V4 behavior remains intact.
## Related Local References
- [V4.1 Implementation Plan](implementation_plan_v4_1.md)
- [V4 Scope Boundary](../ver4/scope_boundary_v4.md)
- [V4 Architecture](../ver4/architecture_v4.md)
- [V4 Schema](../ver4/schema_v4.md)
-258
View File
@@ -1,258 +0,0 @@
# Implementation Plan (Version 4.2)
## Goal
Make processing evidence precise, append-only, secret-safe, and exportable while preserving every existing record and creating a provider-neutral home for future OCR/layout artifacts.
## Implementation Principles
- Implement the [Digital Evidence and AI Processing Provenance invariant](../invariant/ai_evidence_and_provenance.md), not a provider-specific approximation of it.
- Capture transport evidence before SDK parsing.
- Keep exact evidence separate from parsed and normalized representations.
- Prefer additive schema evolution and explicit compatibility behavior.
- Reference source content by digest rather than duplicating it in request JSON.
- Use allowlists for safe metadata capture.
- Keep persistence and evidence semantics behind service boundaries.
- Do not change the default model until a representative benchmark supports that decision.
## Current-State Gaps
| Current behavior | Gap to close |
| --- | --- |
| `Source` stores original file path, digest, and size. | Media type and image/page geometry used for an execution are not frozen with that execution. |
| `Job` stores prompt text/hash, requested model after resolution, temperature, and `top_p`. | The complete effective request structure, omitted-versus-explicit parameter state, routing constraints, and software versions are not frozen. |
| `JobSource.raw_api_response` stores `model_dump()` output from the OpenRouter SDK. | The exact HTTP body can be normalized by OpenRouter and filtered again by the SDK before persistence. |
| `JobSource.ai_metadata` stores finish reason and basic token counts. | Detailed accounting remains only in the SDK snapshot and is not a substitute for exact evidence. |
| Provider exceptions become application errors. | Safe HTTP error bodies, statuses, headers, and no-response distinctions are not persisted. |
| Worker logs elapsed time. | Execution duration is not stored on `JobSource`. |
| Source Detail displays AI metadata and the SDK snapshot. | The UI does not identify evidence layers or expose request/transport/software provenance. |
| No generic processing-artifact model exists. | Future OCR geometry would require ad hoc provider fields or an unrelated schema. |
## Expected Project Impact
| Area | Expected impact |
| --- | --- |
| Database models and upgrades | Add execution-specification, transport-evidence, timing, software-context, and generic artifact storage without removing existing columns. |
| OpenRouter adapter | Introduce a transport boundary that can capture exact body/status/safe headers before typed SDK parsing, or use supported SDK hooks that expose the unparsed response reliably. |
| Provider contract | Return structured evidence for success and failure without leaking provider-specific transport concerns into workflow orchestration. |
| Source and workflow services | Persist one append-only execution outcome and its artifacts transactionally; retain compatibility projections. |
| UI | Label and inspect evidence layers; export safe evidence packages through service operations. |
| Benchmarking | Add a private manifest and repeatable evaluator using the literal-transcription methodology. |
| Tests and documentation | Add compatibility, capture, security, integrity, export, and benchmark-scoring coverage; correct overstated V4 evidence language. |
## Proposed Data Design
Exact names should be confirmed against existing conventions before migration code is written. The design should provide the following logical records.
### 1. Execution Evidence
Extend `JobSource` or associate it one-to-one with a new execution-evidence record containing:
- Request manifest JSON and manifest schema version.
- Transport status, body bytes or exact decoded body plus encoding/content type, and safe headers.
- Parsed SDK snapshot retained separately from transport content.
- Application, adapter, SDK, and runtime version metadata.
- Start, finish, and duration values.
- Router/provider request and generation identifiers when available.
- Failure phase and whether an HTTP response was received.
The implementation should evaluate a companion table rather than continuing to widen `JobSource`. A companion record better isolates large/optional evidence and permits clear one-to-one compatibility semantics.
### 2. Generic Processing Artifact
Add a one-to-many artifact model associated with a source and, when applicable, a producing execution:
- Stable artifact UUID.
- `source_id` and optional execution/`job_source_id`.
- Semantic artifact type.
- Media/serialization format.
- Schema name and version.
- Producer and producer version.
- Inline JSON payload or external location.
- Payload digest and byte size.
- Coordinate-system metadata when relevant.
- Creation timestamp.
Enforce exactly one content location: inline payload or external reference. An external artifact must be written durably and hashed before its database record commits.
### 3. Compatibility Projections
- Keep `JobSource.raw_api_response` unchanged for existing and new compatibility reads until a later deprecation decision.
- Keep `JobSource.ai_metadata` for indexed/display-ready normalized values.
- Keep `Source.raw_transcription` as the latest successful machine-output projection while treating per-execution `JobSource.raw_transcription` as history.
- Document that older rows have an SDK snapshot but no exact transport capture.
## Implementation Phases
### 1. Correct Terminology and Define Typed Contracts
- Add typed domain models for request manifests, software context, transport metadata, failure phase, and artifact descriptors.
- Version every persisted JSON contract from its first release.
- Define the safe response-header allowlist. Begin with correlation, content type/encoding, date, retry/rate-limit, and router-specific generation identifiers only when documented and non-secret.
- Define size limits and external-storage thresholds for exact bodies and artifacts.
- Correct `docs/ver4/schema_v4.md` under “Page-Level Execution and AI Outputs” so the existing column is described as an SDK-serialized OpenRouter response snapshot, not a complete provider envelope, exact HTTP body, or native upstream-provider response. Apply the same terminology to architecture and UI schema references.
- Add serialization and secret-rejection unit tests before provider changes.
### 2. Add Additive Persistence and Upgrade Behavior
- Add the selected execution-evidence and artifact models.
- Add foreign keys, uniqueness constraints, and indexes for source/execution lookup.
- Implement idempotent upgrades following the repository's existing schema-upgrade policy.
- Do not populate exact response fields for historical rows.
- Do not write a capture-time classification onto historical rows during migration. Compatibility reads may describe a populated legacy `raw_api_response` as an SDK snapshot, but exports must identify that description as a later compatibility interpretation rather than execution-time metadata.
- Verify JSON portability and large-payload behavior for SQLite and PostgreSQL.
- Add upgrade tests starting from a representative pre-V4.2 schema.
### 3. Build Secret-Safe Request Manifests
- Build the manifest from the concrete outgoing request body immediately before transport, not from a narrower typed projection that may discard unrecognized request fields.
- Replace each image payload in that concrete representation with a source reference containing source UUID, digest, byte size, media type, dimensions, and transformation identity.
- Store exact prompt content and preserve omitted-versus-explicit parameter state.
- Include requested model, routing preferences, response-format requirements, and timeout/retry policy.
- Record application version/commit when available, adapter contract version, SDK package/version, and request-manifest schema version.
- Hash the canonical manifest representation for integrity checks.
- Test that credentials and embedded image data cannot enter the persisted manifest.
- Test that every field actually sent to the provider, including routing and future provider options, is represented or explicitly excluded by the manifest transform.
### 4. Capture OpenRouter Transport Evidence
- Evaluate the installed OpenRouter SDK hooks/client injection first.
- If hooks cannot expose an exact stable response before typed parsing, implement the non-streaming OpenRouter call through the existing async HTTP client boundary while retaining typed validation in the adapter.
- Read the response body once, preserve it exactly, then parse and normalize it.
- Store status, content type/encoding, allowlisted headers, request/generation ID, and timing.
- Maintain current authentication, referer/title headers, timeout behavior, and error classification.
- Explicitly document that the captured body is the OpenRouter-normalized transport response, not Gemini/Anthropic/OpenAI native upstream JSON.
- Add fixture-based tests proving unknown response fields survive transport capture even if a typed parser ignores them.
### 5. Preserve Failure Evidence
- Return or raise a typed provider failure that carries safe evidence separately from its user-facing error.
- Persist non-success status/body/allowlisted headers before marking an execution failed.
- Represent DNS/connect/TLS/local timeout failures as no-response outcomes with a failure phase and safe diagnostic category.
- Preserve response-validation failures with both the exact body and validation details.
- Keep transcription-quality rejection distinct from provider failure because a valid provider response was received.
- Ensure error strings and logs do not contain authorization data or embedded image payloads.
- Add tests for 4xx, 5xx, malformed JSON, schema mismatch, timeout, connection failure, and quality rejection.
### 6. Make Execution History Reliably Append-Only
- Confirm retry behavior creates a distinct execution attempt rather than reusing and overwriting a completed evidence record.
- Separate queue linkage from execution-attempt identity; the current update-in-place behavior cannot serve as append-only execution history.
- Assign each attempt a deterministic, monotonically increasing attempt number scoped to its Job and Source, enforced by a database uniqueness constraint.
- Update the latest-transcription projection only after a successful attempt.
- Never update prior response bodies, manifests, timings, or artifacts during a retry.
- Select the latest attempt and latest successful attempt by the persisted attempt number with a stable identifier as a defensive secondary key, never by timestamp alone.
- Add service/workflow tests covering retries, partial success, interrupted jobs, and historical projection behavior.
### 7. Add Generic Artifact Persistence
- Implement service operations to create, read, list, verify, export, and, only under explicit retention policy, delete artifacts.
- Validate semantic type, schema/version, digest, media type, and coordinate metadata.
- Support JSON artifacts inline initially when within the agreed size threshold.
- Support external artifacts through a constrained application-data root with atomic write, digest verification, and explicit missing-file errors.
- Add a provider-neutral example fixture representing OCR words/lines with polygons and confidence values.
- Do not integrate a live OCR vendor in this phase.
### 8. Add Evidence Inspection and Export
- Rename the current Source Detail label to identify historical values as an OpenRouter SDK Response Snapshot.
- Add separate sections for Request Manifest, Transport Response, Normalized Metadata, Software Context, and Derived Artifacts.
- Show an explicit “not captured for this historical execution” state instead of an empty object.
- Keep large bodies collapsed by default and avoid rendering embedded source data.
- Add a service-owned export that packages a versioned manifest, evidence JSON/body files, artifact content or references, and digest inventory.
- Exclude secrets and machine-local paths that are not required to interpret the evidence.
- Add UI and export tests for new, historical, failed, and large-evidence records.
### 9. Establish the Private Benchmark
- Select a small initial corpus, then expand only when it exposes meaningful differences.
- Stratify examples by printed/typed text, handwriting style, degradation, layout complexity, language, and editorial anomaly.
- Reference existing Source UUIDs and digests in a private manifest; do not copy family documents into public test fixtures.
- Create manually reviewed reference transcriptions following the invariant methodology.
- Implement or adopt existing project-compatible CER/WER calculations without changing dependencies unless justified.
- Score omissions, inventions, silent modernization, uncertainty markup, and layout fidelity separately from CER/WER.
- Record cost and latency from preserved execution evidence.
- Run the current `google/gemini-2.5-flash` configuration as the baseline before testing alternatives.
- Treat results as model-version/route/corpus specific and preserve each comparison run.
### 10. Verify, Migrate, and Align Documentation
- Run the smallest focused model, provider, service, workflow, UI, upgrade, and export test groups first.
- Run broader regression tests only after focused validation passes.
- Execute all destructive tests through `tools/run_destructive_tests.py`.
- Verify backup creation and required restoration behavior before any test touching real application data.
- Confirm existing Source Detail records remain readable after upgrade.
- Update V4 architecture, schema, requirements, and UI schema mappings to point to V4.2 semantics.
- Record any deliberate deviation from this plan in the V4.2 scope before release.
## Recommended Delivery Order
1. Typed/versioned evidence contracts and terminology.
2. Additive execution-evidence persistence.
3. Secret-safe request manifests.
4. Exact OpenRouter transport capture.
5. Failure evidence and append-only retry semantics.
6. Generic artifact persistence.
7. Inspection and export.
8. Private benchmark tooling and baseline run.
9. Migration, regression verification, and documentation alignment.
## Key Implementation Decisions to Resolve
1. Whether execution evidence is a one-to-one companion to `JobSource` or part of a new execution-attempt model required for append-only retries.
2. Whether exact response bodies remain database values at expected sizes or move to hashed external files above a threshold.
3. The canonical JSON algorithm used to hash request manifests.
4. The safe-header allowlist supported by OpenRouter and future adapters.
5. The application version identity available in local, packaged, and uncommitted development builds.
6. The initial inline/external artifact size threshold and application-data root.
7. Whether evidence exports include original source binaries by default, optionally, or only by reference.
8. The minimum private benchmark corpus size and review process before model comparisons influence defaults.
These decisions must be settled before their corresponding implementation phase; they do not weaken the invariant or expand V4.2 into live OCR integration.
## Resolved Implementation Decisions
1. `JobSource` remains queue linkage and a compatibility projection; immutable retries use a one-to-many
`ExecutionAttempt` model with a unique `(job_id, source_id, attempt_number)` constraint.
2. Exact OpenRouter response bytes remain database values for V4.2. Generic artifacts use inline canonical JSON up
to 1 MiB by default and constrained, atomically written external files above that threshold.
3. Request manifests use `transcription-canonical-json-v1`: UTF-8 JSON with sorted keys, compact separators,
preserved Unicode, and non-finite numbers rejected.
4. Safe response headers are explicitly allowlisted in the evidence contract; all others are discarded before
persistence.
5. Software identity records the package version, optional `TRANSCRIPTION_COMMIT`, adapter contract version,
OpenRouter SDK version, and Python version.
6. The artifact root defaults to `data/artifacts` and stores source-scoped relative references.
7. Evidence exports include source identity and digest by reference, not original source binaries.
8. The benchmark manifest is private and digest-referenced. Corpus size remains archive-dependent, but every run
uses preserved execution-attempt identity and the fixed literal scoring contract.
## Done When
- Every V4.2 acceptance criterion is satisfied by focused tests or an explicit demonstration.
- Existing SDK snapshots retain their content and are labeled accurately.
- New successful and failed calls preserve secret-safe provider-boundary evidence.
- Unknown transport fields survive even when the typed SDK/parser does not recognize them.
- Retries cannot overwrite prior execution evidence.
- A generic versioned artifact can represent OCR geometry and pass integrity verification.
- Evidence can be safely inspected and exported with schema identities and digests.
- The current model has a reproducible private benchmark baseline.
- No credential or embedded source payload appears in persisted manifests, safe headers, logs, or exports.
- Existing V4.1 behavior remains compatible.
## Out of Scope
- Live OCR/document-AI provider integration.
- Automatic model switching.
- Archive-wide reprocessing.
- Native upstream-provider response capture through OpenRouter when OpenRouter does not expose it.
- Guarantees of deterministic hosted-model output.
## Related Local References
- [V4.2 Scope Boundary](scope_boundary_v4_2.md)
- [Digital Evidence and AI Processing Provenance](../invariant/ai_evidence_and_provenance.md)
- [V4 Architecture](../ver4/architecture_v4.md)
- [V4 Schema](../ver4/schema_v4.md)
- [V4 Requirements](../ver4/requirements_v4.md)
- [Draft V4.3 Implementation Plan](../ver4.3/implementation_plan_v4_3.md)
-163
View File
@@ -1,163 +0,0 @@
# V4.2 Scope Boundary
This document defines the boundary for the digital-evidence and AI-provenance revision that follows V4.1 and precedes the planned V4.3 settings work. V4 remains the architecture baseline; V4.2 makes the existing evidence claims precise and adds a provider-neutral foundation for future processing artifacts.
## Purpose
- Align the application with the [Digital Evidence and AI Processing Provenance invariant](../invariant/ai_evidence_and_provenance.md).
- Preserve provider-boundary evidence before SDK parsing can remove unknown fields.
- Make successful and failed processing attempts inspectable without storing secrets.
- Support future OCR and layout outputs without coupling the database to one vendor.
- Establish a repeatable method for comparing transcription models against this archive.
## In Scope
### 1. Evidence Terminology and Existing-Data Compatibility
- Define transport response, router-normalized response, SDK response, normalized metadata, and derived artifact consistently in code, schema documentation, and UI labels.
- Treat existing `JobSource.raw_api_response` values as historical SDK response snapshots.
- Preserve every existing `Job`, `Source`, and `JobSource` row.
- Use additive migrations and compatibility reads; do not reinterpret previously stored values as exact transport captures.
- Correct the “Page-Level Execution and AI Outputs” rule in `docs/ver4/schema_v4.md` that currently describes `JOB_SOURCE` as storing a complete provider response envelope. The corrected rule must identify `raw_api_response` as an SDK-serialized OpenRouter response snapshot and state that it is neither the exact HTTP body nor the native upstream-provider response.
### 2. Secret-Safe Request Manifests
- Persist the effective request specification for each page execution without storing credentials or duplicate base64 media.
- Include requested provider/model, routing constraints, prompt content and hash, explicitly supplied parameters, source digest, media type, dimensions when known, and page identity.
- Distinguish an omitted optional parameter from an explicitly supplied null or value.
- Record application, provider-adapter, Python client, and relevant schema versions.
- Use source or derivative references in place of embedded media bytes.
### 3. Provider-Boundary Response Capture
- Capture the exact HTTP response body before OpenRouter SDK parsing for non-streaming transcription calls.
- Store HTTP status and an explicit allowlist of safe response headers.
- Store router request/generation identifiers and resolved model/provider-routing metadata when exposed.
- Preserve the current parsed SDK snapshot and normalized metadata where useful.
- Keep exact body, parsed representation, and normalized fields distinguishable.
### 4. Failure Evidence and Timing
- Create or update a page execution record for every attempted provider call.
- Persist safe response evidence for non-success HTTP responses.
- Distinguish HTTP response failures, connection failures, local timeouts, response-validation failures, and transcription-quality failures.
- Store execution start/end times or duration using a clearly defined clock policy.
- Do not collapse a provider error body into only a generic user-facing message.
### 5. Generic Processing Artifacts
- Add a provider-neutral representation for versioned derived artifacts.
- Support inline JSON and externally stored payloads with a digest and stable reference.
- Record artifact type, format, schema/version, producer/version, source, producing execution, and creation time.
- Define coordinate-system metadata sufficient for word, line, block, or page geometry.
- Permit future OCR/layout/confidence results without implementing a vendor-specific table for each provider.
### 6. Evidence Inspection and Export
- Expand Source Detail and/or Job Detail to identify the evidence layer being displayed.
- Provide readable JSON inspection for request manifests, transport metadata, parsed responses, normalized metadata, and derived artifacts.
- Provide a safe export containing evidence content or references, relationships, schema versions, and digests.
- Clearly label evidence that was not captured for historical records.
- Do not display or export credentials, unrestricted headers, or embedded base64 source media.
### 7. Representative-Corpus Benchmark Protocol
- Define a private benchmark manifest referencing source digests rather than duplicating archival media.
- Include representative printed, typed, handwritten, degraded, tabular, and spatially complex pages.
- Pair each benchmark item with a manually reviewed literal transcription.
- Score character error rate, word error rate, omissions, inventions, silent normalization, uncertainty handling, layout fidelity, cost, and latency.
- Preserve the complete execution provenance for every benchmark run.
- Keep the current model as a baseline; do not change the application default solely from vendor benchmarks.
### 8. Migration, Integrity, and Verification
- Provide non-destructive upgrade behavior for supported SQLite and PostgreSQL deployments.
- Backfill only facts that can be derived reliably from existing records.
- Mark unavailable historical evidence as unavailable rather than fabricating it.
- Add digest, serialization, header-allowlist, failure-path, compatibility, artifact, export, and UI inspection tests.
- Run destructive tests only through the repository's required backup-and-restore wrapper.
## Out of Scope
- Selecting or declaring a permanent best transcription model.
- Changing the default transcription model without benchmark evidence and a separate decision.
- Integrating Azure Document Intelligence, Google Document AI, Transkribus, Mistral OCR, or another OCR provider in V4.2.
- Generating bounding boxes retroactively for existing transcriptions.
- Bulk reprocessing the archive.
- Packet capture, TLS evidence, full unrestricted request/response headers, or credential retention.
- Storing duplicate base64 source images in request manifests.
- Guaranteeing byte-identical reproduction from nondeterministic or updated hosted models.
- Automatic entity extraction, biography generation, or genealogical inference.
- Replacing the relational database with an event store or content-addressed object store.
- Destructive renaming or removal of `raw_api_response`.
## Locked Design Decisions
### A. The Original Source Is Primary Evidence
- Original uploaded bytes and their digest remain authoritative.
- Processing derivatives and outputs are independently identified derived evidence.
- Future OCR/layout work reuses the original or a documented derivative.
### B. Evidence Is Layered
- Exact transport evidence, SDK-parsed objects, normalized metadata, and transcription text serve different purposes.
- One representation must not silently stand in for another.
- UI and export labels name the stored evidence layer.
### C. History Is Append-Only
- A retry or reprocessing attempt creates new execution evidence.
- Convenience caches may change, but historical execution output does not.
- Human revisions remain separate from machine output.
### D. Capture Is Secret-Safe by Construction
- Safe headers are allowlisted.
- Authorization, cookies, API keys, and unrestricted headers are never persisted.
- Request manifests reference source digests instead of embedding source bytes.
### E. Derived Artifacts Are Generic and Versioned
- Artifact storage is not limited to bounding boxes.
- Coordinate metadata declares units, origin, dimensions, and transformations.
- Provider-specific payloads may be retained without making provider-specific fields the durable application contract.
### F. Existing Evidence Keeps Its Original Meaning
- Existing `raw_api_response` data remains an SDK response snapshot.
- A migration may label or classify it but may not claim that missing transport data was captured.
- Historical nulls and absent fields remain distinguishable from new explicitly captured values.
## Data and Compatibility Policy
- All schema changes are additive in V4.2.
- Existing source files, hashes, transcriptions, revisions, prompts, jobs, and relationships remain valid.
- Compatibility reads continue to display historical SDK snapshots.
- Large derived artifacts may be stored outside the database when the database retains a stable reference, digest, media type, and schema identity.
- JSON evidence must remain portable across SQLite and PostgreSQL.
- Exports use explicit schema versions so later releases can interpret older packages.
## Acceptance Criteria
1. A new execution can be traced from its source digest through its frozen request manifest, transport response, parsed/normalized data, and derived outputs.
2. Exact response content is captured before SDK parsing and is clearly distinguished from the existing SDK snapshot.
3. Failed HTTP calls retain safe provider evidence; calls with no response record that fact explicitly.
4. Omitted parameters remain distinguishable from explicit values.
5. No persisted request, header set, UI display, log, or export contains API credentials.
6. Retrying or reprocessing does not overwrite prior execution evidence.
7. Historical records remain readable and are not mislabeled as exact transport captures.
8. A versioned generic artifact can represent OCR/layout JSON and its coordinate system without a provider-specific schema change.
9. Evidence exports include relationships, schema identities, and digests sufficient for independent integrity checks.
10. The benchmark protocol can compare the current baseline with another model on the same private corpus and scoring rules.
11. Additive migrations and focused tests work across the supported persistence model.
12. All destructive-test runs comply with the backup-and-restore protocol.
## Related Local References
- [V4.2 Implementation Plan](implementation_plan_v4_2.md)
- [Digital Evidence and AI Processing Provenance](../invariant/ai_evidence_and_provenance.md)
- [V4 Architecture](../ver4/architecture_v4.md)
- [V4 Schema](../ver4/schema_v4.md)
- [V4 Requirements](../ver4/requirements_v4.md)
- [Transcription Methodology](../invariant/transcription_methodology.md)
-121
View File
@@ -1,121 +0,0 @@
# Implementation Plan (Version 4.3)
## Goal
Deliver constrained, installation-local application settings while preserving the completed V4.2 behavioral baseline and historical provenance.
## Planning Constraints
- V4, V4.1, and V4.2 remain the behavioral baseline.
- Settings must use explicit domain operations rather than direct database, environment-file, or arbitrary filesystem access from UI pages.
- Prompt changes must preserve historical Job provenance and use a defined safe-write policy.
- Source Page Reordering is excluded.
- Database, integration, and UI tests must use confirmed isolated test data and must never modify `data/transcription.db`.
- Potentially destructive tests must run only through `tools/run_destructive_tests.py`.
## Expected Project Impact
| Area | Expected impact |
| --- | --- |
| Documents service | Expand controlled Document Type maintenance operations. |
| People service | Expand controlled Person Role maintenance operations. |
| Prompt adapter/service | Add constrained listing, reading, validation, atomic writing, backup, and explicit recovery of existing prompt artifacts. |
| UI composition/navigation | Register Settings routes and navigation without moving persistence into UI code. |
| Tests | Add isolated registry lifecycle, prompt safety, and UI workflow coverage. |
## Implementation Phases
### 1. Define Service Contracts
- Define Document Type maintenance commands for create, relabel, activate, deactivate, and delete-if-unreferenced.
- Define Person Role maintenance commands for create, relabel, activate, deactivate, and delete-if-unreferenced.
- Define a Prompt Store interface for constrained list, read, write, backup-status, and explicit recovery behavior.
- Map validation, conflict, not-found, dependency, and filesystem failures to existing `AppError` categories.
### 2. Expand Registry Maintenance Services
- Reuse existing Document and People service ownership.
- Add explicit write methods rather than passing UI-mutated ORM objects directly where practical.
- Normalize Document Type labels and reject case-insensitive duplicates deterministically.
- Keep Person Role stable-code validation and duplicate rejection.
- Permit deletion only after a service-owned reference check proves the entry is unreferenced.
- Reject deletion of referenced entries deterministically without partial mutation.
- Permit label changes whether or not an entry is referenced.
- Preserve inactive entries for historical reads.
- Order Document Types alphabetically by normalized label.
- Order Person Roles deterministically by label and then code without adding a schema field.
- Add service tests for create, relabel, activation, deactivation, duplicates, immutable codes, ordering, allowed deletion, and blocked referenced deletion.
### 3. Add Constrained Prompt Storage
- Place filesystem access behind a dedicated Prompt Store/service boundary.
- Resolve all filenames directly beneath the configured prompt root and reject traversal.
- Permit only existing files with the agreed Markdown extension and reject empty content.
- Exclude prompt creation and deletion.
- Write new content to a sibling temporary file, flush and sync it, preserve the active file as the sole previous-version backup, and atomically replace the active file.
- Expose explicit backup recovery through the same filename validation and safe-write path; never perform automatic rollback.
- Clean up temporary files after failed writes while preserving the active prompt and any valid backup.
- Preserve file encoding and provide explicit failures for read-only or unavailable storage.
- Do not modify any Job row when prompt defaults change.
- Add unit tests for valid reads/writes, traversal, invalid names, nonexistent-file creation attempts, empty content, atomic replacement failures, single-backup rotation, explicit recovery, filesystem failures, and unchanged Job provenance.
### 4. Build the Settings UI
- Register a Settings landing page and navigation entry.
- Add separate pages or panels for Document Types, Person Roles, and Prompts.
- Keep pages responsible for orchestration and notifications only.
- Use service callbacks for all mutations.
- Explain inactive historical entries and future-only prompt effects in the UI.
- Present deletion only for unreferenced registry entries and preserve clear conflict feedback if references appear before submission.
- Present prompt backup availability and recovery as an explicit operator action.
- Do not render raw environment values or secrets.
- Add no settings API routes.
### 5. Verification and Rollout
- Confirm every database, integration, and UI test is configured for an isolated test database before execution.
- Never run those tests against live data and never modify or replace `data/transcription.db`.
- Invoke potentially destructive tests only through `tools/run_destructive_tests.py`.
- Run focused service tests before UI integration tests.
- Verify inactive registry behavior in both historical display and create/edit selectors.
- Verify referenced entries can be relabeled or deactivated but not deleted.
- Verify unreferenced entries can be deleted.
- Verify prompt changes are picked up by newly created Jobs while historical Jobs retain frozen content/hash.
- Run the relevant regression suite.
## Migration and Compatibility Notes
- Existing registry records remain valid.
- Prompt editing changes mutable application files, not database provenance already captured on Jobs.
- V4.3 must not require users to recreate existing Sources, Documents, People, roles, or types.
- Person Role ordering requires no schema migration.
- Registry deletion introduces no cascade behavior; references always block deletion.
## Delivery Order
1. Implement registry maintenance service operations.
2. Implement the Prompt Store and safety policy.
3. Build Settings pages.
4. Run isolated integration and regression verification.
## Done Criteria
- All V4.3 acceptance criteria are testable and satisfied.
- Settings mutations cross explicit service or adapter boundaries.
- Document Types use UUID-only identity and unique labels; Person Role codes cannot be accidentally changed.
- Referenced registry entries can be relabeled or deactivated but cannot be deleted.
- Unreferenced registry entries can be deleted without cascade behavior.
- Prompt writes cannot escape the configured directory or rewrite historical provenance.
- Prompt writes are atomic, retain one backup, and support explicit recovery.
- No secret or raw environment editor exists.
- No settings API surface exists.
- V4.1 and V4.2 workflows remain intact.
- Verification does not touch live data or `data/transcription.db`.
## Related Local References
- [V4.3 Scope Boundary](scope_boundary_v4_3.md)
- [V4.2 Implementation Plan](../ver4.2/implementation_plan_v4_2.md)
- [V4.1 Implementation Plan](../ver4.1/implementation_plan_v4_1.md)
- [V4 Implementation Plan](../ver4/implementation_plan_v4.md)
- [V4 Error Handling Policy](../ver4/error_handling_v4.md)
-154
View File
@@ -1,154 +0,0 @@
# V4.3 Scope Boundary
This document defines the frozen boundary for the constrained-settings revision that follows the completed V4.2 evidence-and-provenance work. V4, V4.1, and V4.2 remain the behavioral and architecture baseline.
## Purpose
- Provide a constrained Settings area for safe maintenance of selected application-managed configuration.
- Avoid exposing secrets, restart-sensitive settings, or unrestricted filesystem editing through the UI.
## In Scope
### 1. Settings Navigation
- Add a Settings entry to application navigation.
- Provide separate, clearly described settings areas rather than a raw configuration editor.
- Restrict V4.3 settings to application-managed values that can be validated and safely changed at runtime.
### 2. Document Type Maintenance
- List active and inactive Document Types.
- Add new types with a unique user-facing label.
- Edit labels and active state.
- Activate or deactivate types without invalidating historical Documents.
- Allow deletion only when no Document references the type.
- Allow label changes regardless of whether the type is referenced.
- Display types alphabetically by label.
### 3. Person Role Maintenance
- List active and inactive Person Roles.
- Add new roles with a stable unique code and user-facing label.
- Edit mutable labels.
- Activate or deactivate roles without invalidating historical links.
- Do not allow changing a stable code after creation.
- Order roles deterministically by label and then code; do not add persisted role sort order.
- Allow deletion only when no document-person link references the role.
- Allow label changes regardless of whether the role is referenced.
### 4. Prompt Maintenance
- List prompt markdown files from the configured prompt directory.
- View a prompt with a concise explanation of its purpose and use.
- Edit an existing prompt as plain markdown text.
- Validate the filename boundary and reject empty prompt content.
- Save changes explicitly and report filesystem failures.
- Preserve submission-time prompt text and hash already frozen on existing Jobs.
- Edit existing prompt files only; prompt creation and deletion are excluded.
- Save through a sibling temporary file, flush and sync file content, retain one previous-version backup, and atomically replace the active file.
- Provide an explicit recovery operation that restores the retained backup through the same safe-write path; do not silently roll back a failed or unwanted edit.
### 5. Deployment Boundary
- Settings changes apply only to the current installation.
- V4.3 adds no settings API endpoints.
- Service contracts must remain independent of the UI so a separately authorized API can be considered later.
## Out of Scope
- Viewing or editing raw `.env` files.
- Displaying or changing provider API keys and other secrets.
- Editing host, port, database connection, upload paths, or other restart-sensitive runtime settings.
- Arbitrary file browsing or arbitrary prompt paths.
- Runtime theme/CSS editing.
- Installing themes or plugins.
- Source page renumbering or reordering.
- Source movement between Documents.
- Automatic ordering based on filenames, OCR, or image content.
- Prompt creation, deletion, and multi-version history.
- Persisted sort-order maintenance for Person Roles.
- Settings read or write API endpoints.
- FamilySearch API synchronization.
- A generic external-reference registry.
- Ancestry references and Google Maps links.
## Locked Design Decisions
### A. Registry Identity and Lifecycle
- Document Types use UUID identity and case-insensitively unique labels; no separate code is exposed or stored.
- Person Role codes remain stable identifiers.
- Labels and active state remain mutable.
- Historical references remain valid when a registry entry is inactive.
- Labels may be updated for referenced and unreferenced entries.
- Unreferenced entries may be deleted; referenced entries may only be deactivated.
### B. No Raw Environment Editor
- `.env` may contain secrets and values that are not safely reloadable.
- V4.3 exposes only purpose-built forms backed by explicit validation and service methods.
### C. Prompt Editing Is Constrained
- Prompt maintenance is limited to direct children of the configured prompt directory.
- Existing Job provenance is never rewritten when a prompt file changes.
- The UI must distinguish editing the default for future submissions from inspecting historical Job prompts.
### D. Prompt Writes Are Atomic and Recoverable
- Writes use a sibling temporary file and atomic replacement so readers observe either the old or new complete prompt.
- The immediately previous prompt version is retained as the sole backup.
- Recovery is an explicit operator action and uses the same validated safe-write path.
- Prompt creation and deletion are not available in V4.3.
### E. Person Role Ordering Is Deterministic, Not Persisted
- Person Roles are ordered by label and then stable code.
- V4.3 does not add a `sort_order` field to Person Roles.
- Document Types use alphabetical label ordering and have no persisted sort order.
### F. Settings Are Installation-Local
- V4.3 provides Settings through the local application UI and domain services only.
- No settings API surface is introduced.
## Data and Compatibility Policy
- V4.3 does not rewrite existing Documents, document-person links, Jobs, Sources, execution evidence, or prompt provenance.
- Deactivation preserves referenced registry entries for historical display while excluding them from default create selectors.
- Deletion checks are performed at the service boundary and must fail deterministically when references exist.
- Prompt files are constrained to existing Markdown files that are direct children of the configured prompt root.
- Settings UI code performs no direct database, environment-file, or arbitrary filesystem mutations.
- Source page numbering and ordering behavior is unchanged.
## Acceptance Criteria
1. Document Type UUID identity and Person Role stable codes preserve historical references.
2. Inactive registry entries remain visible on historical records but are excluded from default create selectors.
3. Labels can be changed for referenced or unreferenced registry entries.
4. An unreferenced Document Type or Person Role can be deleted, while deletion of a referenced entry fails without partial mutation.
5. Document Types use alphabetical label ordering; Person Roles use deterministic label/code ordering.
6. Prompt edits are restricted to existing Markdown files directly beneath the configured prompt directory.
7. Prompt saves use atomic replacement, retain exactly one previous-version backup, and support explicit recovery.
8. A prompt edit affects future Jobs only and leaves stored Job provenance unchanged.
9. No Settings page exposes secrets, unrestricted filesystem access, or a settings API.
10. Focused service and UI tests pass without regressing V4.1 or V4.2 workflows.
11. Database, integration, and UI tests use confirmed isolated test data and never modify `data/transcription.db`; potentially destructive tests run only through `tools/run_destructive_tests.py`.
## Scope Freeze Gate
V4.3 is sufficiently frozen to begin implementation:
- V4.2 is the completed behavioral baseline.
- Registry lifecycle and ordering behavior are resolved.
- Prompt lifecycle, atomic-write, backup, and recovery behavior are resolved.
- The installation-local deployment boundary is resolved.
- The implementation plan is a committed delivery plan.
## Related Local References
- [V4.3 Implementation Plan](implementation_plan_v4_3.md)
- [V4.2 Scope Boundary](../ver4.2/scope_boundary_v4_2.md)
- [V4.1 Scope Boundary](../ver4.1/scope_boundary_v4_1.md)
- [V4 Architecture](../ver4/architecture_v4.md)
- [V4 Schema](../ver4/schema_v4.md)
-189
View File
@@ -1,189 +0,0 @@
# Implementation Plan (Version 4.4)
## Goal
Deliver hidden semantic identity for built-in registries, a single atomic Linked People workflow, and safe browser-native printing of archival Documents and their current transcriptions.
## Planning Constraints
- V4.3 is the completed implementation baseline.
- V4.4 may replace V4/V4.3 registry and document-person contracts only as specified by the V4.4 scope.
- Semantic keys are internal and immutable; UI and public API contracts use UUIDs and labels.
- Document and link edits must not partially commit.
- Print output must not execute stored text or expose machine-local source paths.
- Source page reordering remains excluded.
- Database, integration, and UI tests must use confirmed isolated data and never modify `data/transcription.db`.
- Potentially destructive tests must run only through `tools/run_destructive_tests.py`.
## Expected Project Impact
| Area | Expected impact |
| --- | --- |
| Models and schema bootstrap | Add nullable unique semantic keys, simplify document-person identity, and seed frozen built-ins. |
| Document service | Maintain built-in Document Types, usage summaries, UUID assignment, and atomic Document/link writes. |
| People service | Maintain built-in Person Roles, link summaries, UUID-only role assignment, and one-person-per-document enforcement. |
| V4 document API | Remove role-code selectors and compatibility role fields; enforce UUID-only relationship writes. |
| Settings UI | Use matching table workflows for Document Types and Person Roles. |
| Document Create/Edit | Replace role-specific multiselects with one staged Linked People table and inline editor. |
| Document Detail/printing | Add format selection, print preview, safe Source media rendering, print CSS, and job metadata. |
| Tests and documentation | Replace superseded cardinality/identity assertions and add isolated registry, editor, transaction, and print coverage. |
## Implementation Phases
### 1. Align Durable Registry Contracts
- Add nullable, unique `semantic_key` fields to `DocumentType` and `PersonRole`.
- Keep UUIDs as primary and foreign-key identity.
- Add normalized-label storage and uniqueness to Person Roles using the same trim and case-normalization policy as Document Types.
- Remove the user-created Person Role code contract.
- Define built-in detection as `semantic_key is not null`.
- Centralize the frozen built-in definitions in one domain-owned location.
- Seed six Document Types and three Person Roles idempotently.
- Ensure label edits never change semantic keys.
- Reject deletion of every built-in before checking references.
- Continue blocking deletion of referenced custom entries.
- Return deterministic validation, conflict, dependency, and not-found errors through existing error categories.
### 2. Establish the Clean Schema
- Remove legacy `DocumentPerson.role` compatibility storage and the fixed `DocumentPersonRole` enum.
- Make `DocumentPerson.role_id` required.
- Replace role-specific uniqueness with a unique `(document_id, person_id)` constraint.
- Remove obsolete Document Type and Person Role migration paths that exist only for disposable development data.
- Keep fresh schema creation and built-in seeding portable across SQLite and PostgreSQL.
- Make configured development-database recreation a separate operator-confirmed step that displays the resolved target path rather than assuming `app.db` or `data/transcription.db`.
- Never invoke recreation from application startup or test setup.
- Add isolated schema tests for fresh creation, seed idempotence, semantic-key uniqueness, normalized-label uniqueness, required roles, and link uniqueness.
### 3. Refine Registry Services and API Contracts
- Add summary queries for Document counts and Person Role link counts without per-row queries.
- Order both registries by normalized label with UUID as a deterministic tie-breaker.
- Expose built-in status as a derived read value where the Settings UI needs it.
- Keep semantic-key lookup behind service methods for application-owned behavior such as resolving authors.
- Ensure create operations always produce custom entries with null semantic keys.
- Ensure update operations accept only label and active state.
- Remove `role_code` request alternatives and compatibility role responses from the V4 document API.
- Require `role_id` for document-person creation and updates.
- Add service and API tests for hidden semantic identity, relabeling, activation, built-in protection, custom deletion, counts, ordering, UUID-only writes, and conflicts.
### 4. Build Matching Settings Tables
- Retain the existing Document Types table workflow and add the Built-in column.
- Replace the current per-row Person Role controls with the same selection-based table pattern.
- Render the agreed columns and usage counts.
- Keep labels as the only registry text shown in selectors.
- Add creates custom entries only.
- Edit dialogs expose label and active state only.
- Delete reports protected-built-in and referenced-custom conflicts clearly.
- Avoid direct persistence queries from the Settings page.
- Add component-level UI assertions for columns, actions, label-only selectors, and immutable built-in presentation.
### 5. Add a Staged Linked People Editor
- Introduce a small typed UI-state model for staged `(person_id, role_id)` rows rather than storing raw widget values.
- Share the editor component between Create Document and Edit Document.
- Render a multi-selection table with Person and Role labels.
- Add an inline editor whose mode is explicitly Add or Edit.
- Disable already-linked People when adding; retain the edited Person as an option during Edit.
- Require exactly one row for Edit and allow one or more rows for Delete.
- Save and Delete mutate only staged UI state.
- Cancel discards only the active inline edit.
- Preserve inactive-role historical rows in Edit while restricting new assignments and changes to active roles.
- Preserve `person_id` preselection by staging that Person with the active built-in `author` role, with warning behavior for invalid or unavailable selections.
- Preserve the `return_to=jobs_new` success path.
- Keep navigation to Person creation separate; V4.4 does not add an embedded Person editor.
- Add UI tests for staging, duplicate prevention, selection rules, inactive roles, cancel behavior, and both Document forms.
### 6. Persist Document and Links Atomically
- Add service commands for Create Document with complete links and Update Document with complete links.
- Validate Document Type, every Person, every Person Role, active assignment rules, and duplicate People before mutation.
- Compute deterministic add, update, and remove deltas for Edit.
- Apply Document and link mutations in one database transaction and commit once.
- Roll back the complete operation on any validation, conflict, or persistence failure.
- Return the persisted Document detail required by the UI after success.
- Reuse these commands from UI orchestration rather than sequencing independent service commits.
- Add failure-injection tests proving no partial Document or link mutation survives.
### 7. Define a Print Projection
- Add a read-only service projection containing:
- Document title and selected archival metadata.
- Authors resolved by the `author` semantic key.
- Notes.
- Ordered Sources with application media URLs and current transcription text.
- Ordered Job metadata.
- Load the projection with bounded queries and deterministic ordering.
- Use non-null `revised_text`, including an intentionally empty revision; otherwise fall back to `raw_transcription`.
- Map empty or whitespace-only current text to the explicit unavailable state without falling back past an intentional revision.
- Represent unavailable text and optional metadata explicitly.
- Do not expose semantic keys, direct file paths, full prompts, provider evidence, or raw API responses.
- Keep the projection independent of NiceGUI rendering so formatting tests can use plain typed values.
### 8. Build Print Preview and Styles
- Add a Print action to Document Detail.
- Open a dedicated persisted-Document print route with a Facsimile/Text-only format choice.
- Render the exact content order frozen in the scope.
- Keep print metadata tables content-sized, with a non-wrapping label column and wider wrapping value columns.
- Render stored Notes and transcription as escaped text.
- For Text-only mode, normalize whitespace by joining single line breaks inside paragraphs while preserving blank-line paragraph boundaries.
- For Facsimile mode, preserve line breaks and use a two-column Source layout.
- Start each Facsimile Source on a new printed sheet with CSS page breaks.
- Allow long transcription content to continue rather than clipping it.
- Fetch images through an application-controlled Source media route.
- Add print-only CSS that hides navigation, controls, and non-document chrome.
- Invoke the browser print dialog only from an explicit user action.
- Add rendering tests for both modes, missing data, long text, special characters, image URLs, and page ordering.
### 9. Align Documentation and Verification
- Update V4 architecture, requirements, schema, and Document UI contracts to reflect:
- UUID plus hidden semantic-key registries.
- Built-in protection.
- One Person per Document.
- UUID-only role API writes.
- Atomic Document/link synchronization.
- Browser-native print projection and formats.
- Confirm every database, integration, and UI test target is isolated before execution.
- Run focused registry and service tests first.
- Run schema tests only through the destructive-test wrapper when they are potentially destructive.
- Run Linked People UI and print rendering tests against isolated fixtures.
- Run the broader non-external regression suite after focused coverage passes.
- Verify that `data/transcription.db` was not changed by test execution.
## Delivery Order
1. Registry and clean-schema contracts.
2. Registry services, API changes, and Settings tables.
3. Atomic Document/link service commands.
4. Shared staged Linked People editor.
5. Print projection.
6. Print preview and styles.
7. Documentation alignment and regression verification.
## Done Criteria
- All V4.4 acceptance criteria are implemented and testable.
- UI and API contracts use UUID identity and never expose semantic keys.
- Built-in registries retain meaning after relabeling and cannot be deleted.
- Custom registries retain reference-aware deletion.
- Both Document forms use one Linked People table.
- One Person cannot be linked twice to the same Document.
- Main Document saves are atomic across fields and relationships.
- Print preview provides both frozen formats and content sections.
- Print output uses current human-preferred text, deterministic ordering, escaped content, and application media URLs.
- Job metadata lists every Job oldest-to-newest and ends with Status.
- Source page reordering and server-generated PDFs are not introduced.
- Verification uses isolated data and does not modify `data/transcription.db`.
## Related Local References
- [V4.4 Scope Boundary](scope_boundary_v4_4.md)
- [V4.3 Scope Boundary](../ver4.3/scope_boundary_v4_3.md)
- [V4.3 Implementation Plan](../ver4.3/implementation_plan_v4_3.md)
- [V4 Architecture](../ver4/architecture_v4.md)
- [V4 Schema](../ver4/schema_v4.md)
- [V4 Requirements](../ver4/requirements_v4.md)
- [V4 Error Handling Policy](../ver4/error_handling_v4.md)
-233
View File
@@ -1,233 +0,0 @@
# V4.4 Scope Boundary
This document defines the frozen boundary for the semantic-registry, linked-people, and document-printing revision that follows the completed V4.3 Settings work. V4 through V4.3 remain the architecture and behavioral baseline except where this document explicitly replaces a registry or document-person contract.
## Purpose
- Keep registry identifiers stable without exposing duplicate machine codes in Settings tables or selectors.
- Replace role-specific person selectors with one coherent Linked People editor.
- Provide an archival print view containing document metadata, source pages, current transcription text, and transcription-job metadata.
## In Scope
### 1. Semantic Registry Identity
- `DocumentType` and `PersonRole` use UUIDs as their canonical record and relationship identity.
- Both registries may carry a nullable, unique, immutable `semantic_key` used only for application-defined built-ins.
- Semantic keys are internal implementation details. Settings tables, selectors, and public API payloads do not display or accept them.
- Labels are trimmed, case-insensitively unique, editable, and used for all user-facing display.
- Active state remains editable. Inactive entries remain valid for historical records and are excluded from default assignment selectors.
- Built-in status is derived from the presence of a semantic key and is displayed as a read-only Yes/No value.
- Built-in entries cannot be deleted or converted to custom entries.
- Custom entries have no semantic key and may be deleted only when unreferenced.
- Users may create custom entries but cannot create, change, or assign semantic keys through the UI or API.
### 2. Built-In Document Types
- Seed these built-in semantic keys and initial labels:
| Semantic key | Initial label |
| --- | --- |
| `book` | Book |
| `letter` | Letter |
| `postcard` | Postcard |
| `photo` | Photo |
| `journal` | Journal |
| `form` | Form |
- The Document Types Settings table contains Select, Label, Documents, Active, and Built-in columns.
- Document Types are ordered alphabetically by normalized label.
- Add, Edit, and Delete actions operate on table selection.
- Add creates a custom type. Edit changes only label and active state.
- The Documents count is the number of Documents referencing the type.
- Document Type selectors display labels only and submit UUIDs.
### 3. Built-In Person Roles
- Seed these built-in semantic keys and initial labels:
| Semantic key | Initial label |
| --- | --- |
| `author` | Author |
| `recipient` | Recipient |
| `mentioned` | Mentioned |
- Application behavior that requires authorship resolves the built-in `author` semantic key rather than matching a mutable label.
- The Person Roles Settings table contains Select, Label, Links, Active, and Built-in columns.
- Person Roles are ordered alphabetically by normalized label.
- Add, Edit, and Delete actions operate on table selection.
- Add creates a custom role. Edit changes only label and active state.
- The Links count is the number of document-person relationships referencing the role.
- Person Role selectors display labels only and submit UUIDs.
### 4. Linked People Editor
- Replace the separate role-specific person selectors on both Create Document and Edit Document with one Linked People table.
- The table contains Select, Person, and Role columns.
- Add opens an inline editor beneath the table with Person and Person Role selectors.
- Edit requires exactly one selected row and loads it into the inline editor.
- Save stages the inline addition or edit in the table.
- Cancel exits the inline editor without changing the staged link set.
- Delete stages removal of one or more selected rows.
- A Person may be linked to a Document only once, regardless of role.
- Every link has exactly one Person Role.
- Already-linked People are unavailable when adding another row.
- Existing links using inactive roles remain visible and unchanged unless explicitly edited.
- Only active roles are available for new links or role changes.
- Create Document preserves the existing `person_id` preselection workflow by staging that Person with the active built-in `author` role. An invalid Person or unavailable Author role produces a warning rather than an invalid link.
- Create Document preserves the existing `return_to=jobs_new` success path.
- Linked People changes remain staged until the main Create Document or Save Changes action.
- The Document and its complete staged link set are persisted atomically. A conflict or validation failure leaves both unchanged.
- The API and service contracts identify roles by `role_id`; role-code selectors and compatibility role strings are removed.
- Persistence enforces uniqueness on `(document_id, person_id)`.
### 5. Document Print View
- Add a Print action to Document Detail.
- The action opens a dedicated print-preview page for the persisted Document.
- The preview offers two formats:
- **Facsimile:** source image on the left and current transcription on the right. Original transcription line breaks are preserved, and each Source begins on a new printed sheet.
- **Text only:** no source images. Single line breaks inside a paragraph are reflowed as spaces, while blank-line paragraph boundaries remain.
- Both formats use browser printing through a dedicated print stylesheet and the browser print dialog.
- Server-generated PDF files are not part of V4.4; users may select the browser's Save as PDF destination.
- Sources are ordered by existing `page_number`, with UUID as a deterministic tie-breaker.
- The current transcription for each Source is the non-null `revised_text`, including an intentionally empty revision, otherwise the latest successful machine-output projection in `raw_transcription`.
- Empty or whitespace-only current text displays the explicit unavailable message rather than falling back past an intentional revision.
- A Source with no current transcription displays an explicit unavailable message.
- Transcription and Notes content is treated as text and escaped; model output is not executed as arbitrary HTML.
- Facsimile images use an application-controlled Source media route. Generated markup does not expose direct machine-local file paths.
### 6. Printed Content Contract
The print view contains, in this order:
1. Document title using the Document name.
2. Archival Metadata table:
- Author, containing People linked through the built-in `author` role.
- Document Type.
- Date.
- Location Created.
- Archival Identifier.
3. Notes.
4. Document section containing one numbered section per Source.
5. Transcription Job Metadata table.
Empty metadata values remain visible as `Not set`. Empty Notes display `No notes recorded`.
The job metadata table:
- Lists field names in the first column and adds one column for every Job associated with the Document.
- Orders Job columns from oldest to newest by creation date, then UUID.
- Includes every Job status: `queued`, `processing`, `transcribed`, `completed`, `partial_success`, and `failed`.
- Contains these rows in order:
- Job ID.
- Date, using the Job creation/submission timestamp with timezone.
- Provider.
- Model.
- Prompt, using the frozen prompt filename/name rather than full prompt content.
- Retry Count.
- Status as the final row.
- Displays `Not set` for unavailable optional metadata.
### 7. Clean Development Schema
- V4.4 does not require preservation or migration of rows in the operator-configured development database.
- Implementation may recreate the configured development database, including `data/transcription.db` when it is the explicitly selected target, only through a separate operator-confirmed action that identifies the resolved path. Startup and test execution never delete it automatically.
- Fresh schema creation seeds the agreed built-in Document Types and Person Roles idempotently.
- No test may use, modify, replace, or restore `data/transcription.db`.
- Database, integration, and UI tests use confirmed isolated databases.
- Potentially destructive tests run only through `tools/run_destructive_tests.py`.
## Out of Scope
- User creation, editing, deletion, or direct display of semantic keys.
- Treating custom registry entries as built-ins.
- Additional built-in Document Types or Person Roles beyond the frozen lists.
- Assigning more than one role to the same Person on the same Document.
- Preserving multiple historical links that violate the new one-person-per-document constraint.
- Source page renumbering or reordering.
- Printing unsaved Create/Edit Document state.
- Print actions on Job Detail or other pages.
- Batch printing multiple Documents.
- Server-side PDF generation or PDF file storage.
- Markdown, DOCX, or evidence-package export through the print feature.
- User-editable print templates, fonts, margins, headers, or footers.
- Full frozen prompt content, prompt hashes, transport evidence, API responses, or execution-attempt details in the print footer.
- Rendering transcription text as unrestricted Markdown or HTML.
- Pixel-identical pagination across browsers and printer drivers.
## Locked Design Decisions
### A. UUID Identifies the Row; Semantic Key Identifies Built-In Meaning
- UUIDs remain the only relationship and API identity.
- A hidden semantic key permits reliable built-in behavior after a label is renamed.
- Mutable labels are never used to infer built-in meaning.
### B. Built-Ins Are Protected but Mutable in Presentation
- Built-in labels and active state may change.
- Built-in semantic identity cannot change, and built-ins cannot be deleted.
- Custom entries remain reference-aware and deletable when unreferenced.
### C. Linked People Is a Single-Role Relationship
- One `(document_id, person_id)` row represents the complete relationship.
- Changing a role updates that row rather than adding another relationship.
- The main Document save owns one atomic Document-and-links transaction.
### D. Printing Uses the Current Human-Preferred Text
- Human-revised text takes precedence over the latest successful machine-output projection.
- Job metadata provides processing context but does not claim that a later human revision is raw output from a listed Job.
### E. Printing Is Browser-Native
- A print-specific HTML view and CSS support physical printing and browser Save as PDF.
- Source media is served through application-controlled routes, and all textual content is escaped.
### F. Source Order Is Read-Only in V4.4
- Print order follows existing page numbers.
- Source page reordering remains explicitly excluded.
## Acceptance Criteria
1. Registry selectors and Settings forms never display a machine code or semantic key.
2. Document Types and Person Roles use UUID relationship identity and case-insensitively unique labels.
3. The six Document Type and three Person Role built-ins are seeded with immutable internal semantic keys.
4. Built-ins may be relabeled or disabled but cannot be deleted.
5. Unreferenced custom entries may be deleted; referenced custom entries may only be relabeled or disabled.
6. Settings tables show the agreed columns, alphabetical label order, usage counts, and selection-based actions.
7. Create and Edit Document use one Linked People table with inline staged Add/Edit/Save/Cancel and multi-row Delete.
8. The same Person cannot be staged or persisted twice for one Document, even under different roles.
9. Document fields and Linked People changes commit atomically.
10. Historical inactive roles remain displayable, while only active roles are assignable.
11. Document Detail opens a print preview with Facsimile and Text-only formats.
12. Print pages use current revised text when available and deterministic Source ordering.
13. Printed archival metadata resolves authors through the hidden `author` semantic key after any label change.
14. The job table contains one oldest-to-newest column per Job and ends with the Status row.
15. Print output escapes stored text and does not disclose direct local source paths.
16. Source reordering, server PDF generation, and print-template editing are absent.
17. Database, integration, and UI verification uses isolated test data and never modifies `data/transcription.db`.
## Scope Freeze Gate
V4.4 is sufficiently frozen to begin implementation:
- Built-in registry identity, membership, lifecycle, display, and selector behavior are resolved.
- Linked People selection, editing, uniqueness, inactive-role, staging, and transaction behavior are resolved.
- Print entry point, formats, content order, transcription precedence, page order, job metadata, and output mechanism are resolved.
- Clean development-schema and destructive-test boundaries are resolved.
- Source page reordering remains excluded.
Any expansion of the built-in catalogs, relationship cardinality, print formats, export formats, or print customization requires an explicit V4.4 scope amendment or a later revision.
## Related Local References
- [V4.4 Implementation Plan](implementation_plan_v4_4.md)
- [V4.3 Scope Boundary](../ver4.3/scope_boundary_v4_3.md)
- [V4.3 Implementation Plan](../ver4.3/implementation_plan_v4_3.md)
- [V4 Architecture](../ver4/architecture_v4.md)
- [V4 Schema](../ver4/schema_v4.md)
- [V4 Requirements](../ver4/requirements_v4.md)
-263
View File
@@ -1,263 +0,0 @@
# Implementation Plan (Version 4.5)
## Goal
Normalize metadata-directed image orientation for provider input, improve transcription-medium instructions and deterministic quality warnings, and support user-initiated single-Source retranscription with approved alternate models and explicit candidate promotion.
## Planning Status
- V4.4 is the completed implementation baseline.
- The V4.5 scope is frozen and sufficiently detailed to begin implementation.
- Scope additions require an explicit amendment or a later revision.
## Planning Constraints
- Original uploaded Source files remain immutable.
- Provider input must remain traceable to the original Source and any normalized derivative.
- Normalization is limited to recognized orientation metadata; no enhancement pipeline is introduced.
- Quality warnings never mutate transcription text or trigger paid requests automatically.
- Retranscription creates new immutable Job and execution evidence.
- A retranscription result remains a candidate until explicitly promoted.
- Human revision remains separate from and takes precedence over machine selection.
- Provider work occurs outside database transactions.
- Database, integration, and UI tests use confirmed isolated data and never modify `data/transcription.db`.
- Potentially destructive tests run only through `tools/run_destructive_tests.py`.
## Expected Project Impact
| Area | Expected impact |
| --- | --- |
| Configuration | Add a validated provider-model allowlist while retaining one default model. |
| Image processing | Add metadata-directed orientation normalization and model-input artifact creation. |
| Evidence model | Link each request to the exact model-input artifact and record preferred machine-attempt provenance. |
| Prompt contract | Add explicit body-medium classification and structured-layout rules. |
| Quality service | Add deterministic, non-mutating warnings for known output defects. |
| Job creation | Support a Source-locked retranscription Job and approved model selection. |
| Worker workflows | Preserve candidates without automatically replacing preferred machine text. |
| Source Detail | Add Retranscribe Source, candidate summaries, comparison, promotion, and normalization indicators. |
| Tests and documentation | Add isolated normalization, configuration, warning, retranscription, promotion, and UI coverage. |
## Implementation Phases
### 1. Align Configuration Contracts
- Add `provider_models` as an immutable validated collection in Settings.
- Parse `PROVIDER_MODELS` using the standard Pydantic-settings JSON representation.
- Preserve `PROVIDER_MODEL` as the default.
- If the allowlist is omitted, derive a one-entry list from the default.
- Normalize whitespace, reject empty values, and deduplicate while preserving the relative order of non-default values.
- Ensure the default appears exactly once and first in the effective selector order.
- Validate a submitted model against the allowlist in the service or workflow boundary, not only in the UI.
- Document `.env.example` behavior without adding real credentials.
- Add configuration tests for omitted, valid, duplicate, malformed, and empty model lists.
### 2. Define Orientation-Normalized Artifacts
- Reuse `ProcessingArtifact` for the model-input derivative and transformation metadata.
- Define a versioned orientation-normalization artifact schema containing:
- Original Source identity and digest.
- Original orientation value.
- Applied rotation.
- Original and derivative dimensions, media types, byte sizes, and digests.
- Processor name and version.
- Add one domain-owned orientation-normalization service or adapter; keep image-library details out of UI and provider modules.
- Apply recognized metadata orientation physically to raster pixels.
- Reset or remove orientation metadata on the derivative.
- Store derivatives under application-managed artifact storage with safe relative references.
- Avoid creating a derivative when no supported transformation is required.
- Return a typed provider-input reference that identifies whether the request uses the original or a derivative.
- Never mutate or delete the original Source as part of normalization.
### 3. Integrate Normalization with Provider Input
- Resolve the exact model input before constructing the request manifest.
- Use the normalized derivative when orientation metadata requires it; otherwise use the original Source.
- Extend the existing `SourceEvidenceReference` or associated artifact reference so the manifest identifies:
- Original Source.
- Derivative artifact when present.
- Transformation schema and digest.
- Ensure normalization completes and persists before provider network work begins.
- Bind the immutable model-input artifact reference to the exact ExecutionAttempt that consumed it before terminal attempt persistence completes.
- If normalized bytes are reused, retain attempt-specific association while preserving one content identity and digest.
- If normalization fails, do not send a provider request.
- Preserve safe failure evidence and an actionable error category.
- Confirm provider payload loading and evidence hashing read the same resolved bytes.
### 4. Revise the Transcription Prompt
- Align the prompt with the durable medium rules in `docs/invariant/transcription_methodology.md`.
- Add the four frozen document-body markers.
- Define operational differences among handwritten, typewritten, typeset, and mixed content.
- State that mechanical typewriter variation is not handwriting.
- Require exactly one body marker.
- Prohibit repeated whole-line handwriting wrappers after a whole-body handwritten marker.
- Permit localized handwriting markers only for actual annotations, signatures, or mixed-body portions.
- Add layout instructions for tables of contents, tables, forms, columns, captions, marginalia, page numbers, dotted leaders, and associated references.
- Require plain-text characters rather than HTML entities.
- Retain verbatim, uncertainty, damage, deletion, insertion, and line-break-hyphenation rules.
- Update prompt fixtures and prompt-hash expectations without rewriting historical Job prompt evidence.
### 5. Add Deterministic Quality Analysis
- Introduce a small typed warning model with stable warning codes and human-readable detail.
- Analyze successful output without modifying it.
- Implement warnings for:
- Unicode replacement characters.
- Multiple document-body markers.
- Whole-body handwritten plus repeated line-level handwriting wrappers.
- Likely unresolved HTML entities.
- Keep warning rules deterministic and provider-independent.
- Persist warnings once as an immutable, versioned `ProcessingArtifact` tied to the successful ExecutionAttempt.
- Source Detail renders stored warnings and never recomputes historical attempts under newer warning rules.
- Make warning analysis idempotent and versioned so newer rules apply only to newly analyzed attempts unless a separate future reanalysis workflow is introduced.
- Do not implement confidence scoring or automatic retries.
### 6. Model Retranscription and Selection State
- Add a durable Job purpose or equivalent discriminator for normal transcription versus Source retranscription.
- Ensure a retranscription Job contains exactly one JobSource for the locked Source.
- Add durable preferred-machine-attempt provenance for each Source.
- Retain `Source.raw_transcription` as the preferred-machine-text projection for compatibility.
- Define legacy behavior for Sources whose current projection predates ExecutionAttempt provenance.
- On the first successful result with no preferred machine output, select the successful attempt automatically regardless of Job purpose.
- Once preferred provenance exists, preserve every later successful result as an unselected candidate regardless of Job purpose.
- Prevent normal and retranscription workflows from writing `Source.raw_transcription` directly when preferred provenance already exists.
- Add a promotion command that:
- Loads the Source and successful ExecutionAttempt.
- Verifies ownership and successful text.
- Updates preferred-attempt provenance and `Source.raw_transcription` in one transaction.
- Leaves `Source.revised_text` and all attempts unchanged.
- Reject failed, unrelated, missing, or textless candidates deterministically.
### 7. Add the Retranscribe Source Workflow
- Add a Source Detail **Retranscribe Source** action.
- Navigate to Create Processing Job with an explicit `source_id` query parameter.
- Load the Source and derive its Document server-side.
- Render Source identity and filename as locked context.
- Render Provider from Settings as read-only.
- Render Model as a selector using the effective allowlist and default.
- Reuse the frozen default prompt unless a later scope addition explicitly allows prompt selection.
- Create a new queued retranscription Job and one pending JobSource atomically.
- Notify the worker only after the transaction commits.
- Preserve current preferred machine text and human revision throughout queueing, processing, success, and failure.
- Return to the new Job Detail after successful creation.
### 8. Adapt Worker Success Semantics
- Select the first successful result automatically for Sources without preferred machine output, including a successful retranscription after earlier failures.
- For every later success, persist JobSource and ExecutionAttempt text without replacing the Source projection, regardless of normal or retranscription Job purpose.
- Run deterministic quality analysis after successful normalization of provider output.
- Persist candidate warnings with the attempt.
- Keep all terminal Job status and execution evidence updates atomic according to the existing workflow boundary.
- Ensure a failed retranscription cannot clear or change preferred machine text.
### 9. Build Candidate Review and Promotion UI
- Extend Source Detail with concise machine-output sections:
- Preferred machine transcription.
- Human revision.
- Candidate machine transcriptions.
- List candidates with creation date, provider, model, Job ID, status, and warning indicator.
- Default to a compact candidate list; do not render every full transcript simultaneously.
- When no successful machine result exists, render an explicit empty state without comparison controls.
- When preferred output exists without candidates, render an explicit no-candidates state.
- Allow one candidate to be opened for comparison with the preferred machine transcription.
- Label both sides with provider, model, Job ID, and date.
- Add **Use this transcription** only for a successful unselected candidate.
- Require explicit confirmation before promotion.
- After promotion, refresh Source Detail and retain the former preferred result in attempt history.
- If a human revision exists, explain that promotion changes machine selection but not the human-preferred displayed/printed text.
- When orientation normalization occurred, show a compact indicator and link to transformation evidence; do not require routine side-by-side image display.
### 10. Align API and Service Contracts
- Keep arbitrary provider and model identifiers out of public write contracts.
- If an API is added for retranscription, accept Source UUID and one configured model identifier and validate both server-side.
- If an API is added for promotion, accept Source UUID and ExecutionAttempt UUID and validate their relationship.
- Return stable validation, conflict, not-found, provider, and persistence errors through the existing taxonomy.
- Keep filesystem paths, raw credentials, and unrestricted artifact references out of responses.
### 11. Verification
- Add pure unit tests for:
- Orientation metadata interpretation.
- No-op versus transformed input selection.
- Derivative metadata and hashing.
- Model allowlist normalization.
- Prompt marker rules.
- Every quality-warning code.
- Add isolated service and workflow tests for:
- Original Source immutability.
- Normalization failure before provider invocation.
- Request manifests referencing exact provider-input bytes.
- Single-Source retranscription Job creation.
- Candidate preservation.
- Initial automatic selection.
- Explicit candidate promotion.
- Atomic rollback on invalid promotion.
- Human revision preservation.
- Add UI tests for:
- Retranscribe Source navigation.
- Locked Source and Document context.
- Read-only Provider and allowlisted Model selector.
- Candidate list, warnings, comparison, confirmation, and promotion.
- Normalization indicator without mandatory image comparison.
- Use fixed image fixtures with known EXIF orientation and hashes.
- Use fake providers only; no focused or regression test sends an external provider request.
- Confirm all database, integration, and UI targets use isolated databases.
- Run potentially destructive schema tests only through `tools/run_destructive_tests.py`.
- Run the broader non-external regression suite after focused coverage passes.
- Verify that `data/transcription.db` and curated Source files were not changed by test execution.
### 12. Align Authoritative Documentation
- Update V4 architecture, requirements, and schema for:
- Original versus model-input artifacts.
- Preferred machine-attempt provenance.
- Retranscription Job purpose.
- Candidate and promotion semantics.
- Quality-warning evidence.
- Update Source and Job UI contracts after implementation behavior is accepted.
- Update the transcription methodology and evidence invariant only where V4.5 establishes a durable cross-version rule.
- Keep historical prompt and execution evidence immutable.
## Delivery Order
1. Configuration and model allowlist.
2. Orientation artifact schema and normalization adapter.
3. Provider-input and evidence integration.
4. Prompt revision and deterministic warning analysis.
5. Retranscription and preferred-attempt persistence.
6. Worker candidate semantics.
7. Retranscription creation UI.
8. Candidate comparison and promotion UI.
9. Authoritative documentation and regression verification.
## Done Criteria
- All frozen V4.5 acceptance criteria are implemented and testable.
- EXIF-oriented images are physically upright for provider processing while originals remain byte-for-byte unchanged.
- Every provider request identifies its exact original and normalized inputs.
- The prompt classifies body medium consistently and avoids redundant handwriting wrappers.
- Quality defects produce deterministic warnings without silent rewriting or automatic cost.
- Source Detail can create one-Source retranscription Jobs using configured alternate models.
- Retranscription results remain candidates until explicitly promoted.
- Promotion updates exact preferred-attempt provenance and the compatibility projection atomically.
- Human revisions remain unchanged and retain display/print precedence.
- Previous Jobs and attempts remain immutable and inspectable.
- No manual image editor, visual orientation inference, confidence percentage, automatic retry, or arbitrary model entry is introduced.
- Verification uses isolated data, fake providers, and does not modify `data/transcription.db` or curated Source files.
## Related Local References
- [V4.5 Scope Boundary](scope_boundary_v4_5.md)
- [V4.4 Scope Boundary](../ver4.4/scope_boundary_v4_4.md)
- [V4.4 Implementation Plan](../ver4.4/implementation_plan_v4_4.md)
- [V4.2 Evidence and Provenance Scope](../ver4.2/scope_boundary_v4_2.md)
- [V4 Architecture](../ver4/architecture_v4.md)
- [V4 Schema](../ver4/schema_v4.md)
- [V4 Requirements](../ver4/requirements_v4.md)
- [V4 Error Handling Policy](../ver4/error_handling_v4.md)
- [Transcription Methodology](../invariant/transcription_methodology.md)
- [AI Evidence and Provenance Invariant](../invariant/ai_evidence_and_provenance.md)
-226
View File
@@ -1,226 +0,0 @@
# V4.5 Scope Boundary
This document defines the frozen boundary for transcription input normalization and selective transcription-quality improvement after the completed V4.4 revision. V4 through V4.4 remain the architecture and behavioral baseline except where this document explicitly changes Source processing, transcription selection, or Source Detail behavior.
## Purpose
- Ensure model inputs are physically upright when curated image files rely on orientation metadata.
- Distinguish typewritten, typeset, handwritten, and mixed document bodies consistently.
- Let the user selectively retranscribe an unsatisfactory Source with an approved alternate vision model.
- Preserve every machine result while allowing the user to choose which result is the preferred machine transcription.
## In Scope
### 1. Metadata-Driven Orientation Normalization
- The original uploaded Source remains immutable archival evidence.
- Before a supported raster image is sent to a transcription provider, the application reads recognized orientation metadata.
- When the metadata requires rotation, the application creates a physically upright model-input derivative and resets or removes the derivative's orientation metadata.
- The provider receives the normalized derivative rather than upside-down stored pixels.
- When no supported orientation transformation is required, the original Source may remain the provider input.
- The derivative records:
- Source UUID.
- Original and derivative SHA-256 digests and byte sizes.
- Original and derivative dimensions and media types.
- Applied orientation transformation.
- Transformation implementation and version.
- Creation timestamp.
- The derivative uses the existing processing-artifact and evidence architecture rather than replacing the Source file.
- PDF orientation, visual orientation inference, manual rotation controls, deskewing, cropping, contrast changes, and general image enhancement are not part of V4.5.
### 2. Transcription Medium Contract
- The durable document-medium rules are defined in the cross-version [Transcription Methodology](../invariant/transcription_methodology.md).
- The transcription prompt distinguishes these document-body media:
- `[document body handwritten]`
- `[document body typewritten]`
- `[document body typeset]`
- `[document body mixed]`
- A typewriter's uneven impressions, monospaced characters, or mechanical defects do not by themselves indicate handwriting.
- The transcript contains exactly one applicable document-body marker.
- A wholly handwritten body uses the one body marker rather than wrapping every line in `[handwritten: ...]`.
- Typewritten and typeset bodies do not use handwriting wrappers unless a genuinely handwritten annotation or signature appears.
- A mixed body may use localized handwriting markers only for the handwritten portions.
- Stored transcription output remains plain text. Model-generated HTML entities are not required for ordinary characters.
- The prompt includes layout guidance for tables of contents, tables, forms, columns, captions, marginalia, page numbers, dotted leaders, and associated page references.
- Line-break hyphenation rules continue to preserve intentional hyphens while rejoining words split only by line wrapping.
### 3. Quality Warnings
- The application evaluates successful machine output for deterministic warning conditions, including:
- Unicode replacement characters such as ``.
- A whole-body handwritten marker combined with repeated whole-line handwriting wrappers.
- More than one document-body medium marker.
- Unresolved HTML entities in otherwise plain transcription text.
- Warnings do not silently rewrite model output.
- Warnings do not automatically trigger another paid provider request.
- Source Detail displays warnings with the relevant machine result so the user can decide whether to revise or retranscribe it.
- V4.5 does not assign or display a transcription-confidence percentage. Provider self-assessments and token probabilities are not treated as calibrated transcription confidence.
### 4. Source Retranscription Entry Point
- Source Detail adds a **Retranscribe Source** action.
- The action opens Create Processing Job with the Source preselected and locked.
- The Source's existing Document is derived from its relationship and cannot be changed in this flow.
- The new Job contains a `JobSource` only for the selected Source; it does not retranscribe every Source in the Document.
- Provider is populated from the configured `PROVIDER` value and is read-only while only one provider is configured.
- Model is selected from an operator-configured allowlist.
- The configured default model is initially selected.
- Creating the Job freezes the selected provider, model, prompt, prompt hash, parameters, and Source evidence according to the existing provenance contract.
- Retranscription creates a new Job and new execution evidence. It does not reuse, mutate, or erase a previous Job.
### 5. Configured Vision-Model Allowlist
- `PROVIDER_MODEL` remains the default transcription model.
- `PROVIDER_MODELS` defines the models available in the Create Processing Job model selector.
- The environment representation is a JSON array, for example:
```dotenv
PROVIDER=openrouter
PROVIDER_MODEL=google/gemini-2.5-flash
PROVIDER_MODELS=["google/gemini-2.5-flash","google/gemini-2.5-pro","anthropic/claude-sonnet-4"]
```
- If `PROVIDER_MODELS` is omitted, the selector contains only `PROVIDER_MODEL`.
- The default model appears exactly once and first; remaining configured models retain their relative order.
- Empty, malformed, or duplicate values produce deterministic configuration validation.
- The UI never accepts an arbitrary model identifier outside the configured allowlist.
- The allowlist controls availability, not claims of quality, price, or provider compatibility. The operator is responsible for configuring models supported by the selected provider.
### 6. Candidate Machine Transcriptions
- The first successful transcription becomes preferred automatically whenever the Source has no preferred machine output, regardless of whether it came from an initial or retranscription Job.
- Once a Source has a preferred machine transcription, every later successful result is stored as a candidate regardless of Job purpose and cannot replace the preferred result automatically.
- Every candidate remains associated with its immutable Job, JobSource, ExecutionAttempt, provider, model, prompt, parameters, timestamps, warnings, and normalized-input evidence.
- Source Detail presents:
- The current preferred machine transcription.
- Available successful candidates with date, provider, model, Job ID, and warning state.
- A comparison between the current preferred machine transcription and one selected candidate.
- A **Use this transcription** action for a successful candidate.
- Before any successful result exists, Source Detail displays an explicit no-machine-transcription state.
- When a preferred result exists but no candidates exist, Source Detail omits comparison controls and displays an explicit no-candidates state.
- Promoting a candidate:
- Verifies that the successful execution belongs to the Source.
- Records the selected execution as the preferred machine-output provenance.
- Updates `Source.raw_transcription` as the preferred-machine-text projection.
- Does not modify `Source.revised_text`.
- Does not delete or alter any previous machine result.
- If a human revision exists, it remains the current human-preferred text used by normal display and printing after a machine candidate is promoted.
### 7. Evidence and Failure Behavior
- The original Source and every normalized derivative are content-addressed and traceable.
- Every provider request identifies the exact original Source and model-input artifact used.
- Every execution attempt records the exact immutable model-input artifact it consumed, including when normalized bytes are reused.
- Every retranscription attempt follows the V4.2 immutable execution-attempt contract.
- A normalization failure prevents the provider request and produces an explicit actionable error.
- A provider or persistence failure leaves the current preferred machine transcription and human revision unchanged.
- Candidate promotion is atomic: provenance selection and the preferred-machine-text projection either both commit or both remain unchanged.
- Provider network work occurs outside database transactions.
### 8. Source Detail Terminology
- **Original Source** means the immutable uploaded file.
- **Model input** means the original Source or normalized derivative actually sent to the provider.
- **Machine attempt** means one immutable provider execution.
- **Candidate transcription** means a successful machine result not currently selected as preferred.
- **Preferred machine transcription** means the selected machine result projected through `Source.raw_transcription`.
- **Human revision** means `Source.revised_text`, which remains independent of every machine result.
- These distinctions use concise labels and progressive disclosure; routine satisfactory Sources do not display a mandatory side-by-side original/normalized image comparison.
- When normalization occurred, Source Detail displays an orientation-normalized indicator and makes transformation evidence inspectable through the existing evidence UI.
## Out of Scope
- Manual image rotation or image-editing controls.
- Visual orientation detection when metadata is absent or incorrect.
- Deskewing, cropping, contrast normalization, denoising, sharpening, or restoration.
- Replacing or modifying original Source files.
- Automatically retranscribing every Source.
- Automatic provider retries triggered by quality warnings.
- Arbitrary model identifiers entered by users.
- Multiple provider selection in the UI.
- Model benchmarking, pricing recommendations, or automatic model ranking.
- A provider-independent transcription-confidence percentage.
- Silent cleanup or rewriting of model output.
- Deleting unsuccessful, superseded, or unselected machine attempts.
- Promoting a machine candidate over a human revision.
- Source page renumbering or reordering.
## Locked Design Decisions
### A. Curated Originals Remain Authoritative
- V4.5 corrects metadata-directed orientation only for model processing.
- The archival upload is never replaced by the normalized derivative.
### B. Orientation Is Automatic and Metadata-Driven
- No manual orientation workflow is introduced.
- V4.5 does not guess orientation from page content.
### C. Selective Retranscription Replaces Automatic Escalation
- The normal default model remains efficient for satisfactory Sources.
- The user explicitly chooses when an alternate approved model is worth another provider request.
- Quality warnings inform that choice but never incur cost automatically.
### D. Retranscription Produces Candidates
- Alternate results remain immutable and comparable.
- The user explicitly promotes the preferred machine result.
- Human revision remains a separate, higher-precedence layer.
### E. Model Choice Is Operator-Controlled
- Environment configuration defines the finite allowed model set.
- Job records freeze the actual selected model and request parameters.
### F. Confidence Is Evidence-Based, Not Invented
- V4.5 does not present model self-rating as objective confidence.
- Review uses visible output, deterministic warnings, provenance, and human judgment.
## Acceptance Criteria
1. A JPEG with EXIF Orientation 3 produces an upright model-input derivative while the original bytes remain unchanged.
2. A Source requiring no recognized orientation transformation is not unnecessarily altered.
3. Orientation transformation metadata and hashes identify the exact provider input.
4. The prompt distinguishes handwritten, typewritten, typeset, and mixed bodies with exactly one body marker.
5. Typewritten text is not wrapped line by line as handwriting.
6. Tables of contents retain row associations and page references without handwriting wrappers.
7. Deterministic warnings identify replacement characters and contradictory body markers without changing output.
8. Source Detail provides Retranscribe Source for an existing Source.
9. Create Processing Job locks the Source and Document, uses configured Provider, and restricts Model to the configured allowlist.
10. Retranscription creates a new single-Source Job with complete frozen request and execution evidence.
11. The first successful result is selected automatically; every later successful result remains a candidate and cannot replace the preferred machine transcription automatically.
12. Source Detail can compare the preferred machine transcription with one candidate and promote that candidate explicitly.
13. Candidate promotion records exact successful-execution provenance and updates the machine-text projection atomically.
14. Candidate promotion never changes or clears a human revision.
15. Earlier machine attempts remain inspectable after retranscription and promotion.
16. No confidence percentage, manual image editor, automatic quality retry, or arbitrary model input is introduced.
17. Database, integration, and UI verification uses isolated test data and never modifies `data/transcription.db`.
## Scope Freeze Gate
V4.5 is sufficiently frozen to begin implementation:
- Orientation behavior and preservation rules are resolved.
- Prompt medium categories and marker behavior are resolved.
- Warning behavior and the absence of automatic retry are resolved.
- Retranscription entry point, single-Source scope, and model configuration are resolved.
- Candidate preservation, comparison, promotion, and human-revision precedence are resolved.
- The broader future-feature list has been reviewed, and no additional V4.5 features are required.
Any expansion into manual image editing, visual orientation inference, automatic retries, multiple providers, confidence scoring, or additional processing features requires an explicit V4.5 scope amendment or a later revision.
## Related Local References
- [V4.5 Implementation Plan](implementation_plan_v4_5.md)
- [V4.4 Scope Boundary](../ver4.4/scope_boundary_v4_4.md)
- [V4.4 Implementation Plan](../ver4.4/implementation_plan_v4_4.md)
- [V4.2 Evidence and Provenance Scope](../ver4.2/scope_boundary_v4_2.md)
- [V4 Architecture](../ver4/architecture_v4.md)
- [V4 Schema](../ver4/schema_v4.md)
- [V4 Requirements](../ver4/requirements_v4.md)
- [Transcription Methodology](../invariant/transcription_methodology.md)
- [AI Evidence and Provenance Invariant](../invariant/ai_evidence_and_provenance.md)
-295
View File
@@ -1,295 +0,0 @@
# Implementation Plan (Version 4.6)
## Goal
Pay down the defects, duplication, and structural drift identified in the [Architecture & Code Review Report](../architecture_code_review_2026-08-17.md) without changing any observable behavior. Re-level the database schema from current SQLModel metadata, correct read amplification and missing indexes, consolidate duplicated service and UI code, restore the project's own documented boundaries, and make `ty` a real quality gate.
## Planning Status
- V4.5 is the completed implementation baseline.
- The V4.6 scope is frozen and sufficiently detailed to begin implementation.
- Every task traces to a review finding ID. A change without a finding ID is a scope addition and requires an explicit amendment.
- The `SourceService` split ([MED-14]) is deferred to V4.7 by decision, not by omission.
## Planning Constraints
- **Behavior is preserved exactly.** All 264 pre-existing tests must still pass. A test that must change is evidence the change is not remediation.
- The application targets **SQLite only** in V4.6. PostgreSQL is unblocked but not enabled.
- The deployment is **single user, single process, single worker**. Forward-compatible code is written where cheap and dialect-guarded.
- The schema is re-leveled from metadata. **No Alembic, no revision directory, no history table, no down path.**
- All schema-affecting changes land in **one pass**; partial application is not a valid state.
- The data migration script is authored **last**, against the final schema and final loading strategy.
- Uploaded Source files, portraits, and artifact files on disk are never modified.
- Original Source files remain immutable; every V4.2V4.5 evidence and provenance contract is preserved.
- Provider network work continues to occur outside database transactions.
- Database, integration, and UI tests use confirmed isolated data and never modify `data/transcription.db`.
- Potentially destructive tests run only through `tools/run_destructive_tests.py`.
## Expected Project Impact
| Area | Expected impact |
| --- | --- |
| Dead code | Remove `app_state.py`, `services/transcription.py`, legacy aliases, `ServiceBase.queue`, and a duplicate queued-job query. |
| Persistence | Delete the hand-rolled DDL chain; generate schema from metadata with correct indexes, FK ordering, and loading strategy. |
| Query behavior | Bounded queued-job poll, SQL-side filtering, bounded navigation queries, explicit eager loads. |
| Worker | Provider client and service bundle live for the worker's lifetime rather than per job. |
| Configuration | Remove the timeout cap and the silently-ignored `DATABASE_URL`; resolve dead settings; replace frozen-model mutation. |
| Service layer | Generic registry service, shared not-found guard, single media-storage implementation (~400 lines removed). |
| UI layer | Fix three boundary violations; extract duplicated components (~500 lines removed); externalize the SVG asset. |
| Async I/O | Move filesystem, hashing, and image work off the event loop. |
| Tooling | `ruff` and `ty` both reach zero and gate on pre-commit. |
| Data | One-time migration of backed-up V4.5 data into the re-leveled schema. |
| Tests and documentation | Add index, FK-cycle, claim-boundedness, client-reuse, and registry-parity coverage; correct the stale instruction path. |
## Implementation Phases
### 1. Deletions and Quick Wins
Independent of every other phase. Land first to shrink the surface everything else must consider.
- Delete `src/transcription/app_state.py` and confirm zero importers remain in `src`, `tests`, and `tools` ([HIGH-01]).
- Delete `src/transcription/services/transcription.py` and standardize every `build_prompt_execution` import on `services/sources.py` ([MED-05]).
- Delete the legacy compatibility aliases in `services/store.py:35,382,383` ([MED-05]).
- Delete `ServiceBase.queue` and its unparameterized `asyncio.Queue` ([MED-07]).
- Delete `db/operations.py:get_next_queued_job` as a divergent duplicate of the live implementation ([CRIT-01]).
- Resolve `sqlite_check_same_thread` and `worker_retry_backoff_seconds`: wire each to real behavior or delete it together with its test ([MED-02]).
- Remove `DATABASE_URL` from `docker-compose.yml` and document the real `DATABASE__DRIVER` / `DATABASE__PATH` nested names in `.env.example` ([MED-10]).
- Correct the stale path in `.github/instructions/services.instructions.md:10` to `src/transcription/db/models.py` ([LOW-02]).
- Remove the discarded `load_docs` parameter from `list_jobs` ([LOW-03]).
- Validate the `getattr` result in `resolve_worker_notifier` ([LOW-04]).
- Move `VIBESCRIBE_LOGO_SVG` to `ui/static/vibescribe_logo.svg` and load it through a `read_svg` sibling of `ui/resources.py:read_css` ([MED-09]).
- Run `ruff check --fix` and resolve the remainder by hand ([LOW-01]).
- Route `people_page.py:504` through `error_presenter.show_error` ([LOW-07]).
- Cancel the auto-refresh timer rather than only deactivating it, and name its interval constant ([LOW-06]).
**Verification:** full suite green, `ruff check` reports zero, no import of a deleted symbol remains.
### 2. Schema Re-Level — Single Pass
This phase is atomic. Every task below regenerates the same schema and must be verified together.
- Delete `upgrade_schema` and `_upgrade_*` (`db/operations.py:25-109`) and their tests (`tests/test_db.py:109-172`) ([HIGH-05]).
- Confirm `create_all()` remains gated by `Settings.should_bootstrap_schema` (`config.py:140-145`) ([HIGH-05]).
- Declare the composite index in the model: `Index("ix_job_status_date_created", "status", "date_created")`, plus `index=True` on the foreign keys the worker and detail pages filter on ([HIGH-04]).
- Declare `Source.preferred_execution_attempt_id`'s foreign key with `use_alter=True` and an explicit constraint name, breaking the `source` / `job_source` / `execution_attempt` cycle ([HIGH-08]).
- Flip relationship loading from bidirectional `lazy="selectin"` to `lazy="raise"`, model by model ([CRIT-02]):
- Work one model at a time with the suite as the safety net.
- Where a test fails with a lazy-load error, add an explicit `selectinload()` to the *service query* that feeds it — never restore the model-level default.
- Where an existing explicit `selectinload()` proves redundant, delete it; this is the primary source of the ~160 `ty` diagnostics addressed in Phase 6.
- Follow the two correct precedents already in the codebase: `Source.processing_artifacts:279` and `JobSource.execution_attempts:336`.
- Rebuild the development database from empty. Do **not** attempt to upgrade the existing file.
**Verification:**
- A test asserts the composite `Job` index and the hot foreign-key indexes exist in a freshly created schema.
- A test compiles the metadata against the PostgreSQL dialect and asserts **no** unresolvable-cycle warning is emitted.
- A test asserts `preferred_execution_attempt_id`'s column type matches the model declaration.
- Full suite green under `lazy="raise"`.
- No raw `ALTER TABLE` or `CREATE INDEX` string remains anywhere in `src`.
**Rollback:** this phase reverts as a unit. A partially applied schema pass is not a valid state.
### 3. Worker and Provider Reliability
Depends on Phase 2, because the claim query's cost profile is only correct once eager-loading defaults are fixed.
- Add `.limit(1)` to the queued-job selection and remove its eager-load options from the hot poll ([CRIT-01]).
- Convert the read-then-write claim into an atomic `QUEUED``PROCESSING` transition in one transaction ([CRIT-01]):
- Write the dialect-guarded `with_for_update(skip_locked=True)` branch for the multi-user direction.
- On SQLite, the claim executes as a bounded single-writer transaction.
- Load the eager relationships in a **second** query after the claim succeeds.
- Update the stale comment at `workflows.py:193-194` to describe the actual guarantee rather than the known hazard.
- Hoist `ServiceBundle` and the provider client out of the per-job body in `worker.py:157-174` to worker-loop scope; `aclose()` the client once at loop shutdown, not once per job ([HIGH-02]).
- Add `ServiceBundle.from_session_factory(...)`, replacing the three duplicated instantiation blocks at `app.py:45-50`, `worker.py:160-165`, and `services/__init__.py:19-22`. Have `_recover_stale_processing_jobs` (`app.py:73-84`) use the bundle built five lines earlier ([MED-06]).
- Remove `le=20.0` from `worker_provider_timeout_seconds` (`config.py:110`), raise the default to a realistic vision-transcription duration, and pass an explicit `httpx.Timeout` to the OpenRouter `AsyncClient` (`openrouter.py:198`) ([HIGH-03]).
- Extend the `TranscriptionProvider` Protocol to declare `aclose` and the evidence attributes; delete the per-call `inspect.signature(adapter.transcribe).parameters` reflection at `sources.py:1237` and the associated untyped kwargs dict ([MED-03]).
**Verification:**
- A test asserts the emitted claim SQL contains `LIMIT` and no `selectinload` join.
- A test asserts the worker processes two consecutive jobs against the same provider client instance.
- A test asserts a timeout value above 20 seconds is accepted by `Settings`.
- A test asserts the transcription call path resolves `requested_model` through the Protocol without reflection.
### 4. Service Layer Consolidation
Depends on Phase 2 only for the loading strategy; otherwise independent of Phase 3.
- Introduce `services/registry.py` with a generic `RegistryService[ModelT]` owning list, summaries with counts, create with `IntegrityError` → conflict mapping, read with not-found, update, delete with built-in and referenced guards, and `is_referenced` ([MED-11]):
- Define label normalization, the casefold key, and the summary shape once.
- Reduce `DocumentService`'s document-type methods (`documents.py:350-500`) and `PeopleService`'s person-role methods (`people.py:214-378`) to subclasses declaring model, error class, reference query, and noun.
- Preserve every existing user-facing message, error category, and suggestion string verbatim; template the noun only.
- Add `ServiceBase._get_or_raise(...)` and adopt it at all 38 not-found sites, including `documents.py:174,210,291`, which currently bypass the local `_get_document_or_raise` helper. Delete the now-redundant local helper ([MED-12]).
- Introduce `services/media_storage.py` as the single validate → hash → `mkdir` → write → wrap-`OSError` implementation, replacing `store.py:319-379`, `people.py:596-631`, and `ui/homepage_store.py:31-44`. Wrap the write in `asyncio.to_thread` ([MED-13], [MED-01]).
- Move `source_mime_type` out of `services/sources.py` into a shared module so `documents.py:24` no longer imports a sibling service, restoring the independence rule at `services.instructions.md:13` ([MED-14], partial).
- Correct the four query inefficiencies in `sources.py` ([LOW-08]):
- `list_sources_detail:338-343` — move the `job_id` filter from Python into a SQL join on `JobSource`.
- `read_source_navigation:233-244` — replace the full ordered-id scan with two `LIMIT 1` queries.
- `list_processing_artifacts:961` — add a `limit` parameter matching its summary sibling.
- `build_evidence_export:1012-1013` — move artifact integrity hashing into `asyncio.to_thread`.
**Verification:**
- Existing `DocumentType` and `PersonRole` tests pass **unchanged** against the shared implementation. This is the primary proof that behavior is preserved.
- A test asserts `list_sources_detail` filtered by `job_id` emits a join rather than loading the full table.
- No module in `services/` imports another concrete service module.
### 5. UI Boundaries and Duplication
Independent of Phases 24 except where a service signature changes.
- Fix the three `ui.instructions.md` violations ([HIGH-07]):
- Add a `JobService` or workflow method that owns `session_scope` internally; remove the import and transaction management from `jobs_page.py:17,185-192`.
- Have `SourceService` return a plain `transport_body_deferred: bool` on a read model; remove `sqlalchemy.inspect` from `sources_page.py:13,439`.
- Pass a ready media URL into `document_panzoom`, or delete the component — it is exported from `components/__init__.py` but used by no page ([HIGH-07]).
- Extract the duplication catalogued in review §4, highest value first:
- `ui/components/confirm_delete.py` — the blocked-deps card plus confirm/cancel row, from four pages (~120 lines).
- `ui/components/media_urls.py` — pure upload-URL resolution taking `upload_dir` and `base_url`, from three call sites (~110 lines).
- `ui/components/guards.py` — parse → error label → return, from nine call sites (~90 lines).
- `build_table` adoption for the remaining hand-rolled `ui.table` instances, adding selection and no-search options as needed (~70 lines).
- `ui/components/upload_panel.py` — file-picker wiring, from three pages (~50 lines).
- `ui/components/formatters.py``_parse_uuid` (five copies) and `_parse_iso_date` (two copies) (~49 lines).
- A shared page-helper for `_resolve_runtime_settings(request)` (three copies, ~18 lines).
- Annotate untyped handler parameters and replace loosely-typed dict returns with read models ([LOW-05]).
**Verification:** UI page tests pass unchanged; no page module imports `session_scope`, `sqlalchemy.inspect`, or `get_settings`.
### 6. Async I/O and Configuration Hygiene
- Wrap the remaining blocking work in `asyncio.to_thread`: Pillow orientation normalization, artifact writes, and evidence hashing not already covered by Phase 4 ([MED-01]).
- Replace `functools.cache` on the engine and session factories with an explicit URL-keyed registry supporting targeted eviction, removing the cross-test and cross-tenant coupling and restoring a visible call signature ([MED-04]).
- Replace `object.__setattr__` in `normalize_provider_models` (`config.py:130,137`) with `model_copy(update=...)` or a computed property.
- Add `onupdate` to the `updated_at` / `date_updated` columns that are expected to track modification, so they stop being stale on the update paths that do not set them by hand. Remove the now-redundant manual assignment at `jobs.py:166` and its siblings.
- Surface the exception currently swallowed to `None` in the ORM model property at `models.py:227` ([MED-08]).
**Note:** the `onupdate` change is schema-affecting in principle but not in emitted DDL, since `onupdate` is a Python-side default. If implementation reveals it alters generated DDL, it moves into Phase 2 and Phase 2 is re-verified.
**Verification:** a test asserts an update through a service advances `updated_at`; a test asserts two different database URLs produce two distinct engines and that evicting one leaves the other intact.
### 7. Type Checking and Tooling Gate
Depends on Phase 2, which is expected to remove most diagnostics by deleting redundant eager loads.
- Re-baseline `ty check` after Phase 2 and measure the remaining diagnostic count ([HIGH-06]).
- Convert every surviving `# pyright: ignore[...]` to `# ty: ignore[...]`, since `ty` does not honor pyright directives ([HIGH-06]).
- Fix the two real bugs currently hidden in the noise ([HIGH-06]):
- `tests/ui/test_sources_page.py:25` constructs `Source(...)` without the required `document_id`.
- `tools/run_destructive_tests.py:76,80` uses `fcntl`, which does not exist on Windows; use a cross-platform lock or guard by platform.
- Drive `ty check` to zero diagnostics and wire it into the existing pre-commit setup as a blocking gate.
- Configure `asyncio_default_fixture_loop_scope` explicitly so pytest-asyncio behavior does not change on upgrade.
**Verification:** `ty check` and `ruff check` both report zero; pre-commit fails when either regresses; `tools/run_destructive_tests.py` runs on Windows.
### 8. Data Migration
The final phase. Authored against the completed schema and the completed loading strategy.
- Write a one-time script under `tools/` that reads the backed-up V4.5 database and writes into the re-leveled schema (review §1a, "Items Added During Scoping").
- Because `lazy="raise"` is in force, every relationship traversal in the script carries an explicit eager load. This is the reason the script is written last.
- Preserve identity: UUIDs, digests, timestamps, attempt numbers, and `preferred_execution_attempt_id` selections carry across unchanged.
- Do not reinterpret, normalize, or regenerate any `ExecutionAttempt` or `ProcessingArtifact` evidence.
- Do not modify any on-disk Source file, portrait, or artifact file.
- The script is idempotent, is never invoked from application startup, and never runs in the test suite.
**Verification:** post-migration row counts match the backup for every table (`document` 8, `document_person` 11, `document_type` 7, `execution_attempt` 80, `job` 11, `job_source` 79, `person` 5, `person_role` 3, `processing_artifact` 2, `source` 76); artifact integrity verification passes for every migrated artifact; on-disk file hashes are unchanged.
## Sequencing Constraint
```mermaid
graph TD
P1[1. Deletions & Quick Wins]
P2[2. Schema Re-Level<br/>SINGLE ATOMIC PASS]
P3[3. Worker & Provider]
P4[4. Service Consolidation]
P5[5. UI Boundaries & Duplication]
P6[6. Async I/O & Config]
P7[7. Type-Check Gate]
P8[8. Data Migration]
P1 --> P2
P2 --> P3
P2 --> P4
P2 --> P7
P1 --> P5
P4 --> P5
P4 --> P6
P3 --> P8
P5 --> P8
P6 --> P8
P7 --> P8
```
The binding constraints are:
1. **Phase 2 is indivisible.** `create_all` from metadata, the indexes, `use_alter`, and the `lazy` flip all regenerate the same schema. They land together or not at all.
2. **Phase 7 follows Phase 2.** Measuring the `ty` baseline before the redundant eager loads are deleted would chase diagnostics that Phase 2 removes for free.
3. **Phase 8 is last.** The migration script must be written against the final schema and the final loading strategy.
## Test Strategy
- **The existing suite is the contract.** 264 tests pass today and must pass at every phase boundary. A test that requires modification is treated as a defect in that test, justified individually in the commit, and never as license to change behavior.
- **Registry parity is the key proof.** The `DocumentType` and `PersonRole` tests must pass *unchanged* against the shared `RegistryService`. If they need edits, the abstraction is wrong.
- **New tests are structural, not behavioral.** They assert schema shape, emitted SQL, dialect compatibility, and object lifetime — properties the current suite does not cover and that the review found were the reason these defects survived.
- New coverage to add:
| Assertion | Finding |
| :--- | :--- |
| Composite `Job` index and hot FK indexes exist in a fresh schema | [HIGH-04] |
| PostgreSQL-dialect metadata compilation emits no cycle warning | [HIGH-08] |
| `preferred_execution_attempt_id` column type matches the model | [HIGH-05] |
| Full suite passes under `lazy="raise"` | [CRIT-02] |
| Claim SQL contains `LIMIT` and no eager-load join | [CRIT-01] |
| Provider client instance is reused across two consecutive jobs | [HIGH-02] |
| `Settings` accepts a provider timeout above 20 seconds | [HIGH-03] |
| `list_sources_detail` emits a join rather than a full-table load | [LOW-08] |
| An update through a service advances `updated_at` | [SQLModel §3] |
| Distinct database URLs yield distinct, individually evictable engines | [MED-04] |
| Post-migration row counts match the backup | [Phase 8] |
- All database, integration, and UI tests continue to use isolated data and never touch `data/transcription.db`.
- Destructive tests continue to run only through `tools/run_destructive_tests.py`, which must first be made to run on Windows.
## Risks
| Risk | Likelihood | Impact | Mitigation |
| :--- | :--- | :--- | :--- |
| The `lazy="raise"` flip surfaces load paths the tests do not cover, breaking a UI page at runtime | High | Medium | Flip one model at a time; exercise every page manually at the phase boundary; `lazy="raise"` fails loudly rather than silently, which is the point |
| Phase 2 is partially applied and leaves an inconsistent schema | Medium | High | Treat Phase 2 as one commit; rebuild from empty rather than upgrading; verify all four schema assertions before proceeding |
| `RegistryService` generalization subtly changes a user-facing message or error category | Medium | Medium | Preserve message strings verbatim, templating only the noun; require the existing registry tests to pass unchanged |
| The atomic claim behaves differently on SQLite than the `FOR UPDATE SKIP LOCKED` path it is written to support | Medium | Low | Single worker in V4.6 means the SQLite path is the only one exercised; the Postgres branch is dialect-guarded and explicitly unverified until the cutover |
| Removing the timeout cap allows a pathological hang | Low | Medium | Pair the removal with an explicit `httpx.Timeout` so the client, not the config bound, enforces the ceiling |
| The migration script loses or reinterprets evidence | Low | High | Verify row counts per table, verify artifact integrity hashes post-migration, and never touch on-disk files |
| Remediation quietly becomes feature work | Medium | Medium | Every commit cites a finding ID; anything without one is recorded for a later revision |
| `ty` cannot reach zero without unsound suppressions | Medium | Low | Suppressions are acceptable where SQLModel typing is genuinely unrepresentable, but each must be `# ty: ignore[<rule>]` with a specific rule, never blanket |
## Delivery Order
1. Phase 1 — Deletions and Quick Wins
2. Phase 2 — Schema Re-Level (single atomic pass)
3. Phase 3 — Worker and Provider Reliability
4. Phase 4 — Service Layer Consolidation
5. Phase 5 — UI Boundaries and Duplication
6. Phase 6 — Async I/O and Configuration Hygiene
7. Phase 7 — Type Checking and Tooling Gate
8. Phase 8 — Data Migration
## Done Criteria
V4.6 is complete when every acceptance criterion in the [V4.6 Scope Boundary](scope_boundary_v4_6.md) is satisfied, specifically:
- All 264 pre-existing tests pass, with every modified test individually justified.
- `ruff check` and `ty check` both report zero and gate on pre-commit.
- No hand-rolled DDL, dead module, dead setting, or duplicate implementation identified in the review remains.
- The schema is generated from metadata, correctly indexed, cycle-free under the PostgreSQL dialect, and free of bidirectional `lazy="selectin"`.
- Roughly 900 lines of duplication are removed across the service and UI layers.
- The backed-up V4.5 data is restored into the re-leveled schema with matching row counts and unmodified on-disk files.
- No new user-facing feature exists that did not exist in V4.5.
## Related Local References
- [V4.6 Scope Boundary](scope_boundary_v4_6.md)
- [Architecture & Code Review Report](../architecture_code_review_2026-08-17.md)
- [V4.5 Scope Boundary](../ver4.5/scope_boundary_v4_5.md)
- [V4.5 Implementation Plan](../ver4.5/implementation_plan_v4_5.md)
- [V4.2 Evidence and Provenance Scope](../ver4.2/scope_boundary_v4_2.md)
- [V4 Architecture](../ver4/architecture_v4.md)
- [V4 Schema](../ver4/schema_v4.md)
- [Transcription Methodology](../invariant/transcription_methodology.md)
- [AI Evidence and Provenance Invariant](../invariant/ai_evidence_and_provenance.md)
-489
View File
@@ -1,489 +0,0 @@
# V4.6 Implementation Review Log
Working record kept during the V4.6 remediation release and the V4.7 / V4.8 planning that followed.
This file is the canonical reference for citations of the form **`review log [N]`** in the V4.6, V4.7, and V4.8 planning documents. The numbers below are those `N` values.
The log was maintained live in a session-scoped database and exported here so the citations remain resolvable in later sessions. It is a historical record: entries are not rewritten after the fact, so some capture reasoning that was later revised. Where an entry conflicts with a committed planning document, **the planning document wins**.
## Legend
| Field | Meaning |
| :--- | :--- |
| `kind` | `question` - needed a decision; `comment` - observation; `deviation` - departure from plan; `risk` - identified hazard |
| `status` | `open` - unresolved; `answered` - resolved by a decision; `noted` - recorded, no action required |
| `finding` | Finding ID in [architecture_code_review_2026-08-17.md](../architecture_code_review_2026-08-17.md), where one applies |
**70 entries** - 8 open, 30 answered, 32 noted.
## Still Open
These carry forward. Most are scoped into V4.7; see [V4.7 scope boundary](../ver4.7/scope_boundary_v4_7.md).
| ID | Finding | Summary | Disposition |
| :--- | :--- | :--- | :--- |
| [8] | - | handle_worker_exceptions swallows everything | V4.7 Phase 5 |
| [18] | - | One flaky failure observed once, then five clean full runs | Watch item - no action |
| [40] | HIGH-06 | No CI workflow enforces the gate | V4.7 Phase 6 |
| [45] | n/a | The same JobSourceStatus enum is persisted two different ways | V4.7 Phase 2 (absorbed into the enum migration) |
| [50] | HIGH-03 | Worst-case stall latency is now 60s (2 x 30s), down from 360s | Accepted risk - mitigated by 30s timeout and max_retries=1 |
| [53] | HIGH-03 | PROVIDER_MODELS still offers two models that cannot finish a dense page within 30s | Operator judgement - deliberately left open |
| [54] | n/a | Run-time telemetry is captured but has no aggregate view | V4.8, gated on V4.7 Phase 4 |
| [55] | n/a | duration_ms measures end-to-end page processing, not provider latency | V4.7 Phase 4 |
## Full Log
### V4.6 Phase 1 - deletions and quick wins
#### [1] worker_retry_backoff_seconds deleted, not wired
*deviation* - `LOW-05` - **noted**
No backoff behavior existed anywhere in the codebase. Wiring it would have been a new feature, which V4.6 forbids. Deleted the setting and its test instead.
#### [2] sqlite_check_same_thread wired, not deleted
*deviation* - `MED-02` - **noted**
Opposite call from the one above: the engine hardcoded the setting's own default value, so wiring it through preserved behavior exactly.
### V4.6 Phase 2 - schema re-level
#### [3] Dev DB safe to discard?
*question* - **answered**
Asked before the atomic schema re-level. You chose "Safe to discard, proceed". The old file was moved to data/transcription.db.pre-v46.bak rather than deleted, because Phase 8 needs it as the migration source.
#### [4] Guard test found 3 relationships the review missed
*comment* - `CRIT-02` - **noted**
The new lazy-load regression test caught ExecutionAttempt.job_source, ProcessingArtifact.execution_attempt, and ProcessingArtifact.source declaring no lazy strategy at all, so they silently defaulted to "select". The review had only catalogued the 16 explicit selectinload cases.
#### [5] lazy="raise" not smoke-tested in a browser
*risk* - `CRIT-02` - **answered**
Every relationship access was audited against its feeding service method and all resolve to detail variants with complete eager loads, and the suite is green. But no manual UI walkthrough was done. A missed path would raise at render time rather than silently N+1. [PARTIALLY RESOLVED] Post-migration smoke test against the real 282-row corpus: a service-layer walk over all 8 documents and every source exercised list_documents, list_jobs, read_document, list_sources, read_source_navigation and list_processing_artifacts with no lazy-load error, and all six /ui pages returned HTTP 200. NiceGUI renders over websocket, so this is not a substitute for clicking through a live browser session, but every query path is now exercised against real data.
### V4.6 Phase 3 - worker and provider reliability
#### [6] httpx default timeout was the real bug
*comment* - `HIGH-03` - **noted**
The review said the 20s cap was too low. The actual defect was larger: httpx.AsyncClient was built with no timeout at all, so every phase defaulted to 5s and the outer asyncio.wait_for could never bind. Real read budget was 5s, not 20s.
#### [7] Your .env still pins WORKER_PROVIDER_TIMEOUT_SECONDS=20
*question* - `HIGH-03` - **answered**
I deliberately did not edit your .env. The new default is 180s but your local file overrides it. Do you want it raised, and to what value? [RESOLVED 2026-08-17] .env lines 51-52 replaced: the stale "[0-20]" comment is gone and WORKER_PROVIDER_TIMEOUT_SECONDS is now 180.0. Verified effective value via get_settings() = 180.0. .env is gitignored, so this is a local-only change with no commit. [REVISED 2026-08-18] User challenged 180.0 as too long to feel responsive. Queried the 80 migrated execution_attempt rows: 77 succeeded with max 18.5s, median 4.9s, p95 18.5s. The only 3 attempts over 20s are exactly the 3 local_timeout FAILURES (20.4/20.8/22.0s), so the distribution is CENSORED - the true duration of those 3 is unknown. Set to 60.0: ~3.2x the slowest observed success, while keeping time-to-discovery of a stall at one minute rather than three. User was right; evidence does not support 180.
#### [8] handle_worker_exceptions swallows everything
*comment* - **open**
A stale monkeypatched signature raised TypeError, which the handler suppressed, spinning run_worker_loop forever. Same hazard exists in production code: any programming error inside the loop becomes an infinite silent retry. Out of V4.6 scope.
### V4.6 Phase 4 - service layer consolidation
#### [9] Registry consolidation did not reduce line count much
*comment* - `MED-11` - **noted**
Public method names had to survive so registry tests could pass unchanged, so each service keeps thin delegating wrappers. Net production code is down ~150 lines overall, but the registry work itself is roughly break-even. The win is single-source-of-truth behavior, not brevity.
#### [10] MED-14 done only for documents.py
*deviation* - `MED-14` - **answered**
store.py and workflows.py still import sources.py. Both are orchestration modules, which services.instructions.md:75-77 permits, so I scoped the boundary test to service-class modules only. Flagging in case you read the rule more strictly. [DEFERRED TO V4.7 by user, 2026-08-17] The MED-14 boundary test stays scoped to service-class modules for V4.6; store.py and workflows.py continue to import sources.py under the Service Composition allowance. V4.7 should decide whether to widen the rule and the AST guard.
#### [11] Two behavior changes in the media writer
*deviation* - `MED-13` - **answered**
mkdir failures now raise a domain error instead of escaping as raw OSError, and homepage image writes gained error handling they never had. Both are strictly better but are behavior changes, not pure refactors. [VERIFIED 2026-08-17] Grepped all of src/ for "except OSError": every hit is either the new media_storage.py:44 wrapper itself, a prompts.py file read, a normalization.py decode guard, or a best-effort cleanup/unlink (store.py:325-330, documents.py:259, jobs.py:349, sources.py:504). No caller wraps a media WRITE in except OSError, so translating mkdir failures into a domain error changes no existing handler behaviour.
#### [12] DocumentService print path no longer raises TranscriptionError
*deviation* - `MED-14` - **answered**
_print_media_type raises DocumentError for an unsupported extension where it previously raised TranscriptionError. Only reachable with a corrupt stored filename. Changed because a DocumentService emitting a transcription error is itself the boundary leak MED-14 is about. [VERIFIED 2026-08-17] Grepped all of src/ for "except TranscriptionError": the single hit is store.py:353, which wraps validate_source_content on the UPLOAD path and is unrelated to the print projection in documents.py. No handler anywhere catches TranscriptionError around _print_media_type, so the class change is behaviourally inert. Also clarified to the user that this item is a disclosure of a change already made, not an outstanding error, and is unrelated to .env.
### V4.6 Phase 5 - UI boundaries and duplication
#### [13] document_panzoom: delete or fix?
*question* - `HIGH-07` - **answered**
Unused 170-line component plus ~35 lines of CSS, superseded in practice by dark_room_viewer. You chose to delete it now and rebuild it cleanly in the next revision alongside other photo/image features. ACTION FOR V4.7: pan-zoom must be reintroduced.
### V4.6 Phase 5 - UI boundaries and duplication
#### [14] sources_page transport_body deferral question (answered; premise corrected)
*question* - **answered**
Initially believed the deferred-body branch was dead. Re-check showed _transport_display() is called with latest_attempt from read_latest_execution_attempt(), which DOES defer transport_body. Current behavior is already "Omitted from Source Detail". Fix is therefore a pure boundary move: read_latest_execution_attempt returns a LatestExecutionAttempt read model carrying transport_body_deferred: bool, and sources_page drops sqlalchemy.inspect. No behavior change. User preference recorded: simplest, most supportable, most robust; full bytes remain persisted and retrievable via Export Evidence.
#### [15] Guard-message ordering changed on two Source routes
*deviation* - **noted**
sources_page previously parsed the route id BEFORE rendering the navigation header, then rendered the invalid-id message after it. Adopting the shared parsed_record_id() helper moved the nav header above the parse. Net rendered output is identical; only the internal call order changed.
#### [16] Settings-page registry tables now render inside build_table
*deviation* - **answered**
The two label-registry tables on the settings page were hand-rolled ui.table calls. They now go through build_table via a new components/table/registry.py. build_table wraps its table in a ui.column, so the tables gain one extra container div. Search is disabled and rows-per-page stays 0, so visible behavior is unchanged. | RESOLVED (user directed consolidation): build_table gained row_key; linked_people.py converted; print_preview_page.py shares a local _render_print_table helper (print tables intentionally bypass build_table - no pagination, no search). New AST guard test_only_the_designated_owners_construct_a_raw_table pins ui.table() to exactly table/common.py and print_preview_page.py.
#### [17] Two hand-rolled ui.table instances deliberately left alone
*comment* - **noted**
print_preview_page.py has two print-layout tables and linked_people.py has a component-local editor table. Neither wants build_table search or pagination, so converting them would add indirection without removing duplication. Flagging in case you want them unified later. | CLOSED AS ENVIRONMENTAL: unreproduced after ~54 sequential full-suite runs (incl. a dedicated 25-run soak with -rA traceback capture, 0 failures) plus 5 concurrent-process runs (2x tests/ui, 3x full suite). Not attributable to any V4.6 change; the single observed failure occurred immediately after a burst of bulk file rewrites. No code change made. Re-open if it recurs.
#### [18] One flaky failure observed once, then five clean full runs
*risk* - **open**
tests/ui/test_jobs_page.py::test_job_delete_page_allows_deletion_for_queued_or_completed_job failed once and passed on every subsequent run (5 consecutive full-suite runs, 275 passed / 4 skipped). This matches the known pre-existing aiosqlite event-loop teardown noise that lands on a random test. Not introduced by Phase 5, but worth confirming during Phase 6/7.
#### [19] store.create_document_job / create_job_for_document now own their session
*comment* - **noted**
To remove session_scope from jobs_page, both orchestration functions accept an optional session plus an optional session_factory and open their own scope when neither is supplied. Existing callers that pass a session are unaffected; tests pass unchanged.
#### [20] Upload accept lists are now derived from SOURCE_EXTENSIONS
*comment* - **noted**
The job upload picker previously hard-coded .jpg,.jpeg,.png,.tif,.tiff,.pdf. It now derives the list from services.source_media.SOURCE_EXTENSIONS, so adding a Source format in one place updates the picker. The portrait and homepage pickers share a separate IMAGE_UPLOAD_EXTENSIONS list because they accept gif/webp/bmp, which are not valid Source formats.
### V4.6 Phase 6 - async I/O and configuration hygiene
#### [21] model_copy(update=...) rejected by pydantic-settings
*deviation* - `MED-04` - **noted**
Plan offered "model_copy(update=...) or a computed property" to replace object.__setattr__ in normalize_provider_models. model_copy failed: pydantic-settings warns "A custom validator is returning a value other than self ... isn't supported when validating via __init__" and 3 config tests failed. A computed property would have required renaming the env-facing provider_models field. Implemented as a model_validator(mode="before") over the raw input dict instead, so the derived value is produced by normal construction with no frozen-instance mutation. All 23 config tests pass.
#### [22] provider_model is now trimmed
*deviation* - **noted**
The old object.__setattr__ path assigned provider_model without stripping whitespace; only the provider_models tuple entries were stripped. The before-validator now strips provider_model too. This is a behavior change, judged a correctness improvement since an untrimmed model id would be sent to the provider. No test asserted the old behavior.
#### [23] onupdate confirmed DDL-neutral
*comment* - **noted**
Plan said onupdate moves to Phase 2 if it alters emitted DDL. Verified by hashing CreateTable output for every table on both the sqlite and postgresql dialects before and after the change: identical (b33ad56a...). onupdate stays in Phase 6; Phase 2 does not need re-verification.
#### [24] No-op updates no longer bump the timestamp
*deviation* - **noted**
Removing the 10 manual "updated_at = datetime.now(UTC)" assignments means an update call that changes nothing no longer marks the row dirty, so onupdate does not fire and the timestamp stays put. Previously the manual assignment always bumped it. Judged more correct for a column that is supposed to track modification, but it is an observable change for any caller that relied on update-as-touch.
#### [25] Homepage markdown I/O left unwrapped
*question* - `MED-01` - **answered**
ui/homepage_store.py reads and writes a single small local markdown file synchronously from home_page.py handlers. MED-01 names Pillow normalization, artifact writes, and evidence hashing; this is none of those and the payload is trivial. Left unwrapped to avoid scope creep. Flagging in case you want it wrapped anyway. [RESOLVED 2026-08-17] User decision: leave it synchronous. No change made.
#### [26] Evidence manifest hashing left on the loop
*comment* - `MED-01` - **noted**
providers/evidence.py digest() hashes a small in-memory JSON manifest (microseconds), so it was left inline. The hashing that actually mattered was over page-sized image bytes: the derivative digest is now precomputed inside normalize_orientation (already off-loop) and the artifact digest now shares the same worker-thread hop as the write.
#### [27] dispose_engine on an unknown URL changed behavior
*comment* - `MED-04` - **noted**
The old functools.cache version called get_engine(url) inside dispose_engine, which would construct an engine just to dispose it, and then cache_clear() wiped every other engine too. The registry version pops only the requested URL and no-ops on an unknown one. Covered by tests/test_engine_registry.py.
#### [28] Added homepage_dir setting (user-approved scope addition)
*deviation* - **answered**
ui/homepage_store.py was the only storage path in the codebase derived from Path(__file__).parents[3] rather than from Settings, making it unconfigurable and wrong under a wheel install (it would resolve into site-packages). Not tied to a review finding ID, so it is a deliberate scope addition, approved by the user in-flight. Added Settings.homepage_dir (default ./data/homepage) and rewrote the module to resolve from Settings, with an optional settings parameter on every function. Covered by tests/ui/test_homepage_store.py.
#### [29] homepage default is now CWD-relative
*risk* - **answered**
The old default resolved to <repo>/data/homepage regardless of working directory. The new default Path("./data/homepage") is relative to the process CWD, matching artifact_dir and upload_dir. Running the app from the repo root gives the identical location; running it from elsewhere does not. Consistent with every other storage root, but worth confirming against your deployment/launch scripts. [RESOLVED 2026-08-17] User confirmed the app is only ever launched from the repo root, so CWD-relative ./data/homepage and the old repo-anchored path are identical. Verified live: resolves to C:\GitHub\transcription\data\homepage containing the real homepage.md and portrait. No change needed. Revisit only if a service or scheduled task with its own working directory is introduced.
#### [30] Homepage markdown I/O stays synchronous
*comment* - `MED-01` - **answered**
User question resolved: the async-wrapping question was dropped as negligible (one small local markdown file). The underlying concern turned out to be the hardcoded storage path, addressed separately via Settings.homepage_dir.
### V4.6 Phase 7 - type checking and quality gate
#### [31] selectinload varargs is not equivalent to chaining
*deviation* - `HIGH-06` - **noted**
selectinload(A.b, B.c) and selectinload(A.b).selectinload(B.c) produce an identical .path but the varargs form applies the selectin strategy ONLY to the last element. With lazy="raise" everywhere (Phase 2) the varargs form raises InvalidRequestError at render time. Cost 12 test failures before it was caught. Documented in the db/loading.py docstring.
#### [32] New module src/transcription/db/loading.py
*deviation* - `HIGH-06` - **noted**
Rather than sprinkle 42 suppressions, the SQLModel-field to QueryableAttribute reinterpretation now has one documented home: orm_attribute(), selectinload(), defer(). All 42 "# pyright: ignore[reportArgumentType]" comments in documents/jobs/people/sources were removed as a result.
#### [33] transaction_scope no longer accepts or yields AsyncSessionTransaction
*deviation* - `HIGH-06` - **noted**
AsyncSessionTransaction appeared nowhere outside db/session.py; no caller ever passed one, and sessionmaker.begin() was verified at runtime to yield an AsyncSession. The branch was also latently buggy: services call .exec() which a transaction object does not have. Removing the union cleared 7 downstream workflows.py diagnostics.
#### [34] RegistryService is now bound by a RegistryEntry Protocol
*deviation* - `HIGH-06` - **noted**
RegistryService[ModelT: SQLModel] gave ty no visibility into id/label/normalized_label/is_active. A structural Protocol replaces the bare SQLModel bound - a genuine typing improvement rather than a suppression. Cleared 9 diagnostics.
#### [35] normalization.py now uses isinstance(image, TiffImageFile) instead of image.format == "TIFF"
*deviation* - `HIGH-06` - **noted**
tag_v2 only exists on TiffImageFile. The isinstance check is semantically equivalent and types correctly.
#### [36] linked_people.render switched from @ui.refreshable to @ui.refreshable_method
*deviation* - `HIGH-06` - **noted**
refreshable_method is the NiceGUI API intended for bound methods; the plain decorator mistyped self. render.refresh() call sites are unchanged.
#### [37] read_source_navigation now wraps literal bounds in sqlalchemy.literal()
*deviation* - `HIGH-06` - **noted**
tuple_() rejects raw Python values under typing. literal() is the correct explicit coercion and preserves the emitted SQL.
#### [38] openrouter capturing client re-raises ResponseNotRead for a sync stream
*deviation* - `HIGH-06` - **noted**
response.stream is typed SyncByteStream | AsyncByteStream. The narrowing guard re-raises rather than silently mis-wrapping, which is the honest behavior on an async client.
#### [39] No pre-commit config existed; one was created
*comment* - `HIGH-06` - **noted**
The plan said "wire it into the existing pre-commit setup", but there was no .pre-commit-config.yaml (pre-commit was only a dev dependency, and there are no CI workflows either). A local-repo config with blocking ruff and ty hooks was created and negative-tested. NOTE: hooks use language: system, so the venv Scripts dir must be on PATH.
#### [40] No CI workflow enforces the gate
*risk* - `HIGH-06` - **open**
.github/workflows/ is empty, so ruff/ty/pytest are only enforced locally via pre-commit, and only if the developer has installed the hooks (pre-commit install). Consider adding a CI workflow in a later release.
#### [41] ty check driven from 207 diagnostics to 0
*comment* - `HIGH-06` - **noted**
Two real bugs were fixed en route: tools/run_destructive_tests.py imported ctypes.wintypes at module scope (raising on non-Windows) and used fcntl unconditionally; tests/ui/test_sources_page.py constructed Source(...) without the required document_id. Only two suppressions remain in the whole tree: one "# ty: ignore[invalid-assignment]" in tests/test_prompts.py which deliberately assigns to a frozen field to assert ValidationError.
#### [42] asyncio_default_fixture_loop_scope pinned to "function"
*comment* - `HIGH-06` - **noted**
Set explicitly in pyproject.toml so pytest-asyncio behavior does not shift on upgrade.
### V4.6 Phase 8 - data migration
#### [43] V4.6 re-level changed no columns at all
*comment* - `review 1a` - **noted**
Diffing the backup schema against the current SQLModel metadata showed identical table sets and identical column sets for all 10 tables. What V4.6 actually changed is index coverage (9 new indexes: ix_document_document_type_id, ix_document_person_document_id, ix_document_person_person_id, ix_document_person_role_id, ix_job_document_id, ix_job_source_job_id, ix_job_source_source_id, ix_job_status_date_created, ix_source_document_id - none lost), the use_alter break in the FK cycle, and the relationship loading strategy. The migration is therefore a faithful FK-ordered row copy rather than a transformation.
#### [44] Migration reads the backup with raw sqlite3, not the ORM
*deviation* - `review 1a` - **noted**
The plan anticipated ORM reads carrying explicit eager loads under lazy="raise". Reading raw rows is strictly safer: the V4.5 file is not guaranteed to satisfy the V4.6 mappers, and no relationship is ever traversed, so lazy="raise" cannot bite at all. Writes still go through SQLAlchemy Core against the live metadata, so the script will work against PostgreSQL unchanged.
#### [45] The same JobSourceStatus enum is persisted two different ways
*risk* - **open**
job_source.status declares values_callable and stores lowercase VALUES ("transcribed"); execution_attempt.status does not and stores uppercase NAMES ("TRANSCRIBED"). Both columns use the identical JobSourceStatus enum. This is a genuine latent inconsistency: any raw SQL, reporting query, or future cross-dialect move has to know which spelling each column uses. It is NOT a finding in the review, so under the pure-remediation rule I did not change it - the migration accepts either spelling and round-trips both faithfully. RECOMMEND scheduling this for V4.7.
#### [46] Should the migration be applied to the live data/transcription.db?
*question* - **answered**
The script is fully verified against a throwaway target: 282 rows copied, every table byte-identical to the backup cell-for-cell, idempotent re-run inserts 0, artifact integrity passes, no on-disk file touched. The live data/transcription.db currently holds only bootstrap seed rows (document_type 6, person_role 3) whose UUIDs differ from the backup, so a straight migration would ADD the backup rows alongside the seeds and likely trip the normalized_label uniqueness constraint. Applying cleanly requires replacing the live file. Awaiting user decision. [RESOLVED] User chose to back up and replace. data/transcription.db.seed-20260817-200555.bak holds the old seed file; a fresh DB was created and all 282 rows migrated with artifact integrity verified.
#### [47] Provider timeout set to 60s on evidence, not on the review's suggested figure
*comment* - `HIGH-03` - **noted**
The review recommended 120s and the V4.6 plan used 180s, both chosen without data. The migrated corpus provides data: 77/80 attempts succeeded, all within 18.5s. 60s is the smallest value with real headroom that still surfaces a stall quickly. Revisit only if a genuine local_timeout occurs at 60s.
#### [48] Three historical local_timeout failures are worth re-running
*deviation* - **answered**
All 3 FAILED execution_attempts were killed by the old 20s ceiling and carry response_received=1, meaning a response had begun arriving when the budget expired. With the ceiling now at 60s these three pages may well succeed on a retry. Their evidence rows were migrated unchanged, so the originals are preserved either way. | RESOLVED 2026-08-17: not 3 pages but ONE page (source 302aa684) x 3 models. Re-ran each model 2x with a 300s uncensored ceiling: gemini-flash 9.0/22.5s, claude-opus-5 27.0/27.6s, gpt-5.6 64.5/75.3s. All 6 succeeded - no hangs. gpt-5.6 exceeds the 60s value that was set, so .env raised to 120.0 (~1.6x slowest success). Historical stats were ~96% gemini-flash and understated the budget.
#### [49] Reporting gap: 28 review_log entries were never surfaced to the user
*risk* - **answered**
My end-of-run summaries filtered on status IN (open, answered), which silently excluded every entry recorded as "noted" - 28 of 46. The user caught this. All 28 are now presented. Lesson: "noted" is not the same as "reported".
### Post-V4.6 - timeout calibration and tuning
#### [50] Worst-case stall latency is now 60s (2 x 30s), down from 360s
*risk* - `HIGH-03` - **open**
Superseded by the 2026-08-18 calibration: WORKER_MAX_RETRIES=1 and WORKER_PROVIDER_TIMEOUT_SECONDS=30.0 give a worst case of 60s. The underlying concern stands but is much reduced: handle_worker_exceptions (review_log id 8) still swallows every exception, so a programming error would burn 2 attempts silently with no UI feedback. Keep id 8 as the real fix.
#### [51] Removed WORKER_RETRY_BACKOFF_SECONDS from .env
*comment* - `HIGH-03` - **noted**
The setting was deleted from Settings in Phase 1 (LOW-05). Because Settings uses extra="ignore" it sat in .env silently inert, which is exactly the DATABASE_URL trap the review flagged. Removed from .env so the file matches the model. No behavior change.
#### [52] Timeout set to 30.0s and max_retries to 1 by user decision
*comment* - `HIGH-03` - **answered**
Full dropdown measured twice on the densest page in the corpus with a 300s uncensored ceiling. Fast cluster: gemini-2.5-flash 9.0/22.5, gpt-4o 21.6/22.9, claude-sonnet-4 26.2/26.7, claude-opus-5 27.0/27.6. Slow cluster: gemini-2.5-pro 49.5/79.4, gpt-5.6 64.5/75.3. User chose 30.0s at the low edge of the 27.6-49.5s gap because dense forms are <10 of ~3k documents and ejecting a stalled outlier is preferred over waiting. Worst case is now 2x30=60s. Agent recommended 40s for margin; user declined with stated rationale. Accepted.
#### [53] PROVIDER_MODELS still offers two models that cannot finish a dense page within 30s
*risk* - `HIGH-03` - **open**
gemini-2.5-pro and gpt-5.6 remain selectable in the jobs page dropdown (ui/pages/jobs_page.py:146 reads settings.provider_models). Both exceed 30s on dense forms by design of the chosen budget, though both should still succeed on the ~99.7% of pages that are not dense forms. Left in the list deliberately - not removed - so the user retains them for quality comparison. Revisit if dense-form failures become noisy.
### Post-V4.6 - V4.7 candidates identified
#### [54] Run-time telemetry is captured but has no aggregate view
*comment* - **open**
execution_attempt.duration_ms is a required non-null field written on all three paths in services/workflows.py (success 278, TimeoutError 295, general failure 330); failures use a monotonic clock, so timeout durations are trustworthy. started_at/finished_at are also stored, and normalized_metadata.usage carries token counts on the same row, so tokens/sec is already derivable per attempt. Gaps: (1) sources_page.py:400 renders it raw as "27612 ms" rather than seconds; (2) it is only visible for the latest attempt of one source at a time - there is no rollup, so answering "which model is slow" required hand-written SQL against the database. A small model-performance rollup is a V4.7 candidate.
#### [55] duration_ms measures end-to-end page processing, not provider latency
*comment* - **open**
services/workflows.py:221 sets monotonic_started_at BEFORE provider_input preparation (image normalization, artifact persistence, session.commit() at line 228), and line 251 computes elapsed_seconds from it. But the asyncio.wait_for timeout at lines 240-249 wraps ONLY _call_transcriber. So duration_ms covers a strictly wider window than the budget that governs it. Empirical proof: the three historical local_timeout rows recorded 20.4/20.8/22.0s against a 20.0s timeout, i.e. roughly 0.4-2.0s of non-provider work is folded in. Consequence: duration_ms cannot be used to isolate provider performance, and any model-performance rollup built on it would be polluted by preprocessing time that varies with image size. V4.7 candidate: record provider latency as a separate column, or move monotonic_started_at to just before the wait_for.
### V4.7 planning
#### [56] Release split agreed: V4.7 = architectural cleanup, V4.8 = features
*comment* - `MED-14` - **answered**
User asked whether the sources.py decomposition (architectural) should be separated from pan-zoom and photo work (features). Agreed and documented. Rationale: V4.6 succeeded because it had a binary gate - behavior-identical, suite unchanged. A refactor can be held to that standard; features cannot, since they require new tests. Bundling them destroys the ability to attribute a test delta to a bug versus expected new behavior. The two also touch disjoint trees under different instruction files (services vs ui). Created docs/ver4.7/scope_boundary_v4_7.md, docs/ver4.7/implementation_plan_v4_7.md, docs/ver4.8/feature_backlog_v4_8.md. All cross-links verified.
#### [57] Rejected the original proposal to move update_job_source_transcription to workflows.py
*deviation* - `MED-14` - **noted**
The V4.6 Phase 5 deferral note suggested moving it as orchestration. Rejected in the V4.7 boundary. services.instructions.md:63-65 requires transcript updates and the paired terminal status change to commit or roll back together, and the method writes JobSource plus ExecutionAttempt in one session scope, deriving attempt_number from ExecutionAttempt at lines 595-600. Line 72 assigns session-aware write helpers to services and commit-boundary control to orchestration, so moving a multi-table write into workflows.py inverts the stated architecture. Scope reduced from three moves to two (artifacts.py, evidence.py); expected sources.py ~930 lines rather than the deeper cut originally implied.
#### [58] Pan-zoom renumbered from V4.7 to V4.8, intent preserved
*comment* - `HIGH-07` - **noted**
Commit 6a3ee26 states pan-zoom would return "in V4.7 alongside the other photo/image work". It now sits in the V4.8 backlog. The commit intent was grouping with the photo work, not the specific number, and that grouping is preserved. Recorded in the V4.8 backlog so the git history is not silently contradicted. Also noted there: document_panzoom.py was exported but wired to no page, so no user has seen it - which makes reintroduction a new feature rather than a restoration, and is what puts it on the feature side of the split.
#### [59] Service ownership model for ExecutionAttempt / ProcessingArtifact is undecided
*question* - `MED-14` - **answered**
services.instructions.md:11 says "1 service class per data model" but there are 10 persisted models and 6 service classes. The 4 unnamed models landed arbitrarily: DocumentPerson->people.py, and JobSource + ExecutionAttempt + ProcessingArtifact all -> sources.py. Line count tracks model count: jobs.py 438L (1 model), documents.py 473L (1+registry), people.py 554L (2+registry), sources.py 1389L (4). Evidence gathered 2026-08-18: both tables were added in V4.2 commit 6bd4cbb ("Updated what ai_raw_response data is being captured"), i.e. AFTER the 4-component design. ExecutionAttempt is docstringed "Immutable evidence for one provider call attempt" (request manifest, transport body, router ids, sdk snapshot, software context, timing); ProcessingArtifact is "Provider-neutral, versioned output derived from a Source" with a XOR CheckConstraint on inline vs external content, and a NULLABLE execution_attempt_id, so an artifact can exist with no attempt. Both are append-only provenance, not mutable domain entities. Four options were put to the user (evidence-as-own-subsystem; evidence split with ProcessingArtifact under Source; strict lifecycle into JobService/SourceService; draft the revised instructions first). USER DEFERRED - continuing the discussion interactively, formulating further questions. Do not proceed with V4.7 Phase 1/2 until this is settled, since the chosen model determines the module split.
#### [60] V4.7 boundary overstates the case against moving update_job_source_transcription
*deviation* - `MED-14` - **answered**
The scope boundary as written says moving it to workflows.py would violate services.instructions.md. On re-reading, lines 38-47 (the _finalize contract: commit when service-owned, flush when caller-owned) and line 72 describe exactly the mechanism that makes a multi-service atomic write safe, so the document permits it. The honest objection is weaker: keeping the paired JobSource + ExecutionAttempt write in one method makes atomicity enforced by locality, whereas splitting it makes atomicity depend on every future caller sharing the session correctly. That is a robustness argument, not a rule violation. Correct the wording in docs/ver4.7/scope_boundary_v4_7.md section 1 before that document is treated as frozen.
### V4.7 design - evidence model simplification
#### [61] job_source duplicates execution_attempt columns byte-for-byte
*comment* - `MED-14` - **answered**
Measured on live DB: job_source.raw_transcription 77/77 identical to latest attempt; ai_metadata vs normalized_metadata 77/77; raw_api_response vs sdk_response_snapshot 77/77; error_detail 2/2. The table split itself is justified by cardinality (1:N attempts) but the 1:N is exercised in only 1 of 79 job_source rows. The 4 duplicated columns are an undocumented, unenforced denormalized cache.
#### [62] status mismatch cross-confirms the enum persistence defect
*risk* - `45` - **answered**
job_source.status vs execution_attempt.status compared 0/79 identical - job_source stores lowercase (transcribed), execution_attempt stores uppercase (TRANSCRIBED). Independent confirmation of finding [45].
#### [63] Orientation normalization was NOT a red herring - proven visually
*comment* - `ProcessingArtifact` - **answered**
59 of 79 source images carry EXIF orientation=3 (rotate 180). Rendered the exact page the user described (Pioneer Days page 00, a typed table of contents): raw decoded pixels are genuinely upside down; the 180-rotated version is upright. So sending raw bytes did send an inverted page to the model. resolve_provider_input (workflows.py:226 -> sources.py:850) is on the current hot path, so normalization now runs, but only 1 orientation artifact exists - the other 58 rotated pages were transcribed before normalization was wired.
#### [64] job_source cannot be deleted - it is the work queue, not just a junction
*risk* - `MED-14` - **answered**
store.py:249,313 create JobSource with status=PENDING at job creation, before any provider call. workflows.py:438-440 selects pending work by status != TRANSCRIBED. jobs.py:411-424 retry mutates FAILED back to PENDING and clears fields. jobs.py:378-384 cancel writes FAILED/Cancelled by user with NO provider call, so no execution_attempt row could exist to carry it. An append-only table cannot express queued-not-yet-attempted or cancelled-before-call. Recommend STRIP not DELETE: keep (id, job_id, source_id, status); drop raw_transcription, ai_metadata, raw_api_response, executed_at.
#### [65] Entire artifact subsystem has executed exactly once
*comment* - `ProcessingArtifact` - **answered**
Only 2 rows exist, both from the same job on 2026-08-16 (13:59:57 orientation, 14:00:07 quality). workflows.py:560-571 writes a transcription_quality_warnings artifact on EVERY successful page, yet 77 successful transcriptions produced 1 row - so the code path postdates nearly all data. Ingest already copies bytes via media_storage.py:57 write_bytes, so normalize-at-upload is viable. Caveat: Pillow re-encodes JPEG at quality 95, a permanent generational loss for an archival corpus - recommend retaining original bytes as a sibling file.
#### [66] Retry history already exists in execution_attempt - no resubmitted flag needed
*comment* - `MED-14` - **answered**
ExecutionAttempt UniqueConstraint(job_id, source_id, attempt_number) at models.py:389 already implements keep-the-failed-row-and-add-a-new-one. Proven in live data: attempt 1 FAILED/local_timeout 20.4s and attempt 2 TRANSCRIBED 13.6s both retained. Adding a second job_source row would duplicate that and break the one-row-per-(job,page) assumption in read_job_source_for_job and sources.py:570-574, where uniqueness is enforced in CODE not by a DB constraint.
#### [67] Add JobSourceStatus.CANCELLED to retire job_source.error_detail
*comment* - `MED-14` - **answered**
Cancel currently overloads FAILED plus free text Cancelled by user (jobs.py:378-384). A distinct CANCELLED status separates user cancellation from genuine provider failure and removes the last consumer of job_source.error_detail, reducing job_source from 9 columns to 4: id, job_id, source_id, status.
#### [68] Lossless 180-degree JPEG rotation is viable for 57 of 58 rotated images
*comment* - `ProcessingArtifact` - **answered**
Pillow always round-trips through decoded pixels (normalization.py:76-86), so quality=95 re-encode loss is inherent to the library, not required by the task. A 180 rotation is expressible as a lossless DCT transform when both dimensions are multiples of the 16px MCU. Measured across the corpus: 57/58 qualify; the sole exception is 2306x2019. Alternative that needs no new dependency: normalize at upload and retain the original bytes as the archival master.
#### [69] Quantization-table reuse beats both current settings and the lossless-DCT route
*comment* - `ProcessingArtifact` - **answered**
Measured single-generation rotate-and-restore on 5 rotated JPEGs. Current settings (quality=95, subsampling=0, normalization.py:84-85): PSNR 50.0-53.5 dB, file size +38 percent. Reusing the source quantization tables and subsampling (qtables=im.quantization, subsampling=JpegImagePlugin.get_sampling(im), optimize=True): PSNR 51.5-55.0 dB, max channel delta 7-9/255, file size slightly SMALLER (636KB->595KB). Better on quality and size simultaneously. Critically it works at any dimensions, so the 2306x2019 MCU-misaligned outlier needs no rejection path - the edge case only exists on the lossless-jpegtran route, which would also require an external C binary. Recommend Pillow with qtables reuse; drop the lossless-DCT option.
#### [70] Evidence-model simplification decisions settled by user
*comment* - `DECISIONS` - **answered**
1) job_source is STRIPPED not deleted - keeps id, job_id, source_id, status (9 columns to 4). Drop raw_transcription, ai_metadata, raw_api_response, executed_at, error_detail. All evidence reads move to execution_attempt. 2) Add JobSourceStatus.CANCELLED so cancel no longer overloads FAILED plus free text, retiring error_detail. 3) Retry keeps its current FAILED-to-PENDING reset - history already lives in execution_attempt via UniqueConstraint(job_id, source_id, attempt_number). 4) PENDING-at-job-creation is unchanged. 5) ProcessingArtifact table REMOVED; orientation normalization moves to upload/ingest; transcription_quality_warnings payload folds into execution_attempt.normalized_metadata. 6) Rotation uses Pillow with qtables + subsampling reuse (visually lossless, ~52 dB PSNR, no size growth, no external dependency, no MCU rejection path). No archival master retained. 7) One-time backfill of the 58 already-ingested EXIF-orientation-3 images.
## Related Local References
- [Architecture & Code Review Report](../architecture_code_review_2026-08-17.md) - finding IDs
- [V4.6 Scope Boundary](scope_boundary_v4_6.md)
- [V4.6 Implementation Plan](implementation_plan_v4_6.md)
- [V4.7 Scope Boundary](../ver4.7/scope_boundary_v4_7.md)
- [V4.7 Implementation Plan](../ver4.7/implementation_plan_v4_7.md)
- [V4.8 Feature Backlog](../ver4.8/feature_backlog_v4_8.md)
-205
View File
@@ -1,205 +0,0 @@
# V4.6 Scope Boundary
This document defines the frozen boundary for V4.6, a **pure remediation release**. V4 through V4.5 remain the architecture and behavioral baseline. V4.6 introduces **no new user-facing features**; it pays down the defects, duplication, and structural drift catalogued in [Architecture & Code Review Report](../architecture_code_review_2026-08-17.md).
Every item in scope is traceable to a review finding ID. Any change that cannot be traced to a finding ID is out of scope.
## Purpose
- Remove dead code, dead configuration, and duplicate implementations that create maintenance drift.
- Re-level the database schema from current SQLModel metadata, ending hand-rolled DDL while the schema is still pre-production.
- Correct read amplification, missing indexes, and query patterns that scale with table size rather than result size.
- Restore the boundaries the project already wrote down in `.github/instructions/services.instructions.md` and `ui.instructions.md`.
- Make `ty` usable as a real quality gate.
- Preserve every existing behavior, evidence guarantee, and provenance contract established in V4 through V4.5.
## Confirmed Operating Context
These answers are frozen for V4.6 and govern every decision below.
| Question | Answer |
| :--- | :--- |
| Database | **SQLite only.** PostgreSQL remains the intended destination but is deferred beyond V4.6. `JSONBCompat` and the Postgres drivers are retained. |
| Topology | **Single user, single process, single worker.** A multi-user server is the stated direction, so forward-compatibility work is retained where it is cheap. |
| Schema evolution | **Re-level from current metadata.** No Alembic, no migration framework, no `_upgrade_*` chain. |
| Existing data | The development database is rebuilt from scratch during implementation and migrated from backup as the final step. |
| Release character | **Pure remediation.** No new features. |
| Scope band | Critical through Low, inclusive. |
## In Scope
### 1. Dead Code and Dead Configuration Removal
- `src/transcription/app_state.py` is deleted. It has zero importers and contains a guaranteed `TypeError` ([HIGH-01]).
- `src/transcription/services/transcription.py` is deleted; `build_prompt_execution` has exactly one import path ([MED-05]).
- The legacy compatibility aliases in `services/store.py` are deleted ([MED-05]).
- `ServiceBase.queue` is deleted; no service allocates an unused `asyncio.Queue` ([MED-07]).
- `sqlite_check_same_thread` and `worker_retry_backoff_seconds` are either wired to real behavior or deleted, along with their tests ([MED-02]).
- `db/operations.py:get_next_queued_job` is deleted as a divergent duplicate ([CRIT-01]).
- `DATABASE_URL` is removed from `docker-compose.yml`, and the real nested `DATABASE__*` names are documented. The application never silently ignores a database configuration variable ([MED-10]).
- `document_panzoom` is either fixed or deleted; it is exported but referenced by no page ([HIGH-07]).
### 2. Schema Re-Level
The following changes are schema-affecting and land as **one single pass** against a database rebuilt from empty.
- `upgrade_schema` and the three `_upgrade_*` functions (`db/operations.py:25-109`) are deleted, along with their tests (`tests/test_db.py:109-172`) ([HIGH-05]).
- The schema is generated exclusively from SQLModel metadata via `create_all()`, gated by the existing `Settings.should_bootstrap_schema` ([HIGH-05]).
- The hand-written `CHAR(32)` column for `preferred_execution_attempt_id` ceases to exist; the column type is whatever the model declares ([HIGH-05]).
- A composite index on `Job.status, Job.date_created` is declared in the model, plus `index=True` on the foreign keys the worker and detail pages filter on ([HIGH-04]).
- `Source.preferred_execution_attempt_id` declares its foreign key with `use_alter=True`, resolving the `source` / `job_source` / `execution_attempt` cycle so `create_all` will succeed on PostgreSQL when that cutover is taken ([HIGH-08]).
- Relationship loading defaults change from bidirectional `lazy="selectin"` to `lazy="raise"`, with per-query `selectinload()` retained or added where a load path genuinely requires it ([CRIT-02]).
No migration script runs against a populated database. No history table, revision directory, or down path is introduced.
### 3. Data Migration
- A one-time script under `tools/` migrates the user's backed-up V4.5 data into the re-leveled schema.
- The script is authored **after** the `lazy="raise"` flip is complete, so that every relationship it traverses carries an explicit eager load.
- The script is idempotent, is never invoked automatically at startup, and never runs as part of the test suite.
- Uploaded Source files, portraits, and artifact files on disk are preserved unchanged; only database rows are rewritten.
- This is the **final** step of V4.6.
### 4. Worker and Provider Reliability
- `read_next_queued_job` gains `LIMIT 1` and stops materializing the entire queue plus its eager graph on every poll ([CRIT-01]).
- The claim becomes an atomic `QUEUED``PROCESSING` transition. On SQLite this is a bounded single-writer transaction; the `FOR UPDATE SKIP LOCKED` path is written and dialect-guarded for the multi-user direction but is not exercised in V4.6 ([CRIT-01]).
- Eager relationships are loaded in a second query after the claim succeeds, keeping the hot poll a single narrow row ([CRIT-01]).
- `ServiceBundle` and the provider client are hoisted to worker-loop scope so the HTTP connection pool and TLS session survive across jobs ([HIGH-02], [MED-06]).
- `ServiceBundle` gains a `from_session_factory` constructor, replacing three duplicated instantiation blocks ([MED-06]).
- The `le=20.0` cap on `worker_provider_timeout_seconds` is removed, the default is raised, and an explicit `httpx.Timeout` is passed to the OpenRouter client ([HIGH-03]).
- The `TranscriptionProvider` Protocol is extended to cover `aclose` and the evidence attributes; the per-call `inspect.signature` reflection at `sources.py:1237` is deleted ([MED-03]).
### 5. Service Layer Consolidation
- A generic `RegistryService[ModelT]` owns list, summaries, create, read, update, delete, and reference-check for semantic-key registries. `DocumentType` and `PersonRole` become thin subclasses declaring their model, error class, reference query, and noun ([MED-11]).
- Label normalization, the casefold key, and the registry summary shape are defined once ([MED-11]).
- `ServiceBase` gains `_get_or_raise`, and all 38 hand-written not-found guards adopt it, including the three in `documents.py` that already bypass the local helper ([MED-12]).
- `services/media_storage.py` becomes the single implementation of validate → hash → write → wrap-error, replacing `store_source_file`, `store_person_portrait`, and the homepage image writer ([MED-13]).
- `source_mime_type` moves out of `services/sources.py` to a shared module so `documents.py` no longer imports a sibling service ([MED-14], partial).
- The four query inefficiencies in `sources.py` are corrected: the `job_id` filter moves into SQL, navigation uses two bounded queries, `list_processing_artifacts` gains a `limit`, and artifact re-hashing moves off the event loop ([LOW-08]).
### 6. Async I/O and Configuration Hygiene
- Blocking filesystem and CPU work — media writes, artifact writes, integrity hashing, and Pillow orientation normalization — is wrapped in `asyncio.to_thread` ([MED-01]).
- `functools.cache` on the engine and session factories is replaced with an explicit URL-keyed registry supporting targeted eviction ([MED-04]).
- The `object.__setattr__` mutation of a frozen `Settings` model in `normalize_provider_models` is replaced with `model_copy(update=...)` or a computed property ([Pydantic V2 §3]).
- `models.py` timestamp columns that are expected to track modification gain `onupdate`, so `updated_at` and `date_updated` stop being stale on paths that do not set them by hand ([SQLModel §3]).
- The exception swallowed to `None` in an ORM model property is surfaced ([MED-08]).
### 7. UI Boundary and Duplication
- The three `ui.instructions.md` violations are corrected ([HIGH-07]):
- `jobs_page.py` no longer imports `session_scope` or manages transactions; a service or workflow method owns the session.
- `sources_page.py` no longer imports `sqlalchemy.inspect`; the service returns a plain `transport_body_deferred` flag on a read model.
- `document_panzoom` no longer calls `get_settings()`; a ready media URL is passed in.
- The duplication catalogued in the review's §4 is extracted, highest value first: `confirm_delete`, `media_urls`, `guards`, `formatters`, `upload_panel`, and the hand-rolled tables that should use `build_table` (~500 lines).
- The 23KB inline SVG moves to `ui/static/` and is loaded through an `importlib.resources` reader alongside the existing `read_css` ([MED-09]).
- `people_page.py:504` routes its error through `error_presenter.show_error` like every sibling handler ([LOW-07]).
- Untyped handler parameters and loosely-typed dict returns are annotated ([LOW-05]).
- The auto-refresh timer is cancelled rather than only deactivated, and its interval becomes a named constant ([LOW-06]).
### 8. Type Checking and Tooling Gate
- The codebase standardizes on `ty`. Remaining suppressions are converted from `# pyright: ignore[...]` to `# ty: ignore[...]` ([HIGH-06]).
- The `lazy="raise"` flip in §2 is expected to eliminate most of the ~160 `selectinload` diagnostics by removing redundant eager loads.
- `ty check` reaches zero diagnostics and is wired into the existing pre-commit setup as a gate ([HIGH-06]).
- The two real bugs currently hidden in the diagnostic noise are fixed: `tests/ui/test_sources_page.py:25` constructs `Source(...)` without the required `document_id`, and `tools/run_destructive_tests.py:76,80` uses `fcntl`, which does not exist on the Windows development platform ([HIGH-06]).
- `ruff check` reaches zero errors ([LOW-01]).
- `asyncio_default_fixture_loop_scope` is configured explicitly so pytest-asyncio behavior does not change on upgrade ([Testing §3]).
- The stale path in `.github/instructions/services.instructions.md:10` is corrected to `src/transcription/db/models.py` ([LOW-02]).
- `list_jobs` stops accepting and discarding `load_docs` ([LOW-03]).
- `resolve_worker_notifier` validates its `getattr` result ([LOW-04]).
## Out of Scope
- Any new user-facing feature, page, action, or field.
- PostgreSQL enablement, Postgres-backed CI, or a Postgres cutover. The `use_alter` fix unblocks it; it does not perform it.
- Alembic or any migration framework, revision directory, history table, or down path.
- Multi-worker or multi-process execution. Forward-compatible code paths are written but not enabled or exercised.
- Concurrency limits, backpressure, or parallel job processing. Jobs remain strictly serial.
- **Splitting `SourceService` into per-model services and relocating `update_job_source_transcription` to `workflows.py` ([MED-14]). Deferred to V4.7.** It touches the transcription write path and cannot safely share a release with the schema re-level.
- Any change to transcription prompt content, medium markers, quality-warning rules, or the retranscription workflow established in V4.5.
- Any change to the evidence, provenance, or immutability contracts established in V4.2 through V4.5.
- Deleting, rewriting, or reinterpreting existing `ExecutionAttempt` or `ProcessingArtifact` evidence during data migration.
- Rewriting the UI table architecture, theme system, or CSS conventions beyond removing duplication.
- Performance work not traceable to a review finding.
## Locked Design Decisions
### A. Remediation Only
Every change traces to a review finding ID. A desirable improvement discovered during implementation that has no finding ID is recorded for a later revision rather than absorbed.
### B. Re-Level, Do Not Migrate
The schema is pre-production and the data is disposable and backed up. Deleting the hand-rolled upgrade chain and regenerating from metadata is correct precisely because this window will not exist again. A migration framework is the right answer once the schema stabilizes, and V4.6 deliberately does not pretend that moment has arrived.
### C. One Schema Pass
The re-level, the indexes, the `use_alter` fix, and the `lazy="raise"` flip all regenerate the same schema. They land together, are verified together, and are reverted together if verification fails. Partial application is not a valid state.
### D. Data Migration Is Last
The migration script is written against the final schema and the final loading strategy. Writing it earlier guarantees rework and risks it carrying implicit lazy loads that `lazy="raise"` will later reject.
### E. Forward Compatibility Where It Is Cheap
Single-process operation makes the atomic job claim non-urgent, not wrong. Where the correct multi-user implementation costs little more than the single-user one, V4.6 writes the correct one and guards it by dialect. Where it costs substantially more, V4.6 defers it and documents the assumption.
### F. Behavior Is Preserved Exactly
A pure-remediation release that changes observable behavior has failed. The existing test suite is the contract: 264 passing tests must still pass, and any test that must change is treated as evidence that the change is not remediation.
### G. The Instruction Files Are the Standard
Most findings are deviations from rules the project already wrote down. V4.6 restores conformance to `services.instructions.md` and `ui.instructions.md` rather than inventing new conventions — except where a rule is itself wrong, in which case the rule is corrected explicitly.
## Acceptance Criteria
1. `app_state.py`, `services/transcription.py`, the `store.py` aliases, `ServiceBase.queue`, and `db/operations.py:get_next_queued_job` no longer exist, and the full suite passes without them.
2. `upgrade_schema` and the three `_upgrade_*` functions no longer exist; no raw `ALTER TABLE` or `CREATE INDEX` string appears in `src`.
3. A database created from empty by `create_all()` contains the composite `Job` index, indexed hot foreign keys, and a `preferred_execution_attempt_id` column whose type matches the model declaration.
4. Compiling the metadata against the PostgreSQL dialect emits **no** unresolvable-cycle warning.
5. No `Relationship` in `db/models.py` uses `lazy="selectin"` as a bidirectional default; every load path that requires eager loading declares it per query, and the suite passes under `lazy="raise"`.
6. `read_next_queued_job` returns at most one row and issues no eager-load queries; a test asserts the emitted SQL contains `LIMIT`.
7. The worker processes two consecutive jobs against a single provider client instance; a test asserts the client is not reconstructed between jobs.
8. `worker_provider_timeout_seconds` accepts a value above 20 seconds, and the OpenRouter client receives an explicit `httpx.Timeout`.
9. `inspect.signature` no longer appears in the transcription call path.
10. `DocumentType` and `PersonRole` CRUD is served by one shared implementation; the existing registry tests for both pass unchanged.
11. `ServiceBase._get_or_raise` is the only place a `NOT_FOUND` guard is written for an entity fetched by id.
12. One media-storage implementation serves Source files, portraits, and homepage images, and its write is off the event loop.
13. `list_sources_detail` filters by `job_id` in SQL; `read_source_navigation` issues bounded queries; `list_processing_artifacts` accepts a `limit`.
14. No page imports `session_scope`, `sqlalchemy.inspect`, or `get_settings`.
15. The 23KB SVG literal no longer appears in any `.py` file.
16. `ruff check` reports zero errors.
17. `ty check` reports zero diagnostics and runs as a pre-commit gate.
18. `tools/run_destructive_tests.py` runs on Windows.
19. All 264 pre-existing tests still pass. Any test modified during V4.6 is individually justified as a test defect rather than a behavior change.
20. The migration script restores the backed-up V4.5 data into the re-leveled schema with row counts matching the backup, and no on-disk Source file, portrait, or artifact is modified.
21. No new user-facing feature, page, action, or field exists in V4.6 that did not exist in V4.5.
22. Database, integration, and UI verification uses isolated test data and never modifies `data/transcription.db`.
## Scope Freeze Gate
V4.6 is sufficiently frozen to begin implementation:
- The operating context — SQLite, single process, disposable data — is confirmed and its consequences for severity are resolved.
- The schema strategy is resolved: re-level, no Alembic, one pass, migration last.
- The severity band is resolved: Critical through Low, inclusive.
- The service-layer consolidation set is resolved, and the `SourceService` split is explicitly deferred to V4.7.
- The release character is resolved: pure remediation, no new features.
Any expansion into PostgreSQL enablement, multi-worker execution, a migration framework, the `SourceService` split, or any new feature requires an explicit V4.6 scope amendment or a later revision.
## Related Local References
- [V4.6 Implementation Plan](implementation_plan_v4_6.md)
- [Architecture & Code Review Report](../architecture_code_review_2026-08-17.md)
- [V4.5 Scope Boundary](../ver4.5/scope_boundary_v4_5.md)
- [V4.5 Implementation Plan](../ver4.5/implementation_plan_v4_5.md)
- [V4.2 Evidence and Provenance Scope](../ver4.2/scope_boundary_v4_2.md)
- [V4 Architecture](../ver4/architecture_v4.md)
- [V4 Schema](../ver4/schema_v4.md)
- [Transcription Methodology](../invariant/transcription_methodology.md)
- [AI Evidence and Provenance Invariant](../invariant/ai_evidence_and_provenance.md)
-208
View File
@@ -1,208 +0,0 @@
# Implementation Plan (Version 4.7)
## Goal
Collapse the duplicated evidence model, remove the `ProcessingArtifact` subsystem and move orientation normalization to ingest, complete the `SourceService` decomposition deferred from V4.6 ([MED-14]), correct the run-time measurement window, and close the remaining correctness and tooling items opened during V4.6. No new user-facing behavior.
## Planning Status
**Frozen.** The boundary is [`scope_boundary_v4_7.md`](scope_boundary_v4_7.md). Feature work is parked in [`../ver4.8/feature_backlog_v4_8.md`](../ver4.8/feature_backlog_v4_8.md).
## Planning Constraints
- Every change traces to a finding ID or a V4.6 review-log entry.
- `ruff check` clean and `ty check` at **0 diagnostics** at the end of every phase, matching the V4.6 exit state.
- The full suite passes at the end of every phase.
- **Test changes are expected in Phases 1 and 2.** This differs from V4.6, where the mechanical moves required no test logic changes. Measured blast radius: 33 references to the removed `job_source` evidence fields across 8 test files, and 10 `ProcessingArtifact` references across 3. Only Phase 4 retains the "no test logic changes" rule.
- Use `.\.venv\Scripts\python.exe -m pytest` (the system interpreter has no packages).
- Back up `data/transcription.db` **and** `data/documents/` before running any migration step. The image backfill rewrites files in place.
- One phase, one commit.
## Expected Project Impact
| Area | Before | After |
| :--- | :--- | :--- |
| `job_source` columns | 9 | **4** - `id`, `job_id`, `source_id`, `status` |
| `JobSourceStatus` on disk | two spellings across two tables | one spelling, plus a new `CANCELLED` member |
| `processing_artifact` | table, model, ~283 lines of service code, 2 rows | removed |
| Orientation normalization | derived per transcription, artifact-backed | applied once at ingest, no derivative |
| Stored image rotation | `quality=95, subsampling=0`, +38% size | qtables reuse, ~6% smaller, higher PSNR |
| `services/sources.py` | 1,389 lines, 4 domain models | ~900 lines, `Source` + `JobSource` |
| `services/evidence.py` | does not exist | ~174 lines, `ExecutionAttempt` reads and export |
| `services/artifacts.py` | planned | **cancelled** - deleted rather than extracted |
| `duration_ms` | provider call + normalization + artifact write + commit | the operation the timeout governs |
| Worker loop errors | every `Exception` logged and suppressed | programming errors distinguishable from provider faults |
| Quality gate | local pre-commit only, inert until installed | enforced in CI |
## Migration Handling
All schema and data changes are delivered by a single idempotent `tools/migrate_v46_to_v47.py`, following the `tools/migrate_v45_to_v46.py` conventions: never invoked at startup, never run by the test suite.
The tool is **built incrementally** - Phase 1 creates it with its own step, Phase 2 appends the next - and is **run at the end of each of those phases** so the live database stays usable at every phase boundary. Idempotency is what makes re-running safe.
Steps, in execution order:
1. Rotate the 58 stored images carrying EXIF orientation 3, in place, using qtables reuse; strip the orientation tag.
2. Drop the `processing_artifact` table and remove its external artifact files.
3. Normalize `execution_attempt.status` to the single declared spelling ([45]).
4. Drop `job_source.raw_transcription`, `ai_metadata`, `raw_api_response`, `executed_at`, `error_detail`.
`tools/migrate_v45_to_v46.py` deliberately accepts both enum spellings because it reads historical backups. Leave that tolerance in place.
## Implementation Phases
### 1. Ingest Normalization and ProcessingArtifact Removal
Deletion comes first, so that Phase 4 never restructures code that is on its way out.
Tasks:
1. Move orientation normalization into the ingest path in `media_storage`, ahead of `write_bytes` (`media_storage.py:57`). Rotate, strip the EXIF orientation tag, then store.
2. Change the encode settings at `normalization.py:84-85` from `quality=95, subsampling=0` to `qtables=im.quantization`, `subsampling=JpegImagePlugin.get_sampling(im)`, `optimize=True`. Import `JpegImagePlugin` explicitly - it is not reachable as an attribute of `PIL.Image`.
3. Delete `resolve_provider_input` and its call at `workflows.py:226`. The stored file is now already upright, so the transcription path reads it directly.
4. Fold the `transcription_quality_warnings` payload (`workflows.py:560-571`) into `execution_attempt.normalized_metadata`.
5. Delete the artifact cluster from `sources.py` (lines 732-1015) and the artifact branch of `build_evidence_export`.
6. Delete the `ProcessingArtifact` model and the `CheckConstraint`.
7. Remove the "Orientation normalized" badge at `sources_page.py:174-178`, the artifact evidence dump at line 417, and the quality-warnings render at line 661.
8. Update `test_normalization.py`, `test_v42_evidence.py`, and `test_db.py`. Normalization tests should now assert on ingest behavior rather than on artifact creation.
9. Create `tools/migrate_v46_to_v47.py` with steps 1 and 2. Back up, run, verify.
Verification: re-check EXIF orientation across `data/documents/` - no stored image should report orientation 3, 6, or 8. Spot-check one backfilled page visually.
Exit: suite green, `ty check` at 0, `processing_artifact` gone from schema and code.
### 2. Evidence Model Simplification
Tasks:
1. Add `JobSourceStatus.CANCELLED`. Update `jobs.py:378-384` to write it instead of `FAILED` plus `"Cancelled by user"`.
2. Declare one spelling for `JobSourceStatus` across both `job_source.status` and `execution_attempt.status` ([45]). `job_source.status` already declares `values_callable`; `execution_attempt.status` does not.
3. Remove `raw_transcription`, `ai_metadata`, `raw_api_response`, `executed_at`, and `error_detail` from the `JobSource` model.
4. Redirect every read to `execution_attempt`:
- `transcript.py:103-119` sorts by `executed_at` - sort by `ExecutionAttempt.finished_at`.
- `sources_page.py:390-393` reads `ai_metadata` and `raw_api_response`.
- `sources_page.py:337-375` renders status, executed time, and error detail.
- `models.py:266-278` and `models.py:328-343` derive transcript and error from `job_sources`.
5. Shrink `update_job_source_transcription` (`sources.py:524-679`) to write only the surviving `JobSource` columns. Keep the method in `sources.py` and keep both writes in one session scope.
6. Confirm `jobs.py:411-424` retry still works unchanged. It resets `FAILED` to `PENDING`; the attempt history it appears to discard is preserved by `ExecutionAttempt`'s unique constraint.
7. Confirm `workflows.py:432-442` work selection is unaffected. It filters `status != TRANSCRIBED` within `job.job_sources`, so `CANCELLED` pages are excluded from a re-run only if that is the intent - **decide explicitly** whether cancelled pages should be re-attempted, and encode the answer in the filter rather than leaving it implicit.
8. Update the 33 affected test references across the 8 files identified.
9. Append migration steps 3 and 4. Back up, run, verify row counts before and after.
Exit: suite green, `ty check` at 0, `job_source` at 4 columns.
### 3. Stage B - Extract `services/evidence.py` ([MED-14])
Read-side only, now applied to a substantially smaller `sources.py`.
Move:
- `read_latest_execution_attempt` (216-245), including the `LatestExecutionAttempt` read model
- `promote_machine_attempt` (679-710)
- `list_execution_attempts` (710-732)
- `build_evidence_export` (1015-1107)
Tasks:
1. Create `services/evidence.py` with an `EvidenceService(ServiceBase)` following the `DocumentService` conventions.
2. Move the methods and the `LatestExecutionAttempt` dataclass verbatim. Preserve signatures, keyword-only arguments, error types, and `_session_scope` usage exactly.
3. Preserve every explicit `selectinload()` chain. **Chain, never varargs** - `selectinload(A.b).selectinload(B.c)` and `selectinload(A.b, B.c)` produce an identical `.path` but are not equivalent, and under the `lazy="raise"` default set in V4.6 the varargs form raises at render time. See `db/loading.py`.
4. Update `ServiceBundle` to construct and expose the new service via the `from_session_factory` constructor added in V4.6.
5. Update call sites in `sources_page.py` and `workflows.py`.
6. Run `ruff check --fix` **in the same pass** as the import edits - autofix removes imports that are unused at that moment.
7. **Revise `.github/instructions/services.instructions.md` to describe the boundaries this decomposition actually produced** (review log [59]). Do this *after* the move, not before - the refactor is the empirical test of the rule, and a rule written in advance would have to be bent to fit. Known defects to correct:
- **Line 11, `1 service class per data model`** - the rule is table-shaped rather than aggregate-shaped, and is the measured cause of `sources.py` reaching 1,389 lines. Replace with aggregate ownership.
- **No home for junction tables.** The rule names the four core components (Document, Source, Job, Person) but is silent on `job_source` and `document_person`, where they intersect. Add an explicit model-ownership table naming the owning service for every model, including junctions and `ExecutionAttempt`.
- **Lines 30-32**, mandatory CRUD for every model, is already false: `prompts.py` does not comply. Soften to describe intent rather than mandate a method set.
- **Line 13 vs lines 75-77** read as contradictory on whether a service may touch more than one table. Reword the composition section so the ownership rule and the multi-table-operation guidance agree.
- **Line 77** typo: `picutre`.
Note that line numbers above are pre-Phase-1 positions and will have shifted. Locate by symbol, not by line.
Exit: suite green **with no test logic changes** beyond import paths, `ty check` at 0, and `services.instructions.md` consistent with the post-refactor module layout. Walk every `/ui/*` page and confirm a 200, since `lazy="raise"` turns a missed eager load into a runtime error rather than a slow query.
### 4. Run-Time Measurement Window (review log [55])
Tasks:
1. In `workflows.py`, make the recorded duration cover only the operation the `wait_for` at lines 240-249 governs. Either move `monotonic_started_at` (line 221) to immediately before the `wait_for`, or capture provider latency separately and record that.
2. Apply the same treatment to all three write sites: success (line 278), `TimeoutError` (line 295), and general failure (line 330). The failure paths must keep using the monotonic clock.
3. If preprocessing time is still worth keeping, record it as its own value rather than folding it into `duration_ms`.
4. Update `sources_page.py:400`, which renders the raw integer as `"27612 ms"`.
Phase 1 already removes normalization and the artifact write from this window, which narrows the gap but does not close it - the `session.commit()` at line 228 remains inside it.
Verification: a recorded timeout duration should sit at or just under the configured budget, not 0.4-2.0 s above it as in the three historical `local_timeout` rows.
### 5. Worker Exception Handling (review log [8])
`worker.py:96-106`.
Tasks:
1. Separate genuinely retriable faults from programming errors. `classify_unexpected_error` is already called at line 101 and its result is currently only logged.
2. Ensure a non-retriable error reaches a terminal state instead of being retried.
3. Keep terminal-state and retry persistence atomic per `services.instructions.md:63-65`.
4. Add a test that a deliberate programming error in the loop does not silently retry.
Context: with `WORKER_MAX_RETRIES=1` and a 30 s timeout the worst-case silent burn is 60 s, down from 360 s, so this is no longer urgent - but it remains the real fix behind that risk.
### 6. CI Enforcement ([HIGH-06], review log [40])
Tasks:
1. Add a workflow under `.github/workflows/` running `ruff check`, `ty check`, and `pytest` on push and pull request.
2. Use the same commands as `.pre-commit-config.yaml` so local and CI gates cannot drift.
3. Confirm the 4 tests that skip without `OPENROUTER_API_KEY` skip cleanly in CI rather than failing.
4. Negative-test the workflow by pushing a deliberate lint error on a scratch branch.
## Sequencing Constraints
- **Phase 1 before Phase 3.** Code scheduled for deletion is never extracted first. This is why the previously planned `services/artifacts.py` is cancelled.
- **Phase 1 before Phase 2.** Both touch `workflows.py` write paths; separating them keeps any regression attributable.
- **Phase 4 before any V4.8 telemetry work.** A model-performance rollup built on the current `duration_ms` would chart preprocessing mixed with provider latency.
## Test Strategy
- **Phase 1:** normalization tests move from asserting artifact creation to asserting ingest behavior. Add a test that a stored image never retains EXIF orientation 3, 6, or 8.
- **Phase 2:** assert that evidence reads resolve through `execution_attempt`; assert `CANCELLED` is distinguishable from `FAILED`; verify both status columns round-trip identically and existing rows read back correctly after the fix-up.
- **Phase 3:** the suite passes with **no test logic changes**. Only import paths update. A required behavioral change signals the move was not mechanical - stop and re-examine.
- **Phase 4:** assert the recorded duration is bounded by the configured timeout.
- **Phase 5:** new test that a programming error does not silently retry.
- **Phase 6:** CI must fail on an injected lint error.
## Risks
| Risk | Mitigation |
| :--- | :--- |
| The image backfill corrupts originals - it rewrites files in place with no archival master | Back up `data/documents/` before running; idempotent step keyed on the EXIF tag so a second run is a no-op; visually spot-check a backfilled page |
| Dropping `job_source` columns loses data that turns out not to be duplicated | Verified 77/77 identical on all three evidence columns before dropping; re-run that comparison inside the migration and abort on any mismatch |
| A read still expects a removed `job_source` column and fails only at render time | Grep-driven checklist in Phase 2 task 4; walk every `/ui/*` page after the phase |
| Cancelled pages are silently re-attempted, or silently never re-attempted | Phase 2 task 7 forces an explicit decision in the work-selection filter |
| A moved query loses an eager load and trips `lazy="raise"` at render time | Preserve `selectinload` chains verbatim; walk every `/ui/*` page after Phase 3 |
| `ruff check --fix` deletes an import mid-move | Edit imports and usages in the same pass, as in V4.6 |
| Circular imports between `sources.py` and `evidence.py` | Composition is sanctioned by `services.instructions.md:75-77`; keep the dependency one-directional |
| Scope creep into refactoring `update_job_source_transcription` | Out of scope; the boundary records why |
## Delivery Order
1. Ingest normalization and `ProcessingArtifact` removal
2. Evidence model simplification
3. Stage B - `evidence.py`, then revise `services.instructions.md` to match
4. Measurement window
5. Worker exception handling
6. CI enforcement
## Done Criteria
Every item in the scope boundary's Acceptance Criteria is satisfied, the suite is green, `ty check` reports 0 diagnostics, and no user-facing behavior has changed.
## Related Local References
- [V4.7 Scope Boundary](scope_boundary_v4_7.md)
- [Architecture & Code Review Report](../architecture_code_review_2026-08-17.md)
- [V4.6 Review Log](../ver4.6/review_log_v4_6.md) - resolves the `review log [N]` citations used throughout this document
- [V4.6 Implementation Plan](../ver4.6/implementation_plan_v4_6.md)
- [V4.8 Feature Backlog](../ver4.8/feature_backlog_v4_8.md)
- `.github/instructions/services.instructions.md`
- `src/transcription/db/loading.py` - the `selectinload` varargs trap
-415
View File
@@ -1,415 +0,0 @@
# V4.7 Implementation Review Log
Working record kept during the V4.7 architectural cleanup release.
This file is the canonical reference for citations of the form **`review log [N]`** in V4.7 and later planning documents. The numbers below are those `N` values. They are independent of the [V4.6 log](../ver4.6/review_log_v4_6.md), which has its own numbering.
The log was maintained live in a session-scoped database and exported here so the citations remain resolvable in later sessions. It is a historical record: entries are not rewritten after the fact, except where a later phase resolved an entry that was open at the time, in which case the resolution is appended to the body and marked `RESOLVED:`. Where an entry conflicts with a committed planning document, **the planning document wins**.
## Legend
| Field | Meaning |
| :--- | :--- |
| `kind` | `question` - needed a decision; `comment` - observation; `deviation` - departure from plan; `risk` - identified hazard |
| `status` | `open` - unresolved; `answered` - resolved by a decision; `noted` - recorded, no action required |
| `finding` | Finding ID in [architecture_code_review_2026-08-17.md](../architecture_code_review_2026-08-17.md), where one applies. Most V4.7 entries have none, because V4.7 works from the [implementation plan](implementation_plan_v4_7.md) rather than from that pre-V4.6 snapshot. Where an entry carries a one-line summary instead, it appears as a bold lead-in to the body. |
**50 entries** - 1 open, 18 answered, 31 noted.
## Still Open
These carry forward past V4.7.
| ID | Finding | Summary | Disposition |
| :--- | :--- | :--- | :--- |
| [32] | - | /ui/documents/{id}/sources redirects to /sources, dropping the /ui prefix | Pre-existing and outside the V4.7 scope boundary - deliberately left unfixed |
## Full Log
### Phase 0 - baseline and backups
#### [1] Backups taken and verified
*comment* - **noted**
data/transcription.db and data/documents/ copied to C:\GitHub\_backups\transcription_v47_20260818-092616. All 76 document files SHA256-identical to source. A consistent SQLite snapshot (transcription.consistent.db) was also produced via the sqlite3 backup API because the live DB file is locked by a running app process, making a plain file copy potentially torn.
#### [2] The application appears to be running and holds data/transcription.db
*risk* - **answered**
Two python processes started 2026-08-18 04:58 (PID 14340 is .venv python). Get-FileHash on data/transcription.db failed with a sharing violation. The Phase 1 migration rewrites data/documents/ JPEGs in place and later phases ALTER the live schema; both are unsafe while the app is running. Need the user to stop the app before any migration run. RESOLVED: the user stopped the app before the migration ran; the DB accepted an immediate write lock beforehand.
#### [3] Working tree is clean; the V4.7 doc edits are already committed
*deviation* - **noted**
The brief said both V4.7 docs have uncommitted edits on disk. git status --short is empty and HEAD is 246d7f9 "V4.7 final scope changes", which contains them. Nothing was reverted or stashed; the working tree content matches what the brief described.
#### [4] Baseline is not fully green: 1 pre-existing test failure caused by .env leakage
*risk* - **answered**
tests/test_config.py::TestWorkerReliabilitySettings::test_worker_retry_defaults asserts worker_max_retries == 0 but the local .env sets WORKER_MAX_RETRIES=1, and _make_settings() reads .env. ruff and ty are both clean. This is a test-isolation defect predating V4.7 (introduced when the 2026-08-18 calibration wrote WORKER_MAX_RETRIES into .env), not a code defect. It would pass in CI, where no .env exists, which makes it a latent local/CI divergence relevant to Phase 6. Awaiting a decision on whether to fix it. RESOLVED: tests/test_config.py::_make_settings now builds Settings with _env_file=None, isolating it from the local .env. The local/CI divergence this entry flagged is gone, and CI run 182 confirms the same 295 passed / 4 skipped result on Linux with no .env of the developer's.
### Phase 1 - artifact removal and ingest rotation
#### [5] Where does ingest-time rotation actually belong: media_storage or store.py?
*question* - **answered**
Plan Phase 1 task 1 says rotate inside media_storage ahead of write_bytes. But store.py computes file_hash and file_size_bytes from the ORIGINAL file_bytes (store.py:93, 168) independently of write_media_bytes, which returns only a Path. Rotating inside write_media_bytes would leave Source.file_hash and file_size_bytes describing bytes that were never stored. write_media_bytes is also shared with person portraits and homepage images. Recommend rotating the bytes once at the Source-ingest boundary (store_source_file or its two callers in store.py) so hash, size and stored file all describe the same upright bytes, and leaving media_storage a generic byte writer. [ANSWERED 2026-08-18 by user] Rotate at the Source-ingest boundary, before file_hash/file_size_bytes are computed, so the hash and size describe the stored upright bytes. media_storage stays a generic byte writer. This is a deliberate deviation from Phase 1 task 1 as written.
#### [6] Fate of artifact_dir / artifact_inline_threshold_bytes settings and data/artifacts/
*question* - **answered**
Removing ProcessingArtifact orphans Settings.artifact_dir and Settings.artifact_inline_threshold_bytes (config.py:104,106), the JobService artifact deletion path (jobs.py:307-350), and the on-disk data/artifacts/ tree. config.py is outside the services instruction file. Proposal: delete both settings and the jobs.py deletion path as part of the same removal, and have the migration delete the external artifact files (migration step 2 already says so). Confirm. [ANSWERED 2026-08-18 by user] Delete both Settings.artifact_dir and Settings.artifact_inline_threshold_bytes.
#### [7] Phase 1 blast radius is wider than the plan task list
*comment* - **noted**
Beyond the listed sites, ProcessingArtifact is also referenced by: sources.py delete guards (lines 296-300 and 450-462, which block Source deletion when artifacts exist), jobs.py job-deletion artifact cleanup (307-350), Source.processing_artifacts and ExecutionAttempt.artifacts relationships (models.py:322,431), the selectinload(ExecutionAttempt.artifacts) in list_execution_attempts (sources.py:719), and the ProcessingArtifact lookup that validates model_input_artifact_id inside update_job_source_transcription (sources.py:648-664). All must go with the model.
#### [10] Quality-warnings render: delete (per plan) or re-point at normalized_metadata?
*question* - **answered**
Plan Phase 1 task 7 says delete the render at sources_page.py:661, but task 4 folds the same payload into execution_attempt.normalized_metadata and decision A forbids user-facing change. The render has never fired in practice because it reads attempt.artifacts and only 2 artifact rows exist. [ANSWERED 2026-08-18 by user] Keep the display and re-point it at normalized_metadata. Deviation from Phase 1 task 7 as written; task 4 now has a consumer.
#### [11] SourceEvidenceReference.derivative_id / transformation kept but no longer populated
*deviation* - **noted**
With ProcessingArtifact gone there is no derivative to reference, so both fields are always None. They were left in place rather than removed: RequestManifest is a frozen, versioned evidence contract whose canonical bytes feed request_manifest_sha256, so removing fields would change the digest of every future manifest and arguably require a schema_version bump - cost out of proportion to deleting two optional fields. Raise if you would rather see the contract cleaned up.
#### [12] The session.commit() at workflows.py:228 is removed with resolve_provider_input
*deviation* - **noted**
That commit existed to make the artifact row written during provider-input resolution durable before the provider call. With no artifact write there is nothing pending to commit - the PROCESSING claim was already committed at line 198 / 426 - so the call is removed rather than left as a no-op. This also removes one of the two things Phase 4 has to get out of the duration measurement window.
#### [13] The image backfill must also update source.file_hash and file_size_bytes
*risk* - **answered**
Migration step 1 as written only rotates the stored JPEGs and strips the EXIF tag. But source.file_hash and source.file_size_bytes were computed from the pre-rotation bytes, and after Phase 1 the transcription path derives the evidence digest (SourceEvidenceReference.digest_sha256) straight from source.file_hash. Rotating the file without updating the row would make every backfilled Source advertise a digest that does not match the bytes actually sent to the provider - the exact class of defect the evidence model exists to prevent. The migration therefore rewrites both columns for each rotated image in the same transaction. Not a change of intent, an omission in the step description.
#### [14] Orientation normalization must not change what ingest accepts
*deviation* - **noted**
**Undecodable upload bytes are a normalization no-op, not a rejection**
validate_source_content only checks emptiness and filename; it never decoded the image, so bytes that Pillow cannot open (e.g. the b"image-bytes" fixture in tests/services/test_store.py) were accepted and stored. Moving rotation into store_source_file initially turned that into an OrientationNormalizationError, i.e. a user-facing rejection of previously accepted uploads. Decision A forbids user-facing change, so Image.open failure now logs and returns None; the error is retained only for a decode that succeeded and a rewrite that then failed.
#### [15] Artifact-only tests removed with the subsystem
*deviation* - **noted**
**Two tests deleted rather than rewritten**
tests/test_v42_evidence.py::test_large_json_artifact_uses_constrained_atomic_storage and ::test_rejects_inline_artifact_with_incorrect_integrity exercised only external artifact storage and inline artifact integrity. Both behaviours are deleted by Phase 1, so the tests have no surviving subject. tests/ui/test_sources_page.py lost one assertion ("Derived Artifacts"), and tests/test_db.py lost the processing_artifact table assertion.
#### [16] Fate of the superseded v4.5 to v4.6 migration tool
*question* - **answered**
**tools/migrate_v45_to_v46.py no longer type-checks**
The old migration references ProcessingArtifact (line 157), SourceService._verify_artifacts_integrity (line 169), and carries processing_artifact: 2 in EXPECTED_SOURCE_COUNTS (line 83). All three are gone. It is currently the only remaining ty failure. The plan says to keep its enum-spelling tolerance but does not address this. Options: delete the completed one-time tool; or strip the artifact code path from it. RESOLVED: the completed one-time tool was deleted (Phase 1). tools/ now contains only migrate_v46_to_v47.py, and ty is clean.
#### [17] tools/migrate_v45_to_v46.py removed rather than repaired
*question* - **answered**
**Superseded v4.5 to v4.6 migration deleted**
User decision. The migration is complete, the live database is already V4.6, and after V4.7 it would restore a V4.5 backup into a schema that no longer matches (job_source is stripped in Phase 2). Two doc references remain, both citing it only as a conventions template: implementation_plan_v4_7.md lines 39 and 50, scope_boundary_v4_7.md line 160. Line 50 (enum-spelling tolerance) is now moot. Recoverable from git history if ever needed.
#### [18] DEFAULT_ARTIFACT_DIR constant in migrate_v46_to_v47.py
*deviation* - **noted**
**Migration records the deleted artifact_dir default itself**
Step 2 must delete external artifact files, but Settings.artifact_dir was deleted in the same phase. The migration therefore carries the historical V4.6 default (data/artifacts) as its own constant with a --artifact-dir override, rather than depending on a setting that no longer exists. One external file was present and removed; the directory is now empty.
#### [19] Phase 1 migration outcome
*risk* - **answered**
**Migration executed and verified against the live corpus**
Ran after the user stopped the app and after a fresh pre-migration backup to C:\GitHub\_backups\transcription_v47_premigration_20260818-101232. Result: 58 images rotated, 18 already upright, 0 missing; processing_artifact dropped (2 rows) and its 1 external file removed. Verification: no stored image reports orientation 3/6/8; source.file_hash and file_size_bytes match every file on disk (0 mismatches over 76); 58 of 76 files differ from the backup; PSNR against the un-rotated backup is 50.3 / 51.1 / 56.1 dB (min/median/max) across the 57 JPEGs, allowing for the -6 percent size reduction; a re-run reports rotated=0 and table already absent, confirming idempotency. Visual spot-check of 1547e555 confirmed the page was genuinely stored upside down and is now upright.
### Phase 2 - evidence model simplification
#### [8] Should CANCELLED pages be re-attempted when a job is re-run?
*question* - **answered**
Plan Phase 2 task 7. _resolve_job_sources (workflows.py:432-442) selects work by status != TRANSCRIBED, so once CANCELLED exists as a distinct status a re-run would silently pick cancelled pages back up. Options: (a) exclude CANCELLED from work selection, so cancelling is sticky and a page must be explicitly re-queued; (b) include it, so re-running a job means "do everything not yet transcribed"; (c) clear CANCELLED back to PENDING in the existing retry path (jobs.py:411-424) and exclude it from work selection, which makes re-attempt an explicit user action through the retry button. Decision required before Phase 2 task 1. RESOLVED: re-attempt them. Resubmit accepts FAILED and CANCELLED. Rationale: today cancel writes FAILED, so resubmit already resets cancelled pages to PENDING; introducing a distinct CANCELLED status without widening the resubmit filter would silently make cancelled work unrecoverable, a user-facing regression that decision A forbids. The decision is encoded in the resubmit candidate filter, which is the real decision point - _resolve_job_sources only ever sees these rows after resubmit has already set PENDING. UI copy on both the cancel and resubmit pages is updated to match.
#### [20] Plan task 4 targets a dead module
*question* - **answered**
**ui/components/transcript.py deleted instead of redirected**
Phase 2 task 4 directs transcript.py:103-119 to sort by ExecutionAttempt.finished_at instead of job_source.executed_at. Investigation showed the module is entirely unreferenced: no import of transcription.ui.components.transcript exists in src, tests, or docs, and both public functions (render_original_transcription_card, render_revision_row) have zero callers. Rewriting it would mean maintaining unreachable code against the new evidence model. User decision: delete the module. Recoverable from git history.
#### [21] Defect [45] fixed by declaring one enum spelling
*deviation* - **noted**
**execution_attempt.status gains values_callable**
ExecutionAttempt.status was a bare JobSourceStatus annotation, so SQLAlchemy persisted enum names (TRANSCRIBED) while job_source.status persisted values (transcribed) via values_callable. That is why the two columns matched on 0 of 79 rows. execution_attempt.status now declares the identical SAEnum with values_callable and native_enum=False. Existing rows carry the old spelling and are rewritten by migration step 3.
#### [22] Dead property made more expensive by the evidence move
*question* - **answered**
**Job.error_detail deleted rather than re-derived**
Plan task 4 lists models.py:266-278 (Job.error_detail) for redirection. A full-repo search found zero readers: JobTableRow has no such field and the job detail page never calls it. Re-deriving it from ExecutionAttempt would require a two-level eager load (job_sources -> execution_attempts) on every Job, across a lazy=raise then lazy=noload chain, where a missing load returns an empty list and the property would silently answer None instead of raising. No information is lost: error_detail survives on ExecutionAttempt and is reachable via list_execution_attempts and read_latest_execution_attempt. A future job-level failure view should query attempts directly anyway, since first-error-across-pages is the wrong shape for a partial-success job. User decision: delete.
#### [23] Replacement ordering key after executed_at is dropped
*question* - **answered**
**Source.latest_job_source orders by Job.date_created**
JobSource retains only id, job_id, source_id and status, so max(job_sources, key=executed_at) needs a key from a neighbour. Job.date_created is chosen over the latest ExecutionAttempt.finished_at: it is always present (a PENDING page has no attempt at all), it is already eager-loaded by read_source_detail, and since (job_id, source_id) is unique per source the ordering is exactly most recent job. The two differ only when a job created earlier finishes later, which the single-worker queue does not produce. User decision.
#### [24] The "Cancelled by user" string has no home after job_source is stripped
*deviation* - **noted**
**Cancel no longer records a reason string**
cancel_job previously wrote error_detail="Cancelled by user" onto job_source. That column is gone, and cancel deliberately makes no provider call so it writes no ExecutionAttempt. The reason is now carried by JobSourceStatus.CANCELLED itself, which is strictly more precise than a free-text string. UI copy on the cancel page was updated to say "cancelled" and to state that cancelled sources can be resubmitted.
#### [25] jobs_page "Failed Sources" became "Resubmittable Sources"
*deviation* - **noted**
**Resubmit UI counter renamed**
The resubmit candidate filter now accepts FAILED and CANCELLED per the user decision in entry 8, so the page counter had to count both. Renamed the metadata row and the blocked-error message accordingly.
#### [26] sources_page no longer renders ai_metadata/raw_api_response when no attempt exists
*comment* - **noted**
**Legacy job_source evidence fallback deleted from the detail page**
The "no ExecutionAttempt" branch of _render_provider_evidence used to fall back to the job_source JSON columns for historical rows. Those columns are gone, so the branch now renders only the empty state. Verified against the evidence baseline: all 77 successful transcriptions have a matching execution_attempt row, so no live row loses its evidence display.
#### [27] latest_error_detail reads through job_sources -> execution_attempts
*risk* - **noted**
**Model properties now require a two-level eager load**
Source.latest_error_detail feeds a visible "Error Detail" column on the sources table. Because JobSource.execution_attempts is lazy="noload" it returns empty rather than raising when not loaded, so a caller that forgets the chained selectinload gets a silent blank instead of an error. list_sources_detail and the model-property test were both updated to chain selectinload(...).selectinload(orm_attribute(...)). Any new caller must do the same.
#### [28] job_source.status and execution_attempt.status now agree on every row
*comment* - **noted**
**Defect [45] verified fixed against the live database**
Before: 0/79 rows matched, because execution_attempt persisted enum names and job_source persisted values. After migration step 3: 79/80 join rows agree. The single disagreement is job_source 09cd5f77 which has two attempts - attempt 1 failed, attempt 2 transcribed - so the queue row correctly reflects the final outcome while the history preserves the failure. Comparing job_source against its LATEST attempt gives 79/79.
#### [29] list_sources_detail resolves latest_status and latest_error_detail for all 76 rows
*comment* - **noted**
**Two-level eager load verified against live data, not just tests**
Ran SourceService.list_sources_detail against the migrated production database: 76 sources, 75 transcribed / 1 failed, and the one failed row still exposes latest_error_detail - now read from execution_attempt rather than the dropped job_source column. This closes the silent-blank risk recorded in entry 27 for the shipped call path.
### Phase 3 - evidence service extraction and the ownership rule
#### [9] Junction ownership: which service owns job_source and document_person?
*question* - `MED-14` - **answered**
services.instructions.md names four core components (Document, Source, Job, Person) and is silent on the two junctions, which is exactly where two owners intersect. Candidate tie-break rules: (a) the junction belongs to the service that creates its rows; (b) it belongs to the aggregate whose lifecycle it shares (job_source dies with the Job, document_person dies with the Document); (c) it belongs to the side that reads it most. These do not agree for job_source: it is created by store.py orchestration, its lifecycle is the Job, and it is read predominantly through Source pages, which is how it ended up in sources.py. Decision required at Phase 3 task 7. RESOLVED: measurement showed document_person has a single writer (people.py, every create/delete/sync) and needs no tie-break; documents.py only eager-loads through it. job_source is genuinely contested between sources.py (row existence + per-page outcome) and jobs.py (job-lifecycle status transitions). User selected the LIFECYCLE rule: the service that creates and deletes rows owns the junction, so job_source -> SourceService. Two scoped carve-outs written into the rule: (1) cascade deletion of junction rows when a service deletes its own aggregate root (JobService.delete_job_with_guardrails); (2) status transitions that create and delete nothing (cancel_job, resubmit_failed_sources), because those are Job lifecycle events. No code was moved.
#### [30] Where should the shared transcription error hierarchy live?
*question* - **answered**
**Extraction immediately violated the existing no-sibling-import rule**
tests/test_service_boundaries.py enforces services.instructions.md:13 - a service module must not import a sibling. evidence.py needed TranscriptionNotFoundError, which sources.py also raises, so the extraction failed the rule on the first run. Measured ownership: CandidatePromotionError is now raised only in evidence.py; PromptLoadError and SourceDeleteBlockedError only in sources.py; TranscriptionNotFoundError in both; TranscriptionError is the shared base, caught by store.py. User chose to move the whole five-class hierarchy to a neutral services/errors.py: one obvious home, one import path, and the exception a caller catches no longer changes when an operation moves between services.
#### [31] Two test bundles broke on adding a fifth service, not on the refactor itself
*comment* - **noted**
**ServiceBundle default factories silently bind to the real database**
test_v45_candidates and test_workflows_reliability constructed ServiceBundle(...) field by field. Adding the evidence field meant it fell back to field(default_factory=EvidenceService), which resolves the process-global session factory rather than the test one - so the tests silently queried the wrong database instead of failing loudly. Both were changed to ServiceBundle.from_session_factory(...), which is immune to future additions. This is the same global-singleton hazard recorded in the 2026-08-17 review at line 272.
#### [32] /ui/documents/{id}/sources redirects to /sources, dropping the /ui prefix
*risk* - **open**
**Pre-existing broken redirect found during the UI walk**
The Phase 3 exit criterion requires walking every /ui/* page. 24 of 25 routes return 200. documents_page.py returns RedirectResponse(url=f"/sources?document_id=...") without the /ui mount prefix, so following the 307 lands on a 404. Confirmed pre-existing: documents_page.py has no uncommitted diff and was last touched in 6a3ee26, well before V4.7. Out of the V4.7 scope boundary, so NOT fixed - raised for the user to decide.
#### [33] Instruction-file defects corrected
*comment* - **noted**
**services.instructions.md rewritten after the decomposition, per the mandated order**
All five defects from plan Phase 3 task 7 fixed. (a) Line 11 "1 service class per data model" replaced with one service class per AGGREGATE, with DocumentType-under-DocumentService as the worked example; this is the measured cause of sources.py reaching 1,389 lines. (b) Added a Model Ownership section with a table covering every model plus an explicit junction-table rule, which the file previously had no home for. (c) The mandatory-CRUD rule (old lines 30-32) was already false: prompts.py, quality.py, normalization.py, media_storage.py and source_media.py define no service class at all, EvidenceService deliberately exposes no create/delete because ExecutionAttempt is append-only, and RegistryService uses generic <op>_entry naming. Softened to intent plus an explicit "do not add unused CRUD to satisfy symmetry". (d) Old line 13 (services fully independent) read as contradicting old lines 75-77 (compose across tables); reworded to separate READING across models via eager loads from the owning root, which is allowed, from IMPORTING another service, which is not. (e) Typo "picutre" removed. Also recorded the real enforcement mechanism: tests/test_service_boundaries.py, and errors.py as the neutral shared-type home.
#### [34] Line-number citation removed from the boundary test
*risk* - **noted**
**test_service_boundaries.py cited the rule by line number**
The test docstring pinned .github/instructions/services.instructions.md:13. Rewriting the file invalidated that anchor. Replaced with a section-name citation ("Structure") so future edits to the instruction file cannot silently desynchronise the test docstring. errors.py was also added to the docstring list of neutral modules.
### Phase 4 - measurement window
#### [35] Both cited offenders were already deleted
*deviation* - **noted**
**Phase 4 premise partly overtaken by Phase 1**
The plan states the session.commit() at line 228 "remains inside" the measurement window. Diffed against f86c0ff~1: at V4.6 the window held resolve_provider_input (async; normalization + artifact write + DB work) and that commit. Phase 1 deleted both. What remains between the clock and the wait_for is build_provider_input, now pure field copying because normalization moved to ingest and file_hash is already stored. Measured at 6.2 us per call with zero awaits, so it cannot yield to the event loop. Plan tasks 1-2 were therefore already satisfied in substance; the clock was still moved to make the property structural rather than incidental.
#### [36] No preprocessing left to record separately
*deviation* - **noted**
**Plan task 3 declined**
Task 3 offered recording preprocessing time as its own value. After Phase 1 there is no preprocessing in the window: 6.2 us of attribute copying. Adding a preprocessing_ms column to measure that is unnecessary complexity and was declined under the guiding principle. Raised rather than decided silently.
#### [37] Undocumented 475ms contributor the plan did not identify
*risk* - **answered**
**Lazy provider construction was inside the timed region**
The regression test measured 890ms where ~200ms was expected. Cause: services.sources.provider is a lazy property, and it appears as an argument expression to _call_transcriber, so it is evaluated after the clock starts but before wait_for begins timing. Measured 475ms to construct OpenRouterTranscriptionProvider on first access and 0.001ms after. The first attempt of every worker process therefore booked ~0.5s of HTTP client construction as provider latency. This plausibly accounts for the low end of the historical 0.4-2.0s local_timeout overshoot, and Phase 1 did not touch it. The property is loop-invariant, so it was hoisted above the per-source loop, which also removes the repeated attribute lookup from the two evidence-capture sites.
#### [38] test_timeout_duration_excludes_pre_call_setup
*comment* - **noted**
**Regression guard added**
New test in tests/services/test_workflows_reliability.py simulates 400ms of blocking setup against a 200ms provider budget and asserts the recorded duration_ms sits near the budget and well clear of budget+setup. Verified to fail on the pre-fix code (625 < 540 assertion error) and pass after, so it is a real guard rather than a tautology. This is the plan Phase 4 verification criterion expressed as a test.
#### [39] sources_page.py no longer prints raw milliseconds
*comment* - **noted**
**Duration render scaled**
Plan task 4. _format_duration renders >=1s as "27.6 s" and below that as "612 ms", per user selection. No test asserted the old format.
### Phase 5 - worker fault containment
#### [40] Probed behaviour: the defect is a stranded job, not a silent retry
*deviation* - **noted**
**Plan task 4 describes a failure mode that does not occur**
The plan asks for a test that a deliberate programming error "does not silently retry". Probed empirically with an injected AttributeError. Mode A, error raised after the claim commits (inside advance_job): raised exactly ONCE, job left at PROCESSING, retry_count 0, and never re-claimed because claim_next_queued_job filters status == QUEUED. That is a permanently stranded job with one swallowed log line, not a retry. advance_job PROCESSING branch, commented "Recover mid-flight jobs", is unreachable from the worker for the same reason. Mode B, error raised before or during the claim: 20 raises in 1.2s, an unbounded hot spin at the poll interval. The plan context says worst-case silent burn is 60s under WORKER_MAX_RETRIES=1, but Mode B never reaches the per-job retry machinery so nothing caps it. Both modes share the root cause the plan correctly identifies.
#### [41] Flag set in 9 places, read in none
*comment* - **noted**
**retriable was decorative**
Measured across src/: retriable is assigned at errors.py:40/47/79, sources.py:877/884, store.py:127/205/366, workflows.py:284/580/593/606 and read nowhere. classify_unexpected_error already returns retriable=False, so the classification existed and was discarded. Phase 5 makes it load-bearing in two places.
#### [42] User chose: stop the worker loop
*question* - **answered**
**Loop policy for a non-retriable error with no job to mark**
Mode B has no claimed job, so there is no row to mark FAILED and no reason to expect the next poll to differ. Options offered were stop the loop, circuit-breaker after N consecutive failures, or exponential backoff. User selected stopping the loop, logged at CRITICAL, returning cleanly so the exception does not surface only at app shutdown via worker_consumer_lifespan wait_for.
#### [43] User chose: mark FAILED and keep going
*question* - **answered**
**Loop policy for a non-retriable error where the job CAN be marked failed**
Distinct from entry 42 and not covered by it. Mode A can contain the failure on the job row, so stopping the loop would let one poison job halt transcription for every other job. User selected containment: mark the job FAILED, which is visible in the UI and resubmittable, and continue polling.
#### [44] Containment write uses its own transaction
*risk* - **noted**
**Terminal write runs on a possibly dirty session**
_advance_job_with_containment rolls back the caller session before marking the job FAILED, and calls update_job_state with no session so the service owns and commits its own transaction. This satisfies plan task 3 atomicity: the terminal write cannot be left half-applied by whatever failure poisoned the caller session.
#### [45] test_run_worker_loop_survives_process_next_exception replaced
*deviation* - **noted**
**An existing test encoded the defective behaviour**
That test asserted the loop SURVIVES a RuntimeError and continues, which is exactly the Mode B defect. It was replaced by test_run_worker_loop_stops_on_non_retriable_exception, plus a new test_run_worker_loop_survives_retriable_exception so suppression of genuinely transient faults stays covered. Unlike Phase 3, changing test logic here is the point of the phase. Both new guards plus the Mode A guard were verified to FAIL on pre-fix code: the Mode B test times out, which is the infinite spin made visible.
### Phase 6 - CI enforcement
#### [46] Remote is Gitea 1.27.2, not GitHub
*comment* - **noted**
**The plan assumes GitHub Actions**
Remote is bbchops/transcription on Gitea 1.27.2, which reads .github/workflows/ and proxies actions/checkout@v4 to GitHub. Workflow syntax needed no change. Note the remote default branch is traumatized, not main.
#### [47] Runner availability cannot be confirmed via the API
*risk* - **noted**
**Repo-scoped runner list returns 0; admin endpoint returns 403**
Existence of CI could not be asserted by query on this host. Proven instead by observation: an instance-level runner named docker-runner executed the jobs. Anyone re-verifying this must trigger a run rather than trust the runner API.
#### [48] CI writes a .env file instead of exporting an env var
*deviation* - **noted**
**Settings reads the .env file; the external-test skip guard reads os.getenv**
The two read different sources, and locally both conditions hold at once, which is why 4 tests skip. Measured in CI: no .env = 115 failed / 18 errors; exported dummy var = 3 failed (externals un-skip and hit the network); written .env file = the exact local baseline. Only openrouter_api_key is required.
#### [49] CI invokes pre-commit rather than repeating ruff/ty commands
*deviation* - **noted**
**Plan task 2 asks CI to run the same checks as local**
Satisfied structurally rather than by copying command strings: CI runs uv run pre-commit run --all-files, so the checks have a single definition in .pre-commit-config.yaml and CI cannot drift from local. Hooks are language: system and uv run puts .venv on PATH.
#### [50] Platform-dependent prompt name guard, caught by CI on its first green run
*deviation* - **noted**
**The direct-child name guard relied on Path(name).name != name**
On POSIX, backslash is an ordinary filename character, so nested\prompt.md passed the direct-child guard and failed later as NOT_FOUND instead of VALIDATION. Windows can never reproduce it. No traversal was possible because the path.parent != root check still held, so severity is a wrong error category plus a red gate. Fixed by rejecting / and \ explicitly, matching the ^[^/\\]+$ pattern config.PromptFilename already used. User approved the code fix over weakening the test.
-195
View File
@@ -1,195 +0,0 @@
# V4.7 Scope Boundary
This document defines the frozen boundary for V4.7, an **architectural cleanup and evidence-model re-alignment release**. V4.6 remains the behavioral baseline. V4.7 introduces **no new user-facing features**; it completes the structural work V4.6 deferred, simplifies the evidence model down to what the application actually uses, and closes the correctness items opened during V4.6 implementation.
Every item in scope is traceable either to a review finding ID in [Architecture & Code Review Report](../architecture_code_review_2026-08-17.md) or to a numbered entry in the V4.6 implementation review log. Any change that cannot be traced to one of those is out of scope.
All image *presentation*, media, and telemetry-presentation work is deferred to V4.8. See [`ver4.8/feature_backlog_v4_8.md`](../ver4.8/feature_backlog_v4_8.md).
## Purpose
- Collapse the duplicated evidence model so that `job_source` records **membership and queue state** and `execution_attempt` records **evidence**, with no overlap.
- Remove the `ProcessingArtifact` subsystem, which has executed exactly once in the application's history, and move orientation normalization to ingest where it belongs.
- Complete [MED-14] by decomposing `SourceService`, which still owns four domain models.
- Correct the run-time measurement window so provider latency can be trusted before anything is built on top of it.
- Stop the worker from silently swallowing programming errors.
- End the dual-spelling persistence of `JobSourceStatus`.
- Make the `ruff` / `ty` gate enforceable in CI rather than only on a developer machine that has run `pre-commit install`.
## Confirmed Operating Context
These answers are frozen for V4.7 and govern every decision below.
| Question | Answer |
| :--- | :--- |
| Database | **SQLite only.** PostgreSQL remains the intended destination. The V4.6 re-level already resolved the FK cycle with `use_alter=True`. |
| Topology | **Single user, single process, single worker.** Unchanged from V4.6. |
| Schema evolution | **Re-level from current metadata**, exactly as V4.6. No Alembic, no `_upgrade_*` chain. |
| Existing data | The live database is populated. V4.7 is **schema-affecting**: three structural changes plus a one-time image backfill, delivered by a single `tools/migrate_v46_to_v47.py`. |
| Release character | **Architectural cleanup and evidence-model re-alignment.** No new features. |
| Image fidelity | **Visually lossless is sufficient.** Measured at 51.5-55.0 dB PSNR for a single re-encode generation. Bit-exact preservation was considered and rejected as unnecessary complexity. |
| Provider settings | `WORKER_PROVIDER_TIMEOUT_SECONDS=30.0`, `WORKER_MAX_RETRIES=1`, calibrated 2026-08-18. Not revisited in V4.7. |
## Evidence Gathered
The decisions below rest on measurements taken against the live database on 2026-08-18, not on inspection alone.
| Measurement | Result |
| :--- | :--- |
| `job_source.raw_transcription` vs latest attempt | **77/77 identical** |
| `job_source.ai_metadata` vs `normalized_metadata` | **77/77 identical** |
| `job_source.raw_api_response` vs `sdk_response_snapshot` | **77/77 identical** |
| `job_source.error_detail` vs attempt `error_detail` | 2/2 identical |
| `job_source.status` vs `execution_attempt.status` | 0/79 textually identical - the dual-spelling defect [45] |
| `job_source` rows with more than one attempt | 1 of 79 |
| `processing_artifact` rows in existence | **2**, both from one job on 2026-08-16, against 77 successful transcriptions |
| Source images carrying EXIF orientation 3 | **58 of 79**, of which only 1 was ever normalized |
## In Scope
### 1. Evidence Model Simplification
`job_source` began as the many-to-many link between `job` and `source` and accreted response-capture fields over time. `execution_attempt`, added later in V4.2 (commit `6bd4cbb`), captures the same information in more detail. The measurements above show the overlap is total, not partial.
**`job_source` is stripped, not deleted.** It cannot be folded into `execution_attempt`, because it carries state that exists when no provider call has occurred:
- `store.py:249,313` create rows with `status=PENDING` **at job creation**, before any call.
- `workflows.py:432-442` selects work by `status != TRANSCRIBED` on `job.job_sources`.
- `jobs.py:378-384` cancel writes a terminal state with **no provider call at all**, so no attempt row could carry it.
An append-only evidence table cannot express "queued, not yet attempted" or "cancelled before any call". The junction survives; the duplicated evidence does not.
**Retained:** `id`, `job_id`, `source_id`, `status`.
**Removed:** `raw_transcription`, `ai_metadata`, `raw_api_response`, `executed_at`, `error_detail`.
**Added:** `JobSourceStatus.CANCELLED`, so cancellation stops overloading `FAILED` plus the free-text string `"Cancelled by user"`. This is what retires `error_detail`.
**Unchanged:** the retry reset at `jobs.py:411-424`. Flipping `FAILED` back to `PENDING` loses no history, because `ExecutionAttempt`'s `UniqueConstraint(job_id, source_id, attempt_number)` (`models.py:389`) already preserves every prior attempt. This is confirmed in live data: attempt 1 `FAILED`/`local_timeout` and attempt 2 `TRANSCRIBED` are both retained. Adding a second `job_source` row per retry would duplicate that mechanism and break the one-row-per-`(job, page)` assumption in `read_job_source_for_job` and `sources.py:570-574` - where uniqueness is enforced **in code, not by a database constraint**.
This item absorbs review log [45], since both changes rewrite `JobSourceStatus` persistence and must land as one migration.
### 2. ProcessingArtifact Removal and Ingest Normalization
`ProcessingArtifact` is a generic container for derived data products, with a `CheckConstraint` enforcing that content is either inline JSON or an external file, never both. Two rows exist. The quality-warnings path at `workflows.py:560-571` writes one on **every** successful page, yet 77 successful transcriptions produced a single row, so the subsystem postdates nearly all data and has effectively never run.
Orientation normalization itself is **not** dispensable and was **not** a red herring. 58 of 79 stored images carry EXIF orientation 3, and the raw decoded pixels of the page that prompted the original investigation are genuinely upside down. Sending those bytes unrotated sends an inverted page to the model.
The fix is to normalize at ingest rather than derive at transcription time:
- Rotate on upload, in `media_storage`, before the image is stored. Every stored byte is then already upright and no derivative needs to exist.
- Use Pillow with `qtables=im.quantization`, `subsampling=JpegImagePlugin.get_sampling(im)`, `optimize=True`. Measured against the current `quality=95, subsampling=0` settings at `normalization.py:84-85`, this is **better on both axes**: 51.5-55.0 dB PSNR versus 50.0-53.5 dB, and roughly 6% smaller output versus 38% larger.
- Strip the EXIF orientation tag after rotating.
- No archival master is retained. No external `jpegtran` dependency is introduced. No MCU-alignment rejection path is needed, because Pillow handles any dimensions - including the single 2306x2019 outlier.
Then delete: the `processing_artifact` table, the `ProcessingArtifact` model, the ~283-line artifact cluster in `sources.py` (lines 732-1015), `resolve_provider_input`, and the artifact branch of `build_evidence_export`. The `transcription_quality_warnings` payload folds into `execution_attempt.normalized_metadata`.
Deleting stored images is not involved; the 58 already-ingested rotated images are rotated **in place** by the migration. No live integrity check is invalidated: `Source` has no digest column, and the only stored digests are `ExecutionAttempt.request_manifest_sha256` - a hash of the request manifest, correct as history - and `ProcessingArtifact.payload_sha256`, which is removed with the table.
### 3. SourceService Decomposition ([MED-14])
`services/sources.py` is **1,389 lines** and `SourceService` owns `Source`, `JobSource`, `ExecutionAttempt`, and `ProcessingArtifact`.
Item 2 removes the `ProcessingArtifact` responsibility by **deletion rather than extraction**. The previously planned `services/artifacts.py` is therefore cancelled - extracting ~283 lines into a new module and then deleting that module would be wasted work.
What remains is the `ExecutionAttempt` cluster, moved to **`services/evidence.py` (~174 lines)**: `read_latest_execution_attempt` (216-245) with its `LatestExecutionAttempt` read model, `promote_machine_attempt` (679-710), `list_execution_attempts` (710-732), and `build_evidence_export` (1015-1107).
**`update_job_source_transcription` stays in `sources.py`.** The V4.6 deferral note proposed moving it to `workflows.py` as orchestration; that proposal is not adopted. The method writes `JobSource` and `ExecutionAttempt` inside one session scope and derives `attempt_number` at lines 595-600, and `services.instructions.md:63-65` requires the transcript update and the paired terminal status change to commit or roll back together. `services.instructions.md:72` assigns session-aware write helpers to services and commit-boundary control to orchestration, so the current placement already satisfies the instruction file. Splitting the two writes across modules is the most plausible way that atomicity later gets broken. The method will shrink under item 1, since several of the fields it writes cease to exist.
Expected result: `sources.py` lands near **900 lines**.
**`.github/instructions/services.instructions.md` is revised as part of this item** (review log [59]), after the move rather than before. The instruction file's line 11 rule, `1 service class per data model`, is table-shaped rather than aggregate-shaped and is the measured cause of the 1,389-line module this item exists to break up; leaving it unchanged would license the same growth again. The file is also silent on `job_source` and `document_person`, the junctions where the four core components intersect, so ownership of those has never been written down. Sequencing the revision after the decomposition makes the refactor the empirical test of the rule: if the new rule is right, the resulting module boundaries follow from it, and if the code has to be bent to fit, the rule is wrong.
### 4. Run-Time Measurement Window (review log [55])
`services/workflows.py:221` sets `monotonic_started_at` **before** provider-input preparation and the `session.commit()` at line 228. Line 251 computes `elapsed_seconds` from it. But the `asyncio.wait_for` timeout at lines 240-249 wraps **only** `_call_transcriber`.
`duration_ms` therefore measures a strictly wider window than the budget that governs it. This is observable in the migrated data: three historical `local_timeout` rows recorded 20.4 / 20.8 / 22.0 s against a 20.0 s timeout.
In scope: either record provider latency as a distinct value, or move `monotonic_started_at` to immediately before the `wait_for`. Whichever is chosen, the resulting figure must be the quantity the timeout actually governs. Item 2 also removes normalization from this window entirely, which shrinks the discrepancy but does not by itself fix it.
This item **must land before any V4.8 telemetry presentation work**.
### 5. Worker Exception Handling (review log [8])
`worker.py:96-106`, `handle_worker_exceptions`, catches bare `Exception`, logs it, and suppresses it. A programming error inside the worker loop is therefore indistinguishable from a transient provider fault and is retried silently with no UI signal.
In scope: distinguish genuinely retriable faults from programming errors, and ensure a non-retriable error surfaces rather than looping. Retry counting and terminal-state transitions remain governed by `services.instructions.md:63-65`.
### 6. CI Enforcement of the Quality Gate ([HIGH-06], review log [40])
`.github/workflows/` is empty. The `ruff check` and `ty check` gate established in V4.6 Phase 7 exists only in `.pre-commit-config.yaml`, which is inert until a developer runs `pre-commit install`.
In scope: a CI workflow running `ruff check`, `ty check`, and `pytest` on push and pull request, using the same commands as the local hooks so the two cannot drift.
## Out of Scope
- **All image and media presentation work.** Pan and zoom on Source Detail, the homepage gallery, multi-portrait support, image descriptions, and background wallpaper are V4.8.
- **The model-performance rollup** (review log [54]). It depends on item 4 and is a new user-facing view.
- **Reducing `update_job_source_transcription`.** See section 3.
- **Bit-exact image preservation.** Considered and rejected; see Confirmed Operating Context.
- **PostgreSQL cutover.**
- **Re-tuning `WORKER_PROVIDER_TIMEOUT_SECONDS` or `WORKER_MAX_RETRIES`.** Calibrated 2026-08-18 against measured per-model durations.
- **Removing slow models from `PROVIDER_MODELS`** (review log [53]). A configuration judgement, deliberately left with the operator.
- **Any new feature.**
## Locked Design Decisions
### A. Cleanup Only
V4.7 changes structure and correctness. It does not change what the application does for a user. If a change would be visible on a page as new capability, it belongs in V4.8.
### B. One Home Per Fact
After V4.7, any given piece of evidence is stored in exactly one place. `job_source` holds membership and state; `execution_attempt` holds evidence. Denormalized convenience copies are not reintroduced, and if a read becomes awkward the fix is a query or a read model, not a duplicated column.
### C. Delete Before Refactor
Item 2 deletes the artifact subsystem before item 3 restructures what remains. Code scheduled for deletion is never extracted, renamed, or moved first.
### D. Simplicity Over Edge-Case Management
Where two approaches both satisfy the requirement, the one with fewer moving parts wins. This is why rotation uses Pillow rather than a lossless DCT transform, and why no archival master is kept.
### E. Measurement Before Presentation
Item 4 precedes all V4.8 telemetry work. A dashboard built on a conflated metric looks authoritative and quietly misleads.
### F. One Migration, Backed Up
All schema and data changes land in a single `tools/migrate_v46_to_v47.py`: idempotent, never invoked at startup, never run by the test suite, following the `tools/migrate_v45_to_v46.py` conventions. `data/transcription.db` **and** `data/documents/` are backed up before it runs, because the image backfill rewrites files in place.
### G. The Instruction Files Are the Standard
`.github/instructions/services.instructions.md` and `ui.instructions.md` govern. Where this document and an instruction file disagree, the instruction file wins.
## Acceptance Criteria
- `job_source` carries exactly `id`, `job_id`, `source_id`, `status`; every evidence read resolves through `execution_attempt`.
- `JobSourceStatus.CANCELLED` exists and cancellation no longer writes free text into a removed column.
- `job_source.status` and `execution_attempt.status` persist with one spelling, and existing rows are consistent.
- The `processing_artifact` table, its model, and its service cluster no longer exist.
- Newly uploaded images are stored upright with no EXIF orientation tag, and the 58 pre-existing rotated images have been backfilled.
- `services/sources.py` is materially smaller, with `ExecutionAttempt` responsibilities in `services/evidence.py` and `update_job_source_transcription` unmoved.
- `services.instructions.md` states an aggregate-shaped ownership rule, names an owning service for every model including the junctions, and no longer contradicts itself on multi-table operations.
- The recorded duration reflects only the operation the timeout governs.
- A programming error in the worker loop is distinguishable from a provider fault.
- `ruff check` reports no findings; `ty check` reports **0 diagnostics**, the V4.6 exit state.
- The full test suite passes.
- CI runs the same `ruff` / `ty` / `pytest` gate as the local hooks.
- No new user-facing behavior.
## Scope Freeze Gate
This boundary is frozen. Adding an item requires a finding ID or a review-log entry, and an explicit note recording the addition.
## Related Local References
- [V4.7 Implementation Plan](implementation_plan_v4_7.md)
- [Architecture & Code Review Report](../architecture_code_review_2026-08-17.md) - finding IDs
- [V4.6 Review Log](../ver4.6/review_log_v4_6.md) - resolves the `review log [N]` citations used throughout this document
- [V4.6 Scope Boundary](../ver4.6/scope_boundary_v4_6.md) - the baseline this release builds on
- [V4.6 Implementation Plan](../ver4.6/implementation_plan_v4_6.md)
- [V4.8 Feature Backlog](../ver4.8/feature_backlog_v4_8.md) - where deferred feature work is parked
- `.github/instructions/services.instructions.md`
- `.github/instructions/ui.instructions.md`
-111
View File
@@ -1,111 +0,0 @@
# V4.8 Feature Backlog
**Status: not scoped.** This is a parking document, not a frozen boundary. It records feature work deferred out of V4.6 and V4.7 together with the evidence gathered so far, so that scoping V4.8 does not start from a blank page.
V4.8 is the first release since V4.5 to add **new user-facing behavior**. V4.6 was pure remediation and V4.7 is architectural cleanup; both were held to "no new features." That constraint ends here, which means V4.8 needs a different verification gate: V4.6 and V4.7 could be validated by "the suite still passes unchanged," and V4.8 cannot.
## Dependency on V4.7
**The model-performance rollup below must not begin until V4.7 Phase 4 lands.** `duration_ms` currently measures provider call *plus* image normalization, artifact persistence, and a DB commit, while the timeout governs only the provider call. A rollup built on it would chart preprocessing time mixed with provider latency and look authoritative while quietly misleading. V4.7 Phase 1 removes normalization and artifact persistence from that window, but the commit remains inside it until Phase 4. See [V4.7 scope boundary section 4](../ver4.7/scope_boundary_v4_7.md).
## Candidate Features
### 1. Pan and Zoom on Source Detail
**Practicality: high. Effort: S.**
`ui/components/document_panzoom.py` existed and was **deleted in V4.6 Phase 5** (`6a3ee26`) because it was exported but wired to no page. It is 136 lines and recoverable:
```
git show 6a3ee26^:src/transcription/ui/components/document_panzoom.py
```
It already handled both images and PDFs (the latter via an iframe).
Two things must change on reintroduction - this is not a straight revert:
- It loaded Panzoom from the **unpkg CDN**. For an archival application the library should be vendored locally, otherwise the viewer breaks offline and depends on a third party staying available.
- It carried its own `_document_url()` helper. V4.6 Phase 5 extracted exactly that logic into `ui/components/media_urls.py` as `resolve_media_url`. Reintroducing the old helper would recreate the duplication Phase 5 removed.
Scope note: apply it to **Source Detail only**. `dark_room_viewer` (`ui/components/viewers.py`) is shared by four pages - `sources_page.py:268`, `home_page.py:25` and `:88`, `people_page.py:453`, `documents_page.py:524` - so a flag on it would leak pan-zoom into the homepage and document detail, which is not wanted. Add a separate component and use it only at `sources_page.py:268`.
Numbering note: the Phase 5 commit message states pan-zoom would return "in V4.7 alongside the other photo/image work." Moving it to V4.8 preserves that **intent** - it stays grouped with the photo work - and changes only the release number.
### 2. Homepage Image Gallery
**Practicality: high. Effort: S. Recommended first feature.**
The storage layer is already built:
- `ui/homepage_store.py:82` `list_homepage_images()` already returns **every** stored image, sorted by modification time.
- `store_homepage_image()` already accumulates files rather than overwriting.
- Today the UI calls only `latest_homepage_image()` and displays one image. `list_homepage_images()` is currently exercised **only by tests**.
So multi-image upload is effectively done; what is missing is presentation. NiceGUI 3.13.0 provides `ui.carousel` for left/right navigation and `ui.timer` for rotation.
Sub-items:
- Multi-image display with left/right navigation - small, mostly wiring.
- Optional slideshow rotating every ~10 minutes.
**Performance caveat:** `list_homepage_images()` performs a directory scan with a `stat()` per file on every call, and `home_page.py` already performs blocking I/O in the page handler (V4.6 review log [25], which was deliberately left alone). A rotating timer that re-enumerates on every tick would repeat that scan indefinitely. Enumerate once at page load and cache the list.
### 3. Multiple Person Portraits
**Practicality: medium. Effort: M/L. Defer behind item 2.**
`Person.portrait_path` is a **single string column**. Supporting multiple portraits requires a new table, a data migration, and upload UI - a materially larger job than item 2, and a different one.
### 4. Image Descriptions
**Practicality: medium, conditional. Effort: M.**
Homepage images are **filesystem-only with no metadata store**, so a caption has nowhere to live today. This needs either a sidecar JSON file or a real table.
This is cheap **only if** item 3 is being done at the same time, since both need the same metadata layer. Designing that layer twice would be wasteful; design it once or not at all.
### 5. Model-Performance Rollup (V4.6 review log [54])
**Practicality: high, but blocked. Effort: M.**
Run-time telemetry is already captured and is per page: `execution_attempt.duration_ms` is a required non-null field written on all three paths in `workflows.py` (success 278, `TimeoutError` 295, general failure 330), with failures using a monotonic clock. Verified against the live database: 80 rows across 80 distinct (job, source, attempt) combinations, one row per page - the largest job has 60 attempts across 60 distinct pages - and zero nulls. Token counts live on the same row in `normalized_metadata.usage`, so tokens-per-second is already derivable without a join.
What is missing is **aggregation**. The figure is visible only for the latest attempt of one source at a time (`sources_page.py:400`), rendered raw as `"27612 ms"`. There is no rollup by model, prompt, or document.
The gap is concrete: calibrating the provider timeout on 2026-08-18 required hand-written SQL against the database, because the application could not answer "which model is slow."
Proposed shape: median / p95 / max duration, tokens per second, and a timeout rate, grouped by model. **Blocked on V4.7 Phase 4.**
### 6. Desaturated Background Wallpaper
**Practicality: low. Recommendation: do not build, or gate behind a setting defaulted off.**
Trivial to implement (`ui.add_css` with a CSS `filter`), but this is a dense archival data application - transcripts, JSON evidence panels, data tables. A background image behind all of that costs contrast and legibility on every page, for aesthetic gain only.
## Suggested Grouping
If V4.8 is scoped as one release, the natural split is:
**Track A - image experience:** items 1 and 2. Both are small, both are self-contained UI work, and item 2's storage layer already exists. This is the highest value for the least risk.
**Track B - metadata layer:** items 3 and 4 together, since they share a table. Only worth starting if both are wanted.
**Track C - telemetry:** item 5, gated on V4.7 Phase 4.
Item 6 is not recommended.
## Open Questions for Scoping
- Should Track B happen at all, or is one portrait per person sufficient?
- Should the slideshow interval be configurable, or fixed?
- Should the model-performance rollup be its own page, or a panel on an existing one?
- Should vendored Panzoom be committed to the repository, or fetched at build time?
## Related Local References
- [V4.7 Scope Boundary](../ver4.7/scope_boundary_v4_7.md) - the blocking dependency for item 5
- [V4.6 Scope Boundary](../ver4.6/scope_boundary_v4_6.md)
- [Architecture & Code Review Report](../architecture_code_review_2026-08-17.md)
- `.github/instructions/ui.instructions.md`
- `src/transcription/ui/homepage_store.py` - existing multi-image storage
- `src/transcription/ui/components/media_urls.py` - canonical URL resolution
-213
View File
@@ -1,213 +0,0 @@
# System Architecture (Version 4)
This document describes the production architecture of the document transcription system.
## Architecture Objectives
- Preserve original source material, per-execution machine output, and separate human revision.
- Support batching one or more images into ordered multi-page documents.
- Capture submission-time prompt provenance and a per-page OpenRouter SDK response snapshot.
- Execute page transcription concurrently with bounded `asyncio` workers.
- Maintain relational portability across SQLite and PostgreSQL.
- Keep operator workflows cross-platform and Python-driven.
- Support one role-bearing link per Person and Document through an extensible role registry.
- Support registry-driven document classification with protected semantic built-ins.
## Core Capabilities
- Ingest one or more images into sequential `Source` pages under a `Document`.
- Execute asynchronous vision transcription with bounded worker concurrency.
- Preserve original source files with SHA-256 digests and byte sizes.
- Freeze prompt text, prompt hash, model, and explicitly configured sampling parameters on each `Job`.
- Preserve page-level machine output, normalized metadata, and an SDK-serialized OpenRouter response snapshot on `JobSource`.
- Organize historical `Person` records through UUID-identified Document links and extensible roles.
- Classify Documents through a UUID-identified registry with hidden semantic built-ins and unique labels.
- Maintain human revision separately from machine-generated text.
- Isolate page failures so multi-page jobs can complete with partial success.
- Operate across supported platforms through Python-based application and maintenance tooling.
V4.2 extends this baseline with immutable execution attempts, exact OpenRouter transport evidence, safe
versioned exports, and provider-neutral derived-artifact provenance. `JobSource` remains the mutable queue and
compatibility projection; `ExecutionAttempt` is the authoritative append-only processing history. See the
[V4.2 Scope Boundary](../ver4.2/scope_boundary_v4_2.md).
## Technical Stack
- **Runtime:** Python 3.12 or later.
- **Web application:** FastAPI and NiceGUI.
- **Persistence:** SQLModel and SQLAlchemy, with SQLite and PostgreSQL support.
- **Validation and settings:** Pydantic V2 and pydantic-settings.
- **Concurrency:** Python `asyncio` workers.
- **Vision integration:** OpenRouter through the application's provider adapter.
- **Testing and quality:** pytest, pytest-asyncio, Ruff, and ty.
## Runtime Topology
The runtime operates as an asynchronous Python application:
- FastAPI + NiceGUI web application process.
- In-process `asyncio` worker engine for transcription execution.
- Relational persistence via SQLModel / SQLAlchemy.
- Pydantic V2 validation across API payloads, prompt configuration, and structured metadata.
^^^mermaid
flowchart LR
U[Browser User] --> A[FastAPI + NiceGUI App]
A --> W[Asyncio Worker Engine]
A --> DB[(Relational DB)]
W --> P[Vision Provider APIs]
W --> DB
^^^
## Lifecycle Ownership
Application lifespan owns runtime setup and teardown:
- Initialize logging, settings, directories, and prompt configuration.
- Manage asynchronous database engine connection pools.
- Execute database bootstrap or migrations.
- Recover stale or interrupted jobs on startup.
- Manage graceful shutdown of active background tasks.
## Layered Module Structure
### Interface Layer
- `src/transcription/ui/**`
- `src/transcription/api/**`
Responsibilities:
- Render document, source, person, job, and classification views.
- Accept user input for uploads, editing, linking, and revisions.
- Present structured validation and conflict feedback.
### Application and Async Worker Layer
- `src/transcription/services/workflows.py`
- `src/transcription/worker.py`
Responsibilities:
- Orchestrate uploads, job creation, and status transitions.
- Execute per-page provider calls through bounded concurrency.
- Persist page-level outcomes and update aggregate job state.
### Domain and Service Layer
- `src/transcription/db/models.py`
- `src/transcription/services/documents.py`
- `src/transcription/services/sources.py`
- `src/transcription/services/jobs.py`
- `src/transcription/services/people.py`
- `src/transcription/services/workflows.py`
Responsibilities:
- Keep one primary service boundary per aggregate: Documents, Sources, Jobs, and People.
- Documents own document records and the document-type registry.
- Sources own source records, revisions, source media formats, MIME resolution, and page execution evidence.
- Jobs own job lifecycle state and transitions.
- People own person records, relationship roles, document-person links, and portrait media.
- Apply deterministic conflict handling for relationship-role writes.
- Synchronize each Document's complete Person link set in the same transaction as Document fields.
- Resolve and validate registry records by UUID; use hidden semantic keys only for application-owned built-in behavior.
### Source Media Policy
- `services/sources.py` is the single authority for accepted Source extensions and canonical MIME types.
- Storage and provider payload loading must call the same Source validation functions.
- Supported Source formats are JPEG, PNG, TIFF, and PDF.
- Upload is an interface action, not a domain aggregate. Service names, errors, and workflow variables use
`Source` terminology; compatibility aliases may remain temporarily at old import boundaries.
### Infrastructure Layer
- `src/transcription/db/**`
- `src/transcription/providers/**`
Responsibilities:
- Provide async database sessions and engine configuration.
- Provide provider adapters for vision model execution.
## Core Workflows
### 1. Multi-Page Transcription
1. User uploads one or more images for a `Document`.
2. System stores files, hashes them, creates ordered `Source` rows, and creates a `Job`.
3. Worker claims the job, marks it `processing`, resolves metadata-directed orientation, and sends either the
immutable original or an exact normalized derivative to the provider.
4. Each provider call appends an `ExecutionAttempt` with its request manifest, transport evidence, SDK snapshot,
normalized metadata, timing, and outcome.
5. The linked `JobSource` is updated as a compatibility projection. The first successful attempt establishes
`Source.preferred_execution_attempt_id` and `Source.raw_transcription`; later successes remain candidates.
6. Aggregate status becomes `completed`, `partial_success`, or `failed`.
### 2. Document-Person Relationship Management
1. User opens Document Create or Edit.
2. UI loads one Linked People table containing Person and Role.
3. Add, Edit, and Delete operations change staged UI state only.
4. Service validates the complete desired set and computes deterministic add, update, and remove deltas.
5. Document fields and links commit once in one transaction; any failure leaves both unchanged.
### 3. Document Type Management
1. User selects a registry-backed document type for a document.
2. Service resolves the Document Type UUID.
3. Persistence stores the `document_type_id` reference.
4. Inactive types remain valid for historical rows but are excluded from default selectors.
### 4. Document Printing
1. User opens Print from persisted Document Detail.
2. Service builds a safe projection containing archival metadata, semantic Author links, ordered Sources, current text,
and oldest-to-newest Job metadata.
3. The preview renders Facsimile or Text-only HTML without exposing local file paths.
4. An explicit action opens the browser print dialog; browser Save as PDF remains available.
## V4 Domain Rules
- `JobSource.raw_transcription` preserves page output for its Job execution.
- `Source.raw_transcription` is the selected preferred-machine-output projection for a page.
- `Source.preferred_execution_attempt_id` identifies its exact immutable provenance; candidate promotion updates
both fields atomically.
- Human corrections occur only in `Source.revised_text`.
- Prompt and parameter provenance is frozen on `Job` at submission time.
- The SDK-serialized OpenRouter response snapshot is stored on `JobSource` for each successful page execution.
- Every V4.2 provider call appends a distinct `ExecutionAttempt`; retries never rewrite earlier attempts.
- Exact response bytes identify the OpenRouter HTTP boundary and are not labeled as native upstream-provider JSON.
- Generic `ProcessingArtifact` records use versioned schemas, digests, and one inline or external content location.
- Orientation-normalized model inputs and deterministic quality warnings are versioned `ProcessingArtifact` evidence
attached to the consuming `ExecutionAttempt`.
- A `retranscription` Job contains one locked existing Source and freezes one configured allowlisted model.
- `DocumentPerson` links are unique for `(document_id, person_id)` and require one `role_id`.
- Relationship mutations are deterministic, set-based, and atomic with Document writes.
- `DocumentType.id` and `PersonRole.id` are canonical relationship identities; unique labels may evolve.
- Nullable immutable `semantic_key` values identify protected application-defined built-ins and are never public selectors.
- Current printable text uses non-null `Source.revised_text`; otherwise it uses `Source.raw_transcription`.
## Data Model Summary
- `Document` has one `DocumentType`, many `Source` pages, many `Job` runs, and many `Person` records through `DocumentPerson`.
- `Source` belongs to one `Document` and may participate in many `JobSource` executions.
- `Job` has many `JobSource` rows.
- `PersonRole` defines available relationship roles; `DocumentType` and `PersonRole` may carry hidden semantic identity.
## Test Strategy
- Unit tests for models, validation, hashing, and registry resolution.
- Service tests for registry protection, atomic link synchronization, uniqueness conflicts, and print projections.
- Async workflow tests for page isolation, partial failure handling, and stored evidence.
- UI integration tests for Linked People staging, registry selection, and safe print rendering.
## Related Local References
- [System Overview](index_v4.md)
- [System Requirements](requirements_v4.md)
- [Data Model](schema_v4.md)
- [Error Handling Policy](error_handling_v4.md)
- [Error Handling Invariant](../invariant/error_handling.md)
- [Digital Evidence and AI Processing Provenance](../invariant/ai_evidence_and_provenance.md)
-115
View File
@@ -1,115 +0,0 @@
# Error Handling Policy (Version 4)
This document defines the Version 4 taxonomy, contracts, and framework behavior used to satisfy the cross-version [Error Handling invariant](../invariant/error_handling.md).
## Invariant Alignment
Version 4 implements the invariant through:
- The shared error taxonomy below.
- Structured error envelopes with correlation IDs.
- Page-level failure isolation and explicit aggregate job status.
- Atomic relationship and classification writes.
- Consistent translation across API, UI, service, worker, persistence, and provider boundaries.
- Bounded retry guidance based on category and idempotency.
## Scope and Authority
This policy governs error behavior across:
- NiceGUI pages
- FastAPI routes
- Service-layer orchestration
- `asyncio` worker tasks
- Database interactions
- Provider adapters
## Error Taxonomy
| Category | Definition | Retriable |
| --- | --- | --- |
| `validation_error` | Payload, parameter, or schema validation failure | no |
| `user_input_error` | Unacceptable file, invalid selection, or malformed request from the operator | no |
| `not_found_error` | Requested `Document`, `Source`, `Person`, `Job`, role, or type does not exist | no |
| `conflict_error` | Operation violates uniqueness or relationship-write policy | no |
| `external_provider_error` | Provider API failure, rate limit, or execution problem | yes |
| `infrastructure_transient_error` | Temporary DB, file-system, or network instability | yes |
| `infrastructure_persistent_error` | Persistent configuration, credential, or database availability failure | no |
| `internal_unexpected_error` | Uncaught exception or logic defect | no |
## Async Batch and Page-Level Error Behavior
In multi-page `asyncio` processing:
1. Exceptions from individual page calls are trapped within the page task wrapper.
2. Failed page detail is written to `JobSource.error_detail` and the page state becomes `failed`.
3. Aggregate job status is derived from page outcomes:
- all pages succeed -> `completed`
- some succeed and some fail -> `partial_success`
- all fail -> `failed`
4. Successful pages remain valid even when sister pages fail.
## Relationship and Classification Conflict Behavior
When relationship or document-type writes fail policy checks:
1. Reject the full write operation.
2. Return structured conflict detail including target identifiers and the violated rule.
3. Preserve existing persisted relationships unchanged.
## API Error Response Contract
API error responses return a structured envelope:
^^^json
{
"error_id": "err_uuid_12345",
"category": "conflict_error",
"message": "Relationship write conflicts with existing links.",
"suggestion": "Adjust the requested relationship links and retry.",
"details": {
"document_id": "...",
"person_id": "...",
"attempted_role": "recipient",
"operation": "add_link",
"conflict_reason": "duplicate document-person-role link"
},
"timestamp": "2026-08-10T15:00:00Z"
}
^^^
HTTP status mappings:
- `validation_error`, `user_input_error` -> `400`
- `not_found_error` -> `404`
- `conflict_error` -> `409`
- `external_provider_error` -> `502` or `503`
- `infrastructure_transient_error` -> `503`
- `infrastructure_persistent_error`, `internal_unexpected_error` -> `500`
## UI Error Presentation Rules
- Display concise failure summaries with the next action the operator can take.
- Keep form state in context when feasible.
- Distinguish validation issues, conflict issues, provider failures, and infrastructure failures.
- For bulk relationship updates, identify the specific role or person that caused a conflict.
## Logging and Audit Expectations
- Log worker failures with correlation IDs and provider context.
- Log relationship and classification conflicts with machine-readable detail.
- Log persisted provider errors and page-level execution failures.
## Retry Guidance
- Do not auto-retry validation or conflict failures.
- Permit user-driven retry after the input or selection changes.
- Allow bounded retry for transient provider or infrastructure failures when the operation is idempotent.
## Related Local References
- [Error Handling Invariant](../invariant/error_handling.md)
- [System Overview](index_v4.md)
- [System Requirements](requirements_v4.md)
- [Data Model](schema_v4.md)
- [System Architecture](architecture_v4.md)
-102
View File
@@ -1,102 +0,0 @@
# Implementation Plan (Version 4)
## Goal
Implement the Version 4 project definition from the current repository state while preserving existing data by default.
## Migration Policy
- Database changes are non-destructive by default.
- Exception: the legacy `document_type` text field may be replaced by a `document_type_id` reference without migrating existing text values.
- Exception: `document_person` links may be recreated manually.
## Current Project Impact
- `src/transcription/db/models.py` requires full schema alignment with the V4 core documents.
- `src/transcription/services/documents.py` requires set-based document-person sync and document-type resolution.
- API modules require additive role-aware relationship behavior and document-type selection behavior.
- UI pages require grouped role displays, multi-role editing, and registry-backed document-type selection.
- Existing tests require updates for role enforcement, document-type selection, and regression safety.
## Implementation Phases
### 1. Finalize the Transition Documents
- Confirm the reset scope.
- Confirm the database exception policy.
- Keep core V4 documents as the only authoritative product definition.
### 2. Align the Persistence Layer
- Update SQLModel definitions to match the final V4 schema.
- Add `person_role` and `document_type` support.
- Replace legacy document-type storage with `document_type_id`.
- Apply the accepted manual exception strategy for `document_type` and `document_person` data.
- Preserve all other data structures non-destructively.
### 3. Update Services and Write Semantics
- Organize service ownership around Documents, Sources, Jobs, and People.
- Centralize Source extension and MIME policy in the Sources service.
- Treat upload as an interface action and remove it from domain service naming where compatibility permits.
- Implement set-based synchronization for document-person updates.
- Implement deterministic uniqueness and relationship-write conflict checks.
- Remove suggestion-related service behavior.
- Add document-type resolution and validation by UUID.
### 4. Update API Contracts
- Keep API evolution additive.
- Add role-aware relationship retrieval and write behavior.
- Add document-type catalog retrieval and UUID-based selection for document writes.
- Remove suggestion-related API surfaces from the V4 target state.
### 5. Update UI Workflows
- Replace single-person link editing with grouped multi-role editing.
- Render grouped role links on document and person detail views.
- Replace free-text document type entry with registry-backed selection.
- Preserve clear validation and conflict messaging.
### 6. Verification and Hardening
- Add or update service tests for many-per-role behavior, uniqueness conflict handling, and set-based sync correctness.
- Add API tests for relationship behavior and document-type selection.
- Add UI tests or walkthrough coverage for grouped roles and type selection.
- Add regression coverage for delete and cleanup semantics.
- Enforce backup-first test execution for AI-run unit tests: backup `./data` before tests, then always prompt for restore after successful tests.
- Keep restore confirmation-gated by default so code and test outcomes can be reviewed before data is reverted.
## Done When
- Core V4 documents and code paths agree on the final project definition.
- Relationship-role writes are deterministic and non-destructive.
- Relationship-write conflict rules are enforced consistently.
- Document type selection is registry-backed.
- The accepted manual exceptions for `document_type` and `document_person` are completed.
- The focused test coverage passes.
## Out of Scope
- Suggested/asserted relationship state.
- Suggestion review or extraction workflows.
- Global person entity-resolution engine.
- Automated semantic document-type classification.
## Delivery Order Recommendation
1. Freeze scope boundary and implementation plan.
2. Freeze core V4 documents.
3. Align persistence models.
4. Align services and API behavior.
5. Align UI behavior.
6. Run focused verification and regression checks.
## Related Local References
- [V4 Scope Boundary](scope_boundary_v4.md)
- [System Overview](index_v4.md)
- [System Requirements](requirements_v4.md)
- [Data Model](schema_v4.md)
- [System Architecture](architecture_v4.md)
- [Error Handling Policy](error_handling_v4.md)
-32
View File
@@ -1,32 +0,0 @@
# Document Transcription System Overview (Version 4)
Version 4 is the architecture baseline for the personal-scale application used to transcribe, organize, and preserve historical documents, source images, and related people records.
## Recommended Reading Order
1. [System Architecture](architecture_v4.md) for capabilities, technical stack, runtime structure, workflows, and component ownership.
2. [System Requirements](requirements_v4.md) for the verifiable V4 contract.
3. [Data Model](schema_v4.md) for entities, relationships, constraints, and persistence rules.
4. [Error Handling Policy](error_handling_v4.md) for the V4 taxonomy and boundary contracts.
## Cross-Version Invariants
- [Historical Document Transcription Design Intent](../invariant/intent.md)
- [Transcription Methodology](../invariant/transcription_methodology.md)
- [Error Handling](../invariant/error_handling.md)
- [Digital Evidence and AI Processing Provenance](../invariant/ai_evidence_and_provenance.md)
- [UI Style Guide](../invariant/ui_style_guide.md)
## V4 Transition Documents
- [Scope Boundary](scope_boundary_v4.md)
- [Implementation Plan](implementation_plan_v4.md)
## Incremental Revisions
- [V4.1 Scope](../ver4.1/scope_boundary_v4_1.md) and [Implementation Plan](../ver4.1/implementation_plan_v4_1.md)
- [V4.2 Evidence and Provenance Scope](../ver4.2/scope_boundary_v4_2.md) and [Implementation Plan](../ver4.2/implementation_plan_v4_2.md)
- [V4.3 Settings Scope](../ver4.3/scope_boundary_v4_3.md) and [Implementation Plan](../ver4.3/implementation_plan_v4_3.md)
- [V4.4 Semantic Registries, Linked People, and Printing Scope](../ver4.4/scope_boundary_v4_4.md) and [Implementation Plan](../ver4.4/implementation_plan_v4_4.md)
- [V4.5 Transcription Input Normalization and Quality Scope](../ver4.5/scope_boundary_v4_5.md) and [Implementation Plan](../ver4.5/implementation_plan_v4_5.md)
- [V4.6 Architecture Conformance and Reliability Scope](../ver4.6/scope_boundary_v4_6.md) and [Implementation Plan](../ver4.6/implementation_plan_v4_6.md)
-60
View File
@@ -1,60 +0,0 @@
# Document Transcription System Requirements (Version 4)
This document defines the baseline requirements for the document transcription system.
## Requirements Model
| ID | Category | Requirement | Verify Method |
| --- | --- | --- | --- |
| REQ-0 | System | Provide end-to-end multi-page document transcription with persistent, inspectable async job states. | demonstration |
| REQ-1 | Functional | Allow users to upload one or more images as ordered `Source` pages under a `Document`. | test |
| REQ-2 | Functional | Process page transcription asynchronously using an `asyncio` worker pool bounded by rate limits. | test |
| REQ-3 | Functional | Persist submission-time request provenance and accurately labeled page-level SDK evidence; V4.2 adds exact OpenRouter-boundary transport evidence for new attempts. | test |
| REQ-4 | Functional | Support job states `queued`, `processing`, `completed`, `partial_success`, and `failed`, plus page states `pending`, `transcribed`, and `failed`. | inspection |
| REQ-5 | Functional | Allow users to manage historical `Person` records and link each Person to a Document once with exactly one role. | test |
| REQ-6 | Functional | Support an extensible role taxonomy for document-person relationships. | inspection |
| REQ-7 | Policy Constraint | Enforce deterministic relationship-role writes with uniqueness on `(document_id, person_id)` and explicit conflict responses for duplicate Person links. | test |
| REQ-8 | Functional | Use set-based synchronization for document-person mutations so updates add and remove only the intended links. | test |
| REQ-9 | Functional | Maintain selected machine output and exact attempt provenance on `Source` while permitting independent human edits on `Source.revised_text`. | test |
| REQ-10 | Functional | Support a UUID-identified `DocumentType` taxonomy with unique user-facing labels and active/inactive lifecycle control. | test |
| REQ-11 | Data Constraint | Store `Document` type as a controlled reference to `DocumentType`. | test |
| REQ-12 | Interface | Render multi-page transcriptions sequentially by `page_number` with document, people, and document-type metadata. | demonstration |
| REQ-13 | Interface | Document create/edit UI must provide one staged Linked People table and select active registry entries by UUID and label. | demonstration |
| REQ-14 | API Constraint | Expose additive, role-aware retrieval and write behavior for document-person links and UUID-based selection for document types. | test |
| REQ-15 | Data Constraint | Calculate and store cryptographic file hashes (SHA-256) and file sizes for uploaded source images. | test |
| REQ-16 | Data Constraint | Preserve a portable relational model across supported backends using SQLModel, SQLAlchemy, SQLite, and PostgreSQL. | inspection |
| REQ-17 | Reliability | Ensure delete and update flows for documents, people, and relationship links remain deterministic and safe. | test |
| REQ-18 | Operations Constraint | Keep canonical development, testing, restore, and recovery workflows OS-independent; for AI-run unit tests, require a pre-test backup of `./data` and an always-shown post-success confirmation prompt before any restore action. | inspection |
| REQ-19 | Quality | Provide automated coverage for async transcription workflows, relationship-role enforcement, document-type selection, and regression behavior. | test |
| REQ-20 | Data Constraint | Permit hidden immutable semantic keys only on protected built-in Document Types and Person Roles while retaining UUID as relationship identity. | test |
| REQ-21 | Reliability | Persist Document fields and their complete Linked People set atomically. | test |
| REQ-22 | Interface | Provide safe browser-native Facsimile and Text-only print views from persisted Document Detail. | demonstration |
| REQ-23 | Security | Escape stored print text and serve Source images through record-validated application routes without disclosing local paths. | test |
| REQ-24 | Functional | Print current human-preferred Source text, semantic Author metadata, deterministic Source order, and oldest-to-newest Job metadata. | test |
| REQ-25 | Quality | Physically apply recognized raster orientation metadata to provider-input derivatives without changing original Source bytes. | test |
| REQ-26 | Quality | Persist deterministic, non-mutating output warnings without automatic paid retries. | test |
| REQ-27 | Functional | Create one-Source retranscription Jobs from a configured model allowlist and preserve later successes as candidates until explicit promotion. | test |
## Clarifying Constraints
1. `DocumentType.id` and `PersonRole.id` are their public and relationship identities; labels are unique ignoring case and surrounding whitespace.
2. Nullable `semantic_key` values identify protected application built-ins, remain internal, and never change.
3. Relationship-write policy and conflict handling must be consistent across UI, API, services, and persistence.
4. One Person may appear only once per Document and every link has exactly one role.
5. Relationship conflicts must fail deterministically without partial Document or link mutation.
6. Source page reordering and server-generated PDF files remain outside this revision.
## Element Satisfaction Mapping
- UI (NiceGUI): Satisfies REQ-0, REQ-1, REQ-5, REQ-9, REQ-12, REQ-13, REQ-22, REQ-24.
- API (FastAPI): Satisfies REQ-1, REQ-4, REQ-5, REQ-7, REQ-8, REQ-14, REQ-23.
- Worker (`asyncio`): Satisfies REQ-2, REQ-3, REQ-4.
- Persistence (SQLModel / SQLAlchemy): Satisfies REQ-3, REQ-7, REQ-9, REQ-10, REQ-11, REQ-15, REQ-16, REQ-17, REQ-20, REQ-21.
- Test Suite: Verifies all test-marked requirements and satisfies REQ-19.
## Related Local References
- [System Overview](index_v4.md)
- [System Architecture](architecture_v4.md)
- [Data Model](schema_v4.md)
- [Error Handling Policy](error_handling_v4.md)
-258
View File
@@ -1,258 +0,0 @@
# Database Schema (Version 4)
This document defines the relational schema for the document transcription system.
## Entity Relationship Diagram
```mermaid
erDiagram
DOCUMENT_TYPE {
UUID id PK
TEXT semantic_key UK
TEXT label
TEXT normalized_label
BOOLEAN is_active
TIMESTAMPTZ created_at
TIMESTAMPTZ updated_at
}
PERSON_ROLE {
UUID id PK
TEXT semantic_key UK
TEXT label
TEXT normalized_label
BOOLEAN is_active
TIMESTAMPTZ created_at
TIMESTAMPTZ updated_at
}
PERSON {
UUID id PK
TEXT full_name
TEXT display_name
TEXT maiden_name
DATE birth_date
TEXT birth_date_raw
TEXT birth_place
DATE death_date
TEXT death_date_raw
TEXT death_place
TEXT biography
TEXT portrait_path
JSONB metadata
TIMESTAMPTZ created_at
TIMESTAMPTZ updated_at
}
DOCUMENT {
UUID id PK
UUID document_type_id FK
TEXT name
DATE document_date
TEXT document_date_raw
TEXT location_created
TEXT notes
TEXT archive_identifier
TIMESTAMPTZ created_at
TIMESTAMPTZ updated_at
}
DOCUMENT_PERSON {
UUID id PK
UUID document_id FK
UUID person_id FK
UUID role_id FK
TIMESTAMPTZ created_at
TIMESTAMPTZ updated_at
}
JOB {
UUID id PK
UUID document_id FK
VARCHAR status
INTEGER retry_count
VARCHAR purpose
TEXT provider
TEXT model
TEXT prompt_name
TEXT prompt_hash
TEXT system_prompt
TEXT user_prompt
FLOAT temperature
FLOAT top_p
TIMESTAMPTZ date_created
TIMESTAMPTZ date_updated
}
SOURCE {
UUID id PK
UUID document_id FK
INTEGER page_number
TEXT upload_name
TEXT filename
TEXT file_path
TEXT file_hash
BIGINT file_size_bytes
TEXT raw_transcription
UUID preferred_execution_attempt_id FK
TEXT revised_text
TIMESTAMPTZ date_uploaded
TIMESTAMPTZ date_revised
}
JOB_SOURCE {
UUID id PK
UUID job_id FK
UUID source_id FK
VARCHAR status
TEXT raw_transcription
JSONB ai_metadata
JSONB raw_api_response
TEXT error_detail
TIMESTAMPTZ executed_at
}
EXECUTION_ATTEMPT {
UUID id PK
UUID job_source_id FK
UUID job_id FK
UUID source_id FK
INTEGER attempt_number
VARCHAR status
JSONB request_manifest
TEXT request_manifest_sha256
INTEGER transport_status_code
BINARY transport_body
JSONB transport_safe_headers
JSONB sdk_response_snapshot
JSONB normalized_metadata
JSONB software_context
TEXT raw_transcription
TEXT failure_phase
TIMESTAMPTZ started_at
TIMESTAMPTZ finished_at
INTEGER duration_ms
}
PROCESSING_ARTIFACT {
UUID id PK
UUID source_id FK
UUID execution_attempt_id FK
TEXT artifact_type
TEXT media_type
TEXT schema_name
TEXT schema_version
TEXT producer
TEXT producer_version
JSONB inline_payload
TEXT external_reference
TEXT payload_sha256
BIGINT byte_size
JSONB coordinate_metadata
TIMESTAMPTZ created_at
}
DOCUMENT_TYPE ||--o{ DOCUMENT : classifies
DOCUMENT ||--o{ DOCUMENT_PERSON : has_people
PERSON ||--o{ DOCUMENT_PERSON : appears_in
PERSON_ROLE ||--o{ DOCUMENT_PERSON : labels
DOCUMENT ||--o{ JOB : has_jobs
DOCUMENT ||--o{ SOURCE : contains_pages
JOB ||--o{ JOB_SOURCE : executes
SOURCE ||--o{ JOB_SOURCE : processed_in
JOB_SOURCE ||--o{ EXECUTION_ATTEMPT : projects
SOURCE ||--o{ PROCESSING_ARTIFACT : derives
EXECUTION_ATTEMPT ||--o{ PROCESSING_ARTIFACT : produces
```
## Domain Invariants and Provenance Rules
### Page-Level Execution and AI Outputs
- Every single page execution by an AI model produces a dedicated `JOB_SOURCE` record.
- Every `JOB` stores the frozen prompt identifier, prompt text, and hyperparameters used at submission time.
- `JOB_SOURCE.raw_api_response` is a compatibility projection containing an SDK-serialized OpenRouter response
snapshot. It is neither the exact HTTP body nor the native upstream-provider response.
- Every new provider call creates an immutable `EXECUTION_ATTEMPT` containing the frozen request manifest,
exact OpenRouter-boundary response bytes when received, safe transport metadata, SDK snapshot, normalized
metadata, timing, and outcome.
- `EXECUTION_ATTEMPT(job_id, source_id, attempt_number)` is unique; retries increment the persisted attempt number.
- Historical `JOB_SOURCE` rows without an `EXECUTION_ATTEMPT` remain SDK snapshots and are explicitly labeled as
lacking transport evidence.
- `SOURCE.raw_transcription` caches the explicitly selected preferred machine output for that page.
- `SOURCE.preferred_execution_attempt_id` records exact successful-attempt provenance. Legacy projections may remain
null until a new successful result is selected.
### Generic Processing Artifacts
- `PROCESSING_ARTIFACT` stores provider-neutral versioned derived outputs.
- Exactly one of `inline_payload` and `external_reference` is populated.
- Externally stored artifacts use application-managed relative references and are verified by SHA-256 and byte size.
- Coordinate metadata declares units, origin, dimensions, and transformations when geometry is present.
- Orientation-normalized binary model inputs and JSON quality-warning results use distinct versioned artifact types
and are attached to the exact consuming `EXECUTION_ATTEMPT`.
### Image Storage and Integrity
- Binary images are stored on disk; `SOURCE.file_path` stores the persisted path.
- `SOURCE.file_hash` stores a SHA-256 digest.
- `SOURCE.file_size_bytes` stores the original file size.
### Page Ordering and Revisions
- `SOURCE.page_number` dictates page ordering within a document.
- `SOURCE.raw_transcription` changes only through first-success selection or explicit candidate promotion.
- `SOURCE.revised_text` stores human edits and is the preferred display value when present.
### Semantic Registry Governance
- `DOCUMENT_TYPE.id` and `PERSON_ROLE.id` are the only relationship and public API identities.
- Nullable unique `semantic_key` values identify application-defined built-ins and are immutable after creation.
- Semantic keys are internal and are never accepted from Settings or public relationship APIs.
- A non-null semantic key marks a protected built-in; built-ins may be relabeled or disabled but not deleted.
- Custom entries have null semantic keys and may be deleted only when unreferenced.
- Labels are mutable display text and are unique after trimming and case normalization.
- Inactive entries remain valid for historical rows but are excluded from new-assignment selectors.
### Document-Person Role Governance
- Documents support zero or one relationship for each Person.
- Relationship roles are defined by `PERSON_ROLE` rather than hardcoded columns.
- `DOCUMENT_PERSON.role_id` is required.
- `DOCUMENT_PERSON` must be unique for `(document_id, person_id)`.
- Complete link sets and Document fields are validated and persisted in one atomic transaction.
- Existing inactive roles may remain unchanged; new or changed assignments require active roles.
### Document Type Governance
- Every document type is defined by `DOCUMENT_TYPE`.
- `DOCUMENT_TYPE.id` is the relationship identity; hidden semantic keys identify protected built-in meaning.
- `DOCUMENT_TYPE.label` is mutable display text and is unique after trimming and case normalization.
- `DOCUMENT_TYPE.normalized_label` stores the normalized uniqueness key.
- Inactive types remain valid for historical rows but should be excluded from default selection UIs.
## Constraint Summary
- `DOCUMENT_TYPE.normalized_label` is unique.
- `DOCUMENT_TYPE.semantic_key` is nullable and unique.
- `PERSON_ROLE.normalized_label` is unique.
- `PERSON_ROLE.semantic_key` is nullable and unique.
- `DOCUMENT_PERSON(document_id, person_id)` is unique.
## Indexing Guidance
- `document(document_type_id)`
- `document_person(document_id)`
- `document_person(person_id)`
- `document_person(role_id)`
- `source(document_id, page_number)`
- `job(document_id, status)`
- `job_source(job_id)`
- `job_source(source_id)`
## Related Local References
- [System Overview](index_v4.md)
- [System Architecture](architecture_v4.md)
- [System Requirements](requirements_v4.md)
- [Error Handling Policy](error_handling_v4.md)
-89
View File
@@ -1,89 +0,0 @@
# V4 Scope Boundary
This document defines the scope for the transition from the current repository state to the Version 4 project definition.
## Purpose
Define what this revision includes, what it intentionally excludes, and what migration rules govern the transition work.
## In Scope
### 1. Relationship Model
- Extensible role taxonomy for document-person relationships.
- Many-to-many document-person links with many people per role.
- Set-based add/remove synchronization for document-person updates.
### 2. Document Type Governance
- Registry-driven `DocumentType` model with UUID identity, unique labels, and controlled selection.
- Minimal rollout for the current corpus with no alias helper table.
### 3. UI and API Behavior
- Grouped role links on document and person views.
- Multi-role relationship editing on document create/edit flows.
- Role-aware API retrieval and write behavior.
- Additive API evolution with explicit deprecations.
### 4. Verification
- Tests for many-per-role behavior.
- Tests for set-based relationship mutation behavior.
- Tests for document and person delete/link cleanup regressions.
## Out of Scope
- Suggested versus asserted relationship states.
- Suggestion storage, review, acceptance, or rejection workflows.
- Automatic relationship extraction or recommendation features.
- Full entity resolution or identity merge across all people.
- Automated semantic document type classification.
- Redesign of the core transcription execution model.
## Locked Design Decisions
### A. Role Extensibility Mechanism
- Use registry tables for relationship roles.
### B. API Compatibility Strategy
- Use additive API evolution.
- In development mode, the current revision is authoritative.
- Deprecations should be explicit and short-lived.
### C. Document Type Rollout Strategy
- Use a minimal registry rollout for the current corpus.
- Do not introduce a `document_type_alias` helper table.
### D. Database Change Policy
- Future schema changes are non-destructive by default.
- Exception: `document_type` text may be replaced by `document_type_id` without migrating the legacy text values.
- Exception: `document_person` links may be recreated manually.
## Compatibility and Rollout
- Preserve existing repository behavior where unaffected by the V4 scope.
- Treat scope boundary and implementation plan as the only transition documents.
- Treat core V4 documents as the authoritative project definition once rewritten.
## Exit Criteria for Scope Freeze
V4 scope is considered frozen when:
- Relationship model and document-type governance are approved.
- Relationship model and document-type governance are approved.
- Additive API change list and deprecation schedule are approved.
- Migration exceptions are explicitly acknowledged.
## Core V4 Documents
1. `docs/ver4/index_v4.md`
2. `docs/ver4/requirements_v4.md`
3. `docs/ver4/schema_v4.md`
4. `docs/ver4/architecture_v4.md`
5. `docs/ver4/error_handling_v4.md`
6. `docs/ver4/implementation_plan_v4.md`
+5
View File
@@ -15,12 +15,17 @@ dependencies = [
"aiosqlite>=0.21.0", "aiosqlite>=0.21.0",
"asyncpg>=0.31.0", "asyncpg>=0.31.0",
"fastapi>=0.138.0", "fastapi>=0.138.0",
# Exact pin, deliberate. NiceGUI 3.x minor releases ship Quasar/Vue changes that
# break component props and styling, and tests/ui/ cannot detect visual regressions.
# Hold through the current release stabilization; revisit as a scheduled upgrade.
# See docs/production-runbook.md, "Dependency upgrade policy".
"nicegui==3.13.0", "nicegui==3.13.0",
"openrouter>=0.7.0", "openrouter>=0.7.0",
"pillow>=10.0.0", "pillow>=10.0.0",
"psycopg2-binary>=2.9.12", "psycopg2-binary>=2.9.12",
"pydantic>=2.13.4", "pydantic>=2.13.4",
"pydantic-settings>=2.9.1", "pydantic-settings>=2.9.1",
"python-gedcom>=1.1.0",
"sqlmodel>=0.0.25", "sqlmodel>=0.0.25",
] ]
Binary file not shown.
+1
View File
@@ -26,6 +26,7 @@ extend-select = [
"E", "W", # https://docs.astral.sh/ruff/rules/#pycodestyle-e-w "E", "W", # https://docs.astral.sh/ruff/rules/#pycodestyle-e-w
"F", # https://docs.astral.sh/ruff/rules/#pyflakes-f "F", # https://docs.astral.sh/ruff/rules/#pyflakes-f
"FURB", # https://docs.astral.sh/ruff/rules/#refurb-furb "FURB", # https://docs.astral.sh/ruff/rules/#refurb-furb
"G", # https://docs.astral.sh/ruff/rules/#flake8-logging-format-g
"I", # https://docs.astral.sh/ruff/rules/#isort-i "I", # https://docs.astral.sh/ruff/rules/#isort-i
"N", # https://docs.astral.sh/ruff/rules/#pep8-naming-n "N", # https://docs.astral.sh/ruff/rules/#pep8-naming-n
"PD", # https://docs.astral.sh/ruff/rules/#pandas-vet-pd "PD", # https://docs.astral.sh/ruff/rules/#pandas-vet-pd
@@ -1,4 +1,4 @@
"""Additive V4 API routes for relationship and classification registries.""" """API routes for relationship and classification registries."""
from __future__ import annotations from __future__ import annotations
@@ -20,7 +20,7 @@ from transcription.db.models import PersonRole
from transcription.services import DocumentService from transcription.services import DocumentService
from transcription.services import PeopleService from transcription.services import PeopleService
router = APIRouter(prefix="/api/v4", tags=["v4-documents"]) router = APIRouter(prefix="/api", tags=["documents"])
class ApiModel(BaseModel): class ApiModel(BaseModel):
@@ -72,19 +72,17 @@ class DocumentPeopleResponse(ApiModel):
def _document_type_to_read(item: DocumentType) -> DocumentTypeRead: def _document_type_to_read(item: DocumentType) -> DocumentTypeRead:
return DocumentTypeRead( item_id, label, is_active = _registry_read_values(item)
id=item.id, return DocumentTypeRead(id=item_id, label=label, is_active=is_active)
label=item.label,
is_active=item.is_active,
)
def _person_role_to_read(item: PersonRole) -> PersonRoleRead: def _person_role_to_read(item: PersonRole) -> PersonRoleRead:
return PersonRoleRead( item_id, label, is_active = _registry_read_values(item)
id=item.id, return PersonRoleRead(id=item_id, label=label, is_active=is_active)
label=item.label,
is_active=item.is_active,
) def _registry_read_values(item: DocumentType | PersonRole) -> tuple[UUID, str, bool]:
return item.id, item.label, item.is_active
def _document_person_to_read(item: DocumentPerson) -> DocumentPersonRead: def _document_person_to_read(item: DocumentPerson) -> DocumentPersonRead:
+2
View File
@@ -21,7 +21,9 @@ _STATUS_BY_CATEGORY: dict[ErrorCategory, int] = {
ErrorCategory.NOT_FOUND: 404, ErrorCategory.NOT_FOUND: 404,
ErrorCategory.CONFLICT: 409, ErrorCategory.CONFLICT: 409,
ErrorCategory.EXTERNAL_PROVIDER: 503, ErrorCategory.EXTERNAL_PROVIDER: 503,
ErrorCategory.EXTERNAL_TIMEOUT: 503,
ErrorCategory.INFRA_TRANSIENT: 503, ErrorCategory.INFRA_TRANSIENT: 503,
ErrorCategory.PROCESSING: 500,
ErrorCategory.INFRA_PERSISTENT: 500, ErrorCategory.INFRA_PERSISTENT: 500,
ErrorCategory.INTERNAL_UNEXPECTED: 500, ErrorCategory.INTERNAL_UNEXPECTED: 500,
} }
+33 -5
View File
@@ -1,16 +1,44 @@
"""Health endpoint routes.""" """Health endpoint routes."""
from typing import NotRequired
from typing import TypedDict
from fastapi import APIRouter from fastapi import APIRouter
from fastapi import Request
from transcription.worker import resolve_worker_health
router = APIRouter() router = APIRouter()
def healthz() -> dict[str, str]: class WorkerHealthPayload(TypedDict):
"""Return a simple health status payload.""" state: str
return {"status": "ok"} error_id: NotRequired[str]
error_category: NotRequired[str]
class HealthPayload(TypedDict):
status: str
worker: WorkerHealthPayload
def healthz(request: Request) -> HealthPayload:
"""Return health status with worker-liveness signal."""
worker = resolve_worker_health(request.app.state)
payload: HealthPayload = {
"status": "ok",
"worker": {
"state": worker.state,
},
}
if worker.error_id is not None:
payload["worker"]["error_id"] = worker.error_id
if worker.error_category is not None:
payload["worker"]["error_category"] = worker.error_category
return payload
@router.get("/healthz") @router.get("/healthz")
def healthz_route() -> dict[str, str]: def healthz_route(request: Request) -> HealthPayload:
"""Route wrapper for health status payload.""" """Route wrapper for health status payload."""
return healthz() return healthz(request)
@@ -1,4 +1,4 @@
"""Safe media route for V4.4 Document print previews.""" """Safe media route for Document print previews."""
from __future__ import annotations from __future__ import annotations
@@ -15,7 +15,7 @@ from fastapi.responses import FileResponse
from transcription.services.source_media import SOURCE_MIME_TYPES from transcription.services.source_media import SOURCE_MIME_TYPES
from transcription.services.sources import SourceService from transcription.services.sources import SourceService
router = APIRouter(prefix="/api/v4", tags=["v4-print"]) router = APIRouter(prefix="/api", tags=["print"])
def get_source_service(request: Request) -> SourceService: def get_source_service(request: Request) -> SourceService:
@@ -39,8 +39,8 @@ async def read_document_source_media(
if source.document_id != document_id: if source.document_id != document_id:
raise HTTPException(status_code=404, detail="Source not found for Document") raise HTTPException(status_code=404, detail="Source not found for Document")
path = Path(source.file_path).resolve()
upload_root = service.settings.upload_dir.resolve() upload_root = service.settings.upload_dir.resolve()
path = (upload_root / Path(source.file_path)).resolve()
try: try:
path.relative_to(upload_root) path.relative_to(upload_root)
except ValueError as exc: except ValueError as exc:
+24 -14
View File
@@ -14,16 +14,18 @@ from fastapi import status
from fastapi.responses import RedirectResponse from fastapi.responses import RedirectResponse
from fastapi.staticfiles import StaticFiles from fastapi.staticfiles import StaticFiles
from .api.documents_api import router as documents_router
from .api.errors import register_error_handlers from .api.errors import register_error_handlers
from .api.health import router as health_router from .api.health import router as health_router
from .api.v4_documents import router as v4_documents_router from .api.print_api import router as print_router
from .api.v4_print import router as v4_print_router
from .config import Settings from .config import Settings
from .config import configure_logging from .config import configure_logging
from .config import get_settings from .config import get_settings
from .db import create_all from .db import create_all
from .db import dispose_database_runtime from .db import dispose_database_runtime
from .db import initialize_database_runtime from .db import initialize_database_runtime
from .db import reconcile_canonical_media_paths
from .db import reconcile_legacy_job_source_columns
from .services import ServiceBundle from .services import ServiceBundle
from .ui import register_pages from .ui import register_pages
from .worker import worker_consumer_lifespan from .worker import worker_consumer_lifespan
@@ -42,33 +44,41 @@ async def _lifespan(app: FastAPI):
if settings.should_bootstrap_schema: if settings.should_bootstrap_schema:
await create_all(engine=app.state.runtime.engine) await create_all(engine=app.state.runtime.engine)
await reconcile_legacy_job_source_columns(engine=app.state.runtime.engine)
await reconcile_canonical_media_paths(engine=app.state.runtime.engine)
settings.upload_dir.mkdir(parents=True, exist_ok=True) settings.upload_dir.mkdir(parents=True, exist_ok=True)
settings.prompt_dir.mkdir(parents=True, exist_ok=True) settings.prompt_dir.mkdir(parents=True, exist_ok=True)
settings.log_dir.mkdir(parents=True, exist_ok=True)
await _recover_stale_processing_jobs(app) await _recover_stale_processing_jobs(app)
async with AsyncExitStack() as stack: async with AsyncExitStack() as stack:
stack.push_async_callback(dispose_database_runtime) stack.push_async_callback(dispose_database_runtime)
stop_event, worker_notifier = await stack.enter_async_context( if settings.run_embedded_worker:
worker_consumer_lifespan( stop_event, worker_notifier, worker_health = await stack.enter_async_context(
session_factory=app.state.runtime.session_factory, worker_consumer_lifespan(
poll_interval_seconds=1.0, session_factory=app.state.runtime.session_factory,
poll_interval_seconds=settings.worker_poll_interval_seconds,
shutdown_timeout_seconds=(
settings.worker_provider_timeout_seconds + settings.worker_shutdown_grace_seconds
),
)
) )
) app.state.worker_stop_event = stop_event
app.state.worker_stop_event = stop_event app.state.worker_notifier = worker_notifier
app.state.worker_notifier = worker_notifier app.state.worker_health = worker_health
yield yield
async def _recover_stale_processing_jobs(app: FastAPI) -> None: async def _recover_stale_processing_jobs(app: FastAPI) -> None:
"""Re-queue stale processing jobs at startup. """Re-queue stale processing jobs at startup.
Any job left in PROCESSING longer than the configured provider timeout is Any job left in PROCESSING longer than the stale-job threshold is assumed
assumed orphaned and moved back to QUEUED before the worker starts. orphaned and moved back to QUEUED before the worker starts.
""" """
settings = app.state.settings settings = app.state.settings
stale_before = datetime.now(UTC) - timedelta(seconds=settings.worker_provider_timeout_seconds) stale_before = datetime.now(UTC) - timedelta(seconds=settings.worker_stale_job_seconds)
recovered = await app.state.services.jobs.requeue_stale_processing_jobs(stale_before=stale_before) recovered = await app.state.services.jobs.requeue_stale_processing_jobs(stale_before=stale_before)
if recovered > 0: if recovered > 0:
logger.warning("Recovered %s stale processing job(s) at startup", recovered) logger.warning("Recovered %s stale processing job(s) at startup", recovered)
@@ -95,7 +105,7 @@ def create_app(settings: Settings | None = None) -> FastAPI:
register_error_handlers(app) register_error_handlers(app)
app.include_router(health_router) app.include_router(health_router)
app.include_router(v4_documents_router) app.include_router(documents_router)
app.include_router(v4_print_router) app.include_router(print_router)
register_pages(app) register_pages(app)
return app return app
+8 -19
View File
@@ -1,4 +1,11 @@
"""Private-corpus benchmark contracts and deterministic text scoring.""" """Private-corpus benchmark contracts and deterministic text scoring.
This module is retained as the implementation of the evaluation policy in
`docs/invariant/ai_evidence_and_provenance.md` §5. Application runtime paths do
not call it directly, but preserved execution attempts and manually reviewed
references need a deterministic scorer that remains importable for tests and
operator tooling.
"""
from __future__ import annotations from __future__ import annotations
@@ -13,24 +20,6 @@ class BenchmarkModel(BaseModel):
model_config = ConfigDict(extra="forbid", frozen=True) model_config = ConfigDict(extra="forbid", frozen=True)
class BenchmarkItem(BenchmarkModel):
"""One private benchmark item referenced by archival identity."""
source_id: UUID
source_digest_sha256: str = Field(pattern=r"^[0-9a-f]{64}$")
categories: frozenset[str] = Field(min_length=1)
reference_transcription: str = Field(min_length=1)
class BenchmarkManifest(BenchmarkModel):
"""Versioned private benchmark definition without copied source media."""
schema_name: str = "transcription.private-benchmark"
schema_version: str = "1"
name: str = Field(min_length=1)
items: tuple[BenchmarkItem, ...] = Field(min_length=1)
class EditorialAssessment(BenchmarkModel): class EditorialAssessment(BenchmarkModel):
"""Manually reviewed errors not represented adequately by CER or WER.""" """Manually reviewed errors not represented adequately by CER or WER."""
+78 -10
View File
@@ -1,11 +1,13 @@
"""Centralized application configuration. """Centralized application configuration.
All settings are loaded from environment variables (or a .env file) All settings are loaded from environment variables (or an env file)
once at startup. Provider-specific defaults (model names, base URLs) once at startup. Provider-specific defaults (model names, base URLs)
are resolved by the provider adapters, not here. are resolved by the provider adapters, not here.
""" """
import copy
import logging.config import logging.config
import os
from collections.abc import Sequence from collections.abc import Sequence
from enum import StrEnum from enum import StrEnum
from functools import cache from functools import cache
@@ -25,6 +27,16 @@ from pydantic_settings import BaseSettings
from pydantic_settings import SettingsConfigDict from pydantic_settings import SettingsConfigDict
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
PROJECT_ROOT = Path(__file__).resolve().parents[2]
DEFAULT_ENV_FILE_NAME = ".env.production"
def resolve_settings_env_file_path() -> Path:
"""Resolve the runtime env file independent of the process working directory."""
override = os.getenv("ENV_FILE", "").strip()
if override:
return Path(override)
return PROJECT_ROOT / DEFAULT_ENV_FILE_NAME
class Provider(StrEnum): class Provider(StrEnum):
@@ -36,13 +48,14 @@ PromptFilename = Annotated[str, StringConstraints(strip_whitespace=True, min_len
Probability = Annotated[float, Field(ge=0.0, le=1.0)] Probability = Annotated[float, Field(ge=0.0, le=1.0)]
Temperature = Annotated[float, Field(ge=0.0, le=2.0)] Temperature = Annotated[float, Field(ge=0.0, le=2.0)]
DEFAULT_PROVIDER_MODEL = "google/gemini-2.5-flash" DEFAULT_PROVIDER_MODEL = "google/gemini-2.5-flash"
WORKER_STALE_TIMEOUT_MULTIPLIER = 3.0
class SqliteSettings(BaseModel): class SqliteSettings(BaseModel):
model_config = ConfigDict(extra="forbid", frozen=True) model_config = ConfigDict(extra="forbid", frozen=True)
driver: Literal["sqlite"] = "sqlite" driver: Literal["sqlite"] = "sqlite"
path: NonEmptyStr = "app.db" path: NonEmptyStr = "./data/transcription.db"
class PostgresSettings(BaseModel): class PostgresSettings(BaseModel):
@@ -64,7 +77,7 @@ DatabaseSettings = Annotated[
class Settings(BaseSettings): class Settings(BaseSettings):
model_config = SettingsConfigDict( model_config = SettingsConfigDict(
env_file=".env", env_file=None,
env_file_encoding="utf-8", env_file_encoding="utf-8",
extra="ignore", extra="ignore",
env_nested_delimiter="__", env_nested_delimiter="__",
@@ -73,11 +86,20 @@ class Settings(BaseSettings):
frozen=True, frozen=True,
) )
def __init__(self, /, **values: Any) -> None:
if "_env_file" not in values:
values["_env_file"] = resolve_settings_env_file_path()
super().__init__(**values)
# --- NiceGUI Server --- # --- NiceGUI Server ---
host: str = "0.0.0.0" host: str = "0.0.0.0"
port: int = 8000 port: int = 8000
log_level: Literal["critical", "error", "warning", "info", "debug", "trace"] = "info" log_level: Literal["critical", "error", "warning", "info", "debug", "trace"] = "info"
reload: bool = False reload: bool = False
log_dir: Path = Path("./data/logs")
log_file_name: NonEmptyStr = "transcription.log"
log_file_max_bytes: int = Field(default=10 * 1024 * 1024, gt=0)
log_file_backup_count: int = Field(default=5, ge=1)
# --- AI provider --- # --- AI provider ---
provider: Provider = Provider.OPENROUTER provider: Provider = Provider.OPENROUTER
@@ -92,6 +114,8 @@ class Settings(BaseSettings):
# --- runtime environment --- # --- runtime environment ---
environment: Literal["development", "test", "production"] = "development" environment: Literal["development", "test", "production"] = "development"
transcription_commit: NonEmptyStr | None = None
run_embedded_worker: bool = True
# --- persistence --- # --- persistence ---
database: DatabaseSettings = Field(default_factory=SqliteSettings) database: DatabaseSettings = Field(default_factory=SqliteSettings)
@@ -99,15 +123,18 @@ class Settings(BaseSettings):
sqlite_check_same_thread: bool = False sqlite_check_same_thread: bool = False
# --- filesystem paths --- # --- filesystem paths ---
upload_dir: Path = Path("./uploads") upload_dir: Path = Path("./data")
prompt_dir: Path = Path("./prompts") prompt_dir: Path = Path("./prompts")
homepage_dir: Path = Path("./data/homepage")
# --- worker reliability --- # --- worker reliability ---
worker_max_retries: int = Field(default=0, ge=0) worker_max_retries: int = Field(default=0, ge=0)
# Bounded only from below. Vision transcription of a dense page routinely runs # Bounded only from below. Vision transcription of a dense page routinely runs
# well past twenty seconds, so an upper cap here would silently fail real work. # well past twenty seconds, so an upper cap here would silently fail real work.
worker_provider_timeout_seconds: float = Field(default=180.0, gt=0.0) worker_provider_timeout_seconds: float = Field(default=30.0, gt=0.0)
worker_stale_job_seconds: float = Field(default=90.0, gt=0.0)
worker_retry_backoff_seconds: float = Field(default=1.0, ge=0.0)
worker_shutdown_grace_seconds: float = Field(default=5.0, ge=0.0)
worker_poll_interval_seconds: float = Field(default=1.0, gt=0.0)
worker_min_transcription_chars: int = Field(default=0, ge=0) worker_min_transcription_chars: int = Field(default=0, ge=0)
worker_min_transcription_lines: int = Field(default=0, ge=0) worker_min_transcription_lines: int = Field(default=0, ge=0)
worker_fail_on_finish_reason_length: bool = False worker_fail_on_finish_reason_length: bool = False
@@ -159,6 +186,34 @@ class Settings(BaseSettings):
return {**data, "provider_model": default_model, "provider_models": tuple(deduplicated)} return {**data, "provider_model": default_model, "provider_models": tuple(deduplicated)}
@model_validator(mode="before")
@classmethod
def _derive_worker_stale_job_seconds(cls, data: object) -> object:
"""Default stale-job recovery with margin over one provider timeout."""
if not isinstance(data, dict):
return data
if data.get("worker_stale_job_seconds") is not None:
return data
timeout = data.get("worker_provider_timeout_seconds", 30.0)
if not isinstance(timeout, (str, int, float)):
return data
try:
timeout_seconds = float(timeout)
except ValueError:
return data
return {
**data,
"worker_stale_job_seconds": timeout_seconds * WORKER_STALE_TIMEOUT_MULTIPLIER,
}
@model_validator(mode="after")
def _validate_worker_stale_job_seconds(self) -> "Settings":
"""Reject stale recovery that can fire before one provider timeout expires."""
if self.worker_stale_job_seconds <= self.worker_provider_timeout_seconds:
raise ValueError("WORKER_STALE_JOB_SECONDS must exceed WORKER_PROVIDER_TIMEOUT_SECONDS")
return self
@property @property
def should_bootstrap_schema(self) -> bool: def should_bootstrap_schema(self) -> bool:
"""Return whether startup should auto-create schema for this environment.""" """Return whether startup should auto-create schema for this environment."""
@@ -193,16 +248,24 @@ LOGGING_CONFIG: dict[str, Any] = {
"class": "logging.StreamHandler", "class": "logging.StreamHandler",
"formatter": "standard", "formatter": "standard",
"stream": "ext://sys.stdout", "stream": "ext://sys.stdout",
} },
"file": {
"class": "logging.handlers.RotatingFileHandler",
"formatter": "standard",
"filename": str(Path("./data/logs") / "transcription.log"),
"maxBytes": 10 * 1024 * 1024,
"backupCount": 5,
"encoding": "utf-8",
},
}, },
"root": { "root": {
"level": "INFO", "level": "INFO",
"handlers": ["console"], "handlers": ["console", "file"],
}, },
"loggers": { "loggers": {
"transcription": { "transcription": {
"level": "DEBUG", "level": "DEBUG",
"handlers": ["console"], "handlers": ["console", "file"],
"propagate": False, "propagate": False,
} }
}, },
@@ -211,8 +274,13 @@ LOGGING_CONFIG: dict[str, Any] = {
def configure_logging(settings: Settings | None = None) -> None: def configure_logging(settings: Settings | None = None) -> None:
"""Configure root logging once at startup.""" """Configure root logging once at startup."""
cfg = LOGGING_CONFIG.copy() cfg = copy.deepcopy(LOGGING_CONFIG)
active_settings = settings or get_settings() active_settings = settings or get_settings()
active_settings.log_dir.mkdir(parents=True, exist_ok=True)
file_handler = cfg["handlers"]["file"]
file_handler["filename"] = str(active_settings.log_dir / active_settings.log_file_name)
file_handler["maxBytes"] = active_settings.log_file_max_bytes
file_handler["backupCount"] = active_settings.log_file_backup_count
cfg["loggers"]["transcription"]["level"] = active_settings.log_level.upper() cfg["loggers"]["transcription"]["level"] = active_settings.log_level.upper()
logging.config.dictConfig(cfg) logging.config.dictConfig(cfg)
logger.debug("Logging configured") logger.debug("Logging configured")
+6
View File
@@ -1,4 +1,7 @@
from .operations import create_all from .operations import create_all
from .operations import reconcile_canonical_media_paths
from .operations import reconcile_legacy_job_source_columns
from .operations import reconcile_person_name_columns
from .runtime import dispose_database_runtime from .runtime import dispose_database_runtime
from .runtime import initialize_database_runtime from .runtime import initialize_database_runtime
from .session import session_scope from .session import session_scope
@@ -8,6 +11,9 @@ __all__ = [
"create_all", "create_all",
"dispose_database_runtime", "dispose_database_runtime",
"initialize_database_runtime", "initialize_database_runtime",
"reconcile_canonical_media_paths",
"reconcile_legacy_job_source_columns",
"reconcile_person_name_columns",
"session_scope", "session_scope",
"transaction_scope", "transaction_scope",
] ]
-11
View File
@@ -73,14 +73,3 @@ async def dispose_engine(database_url: str) -> None:
engine = _ENGINES.pop(database_url, None) engine = _ENGINES.pop(database_url, None)
if engine is not None: if engine is not None:
await engine.dispose() await engine.dispose()
async def dispose_all_engines() -> None:
while _ENGINES:
_, engine = _ENGINES.popitem()
await engine.dispose()
async def refresh_engine(database_url: str) -> AsyncEngine:
await dispose_engine(database_url)
return get_engine(database_url)
+610
View File
@@ -0,0 +1,610 @@
from __future__ import annotations
import base64
import json
import shutil
from collections.abc import Sequence
from dataclasses import dataclass
from datetime import UTC
from datetime import date
from datetime import datetime
from pathlib import Path
from typing import Any
from uuid import UUID
from uuid import uuid4
from sqlalchemy import URL
from sqlalchemy import MetaData
from sqlalchemy import Table
from sqlalchemy import bindparam
from sqlalchemy import create_engine
from sqlalchemy import func
from sqlalchemy import inspect as sqlalchemy_inspect
from sqlalchemy import select
from sqlalchemy import text
from sqlalchemy.engine import RowMapping
from sqlalchemy.engine import make_url
from sqlmodel import SQLModel
from transcription.config import Settings
from transcription.config import get_settings
# Register table metadata.
from transcription.db import models as _models # noqa: F401
from transcription.db.engine import get_database_url
EXPORT_TABLE_ORDER = (
"document_type",
"person_role",
"tag",
"document",
"person",
"genealogy_person",
"genealogy_family",
"photo",
"document_person",
"document_tag",
"person_tag",
"job",
"source",
"job_source",
"execution_attempt",
"genealogy_family_child",
"genealogy_citation",
)
BYTES_FIELDS = {"transport_body"}
VERIFICATION_TABLES = EXPORT_TABLE_ORDER
@dataclass(frozen=True)
class MigrationPaths:
source_db_url: str
target_db_url: str
source_upload_dir: Path
target_upload_dir: Path
bundle_dir: Path
def export_bundle(*, source_db_url: str, source_upload_dir: Path, bundle_dir: Path) -> None:
bundle_dir.mkdir(parents=True, exist_ok=True)
export_json = bundle_dir / "database.json"
uploads_bundle_dir = bundle_dir / "uploads"
payload: dict[str, Any] = {
"schema_name": "transcription.export-import",
"schema_version": "1",
"created_at": datetime.now(UTC).isoformat(),
"tables": {},
}
engine = create_engine(source_db_url)
legacy_portrait_rows: Sequence[RowMapping] = ()
try: # noqa: PLR1702
inspector = sqlalchemy_inspect(engine)
source_tables = set(inspector.get_table_names())
metadata = MetaData()
metadata.reflect(bind=engine)
current_metadata = SQLModel.metadata
with engine.connect() as connection:
for table_name in EXPORT_TABLE_ORDER:
if table_name not in source_tables:
payload["tables"][table_name] = []
continue
source_table = metadata.tables[table_name]
target_table = current_metadata.tables[table_name]
export_columns = [column.name for column in target_table.columns if column.name in source_table.columns]
if table_name == "person" and "full_name" in source_table.columns:
for legacy_column in ("full_name",):
if legacy_column not in export_columns:
export_columns.append(legacy_column)
if table_name == "person" and "portrait_path" in source_table.columns:
legacy_portrait_rows = (
connection.execute(
select(source_table.c["id"], source_table.c["portrait_path"]).where(
source_table.c["portrait_path"].is_not(None)
)
)
.mappings()
.all()
)
rows = connection.execute(select(*(source_table.c[name] for name in export_columns))).mappings().all()
payload["tables"][table_name] = [
_serialize_row(row, table_name=table_name, source_upload_dir=source_upload_dir) for row in rows
]
finally:
engine.dispose()
if uploads_bundle_dir.exists():
shutil.rmtree(uploads_bundle_dir)
if source_upload_dir.exists():
shutil.copytree(source_upload_dir, uploads_bundle_dir)
else:
uploads_bundle_dir.mkdir(parents=True, exist_ok=True)
_prepare_photo_payload_and_uploads(
payload=payload,
uploads_bundle_dir=uploads_bundle_dir,
legacy_portrait_rows=legacy_portrait_rows,
)
_relocate_homepage_markdown(uploads_bundle_dir=uploads_bundle_dir)
export_json.write_text(json.dumps(payload, indent=2), encoding="utf-8")
def import_bundle(*, target_db_url: str, target_upload_dir: Path, bundle_dir: Path) -> None:
export_json = bundle_dir / "database.json"
uploads_bundle_dir = bundle_dir / "uploads"
payload = json.loads(export_json.read_text(encoding="utf-8"))
if target_upload_dir.exists():
shutil.rmtree(target_upload_dir)
target_upload_dir.mkdir(parents=True, exist_ok=True)
if uploads_bundle_dir.exists():
shutil.copytree(uploads_bundle_dir, target_upload_dir, dirs_exist_ok=True)
_reset_sqlite_target_file(target_db_url)
_ensure_sqlite_target_parent_exists(target_db_url)
engine = create_engine(target_db_url)
try:
SQLModel.metadata.create_all(engine)
execution_attempt_ids = _collect_execution_attempt_ids(payload)
with engine.begin() as connection:
for table_name in reversed(EXPORT_TABLE_ORDER):
table = SQLModel.metadata.tables[table_name]
connection.execute(table.delete())
deferred_source_preferred_attempt_updates: list[dict[str, Any]] = []
for table_name in EXPORT_TABLE_ORDER:
rows = payload.get("tables", {}).get(table_name, [])
if not rows:
continue
table = SQLModel.metadata.tables[table_name]
if table_name == "source":
prepared_source_rows, updates = _prepare_source_rows_for_import(
rows=rows,
source_table=table,
execution_attempt_ids=execution_attempt_ids,
)
deferred_source_preferred_attempt_updates.extend(updates)
connection.execute(table.insert(), prepared_source_rows)
continue
connection.execute(table.insert(), [_deserialize_row(row, table) for row in rows])
if deferred_source_preferred_attempt_updates:
source_table = SQLModel.metadata.tables["source"]
connection.execute(
source_table.update()
.where(source_table.c.id == bindparam("source_id"))
.values(preferred_execution_attempt_id=bindparam("preferred_execution_attempt_id")),
deferred_source_preferred_attempt_updates,
)
finally:
engine.dispose()
def _collect_execution_attempt_ids(payload: dict[str, Any]) -> set[str]:
execution_attempt_rows = payload.get("tables", {}).get("execution_attempt", [])
return {_normalize_uuid_like(row.get("id")) for row in execution_attempt_rows if row.get("id") is not None}
def _prepare_source_rows_for_import(
*,
rows: list[dict[str, Any]],
source_table: Table,
execution_attempt_ids: set[str],
) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]:
prepared_source_rows: list[dict[str, Any]] = []
updates: list[dict[str, Any]] = []
for row in rows:
source_row = _deserialize_row(row, source_table)
source_id = source_row.get("id")
preferred_attempt_id = source_row.get("preferred_execution_attempt_id")
if (
source_id is not None
and preferred_attempt_id is not None
and _normalize_uuid_like(preferred_attempt_id) in execution_attempt_ids
):
updates.append(
{
"source_id": source_id,
"preferred_execution_attempt_id": preferred_attempt_id,
}
)
source_row["preferred_execution_attempt_id"] = None
prepared_source_rows.append(source_row)
return prepared_source_rows, updates
def _ensure_sqlite_target_parent_exists(target_db_url: str) -> None:
parsed = make_url(target_db_url)
if not parsed.drivername.startswith("sqlite"):
return
database = parsed.database
if not database or database == ":memory:":
return
Path(database).parent.mkdir(parents=True, exist_ok=True)
def _reset_sqlite_target_file(target_db_url: str) -> None:
parsed = make_url(target_db_url)
if not parsed.drivername.startswith("sqlite"):
return
database = parsed.database
if not database or database == ":memory:":
return
target = Path(database)
if target.exists():
target.unlink()
def migrate_via_bundle(paths: MigrationPaths) -> None:
export_bundle(
source_db_url=paths.source_db_url,
source_upload_dir=paths.source_upload_dir,
bundle_dir=paths.bundle_dir,
)
import_bundle(
target_db_url=paths.target_db_url,
target_upload_dir=paths.target_upload_dir,
bundle_dir=paths.bundle_dir,
)
@dataclass(frozen=True)
class MigrationVerificationReport:
source_counts: dict[str, int]
target_counts: dict[str, int]
mismatched_tables: dict[str, dict[str, int]]
integrity_violations: dict[str, int]
success: bool
def to_dict(self) -> dict[str, Any]:
return {
"success": self.success,
"source_counts": self.source_counts,
"target_counts": self.target_counts,
"mismatched_tables": self.mismatched_tables,
"integrity_violations": self.integrity_violations,
}
def sqlite_url_from_path(path: Path) -> str:
return URL.create(drivername="sqlite", database=str(path)).render_as_string(hide_password=False)
def default_sync_db_url(settings: Settings | None = None) -> str:
runtime_settings = settings or get_settings()
return get_database_url(runtime_settings).replace("+aiosqlite", "").replace("+asyncpg", "").replace("+psycopg", "")
def verify_migration(*, source_db_url: str, target_db_url: str) -> MigrationVerificationReport:
source_counts = _table_counts(source_db_url)
target_counts = _table_counts(target_db_url)
mismatched_tables = {
table_name: {"source": source_counts[table_name], "target": target_counts[table_name]}
for table_name in VERIFICATION_TABLES
if source_counts[table_name] != target_counts[table_name]
}
integrity_violations = _integrity_violations(target_db_url)
success = not mismatched_tables and all(count == 0 for count in integrity_violations.values())
return MigrationVerificationReport(
source_counts=source_counts,
target_counts=target_counts,
mismatched_tables=mismatched_tables,
integrity_violations=integrity_violations,
success=success,
)
def _table_counts(db_url: str) -> dict[str, int]:
engine = create_engine(db_url)
try:
metadata = MetaData()
metadata.reflect(bind=engine)
counts: dict[str, int] = {}
with engine.connect() as connection:
for table_name in VERIFICATION_TABLES:
table = metadata.tables.get(table_name)
if table is None:
counts[table_name] = 0
continue
counts[table_name] = int(connection.execute(select(func.count()).select_from(table)).scalar_one())
return counts
finally:
engine.dispose()
def _integrity_violations(db_url: str) -> dict[str, int]:
checks = {
"orphan_source_document": (
"select count(*) from source s left join document d on d.id = s.document_id where d.id is null"
),
"orphan_job_document": (
"select count(*) from job j left join document d on d.id = j.document_id where d.id is null"
),
"orphan_job_source_job": (
"select count(*) from job_source js left join job j on j.id = js.job_id where j.id is null"
),
"orphan_job_source_source": (
"select count(*) from job_source js left join source s on s.id = js.source_id where s.id is null"
),
"orphan_attempt_job_source": (
"select count(*) from execution_attempt ea "
"left join job_source js on js.id = ea.job_source_id "
"where js.id is null"
),
"orphan_attempt_job": (
"select count(*) from execution_attempt ea left join job j on j.id = ea.job_id where j.id is null"
),
"orphan_attempt_source": (
"select count(*) from execution_attempt ea left join source s on s.id = ea.source_id where s.id is null"
),
"duplicate_attempt_numbers": (
"select count(*) from ("
" select job_id, source_id, attempt_number, count(*) as c"
" from execution_attempt"
" group by job_id, source_id, attempt_number"
" having count(*) > 1"
") x"
),
}
engine = create_engine(db_url)
try:
with engine.connect() as connection:
return {
check_name: int(connection.execute(text(query)).scalar_one()) for check_name, query in checks.items()
}
finally:
engine.dispose()
def _serialize_row(row: RowMapping, *, table_name: str, source_upload_dir: Path) -> dict[str, Any]:
serialized: dict[str, Any] = {}
for raw_key, value in row.items():
key = str(raw_key)
serialized_value = _serialize_value(value)
if table_name == "source" and key == "file_path" and isinstance(serialized_value, str):
serialized[key] = _canonical_media_relative_path(
serialized_value,
source_upload_dir=source_upload_dir,
preferred_prefix="documents/",
)
continue
if table_name == "photo" and key == "path" and isinstance(serialized_value, str):
serialized[key] = _canonical_media_relative_path(
serialized_value,
source_upload_dir=source_upload_dir,
preferred_prefix="photos/",
)
continue
if table_name == "person" and key == "full_name" and isinstance(serialized_value, str):
given_names, last_name = _split_legacy_full_name(serialized_value)
serialized["given_names"] = given_names
serialized["last_name"] = last_name
continue
serialized[key] = serialized_value
if table_name == "person":
serialized["given_names"] = str(serialized.get("given_names") or "").strip()
serialized["last_name"] = str(serialized.get("last_name") or "").strip()
return serialized
def _split_legacy_full_name(full_name: str) -> tuple[str, str]:
tokens = [token for token in full_name.strip().split() if token]
if len(tokens) >= 2:
return (" ".join(tokens[:-1]), tokens[-1])
if len(tokens) == 1:
return (tokens[0], tokens[0])
return ("Unknown", "Unknown")
def _serialize_value(value: Any) -> Any:
if isinstance(value, UUID):
return str(value)
if isinstance(value, (datetime, date)):
return value.isoformat()
if isinstance(value, bytes):
return {"encoding": "base64", "data": base64.b64encode(value).decode("ascii")}
if isinstance(value, dict):
return {str(k): _serialize_value(v) for k, v in value.items()}
if isinstance(value, list):
return [_serialize_value(item) for item in value]
return value
def _deserialize_row(row: dict[str, Any], table: Table) -> dict[str, Any]:
deserialized: dict[str, Any] = {}
for key, value in row.items():
if key in BYTES_FIELDS and isinstance(value, dict) and value.get("encoding") == "base64":
deserialized[key] = base64.b64decode(value["data"])
continue
if key in table.columns:
try:
python_type: type[Any] = table.columns[key].type.python_type
except NotImplementedError:
deserialized[key] = value
continue
deserialized[key] = _deserialize_value(python_type, value)
return deserialized
def _deserialize_value(python_type: type[Any], value: Any) -> Any:
if value is None:
return None
if python_type is UUID and isinstance(value, str):
return UUID(value)
if python_type is datetime and isinstance(value, str):
return datetime.fromisoformat(value)
if python_type is date and isinstance(value, str):
return date.fromisoformat(value)
return value
def _normalize_uuid_like(value: Any) -> str:
if isinstance(value, UUID):
return str(value)
if isinstance(value, str):
text_value = value.strip()
try:
return str(UUID(text_value))
except ValueError:
return text_value
return str(value)
def _canonical_media_relative_path(value: str, *, source_upload_dir: Path, preferred_prefix: str) -> str:
normalized = value.strip().replace("\\", "/")
lowered = normalized.casefold()
upload_root = source_upload_dir.resolve().as_posix().casefold().rstrip("/")
if lowered.startswith(upload_root + "/"):
normalized = normalized[len(source_upload_dir.resolve().as_posix()) + 1 :]
lowered = normalized.casefold()
if lowered.startswith("/uploads/"):
normalized = normalized[len("/uploads/") :]
lowered = normalized.casefold()
elif lowered.startswith("uploads/"):
normalized = normalized[len("uploads/") :]
lowered = normalized.casefold()
elif lowered.startswith("data/"):
normalized = normalized[len("data/") :]
lowered = normalized.casefold()
if preferred_prefix == "persons/" and lowered.startswith("portraits/"):
normalized = "persons/" + normalized[len("portraits/") :]
lowered = normalized.casefold()
for prefix in ("documents/", "photos/", "persons/"):
marker = f"/{prefix}"
index = lowered.find(marker)
if index >= 0:
normalized = normalized[index + 1 :]
lowered = normalized.casefold()
break
if not lowered.startswith(preferred_prefix):
return normalized
return Path(normalized).as_posix()
def _prepare_photo_payload_and_uploads( # noqa: PLR0915
*,
payload: dict[str, Any],
uploads_bundle_dir: Path,
legacy_portrait_rows: Sequence[RowMapping],
) -> None:
photo_rows = payload.setdefault("tables", {}).setdefault("photo", [])
photos_dir = uploads_bundle_dir / "photos"
photos_dir.mkdir(parents=True, exist_ok=True)
# Keep only photo rows whose referenced media exists inside the uploads tree.
# This prevents stale/injected rows from blocking legacy backfill.
retained_rows: list[dict[str, Any]] = []
for row in photo_rows:
path_value = row.get("path")
if not isinstance(path_value, str) or not path_value.strip():
continue
canonical_path = _canonical_media_relative_path(
path_value,
source_upload_dir=uploads_bundle_dir,
preferred_prefix="photos/",
)
candidate = uploads_bundle_dir / canonical_path
if not candidate.exists():
continue
row["path"] = canonical_path
retained_rows.append(row)
photo_rows[:] = retained_rows
existing_homepage_rows = [row for row in photo_rows if row.get("person_id") is None]
existing_person_ids = {str(row["person_id"]) for row in photo_rows if row.get("person_id") is not None}
existing_primary_person_ids = {
str(row["person_id"]) for row in photo_rows if row.get("person_id") is not None and bool(row.get("is_primary"))
}
has_homepage_primary = any(bool(row.get("is_primary")) for row in existing_homepage_rows)
now_iso = datetime.now(UTC).isoformat()
for row in legacy_portrait_rows:
portrait_path = row.get("portrait_path")
person_id = row.get("id")
if not isinstance(portrait_path, str) or not portrait_path.strip():
continue
if person_id is None:
continue
canonical = _canonical_media_relative_path(
portrait_path,
source_upload_dir=uploads_bundle_dir,
preferred_prefix="persons/",
)
source_file = uploads_bundle_dir / canonical
if not source_file.exists():
continue
person_key = str(person_id)
if person_key in existing_person_ids:
continue
suffix = Path(canonical).suffix.lower() or ".jpg"
photo_id = str(uuid4())
relative_path = f"photos/{photo_id}{suffix}"
target_file = uploads_bundle_dir / relative_path
target_file.parent.mkdir(parents=True, exist_ok=True)
shutil.copy2(source_file, target_file)
is_primary = person_key not in existing_primary_person_ids
photo_rows.append(
{
"id": photo_id,
"person_id": person_key,
"path": relative_path,
"description": None,
"is_primary": is_primary,
"created_at": now_iso,
"updated_at": now_iso,
}
)
existing_person_ids.add(person_key)
if is_primary:
existing_primary_person_ids.add(person_key)
legacy_homepage_dir = uploads_bundle_dir / "homepage"
if not legacy_homepage_dir.exists():
return
homepage_images = sorted(
[
path
for path in legacy_homepage_dir.iterdir()
if path.is_file()
and path.suffix.lower() in {".jpg", ".jpeg", ".png", ".gif", ".webp", ".bmp", ".tif", ".tiff"}
],
key=lambda path: (path.stat().st_mtime, path.name),
)
if existing_homepage_rows:
return
for index, image_path in enumerate(homepage_images):
photo_id = str(uuid4())
relative_path = f"photos/{photo_id}{image_path.suffix.lower()}"
target_file = uploads_bundle_dir / relative_path
target_file.parent.mkdir(parents=True, exist_ok=True)
shutil.copy2(image_path, target_file)
photo_rows.append(
{
"id": photo_id,
"person_id": None,
"path": relative_path,
"description": None,
"is_primary": (not has_homepage_primary) and index == 0,
"created_at": now_iso,
"updated_at": now_iso,
}
)
def _relocate_homepage_markdown(*, uploads_bundle_dir: Path) -> None:
legacy_markdown = uploads_bundle_dir / "homepage" / "homepage.md"
target_markdown = uploads_bundle_dir / "homepage.md"
if not legacy_markdown.exists() or target_markdown.exists():
return
target_markdown.parent.mkdir(parents=True, exist_ok=True)
shutil.copy2(legacy_markdown, target_markdown)
+356 -32
View File
@@ -46,6 +46,11 @@ def _loaded_attribute(instance: object, attribute: str) -> Any | None:
return state.dict.get(attribute) return state.dict.get(attribute)
def _utc_now_naive() -> datetime:
"""Return a UTC timestamp stored as a naive datetime."""
return datetime.now(UTC).replace(tzinfo=None)
class JSONBCompat(TypeDecorator): class JSONBCompat(TypeDecorator):
"""JSONB for PostgreSQL and JSON for SQLite/testing backends.""" """JSONB for PostgreSQL and JSON for SQLite/testing backends."""
@@ -61,7 +66,6 @@ class JobStatus(StrEnum):
QUEUED = "queued" QUEUED = "queued"
PROCESSING = "processing" PROCESSING = "processing"
TRANSCRIBED = "transcribed" TRANSCRIBED = "transcribed"
COMPLETED = "completed"
PARTIAL_SUCCESS = "partial_success" PARTIAL_SUCCESS = "partial_success"
FAILED = "failed" FAILED = "failed"
@@ -78,6 +82,31 @@ class JobPurpose(StrEnum):
RETRANSCRIPTION = "retranscription" RETRANSCRIPTION = "retranscription"
class MaintenanceJobType(StrEnum):
BACKUP = "backup"
STORAGE_RECONCILIATION = "storage_reconciliation"
GEDCOM_IMPORT = "gedcom_import"
class MaintenanceRunStatus(StrEnum):
QUEUED = "queued"
PROCESSING = "processing"
SUCCEEDED = "succeeded"
FAILED = "failed"
class GenealogyCitationFactType(StrEnum):
BIRTH = "birth"
DEATH = "death"
MARRIAGE = "marriage"
OTHER = "other"
class GenealogyCitationSourceKind(StrEnum):
FAMILYSEARCH_IMPORTED = "familysearch_imported"
TRANSCRIPTION_EVIDENCE = "transcription_evidence"
class DocumentType(SQLModel, table=True): class DocumentType(SQLModel, table=True):
"""Registry of allowed document types.""" """Registry of allowed document types."""
@@ -88,10 +117,10 @@ class DocumentType(SQLModel, table=True):
label: str label: str
normalized_label: str = Field(index=True, unique=True) normalized_label: str = Field(index=True, unique=True)
is_active: bool = True is_active: bool = True
created_at: datetime = Field(default_factory=lambda: datetime.now(UTC)) created_at: datetime = Field(default_factory=_utc_now_naive)
updated_at: datetime = Field( updated_at: datetime = Field(
default_factory=lambda: datetime.now(UTC), default_factory=_utc_now_naive,
sa_column_kwargs={"onupdate": lambda: datetime.now(UTC)}, sa_column_kwargs={"onupdate": _utc_now_naive},
) )
documents: list["Document"] = Relationship( documents: list["Document"] = Relationship(
@@ -109,10 +138,10 @@ class PersonRole(SQLModel, table=True):
label: str label: str
normalized_label: str = Field(index=True, unique=True) normalized_label: str = Field(index=True, unique=True)
is_active: bool = True is_active: bool = True
created_at: datetime = Field(default_factory=lambda: datetime.now(UTC)) created_at: datetime = Field(default_factory=_utc_now_naive)
updated_at: datetime = Field( updated_at: datetime = Field(
default_factory=lambda: datetime.now(UTC), default_factory=_utc_now_naive,
sa_column_kwargs={"onupdate": lambda: datetime.now(UTC)}, sa_column_kwargs={"onupdate": _utc_now_naive},
) )
document_people: list["DocumentPerson"] = Relationship( document_people: list["DocumentPerson"] = Relationship(
@@ -120,6 +149,32 @@ class PersonRole(SQLModel, table=True):
) )
class Tag(SQLModel, table=True):
"""Registry of labels that can be attached to Documents."""
__tablename__ = "tag"
id: UUID = Field(default_factory=uuid4, primary_key=True)
semantic_key: str | None = Field(default=None, index=True, unique=True)
label: str
normalized_label: str = Field(index=True, unique=True)
is_active: bool = True
created_at: datetime = Field(default_factory=_utc_now_naive)
updated_at: datetime = Field(
default_factory=_utc_now_naive,
sa_column_kwargs={"onupdate": _utc_now_naive},
)
document_tags: list["DocumentTag"] = Relationship(
back_populates="tag_ref",
sa_relationship_kwargs={"lazy": "raise"},
)
person_tags: list["PersonTag"] = Relationship(
back_populates="tag_ref",
sa_relationship_kwargs={"lazy": "raise"},
)
class Document(SQLModel, table=True): class Document(SQLModel, table=True):
"""An historical document.""" """An historical document."""
@@ -131,10 +186,10 @@ class Document(SQLModel, table=True):
location_created: str | None = None location_created: str | None = None
notes: str | None = None notes: str | None = None
archive_identifier: str | None = None archive_identifier: str | None = None
created_at: datetime = Field(default_factory=lambda: datetime.now(UTC)) created_at: datetime = Field(default_factory=_utc_now_naive)
updated_at: datetime = Field( updated_at: datetime = Field(
default_factory=lambda: datetime.now(UTC), default_factory=_utc_now_naive,
sa_column_kwargs={"onupdate": lambda: datetime.now(UTC)}, sa_column_kwargs={"onupdate": _utc_now_naive},
) )
jobs: list["Job"] = Relationship(back_populates="document", sa_relationship_kwargs={"lazy": "raise"}) jobs: list["Job"] = Relationship(back_populates="document", sa_relationship_kwargs={"lazy": "raise"})
@@ -142,6 +197,10 @@ class Document(SQLModel, table=True):
document_people: list["DocumentPerson"] = Relationship( document_people: list["DocumentPerson"] = Relationship(
back_populates="document", sa_relationship_kwargs={"lazy": "raise"} back_populates="document", sa_relationship_kwargs={"lazy": "raise"}
) )
document_tags: list["DocumentTag"] = Relationship(
back_populates="document",
sa_relationship_kwargs={"lazy": "raise"},
)
document_type_ref: Optional["DocumentType"] = Relationship( document_type_ref: Optional["DocumentType"] = Relationship(
back_populates="documents", sa_relationship_kwargs={"lazy": "raise"} back_populates="documents", sa_relationship_kwargs={"lazy": "raise"}
) )
@@ -151,9 +210,8 @@ class Person(SQLModel, table=True):
"""A historical person linked to one or more documents.""" """A historical person linked to one or more documents."""
id: UUID = Field(default_factory=uuid4, primary_key=True) id: UUID = Field(default_factory=uuid4, primary_key=True)
full_name: str last_name: str
display_name: str | None = None given_names: str
maiden_name: str | None = None
birth_date: date | None = None birth_date: date | None = None
birth_date_raw: str | None = None birth_date_raw: str | None = None
birth_place: str | None = None birth_place: str | None = None
@@ -161,21 +219,195 @@ class Person(SQLModel, table=True):
death_date_raw: str | None = None death_date_raw: str | None = None
death_place: str | None = None death_place: str | None = None
biography: str | None = None biography: str | None = None
portrait_path: str | None = None
family_search_id: str | None = Field(default=None, unique=True) family_search_id: str | None = Field(default=None, unique=True)
metadata_: dict[str, JsonValue] | None = Field( metadata_: dict[str, JsonValue] | None = Field(
default=None, default=None,
sa_column=Column("metadata", JSONBCompat(), nullable=True), sa_column=Column("metadata", JSONBCompat(), nullable=True),
) )
created_at: datetime = Field(default_factory=lambda: datetime.now(UTC)) created_at: datetime = Field(default_factory=_utc_now_naive)
updated_at: datetime = Field( updated_at: datetime = Field(
default_factory=lambda: datetime.now(UTC), default_factory=_utc_now_naive,
sa_column_kwargs={"onupdate": lambda: datetime.now(UTC)}, sa_column_kwargs={"onupdate": _utc_now_naive},
) )
document_people: list["DocumentPerson"] = Relationship( document_people: list["DocumentPerson"] = Relationship(
back_populates="person", sa_relationship_kwargs={"lazy": "raise"} back_populates="person", sa_relationship_kwargs={"lazy": "raise"}
) )
person_tags: list["PersonTag"] = Relationship(
back_populates="person",
sa_relationship_kwargs={"lazy": "raise"},
)
photos: list["Photo"] = Relationship(
back_populates="person",
sa_relationship_kwargs={"lazy": "raise"},
)
@property
def full_name(self) -> str:
"""Presentation-friendly combined name."""
return f"{self.given_names} {self.last_name}".strip()
class GenealogyPerson(SQLModel, table=True):
"""An individual imported from a GEDCOM export."""
__tablename__ = "genealogy_person"
id: UUID = Field(default_factory=uuid4, primary_key=True)
fs_id: str = Field(index=True, unique=True)
full_name: str
birth_date: date | None = None
birth_date_raw: str | None = None
birth_place: str | None = None
death_date: date | None = None
death_date_raw: str | None = None
death_place: str | None = None
created_at: datetime = Field(default_factory=_utc_now_naive)
updated_at: datetime = Field(
default_factory=_utc_now_naive,
sa_column_kwargs={"onupdate": _utc_now_naive},
)
husband_families: list["GenealogyFamily"] = Relationship(
back_populates="husband",
sa_relationship_kwargs={"lazy": "raise", "foreign_keys": "[GenealogyFamily.husband_id]"},
)
wife_families: list["GenealogyFamily"] = Relationship(
back_populates="wife",
sa_relationship_kwargs={"lazy": "raise", "foreign_keys": "[GenealogyFamily.wife_id]"},
)
child_family_memberships: list["GenealogyFamilyChild"] = Relationship(
back_populates="child",
sa_relationship_kwargs={"lazy": "raise"},
)
citations: list["GenealogyCitation"] = Relationship(
back_populates="genealogy_person",
sa_relationship_kwargs={"lazy": "raise"},
)
class GenealogyFamily(SQLModel, table=True):
"""A family linking two GenealogyPerson records."""
__tablename__ = "genealogy_family"
id: UUID = Field(default_factory=uuid4, primary_key=True)
fs_family_id: str = Field(index=True, unique=True)
husband_id: UUID | None = Field(default=None, foreign_key="genealogy_person.id", index=True)
wife_id: UUID | None = Field(default=None, foreign_key="genealogy_person.id", index=True)
marriage_date: date | None = None
marriage_date_raw: str | None = None
marriage_place: str | None = None
created_at: datetime = Field(default_factory=_utc_now_naive)
updated_at: datetime = Field(
default_factory=_utc_now_naive,
sa_column_kwargs={"onupdate": _utc_now_naive},
)
husband: Optional["GenealogyPerson"] = Relationship(
back_populates="husband_families",
sa_relationship_kwargs={"lazy": "raise", "foreign_keys": "[GenealogyFamily.husband_id]"},
)
wife: Optional["GenealogyPerson"] = Relationship(
back_populates="wife_families",
sa_relationship_kwargs={"lazy": "raise", "foreign_keys": "[GenealogyFamily.wife_id]"},
)
children: list["GenealogyFamilyChild"] = Relationship(
back_populates="family",
sa_relationship_kwargs={"lazy": "raise"},
)
citations: list["GenealogyCitation"] = Relationship(
back_populates="genealogy_family",
sa_relationship_kwargs={"lazy": "raise"},
)
class GenealogyFamilyChild(SQLModel, table=True):
"""Junction table for child membership within a genealogy family."""
__tablename__ = "genealogy_family_child"
id: UUID = Field(default_factory=uuid4, primary_key=True)
family_id: UUID = Field(foreign_key="genealogy_family.id", index=True)
child_id: UUID = Field(foreign_key="genealogy_person.id", index=True)
relationship_type: str | None = None
created_at: datetime = Field(default_factory=_utc_now_naive)
__table_args__ = (UniqueConstraint("family_id", "child_id", name="uq_genealogy_family_child"),)
family: Optional["GenealogyFamily"] = Relationship(
back_populates="children",
sa_relationship_kwargs={"lazy": "raise"},
)
child: Optional["GenealogyPerson"] = Relationship(
back_populates="child_family_memberships",
sa_relationship_kwargs={"lazy": "raise"},
)
class GenealogyCitation(SQLModel, table=True):
"""A source citation attached to a genealogical fact."""
__tablename__ = "genealogy_citation"
id: UUID = Field(default_factory=uuid4, primary_key=True)
genealogy_person_id: UUID | None = Field(default=None, foreign_key="genealogy_person.id", index=True)
genealogy_family_id: UUID | None = Field(default=None, foreign_key="genealogy_family.id", index=True)
fact_type: GenealogyCitationFactType = Field(
sa_column=Column(
SAEnum(
GenealogyCitationFactType,
values_callable=lambda enum_cls: [item.value for item in enum_cls],
native_enum=False,
),
nullable=False,
)
)
raw_citation_text: str
source_kind: GenealogyCitationSourceKind = Field(
sa_column=Column(
SAEnum(
GenealogyCitationSourceKind,
values_callable=lambda enum_cls: [item.value for item in enum_cls],
native_enum=False,
),
nullable=False,
)
)
document_id: UUID | None = Field(default=None, foreign_key="document.id", index=True)
created_at: datetime = Field(default_factory=_utc_now_naive)
genealogy_person: Optional["GenealogyPerson"] = Relationship(
back_populates="citations",
sa_relationship_kwargs={"lazy": "raise"},
)
genealogy_family: Optional["GenealogyFamily"] = Relationship(
back_populates="citations",
sa_relationship_kwargs={"lazy": "raise"},
)
document: Optional["Document"] = Relationship(sa_relationship_kwargs={"lazy": "raise"})
class Photo(SQLModel, table=True):
"""A reusable image record for Person and homepage galleries."""
__tablename__ = "photo"
id: UUID = Field(default_factory=uuid4, primary_key=True)
person_id: UUID | None = Field(default=None, foreign_key="person.id", index=True)
path: str
description: str | None = None
is_primary: bool = False
created_at: datetime = Field(default_factory=_utc_now_naive)
updated_at: datetime = Field(
default_factory=_utc_now_naive,
sa_column_kwargs={"onupdate": _utc_now_naive},
)
person: Optional["Person"] = Relationship(
back_populates="photos",
sa_relationship_kwargs={"lazy": "raise"},
)
class DocumentPerson(SQLModel, table=True): class DocumentPerson(SQLModel, table=True):
@@ -187,10 +419,10 @@ class DocumentPerson(SQLModel, table=True):
document_id: UUID = Field(foreign_key="document.id", index=True) document_id: UUID = Field(foreign_key="document.id", index=True)
person_id: UUID = Field(foreign_key="person.id", index=True) person_id: UUID = Field(foreign_key="person.id", index=True)
role_id: UUID = Field(foreign_key="person_role.id", index=True) role_id: UUID = Field(foreign_key="person_role.id", index=True)
created_at: datetime = Field(default_factory=lambda: datetime.now(UTC)) created_at: datetime = Field(default_factory=_utc_now_naive)
updated_at: datetime = Field( updated_at: datetime = Field(
default_factory=lambda: datetime.now(UTC), default_factory=_utc_now_naive,
sa_column_kwargs={"onupdate": lambda: datetime.now(UTC)}, sa_column_kwargs={"onupdate": _utc_now_naive},
) )
__table_args__ = (UniqueConstraint("document_id", "person_id", name="uq_document_person"),) __table_args__ = (UniqueConstraint("document_id", "person_id", name="uq_document_person"),)
@@ -206,6 +438,58 @@ class DocumentPerson(SQLModel, table=True):
) )
class DocumentTag(SQLModel, table=True):
"""Associates Documents with Tags."""
__tablename__ = "document_tag"
id: UUID = Field(default_factory=uuid4, primary_key=True)
document_id: UUID = Field(foreign_key="document.id", index=True)
tag_id: UUID = Field(foreign_key="tag.id", index=True)
created_at: datetime = Field(default_factory=_utc_now_naive)
updated_at: datetime = Field(
default_factory=_utc_now_naive,
sa_column_kwargs={"onupdate": _utc_now_naive},
)
__table_args__ = (UniqueConstraint("document_id", "tag_id", name="uq_document_tag"),)
document: Optional["Document"] = Relationship(
back_populates="document_tags",
sa_relationship_kwargs={"lazy": "raise"},
)
tag_ref: Optional["Tag"] = Relationship(
back_populates="document_tags",
sa_relationship_kwargs={"lazy": "raise"},
)
class PersonTag(SQLModel, table=True):
"""Associates People with Tags."""
__tablename__ = "person_tag"
id: UUID = Field(default_factory=uuid4, primary_key=True)
person_id: UUID = Field(foreign_key="person.id", index=True)
tag_id: UUID = Field(foreign_key="tag.id", index=True)
created_at: datetime = Field(default_factory=_utc_now_naive)
updated_at: datetime = Field(
default_factory=_utc_now_naive,
sa_column_kwargs={"onupdate": _utc_now_naive},
)
__table_args__ = (UniqueConstraint("person_id", "tag_id", name="uq_person_tag"),)
person: Optional["Person"] = Relationship(
back_populates="person_tags",
sa_relationship_kwargs={"lazy": "raise"},
)
tag_ref: Optional["Tag"] = Relationship(
back_populates="person_tags",
sa_relationship_kwargs={"lazy": "raise"},
)
class Job(SQLModel, table=True): class Job(SQLModel, table=True):
"""A transcription job tied to a single document.""" """A transcription job tied to a single document."""
@@ -237,10 +521,10 @@ class Job(SQLModel, table=True):
default=JobPurpose.TRANSCRIPTION.value, default=JobPurpose.TRANSCRIPTION.value,
), ),
) )
date_created: datetime = Field(default_factory=lambda: datetime.now(UTC)) date_created: datetime = Field(default_factory=_utc_now_naive)
date_updated: datetime = Field( date_updated: datetime = Field(
default_factory=lambda: datetime.now(UTC), default_factory=_utc_now_naive,
sa_column_kwargs={"onupdate": lambda: datetime.now(UTC)}, sa_column_kwargs={"onupdate": _utc_now_naive},
) )
provider: str | None = None provider: str | None = None
model: str | None = None model: str | None = None
@@ -271,6 +555,47 @@ class Job(SQLModel, table=True):
return "unknown" return "unknown"
class MaintenanceRun(SQLModel, table=True):
"""A queued/processed maintenance task execution record."""
__tablename__ = "maintenance_run"
id: UUID = Field(default_factory=uuid4, primary_key=True)
job_type: MaintenanceJobType = Field(
sa_column=Column(
SAEnum(
MaintenanceJobType,
values_callable=lambda enum_cls: [item.value for item in enum_cls],
native_enum=False,
),
nullable=False,
)
)
status: MaintenanceRunStatus = Field(
default=MaintenanceRunStatus.QUEUED,
sa_column=Column(
SAEnum(
MaintenanceRunStatus,
values_callable=lambda enum_cls: [item.value for item in enum_cls],
native_enum=False,
),
nullable=False,
default=MaintenanceRunStatus.QUEUED.value,
),
)
started_at: datetime | None = None
finished_at: datetime | None = None
triggered_by: str | None = None
summary: str | None = None
log_path: str | None = None
error_detail: str | None = None
created_at: datetime = Field(default_factory=_utc_now_naive)
updated_at: datetime = Field(
default_factory=_utc_now_naive,
sa_column_kwargs={"onupdate": _utc_now_naive},
)
class Source(SQLModel, table=True): class Source(SQLModel, table=True):
"""A document source image or PDF page.""" """A document source image or PDF page."""
@@ -299,7 +624,7 @@ class Source(SQLModel, table=True):
), ),
) )
revised_text: str | None = None revised_text: str | None = None
date_uploaded: datetime = Field(default_factory=lambda: datetime.now(UTC)) date_uploaded: datetime = Field(default_factory=_utc_now_naive)
date_revised: datetime | None = None date_revised: datetime | None = None
document: Optional["Document"] = Relationship( document: Optional["Document"] = Relationship(
@@ -321,9 +646,7 @@ class Source(SQLModel, table=True):
""" """
job_sources = _loaded_attribute(self, "job_sources") or () job_sources = _loaded_attribute(self, "job_sources") or ()
dated = [ dated = [
(job, job_source) (job, job_source) for job_source in job_sources if (job := _loaded_attribute(job_source, "job")) is not None
for job_source in job_sources
if (job := _loaded_attribute(job_source, "job")) is not None
] ]
if dated: if dated:
return max(dated, key=lambda pair: pair[0].date_created)[1] return max(dated, key=lambda pair: pair[0].date_created)[1]
@@ -346,10 +669,10 @@ class Source(SQLModel, table=True):
if latest is None: if latest is None:
return None return None
attempts = _loaded_attribute(latest, "execution_attempts") or () attempts = _loaded_attribute(latest, "execution_attempts") or ()
for attempt in sorted(attempts, key=lambda item: item.attempt_number, reverse=True): if not attempts:
if attempt.error_detail: return None
return attempt.error_detail latest_attempt = max(attempts, key=lambda item: item.attempt_number)
return None return latest_attempt.error_detail
@property @property
def document_name(self) -> str | None: def document_name(self) -> str | None:
@@ -361,6 +684,7 @@ class JobSource(SQLModel, table=True):
"""A single AI execution record for one source page.""" """A single AI execution record for one source page."""
__tablename__ = "job_source" __tablename__ = "job_source"
__table_args__ = (UniqueConstraint("job_id", "source_id", name="uq_job_source_job_source"),)
id: UUID = Field(default_factory=uuid4, primary_key=True) id: UUID = Field(default_factory=uuid4, primary_key=True)
job_id: UUID = Field(foreign_key="job.id", index=True) job_id: UUID = Field(foreign_key="job.id", index=True)
@@ -438,7 +762,7 @@ class ExecutionAttempt(SQLModel, table=True):
started_at: datetime started_at: datetime
finished_at: datetime finished_at: datetime
duration_ms: int = Field(ge=0) duration_ms: int = Field(ge=0)
created_at: datetime = Field(default_factory=lambda: datetime.now(UTC)) created_at: datetime = Field(default_factory=_utc_now_naive)
job_source: Optional["JobSource"] = Relationship( job_source: Optional["JobSource"] = Relationship(
back_populates="execution_attempts", sa_relationship_kwargs={"lazy": "raise"} back_populates="execution_attempts", sa_relationship_kwargs={"lazy": "raise"}
+188
View File
@@ -1,7 +1,10 @@
from __future__ import annotations from __future__ import annotations
import logging import logging
from pathlib import Path
from sqlalchemy import inspect as sqlalchemy_inspect
from sqlalchemy import text
from sqlalchemy.ext.asyncio import AsyncEngine from sqlalchemy.ext.asyncio import AsyncEngine
from sqlalchemy.ext.asyncio import async_sessionmaker from sqlalchemy.ext.asyncio import async_sessionmaker
from sqlmodel import SQLModel from sqlmodel import SQLModel
@@ -29,6 +32,191 @@ async def create_all(*, engine: AsyncEngine | None = None) -> None:
logger.debug("Database schema bootstrap complete for database_url=%s", active_engine.url) logger.debug("Database schema bootstrap complete for database_url=%s", active_engine.url)
async def reconcile_legacy_job_source_columns(*, engine: AsyncEngine | None = None) -> int:
"""Remove stale V4.6 ``job_source`` evidence columns from existing databases.
Runtime models define ``job_source`` as a queue/projection table only. If an
older database still carries the retired evidence columns, writes can fail
on stale constraints (for example ``executed_at NOT NULL``).
"""
active_engine = engine or resolve_engine()
if not hasattr(active_engine, "begin"):
return 0
def _reconcile(sync_connection) -> int:
inspector = sqlalchemy_inspect(sync_connection)
table_names = set(inspector.get_table_names())
if "job_source" not in table_names:
return 0
present_columns = {column["name"] for column in inspector.get_columns("job_source")}
dropped = 0
for column_name in (
"raw_transcription",
"ai_metadata",
"raw_api_response",
"error_detail",
"executed_at",
):
if column_name not in present_columns:
continue
sync_connection.execute(text(f'alter table "job_source" drop column "{column_name}"'))
dropped += 1
return dropped
async with active_engine.begin() as connection:
dropped_columns = await connection.run_sync(_reconcile)
if dropped_columns:
logger.warning("Dropped %s legacy job_source column(s) during startup reconciliation", dropped_columns)
return dropped_columns
async def reconcile_canonical_media_paths(*, engine: AsyncEngine | None = None) -> int:
"""Normalize stored media paths to upload-root-relative POSIX form."""
active_engine = engine or resolve_engine()
if not hasattr(active_engine, "begin"):
return 0
def _reconcile(sync_connection) -> int:
rows_changed = 0
inspector = sqlalchemy_inspect(sync_connection)
table_names = set(inspector.get_table_names())
if "source" in table_names:
rows = (
sync_connection.execute(text('select id, file_path from "source" where file_path is not null'))
.mappings()
.all()
)
for row in rows:
original = str(row["file_path"])
normalized = _canonical_relative_path(original, preferred_prefix="documents/")
if normalized is None or normalized == original:
continue
sync_connection.execute(
text('update "source" set file_path = :file_path where id = :id'),
{"id": row["id"], "file_path": normalized},
)
rows_changed += 1
if "photo" in table_names:
rows = sync_connection.execute(text('select id, path from "photo" where path is not null')).mappings().all()
for row in rows:
original = str(row["path"])
normalized = _canonical_relative_path(original, preferred_prefix="photos/")
if normalized is None or normalized == original:
continue
sync_connection.execute(
text('update "photo" set path = :path where id = :id'),
{"id": row["id"], "path": normalized},
)
rows_changed += 1
return rows_changed
async with active_engine.begin() as connection:
rows_changed = await connection.run_sync(_reconcile)
if rows_changed:
logger.warning("Normalized %s media-path row(s) to canonical relative format", rows_changed)
return rows_changed
async def reconcile_person_name_columns(*, engine: AsyncEngine | None = None) -> int:
"""Backfill V5.1 Person name columns on existing databases."""
active_engine = engine or resolve_engine()
if not hasattr(active_engine, "begin"):
return 0
def _reconcile(sync_connection) -> int:
rows_changed = 0
inspector = sqlalchemy_inspect(sync_connection)
table_names = set(inspector.get_table_names())
if "person" not in table_names:
return 0
present_columns = {column["name"] for column in inspector.get_columns("person")}
if "last_name" not in present_columns:
sync_connection.execute(text('alter table "person" add column "last_name" varchar'))
if "given_names" not in present_columns:
sync_connection.execute(text('alter table "person" add column "given_names" varchar'))
query = (
text('select id, full_name, given_names, last_name from "person"')
if "full_name" in present_columns
else text('select id, null as full_name, given_names, last_name from "person"')
)
rows = sync_connection.execute(query).mappings().all()
for row in rows:
given_names = (str(row.get("given_names") or "")).strip()
last_name = (str(row.get("last_name") or "")).strip()
if given_names and last_name:
continue
tokens = [token for token in str(row.get("full_name") or "").split() if token]
if len(tokens) >= 2:
given_names, last_name = (" ".join(tokens[:-1]), tokens[-1])
elif len(tokens) == 1:
given_names = tokens[0]
last_name = tokens[0]
else:
given_names = "Unknown"
last_name = "Unknown"
sync_connection.execute(
text('update "person" set given_names = :given_names, last_name = :last_name where id = :id'),
{
"id": row["id"],
"given_names": given_names,
"last_name": last_name,
},
)
rows_changed += 1
return rows_changed
async with active_engine.begin() as connection:
rows_changed = await connection.run_sync(_reconcile)
if rows_changed:
logger.warning("Backfilled V5.1 name columns for %s person row(s)", rows_changed)
return rows_changed
def _canonical_relative_path(value: str, *, preferred_prefix: str) -> str | None:
normalized = value.strip().replace("\\", "/")
if not normalized:
return None
lowered = normalized.casefold()
if lowered.startswith(("http://", "https://", "data:")):
return None
if lowered.startswith("/uploads/"):
normalized = normalized[len("/uploads/") :]
lowered = normalized.casefold()
elif lowered.startswith("uploads/"):
normalized = normalized[len("uploads/") :]
lowered = normalized.casefold()
elif lowered.startswith("data/"):
normalized = normalized[len("data/") :]
lowered = normalized.casefold()
for prefix in ("documents/", "photos/", "persons/", "portraits/"):
marker = f"/{prefix}"
index = lowered.find(marker)
if index >= 0:
normalized = normalized[index + 1 :]
lowered = normalized.casefold()
break
if lowered.startswith(prefix):
break
if preferred_prefix == "persons/" and lowered.startswith("portraits/"):
normalized = "persons/" + normalized[len("portraits/") :]
lowered = normalized.casefold()
if not lowered.startswith(preferred_prefix):
return None
# Collapse any accidental "." segments while preserving relative semantics.
collapsed = Path(normalized).as_posix()
if collapsed.startswith("../") or collapsed == "..":
return None
return collapsed
async def seed_registry_defaults(*, engine: AsyncEngine | None = None) -> None: async def seed_registry_defaults(*, engine: AsyncEngine | None = None) -> None:
"""Seed default registry rows for role and document type taxonomies.""" """Seed default registry rows for role and document type taxonomies."""
active_engine = engine or resolve_engine() active_engine = engine or resolve_engine()
+1 -2
View File
@@ -50,8 +50,7 @@ def initialize_database_runtime(*, settings: Settings | None = None) -> Database
runtime_url = runtime.engine.url.render_as_string(hide_password=False) runtime_url = runtime.engine.url.render_as_string(hide_password=False)
if runtime_url != database_url: if runtime_url != database_url:
raise RuntimeError( raise RuntimeError(
"Database runtime is already initialized for a different database: " f"Database runtime is already initialized for a different database: {runtime_url!r} != {database_url!r}"
f"{runtime_url!r} != {database_url!r}"
) )
return runtime return runtime
+65 -6
View File
@@ -2,12 +2,15 @@
from __future__ import annotations from __future__ import annotations
import logging
from dataclasses import dataclass from dataclasses import dataclass
from datetime import UTC from datetime import UTC
from datetime import datetime from datetime import datetime
from enum import StrEnum from enum import StrEnum
from uuid import uuid4 from uuid import uuid4
logger = logging.getLogger(__name__)
class ErrorCategory(StrEnum): class ErrorCategory(StrEnum):
"""Stable error categories defined by docs/error_handling.md.""" """Stable error categories defined by docs/error_handling.md."""
@@ -17,6 +20,7 @@ class ErrorCategory(StrEnum):
NOT_FOUND = "not_found_error" NOT_FOUND = "not_found_error"
CONFLICT = "conflict_error" CONFLICT = "conflict_error"
EXTERNAL_PROVIDER = "external_provider_error" EXTERNAL_PROVIDER = "external_provider_error"
EXTERNAL_TIMEOUT = "external_timeout_error"
PROCESSING = "processing_error" PROCESSING = "processing_error"
INFRA_TRANSIENT = "infrastructure_transient_error" INFRA_TRANSIENT = "infrastructure_transient_error"
INFRA_PERSISTENT = "infrastructure_persistent_error" INFRA_PERSISTENT = "infrastructure_persistent_error"
@@ -28,6 +32,11 @@ def new_error_id() -> str:
return uuid4().hex[:8] return uuid4().hex[:8]
def exception_detail(exc: BaseException) -> str:
"""Return internal-only root-cause text for persisted diagnostics."""
return f"{type(exc).__name__}: {exc}"
class AppError(RuntimeError): class AppError(RuntimeError):
"""Base application error carrying user-safe handling metadata.""" """Base application error carrying user-safe handling metadata."""
@@ -39,6 +48,7 @@ class AppError(RuntimeError):
suggestion: str = "Retry once. If it persists, review logs and report the error reference id.", suggestion: str = "Retry once. If it persists, review logs and report the error reference id.",
retriable: bool = False, retriable: bool = False,
error_id: str | None = None, error_id: str | None = None,
detail: str | None = None,
) -> None: ) -> None:
super().__init__(message) super().__init__(message)
self.message = message self.message = message
@@ -46,6 +56,10 @@ class AppError(RuntimeError):
self.suggestion = suggestion self.suggestion = suggestion
self.retriable = retriable self.retriable = retriable
self.error_id = error_id or new_error_id() self.error_id = error_id or new_error_id()
# Internal-only diagnostic text. Persisted to evidence and logs, never rendered
# to users or serialized into API envelopes, because it may embed local
# filesystem paths and other infrastructure detail.
self.detail = detail
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -59,11 +73,28 @@ class ErrorEnvelope:
timestamp: str timestamp: str
def canonical_error_category(error: AppError) -> str:
"""Map internal categories to canonical API/UI envelope categories."""
mapping: dict[ErrorCategory, str] = {
ErrorCategory.VALIDATION: "validation",
ErrorCategory.USER_INPUT: "validation",
ErrorCategory.NOT_FOUND: "not_found",
ErrorCategory.CONFLICT: "conflict",
ErrorCategory.EXTERNAL_PROVIDER: "external",
ErrorCategory.EXTERNAL_TIMEOUT: "timeout",
ErrorCategory.INFRA_TRANSIENT: "timeout",
ErrorCategory.PROCESSING: "internal",
ErrorCategory.INFRA_PERSISTENT: "internal",
ErrorCategory.INTERNAL_UNEXPECTED: "internal",
}
return mapping.get(error.category, "internal")
def build_error_envelope(error: AppError) -> ErrorEnvelope: def build_error_envelope(error: AppError) -> ErrorEnvelope:
"""Build an API-safe response envelope from an AppError.""" """Build an API-safe response envelope from an AppError."""
return ErrorEnvelope( return ErrorEnvelope(
error_id=error.error_id, error_id=error.error_id,
category=error.category.value, category=canonical_error_category(error),
message=error.message, message=error.message,
suggestion=error.suggestion, suggestion=error.suggestion,
timestamp=datetime.now(UTC).isoformat(), timestamp=datetime.now(UTC).isoformat(),
@@ -71,15 +102,43 @@ def build_error_envelope(error: AppError) -> ErrorEnvelope:
def classify_unexpected_error(exc: Exception, *, operation: str) -> AppError: def classify_unexpected_error(exc: Exception, *, operation: str) -> AppError:
"""Normalize unknown exceptions into internal_unexpected_error.""" """Normalize unknown exceptions into internal_unexpected_error.
return AppError(
f"Unexpected error during {operation}: {exc}", The exception text is deliberately excluded from ``message``. ``AppError.message``
is rendered directly to users by the UI error presenter and is serialized into API
responses by :func:`build_error_envelope`, and unexpected exceptions routinely embed
local filesystem paths (SQLAlchemy ``OperationalError`` carries the database path,
``OSError`` carries the storage root). Leaking those is forbidden by
``.github/instructions/error-handling.instructions.md``.
The detail is preserved on ``AppError.detail`` and logged against ``error_id``. That
keeps the root cause in evidence records and operator logs, which are internal, while
keeping it out of user-facing and API-facing text.
"""
error = AppError(
f"Unexpected error during {operation}.",
category=ErrorCategory.INTERNAL_UNEXPECTED, category=ErrorCategory.INTERNAL_UNEXPECTED,
suggestion="Retry once. If it persists, review logs and report the error reference id.", suggestion="Retry once. If it persists, review logs and report the error reference id.",
retriable=False, retriable=False,
detail=exception_detail(exc),
) )
logger.error(
"Unexpected error operation=%s error_id=%s",
operation,
error.error_id,
exc_info=exc,
)
return error
def format_error_detail(error: AppError) -> str: def format_error_detail(error: AppError) -> str:
"""Return a compact persisted failure string for transcript.error_detail.""" """Return a compact persisted failure string for transcript.error_detail.
return f"[{error.category.value}] {error.message} | suggestion={error.suggestion} | error_id={error.error_id}"
This is internal provenance, not user-facing output, so it carries
``AppError.detail`` (the root cause) in addition to the user-safe message.
"""
parts = [f"[{error.category.value}] {error.message}"]
if error.detail:
parts.append(f"detail={error.detail}")
parts.extend((f"suggestion={error.suggestion}", f"error_id={error.error_id}"))
return " | ".join(parts)
+2
View File
@@ -4,6 +4,7 @@ from transcription.config import Provider
from transcription.config import Settings from transcription.config import Settings
from transcription.config import get_settings from transcription.config import get_settings
from transcription.providers.base import ProviderAuthError from transcription.providers.base import ProviderAuthError
from transcription.providers.base import ProviderCallEvidence
from transcription.providers.base import ProviderError from transcription.providers.base import ProviderError
from transcription.providers.base import ProviderResponseError from transcription.providers.base import ProviderResponseError
from transcription.providers.base import TranscriptionMetadata from transcription.providers.base import TranscriptionMetadata
@@ -27,6 +28,7 @@ def get_transcription_provider(*, settings: Settings | None = None) -> Transcrip
__all__ = [ __all__ = [
"OpenRouterTranscriptionProvider", "OpenRouterTranscriptionProvider",
"ProviderAuthError", "ProviderAuthError",
"ProviderCallEvidence",
"ProviderError", "ProviderError",
"ProviderResponseError", "ProviderResponseError",
"RequestManifest", "RequestManifest",
+11 -11
View File
@@ -1,5 +1,6 @@
"""Provider interfaces and validated shared contracts for transcription adapters.""" """Provider interfaces and validated shared contracts for transcription adapters."""
from dataclasses import dataclass
from typing import Protocol from typing import Protocol
from pydantic import BaseModel from pydantic import BaseModel
@@ -37,6 +38,14 @@ class ProviderResponseError(ProviderError):
"""Raised when provider responses are malformed or unusable.""" """Raised when provider responses are malformed or unusable."""
@dataclass(slots=True)
class ProviderCallEvidence:
"""Caller-owned evidence sink for one provider invocation."""
request_manifest: RequestManifest | None = None
transport_evidence: TransportEvidence | None = None
class ProviderUsage(BaseModel): class ProviderUsage(BaseModel):
"""Normalized provider token accounting.""" """Normalized provider token accounting."""
@@ -107,16 +116,6 @@ class TranscriptionProvider(Protocol):
"""Return the resolved model slug this adapter will call.""" """Return the resolved model slug this adapter will call."""
... ...
@property
def current_request_manifest(self) -> RequestManifest | None:
"""Return the manifest for the most recent call, for failure evidence."""
...
@property
def current_transport_evidence(self) -> TransportEvidence | None:
"""Return transport-level evidence for the most recent call."""
...
async def transcribe( async def transcribe(
self, self,
*, *,
@@ -127,8 +126,9 @@ class TranscriptionProvider(Protocol):
top_p: float | None = None, top_p: float | None = None,
source_reference: SourceEvidenceReference | None = None, source_reference: SourceEvidenceReference | None = None,
requested_model: str | None = None, requested_model: str | None = None,
evidence_capture: ProviderCallEvidence | None = None,
) -> TranscriptionResult: ) -> TranscriptionResult:
"""Transcribe the provided image according to the prompt text.""" """Transcribe one source and write failure evidence into the provided capture sink."""
... ...
async def aclose(self) -> None: async def aclose(self) -> None:
+10 -3
View File
@@ -4,7 +4,6 @@ from __future__ import annotations
import hashlib import hashlib
import json import json
import os
import platform import platform
from importlib.metadata import PackageNotFoundError from importlib.metadata import PackageNotFoundError
from importlib.metadata import version from importlib.metadata import version
@@ -17,6 +16,8 @@ from pydantic import ConfigDict
from pydantic import Field from pydantic import Field
from pydantic import JsonValue from pydantic import JsonValue
from transcription.config import Settings
REQUEST_MANIFEST_SCHEMA = "transcription.request-manifest" REQUEST_MANIFEST_SCHEMA = "transcription.request-manifest"
REQUEST_MANIFEST_VERSION = "1" REQUEST_MANIFEST_VERSION = "1"
SOFTWARE_CONTEXT_SCHEMA = "transcription.software-context" SOFTWARE_CONTEXT_SCHEMA = "transcription.software-context"
@@ -141,11 +142,17 @@ def package_version(package: str) -> str:
return "unknown" return "unknown"
def build_software_context(*, adapter_name: str, adapter_version: str, client_library: str) -> SoftwareContext: def build_software_context(
*,
adapter_name: str,
adapter_version: str,
client_library: str,
settings: Settings,
) -> SoftwareContext:
"""Build the runtime software identity for an execution.""" """Build the runtime software identity for an execution."""
return SoftwareContext( return SoftwareContext(
application_version=package_version("transcription"), application_version=package_version("transcription"),
application_commit=os.environ.get("TRANSCRIPTION_COMMIT") or None, application_commit=settings.transcription_commit,
adapter_name=adapter_name, adapter_name=adapter_name,
adapter_version=adapter_version, adapter_version=adapter_version,
client_library=client_library, client_library=client_library,
+109 -96
View File
@@ -3,11 +3,13 @@
from __future__ import annotations from __future__ import annotations
import base64 import base64
import contextvars
import hashlib import hashlib
import json import json
import logging import logging
from collections.abc import AsyncIterator from collections.abc import AsyncIterator
from collections.abc import Callable from collections.abc import Callable
from dataclasses import dataclass
from typing import Annotated from typing import Annotated
from typing import Any from typing import Any
from typing import Literal from typing import Literal
@@ -25,6 +27,7 @@ from pydantic import ValidationError
from transcription.config import Settings from transcription.config import Settings
from transcription.config import get_settings from transcription.config import get_settings
from transcription.providers.base import ProviderAuthError from transcription.providers.base import ProviderAuthError
from transcription.providers.base import ProviderCallEvidence
from transcription.providers.base import ProviderError from transcription.providers.base import ProviderError
from transcription.providers.base import ProviderResponseError from transcription.providers.base import ProviderResponseError
from transcription.providers.base import ProviderUsage from transcription.providers.base import ProviderUsage
@@ -65,19 +68,24 @@ class _CapturingAsyncClient:
def __init__(self, client: httpx.AsyncClient): def __init__(self, client: httpx.AsyncClient):
self._client = client self._client = client
self.last_response: httpx.Response | None = None self._active_capture: contextvars.ContextVar[_TransportCapture | None] = contextvars.ContextVar(
self.last_body: bytes | None = None "openrouter_transport_capture",
default=None,
)
async def send(self, request: httpx.Request, **kwargs: Any) -> httpx.Response: async def send(self, request: httpx.Request, **kwargs: Any) -> httpx.Response:
capture = self._active_capture.get()
response = await self._client.send(request, **kwargs) response = await self._client.send(request, **kwargs)
self.last_response = response if capture is None:
return response
capture.response = response
try: try:
self.last_body = response.content capture.body = response.content
except httpx.ResponseNotRead: except httpx.ResponseNotRead:
stream = response.stream stream = response.stream
if not isinstance(stream, httpx.AsyncByteStream): if not isinstance(stream, httpx.AsyncByteStream):
raise raise
response.stream = _CapturingAsyncByteStream(stream, self._capture_body) response.stream = _CapturingAsyncByteStream(stream, lambda body: self._capture_body(capture, body))
return response return response
def build_request(self, *args: Any, **kwargs: Any) -> httpx.Request: def build_request(self, *args: Any, **kwargs: Any) -> httpx.Request:
@@ -86,12 +94,21 @@ class _CapturingAsyncClient:
async def aclose(self) -> None: async def aclose(self) -> None:
await self._client.aclose() await self._client.aclose()
def reset(self) -> None: def begin_capture(self, capture: _TransportCapture) -> contextvars.Token[_TransportCapture | None]:
self.last_response = None return self._active_capture.set(capture)
self.last_body = None
def _capture_body(self, body: bytes) -> None: def end_capture(self, token: contextvars.Token[_TransportCapture | None]) -> None:
self.last_body = body self._active_capture.reset(token)
@staticmethod
def _capture_body(capture: _TransportCapture, body: bytes) -> None:
capture.body = body
@dataclass(slots=True)
class _TransportCapture:
response: httpx.Response | None = None
body: bytes | None = None
class _ProviderModel(BaseModel): class _ProviderModel(BaseModel):
@@ -195,8 +212,6 @@ class OpenRouterTranscriptionProvider:
self._settings = settings or get_settings() self._settings = settings or get_settings()
self._model = self._settings.provider_model or DEFAULT_OPENROUTER_MODEL self._model = self._settings.provider_model or DEFAULT_OPENROUTER_MODEL
self._capturing_client: _CapturingAsyncClient | None = None self._capturing_client: _CapturingAsyncClient | None = None
self._current_request_manifest: RequestManifest | None = None
self._current_transport_evidence: TransportEvidence | None = None
if client is None: if client is None:
# httpx defaults every phase to 5s, which silently caps provider calls far # httpx defaults every phase to 5s, which silently caps provider calls far
# below worker_provider_timeout_seconds. Track the configured budget instead. # below worker_provider_timeout_seconds. Track the configured budget instead.
@@ -218,18 +233,6 @@ class OpenRouterTranscriptionProvider:
"""Return the resolved OpenRouter model slug.""" """Return the resolved OpenRouter model slug."""
return self._model return self._model
@property
def current_request_manifest(self) -> RequestManifest | None:
return self._current_request_manifest
@property
def current_transport_evidence(self) -> TransportEvidence | None:
if self._current_transport_evidence is not None:
return self._current_transport_evidence
if self._current_request_manifest is None:
return None
return self._captured_transport_evidence()
async def aclose(self) -> None: async def aclose(self) -> None:
if self._capturing_client is not None: if self._capturing_client is not None:
await self._capturing_client.aclose() await self._capturing_client.aclose()
@@ -244,6 +247,7 @@ class OpenRouterTranscriptionProvider:
top_p: float | None = None, top_p: float | None = None,
source_reference: SourceEvidenceReference | None = None, source_reference: SourceEvidenceReference | None = None,
requested_model: str | None = None, requested_model: str | None = None,
evidence_capture: ProviderCallEvidence | None = None,
) -> TranscriptionResult: ) -> TranscriptionResult:
"""Send prompt + image to OpenRouter and return normalized text output.""" """Send prompt + image to OpenRouter and return normalized text output."""
request = self._build_request( request = self._build_request(
@@ -261,79 +265,87 @@ class OpenRouterTranscriptionProvider:
temperature=temperature, temperature=temperature,
top_p=top_p, top_p=top_p,
) )
self._current_request_manifest = manifest if evidence_capture is not None:
self._current_transport_evidence = None evidence_capture.request_manifest = manifest
if self._capturing_client is not None: evidence_capture.transport_evidence = None
self._capturing_client.reset() transport_capture = _TransportCapture()
token = self._capturing_client.begin_capture(transport_capture) if self._capturing_client is not None else None
try: try:
response = await self._client.chat.send_async( try:
**request.model_dump(mode="json", exclude_none=True), response = await self._client.chat.send_async(
retries=None, **request.model_dump(mode="json", exclude_none=True),
) retries=None,
except Exception as exc: )
transport = self._captured_transport_evidence() except Exception as exc:
self._current_transport_evidence = transport transport = self._captured_transport_evidence(transport_capture)
if isinstance(exc, openrouter_errors.UnauthorizedResponseError): if evidence_capture is not None:
raise ProviderAuthError( evidence_capture.transport_evidence = transport
"OpenRouter authentication failed", if isinstance(exc, openrouter_errors.UnauthorizedResponseError):
raise ProviderAuthError(
"OpenRouter authentication failed",
request_manifest=manifest,
transport_evidence=transport,
failure_phase="http_response" if transport.response_received else "connection",
) from exc
failure_phase = (
"response_validation"
if isinstance(exc, openrouter_errors.ResponseValidationError)
else "http_response"
if transport.response_received
else "connection"
)
raise ProviderError(
self._transport_error_message(transport),
request_manifest=manifest, request_manifest=manifest,
transport_evidence=transport, transport_evidence=transport,
failure_phase="http_response" if transport.response_received else "connection", failure_phase=failure_phase,
) from exc ) from exc
failure_phase = (
"response_validation" transport = self._captured_transport_evidence(transport_capture)
if isinstance(exc, openrouter_errors.ResponseValidationError) if evidence_capture is not None:
else "http_response" evidence_capture.transport_evidence = transport
if transport.response_received raw_api_response = self._coerce_raw_response(response)
else "connection" try:
validated_response = OpenRouterResponse.model_validate(raw_api_response)
except ValidationError as exc:
raise ProviderResponseError(
"OpenRouter response failed schema validation",
request_manifest=manifest,
transport_evidence=transport,
failure_phase="response_validation",
) from exc
try:
text = self._extract_text(validated_response)
except ProviderResponseError as exc:
raise ProviderResponseError(
str(exc),
request_manifest=manifest,
transport_evidence=transport,
failure_phase="response_validation",
) from exc
model = validated_response.model or requested_model or self.model
metadata = self._build_metadata(validated_response)
logger.info("OpenRouter transcription completed using model=%s", model)
return TranscriptionResult(
text=text,
provider="openrouter",
prompt_name=None,
prompt_hash=None,
system_prompt=None,
user_prompt=prompt_text,
temperature=temperature,
top_p=top_p,
model=model,
metadata=metadata,
raw_api_response=raw_api_response,
request_manifest=manifest,
transport_evidence=transport,
) )
raise ProviderError( finally:
self._transport_error_message(transport), client = self._capturing_client
request_manifest=manifest, if token is not None and client is not None:
transport_evidence=transport, client.end_capture(token)
failure_phase=failure_phase,
) from exc
transport = self._captured_transport_evidence()
self._current_transport_evidence = transport
raw_api_response = self._coerce_raw_response(response)
try:
validated_response = OpenRouterResponse.model_validate(raw_api_response)
except ValidationError as exc:
raise ProviderResponseError(
"OpenRouter response failed schema validation",
request_manifest=manifest,
transport_evidence=transport,
failure_phase="response_validation",
) from exc
try:
text = self._extract_text(validated_response)
except ProviderResponseError as exc:
raise ProviderResponseError(
str(exc),
request_manifest=manifest,
transport_evidence=transport,
failure_phase="response_validation",
) from exc
model = validated_response.model or requested_model or self.model
metadata = self._build_metadata(validated_response)
logger.info("OpenRouter transcription completed using model=%s", model)
return TranscriptionResult(
text=text,
provider="openrouter",
prompt_name=None,
prompt_hash=None,
system_prompt=None,
user_prompt=prompt_text,
temperature=temperature,
top_p=top_p,
model=model,
metadata=metadata,
raw_api_response=raw_api_response,
request_manifest=manifest,
transport_evidence=transport,
)
def _build_request_manifest( def _build_request_manifest(
self, self,
@@ -345,6 +357,7 @@ class OpenRouterTranscriptionProvider:
top_p: float | None, top_p: float | None,
) -> RequestManifest | None: ) -> RequestManifest | None:
if source_reference is None: if source_reference is None:
logger.warning("OpenRouter request manifest omitted because source evidence reference is missing.")
return None return None
request_payload = request.model_dump(mode="json", exclude_none=True) request_payload = request.model_dump(mode="json", exclude_none=True)
sanitized_request = self._replace_embedded_media(request_payload, source_reference=source_reference) sanitized_request = self._replace_embedded_media(request_payload, source_reference=source_reference)
@@ -369,6 +382,7 @@ class OpenRouterTranscriptionProvider:
adapter_name="openrouter", adapter_name="openrouter",
adapter_version=OPENROUTER_ADAPTER_VERSION, adapter_version=OPENROUTER_ADAPTER_VERSION,
client_library="openrouter", client_library="openrouter",
settings=self._settings,
), ),
) )
@@ -392,16 +406,15 @@ class OpenRouterTranscriptionProvider:
return [self._replace_embedded_media(item, source_reference=source_reference) for item in value] return [self._replace_embedded_media(item, source_reference=source_reference) for item in value]
return value return value
def _captured_transport_evidence(self) -> TransportEvidence: def _captured_transport_evidence(self, capture: _TransportCapture) -> TransportEvidence:
response = self._capturing_client.last_response if self._capturing_client is not None else None response = capture.response
if response is None: if response is None:
return TransportEvidence(response_received=False) return TransportEvidence(response_received=False)
headers = filter_safe_response_headers(response.headers) headers = filter_safe_response_headers(response.headers)
body = self._capturing_client.last_body if self._capturing_client is not None else None
return TransportEvidence( return TransportEvidence(
response_received=True, response_received=True,
status_code=response.status_code, status_code=response.status_code,
body=body, body=capture.body,
safe_headers=headers, safe_headers=headers,
content_type=headers.get("content-type"), content_type=headers.get("content-type"),
content_encoding=headers.get("content-encoding"), content_encoding=headers.get("content-encoding"),
+37
View File
@@ -0,0 +1,37 @@
from __future__ import annotations
import asyncio
from collections.abc import Awaitable
from collections.abc import Callable
from typing import TypeVar
from sqlalchemy.exc import IntegrityError
ResultT = TypeVar("ResultT")
async def run_blocking(func: Callable[..., ResultT], /, *args, **kwargs) -> ResultT:
"""Run blocking CPU/filesystem work on a worker thread."""
return await asyncio.to_thread(func, *args, **kwargs)
async def insert_with_sequence_retry(
*,
max_retries: int,
operation: Callable[[int], Awaitable[ResultT]],
on_conflict: Callable[[int, IntegrityError], None] | None = None,
) -> ResultT:
"""Retry a sequence-based insert operation on unique-key conflicts."""
if max_retries < 1:
raise ValueError("max_retries must be at least 1")
for retry in range(1, max_retries + 1):
try:
return await operation(retry)
except IntegrityError as exc:
if on_conflict is not None:
on_conflict(retry, exc)
if retry == max_retries:
raise
raise RuntimeError("insert_with_sequence_retry exhausted retries without returning or raising")
+8
View File
@@ -11,7 +11,9 @@ from ..config import Settings
from .documents import DocumentService from .documents import DocumentService
from .evidence import EvidenceService from .evidence import EvidenceService
from .jobs import JobService from .jobs import JobService
from .maintenance import MaintenanceService
from .people import PeopleService from .people import PeopleService
from .photos import PhotosService
from .prompts import PromptStore from .prompts import PromptStore
from .sources import SourceService from .sources import SourceService
@@ -19,7 +21,9 @@ __all__ = [
"DocumentService", "DocumentService",
"EvidenceService", "EvidenceService",
"JobService", "JobService",
"MaintenanceService",
"PeopleService", "PeopleService",
"PhotosService",
"PromptStore", "PromptStore",
"ServiceBundle", "ServiceBundle",
"SourceService", "SourceService",
@@ -33,7 +37,9 @@ class ServiceBundle:
documents: DocumentService = field(default_factory=DocumentService) documents: DocumentService = field(default_factory=DocumentService)
sources: SourceService = field(default_factory=SourceService) sources: SourceService = field(default_factory=SourceService)
jobs: JobService = field(default_factory=JobService) jobs: JobService = field(default_factory=JobService)
maintenance: MaintenanceService = field(default_factory=MaintenanceService)
people: PeopleService = field(default_factory=PeopleService) people: PeopleService = field(default_factory=PeopleService)
photos: PhotosService = field(default_factory=PhotosService)
evidence: EvidenceService = field(default_factory=EvidenceService) evidence: EvidenceService = field(default_factory=EvidenceService)
@classmethod @classmethod
@@ -50,7 +56,9 @@ class ServiceBundle:
documents=DocumentService(session_factory=session_factory, settings=settings), documents=DocumentService(session_factory=session_factory, settings=settings),
sources=SourceService(session_factory=session_factory, settings=settings), sources=SourceService(session_factory=session_factory, settings=settings),
jobs=JobService(session_factory=session_factory, settings=settings), jobs=JobService(session_factory=session_factory, settings=settings),
maintenance=MaintenanceService(session_factory=session_factory, settings=settings),
people=PeopleService(session_factory=session_factory, settings=settings), people=PeopleService(session_factory=session_factory, settings=settings),
photos=PhotosService(session_factory=session_factory, settings=settings),
evidence=EvidenceService(session_factory=session_factory, settings=settings), evidence=EvidenceService(session_factory=session_factory, settings=settings),
) )
+165 -10
View File
@@ -20,12 +20,15 @@ from ..db.loading import orm_attribute
from ..db.loading import selectinload from ..db.loading import selectinload
from ..db.models import Document from ..db.models import Document
from ..db.models import DocumentPerson from ..db.models import DocumentPerson
from ..db.models import DocumentTag
from ..db.models import DocumentType from ..db.models import DocumentType
from ..db.models import Tag
from ..db.registries import AUTHOR_ROLE_SEMANTIC_KEY from ..db.registries import AUTHOR_ROLE_SEMANTIC_KEY
from ..errors import AppError from ..errors import AppError
from ..errors import ErrorCategory from ..errors import ErrorCategory
from .base import ServiceBase from .base import ServiceBase
from .registry import RegistryService from .registry import RegistryService
from .registry import RegistrySummary
from .source_media import lookup_source_mime_type from .source_media import lookup_source_mime_type
from .source_media import supported_source_formats from .source_media import supported_source_formats
@@ -52,6 +55,10 @@ class DocumentTypeError(DocumentError):
"""Raised when Document Type maintenance fails.""" """Raised when Document Type maintenance fails."""
class TagError(DocumentError):
"""Raised when Tag maintenance fails."""
class DocumentTypeRegistry(RegistryService[DocumentType]): class DocumentTypeRegistry(RegistryService[DocumentType]):
"""Document Type registry maintenance.""" """Document Type registry maintenance."""
@@ -71,15 +78,29 @@ class DocumentTypeRegistry(RegistryService[DocumentType]):
return col(Document.document_type_id) return col(Document.document_type_id)
@dataclass(frozen=True, slots=True) type DocumentTypeSummary = RegistrySummary
class DocumentTypeSummary:
"""Settings read model for a Document Type and its usage count."""
id: UUID
label: str class TagRegistry(RegistryService[Tag]):
is_active: bool """Tag registry maintenance."""
is_built_in: bool
document_count: int model = Tag
error = TagError
noun = "Tag"
short_noun = "tag"
referenced_retainer = "historical Documents"
def reference_model(self) -> type[SQLModel]:
return DocumentTag
def reference_id_column(self) -> Any:
return col(DocumentTag.id)
def reference_key_column(self) -> Any:
return col(DocumentTag.tag_id)
type TagSummary = RegistrySummary
@dataclass(frozen=True, slots=True) @dataclass(frozen=True, slots=True)
@@ -126,6 +147,7 @@ class DocumentService(ServiceBase):
) -> None: ) -> None:
super().__init__(session_factory, settings) super().__init__(session_factory, settings)
self._document_types = DocumentTypeRegistry(self.session_factory, self.settings) self._document_types = DocumentTypeRegistry(self.session_factory, self.settings)
self._tags = TagRegistry(self.session_factory, self.settings)
async def _validate_document_type(self, *, session: AsyncSession, document: Document) -> None: async def _validate_document_type(self, *, session: AsyncSession, document: Document) -> None:
"""Validate the UUID-backed Document Type reference.""" """Validate the UUID-backed Document Type reference."""
@@ -279,6 +301,9 @@ class DocumentService(ServiceBase):
selectinload(Document.document_people).selectinload(orm_attribute(DocumentPerson.person)), selectinload(Document.document_people).selectinload(orm_attribute(DocumentPerson.person)),
selectinload(Document.document_people).selectinload(orm_attribute(DocumentPerson.role_ref)), selectinload(Document.document_people).selectinload(orm_attribute(DocumentPerson.role_ref)),
selectinload(Document.document_type_ref), selectinload(Document.document_type_ref),
selectinload(Document.document_tags).selectinload(orm_attribute(DocumentTag.tag_ref)),
selectinload(Document.sources),
selectinload(Document.jobs),
) )
result = await _session.exec(query) result = await _session.exec(query)
return result.all() return result.all()
@@ -294,6 +319,7 @@ class DocumentService(ServiceBase):
selectinload(Document.document_people).selectinload(orm_attribute(DocumentPerson.person)), selectinload(Document.document_people).selectinload(orm_attribute(DocumentPerson.person)),
selectinload(Document.document_people).selectinload(orm_attribute(DocumentPerson.role_ref)), selectinload(Document.document_people).selectinload(orm_attribute(DocumentPerson.role_ref)),
selectinload(Document.document_type_ref), selectinload(Document.document_type_ref),
selectinload(Document.document_tags).selectinload(orm_attribute(DocumentTag.tag_ref)),
) )
.where(Document.id == document_id) .where(Document.id == document_id)
.execution_options(populate_existing=True) .execution_options(populate_existing=True)
@@ -369,6 +395,15 @@ class DocumentService(ServiceBase):
"""List configured document types.""" """List configured document types."""
return await self._document_types.list_entries(active_only=active_only, session=session) return await self._document_types.list_entries(active_only=active_only, session=session)
async def list_tags(
self,
*,
active_only: bool = True,
session: AsyncSession | None = None,
) -> Sequence[Tag]:
"""List configured tags."""
return await self._tags.list_entries(active_only=active_only, session=session)
async def list_document_type_summaries( async def list_document_type_summaries(
self, self,
*, *,
@@ -377,16 +412,34 @@ class DocumentService(ServiceBase):
"""List Document Types alphabetically with current usage counts.""" """List Document Types alphabetically with current usage counts."""
rows = await self._document_types.list_entries_with_counts(session=session) rows = await self._document_types.list_entries_with_counts(session=session)
return [ return [
DocumentTypeSummary( RegistrySummary(
id=document_type.id, id=document_type.id,
label=document_type.label, label=document_type.label,
is_active=document_type.is_active, is_active=document_type.is_active,
is_built_in=document_type.semantic_key is not None, is_built_in=document_type.semantic_key is not None,
document_count=document_count, reference_count=document_count,
) )
for document_type, document_count in rows for document_type, document_count in rows
] ]
async def list_tag_summaries(
self,
*,
session: AsyncSession | None = None,
) -> Sequence[TagSummary]:
"""List Tags alphabetically with current usage counts."""
rows = await self._tags.list_entries_with_counts(session=session)
return [
RegistrySummary(
id=tag.id,
label=tag.label,
is_active=tag.is_active,
is_built_in=tag.semantic_key is not None,
reference_count=document_count,
)
for tag, document_count in rows
]
async def create_document_type( async def create_document_type(
self, self,
*, *,
@@ -397,6 +450,16 @@ class DocumentService(ServiceBase):
"""Create a UUID-identified Document Type with a unique label.""" """Create a UUID-identified Document Type with a unique label."""
return await self._document_types.create_entry(label=label, is_active=is_active, session=session) return await self._document_types.create_entry(label=label, is_active=is_active, session=session)
async def create_tag(
self,
*,
label: str,
is_active: bool = True,
session: AsyncSession | None = None,
) -> Tag:
"""Create a UUID-identified Tag with a unique label."""
return await self._tags.create_entry(label=label, is_active=is_active, session=session)
async def read_document_type( async def read_document_type(
self, self,
document_type_id: UUID, document_type_id: UUID,
@@ -406,6 +469,15 @@ class DocumentService(ServiceBase):
"""Read a Document Type by id.""" """Read a Document Type by id."""
return await self._document_types.read_entry(document_type_id, session=session) return await self._document_types.read_entry(document_type_id, session=session)
async def read_tag(
self,
tag_id: UUID,
*,
session: AsyncSession | None = None,
) -> Tag:
"""Read a Tag by id."""
return await self._tags.read_entry(tag_id, session=session)
async def update_document_type( async def update_document_type(
self, self,
document_type_id: UUID, document_type_id: UUID,
@@ -422,6 +494,22 @@ class DocumentService(ServiceBase):
session=session, session=session,
) )
async def update_tag(
self,
tag_id: UUID,
*,
label: str,
is_active: bool,
session: AsyncSession | None = None,
) -> Tag:
"""Update a Tag label and active state."""
return await self._tags.update_entry(
tag_id,
label=label,
is_active=is_active,
session=session,
)
async def delete_document_type( async def delete_document_type(
self, self,
document_type_id: UUID, document_type_id: UUID,
@@ -431,6 +519,15 @@ class DocumentService(ServiceBase):
"""Delete an unreferenced Document Type without cascade behavior.""" """Delete an unreferenced Document Type without cascade behavior."""
await self._document_types.delete_entry(document_type_id, session=session) await self._document_types.delete_entry(document_type_id, session=session)
async def delete_tag(
self,
tag_id: UUID,
*,
session: AsyncSession | None = None,
) -> None:
"""Delete an unreferenced Tag without cascade behavior."""
await self._tags.delete_entry(tag_id, session=session)
async def is_document_type_referenced( async def is_document_type_referenced(
self, self,
document_type_id: UUID, document_type_id: UUID,
@@ -440,6 +537,15 @@ class DocumentService(ServiceBase):
"""Return whether a Document references a Document Type.""" """Return whether a Document references a Document Type."""
return await self._document_types.is_referenced(document_type_id, session=session) return await self._document_types.is_referenced(document_type_id, session=session)
async def is_tag_referenced(
self,
tag_id: UUID,
*,
session: AsyncSession | None = None,
) -> bool:
"""Return whether a Document references a Tag."""
return await self._tags.is_referenced(tag_id, session=session)
async def set_document_type( async def set_document_type(
self, self,
*, *,
@@ -455,6 +561,55 @@ class DocumentService(ServiceBase):
await self._finalize(session=_session, caller_session=session, refresh=(document,)) await self._finalize(session=_session, caller_session=session, refresh=(document,))
return document return document
async def sync_document_tags_by_labels(
self,
*,
document_id: UUID,
labels: Sequence[str],
session: AsyncSession | None = None,
) -> None:
"""Replace a Document's tag set using label-based assignment."""
normalized_labels = [self._tags.normalize_label(label) for label in labels]
deduplicated_labels = list(dict.fromkeys(normalized_labels))
label_keys = [self._tags.label_key(label) for label in deduplicated_labels]
async with self._session_scope(session) as _session:
existing_document = await _session.get(Document, document_id)
if existing_document is None:
raise DocumentError(
f"Document with id {document_id} not found",
category=ErrorCategory.NOT_FOUND,
suggestion="Refresh and select an existing document.",
)
existing_tags = (
(await _session.exec(select(Tag).where(col(Tag.normalized_label).in_(label_keys)))).all()
if label_keys
else []
)
tags_by_key = {tag.normalized_label: tag for tag in existing_tags}
selected_tag_ids: set[UUID] = set()
for label in deduplicated_labels:
key = self._tags.label_key(label)
tag = tags_by_key.get(key)
if tag is None:
tag = await self._tags.create_entry(label=label, is_active=True, session=_session)
tags_by_key[key] = tag
selected_tag_ids.add(tag.id)
links = (await _session.exec(select(DocumentTag).where(DocumentTag.document_id == document_id))).all()
existing_ids = {link.tag_id for link in links}
for link in links:
if link.tag_id not in selected_tag_ids:
await _session.delete(link)
for tag_id in selected_tag_ids - existing_ids:
_session.add(DocumentTag(document_id=document_id, tag_id=tag_id))
await self._finalize(session=_session, caller_session=session)
def _print_media_type(filename: str) -> str: def _print_media_type(filename: str) -> str:
"""Resolve a stored Source filename to its MIME type for print rendering.""" """Resolve a stored Source filename to its MIME type for print rendering."""
+4
View File
@@ -16,6 +16,10 @@ class PromptLoadError(AppError):
"""Raised when prompt artifacts cannot be loaded safely.""" """Raised when prompt artifacts cannot be loaded safely."""
class PromptStoreError(PromptLoadError):
"""Raised when prompt storage validation or persistence fails."""
class TranscriptionError(AppError): class TranscriptionError(AppError):
"""Raised when transcription execution fails.""" """Raised when transcription execution fails."""
+20
View File
@@ -43,6 +43,26 @@ class LatestExecutionAttempt:
class EvidenceService(ServiceBase): class EvidenceService(ServiceBase):
"""Read, project, and export execution attempt evidence.""" """Read, project, and export execution attempt evidence."""
async def read_latest_job_error_category(
self,
*,
job_id: UUID,
session: AsyncSession | None = None,
) -> str | None:
"""Read the latest persisted execution-attempt error category for a job."""
async with self._session_scope(session) as _session:
query = (
select(ExecutionAttempt.error_category)
.where(ExecutionAttempt.job_id == job_id)
.where(col(ExecutionAttempt.error_category).is_not(None))
.order_by(
col(ExecutionAttempt.created_at).desc(),
col(ExecutionAttempt.id).desc(),
)
.limit(1)
)
return (await _session.exec(query)).first()
async def read_latest_execution_attempt( async def read_latest_execution_attempt(
self, self,
*, *,

Some files were not shown because too many files have changed in this diff Show More