unified implementation plan

This commit is contained in:
John Lancaster
2026-07-31 10:28:47 -05:00
parent bbf7fe28c2
commit 6b5b0500b3
3 changed files with 42 additions and 416 deletions
+41 -232
View File
@@ -1,248 +1,57 @@
# Implementation Plan (Version 2) # implementation_plan_v2
This plan defines the path from the V1 baseline to **Version 2 complete**, aligned to the updated multi-image and multi-person relational domain model: ## Goal
* `Document` acts as a logical parent container for physical artifacts, supporting multi-author and multi-recipient relationships via `DocumentPerson`. Replace the current V1 SQLModel schema with the approved V2 schema and make sure every database operation works through the existing async SQLAlchemy/SQLModel session layer.
* `Source` represents an individual image page within a document, maintaining sequential order (`page_number`), cached active machine output (`raw_transcription`), and inline single user revisions (`revised_text`).
* `Job` acts as an overarching batch orchestrator for multi-page async processing tasks.
* `JobSource` records individual point-in-time API executions per image page, storing Pydantic-validated `ai_metadata` and raw REST envelopes (`raw_api_response`).
* **Pydantic V2** acts as the single source of truth for runtime validation, API payload parsing, and PostgreSQL JSONB serialization.
The objective is to complete the V2 scope with production readiness while keeping non-V2 enhancements out of active delivery. Use a fresh database. There will be no migrations, data conversion, legacy compatibility shims, or parallel V1/V2 code paths.
--- ## Current Project Impact
## V2 Completion Definition - `src/transcription/db/models.py` still defines the V1 `Document`, `Source`, `Job`, and `Revision` tables.
- The V2 target adds `Person`, `DocumentPerson`, and `JobSource`, moves revisions onto `Source`, and removes the direct `Source.job_id` relationship.
- The engine, session factory, transaction handling, and PostgreSQL async support already exist and do not need to be rewritten.
- Async CRUD currently lives in `DocumentService`, `JobService`, `TranscriptionService`, and the upload record helper. Their queries and eager-loading options depend on V1 relationships.
- Existing tests cover only part of the schema and CRUD surface.
V2 is complete when all of the following are true: ## Implementation
1. **Functional complete** ### 1. Update the schema
* Multi-image and whole-folder uploads assign sequential page numbers to `Source` records under a single `Document`.
* Batch jobs process pages concurrently using an `asyncio` worker pool with semaphore rate limiting.
* Partial job failures resolve cleanly to `partial_success`, allowing single-page retries without re-running successful pages.
* Multi-author and multi-recipient tagging is supported on `Document`.
- Replace the models in `src/transcription/db/models.py` with the approved V2 tables, enums, relationships, foreign keys, constraints, and indexes.
- Remove `Revision`, `Source.job_id`, and the transcription fields that no longer belong on `Job`.
- Keep `create_all()` as the schema bootstrap for a fresh database.
- Delete `_ensure_sqlite_compat_columns()` and all schema patching from `src/transcription/db/operations.py`.
- Keep the Python models, `docs/schema_v2.md`, and `docs/ddl_v2.sql` consistent.
2. **Data-model complete** ### 2. Align the async CRUD methods
* SQLite is fully replaced with PostgreSQL (using `asyncpg` or `psycopg3`).
* Pydantic V2 models validate all API payloads, database row mappings, and `JSONB` structures.
- Keep the existing `ServiceBase` session and transaction pattern.
- Update document CRUD to load and manage its ordered `Source` rows and `DocumentPerson` links.
- Update job CRUD and queue queries to use `JobSource` instead of `Source.job_id`.
- Add the missing async CRUD operations for `Person`, `Source`, `DocumentPerson`, and `JobSource` using the existing service style. Do not add another repository abstraction.
- Replace revision CRUD with direct updates to `Source.revised_text` and `Source.date_revised`.
- Remove the temporary transcript compatibility aliases instead of redirecting them.
- Update only direct database call sites that construct or query these records; UI and worker feature changes are not part of this work.
3. **Operational complete** ### 3. Verify the schema and CRUD
* Concurrency controls, worker pool metrics, and database connections operate safely under batch load.
- Update the schema bootstrap test to expect `person`, `document`, `document_person`, `source`, `job`, and `job_source`, with no `revision` table.
- Add async create, read, update, delete, list, and filtered-query tests for each entity that exposes those operations.
- Test relationship loading, page ordering, uniqueness constraints, delete behavior, status values, and `JobSource` JSON fields.
- Test both service-owned sessions and caller-provided sessions so flush/commit behavior remains correct.
- Run the focused database and service tests, then the full suite with `uv run pytest`.
4. **Documentation complete** ## Done When
* `schema_v2.md`, `DDL_v2.sql`, Pydantic model contracts are updated and consistent.
---
## Phase 1 — Data Contract Stabilization & Pydantic Baseline
**Goal:** Lock the PostgreSQL schema, DDL, and Pydantic V2 models before refactoring service logic.
### Tasks
1. Finalize DDL for PostgreSQL native types (`UUID`, `TIMESTAMPTZ`, `JSONB`) and junction tables (`document_person`, `job_source`).
2. Build core Pydantic V2 schemas (`Person`, `Document`, `Source`, `Job`, `JobSource`, `PageAIMetadata`).
3. Confirm and document data invariants:
* `source.raw_transcription` and `job_source.raw_transcription` are immutable machine outputs.
* `source.revised_text` holds user edits. UI renders `COALESCE(revised_text, raw_transcription)`.
* Page sequence is strictly ordered by `source.page_number ASC`.
4. Freeze V2 job status values (`queued`, `processing`, `completed`, `partial_success`, `failed`) and page execution status values (`pending`, `transcribed`, `failed`).
### Deliverables
* Canonical `docs/schema_v2.md` and `docs/DDL_v2.sql`.
* Centralized Pydantic validation suite in `models/schemas_v2.py`.
### Exit Criteria
* All database tables, relationships, and JSONB structures have corresponding Pydantic V2 models passing unit validation tests.
---
## Phase 2 — Persistence Layer Transition (SQLite to PostgreSQL)
**Goal:** Replace the SQLite storage layer with an asynchronous PostgreSQL driver (`asyncpg` or `psycopg3`).
### Tasks
1. Configure PostgreSQL database connection pooling and environment configuration.
2. Refactor `services/store.py` / repository layers to execute parameterized async SQL queries (`$1`, `$2`).
3. Implement JSONB serialization and deserialization helpers using Pydantic's `.model_dump_json()` and `.model_validate()`.
4. Implement database bootstrap routines for PostgreSQL table creation and index initialization.
### Deliverables
* PostgreSQL-native database connection and query service modules.
* Integration test suite confirming connection pooling and JSONB CRUD operations.
### Exit Criteria
* All database reads/writes run asynchronously against PostgreSQL with zero remaining SQLite driver dependencies.
---
## Phase 3 — Service Layer & `asyncio` Engine Refactor
**Goal:** Implement batch orchestration and parallel single-image API execution.
### Tasks
1. Refactor upload service to process folder/multi-image input:
* Group files into a single `Document`.
* Create ordered `Source` rows (`page_number = 1..N`).
2. Refactor `services/workflows.py` with `asyncio` worker pools:
* Use `asyncio.Semaphore` to enforce API provider rate limits.
* Issue parallel single-image requests to Vision APIs (OpenAI/Claude).
* Parse API responses directly into Pydantic models (`PageAIMetadata`).
3. Update execution tracking:
* Create a `JobSource` row per page call to record `raw_transcription`, `ai_metadata`, and `raw_api_response`.
* Update active `source.raw_transcription` upon task completion.
* Calculate aggregate batch status (`completed`, `partial_success`, `failed`) on the parent `Job`.
4. Refactor `services/person.py` and `services/documents.py` to handle multi-person roles via `document_person`.
### Deliverables
* Asynchronous batch execution engine in `services/workflows.py`.
* Service routines for multi-person tagging and page-level retries.
### Exit Criteria
* Executing a folder upload of 10+ images processes concurrently, populates page-level `JobSource` entries, and handles partial worker errors without crashing the batch.
---
## Phase 4 — UI & API Contract Alignment
**Goal:** Update API endpoints and frontend/UI views to render multi-page documents and person roles.
### Tasks
1. Update document and job API endpoints to accept batch file arrays and multi-person ID payloads.
2. Update UI document views:
* Render multi-page document transcriptions sequentially by `page_number`.
* Display author and recipient chips/cards linked from `document_person`.
3. Update job detail UI to show page-level execution statuses (`transcribed` vs. `failed`) and provide a "Retry Failed Pages" action for `partial_success` jobs.
4. Align inline page editing controls to update `source.revised_text` and `source.date_revised`.
### Deliverables
* Refactored API routes and UI components supporting multi-page rendering and person management.
### Exit Criteria
* UI successfully displays multi-page document text, allows per-page human revisions, and shows author/recipient metadata.
---
## Phase 5 — Test Suite Realignment & Concurrency Testing
**Goal:** Ensure end-to-end system stability under concurrent async execution and load.
### Tasks
1. Write unit tests for Pydantic models, custom validators, and JSONB conversions.
2. Write integration tests for async database operations:
* CRUD for `Document`, `Person`, `DocumentPerson`, `Source`, `Job`, and `JobSource`.
3. Write mock-backed async workflow tests:
* Verify `asyncio.Semaphore` bounds concurrent tasks properly.
* Validate state transition logic for `completed`, `partial_success`, and `failed` jobs.
* Confirm retry routines process only targeted `JobSource` records marked as `failed`.
4. Re-enable CI quality gates (linting, type checking with Pyright/mypy, pytest).
### Deliverables
* Passing asynchronous test suite covering core workflows, edge cases, and failure recoveries.
### Exit Criteria
* CI pipeline is green with comprehensive coverage across database operations, Pydantic models, and worker queues.
---
## Phase 6 — Reliability, Operations, and Release Readiness
**Goal:** Prepare V2 for production deployment and operator management.
### Tasks
1. Verify structured logging includes `job_id`, `document_id`, `source_id`, and `person_id`.
2. Tune PostgreSQL connection pool limits and `asyncio` concurrency thresholds for production infrastructure.
3. Update operational documentation:
* Review and update `docs/schema_v2.md` as needed.
* Create `docs/runbook_v2.md` detailing PostgreSQL maintenance, JSONB index management, and worker queue monitoring.
* Create `docs/release_checklist_v2.md` for launch sign-off.
### Deliverables
* Updated project documentation and operational runbooks.
* V2 release sign-off checklist.
### Exit Criteria
* All documentation reflects V2 architecture; launch checklist is fully verified.
---
## Requirement Traceability Focus
Maintain evidence against these V2 requirement groups:
* **Batch & Multi-Image Pipeline:** Folder ingestion, page ordering, async worker execution.
* **Database & Persistence:** PostgreSQL, native UUIDs, JSONB execution storage, `asyncpg` pooling.
* **Validation & Schemas:** Pydantic V2 models for DB rows, API requests, and AI vision responses.
* **Attribution & Metadata:** Multi-author and multi-recipient tagging, biographical entity management.
* **Error Recovery:** Partial success states, page-level status flags, isolated retry execution.
---
## Scope Discipline Rule (V2 Focus)
* Only tasks required for V2 scope (PostgreSQL, Pydantic V2, folder/async processing, multi-person roles) enter this plan.
* V3 candidate features (such as side-by-side multi-provider model output comparison) remain strictly in the future backlog.
* Any schema adjustments during implementation require immediate updates to `DDL_v2.sql`, Pydantic models, and `schema_v2.md`.
---
## Technology References
- [FastAPI documentation](https://fastapi.tiangolo.com/)
- [NiceGUI documentation](https://nicegui.io/documentation)
- [PostgreSQL documentation](https://www.postgresql.org/docs/)
- [Python asyncio](https://docs.python.org/3/library/asyncio.html#module-asyncio)
- [Pydantic Validation](https://pydantic.dev/docs/validation/latest/get-started/)
- [Pydantic AI](https://pydantic.dev/docs/ai/overview/)
## Related Local References
- [System Overview](index_v2.md)
- [System Design Intent](intent.md)
- [Transcription Methodology](transcription_methodology.md)
- [System Architecture](architecture_v2.md)
- [System Requirements](requirements_v2.md)
- [Data model](schema_v2.md)
- [Error Handling Policy](error_handling_v2.md)
- Implementation Plan (this document)
- A fresh database is created directly from the V2 SQLModel metadata.
- All async CRUD methods pass against the V2 relationships and fields.
- No code references `Revision`, `Source.job_id`, removed `Job` transcription fields, or compatibility aliases.
- The focused tests and full test suite pass.
## Out of Scope
- Database migrations or preservation of V1 data
- Legacy compatibility code
- Database engine or session-layer rewrites
- UI redesign, batch orchestration, worker concurrency, deployment, and operational runbooks
+1 -1
View File
@@ -19,7 +19,7 @@ Read [architecture_v2.md](architecture_v2.md) first for technical overview and s
## Technical Stack ## Technical Stack
* **Application Web Framework:** FastAPI + NiceGUI * **Application Web Framework:** FastAPI + NiceGUI
* **Persistence Engine:** PostgreSQL 13+ * **Persistence Engine:** PostgreSQL 18+
* **Data Validation & Schemas:** Pydantic V2 * **Data Validation & Schemas:** Pydantic V2
* **Concurrency & Workers:** Python `asyncio` worker pool with `asyncio.Semaphore` * **Concurrency & Workers:** Python `asyncio` worker pool with `asyncio.Semaphore`
* **Vision Providers:** OpenAI (GPT-4o) and Anthropic (Claude 3.5 Sonnet) via native SDKs * **Vision Providers:** OpenAI (GPT-4o) and Anthropic (Claude 3.5 Sonnet) via native SDKs
-183
View File
@@ -1,183 +0,0 @@
# Version 2 Plan
## Purpose
Version 2 updates the existing SQLModel domain schema to support multi-page documents, page-level transcription results, richer document metadata, and author/recipient attribution.
PostgreSQL support is already present in the database runtime. V2 does not require a database-layer rewrite or a general SQLite-to-PostgreSQL migration system. PostgreSQL adoption consists primarily of selecting the existing PostgreSQL settings, provisioning the database, creating the V2 schema, and verifying the application against it.
The main implementation effort is the schema update and the application changes that depend on it.
---
## Current State
- The application uses Python 3.12, Pydantic v2, SQLModel, and async SQLAlchemy sessions.
- The database engine already supports both SQLite and PostgreSQL through `SqliteSettings` and `PostgresSettings`.
- The PostgreSQL async driver is installed and the engine already builds `postgresql+asyncpg` connections.
- Schema bootstrap currently uses `SQLModel.metadata.create_all()`.
- SQLite remains the default local configuration and the current Compose configuration still selects SQLite.
- The current V1 domain contains `Document`, `Source`, `Job`, and `Revision` tables.
- The V2 target is defined in [V2 DB Schema](V2%20DB%20Schema.md) and [V2 PostgreSQL DDL Specification](V2%20PostgreSQL%20DDL%20Specification.md).
---
## V2 Outcomes
1. **V2 schema implemented**
- SQLModel models, relationships, enums, constraints, and indexes match the approved V2 schema.
1. **Page-level batch processing supported**
- A job can process multiple sources and retain an independent result for each source through `JobSource`.
1. **Document metadata expanded**
- Documents support ordered pages, descriptive metadata, and multiple authors and recipients.
1. **Raw provider data retained**
- Complete provider payloads are stored in PostgreSQL `JSONB` without flattening or discarding fields.
- Stored documents remain suitable for a future MongoDB import if one is ever needed.
1. **PostgreSQL enabled through configuration**
- The application starts against a provisioned PostgreSQL database using the existing runtime path.
1. **Existing workflows remain reliable**
- Upload, worker execution, status inspection, and transcription revision work with the new schema.
---
## Non-Goals
- Rewriting the database engine or session layer
- Building a general-purpose SQLite-to-PostgreSQL migration utility
- Rehearsing a production database cutover when no production dataset requires preservation
- Running or integrating MongoDB in V2
- Building MongoDB projections, synchronization, or fallback behavior
- Replacing Python, Pydantic, SQLModel, or SQLAlchemy
- Supporting more than one active human revision per source
If an existing SQLite dataset must be retained, define a small one-time import task separately. It is not part of the default V2 implementation path.
---
## Scope
### A) SQLModel schema update (primary)
- Add `Person`, `DocumentPerson`, and `JobSource` models.
- Expand `Document` with type, date, location, notes, archive identifier, and timestamps.
- Update `Source` with page ordering, active raw transcription, revised text, and revision timestamp.
- Update `Job` for batch execution and the `partial_success` terminal state.
- Replace the standalone `Revision` table with revision fields on `Source`.
- Remove the direct `Source.job_id` relationship; connect sources to jobs through `JobSource`.
- Add role, job status, and job-source status enums.
- Add required uniqueness constraints, foreign-key delete behavior, lookup indexes, and PostgreSQL JSON indexes.
- Keep model definitions aligned with [V2 Python Pydantic Models](V2%20Python%20Pydantic%20Models.md).
### B) Raw document storage
- Store structured AI metadata in `job_source.ai_metadata` as `JSONB`.
- Store the complete raw provider response in `job_source.raw_api_response` as `JSONB`.
- Preserve the original document structure, field names, nested values, and unknown fields in the raw response.
- Keep validation of extracted application fields separate from retention of the raw response.
- Serialize UUIDs and datetimes using portable string representations.
- Use MongoDB Extended JSON representations only if a provider value cannot be represented faithfully in standard JSON.
- Do not store opaque BSON bytes in PostgreSQL unless a future payload contains BSON-only values that cannot be preserved in `JSONB`.
### C) Schema creation and verification
- Use a fresh V2 database during development unless preservation of existing data becomes a requirement.
- Create the schema from SQLModel metadata and verify it against the approved DDL.
- Keep SQLite available for fast unit tests where its behavior is equivalent.
- Add focused PostgreSQL integration tests for native UUIDs, JSON storage, constraints, indexes, and transactions.
- Introduce migration tooling only if V2 must update a populated deployed database in place.
### D) Service and worker alignment
- Update document, source, job, and store operations for the new relationships.
- Create one `JobSource` row per source included in a job.
- Persist page-level status, transcription, AI metadata, raw provider response, and errors on `JobSource`.
- Derive the parent job status from its page results:
- `completed` when all pages succeed
- `partial_success` when successful and failed pages are mixed
- `failed` when all pages fail or a job-level failure prevents execution
- Update `Source.raw_transcription` after a successful page result while keeping the original `JobSource.raw_transcription` immutable.
- Read `Source.revised_text` in preference to `Source.raw_transcription` when presenting active text.
### E) Document and multi-image workflows
- Require a document before associating uploaded sources.
- Support uploading multiple images into one document.
- Preserve page order through `Source.page_number`.
- Allow a job to include one or more sources from the same document.
- Update job details to show the document, each source filename, page order, page status, and page-level errors.
- Define a practical upload limit and split oversized selections into manageable batches if needed.
### F) PostgreSQL configuration
- Provision PostgreSQL for local and deployed environments.
- Configure the existing `Settings.database` field with PostgreSQL host, port, database, user, and password values.
- Update Compose and environment configuration to stop selecting SQLite.
- Decide whether schema bootstrap is enabled for local development or performed as a separate deployment step.
- Run a connectivity and schema smoke test against PostgreSQL.
- Keep uploaded files on a persistent, backup-capable path outside the application image.
---
## Milestones
### M1 - Schema models
- Implement the V2 SQLModel models and enums.
- Implement relationships, constraints, indexes, and JSON column types.
- Update the Pydantic data contracts where model decisions change.
- Add schema-focused tests.
**Exit criteria:** SQLModel metadata represents the approved V2 schema and schema tests pass.
### M2 - Persistence and worker behavior
- Update database operations and services for the V2 entities.
- Implement page-level `JobSource` execution records.
- Preserve complete raw provider responses in `JSONB`.
- Implement aggregate job status calculation.
- Add transaction, partial-success, and failure-isolation tests.
**Exit criteria:** single-page and multi-page jobs persist correct page, raw payload, and aggregate states.
### M3 - Document and upload workflows
- Update document creation and source association flows.
- Add ordered multi-image upload.
- Update job and document detail views for page-level results.
- Add focused UI and service tests.
**Exit criteria:** a user can create a document, upload ordered pages, run a job, and inspect each result.
### M4 - PostgreSQL verification and release
- Switch local or test configuration to the existing PostgreSQL runtime path.
- Create the V2 schema in a fresh PostgreSQL database.
- Run PostgreSQL-specific schema and workflow tests.
- Document startup, backup, and recovery settings.
- Run the final regression suite.
**Exit criteria:** V2 workflows pass against PostgreSQL and release checks are complete.
---
## Risks and Mitigations
- **Model and DDL drift** -> compare generated metadata with the approved schema and test named constraints and indexes.
- **Raw payload loss** -> retain the complete provider response separately from validated and extracted fields.
- **Cross-database differences** -> retain fast SQLite tests but verify PostgreSQL-native UUID, JSON, and index behavior in integration tests.
- **Batch state errors** -> test all-success, mixed-result, and all-failed jobs explicitly.
- **Page ordering errors** -> enforce uniqueness and ordering rules for document pages.
- **Unexpected data-preservation need** -> confirm whether existing SQLite data matters before implementation; add a one-time importer only when required.
- **Worker regressions** -> preserve terminal-state and retry reliability tests while changing persistence ownership.
---
## Suggested First Tasks
1. Update `src/transcription/db/models.py` to represent the approved V2 schema.
2. Add schema tests for tables, columns, relationships, constraints, indexes, and enums.
3. Define and test lossless raw provider response storage in `job_source.raw_api_response`.
4. Update database operations and services to use `JobSource` and source-level revisions.
5. Add page-result aggregation tests before changing the worker workflow.
6. Update document and multi-image upload flows.
7. Select PostgreSQL in configuration and run the integration suite against a fresh V2 database.