generated from john/python-template
unified implementation plan
This commit is contained in:
+41
-232
@@ -1,248 +1,57 @@
|
|||||||
# Implementation Plan (Version 2)
|
# implementation_plan_v2
|
||||||
|
|
||||||
This plan defines the path from the V1 baseline to **Version 2 complete**, aligned to the updated multi-image and multi-person relational domain model:
|
## Goal
|
||||||
|
|
||||||
* `Document` acts as a logical parent container for physical artifacts, supporting multi-author and multi-recipient relationships via `DocumentPerson`.
|
Replace the current V1 SQLModel schema with the approved V2 schema and make sure every database operation works through the existing async SQLAlchemy/SQLModel session layer.
|
||||||
* `Source` represents an individual image page within a document, maintaining sequential order (`page_number`), cached active machine output (`raw_transcription`), and inline single user revisions (`revised_text`).
|
|
||||||
* `Job` acts as an overarching batch orchestrator for multi-page async processing tasks.
|
|
||||||
* `JobSource` records individual point-in-time API executions per image page, storing Pydantic-validated `ai_metadata` and raw REST envelopes (`raw_api_response`).
|
|
||||||
* **Pydantic V2** acts as the single source of truth for runtime validation, API payload parsing, and PostgreSQL JSONB serialization.
|
|
||||||
|
|
||||||
The objective is to complete the V2 scope with production readiness while keeping non-V2 enhancements out of active delivery.
|
Use a fresh database. There will be no migrations, data conversion, legacy compatibility shims, or parallel V1/V2 code paths.
|
||||||
|
|
||||||
---
|
## Current Project Impact
|
||||||
|
|
||||||
## V2 Completion Definition
|
- `src/transcription/db/models.py` still defines the V1 `Document`, `Source`, `Job`, and `Revision` tables.
|
||||||
|
- The V2 target adds `Person`, `DocumentPerson`, and `JobSource`, moves revisions onto `Source`, and removes the direct `Source.job_id` relationship.
|
||||||
|
- The engine, session factory, transaction handling, and PostgreSQL async support already exist and do not need to be rewritten.
|
||||||
|
- Async CRUD currently lives in `DocumentService`, `JobService`, `TranscriptionService`, and the upload record helper. Their queries and eager-loading options depend on V1 relationships.
|
||||||
|
- Existing tests cover only part of the schema and CRUD surface.
|
||||||
|
|
||||||
V2 is complete when all of the following are true:
|
## Implementation
|
||||||
|
|
||||||
1. **Functional complete**
|
### 1. Update the schema
|
||||||
* Multi-image and whole-folder uploads assign sequential page numbers to `Source` records under a single `Document`.
|
|
||||||
* Batch jobs process pages concurrently using an `asyncio` worker pool with semaphore rate limiting.
|
|
||||||
* Partial job failures resolve cleanly to `partial_success`, allowing single-page retries without re-running successful pages.
|
|
||||||
* Multi-author and multi-recipient tagging is supported on `Document`.
|
|
||||||
|
|
||||||
|
- Replace the models in `src/transcription/db/models.py` with the approved V2 tables, enums, relationships, foreign keys, constraints, and indexes.
|
||||||
|
- Remove `Revision`, `Source.job_id`, and the transcription fields that no longer belong on `Job`.
|
||||||
|
- Keep `create_all()` as the schema bootstrap for a fresh database.
|
||||||
|
- Delete `_ensure_sqlite_compat_columns()` and all schema patching from `src/transcription/db/operations.py`.
|
||||||
|
- Keep the Python models, `docs/schema_v2.md`, and `docs/ddl_v2.sql` consistent.
|
||||||
|
|
||||||
2. **Data-model complete**
|
### 2. Align the async CRUD methods
|
||||||
* SQLite is fully replaced with PostgreSQL (using `asyncpg` or `psycopg3`).
|
|
||||||
* Pydantic V2 models validate all API payloads, database row mappings, and `JSONB` structures.
|
|
||||||
|
|
||||||
|
- Keep the existing `ServiceBase` session and transaction pattern.
|
||||||
|
- Update document CRUD to load and manage its ordered `Source` rows and `DocumentPerson` links.
|
||||||
|
- Update job CRUD and queue queries to use `JobSource` instead of `Source.job_id`.
|
||||||
|
- Add the missing async CRUD operations for `Person`, `Source`, `DocumentPerson`, and `JobSource` using the existing service style. Do not add another repository abstraction.
|
||||||
|
- Replace revision CRUD with direct updates to `Source.revised_text` and `Source.date_revised`.
|
||||||
|
- Remove the temporary transcript compatibility aliases instead of redirecting them.
|
||||||
|
- Update only direct database call sites that construct or query these records; UI and worker feature changes are not part of this work.
|
||||||
|
|
||||||
3. **Operational complete**
|
### 3. Verify the schema and CRUD
|
||||||
* Concurrency controls, worker pool metrics, and database connections operate safely under batch load.
|
|
||||||
|
|
||||||
|
- Update the schema bootstrap test to expect `person`, `document`, `document_person`, `source`, `job`, and `job_source`, with no `revision` table.
|
||||||
|
- Add async create, read, update, delete, list, and filtered-query tests for each entity that exposes those operations.
|
||||||
|
- Test relationship loading, page ordering, uniqueness constraints, delete behavior, status values, and `JobSource` JSON fields.
|
||||||
|
- Test both service-owned sessions and caller-provided sessions so flush/commit behavior remains correct.
|
||||||
|
- Run the focused database and service tests, then the full suite with `uv run pytest`.
|
||||||
|
|
||||||
4. **Documentation complete**
|
## Done When
|
||||||
* `schema_v2.md`, `DDL_v2.sql`, Pydantic model contracts are updated and consistent.
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Phase 1 — Data Contract Stabilization & Pydantic Baseline
|
|
||||||
|
|
||||||
**Goal:** Lock the PostgreSQL schema, DDL, and Pydantic V2 models before refactoring service logic.
|
|
||||||
|
|
||||||
### Tasks
|
|
||||||
|
|
||||||
1. Finalize DDL for PostgreSQL native types (`UUID`, `TIMESTAMPTZ`, `JSONB`) and junction tables (`document_person`, `job_source`).
|
|
||||||
2. Build core Pydantic V2 schemas (`Person`, `Document`, `Source`, `Job`, `JobSource`, `PageAIMetadata`).
|
|
||||||
3. Confirm and document data invariants:
|
|
||||||
* `source.raw_transcription` and `job_source.raw_transcription` are immutable machine outputs.
|
|
||||||
* `source.revised_text` holds user edits. UI renders `COALESCE(revised_text, raw_transcription)`.
|
|
||||||
* Page sequence is strictly ordered by `source.page_number ASC`.
|
|
||||||
|
|
||||||
|
|
||||||
4. Freeze V2 job status values (`queued`, `processing`, `completed`, `partial_success`, `failed`) and page execution status values (`pending`, `transcribed`, `failed`).
|
|
||||||
|
|
||||||
### Deliverables
|
|
||||||
|
|
||||||
* Canonical `docs/schema_v2.md` and `docs/DDL_v2.sql`.
|
|
||||||
* Centralized Pydantic validation suite in `models/schemas_v2.py`.
|
|
||||||
|
|
||||||
### Exit Criteria
|
|
||||||
|
|
||||||
* All database tables, relationships, and JSONB structures have corresponding Pydantic V2 models passing unit validation tests.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Phase 2 — Persistence Layer Transition (SQLite to PostgreSQL)
|
|
||||||
|
|
||||||
**Goal:** Replace the SQLite storage layer with an asynchronous PostgreSQL driver (`asyncpg` or `psycopg3`).
|
|
||||||
|
|
||||||
### Tasks
|
|
||||||
|
|
||||||
1. Configure PostgreSQL database connection pooling and environment configuration.
|
|
||||||
2. Refactor `services/store.py` / repository layers to execute parameterized async SQL queries (`$1`, `$2`).
|
|
||||||
3. Implement JSONB serialization and deserialization helpers using Pydantic's `.model_dump_json()` and `.model_validate()`.
|
|
||||||
4. Implement database bootstrap routines for PostgreSQL table creation and index initialization.
|
|
||||||
|
|
||||||
### Deliverables
|
|
||||||
|
|
||||||
* PostgreSQL-native database connection and query service modules.
|
|
||||||
* Integration test suite confirming connection pooling and JSONB CRUD operations.
|
|
||||||
|
|
||||||
### Exit Criteria
|
|
||||||
|
|
||||||
* All database reads/writes run asynchronously against PostgreSQL with zero remaining SQLite driver dependencies.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Phase 3 — Service Layer & `asyncio` Engine Refactor
|
|
||||||
|
|
||||||
**Goal:** Implement batch orchestration and parallel single-image API execution.
|
|
||||||
|
|
||||||
### Tasks
|
|
||||||
|
|
||||||
1. Refactor upload service to process folder/multi-image input:
|
|
||||||
* Group files into a single `Document`.
|
|
||||||
* Create ordered `Source` rows (`page_number = 1..N`).
|
|
||||||
|
|
||||||
|
|
||||||
2. Refactor `services/workflows.py` with `asyncio` worker pools:
|
|
||||||
* Use `asyncio.Semaphore` to enforce API provider rate limits.
|
|
||||||
* Issue parallel single-image requests to Vision APIs (OpenAI/Claude).
|
|
||||||
* Parse API responses directly into Pydantic models (`PageAIMetadata`).
|
|
||||||
|
|
||||||
|
|
||||||
3. Update execution tracking:
|
|
||||||
* Create a `JobSource` row per page call to record `raw_transcription`, `ai_metadata`, and `raw_api_response`.
|
|
||||||
* Update active `source.raw_transcription` upon task completion.
|
|
||||||
* Calculate aggregate batch status (`completed`, `partial_success`, `failed`) on the parent `Job`.
|
|
||||||
|
|
||||||
|
|
||||||
4. Refactor `services/person.py` and `services/documents.py` to handle multi-person roles via `document_person`.
|
|
||||||
|
|
||||||
### Deliverables
|
|
||||||
|
|
||||||
* Asynchronous batch execution engine in `services/workflows.py`.
|
|
||||||
* Service routines for multi-person tagging and page-level retries.
|
|
||||||
|
|
||||||
### Exit Criteria
|
|
||||||
|
|
||||||
* Executing a folder upload of 10+ images processes concurrently, populates page-level `JobSource` entries, and handles partial worker errors without crashing the batch.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Phase 4 — UI & API Contract Alignment
|
|
||||||
|
|
||||||
**Goal:** Update API endpoints and frontend/UI views to render multi-page documents and person roles.
|
|
||||||
|
|
||||||
### Tasks
|
|
||||||
|
|
||||||
1. Update document and job API endpoints to accept batch file arrays and multi-person ID payloads.
|
|
||||||
2. Update UI document views:
|
|
||||||
* Render multi-page document transcriptions sequentially by `page_number`.
|
|
||||||
* Display author and recipient chips/cards linked from `document_person`.
|
|
||||||
|
|
||||||
|
|
||||||
3. Update job detail UI to show page-level execution statuses (`transcribed` vs. `failed`) and provide a "Retry Failed Pages" action for `partial_success` jobs.
|
|
||||||
4. Align inline page editing controls to update `source.revised_text` and `source.date_revised`.
|
|
||||||
|
|
||||||
### Deliverables
|
|
||||||
|
|
||||||
* Refactored API routes and UI components supporting multi-page rendering and person management.
|
|
||||||
|
|
||||||
### Exit Criteria
|
|
||||||
|
|
||||||
* UI successfully displays multi-page document text, allows per-page human revisions, and shows author/recipient metadata.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Phase 5 — Test Suite Realignment & Concurrency Testing
|
|
||||||
|
|
||||||
**Goal:** Ensure end-to-end system stability under concurrent async execution and load.
|
|
||||||
|
|
||||||
### Tasks
|
|
||||||
|
|
||||||
1. Write unit tests for Pydantic models, custom validators, and JSONB conversions.
|
|
||||||
2. Write integration tests for async database operations:
|
|
||||||
* CRUD for `Document`, `Person`, `DocumentPerson`, `Source`, `Job`, and `JobSource`.
|
|
||||||
|
|
||||||
|
|
||||||
3. Write mock-backed async workflow tests:
|
|
||||||
* Verify `asyncio.Semaphore` bounds concurrent tasks properly.
|
|
||||||
* Validate state transition logic for `completed`, `partial_success`, and `failed` jobs.
|
|
||||||
* Confirm retry routines process only targeted `JobSource` records marked as `failed`.
|
|
||||||
|
|
||||||
|
|
||||||
4. Re-enable CI quality gates (linting, type checking with Pyright/mypy, pytest).
|
|
||||||
|
|
||||||
### Deliverables
|
|
||||||
|
|
||||||
* Passing asynchronous test suite covering core workflows, edge cases, and failure recoveries.
|
|
||||||
|
|
||||||
### Exit Criteria
|
|
||||||
|
|
||||||
* CI pipeline is green with comprehensive coverage across database operations, Pydantic models, and worker queues.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Phase 6 — Reliability, Operations, and Release Readiness
|
|
||||||
|
|
||||||
**Goal:** Prepare V2 for production deployment and operator management.
|
|
||||||
|
|
||||||
### Tasks
|
|
||||||
|
|
||||||
1. Verify structured logging includes `job_id`, `document_id`, `source_id`, and `person_id`.
|
|
||||||
2. Tune PostgreSQL connection pool limits and `asyncio` concurrency thresholds for production infrastructure.
|
|
||||||
3. Update operational documentation:
|
|
||||||
* Review and update `docs/schema_v2.md` as needed.
|
|
||||||
* Create `docs/runbook_v2.md` detailing PostgreSQL maintenance, JSONB index management, and worker queue monitoring.
|
|
||||||
* Create `docs/release_checklist_v2.md` for launch sign-off.
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
### Deliverables
|
|
||||||
|
|
||||||
* Updated project documentation and operational runbooks.
|
|
||||||
* V2 release sign-off checklist.
|
|
||||||
|
|
||||||
### Exit Criteria
|
|
||||||
|
|
||||||
* All documentation reflects V2 architecture; launch checklist is fully verified.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Requirement Traceability Focus
|
|
||||||
|
|
||||||
Maintain evidence against these V2 requirement groups:
|
|
||||||
|
|
||||||
* **Batch & Multi-Image Pipeline:** Folder ingestion, page ordering, async worker execution.
|
|
||||||
* **Database & Persistence:** PostgreSQL, native UUIDs, JSONB execution storage, `asyncpg` pooling.
|
|
||||||
* **Validation & Schemas:** Pydantic V2 models for DB rows, API requests, and AI vision responses.
|
|
||||||
* **Attribution & Metadata:** Multi-author and multi-recipient tagging, biographical entity management.
|
|
||||||
* **Error Recovery:** Partial success states, page-level status flags, isolated retry execution.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Scope Discipline Rule (V2 Focus)
|
|
||||||
|
|
||||||
* Only tasks required for V2 scope (PostgreSQL, Pydantic V2, folder/async processing, multi-person roles) enter this plan.
|
|
||||||
* V3 candidate features (such as side-by-side multi-provider model output comparison) remain strictly in the future backlog.
|
|
||||||
* Any schema adjustments during implementation require immediate updates to `DDL_v2.sql`, Pydantic models, and `schema_v2.md`.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Technology References
|
|
||||||
|
|
||||||
- [FastAPI documentation](https://fastapi.tiangolo.com/)
|
|
||||||
- [NiceGUI documentation](https://nicegui.io/documentation)
|
|
||||||
- [PostgreSQL documentation](https://www.postgresql.org/docs/)
|
|
||||||
- [Python asyncio](https://docs.python.org/3/library/asyncio.html#module-asyncio)
|
|
||||||
- [Pydantic Validation](https://pydantic.dev/docs/validation/latest/get-started/)
|
|
||||||
- [Pydantic AI](https://pydantic.dev/docs/ai/overview/)
|
|
||||||
|
|
||||||
## Related Local References
|
|
||||||
|
|
||||||
- [System Overview](index_v2.md)
|
|
||||||
- [System Design Intent](intent.md)
|
|
||||||
- [Transcription Methodology](transcription_methodology.md)
|
|
||||||
- [System Architecture](architecture_v2.md)
|
|
||||||
- [System Requirements](requirements_v2.md)
|
|
||||||
- [Data model](schema_v2.md)
|
|
||||||
- [Error Handling Policy](error_handling_v2.md)
|
|
||||||
- Implementation Plan (this document)
|
|
||||||
|
|
||||||
|
- A fresh database is created directly from the V2 SQLModel metadata.
|
||||||
|
- All async CRUD methods pass against the V2 relationships and fields.
|
||||||
|
- No code references `Revision`, `Source.job_id`, removed `Job` transcription fields, or compatibility aliases.
|
||||||
|
- The focused tests and full test suite pass.
|
||||||
|
|
||||||
|
## Out of Scope
|
||||||
|
|
||||||
|
- Database migrations or preservation of V1 data
|
||||||
|
- Legacy compatibility code
|
||||||
|
- Database engine or session-layer rewrites
|
||||||
|
- UI redesign, batch orchestration, worker concurrency, deployment, and operational runbooks
|
||||||
|
|||||||
+1
-1
@@ -19,7 +19,7 @@ Read [architecture_v2.md](architecture_v2.md) first for technical overview and s
|
|||||||
## Technical Stack
|
## Technical Stack
|
||||||
|
|
||||||
* **Application Web Framework:** FastAPI + NiceGUI
|
* **Application Web Framework:** FastAPI + NiceGUI
|
||||||
* **Persistence Engine:** PostgreSQL 13+
|
* **Persistence Engine:** PostgreSQL 18+
|
||||||
* **Data Validation & Schemas:** Pydantic V2
|
* **Data Validation & Schemas:** Pydantic V2
|
||||||
* **Concurrency & Workers:** Python `asyncio` worker pool with `asyncio.Semaphore`
|
* **Concurrency & Workers:** Python `asyncio` worker pool with `asyncio.Semaphore`
|
||||||
* **Vision Providers:** OpenAI (GPT-4o) and Anthropic (Claude 3.5 Sonnet) via native SDKs
|
* **Vision Providers:** OpenAI (GPT-4o) and Anthropic (Claude 3.5 Sonnet) via native SDKs
|
||||||
|
|||||||
@@ -1,183 +0,0 @@
|
|||||||
# Version 2 Plan
|
|
||||||
|
|
||||||
## Purpose
|
|
||||||
|
|
||||||
Version 2 updates the existing SQLModel domain schema to support multi-page documents, page-level transcription results, richer document metadata, and author/recipient attribution.
|
|
||||||
|
|
||||||
PostgreSQL support is already present in the database runtime. V2 does not require a database-layer rewrite or a general SQLite-to-PostgreSQL migration system. PostgreSQL adoption consists primarily of selecting the existing PostgreSQL settings, provisioning the database, creating the V2 schema, and verifying the application against it.
|
|
||||||
|
|
||||||
The main implementation effort is the schema update and the application changes that depend on it.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Current State
|
|
||||||
|
|
||||||
- The application uses Python 3.12, Pydantic v2, SQLModel, and async SQLAlchemy sessions.
|
|
||||||
- The database engine already supports both SQLite and PostgreSQL through `SqliteSettings` and `PostgresSettings`.
|
|
||||||
- The PostgreSQL async driver is installed and the engine already builds `postgresql+asyncpg` connections.
|
|
||||||
- Schema bootstrap currently uses `SQLModel.metadata.create_all()`.
|
|
||||||
- SQLite remains the default local configuration and the current Compose configuration still selects SQLite.
|
|
||||||
- The current V1 domain contains `Document`, `Source`, `Job`, and `Revision` tables.
|
|
||||||
- The V2 target is defined in [V2 DB Schema](V2%20DB%20Schema.md) and [V2 PostgreSQL DDL Specification](V2%20PostgreSQL%20DDL%20Specification.md).
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## V2 Outcomes
|
|
||||||
|
|
||||||
1. **V2 schema implemented**
|
|
||||||
- SQLModel models, relationships, enums, constraints, and indexes match the approved V2 schema.
|
|
||||||
1. **Page-level batch processing supported**
|
|
||||||
- A job can process multiple sources and retain an independent result for each source through `JobSource`.
|
|
||||||
1. **Document metadata expanded**
|
|
||||||
- Documents support ordered pages, descriptive metadata, and multiple authors and recipients.
|
|
||||||
1. **Raw provider data retained**
|
|
||||||
- Complete provider payloads are stored in PostgreSQL `JSONB` without flattening or discarding fields.
|
|
||||||
- Stored documents remain suitable for a future MongoDB import if one is ever needed.
|
|
||||||
1. **PostgreSQL enabled through configuration**
|
|
||||||
- The application starts against a provisioned PostgreSQL database using the existing runtime path.
|
|
||||||
1. **Existing workflows remain reliable**
|
|
||||||
- Upload, worker execution, status inspection, and transcription revision work with the new schema.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Non-Goals
|
|
||||||
|
|
||||||
- Rewriting the database engine or session layer
|
|
||||||
- Building a general-purpose SQLite-to-PostgreSQL migration utility
|
|
||||||
- Rehearsing a production database cutover when no production dataset requires preservation
|
|
||||||
- Running or integrating MongoDB in V2
|
|
||||||
- Building MongoDB projections, synchronization, or fallback behavior
|
|
||||||
- Replacing Python, Pydantic, SQLModel, or SQLAlchemy
|
|
||||||
- Supporting more than one active human revision per source
|
|
||||||
|
|
||||||
If an existing SQLite dataset must be retained, define a small one-time import task separately. It is not part of the default V2 implementation path.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Scope
|
|
||||||
|
|
||||||
### A) SQLModel schema update (primary)
|
|
||||||
|
|
||||||
- Add `Person`, `DocumentPerson`, and `JobSource` models.
|
|
||||||
- Expand `Document` with type, date, location, notes, archive identifier, and timestamps.
|
|
||||||
- Update `Source` with page ordering, active raw transcription, revised text, and revision timestamp.
|
|
||||||
- Update `Job` for batch execution and the `partial_success` terminal state.
|
|
||||||
- Replace the standalone `Revision` table with revision fields on `Source`.
|
|
||||||
- Remove the direct `Source.job_id` relationship; connect sources to jobs through `JobSource`.
|
|
||||||
- Add role, job status, and job-source status enums.
|
|
||||||
- Add required uniqueness constraints, foreign-key delete behavior, lookup indexes, and PostgreSQL JSON indexes.
|
|
||||||
- Keep model definitions aligned with [V2 Python Pydantic Models](V2%20Python%20Pydantic%20Models.md).
|
|
||||||
|
|
||||||
### B) Raw document storage
|
|
||||||
|
|
||||||
- Store structured AI metadata in `job_source.ai_metadata` as `JSONB`.
|
|
||||||
- Store the complete raw provider response in `job_source.raw_api_response` as `JSONB`.
|
|
||||||
- Preserve the original document structure, field names, nested values, and unknown fields in the raw response.
|
|
||||||
- Keep validation of extracted application fields separate from retention of the raw response.
|
|
||||||
- Serialize UUIDs and datetimes using portable string representations.
|
|
||||||
- Use MongoDB Extended JSON representations only if a provider value cannot be represented faithfully in standard JSON.
|
|
||||||
- Do not store opaque BSON bytes in PostgreSQL unless a future payload contains BSON-only values that cannot be preserved in `JSONB`.
|
|
||||||
|
|
||||||
### C) Schema creation and verification
|
|
||||||
|
|
||||||
- Use a fresh V2 database during development unless preservation of existing data becomes a requirement.
|
|
||||||
- Create the schema from SQLModel metadata and verify it against the approved DDL.
|
|
||||||
- Keep SQLite available for fast unit tests where its behavior is equivalent.
|
|
||||||
- Add focused PostgreSQL integration tests for native UUIDs, JSON storage, constraints, indexes, and transactions.
|
|
||||||
- Introduce migration tooling only if V2 must update a populated deployed database in place.
|
|
||||||
|
|
||||||
### D) Service and worker alignment
|
|
||||||
|
|
||||||
- Update document, source, job, and store operations for the new relationships.
|
|
||||||
- Create one `JobSource` row per source included in a job.
|
|
||||||
- Persist page-level status, transcription, AI metadata, raw provider response, and errors on `JobSource`.
|
|
||||||
- Derive the parent job status from its page results:
|
|
||||||
- `completed` when all pages succeed
|
|
||||||
- `partial_success` when successful and failed pages are mixed
|
|
||||||
- `failed` when all pages fail or a job-level failure prevents execution
|
|
||||||
- Update `Source.raw_transcription` after a successful page result while keeping the original `JobSource.raw_transcription` immutable.
|
|
||||||
- Read `Source.revised_text` in preference to `Source.raw_transcription` when presenting active text.
|
|
||||||
|
|
||||||
### E) Document and multi-image workflows
|
|
||||||
|
|
||||||
- Require a document before associating uploaded sources.
|
|
||||||
- Support uploading multiple images into one document.
|
|
||||||
- Preserve page order through `Source.page_number`.
|
|
||||||
- Allow a job to include one or more sources from the same document.
|
|
||||||
- Update job details to show the document, each source filename, page order, page status, and page-level errors.
|
|
||||||
- Define a practical upload limit and split oversized selections into manageable batches if needed.
|
|
||||||
|
|
||||||
### F) PostgreSQL configuration
|
|
||||||
|
|
||||||
- Provision PostgreSQL for local and deployed environments.
|
|
||||||
- Configure the existing `Settings.database` field with PostgreSQL host, port, database, user, and password values.
|
|
||||||
- Update Compose and environment configuration to stop selecting SQLite.
|
|
||||||
- Decide whether schema bootstrap is enabled for local development or performed as a separate deployment step.
|
|
||||||
- Run a connectivity and schema smoke test against PostgreSQL.
|
|
||||||
- Keep uploaded files on a persistent, backup-capable path outside the application image.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Milestones
|
|
||||||
|
|
||||||
### M1 - Schema models
|
|
||||||
|
|
||||||
- Implement the V2 SQLModel models and enums.
|
|
||||||
- Implement relationships, constraints, indexes, and JSON column types.
|
|
||||||
- Update the Pydantic data contracts where model decisions change.
|
|
||||||
- Add schema-focused tests.
|
|
||||||
|
|
||||||
**Exit criteria:** SQLModel metadata represents the approved V2 schema and schema tests pass.
|
|
||||||
|
|
||||||
### M2 - Persistence and worker behavior
|
|
||||||
|
|
||||||
- Update database operations and services for the V2 entities.
|
|
||||||
- Implement page-level `JobSource` execution records.
|
|
||||||
- Preserve complete raw provider responses in `JSONB`.
|
|
||||||
- Implement aggregate job status calculation.
|
|
||||||
- Add transaction, partial-success, and failure-isolation tests.
|
|
||||||
|
|
||||||
**Exit criteria:** single-page and multi-page jobs persist correct page, raw payload, and aggregate states.
|
|
||||||
|
|
||||||
### M3 - Document and upload workflows
|
|
||||||
|
|
||||||
- Update document creation and source association flows.
|
|
||||||
- Add ordered multi-image upload.
|
|
||||||
- Update job and document detail views for page-level results.
|
|
||||||
- Add focused UI and service tests.
|
|
||||||
|
|
||||||
**Exit criteria:** a user can create a document, upload ordered pages, run a job, and inspect each result.
|
|
||||||
|
|
||||||
### M4 - PostgreSQL verification and release
|
|
||||||
|
|
||||||
- Switch local or test configuration to the existing PostgreSQL runtime path.
|
|
||||||
- Create the V2 schema in a fresh PostgreSQL database.
|
|
||||||
- Run PostgreSQL-specific schema and workflow tests.
|
|
||||||
- Document startup, backup, and recovery settings.
|
|
||||||
- Run the final regression suite.
|
|
||||||
|
|
||||||
**Exit criteria:** V2 workflows pass against PostgreSQL and release checks are complete.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Risks and Mitigations
|
|
||||||
|
|
||||||
- **Model and DDL drift** -> compare generated metadata with the approved schema and test named constraints and indexes.
|
|
||||||
- **Raw payload loss** -> retain the complete provider response separately from validated and extracted fields.
|
|
||||||
- **Cross-database differences** -> retain fast SQLite tests but verify PostgreSQL-native UUID, JSON, and index behavior in integration tests.
|
|
||||||
- **Batch state errors** -> test all-success, mixed-result, and all-failed jobs explicitly.
|
|
||||||
- **Page ordering errors** -> enforce uniqueness and ordering rules for document pages.
|
|
||||||
- **Unexpected data-preservation need** -> confirm whether existing SQLite data matters before implementation; add a one-time importer only when required.
|
|
||||||
- **Worker regressions** -> preserve terminal-state and retry reliability tests while changing persistence ownership.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Suggested First Tasks
|
|
||||||
|
|
||||||
1. Update `src/transcription/db/models.py` to represent the approved V2 schema.
|
|
||||||
2. Add schema tests for tables, columns, relationships, constraints, indexes, and enums.
|
|
||||||
3. Define and test lossless raw provider response storage in `job_source.raw_api_response`.
|
|
||||||
4. Update database operations and services to use `JobSource` and source-level revisions.
|
|
||||||
5. Add page-result aggregation tests before changing the worker workflow.
|
|
||||||
6. Update document and multi-image upload flows.
|
|
||||||
7. Select PostgreSQL in configuration and run the integration suite against a fresh V2 database.
|
|
||||||
Reference in New Issue
Block a user