generated from john/python-template
237 lines
6.1 KiB
Markdown
237 lines
6.1 KiB
Markdown
# Database Schema (Version 4)
|
|
|
|
This document defines the relational schema for the document transcription system.
|
|
|
|
## Entity Relationship Diagram
|
|
|
|
```mermaid
|
|
erDiagram
|
|
DOCUMENT_TYPE {
|
|
UUID id PK
|
|
TEXT label
|
|
TEXT normalized_label
|
|
BOOLEAN is_active
|
|
TIMESTAMPTZ created_at
|
|
TIMESTAMPTZ updated_at
|
|
}
|
|
|
|
PERSON_ROLE {
|
|
UUID id PK
|
|
TEXT code
|
|
TEXT label
|
|
BOOLEAN is_active
|
|
TIMESTAMPTZ created_at
|
|
TIMESTAMPTZ updated_at
|
|
}
|
|
|
|
PERSON {
|
|
UUID id PK
|
|
TEXT full_name
|
|
TEXT display_name
|
|
TEXT maiden_name
|
|
DATE birth_date
|
|
TEXT birth_date_raw
|
|
TEXT birth_place
|
|
DATE death_date
|
|
TEXT death_date_raw
|
|
TEXT death_place
|
|
TEXT biography
|
|
TEXT portrait_path
|
|
JSONB metadata
|
|
TIMESTAMPTZ created_at
|
|
TIMESTAMPTZ updated_at
|
|
}
|
|
|
|
DOCUMENT {
|
|
UUID id PK
|
|
UUID document_type_id FK
|
|
TEXT name
|
|
DATE document_date
|
|
TEXT document_date_raw
|
|
TEXT location_created
|
|
TEXT notes
|
|
TEXT archive_identifier
|
|
TIMESTAMPTZ created_at
|
|
TIMESTAMPTZ updated_at
|
|
}
|
|
|
|
DOCUMENT_PERSON {
|
|
UUID id PK
|
|
UUID document_id FK
|
|
UUID person_id FK
|
|
UUID role_id FK
|
|
TIMESTAMPTZ created_at
|
|
TIMESTAMPTZ updated_at
|
|
}
|
|
|
|
JOB {
|
|
UUID id PK
|
|
UUID document_id FK
|
|
VARCHAR status
|
|
INTEGER retry_count
|
|
TEXT provider
|
|
TEXT model
|
|
TEXT prompt_name
|
|
TEXT prompt_hash
|
|
TEXT system_prompt
|
|
TEXT user_prompt
|
|
FLOAT temperature
|
|
FLOAT top_p
|
|
TIMESTAMPTZ date_created
|
|
TIMESTAMPTZ date_updated
|
|
}
|
|
|
|
SOURCE {
|
|
UUID id PK
|
|
UUID document_id FK
|
|
INTEGER page_number
|
|
TEXT upload_name
|
|
TEXT filename
|
|
TEXT file_path
|
|
TEXT file_hash
|
|
BIGINT file_size_bytes
|
|
TEXT raw_transcription
|
|
TEXT revised_text
|
|
TIMESTAMPTZ date_uploaded
|
|
TIMESTAMPTZ date_revised
|
|
}
|
|
|
|
JOB_SOURCE {
|
|
UUID id PK
|
|
UUID job_id FK
|
|
UUID source_id FK
|
|
VARCHAR status
|
|
TEXT raw_transcription
|
|
JSONB ai_metadata
|
|
JSONB raw_api_response
|
|
TEXT error_detail
|
|
TIMESTAMPTZ executed_at
|
|
}
|
|
|
|
EXECUTION_ATTEMPT {
|
|
UUID id PK
|
|
UUID job_source_id FK
|
|
UUID job_id FK
|
|
UUID source_id FK
|
|
INTEGER attempt_number
|
|
VARCHAR status
|
|
JSONB request_manifest
|
|
TEXT request_manifest_sha256
|
|
INTEGER transport_status_code
|
|
BINARY transport_body
|
|
JSONB transport_safe_headers
|
|
JSONB sdk_response_snapshot
|
|
JSONB normalized_metadata
|
|
JSONB software_context
|
|
TEXT raw_transcription
|
|
TEXT failure_phase
|
|
TIMESTAMPTZ started_at
|
|
TIMESTAMPTZ finished_at
|
|
INTEGER duration_ms
|
|
}
|
|
|
|
PROCESSING_ARTIFACT {
|
|
UUID id PK
|
|
UUID source_id FK
|
|
UUID execution_attempt_id FK
|
|
TEXT artifact_type
|
|
TEXT media_type
|
|
TEXT schema_name
|
|
TEXT schema_version
|
|
TEXT producer
|
|
TEXT producer_version
|
|
JSONB inline_payload
|
|
TEXT external_reference
|
|
TEXT payload_sha256
|
|
BIGINT byte_size
|
|
JSONB coordinate_metadata
|
|
TIMESTAMPTZ created_at
|
|
}
|
|
|
|
DOCUMENT_TYPE ||--o{ DOCUMENT : classifies
|
|
DOCUMENT ||--o{ DOCUMENT_PERSON : has_people
|
|
PERSON ||--o{ DOCUMENT_PERSON : appears_in
|
|
PERSON_ROLE ||--o{ DOCUMENT_PERSON : labels
|
|
DOCUMENT ||--o{ JOB : has_jobs
|
|
DOCUMENT ||--o{ SOURCE : contains_pages
|
|
JOB ||--o{ JOB_SOURCE : executes
|
|
SOURCE ||--o{ JOB_SOURCE : processed_in
|
|
JOB_SOURCE ||--o{ EXECUTION_ATTEMPT : projects
|
|
SOURCE ||--o{ PROCESSING_ARTIFACT : derives
|
|
EXECUTION_ATTEMPT ||--o{ PROCESSING_ARTIFACT : produces
|
|
```
|
|
|
|
## Domain Invariants and Provenance Rules
|
|
|
|
### Page-Level Execution and AI Outputs
|
|
|
|
- Every single page execution by an AI model produces a dedicated `JOB_SOURCE` record.
|
|
- Every `JOB` stores the frozen prompt identifier, prompt text, and hyperparameters used at submission time.
|
|
- `JOB_SOURCE.raw_api_response` is a compatibility projection containing an SDK-serialized OpenRouter response
|
|
snapshot. It is neither the exact HTTP body nor the native upstream-provider response.
|
|
- Every new provider call creates an immutable `EXECUTION_ATTEMPT` containing the frozen request manifest,
|
|
exact OpenRouter-boundary response bytes when received, safe transport metadata, SDK snapshot, normalized
|
|
metadata, timing, and outcome.
|
|
- `EXECUTION_ATTEMPT(job_id, source_id, attempt_number)` is unique; retries increment the persisted attempt number.
|
|
- Historical `JOB_SOURCE` rows without an `EXECUTION_ATTEMPT` remain SDK snapshots and are explicitly labeled as
|
|
lacking transport evidence.
|
|
- `SOURCE.raw_transcription` caches the latest successful machine output for that page.
|
|
|
|
### Generic Processing Artifacts
|
|
|
|
- `PROCESSING_ARTIFACT` stores provider-neutral versioned derived outputs.
|
|
- Exactly one of `inline_payload` and `external_reference` is populated.
|
|
- Externally stored artifacts use application-managed relative references and are verified by SHA-256 and byte size.
|
|
- Coordinate metadata declares units, origin, dimensions, and transformations when geometry is present.
|
|
|
|
### Image Storage and Integrity
|
|
|
|
- Binary images are stored on disk; `SOURCE.file_path` stores the persisted path.
|
|
- `SOURCE.file_hash` stores a SHA-256 digest.
|
|
- `SOURCE.file_size_bytes` stores the original file size.
|
|
|
|
### Page Ordering and Revisions
|
|
|
|
- `SOURCE.page_number` dictates page ordering within a document.
|
|
- `SOURCE.raw_transcription` remains immutable machine output.
|
|
- `SOURCE.revised_text` stores human edits and is the preferred display value when present.
|
|
|
|
### Document-Person Role Governance
|
|
|
|
- Documents support zero, one, or many people per relationship role.
|
|
- Relationship roles are defined by `PERSON_ROLE` rather than hardcoded columns.
|
|
- `DOCUMENT_PERSON` must be unique for `(document_id, person_id, role_id)`.
|
|
- Relationship writes must be deterministic and use explicit add/remove link intent.
|
|
|
|
### Document Type Governance
|
|
|
|
- Every document type is defined by `DOCUMENT_TYPE`.
|
|
- `DOCUMENT_TYPE.id` is the sole machine identity.
|
|
- `DOCUMENT_TYPE.label` is mutable display text and is unique after trimming and case normalization.
|
|
- `DOCUMENT_TYPE.normalized_label` stores the normalized uniqueness key.
|
|
- Inactive types remain valid for historical rows but should be excluded from default selection UIs.
|
|
|
|
## Constraint Summary
|
|
|
|
- `DOCUMENT_TYPE.normalized_label` is unique.
|
|
- `PERSON_ROLE.code` is unique.
|
|
- `DOCUMENT_PERSON(document_id, person_id, role_id)` is unique.
|
|
|
|
## Indexing Guidance
|
|
|
|
- `document(document_type_id)`
|
|
- `document_person(document_id)`
|
|
- `document_person(person_id)`
|
|
- `document_person(role_id)`
|
|
- `source(document_id, page_number)`
|
|
- `job(document_id, status)`
|
|
- `job_source(job_id)`
|
|
- `job_source(source_id)`
|
|
|
|
## Related Local References
|
|
|
|
- [System Overview](index_v4.md)
|
|
- [System Architecture](architecture_v4.md)
|
|
- [System Requirements](requirements_v4.md)
|
|
- [Error Handling Policy](error_handling_v4.md)
|