--- description: Follow these guidelines when editing the services applyTo: 'src/transcription/services/*.py' --- # Services ## Structure - Project core data models are defined in [models](../../src/transcription/db/models.py) - One service class per **aggregate**, not per table. An aggregate is a root model plus the models that have no independent lifecycle of their own. `DocumentType` has no meaning without `Document`, so it belongs to `DocumentService`; it does not get its own service. Splitting per table produces services that must reach across each other for every real operation, which is what line 13 forbids. - Only services interact with the database, and only through async methods. - **A service module must not import another service module.** This is enforced by [test_service_boundaries](../../tests/test_service_boundaries.py). Shared types go in a neutral module that defines no service class (see [errors](../../src/transcription/services/errors.py)). - Not every module in this package is a service. Helper modules that define no `*Service` class (`base`, `errors`, `normalization`, `prompts`, `quality`, `media_storage`, `source_media`) are free-function modules and are exempt from the service rules below. ## Model Ownership Every model has exactly one owning service. The owner defines that model's invariants and is the only service that may **create or delete** its rows. | Model | Owner | | --- | --- | | `Document`, `DocumentType` | `DocumentService` | | `Source`, `JobSource` | `SourceService` | | `Job` | `JobService` | | `Person`, `PersonRole`, `DocumentPerson` | `PeopleService` | | `ExecutionAttempt` | `EvidenceService` | ### Junction tables A junction table is owned by the service that **creates and deletes its rows** — its lifecycle owner. The service on the other side may read through the junction (via `selectinload`) but must not create rows in it. - `document_person` -> `PeopleService`. Every write is there; `DocumentService` only eager-loads through it. - `job_source` -> `SourceService`, which creates the row, records each page's outcome, and deletes it. Two consequences follow, and both are deliberate: - **Cascade deletion is not a violation.** A service deleting the aggregate root it owns may delete junction rows referencing that root, because they cannot outlive it (`JobService.delete_job_with_guardrails`). - **Ownership governs creation and deletion, not every state transition.** `job_source` is both a link and the transcription work queue. `JobService.cancel_job` and `resubmit_failed_sources` transition `job_source.status` across a whole job, because that transition is a Job lifecycle event, not a per-page outcome. They create and delete nothing. `EvidenceService.promote_machine_attempt` writes two fields on `Source` (`preferred_execution_attempt_id`, `raw_transcription`). This is allowed on the same principle: selecting which attempt a Source presents is an evidence decision that happens to land on `Source`. It is scoped to those two projection fields. If a new operation cannot be expressed within one owner, it belongs in an orchestration module, not in a cross-service import. ## Error Handling - Errors used by a single service are defined at the top of that module and inherit from `AppError`. - Errors shared by more than one service go in [errors](../../src/transcription/services/errors.py), which defines no service class and is therefore importable by any of them. - Use a context manager for large `try/except` blocks, like `handle_transcription_errors` in [sources](../../src/transcription/services/sources.py). ## Checklist - [ ] Uses `ServiceBase` for common logic - [ ] Session kwarg for `AsyncSession` to pass a session object into each method - [ ] Services use `self._session_scope` in their methods to pass the session through - Multiple operations on the same object(s) require sharing a session between all the methods used - [ ] Every model the module touches is either owned by it or reached read-only ## CRUD Methods - Name format `_`, for example `create_document` or `update_job`. - Where a service exposes create/read/update/delete for its root model, define them at the top of the class in that order, before derived reads and workflow helpers. - Not every aggregate needs all four. `ExecutionAttempt` is append-only evidence written by `workflows.py`, so `EvidenceService` deliberately exposes reads and no create or delete. Do not add unused CRUD methods to satisfy symmetry. - `RegistryService` is generic across small lookup models and uses `_entry` naming instead. ## Transaction Finalization When a service method accepts an optional `session` kwarg, write methods must use `self._finalize` to finalize the transaction properly according to whether or not they are sharing a session. - If `session` is `None`: the method owns the transaction and should `commit()`. - If `session` is provided: the method must **not** commit; it should `flush()` so IDs and FK values are available to the caller's transaction. - Use `refresh()` on returned ORM objects when the caller needs DB-populated values (defaults, triggers, merged state). Recommended helper behavior: - Inputs: active session object, original `session` arg (or a boolean ownership flag), and an optional list of objects to refresh. - Logic: `commit` when service-owned session, `flush` when caller-owned session, then refresh requested objects. This keeps orchestration functions atomic: they can pass one shared session across multiple services and commit exactly once at the workflow boundary. ## Workflow Transaction Boundaries For multi-step job lifecycles (for example queued transcription jobs), orchestration functions must use explicit transaction phases. Required boundary model: - **Transaction A (claim):** transition `JobStatus.QUEUED -> JobStatus.PROCESSING` and commit immediately. - Perform provider/network work **outside** database transactions. - **Transaction B (terminal success):** write transcript content and set `JobStatus.TRANSCRIBED` in the same shared-session commit. - **Transaction B (terminal failure):** write transcript error detail and set `JobStatus.FAILED` in the same shared-session commit. - **Transaction C (retry path):** write transcript error detail, increment retry count, and set `JobStatus.QUEUED` in one shared-session commit. Atomicity rules: - Never commit transcript updates separately from the paired terminal/retry job status change. - Terminal state (`TRANSCRIBED` or `FAILED`) and transcript row changes must succeed or roll back together. - Retry persistence (`QUEUED` + retry increment + error detail) must succeed or roll back together. Separation of concerns: - Worker modules should stay lightweight and delegate lifecycle transitions to service/workflow orchestration functions. - In `workflows.py`, `process_queued_job` should own one complete attempt lifecycle: `QUEUED -> PROCESSING -> TRANSCRIBED|FAILED`. - In `workflows.py`, `advance_job` should coordinate broader status progression around attempts (for example retry scheduling from `FAILED -> QUEUED`). - Services should expose session-aware write helpers (flush on caller-owned session) so orchestration controls commit boundaries. - Backoff/sleep behavior must run outside transactional scopes. # Service Composition A service method may read across models it does not own, using eager loads from its own aggregate root. What it may not do is import another service. Operations that must **write** models owned by more than one service — uploading a picture, for example — are composed in an orchestration module ([store](../../src/transcription/services/store.py), [workflows](../../src/transcription/services/workflows.py)). Orchestration modules define no service class, may import any service, and own the commit boundary.