V4.6 Phase 4: service layer consolidation

Removes the duplicated registry CRUD, the hand-written not-found raises, and
the three divergent media writers. Behavior is preserved: every existing
Document Type and Person Role test passes unchanged, which is the primary
proof for MED-11.

[MED-11] Generic registry service
- New services/registry.py owns RegistryService[ModelT]: list, list with
  counts, create with IntegrityError -> conflict mapping, read, update,
  delete with built-in and referenced guards, is_referenced, and label
  normalization/casefold keying.
- DocumentTypeRegistry and PersonRoleRegistry declare only the model, error
  class, noun, short noun, retainer phrase, and reference columns.
- DocumentService and PeopleService keep their public method names and
  delegate. Every user-facing message, error category, and suggestion string
  is reproduced verbatim; only the noun is templated.
- Deleted _normalize_registry_label, _document_type_label_key,
  _normalize_role_label, _person_role_label_key,
  _document_type_is_referenced, and _person_role_is_referenced.

[MED-12] Shared not-found lookup
- ServiceBase._get_or_raise(model, id, *, session, error, noun, suggestion,
  options) loads by primary key or raises the caller's error type.
- documents.py: local _get_document_or_raise deleted; replaced by _read_document
  and adopted at read_document, delete_document, and set_document_type, which
  previously bypassed the helper and hand-wrote the raise.
- sources.py: 8 identical Source raises and 1 Job raise collapsed into
  _read_source / _get_or_raise.
- jobs.py and people.py already funneled through local _not_found builders and
  were left alone.

[MED-13][MED-01] Single media writer
- New services/media_storage.py owns validate -> name -> mkdir -> write ->
  wrap OSError. The write runs in asyncio.to_thread, so uploads no longer block
  the event loop.
- store_source_file, store_person_portrait, and store_homepage_image now share
  it and are async. Callers in store.py, people_page.py, and home_page.py await
  them. mkdir failures are now also translated to a domain error instead of
  escaping as a raw OSError.
- homepage_store gains HomepageStorageError so its write reports like the others.

[MED-14, partial] Service independence
- New services/source_media.py owns SOURCE_MIME_TYPES, SOURCE_EXTENSIONS,
  lookup_source_mime_type, and supported_source_formats.
- documents.py no longer imports services/sources.py. Its print projection uses
  the non-raising lookup and raises DocumentError, so DocumentService no longer
  emits a TranscriptionError.
- api/v4_print.py imports the mapping from the policy module.
- store.py and workflows.py still import sources.py; both are orchestration
  modules, which services.instructions.md:75-77 explicitly permits.
- Splitting SourceService itself remains deferred to V4.7.

[LOW-08] Query shape
- list_sources_detail filters job_id with a JOIN on JobSource instead of
  loading every Source and filtering in Python.
- read_source_navigation replaces the full ordered-id scan and .index() with
  two row-value comparisons bounded by LIMIT 1.
- list_processing_artifacts gains the limit parameter its summary sibling
  already had.
- build_evidence_export runs artifact integrity hashing and file reads through
  asyncio.to_thread.

Tests
- tests/test_service_boundaries.py: AST guard asserting no service module
  imports a sibling service module, plus a guard that the scan is non-empty.
- tests/services/test_transcription_service.py: asserts the job_id filter emits
  a JOIN, and that navigation emits exactly two LIMIT queries.
- tests/services/test_store.py: the two storage tests are now async.

Verified: 276 passed, 4 skipped; ruff check clean.
This commit is contained in:
zoltan57
2026-08-17 16:46:15 -05:00
parent 7b9715b3f1
commit 97b3d0fd62
15 changed files with 814 additions and 462 deletions
+6 -4
View File
@@ -114,11 +114,12 @@ async def test_create_document_job_stores_source_under_document_id_directory(asy
assert created_job.prompt_name == "transcribe_document.md"
def test_store_person_portrait_stores_file_under_person_id_directory(tmp_path):
@pytest.mark.asyncio
async def test_store_person_portrait_stores_file_under_person_id_directory(tmp_path):
settings = Settings(openrouter_api_key="test-key", upload_dir=tmp_path)
person_id = uuid4()
stored_path = store_person_portrait(
stored_path = await store_person_portrait(
person_id=person_id,
filename="portrait.png",
file_bytes=b"portrait-bytes",
@@ -144,8 +145,9 @@ def test_source_mime_type_uses_canonical_source_policy(filename, expected_mime_t
assert source_mime_type(filename) == expected_mime_type
def test_source_storage_rejects_unsupported_format(tmp_path):
@pytest.mark.asyncio
async def test_source_storage_rejects_unsupported_format(tmp_path):
settings = Settings(openrouter_api_key="test-key", upload_dir=tmp_path)
with pytest.raises(SourceStorageError):
store_source_file(filename="page.txt", file_bytes=b"text", settings=settings)
await store_source_file(filename="page.txt", file_bytes=b"text", settings=settings)
@@ -3,6 +3,7 @@
from uuid import uuid4
import pytest
from sqlalchemy import event
from transcription.config import Settings
from transcription.db.models import Document
@@ -232,3 +233,112 @@ class TestSourceServiceRevisionUpsert:
with pytest.raises(SourceDeleteBlockedError):
await transcriptions.delete_unlinked_source(source_id=source.id)
@pytest.mark.integration
class TestSourceServiceQueryShape:
"""LOW-08: reads must filter and bound in SQL, not in Python."""
@pytest.mark.asyncio
async def test_list_sources_detail_filters_job_id_with_a_join(self, default_session_factory):
documents = DocumentService(session_factory=default_session_factory)
jobs = JobService(session_factory=default_session_factory)
transcriptions = SourceService(session_factory=default_session_factory)
document = Document(id=uuid4(), name="join-filter")
await documents.create_document(document=document)
job = Job(document_id=document.id, status=JobStatus.QUEUED)
other_job = Job(document_id=document.id, status=JobStatus.QUEUED)
await jobs.create_job(job=job)
await jobs.create_job(job=other_job)
linked = Source(
document_id=document.id,
page_number=1,
upload_name="linked.jpg",
filename="linked.jpg",
file_path="uploads/linked.jpg",
file_hash="7" * 64,
file_size_bytes=1,
)
unlinked = Source(
document_id=document.id,
page_number=2,
upload_name="unlinked.jpg",
filename="unlinked.jpg",
file_path="uploads/unlinked.jpg",
file_hash="8" * 64,
file_size_bytes=1,
)
async with transcriptions._session_scope() as session:
session.add_all((linked, unlinked))
await session.flush()
session.add(JobSource(job_id=job.id, source_id=linked.id, status=JobSourceStatus.PENDING))
session.add(JobSource(job_id=other_job.id, source_id=unlinked.id, status=JobSourceStatus.PENDING))
await session.commit()
statements: list[str] = []
async with transcriptions._session_scope() as session:
bind = session.get_bind()
def capture(_conn, _cursor, statement, *_rest):
statements.append(statement)
event.listen(bind, "before_cursor_execute", capture)
try:
sources = await transcriptions.list_sources_detail(job_id=job.id, session=session)
finally:
event.remove(bind, "before_cursor_execute", capture)
assert [source.id for source in sources] == [linked.id]
primary = next(item for item in statements if item.lstrip().upper().startswith("SELECT"))
assert "JOIN" in primary.upper()
assert "JOBSOURCE" in primary.upper().replace("_", "")
@pytest.mark.asyncio
async def test_read_source_navigation_does_not_scan_every_sibling(self, default_session_factory):
documents = DocumentService(session_factory=default_session_factory)
transcriptions = SourceService(session_factory=default_session_factory)
document = Document(id=uuid4(), name="navigation-bounds")
await documents.create_document(document=document)
pages = [
Source(
document_id=document.id,
page_number=page_number,
upload_name=f"page-{page_number}.jpg",
filename=f"page-{page_number}.jpg",
file_path=f"uploads/page-{page_number}.jpg",
file_hash=str(page_number) * 64,
file_size_bytes=1,
)
for page_number in range(1, 5)
]
async with transcriptions._session_scope() as session:
session.add_all(pages)
await session.commit()
for page in pages:
await session.refresh(page)
statements: list[str] = []
async with transcriptions._session_scope() as session:
bind = session.get_bind()
def capture(_conn, _cursor, statement, *_rest):
statements.append(statement)
event.listen(bind, "before_cursor_execute", capture)
try:
navigation = await transcriptions.read_source_navigation(pages[1].id, session=session)
finally:
event.remove(bind, "before_cursor_execute", capture)
assert navigation.previous_id == pages[0].id
assert navigation.next_id == pages[2].id
adjacency = [item for item in statements if "LIMIT" in item.upper()]
assert len(adjacency) == 2, statements