generated from john/python-template
This commit is contained in:
@@ -123,6 +123,10 @@ Evaluation should:
|
||||
5. Preserve the exact model, endpoint or route, parameters, prompt, source digest, and scoring method for every comparison.
|
||||
6. Treat model rankings as corpus- and version-specific, not permanent declarations of a universal “best” model.
|
||||
|
||||
The deterministic scorer for these comparisons lives in `src/transcription/benchmarking.py`; it is
|
||||
retained as evaluation-policy infrastructure even though application runtime paths do not call it
|
||||
directly.
|
||||
|
||||
Benchmark material containing family records remains private application data unless explicitly approved for publication.
|
||||
|
||||
## 6. Ownership and Change Policy
|
||||
|
||||
Reference in New Issue
Block a user