Sentinel Policy M7: golden evaluation harness (v1.0.685)
What this milestone adds
A curated set of real, hand-labeled policy diffs ("golden cases") plus a harness that runs the real extraction pipeline against them and scores precision/recall/F1. Two run modes:
- replay (free, deterministic, CI-safe): patches
- live (real spend, manual only): calls the real configured LLM.
call_extraction_model with each case's recorded canned response (a real output, captured once, hand-verified against the real text). A regression test of the pipeline and scorer, not an evaluation of the live model.
Never automatic, never CI -- same discipline as M5's extraction endpoint.
New: eval/golden_cases.py, eval/scorer.py, eval/runner.py, eval/cli.py (python -m sentinel_policy.eval.cli run --mode live|replay), EvaluationRun/EvaluationCaseResult tables (migration 0006), GET /v1/eval-runs.
Golden cases run under a dedicated cms-medicare-ncd-eval Source, kept completely separate from real production Policy/PolicyVersion data so re-running the suite never collides with or pollutes real synced NCDs.
Two real bugs found while sourcing real golden data
Building the flagship case required a real, chronologically-adjacent historical diff. Sourcing it surfaced two real, previously-undiscovered production bugs -- fixed as part of this milestone, not filed separately, same "found while planning, real, in-scope" precedent M6 set for its own evidence-verification bugfix.
1. Version-chain linking (policyintel.py)
resolve_policy_version() linked a new version's previous_version_id to "whichever version currently has the highest ordinal in the DB" -- correct only if versions are always ingested in increasing ordinal order. NCD 110 in production had its current (2023) version synced first, then two older historical versions (1999, 2003) backfilled in afterward -- both got wired to the 2023 version instead of chaining 1999 -> 2003 -> 2023. Confirmed live: the "first real LLM extraction" milestone check from M6 was actually diffing 2003-vs-2023 text, not 2003-vs-1999.
Fixed the lookup to pick the nearest lower ordinal. CandidateDiff being unique on the (previous_version_id, version_id) pair (not version_id alone) meant the fix could be purely additive: a one-time repair script (scripts/repair_ncd110_version_chain.py) corrected the three affected pointers and computed two new, correctly-paired CandidateDiffs, without touching or deleting the old, incorrectly- paired CandidateDiff/ExtractionRun/ChangeEvent rows (immutable per LED-001; correction/supersession is M8's deferred workflow, per LED-002).
This exposed a second bug: PolicyVersion.candidate_diff and the /extract endpoint's lookup both assumed one CandidateDiff per version_id, which is no longer true once a version's previous_version_id can be corrected after the fact. Both now disambiguate by the version's current previous_version_id.
2. EXT-008's cross-candidate date-agreement check (extraction.py)
M6 shipped a heuristic: a single response asserting more than one distinct effective date across its reported changes was treated as internal disagreement (DAT-001) and failed the run. Tried and reverted twice against the real, now-correctly-paired 2003-vs-2023 NCD 110 diff:
- Per-response (as shipped in M6): failed on any large diff touching
- Per-section (first fix this milestone): still failed -- a single real
multiple real dates across different sections -- genuinely common, not disagreement.
section can legitimately contain multiple real historical sub-clauses with different real dates (found live: indications_limitations genuinely reports both 2018-02-01 and 2018-02-15 for two different real amendments).
Removed entirely. Each ExtractedChangeCandidate has exactly one effective_date_candidate field -- it structurally cannot self- contradict -- so a cross-candidate agreement check was never a meaningful signal against real, richly-multi-dated CMS data. EXT-008 is not currently implemented; noted as such in extraction.py's docstring.
Golden cases (v1)
- **
ncd110-1999-vs-2003** (flagship, real, non-exhaustive): the real - **
ncd110-2003-vs-2023** (real, non-exhaustive): the same 2003 text - **
ncd108-administrative-republish-true-negative** (synthetic,
October 1, 2003 Medicare coverage expansion for primary-prevention ICD implantation. Canned response captured live, all 8 real facts hand-verified against the real CMS text.
diffed forward against the current version -- a large, real, two-decade span, labeled with only a few high-confidence facts.
exhaustive, CI-only): reuses M4's existing administrative-republish fixture. No real change exists, so any reported fact is a false positive.
Verification
PYTHONPATH=policy/src:policy/tests python -m pytest policy/tests -q- Real-Postgres repair confirmed live: NCD 110's version chain and two
- Real
live-mode extraction confirmed for both real NCD 110 cases
-- 111 tests passing (up from 94 pre-M7).
new correctly-paired CandidateDiffs.
against production, post both bugfixes -- the true end-to-end exit-criterion check for this milestone.