Sentinel Policy M5: structured semantic intelligence (v1.0.678)
What shipped
M5 (EXT-001..008) — the first place an LLM enters the pipeline, and only ever on M4's already-bounded, already-deterministic CandidateDiff output, never a raw document (EXT-001, the product's core invariant: "LLMs may interpret evidence. They may never become the evidence.").
- **
llm.py**: thin Anthropic client wrapper, mirroringfetch.py's - **
extraction_schema.py**:ExtractionOutput(Pydantic, EXT-002), - **
prompts/change_extraction/v1.md**: version-controlled prompt file - **
extraction.py**: orchestrates one bounded extraction per - New tables:
ExtractionRun(EXT-006: model/prompt/schema version, - Manual-trigger only (your explicit decision): new
- Idempotent: a repeat trigger on the same
CandidateDiffreturns the
existing tenacity retry pattern exactly (transient vs permanent error classes, exponential backoff+jitter). Model forced to respond via a single tool call (record_policy_changes) matching the output schema, not free-text.
change_type restricted to a Literal matching the spec's exact 22-value CHANGE_TAXONOMY (now in db/models.py, TAX-001) — the model cannot invent a taxonomy value; Pydantic rejects anything else.
(MAINT-005), not an inline string. Confirmed it actually ships inside the built wheel/Docker image (the Dockerfile only COPYs src/, so pyproject.toml needed an explicit [tool.setuptools.package-data] entry — verified by building a real wheel and inspecting its contents before trusting it).
CandidateDiff. Critically, **the model's evidence_quote is never trusted as a real offset** — the app searches for it verbatim in the actual stored PolicySection.normalized_text and computes the real offsets itself (EXT-004/005/AGENT-005); a candidate whose quote doesn't verify is simply dropped, not persisted. One same-model retry on a schema-invalid response (EXT-007). EXT-008 (model disagreement on critical fields) is satisfied without a redundant second paid call: a single response asserting more than one distinct effective date across its own reported changes is internally inconsistent on a critical field (DAT-001) and routes to review.
attempt count, retained per run), ExtractedFact (only ever persisted post-verification), ExtractionReviewCase (AGENT-017's "inspectable record, not a fabricated success" pattern, same shape as M3's PolicyIdentityReviewCase but for this failure domain — SCHEMA_INVALID_AFTER_RETRY / NO_VERIFIABLE_EVIDENCE / CRITICAL_FIELD_DISAGREEMENT).
POST /v1/policies/{policy_id}/versions/{version_id}/extract. Nothing runs automatically like M4's free diff engine — extraction only happens when explicitly requested, and the endpoint 503s cleanly (not a crash) if SENTINEL_POLICY_ANTHROPIC_API_KEY isn't configured, mirroring require_admin_token's existing fail-closed pattern.
existing ExtractionRun without re-calling the model — a re-run can never silently re-spend money.
Deliberately deferred, honestly
- EXT-007's secondary-provider fallback got a real config seam
- Confidence scoring (CNF-family), the formal
evidence_spanentity,
(SENTINEL_POLICY_EXTRACTION_FALLBACK_MODEL) but not a second real provider integration — there's no second provider account to fail over to today. Same honesty as M2 declining to implement the AMA license flow: the extension point is real, the second integration isn't faked.
and publication decisions are M6. This milestone produces candidate facts; it doesn't decide what gets published.
Verification
PYTHONPATH=policy/src:policy/tests python -m pytest policy/tests -q—- Migration verified against real local PostgreSQL.
- Full pipeline verified end-to-end against real Postgres using the real
- Confirmed the extraction endpoint 503s cleanly with no API key
80 passed (up from 66). All extraction orchestration tests mock the LLM client at the sentinel_policy.extraction.call_extraction_model boundary — no real API calls in CI, matching this project's existing fixture-server discipline for a different I/O boundary. Covers: successful extraction persisting verified facts; an unverifiable quote being dropped without failing silently; schema-invalid-then-valid retry succeeding; schema-invalid-after-retry and LLM-error-after-retry both routing to review; internally-disagreeing effective dates routing to review; idempotent re-trigger not re-invoking the model.
NCD 110 diff already proven live in M4 (a genuine 2003 Medicare coverage-expansion change) — LLM call mocked (no API key provisioned yet), everything else real: real CandidateDiff, real PolicySection text, real offset verification, real persistence, confirmed correct through the actual GET /v1/policies/{id} endpoint.
configured, via a real TestClient call against the real FastAPI app.
Not yet run against a real model
Per your explicit decision, this milestone does not call the real Anthropic API anywhere (tests, this deploy, or production) — a SENTINEL_POLICY_ANTHROPIC_API_KEY has not been provisioned. The real exit-criterion check (a real extraction against the real NCD 110 diff, confirming a sensible change_type and code/date extraction) is the first thing to run once you provide one.