Sentinel Policy: OpenAI provider swap (v1.0.680)
What changed
User decision (2026-08-20): swap M5/M6's LLM provider from Anthropic to OpenAI (gpt-5.1), since only an OpenAI API key was actually available to provision. A straight swap, not a multi-provider abstraction — EXT-007's extraction_fallback_model config seam stays exactly as documented before (a real, unused gap), matching the same honesty as M2's AMA license non-goal.
llm.py: rewritten against theopenaiSDK's Chat Completions API,config.py:SENTINEL_POLICY_ANTHROPIC_API_KEY→SENTINEL_POLICY_OPENAI_API_KEY,api/deps.py'srequire_extraction_configured(the 503-fail-closedpyproject.toml:anthropicdependency replaced withopenai>=1.0,<2ExtractionRun.model_provider's default changed"anthropic"→
forcing structured output via tools/tool_choice (function-calling, the OpenAI equivalent of Anthropic's forced tool_use) rather than Anthropic's Messages API. Same transient/permanent error-classification shape as before, mapped to OpenAI's own exception hierarchy (RateLimitError/APITimeoutError/APIConnectionError → transient; InternalServerError (5xx) → transient; everything else → permanent). One real shape difference handled: OpenAI returns tool-call arguments as a JSON string (Anthropic returns an already-parsed dict) — added explicit json.loads() with its own PermanentLLMError on malformed JSON.
default model claude-sonnet-5 → gpt-5.1.
gate) now checks the renamed field — reconfirmed live via TestClient that it still 503s cleanly with no key configured.
(pinned to the well-established 1.x line, not the very new 3.x major which changed its HTTP internals — chose the safer, more widely-used version).
"openai".
What did NOT need to change
Every layer above llm.py — extraction.py's orchestration, retry/ disagreement logic, evidence verification, confidence.py, ledger.py — is provider-agnostic by design (M5's own architecture kept the SDK call isolated behind call_extraction_model()), so none of it needed touching. All 94 tests, which mock at that exact boundary, passed unchanged.
Verification
PYTHONPATH=policy/src:policy/tests python -m pytest policy/tests -q—- Confirmed live via a real
TestClientcall that the extraction endpoint
94 passed, unchanged (the whole point of the boundary holding).
still 503s cleanly with SENTINEL_POLICY_OPENAI_API_KEY unset.
Verified live (2026-08-20, v1.0.681)
SENTINEL_POLICY_OPENAI_API_KEY was provisioned in production. A real (non-mocked) extraction was triggered via POST /v1/policies/000ef49e.../versions/f85bb8e7.../extract against NCD 110's real 2003 coverage-expansion diff — the same exit-criterion check called out below. Both open questions are now resolved:
gpt-5.1acceptstemperature=0without error — confirmed- Forced function-calling worked reliably — the run
SUCCEEDEDon the
(model_name: "gpt-5.1-2025-11-13", temperature: 0.0 in the run record).
first attempt (attempt_count: 1), returning well-formed structured facts[] with no malformed-JSON fallback triggered.
The run completed in ~12.7s, extracted two facts (device-description rewording and the October 1, 2003 additional-indications expansion, correctly pulling effective_date_candidate: "2003-10-01" out of prose), both verified against real current-version section text (evidence_side: "current"), scored 0.955 and 1.0 by confidence.py, and AUTO_PUBLISHED into ChangeEvent rows with full EvidenceSpan lineage back to the source document. This was the last open item from both M5/M6 and this swap — end-to-end LLM-backed extraction is now confirmed working in production, not just unit-tested against a mock.
One production config gap was found and fixed en route (v1.0.681): docker-compose.yml's policy-web service environment: block never listed SENTINEL_POLICY_OPENAI_API_KEY or the other M5/M6 extraction settings, so the key sat correctly in prod.env but never reached the container (require_extraction_configured kept 503ing). Each service's environment: block is an explicit allowlist — setting a var in prod.env alone doesn't pass it through.