MCP Verify — Round 12 Implementation Report (post-1.0.512 verification)
Implements the full "MCP Verify — Requirements, Round 6" doc, found against ai.drillr/drillr (live v1.0.512), cross-referenced against ai.rams/rams (v1.0.511) and ai.pyrimid/pyrimid (v1.0.510). Preceded by a planning-only pass (see verify_round12_architecture_plan in memory) that diagnosed every P0/P1 item's root cause against current code before any implementation started. All P0 items shipped (R32, R33, R34.1, R30 restated), plus P1 (R35, R23 restated confirmed as non-bug, R11 labeled-set expansion, R22 confirmed already satisfied). R34.2 deliberately diagnosed and left unfixed per the doc's own instruction. 369 tests passing, up from 363 at the start of this round.
Standing order (§0): classifier freeze
No new capability-classifier rules landed this round. R30 (annotations) and R34.1 (removing an existing false positive) are both explicitly exempt per the doc's own text, and that's all that touched detect_tool_capabilities. R34.2 (a real, located gap) was deliberately left unfixed — diagnosed and added to the labeled set instead of shipped as an unmeasured rule.
R32 — Server-card divergence, with a fixture
ai.drillr/drillr's card names itself drillr-agentic (1 tool, search) while live initialize names itself drillr-data (9 tools, zero overlap) and declares api_key auth while live negotiates working OAuth.
R32.1/R32.2 — build_schema_divergence_probe (built round 10 for R25.1) already correctly flagged the tool-membership gap (tools_missing_from_live: ["search"], status error) — confirmed by direct testing before assuming a fix was needed. Extended it with two new comparison dimensions: serverInfo.name/serverInfo.version (card vs. live initialize) and declared-vs-observed auth (card's authentication.schemes vs. oauth_protected_resource status). Added a severity field to the probe's details (critical/high/medium/None) per the doc's tiers — critical for a differing server name or zero tool overlap, high for one-sided tool/auth absence, medium for the original round-10 same-tools-different-params case. Severity feeds the finding/alert text only; per R25.2's original instruction this probe still never feeds a numeric score.
R32.3 — the named finding/alert (server_card_schema_drifted, built round 10) now leads with severity and states which specific dimension diverged (server identity, auth scheme, or parameters), instead of always reading as a generic "schemas disagree" message.
R32.4 — build_provenance_divergence_probe's round-11 fix (readable_source_count < 2 → not_assessed) only guarded at the source level. Drillr's registry side is nominally "readable" (official_registry_probe.direct_match: true) but contributed zero populated fields (title/version/homepage/repository all null) — a field-level instance of the same "unknown reads as pass" pattern, one layer more granular. Added a comparable_field_count guard: fewer than one field with data on both sides now also triggers not_assessed, regardless of whether the sources themselves counted as "readable." Fixing this also surfaced that the round-11 regression test for the "genuine agreement" case had — by accident — been testing the pyrimid fixture, whose registry side turns out to have the identical zero-populated-fields shape; replaced with a synthetic registry payload that genuinely supplies a matching field.
R32.5 — the Capabilities block's "Server card: none" line read server.server_card_url, a persisted, registration-time DB column that was never populated for this server, instead of the live server_card check's actual status/fetched URL (confirmed ok, real payload, real URL in details.url). Now prefers the live check's fetched URL, falling back to the persisted field only if the check never ran.
R33 — One snapshot, one set of counts
R33.1 — confirmed live: action_safety_probe.details.summary (built at validation time inside validate_server_record, then persisted verbatim) said 4 exec / 4 high-risk tools; a fresh recompute of build_tool_security_inventory/summarize_tool_security_inventory against the same tools_list payload with current code said 0 exec / 1 high-risk — matching the page's own displayed table exactly. Root cause: classifier code changed between rounds 10/11 (deployed after this server's last validation) and the persisted probe was never refreshed. Same root-cause shape as R29's stale Server.current_score column, the second time this exact pattern has been found in this remediation effort.
R33.2 — new refresh_action_safety_probe_summary() in main.py: overrides the persisted action_safety_probe.details.summary with the already-fresh security_posture_summary at API-serialization time, in every place that builds a /report-shaped payload (the main endpoint, its partial-fallback variant already returned latest_validation: None so needed no change, and the prewarm-cache code path that duplicates the same payload shape). The underlying ValidationRun database row is never rewritten — build_validation_diff's legitimate use of the frozen historical value for trend comparison between two past runs is untouched. The override is marked explicitly (summary_is_live_recomputed: true) so a reader of the raw JSON can tell.
R33.3 — extended build_snapshot_invariant with optional page_counts/badge_counts/report_counts/policy_counts parameters. A shared snapshot ID across surfaces no longer implies agreement — when two or more surfaces supply counts, every shared key must match, or ok becomes False with a count_mismatches detail naming exactly which key and which surfaces disagreed. Wired real counts into the /report endpoint (detail.security_posture_summary vs. the report's own embedded action_safety_probe summary) as the concrete demonstration — after R33.2's fix these now trivially agree, which is itself the regression guard: if a future change ever reintroduces the split, this catches it immediately.
R34 — Revert the financial misfire, diagnose the sanitization miss
R34.1 — sec_report_search's own description ("share repurchase authorization", "dilution / SBC / buyback") matched "purchase"/"buy" in FINANCIAL_DESCRIPTION_HINTS as plain Python substring containment — "purchase" inside "repurchase", "buy" inside "buyback" — describing search topics about a company's own buyback program, not the tool executing a purchase. Fixed by switching to the same word-boundary matching (_blob_contains_word) already used for TOOL_CAPABILITY_TEXT_HINTS. This surfaced one self-caught regression before shipping: pyrimid_preview's "without buying" (an existing dry-run phrase) stopped matching "buy" under word-boundary rules, because "buying" was a missing inflection alongside the existing "buy"/"buys"/"bought" — added it. Also gated the financial_transaction risk flag on readOnlyHint is not True, reusing this round's own R30 annotation-first principle: a tool the publisher declares read-only cannot be transacting. The financial capability tag itself is unaffected (stays for transparency on ambiguous cases, per round 10's original reasoning).
R34.2 — root cause located precisely: detect_tool_risk_flags's risky-parameter-name detection (a token list including "query" — a different, pre-existing list from the exec-hints round 10 touched) has no entry matching a parameter literally named sql. run_sql's unconstrained SQL parameter therefore never counts toward freeform_input_surface/sanitization scoring. Per the doc's explicit instruction, not fixed this round — documented, and run_sql added to the classifier's labeled test set (deliberately unlabeled for any accusatory class) so the eventual fix lands measured against R24.3/R11's precision numbers, not shipped as another unmeasured rule.
R30 (restated) — extend the annotation veto to exec
Round 5's veto (readOnlyHint: false → never read) only ever covered the read/write axis. All 9 drillr tools declare readOnlyHint: true; a schema-shape exec inference (matched on at least one tool before this fix) survived the veto untouched. Added capabilities.discard("exec") alongside the existing read/write discards when readOnlyHint is True. detect_tool_annotation_conflicts (built round 11) already flags exactly this disagreement as a named finding, so the publisher's claim isn't trusted silently — vetoing the capability just stops double-counting the same disagreement as an outright classification too.
R35 — Verdicts must degrade when alerts fire
Three independent, confirmed bugs — not one bug with three symptoms.
R35.1 — the "Critical alerts" counter (build_production_readiness) counts alerts of severity {critical, high}, unrelated to the "No Critical Risk" badge (which checks tool risk level, a different dimension). Neither was wrong in isolation; the shared word "critical" made them read as contradictory. Relabeled the counter "High/critical-severity alerts" to describe what it actually counts.
R35.2 — infer_server_verdict (the shared, cheap, DB-level gate also used for ?verdict= list filtering) reads only status/score/freshness, deliberately alert-blind to avoid a per-row active_alerts computation at catalog scale. build_production_readiness received active_alerts but, before this fix, only used it to adjust reason text on the top verdict tier — safe_for_evaluation and needs_remediation never moved regardless of active alert count (confirmed live: drillr sat at "Safe for evaluation" with three active high-severity alerts, directly beneath the methodology's own claim that verdicts degrade). Added an explicit downgrade inside build_production_readiness (the one place active_alerts is already a required input): safe_for_production → safe_for_evaluation on any high/critical alert; safe_for_evaluation → needs_remediation on any true critical alert or 2+ high/critical alerts. infer_server_verdict itself is untouched — the downgrade is scoped to the one already-alert-aware call site, not the shared list-filtering gate.
R35.3 — "Write-safe publishing: Ready — No explicit blockers recorded" while write_action_surface_expanded (already computed by build_validation_diff) was true and action_safety_probe was in warning. build_client_remediation_modes's write-safe blocker checklist only populated when write_action_governance.get("safe_to_publish") was False — neither signal was ever read, regardless of safe_to_publish. Both are now wired in as independent blocker conditions; the mode's status goes "blocked" whenever either fires, even when governance alone would say ready.
R23 (restated) — confirmed non-bug, not fixed
The doc itself hedged this one ("OAuth Interop... the rams instance may have been coincidence, not a fix"). Direct fresh recomputation against the frozen drillr fixture (same methodology used for OAuth Interop in round 11): score_request_association/score_prompt_contract/score_resource_contract all correctly return None (excluded) right now — round 10's existing fix already handles this server correctly. The reported 3/4, 2/4, 2/4 are persisted staleness (this server hasn't been revalidated since round 10 shipped), not a live gap. score_oauth_interop correctly returns a real, evidence-based positive value (OAuth genuinely works here). No code change; locked in with a regression test so a future round doesn't re-investigate the same non-bug from the same stale evidence.
R22 — confirmed already satisfied
/v1/build + mcp_verify.regression_checks (built round 9) already provide exactly what this item asks for — version, commit, deploy time, per-assertion pass/fail. This round's new checks (check_r30_restated_read_only_true_vetoes_exec, check_r34_1_financial_hint_is_word_boundary_matched, check_r25_3_one_readable_source_is_not_assessed_not_ok from round 11) are all registered and passing. No further action needed.
R11 — labeled-set scope expanded
The doc's expanded scope ask (tools with/without annotations, annotation conflicts, payment/transfer tools, read-only domain-vocabulary search, query-language parameters) is now substantially covered by this round's own fixture work: drillr contributed 9 read-only, annotated, financial-vocabulary search tools plus one genuine query-language parameter (run_sql's sql, deliberately left unlabeled per R34.2's scope). The one gap called out explicitly in the labeled set's own honesty note: no medical- or legal-vocabulary search tool example yet, real or hand-built. The human-reviewer rubric (verify/eval/ground-truth/rubric.md, scaffolded in an earlier round) needed no changes — it governs a different process (human ground-truth ratings) from the classifier's own labeled precision set (R24.3, in test_capability_classifier.py).
Test suite
369 tests passing (PYTHONPATH=verify/src pytest verify/tests), up from 363 at the start of this round. Labeled classifier set: 36 real tools (up from 27), still short of the plan's 50-tool target — same honesty note as every prior round.
Deferred
Nothing new deferred beyond what round 11 already carried forward (R27, R28, R12/R7/R9/R10 restated, R17/R19/R20/R14c/R15, R11's human-reviewer recruitment) plus R34.2 (deliberately diagnosed, not fixed, per this round's own explicit instruction) and R24.4/R6-style corpus-scale correction logging (not run without explicit confirmation, same posture as every round since R6).