Sentinel Signal

MCP Verify — Round 10 Implementation Report (post-1.0.510 verification)

Source: docs/mcp-verify-round-10-implementation.md

Document Content

MCP Verify — Round 10 Implementation Report (post-1.0.510 verification)

Implements the P0 items from mcp-verify-requirements-round4.md — the four requirements docs in this series build on each other; the "round 10" label here matches this doc set's own internal numbering ("Round 4" in the doc's title, "round 10" in the ongoing implementation-session count). All three P0 items shipped: R24 (classifier accuracy), R25 (schema divergence detection), R23 (complete the R1 audit). P1/P2 items are explicitly deferred — see "Deferred to a follow-up round" below, with reasons.

R24 — Capability classifier accuracy

R24.1 — removed freeform-string-implies-exec inference. The literal bug: query was in the exec-parameter-name hint list, so any tool with a bare search-query string parameter (e.g. pyrimid_browse, a catalog search tool) was classified exec/High. Fixed with a two-tier hint system in verify/src/mcp_verify/validation/service.py:

  • EXECUTION_PARAMETER_HINTS (command, cmd, script, shell) — strong
  • enough alone.

  • EXECUTION_AMBIGUOUS_PARAMETER_HINTS (code, template, expression,
  • statement) — only counts with description corroboration from EXECUTION_DESCRIPTION_HINTS.

**R24.2 — added financial/irreversible as a first-class signal.** New financial capability and financial_transaction risk flag, derived from either (a) a numeric amount/price parameter paired with an identifier parameter, or (b) description text matching an action-verb phrase (purchase, pays via, transfer funds, checkout, etc — never a bare noun; see below). infer_tool_risk_level now treats financial_transaction as high-severity on its own. pyrimid_buy (spends real USDC via x402, irreversibly, on a schema with no amount field at all — only the description states the payment semantics) now classifies financial/high-risk.

Two false positives were found and fixed via the labeled test set itself, before shipping:

  • A currency name ("USDC") in a description is not evidence of an executing
  • financial action — pyrimid_register_affiliate (records a wallet address for future payouts, doesn't move money) was misclassified until "usdc" was removed from the hint list.

  • Same for bare nouns like "invoice" — list_recent_invoices (a paginated
  • read-only listing) was misclassified until "invoice"/"subscription"/ "payment"/"settlement"/"charge" were removed, leaving only unambiguous verb phrases.

R24.3 — labeled set expanded with the 7 real tools from ai.pyrimid/pyrimid in verify/tests/test_capability_classifier.py, including the search tool and the payment tool named in the requirements doc's evidence table. ACCUSATORY_CLASSES now includes financial.

R24.4 — corpus re-run: NOT executed. This is a production-scale write (re-scoring and correction-logging across the full corpus), the same category of action as the round-6 corpus rewrite, which required explicit user confirmation before running. Not run without that confirmation.

R25 — Server-card vs. live-schema divergence

New build_schema_divergence_probe() in service.py, wired into validate_server_record() right after provenance_divergence_probe. Compares, per shared tool name: parameter set membership, required-parameter sets, per-parameter types, and output-schema presence. Reports error on required-set/type mismatches or tools missing from the live surface, ok when nothing diverges, missing when there's no card or no live tools to compare.

On the ai.pyrimid/pyrimid fixture this correctly detects divergence on all 6 shared tools, not just the 2 named in the requirements doc (a systemic camelCase-vs-snake_case naming mismatch across the whole server) — broader evidence of the fix's value than the acceptance bar required.

Per R25.2's explicit instruction, this is a separate finding, not folded into provenance_divergence_score: own REMEDIATION_RULES entry, own build_active_alerts() branch (server_card_schema_drifted), own remediation playbook — reusing the codebase's existing named-finding mechanism rather than inventing a new one or adding a 60th scored dimension.

Per R25.3, build_provenance_divergence_probe()'s output now includes compared_fields: ["title", "version", "homepage", "repository"] so an ok/no-drift verdict states what it actually compared, rather than reading as a blanket all-clear.

R23 — Complete the R1 audit systematically

The finding: four components (Step-Up Auth, Request Association, Interactive Flow Safety, Prompt Contract) scored majority credit (3/4, 3/4, 3/4, 2/4) on a live, confirmed-healthy server whose specific source probe never ran (status missing). Round 8's earlier zero-anchoring fixes only zero-anchored these functions' unconfirmed-handshake branch; a confirmed-healthy server with one unobserved, optional dimension is a distinct case that was never audited.

Grepped for the pattern class (per the standing R16/R18/R20/R25.3 instruction) rather than only fixing the four named instances — found and fixed a fifth, structurally identical bug in score_resource_contract (not named in the requirements doc's evidence table, same shape as score_prompt_contract).

The fix, applied to all five functions in service.py (score_step_up_auth, score_request_association, score_interactive_flow_safety, score_prompt_contract, score_resource_contract): when the source check status is missing and the handshake is confirmed, distinguish two cases instead of one flat default:

  • The server has a clear, independent signal it doesn't claim the relevant
  • capability at all (server.has_oauth False, server.has_prompts False, supports_resource_capability(...) False, or the probe's own advertised_capabilities/elicitation-or-sampling signal is empty) → exclude the dimension (None), matching R17's existing not_assessed/None-exclusion precedent for score_transport_compliance. compute_algorithmic_score_components already drops None values from both the composite's numerator and denominator — no new plumbing needed.

  • The server does claim the capability but nothing was observed → zero-anchor
  • (0.0) — a probe that should have run and didn't is not evidence of correct behavior.

score_step_up_auth's signature changed (now takes server to read has_oauth); the one other call site (scripts/backfill_score_and_capability_corrections.py) was updated and guarded against the new None return with a small _set_or_exclude helper that pops the key rather than crashing on float(None).

SCORE_COMPONENT_ZERO_POINT in main.py (the registry the public /methodology page renders) was updated to describe the new behavior accurately for all five components — and, in the course of doing that, the existing whole-server enforced-ceiling test (test_r1_zero_point_registry_is_an_enforced_ceiling_not_just_documentation) caught an over-correction: prompt_contract_score/resource_contract_score's zero point could not be dropped to a flat 0.0, because a genuinely attempted-and-failed check (error/warning/auth_required without claimed support — a different branch, never part of this bug) still legitimately returns up to 6.0 on a real, totally-dead server. Reverted to 6.0 with updated justification text describing exactly which sub-branch changed.

R23.1 — machine-readable mapping table, generated from code. New ml/registry/generate_score_component_registry.py: parses validation/service.py with ast, finds every def score_*(...) function, and statically extracts the literal check names it reads via checks.get("...")/checks["..."]. Cross-references SCORE_COMPONENT_ZERO_POINT from main.py for zero point + justification. Output: ml/registry/score_component_registry.json (59 components — 49 documented top-level dimensions, 10 internal sub-scores like score_protocol_conformance that feed a composite rather than appearing in the top-level metrics dict directly). verify/tests/test_score_component_registry.py regenerates and diffs against the checked-in file, so a code change that alters a function's check dependencies without regenerating fails CI.

R23.2 — CI assertion joining score to check status. Two layers:

  1. The pre-existing test_r1_zero_point_registry_is_an_enforced_ceiling_not_just_documentation
  2. (runs the real RemoteValidationService against a fully-dead mock server, checks every returned component against its documented zero point) — this already is exactly the systemic assertion R23.2 asks for, for the "everything is broken" case.

  3. New targeted tests in test_score_integrity.py for the specific shape
  4. R23's own evidence table used — one probe missing while the rest of a confirmed-healthy server is fine — covering both the exclusion branch and the zero-anchor branch for all five fixed functions, using the real ai.pyrimid/pyrimid fixture plus small synthetic cases for the "claims the capability but unobserved" branch (which the real fixture doesn't happen to exercise).

R23.3 — verified the failing-server case. task_success_score and installability_score are correctly 4/4 on the confirmed-healthy pyrimid fixture (new test: test_r23_3_task_success_and_installability_are_4_of_4_on_a_genuinely_healthy_server) and correctly ~0 on the broken getvari fixture (pre-existing tests, unchanged).

Test suite

334 tests passing (PYTHONPATH=verify/src pytest verify/tests), up from 319 at the start of this round.

Deferred to a follow-up round

Not started this round — all P1/P2 in the requirements doc, lower priority than the three P0s above, and substantial enough individually (data investigations, a rubric-drafting exercise, a methodology-page rewrite, a 50-tag manual review) that finishing them properly would have meant either rushing them alongside the P0 work or delaying the P0 ship. Flagging each explicitly rather than silently dropping them:

  • R11 — ground-truth rubric doc. Can build the artifact; cannot recruit
  • real reviewers (that's a people/process step, not a code change).

  • R26 — validation-diff rendering bug (utility_coverage_score renders
  • a blank "Latest" value). Needs the actual render function located and the empty-string-vs-None handling traced before fixing, not yet started.

  • R27 — duplicate validation runs at identical timestamps. Needs root
  • cause (double-submission vs. retry without an idempotency key) before fixing — a display-only dedupe would hide a real double-write bug if one exists.

  • R28 — historical Healthy+0-tools rows. The forward-going gate is
  • already covered by R18's not_assessed mechanism; the backfill of existing bad historical rows is a data migration and needs a dry-run count first, not attempted this round.

  • R12 restated — full funnel (discovered → ... → strong live) on the
  • methodology page. Not started.

  • R7 restated — 50-assignment tag spot-check. Not started; the pyrimid
  • database tag named in the doc is still live.

  • R9 restated — framing lint on remaining conclusory verdict labels
  • ("Allow With Approval," "Block For Production," etc). The RISK_FLAG_EVIDENCE_TEXT lint mechanism exists for risk-flag text; whether it (or a parallel lint) covers these four verdict-label phrases specifically was not checked this round.

  • R22 — build endpoint. Likely already satisfied by the existing
  • /v1/build + regression_checks.py from round 9; not re-verified against this doc's specific acceptance text this round.

  • R17, R19, R20, R14c, R15, R10 — unchanged, carried forward.