Sentinel Signal

MCP Verify — Round 13 Implementation Report (post-1.0.513 verification)

Source: docs/mcp-verify-round-13-implementation.md

Document Content

MCP Verify — Round 13 Implementation Report (post-1.0.513 verification)

Implements the full "MCP Verify — Requirements, Round 7" doc, found against ai.moonlings/moonlings (live v1.0.513, snapshot trustsnap_1c5cd31d9ac5235a). Preceded by a planning-only pass (see verify_round13_architecture_plan in memory) that diagnosed every P0/P1 item's root cause against current code before any implementation started. All P0 items shipped (R37, R38, R39), plus P1 (R40, R27, R28). 377 tests passing, up from 368 at the start of this round.

Standing order (§0): classifier freeze

No new capability-classifier rules landed this round. R37 wires an existing signal (declared_non_read_only_tools, built round 11) into downstream consumers that never read it — explicitly exempt under the doc's own text, since it's wiring, not a new rule.

R37 — declared_non_read_only_tools made load-bearing

ai.moonlings/moonlings's start_deep_report ($9.99/call, no confirmation parameter) correctly classified non-read (round 11, R30.3), but declared_non_read_only_tools — already computed in summarize_tool_security_inventory()'s output — sat unread by every consumer downstream of it.

R37.1 — build_action_safety_probe (validation/service.py): the status ladder's "no high-risk/destructive/exec tools" branch returned "ok" unconditionally. Now returns "warning" when a declared non-read-only tool exists without a confirmation signal or safeguard — "warning", not "error", per the plan's own reasoning: a single declared write tool with a clear price disclosure and no other risk flags isn't equivalent to an unconfirmed exec-capable tool. score_action_safety already had an unconditional if probe.status == "warning": return 5.5 branch, so the score cascades correctly (8.5 → 5.5) with no further change.

R37.2 — build_write_action_governance (insights.py): extended risky_action_surface to OR in declared_non_read_only_tools > 0, and blast_radius's medium tier to trigger when a declared write tool has no publish controls. Fixed a pre-existing type bug found alongside this: has_publish_controls = auth_boundary == "..." and (safeguards > 0 or confirmation_signals) could resolve to the literal list confirmation_signals or [] instead of a real bool (Python's and/or don't coerce), newly reachable once this branch became reachable for a case it wasn't before. Wrapped in bool(...).

build_benchmark_tasks (insights.py): the safe_write_flow_with_confirmation benchmark never checked for confirmation at all — a declared write tool with zero confirmation_signals still reported "passes". Now forces "degraded" when a declared-write tool lacks a confirmation signal, checked before the existing passes/degraded/likely_to_fail logic.

Verified end-to-end against the frozen moonlings fixture (test_r37_declared_non_read_only_tool_is_load_bearing_end_to_end): action_safety_probe.status → "warning", write_action_governance.status → "warning", safe_to_publish → False, blast_radius → "medium", and the confirmation benchmark → not "passes".

R39 — missing means one thing

Round 10 (R23) built a None-exclusion mechanism (a score_* function returning None drops that dimension from the composite denominator entirely) intended for genuine not_assessed cases — Verify itself couldn't assess the dimension (owner opt-out, an unreadable comparison source). It was then applied to five components' plain missing status too ("server doesn't claim this optional capability"), producing exactly the inconsistency R39 reports: a server graded on fewer dimensions could outscore one zeroed for an identical gap.

R39.1 — reverted the missing-status None-exclusion in five functions, all now zero-anchoring instead: score_step_up_auth, score_request_association, score_prompt_contract, score_resource_contract, score_interactive_flow_safety (only its "doesn't claim the capability" branch — the metadata-corroboration branch for servers that do claim it, unobserved, is untouched). Three functions were confirmed genuinely not_assessed-only and left unchanged: score_transport_compliance, score_action_safety, score_provenance_divergence — their None branches are gated on a literal probe.status == "not_assessed" comparison, not on missing.

R39.2/R39.3 — ml/registry/generate_score_component_registry.py (built round 10) extended with an AST-derived exclusion_capable (does the function's return type allow | None) and not_assessed_only_exclusion (is every bare return None reachable only through a branch whose condition mentions the literal "not_assessed") field per component. test_score_component_registry.py asserts no component is exclusion_capable without not_assessed_only_exclusion — a CI guardrail against a sixth component acquiring the same bug shape, and a converse pin that exactly {action_safety_score, provenance_divergence_score, transport_compliance_score} remain exclusion-capable. This is a narrower, code-derived audit rather than the doc's full empirical per-status × per-component matrix — a scope decision made during implementation given R39.1's own size; flagged here rather than shipped silently.

R39.4 — verified analytically and empirically that newly_assessed/newly_excluded component-diff tracking (insights.py, round 11) still fires correctly: with five fewer functions capable of returning None, only the three genuinely-not_assessed-capable components can ever trigger it now. No code change needed.

R38 — Badge and blocker eligibility read from the verdict

Confirmed live: "No Critical Risk" and "Write-Safe" badges active alongside a Block-for-Production verdict and 2 active high-severity alerts; "No explicit blockers recorded" printed beneath a blocking verdict.

R38.1/R38.2 — build_trust_badge_states (main.py) now reads production_readiness.code (already present in its payload parameter) and gates both badges: ineligible when the code is needs_remediation or metadata_only. The Write-Safe badge additionally requires that any declared non-read-only tool have a confirmation signal — the same write_confirmation_required_but_absent shape built for R37.2, reused rather than duplicated. Reused production_readiness's own code vocabulary instead of inventing a fourth verdict computation (build_executive_verdict already flags itself as a third, independent one — no reason to add a fourth).

R38.3 — build_client_remediation_modes (insights.py) now accepts production_readiness/active_alerts (moved build_production_readiness's call earlier in insights.py so it's available before this call site). The write_safe mode's blocker checklist previously only reflected its own narrow inputs (auth boundary, confirmation signals, exec tools, surface expansion, probe status) — it could be entirely clean while the signals that actually drove a blocking verdict (e.g. tool_snapshot_changed, auth_mode_changed) were never read. Now appends the driving high/critical alerts as blockers when the verdict is blocking, so an empty checklist can no longer coexist with a blocking verdict.

R38.4 — added a genuine end-to-end contradiction-guard test (test_r38_4_no_positive_badge_or_empty_blockers_alongside_blocking_verdict_and_active_alerts, test_api.py) reusing the existing drifty two-validation-run fixture shape (real tool_snapshot_changed/auth_mode_changed alerts) with the score dropped below the evaluation floor to force a genuinely blocking verdict, hitting both /v1/servers/{ns}/{name} and /v1/servers/{ns}/{name}/badges/trust, and asserting neither badge is active and the write-safe blocker list is non-empty.

R40 — Rank/percentile gated on the verdict, not just Failing status

score_percentile (db/repository.py) already gates on status == "failing" and tool_count <= 0, but a server with a genuinely blocking verdict (needs_remediation) and a non-failing summary status could still show a flattering "Top N% of scored public servers" line two rows beneath its own "Block For Production" decision on the same card. Fixed at the display layer, not in score_percentile's computation — matching R35.2's (round 12) precedent of leaving the cheap DB-level gate verdict-blind and applying the downgrade only where the verdict is already available. render_server_answer_block already receives executive_verdict for the decision text itself; the percentile string is now suppressed in the same place when executive_verdict["decision"] == "Block for production".

R27 — Duplicate validation runs: a claim-time race, not an enqueue-time one

Confirmed worse: runs ~7 seconds apart, identical score. The enqueue-time dedup (prune_duplicate_pending_jobs, the per-tick active_validation_servers check) already prevents most duplicate pending job creation, so the scheduler was investigated first per the plan's own instruction — root cause was one level down, in claiming.

JobRepository.claim_due_jobs (db/jobs.py) issued a plain SELECT ... WHERE status='pending' ... followed by a separate ORM-tracked UPDATE, with no row lock between them. Production runs 3 verify-worker replicas (VERIFY_WORKER_SCALE=3, deploy/ionos/docker-compose.yml), each polling independently. Under READ COMMITTED, two replicas' SELECTs can both see the same still-pending row before either replica's UPDATE commits, so both claim it and both validate the same server — the ~7s gap is the two replicas' independent outbound HTTP round-trips to the live target, not the claim race itself (that window is milliseconds).

Fixed with SELECT ... FOR UPDATE SKIP LOCKED on Postgres (a row locked by one claimer becomes invisible to a concurrent claimer's SELECT, rather than raced with it), gated by dialect — SQLite (the test suite) doesn't support SKIP LOCKED and doesn't need it, since tests run single-threaded and no concurrent claimer ever exists. Matches the existing dialect branch already established in prune_duplicate_pending_jobs. Verified with two new tests compiling the actual claim statement against both dialects (postgresql+psycopg:// engines resolve their dialect from the URL alone, no live connection needed) — confirms FOR UPDATE SKIP LOCKED is present in the Postgres-compiled SQL and absent from the SQLite-compiled SQL.

Live confirmation that duplicates stop requires an actual deploy under real concurrent load — noted the same way TASK-08b's cache fix needed a live two-deploy check; can't be confirmed by code/test review alone.

R28 — Zero-tool servers no longer report Healthy

Confirmed worse: 6 of 7 historical runs for one server. compute_summary_status (validation/service.py) gated "healthy" only on tools_status == "ok" — that only means the tools/list RPC call itself succeeded, not that it returned anything. A server answering with an empty tools array passed every other branch and fell straight through to "healthy".

Added a tool_count parameter (threaded from the one call site that has it, validate_server_record, via len(current_tools)) and a branch returning "unknown" when tool_count <= 0. "unknown" is not a new status invented for this fix: it was already Server.summary_status's DB default, already in SUMMARY_SCORE_MAP, and already a selectable server-list filter option — this is simply the first branch that actually produces it, reserved for "nothing failed, but there's nothing to call healthy either," distinct from "degraded" (something observable regressed) and "failing" (a core flow broke). Backward-compatible: tool_count defaults to None, preserving old behavior for any other caller.

Not done, needs explicit confirmation: a historical backfill of already-persisted ValidationRun.summary_status/Server.current_status rows that were marked "healthy" under the old, buggy rule. This changes live production data (not just going-forward behavior) and per this engagement's standing constraint on corpus-scale operations, it isn't run without the user asking for it explicitly.

Test suite

377 tests passing (PYTHONPATH=verify/src pytest verify/tests), up from 368 at the start of this round.

Deferred / carried forward

P2 items from the requirements doc (R32.2, R7, R41 naming inconsistency) not in scope for this round. The R39.2 scope narrowing (AST-derived audit instead of the full empirical per-status matrix) and the R28 historical backfill are both flagged above rather than silently shipped or silently skipped.