MCP Verify Round 9 implementation notes
Date: 2026-08-05
This documents "MCP Verify — Requirements, Round 3 (post-1.0.508 verification)," a handoff document supplied directly, verified against live production on 2026-08-05 (fixture snapshot trustsnap_35d5b83fe2bced51). It supersedes the prior remediation plan for status/priority; task definitions there remain authoritative where not restated. P0 items (R17, R1/R18, R21, R22) are complete. Of P1, R20 is complete; R19 was investigated to a concrete root-cause hypothesis but not fixed this round (see its section).
R17 — Owner opt-out (new, P0)
Honoring robots.txt (R14b, round 7) revealed that servers can and do disallow Verify's validator -- and the two probes that get skipped for that reason (determinism_probe, transport_compliance_probe) were indistinguishable from probes skipped because initialize had already failed. Both problems the requirements doc named are fixed:
**R17.1 — a distinct not_assessed status.** Added alongside the existing skipped status (R14): skipped means "not attempted because a different check already failed" (still zero-anchored -- that's a real negative signal); not_assessed means "not attempted at the operator's request" (excluded from the composite entirely, never zero-anchored). score_transport_compliance and score_action_safety now return None -- not 0.0 -- when their source check is not_assessed, and compute_algorithmic_score_components's final dict comprehension omits None values entirely, which naturally shrinks compute_current_score's denominator (len(components) * COMPONENT_POINT_MAX) rather than scoring the excluded dimension as a failure.
R17.2 — publication policy for opted-out servers. New persisted Server.owner_opted_out column (migration 0022_owner_opted_out, set in validate_server_record from whether any check in the run has status="not_assessed", reason="owner_opt_out") gates every public surface:
public_display_score()returnsNone-- no score, adverse or not.should_noindex_public_server_page()returnsTrue-- noindexed until the policy is confirmed live.- Found and fixed three independent verdict/decision implementations, not one:
build_production_readiness(insights.py),infer_server_verdict(db/repository.py, used for?verdict=filtering), andbuild_executive_verdict(main.py, the "Block for production" decision text) each now short-circuit to a distinct, non-adversenot_assessed_by_requeststate before their normal scoring/threshold logic runs. This is the same duplicate-verdict-implementation pattern TASK-31 found once before (Round 3) -- three computations of "is this server okay," never fully consolidated, each needing its own guard. - The listing entry itself (name, registry provenance, transport) is untouched -- opted-out servers stay indexed as listings, only the verdict/score layer is suppressed.
R17.3 — opt-out coverage in the funnel. build_public_coverage_stats gained an opted_out_servers count (new homepage funnel card, "Opted out"), reported as its own stage rather than folded into "failing" or any other bucket.
Tests: verify/tests/test_score_integrity.py (test_r17_robots_disallow_produces_not_assessed_not_skipped, test_r17_not_assessed_excludes_dimension_from_composite_denominator, test_r17_score_transport_compliance_excludes_not_assessed, test_r17_2_opted_out_server_gets_no_score_and_no_adverse_verdict, test_r17_3_opted_out_servers_counted_in_coverage_funnel); updated test_r14_robots_txt_disallow_skips_only_the_exploratory_probes (round 7) for the status rename.
R1 (complete) / R18 — the zero-observation class
R1's acceptance criteria from the requirements doc were already satisfied by round 8's fixes (task_success_score ≈ 0, installability_score ≈ 0 on initialize error) -- verified again this round, still true. The two remaining criteria were new:
**R18 — action_safety_probe reported ok on an empty, unconfirmed observation set.** build_action_safety_probe computed "0 high-risk tools" from an empty tool_inventory (because tools_list never confirmed anything) and reported status="ok" unconditionally -- a server with no observable tool surface was passing an action-safety gate. Fixed at the probe itself, not just the score function that reads it: when tool_inventory is empty and is_core_success_from_check_results(checks) is false, the probe now returns not_assessed (reason empty_observation_set), never ok. A genuinely-confirmed empty tool surface (tools_list actually succeeded and returned zero tools) still correctly reads ok -- that's real evidence of nothing to be unsafe with, not a default.
Fixing it at the probe (not just score_action_safety, which round 8 had already patched) means every consumer of action_safety_probe.status is automatically correct, not just the one score function -- confirmed directly: build_connector_publishability_probe's criteria.action_safety reads action_probe.status in {"ok", "warning"}, which now correctly evaluates False for not_assessed with no code change needed there.
**Audited every other build_*_probe function for the same pattern** (advanced_capabilities, tool_snapshot, connector_replay, interactive_flow, step_up_auth, official_registry, provenance_divergence -- the last already fixed in round 8 for the identical bug shape): all already correctly default to missing, not ok, on an empty observation set. action_safety_probe was the one genuine instance beyond the one round 8 had already found.
Also found, per the standing constraint's explicit "grep for it" instruction: two client-compatibility profiles (openai_connectors, claude_desktop in insights.py) counted transport_compliance_probe.status == "missing" as satisfying a compatibility criterion -- inconsistent with the anthropic_remote_mcp profile right next to them, which never included "missing" in its passing set. Fixed both to match (low-impact in practice, since both profiles already hard-gate on initialize/tools_list separately, but real and worth closing).
Documentation: SCORE_COMPONENT_ZERO_POINT entries for transport_compliance_score and action_safety_score now note the not_assessed-excludes-from-denominator case explicitly. Added a new "Opted-out servers" section to /methodology, and extended the "Subscore zero points" section's intro paragraph to describe the third case (excluded, distinct from both credit and zero).
Tests: verify/tests/test_score_integrity.py (test_r18_action_safety_probe_not_assessed_on_empty_unconfirmed_observation, test_r18_publishability_criteria_treat_not_assessed_as_unsatisfied).
R21 — a usable classifier fixture
getvari/vari-mcp can no longer verify R5: its initialize fails and its robots.txt now correctly disallows deeper probing (R17), so tool_security_inventory is empty for reasons unrelated to the classifier.
R21.1: selected ai.dreamlit/mcp from live production (status=healthy&tools_min=10, real API query, no fabricated data) -- Healthy, confirmed not robots-disallowed (probe_noise_resilience.details.validation_disallowed: false), 11 real tools spanning read-only, write, export/bulk-access, and genuinely destructive operations (confirm_publish, unpublish_workflow, flagged destructive_operation). Frozen into verify/eval/fixtures/regression/ai.dreamlit_mcp_{server_detail,validations}.json, documented in that directory's README alongside the existing getvari fixture.
R21.2: expanded the labeled capability set from 6 to 17 real tools (the existing 6 getvari calculators plus the 11 new dreamlit tools, each independently labeled against its real schema, not copied from the classifier's own output). Still well short of the plan's 50-tool target, and still zero genuine true positives for admin/secrets/exec specifically -- neither fixture server happens to have a tool that actually warrants those labels. Stated plainly in the test file's own honesty note, matching the pattern already established for R5/R6/R7's prior honesty notes: this session has no production database query access to pull a genuine random/systematic corpus sample, so 17 real tools plus hand-built representative cases is what's achievable, not a discharge of the 50-tool acceptance criterion.
R21.3: not done this round. Re-verifying R6 (corpus recompute) and R7 (taxonomy) against the new fixture is worth doing but wasn't reached given the scope already covered.
R20 — snapshot invariant honesty
build_snapshot_invariant computed ok = len(present_ids) <= 1 after filtering out every None surface -- with only 0 or 1 of the 4 possible surfaces (page/badge/report/policy) actually non-null, that's vacuously True: comparing a snapshot to nothing read as confirmed consistency, the same "compared nothing to nothing" pattern already fixed once this round in build_provenance_divergence_probe.
Investigated all 6 call sites first, since the reported evidence ("badge": null, "report": null) looked at first read like unreadable surfaces: every real call site only ever passes exactly 2 of the 4 possible kwargs by design (this function always compares the page snapshot against exactly one other surface -- never all four at once), so two null keys on every real response is not itself a defect, it's the shape of a pairwise comparison. The actual defect is narrower and real: nothing previously required at least 2 surfaces to be present before claiming ok. Fixed: ok now requires len(checked_surfaces) >= 2 and len(present_ids) <= 1, and the response gained a checked_surfaces field naming exactly which surfaces this specific call compared -- so "wasn't asked to check this surface" (by this function's design) is now distinguishable in the output from "was asked and it came back unreadable" (which, per the call-site audit, doesn't currently happen in practice, but the field makes the distinction explicit rather than relying on that being true forever).
Tests: verify/tests/test_score_integrity.py::test_r20_snapshot_invariant_never_vacuously_passes_on_one_surface.
R19 — investigated, not fixed this round
The requirements doc's evidence: fields_unavailable returned the identical six entries eleven hours apart on a server validated two hours prior -- partial behaving as the steady state, not load-shedding. Reproduced live during this round: three consecutive GET /v1/servers/ai.dreamlit/mcp/report requests, several seconds apart, all returned "partial": true with the same fast-fallback cache_note.
Traced the mechanism, not yet the full root cause. _get_cached_route_value's miss-fallback path takes refresh_on_miss_fallback=False for every per-server surface (report, policy, badge, server page, trust summary, ledger, management -- 12 of 13 call sites), while the single-key, low-cardinality surfaces (sitemap) use True. This looks like a deliberate capacity-protection choice (spawning a background-refresh thread on every cache miss across tens of thousands of distinct per-server keys during a traffic burst could genuinely overwhelm the process), not an oversight -- but the design's own second-request healing path (a miss-fallback entry is stored backdated past its TTL so the next request sees it as stale and triggers a real background refresh via refresh_builder) evidently isn't resolving within the reproduced window. Whether that's the refresh thread erroring silently (logged server-side as route_cache_refresh_failed, not visible without production log access), a key-matching bug between the read path and the refresh-thread's write path, or something else, wasn't isolated further this round -- doing so without production log access risks guessing at a fix for live caching behavior this session can't verify end-to-end, which is a worse outcome than leaving it open with a concrete, falsifiable hypothesis attached.
What was confirmed already exists and doesn't need building: route_cache_events_counter, a Prometheus counter labeled by cache/event (including every miss_fallback/miss_fallback_stored/miss_fallback_refresh/hit/stale transition), already captures the raw data a partial-response rate would be computed from -- R19's "instrument" ask is largely already satisfied at the collection layer. What's missing, and wasn't built this round: publishing a computed rate (e.g. on /status) and an explicit alerting threshold in the existing alertmanager stack.
Recommended next step for whoever picks this up: get production log access for the mcp-verify-server_report-refresh (per-cache-name) thread's exception logging, or add a temporary debug counter incrementing specifically inside _spawn_route_cache_refresh's except Exception branch and check whether it's firing for the server_report cache. That will distinguish "refresh threads are erroring" from "refresh threads aren't being triggered at all" from "refresh succeeds but something reads the wrong key" -- three different fixes.
Standing constraint: grep sweep
Per "unknown is never a pass ... grep for it rather than fixing instances as they are reported": audited every build_*_probe function (10 total) for the zero-observation-as-ok pattern (found: action_safety_probe, R18 above); audited every .status == "missing" comparison in validation/service.py for the R14/R17 "skipped"/"not_assessed" status-string interaction (found and fixed in round 8: score_session_resume; re-checked this round, no others affected since not_assessed only applies to the two robots-gated probes already handled); audited insights.py's client-compatibility profiles for the same "missing counts as passing" pattern found once already at the anthropic_remote_mcp profile's own (already-correct) implementation (found and fixed: openai_connectors/claude_desktop, R1/R18 section above); fixed build_snapshot_invariant's vacuous-single-surface pass (R20 above).
Verification
PYTHONPATH=verify/src:verify/tests pytest verify/tests -- 313 passed (up from 293 at the start of this round: 5 R17 tests, 2 R18 tests, 1 R22 endpoint test, 1 R20 test, plus the R21.2 fixture expansion exercised through the existing parametrized capability-classifier tests, and one round-7 test updated for the not_assessed status rename). /v1/build is live and returns regression_checks_passing: true for all 6 named checks. verify/tests is now wired into GitHub Actions CI (.github/workflows/ci.yml, new verify-checks job) -- a regression in any of these fixture assertions fails a PR directly, per R22.