MCP Verify Round 7 implementation notes
Date: 2026-08-05
This documents Round 7: the remaining P1/P2 items from the post-1.0.503 remediation plan (R12-R15), plus R16 -- a new finding reported mid-round, folded into this pass as P0 (same severity class as the R2 suppression-marker bug: a null/empty value that means two different things, one of which is safety-critical to distinguish).
R16 -- Fast-fallback responses must not fail open
Reported directly against a live response: a degraded (cache_note: "Fast fallback response...") reply from /report returned tool_security_inventory: [], security_posture_summary: {}, write_action_governance: {}, capability_taxonomy: [], active_alerts: [], remediations: [] -- empty containers, not error states. A consumer reading active_alerts: [] cannot tell "we didn't load this" from "this server has no active alerts," and the fixture this round is built against (getvari/vari-mcp) has five active alerts including two criticals that a degraded response would have silently reported as zero.
Root cause: five "fast fallback" builders (build_fast_report_payload, build_fast_policy_payload, build_fast_badge_metadata, build_fast_server_page, and their shared pattern) construct ServerDetailResponse.model_validate(route_server) directly from the raw Server row instead of the full build_server_insights() pipeline, to protect web capacity when the cached response has expired. The six fields above are computed only by that full pipeline, so on the fast path they silently default to Pydantic's empty container -- indistinguishable from "checked, found nothing."
Traced the actual severity, not just the symptom. The worst instance is /policy: build_policy_export_payload derives allowed_tools/blocked_tools/allow entirely from tool_security_inventory. With that empty, zero tools get blocked and "allow": not blockers evaluates to true -- a policy consumer (the GitHub Action wedge, a gateway's gateway_config) would mechanically permit a server this exact response never inspected. A close second: the badge endpoint's "No Critical Risk" badge (build_trust_badge_states) reads security_posture_summary.risk_distribution.critical, which is 0 on the empty fallback -- rendering a false-positive green badge.
Fix, confirmed with the user before implementing (chose fail-closed over label-only):
- Added
partial: boolandfields_unavailable: list[str]toServerDetailResponse(api/schemas.py) and toDETAIL_PAYLOAD_KEYSso they survivedetail_payload_from_model's shallow view. PARTIAL_FALLBACK_SECURITY_FIELDSnames the six affected fields once; all five fast builders now setpartial=True, fields_unavailable=[...]via the samemodel_copy(update={...})call they already use for the "evaluation only" / "warming" labels.build_policy_export_payload: whenpartial, appends"evidence_unavailable_fail_closed"toblockers(forcingallow: False) and forcesrequires_human_approval: True. A normal, fully-computed response is unaffected.build_trust_badge_states: whenpartial, the"No Critical Risk"badge is forced inactive with an explicit "unavailable, not confirmed clean" reason instead of defaulting to a false green.build_fast_report_payload'scache_notenow says explicitly that empty fields infields_unavailablemean unchecked, not clean.
Tests: verify/tests/test_score_integrity.py (test_r16_fast_fallback_policy_fails_closed_not_open, test_r16_fast_fallback_badge_does_not_claim_no_critical_risk), both asserting the partial case fails closed and the non-partial case is unaffected.
Not changed: the ledger/audit-evidence fast fallback (build_fast_server_ledger_payload, a different shape -- agent invocation counts, not security fields) shows a similar pattern (invocations: [], audit_events: []) but wasn't part of what was reported and is a compliance-log concern rather than a deployment-gating one; flagged here rather than silently left unexamined.
R12 -- Homepage funnel
The homepage showed Indexed 74,748 -> Fresh validations 74,023 -> Strong live 3 with nothing between, and "Fresh validations" read as "passed" when its query (last_validated_at within the freshness window, any status) actually means "attempted." Replaced the flat 3-stat block with a real funnel where each stage is a strict subset of the one before it:
Indexed/Discovered -> Endpoint on record -> Validated at least once -> Handshake + tools succeeded -> Healthy -> Fresh -> Strong live candidates
build_public_coverage_stats() gained four new cheap COUNT() queries against the servers table (endpoint_on_record, handshake_and_tools_ok, healthy_servers, healthy_and_fresh) -- no new joins against ValidationRun, consistent with the function's existing style and this being a page computed behind the build-scoped route cache, not a per-request query. scored_servers (existing field) is reused, relabeled, as the "validated at least once" stage. The pre-existing fresh_validations field (used elsewhere, e.g. the methodology page's coverage table, with an already-accurate description) is left unchanged -- the homepage funnel's "Fresh" stage uses the new healthy_and_fresh field instead, since the funnel needs "fresh" to be a subset of "healthy" (the old field counts any status, which would make the funnel visually non-monotonic).
The 3 (strong live candidates) is untouched -- same query, same thresholds, per the plan's explicit "do not adjust the threshold to move it."
Tests: verify/tests/test_api.py::test_r12_coverage_stats_funnel_stages_are_strict_subsets -- five servers deliberately placed at each funnel boundary, asserts both the exact counts and that the stage values are non-increasing, plus checks the new labels render and the old "Scored servers"/"Fresh validations" labels don't appear on the homepage anymore.
R13 -- Navigation
The 1.0.503 changelog claimed "Simplified public navigation"; the header nav still carried a "Governance:"/"Publishers:"/"Research:"/"Docs:" set of grouped links (10 destinations) plus a footer with 13 more (23 total, close to the reported ~27 once the per-server Agent Commerce block is counted as its own navigable surface).
- **New
MAIN_NAV_ITEMSconstant** -- exactly 5 primary destinations: Search (/), Rankings (/rankings), TrustOps (/trustops), Docs (/methodology), Pricing (/pricing).render_site_nav()renders from this list. The pre-existing "Claim/Manage" link stays as a single, visually distinct CTA (an account/ownership action, not a content destination) rather than a 6th primary item -- this is the "publisher flows under one entry" consolidation: the old header also had a separate "Publishers: Claim / Badges / Profiles" group, which is now gone entirely (Badges and Profiles moved to the footer; Claim/Manage was already the one entry). - **Footer (
render_site_secondary_links()) now carries everything demoted from the header: Gateway, Compare, Trust Index (from the old Governance/Research groups), Badges, Profiles (from Publishers), Security, MCP API (from Docs). Docs Compiler and Agent Commerce are absent from the footer too** -- removed from navigation entirely, per the plan, not merely demoted. Their pages (/docs-compiler,/agent-commerce) still exist and are still linked from within relevant content (e.g. the pricing page's feature list), just not from site-wide nav. - Per-server Agent Commerce block removed.
render_agent_commerce_readiness()is no longer called when building a server-detail page; the "Agent Commerce & Payment Readiness" subsection is gone from the Compatibility tier (44 tier-subsections -> 43). The function itself and its JSON exposure (/report'sagent_commerce_readinessfield) are untouched -- the plan's ask was to stop presenting a signal it called "too weak to assess" on every one of 74k+ profiles, not to delete the underlying data. Not done this round: actually building the "Trust Index as research" home the plan suggests for this data -- that's new surface area (aggregate agent-commerce signal across the corpus), out of scope for "remove the per-server block." - CI enforcement, matching this codebase's established "no separate CI stage, enforce via pytest" pattern (same as R9's copy lint):
test_r13_primary_navigation_is_capped_at_five_itemsassertslen(MAIN_NAV_ITEMS) <= 5and pins the exact 5 labels, so a future addition regresses a test, not just a design intent.
Tests: verify/tests/test_api.py (test_r13_primary_navigation_is_capped_at_five_items, test_r13_navigation_drops_docs_compiler_and_agent_commerce_entirely, test_r13_per_server_agent_commerce_block_is_removed); updated the existing TASK-09 tier-restructure test (test_server_detail_page_regroups_sections_into_five_tiers) for the new 43-subsection count.
R14 -- Exponential backoff and probe discipline
Three independent fixes, all found by reading the actual validation battery rather than assuming the plan's framing was already correct:
1. Exponential backoff for persistently-failing servers. should_validate_server previously revalidated any failing/degraded server at a flat max(1, stale_after_hours/3) cadence (~8h) forever -- a server that's been failing for months got the same aggressive treatment as one that failed once. Added Server.consecutive_failure_count (migration 0021_consecutive_failures), incremented on every failing result and reset to 0 immediately on any non-failing result (RemoteValidationService.validate_server_record, the only writer). compute_failing_server_backoff_hours() doubles the interval per consecutive failure beyond the first (floor unchanged at count=1, so a single blip isn't penalized), capped at a 30-day ceiling so a server is never abandoned entirely. degraded status is unaffected -- it isn't a persistent-failure signal in this codebase's status model, and backing it off wasn't asked for. Owner-initiated revalidation (/revalidate, _queue_server_revalidation) already bypasses the scheduler's cadence gate entirely (it enqueues directly at high priority) -- confirmed, not changed.
2. Skip the rest of the ~26-check battery once initialize fails. is_core_success_from_check_results already treats initialize not in {ok, auth_required} as a hard failure -- but the validator ran the full battery against such a server anyway. Found via a request-count assertion in a new test, not by inspection alone: tools_list, prompts_list/prompt_get, resources_list/resource_read, session_resume_probe, transport_compliance_probe, and (found only because the count assertion didn't match the expected 1) _utility_coverage_probe's tasks_list_probe were all making real network calls regardless of whether initialize ever completed. All are now gated on attempt_downstream_probes (initialize_reachable or not server.remote_url -- the not server.remote_url clause preserves the existing zero-cost, more-specific "no_remote_url" reason for servers with no endpoint at all, rather than replacing it with a less precise "initialize unreachable"). Skipped checks get an explicit status="skipped" (via a new _skip_check() helper) rather than being omitted from the checks dict, so every existing checks["name"] access stays safe, and should_record_optional_failure/compute_summary_status treat skipped correctly (not a false failure, still correctly not "ok" for status computation).
3. Honor robots.txt as an obedience check, not only noise-resilience. _probe_noise_resilience_check already fetched /robots.txt from the origin, but only to test that the server returns a sane response to an unexpected path -- it never parsed the body. Added a minimal robots.txt parser (parse_robots_txt_disallowed_paths, robots_txt_disallows_path -- groups consecutive User-agent: lines per the de facto convention, not a fully spec-compliant parser) and reused the existing fetch to determine whether the site has disallowed mcp-verify. Scoped deliberately: gates only the two probes that go beyond the single legitimate initialize+tools/list handshake this product exists to run -- determinism_probe (2 extra tools/list repeats purely for our own determinism testing) and transport_compliance_probe's deliberately non-compliant bad-protocol-version request. The core handshake itself is never skipped on a robots signal alone; that's the whole point of the product, and "handshake succeeded" isn't optional/adversarial traffic the way a repeat-call determinism probe is.
Tests: verify/tests/test_scheduler.py (test_r14_failing_server_backoff_grows_exponentially_with_a_floor_and_ceiling, test_r14_should_validate_server_backs_off_failing_servers_by_failure_count), verify/tests/test_validation_service.py (test_r14_skips_downstream_probes_and_backs_off_when_initialize_is_unreachable -- asserts exactly 1 POST to the server's endpoint across the whole run, down from what would otherwise be 7+; test_r14_robots_txt_disallow_skips_only_the_exploratory_probes -- asserts the core handshake still runs while the two exploratory probes are skipped).
R15 -- Deploy consistency (verified, not rebuilt)
The original finding (three app versions served simultaneously across different URLs, one with stale cached coverage stats) predates this remediation plan and was already substantially addressed by infrastructure built in an earlier round, confirmed by direct inspection this round rather than assumed:
- Single instance, one build:
deploy/ionos/docker-compose.yml'sverify-webservice has noreplicas/scaledirective;deploy/ionos/deploy.shdoes an idempotent set-or-append ofMCP_VERIFY_BUILD_SHA/MCP_VERIFY_SITE_VERSIONinto the env file (not the old silently-no-opsed), thensystemctl restarts the one service. - Version-consistency check:
deploy.shwaits for/healthz, then runsscripts/verify_public_route_versions.shagainst 6 key routes (/,/pricing,/trust-index, a server detail page,/methodology,/status) with the just-deployed SHA/version as the expected value, and fails the deploy loudly on any mismatch, missingX-MCP-Verify-Buildheader, or a route serving withoutno-storecache headers. - Coverage-snapshot cache invalidation:
build_public_coverage_stats()is not independently cached -- it runs a real-time query every time it's called, embedded inside the same build-scoped-cached index page response as everything else on that page. Its earlier staleness was a symptom of the (now-fixed) stale-MCP_VERIFY_BUILD_SHAroute-cache bug, not a separate caching layer of its own; there was nothing further to fix here once that bug was closed.
Not resolved, flagged rather than silently dropped: verify/fly.toml still references the legacy Fly app sentinel-signal-mcp-index. DNS is confirmed fully cut over to IONOS (no AAAA records, A record points only at the IONOS host), so this app is not serving public traffic regardless of whether it's still running -- meaning it doesn't violate "all instances serve one build" in the sense that matters for the reported symptom (visitors seeing different versions). Whether it's still running and costing money is unconfirmed; decommissioning it requires Fly API/CLI credentials this environment doesn't have. Left as an explicit follow-up for whoever has that access, not silently ignored.
Verification
PYTHONPATH=verify/src:verify/tests pytest verify/tests -- 293 passed (up from 285 at the start of this round: 2 R16 + 1 R12 + 3 R13 + 4 R14 new tests, plus 2 pre-existing tests updated for the intentional 44->43 tier-subsection count and the renamed coverage labels). No live-deploy verification was performed as part of writing this document; that happens as part of the actual IONOS deploy this round ends with.