Sentinel Signal

MCP Verify — Round 15 post-1.0.530 Epovest implementation

Source: docs/mcp-verify-round-15-post-1.0.530-epovest-implementation.md

Document Content

MCP Verify — Round 15 post-1.0.530 Epovest implementation

Target release: 1.0.535.

Source: "MCP Verify — Requirements, Round 15 (post-1.0.530 Epovest verification)", verified against one fixture: Epovest/mcp-server (trustsnap_b37cac4e2b825460) — 61 tools, 42 declared non-read-only, 5 destructive, a real payment tool, 7 of 8 prior runs non-enumerating. Every claim was checked against the live production API response and rendered HTML page for this fixture before any code changed; two of the doc's own evidence points (R75.4's confidence count, R76's live-visible divergence) had already shifted by the time this pass started, since a new validation run landed between the doc's snapshot and this session — noted inline below rather than silently re-derived.

Implemented changes

  • R75.1: change detection now compares the current validation run
  • against the most recent run that actually observed a tool surface, never the immediately-adjacent run. New shared helper, find_previous_enumerating_validation, reuses the existing validation_has_observed_tool_surface predicate (built for round 12's R56.4 healthy-ratio filter) and is applied at every call site that previously did a bare validations[idx + 1] / validations[1] lookup: build_server_insights, build_agent_match_context, build_validation_timeline, build_public_server_reputation, build_server_incident_feed, the hosted-runtime alert helper in main.py, and the /v1/servers/{namespace}/{name}/history and /timeline API endpoints. Confirmed live: a non-enumerating baseline fabricated three false high-severity alerts (auth_mode_changed/write_action_surface_expanded/tool_snapshot_changed) purely because Verify itself failed to enumerate on its own prior run — none was a real change to the server, and they degraded the production verdict and blocked write publishing.

A second, related bug found while fixing this: three routes (the main /servers/{namespace}/{name} page, /v1/servers/.../evidence, and /v1/servers/.../history and /timeline) fetched validation history with full_payload_limit=1 — meaning every row past the newest was guaranteed lightweight (checks={}), so find_previous_enumerating_validation could never find a real baseline at all, silently suppressing every diff-derived field (not just the false ones) rather than fixing the problem. Bumped to full_payload_limit=2, matching the pattern already used at every other list_validation_runs_for_server_id call site in this codebase (and _get_cached_server_detail_response's own default).

  • R75.3: the "Validation history" table (render_history_summary)
  • showed 0 tools for a run that never fetched a tool surface — point.get("tool_count") or 0 collapsed the already-correct None to a false 0. The "Validation timeline" table (already fixed for this class of bug in round 18's R56) showed "not fetched" for the exact same underlying run. Both tables now agree.

  • R75.4: investigated, does not reproduce in code. healthy_ratio_recent
  • (build_history_summary) already filters through validation_has_observed_tool_surface before computing the ratio — an existing test, test_r68_recent_healthy_ratio_excludes_unobserved_tool_surfaces, already pins this. Live re-check found evidence_confidence.reason now reads "Based on 2 recent validations" (not "1" as the requirements doc's snapshot shows) — a second enumerating run landed between the doc being written and this pass starting. No code change made for this item.

  • R75.5: compute_summary_status already returns "unverified" for
  • exactly this case (observation_state in {"auth_gated_or_unfetched", "explicit_empty_surface"}, shipped as R56.3/R68 in the round-13 post-1.0.524 feedback release). scripts/backfill_unverified_validation_status.py already exists and already implements the correction (dry-run-first, --apply requires explicit confirmation) — written in an earlier round but never actually run against production. Ran post-deploy: the actual corpus-wide scope was 3 validation runs on 1 server (github-depixapp/depix-mcp), not the large sweep the planning pass expected from Epovest's own 7 non-enumerating runs. Querying the database directly (bypassing the API's lightweight-row fetch entirely) showed Epovest's persisted ValidationRun.checks were never actually empty or wrong — the "not fetched" appearance in the timeline API was purely list_validation_runs_for_server_id's full_payload_limit optimization choosing not to fetch/display older rows' real checks for query-cost reasons, not a genuine data defect. R75.1's fix (comparing against the last run the API actually fetched) is what corrects Epovest's symptom; the backfill addresses a smaller, genuinely different population — rows where compute_summary_status's pre-R56.3 logic really did persist an incorrect "healthy" label. Applied with explicit user sign-off after reviewing the dry-run output; verified directly against the database afterward that all 3 rows now read unverified. 4 stale materialized payload rows were also cleared so they rebuild from the repaired state.

  • R76: build_publishability_policy_profiles's Claude status
  • computation had admin_refresh_required=False hardcoded (preserved faithfully from the pre-R71.2 ternary, which never referenced it for Claude at all), while claude_gates["admin_refresh_required"] still displayed the real computed value — a reader saw identical field values in both clients' gates tables, but only ChatGPT's status computation actually used the value to downgrade from "ready". Confirmed live: Epovest showed six identical field values and identical summary prose, but "! Compatible with review" for ChatGPT against "Connector-compatible" for Claude. admin_refresh_required originates from tool-schema refresh-breakage signals that apply to any client persisting an installed configuration across schema changes, not an OpenAI-connector- specific concept — both clients now weight it identically. R76.2's "publishability capped by client compatibility" is already structurally guaranteed: both status computations derive their base value from the same client_profiles[key].compatibility field client_readiness_verdicts itself reads, so no additional code was needed once the one real asymmetry was removed.

  • R77.1: extracted the full client-blocker computation (client-
  • compatibility criteria checklist, plus publishability policy gates, plus request-association and transport-compliance extras) out of build_client_remediation_modes into a new shared function, _full_client_blocker_list. build_client_readiness_verdicts (the "Client compatibility" block) previously derived its summary from only the narrower criteria checklist — compatibility can read "compatible" (an empty criteria checklist) while the policy gates are still failing, exactly Epovest's case, which is why "No major blockers detected" rendered directly above a details panel listing six real blockers for the same badge. Both blocks now build their summary text from the identical blocker list (each may still render it differently, per the established R70 summary/checklist text-shape split — they just can no longer disagree about what the blockers are).

  • R78.1/R78.2: added Epovest/mcp-server's 10 flagged tools to the
  • R24.3 labeled set (test_capability_classifier.py) — the first fixture with both correct and incorrect arbitrary_network_egress calls on one server. 4 confirmed incorrect via each tool's real input_schema description: accept_keyword_discovery/dismiss_keyword_discovery/ restore_keyword_discovery (domain must match a prior list_keyword_discoveries result, not a caller-chosen fetch target) and list_sources (domain is a substring filter over stored rows, readOnlyHint: true). Root cause traced but not fixed this round, per the doc's own explicit instruction: the schema-based URL/host parameter heuristic matches on the parameter being named domain, independent of whether the description describes a fetch target versus a lookup key or filter. Recorded as a labeling criterion for R24.3's eventual precision measurement (expected_flags=set(), deliberately not asserting either the current or a hypothetical fixed answer), matching the exact precedent already established for round 12's _FIXED_ENDPOINT_API_WRAPPER_TOOLS.

Not changed, deliberately

  • R78.3's "Disabled-by-default candidates" labeling caveat — recommended
  • in the plan as a lower-risk P1 action, but no fixture evidence this round showed it actively misleading a reader beyond what R78.1/R78.2 already document; deferred rather than adding an unevidenced UI change in the same pass as the classifier-precision groundwork.

  • R72 (registry endpoint) verification, R66 verification, R74 distribution
  • report — all require fixture discovery or a corpus-wide script run, not code changes; carried forward per the plan.

Validation

  • PYTHONPATH=verify/src pytest verify/tests -q — 475 passed (up from
  • 469).

  • python3 scripts/export_openapi.py --check — fresh.
  • pytest (root repo, app/token_service) — unaffected; no files
  • outside verify/, docs/, CHANGELOG.md, and version-artifact files touched.

Live re-verification (post-deploy) — confirmed

Deployed to v1.0.535. Epovest/mcp-server re-fetched from the live public API and rendered HTML page afterward, every claim below independently confirmed, not assumed from a successful deploy:

  • R75.1: validation_timeline[1] (the run that previously showed
  • change_flags: ["auth_mode_changed", "write_surface_expanded", "tool_snapshot_changed"]) now shows change_flags: []. active_alerts is empty (previously 3 false high-severity alerts); production_readiness.code == "safe_for_evaluation" with degraded_by_active_alerts: false (previously degraded to "Needs remediation").

  • R75.3: "Validation history" table's Aug 10, 07:02:26 AM row now
  • shows 61 (was 0); the Aug 09, 06:50:52 PM row shows not fetched (was 0) — both tables now agree.

  • R76: publishability_policy_profiles[].gates.admin_refresh_required
  • currently reads False for both clients on this fixture (the specific divergence isn't live-reproducible right now, same as last round's R76 investigation found for the analogous case) — confirmed via a direct unit test instead (test_r76_admin_refresh_required_downgrades_both_clients_identically) that the fix holds when it's True.

  • R77.1: client_readiness_verdicts[].reason no longer reads "No
  • major blockers detected" — it now shows the same real blocker text client_remediation_modes[].summary shows for the same client (5 real blockers, confirmed identical between the two).

  • R77.2 duplicate search: found a third occurrence of "No major
  • blockers detected" in render_client_profiles ("Client Profiles" block) — assessed as a legitimately different, narrower-scoped display (always paired with its own "N of M requirements met" qualifier directly adjacent, unlike the two blocks that shared one "✓Client-compatible" badge with no distinguishing scope language) and left unchanged. Recorded here rather than silently ignored or reflexively "fixed" without evidence it's the same defect shape.

  • R75.5: backfill dry-run reviewed with the user, applied with
  • explicit sign-off, and verified directly against the database afterward — see the R75.5 section above for the corrected scope finding.