MCP Verify Round 8 implementation notes
Date: 2026-08-05
This documents a re-audit of R1 (zero-anchor the subscores), triggered by an external check reported directly against production after v1.0.508 shipped. The check named two components (task_success_score, installability_score) and asked a specific, falsifiable question: on any server whose initialize returns an error, do they read ~0? They didn't. Verification confirmed the claim, then went further -- the underlying causes turned out to be broader than the two named components, including one bug in the core scoring pipeline itself that affected all 59 score dimensions, not just these two.
What was reported and how it was verified
The report gave three jq checks to run against a live, non-cached response, plus a direct claim: task_success_score's own description is "Can an agent reliably initialize, enumerate tools, and execute core MCP flows?" -- if it still reads non-zero on a server whose initialize fails, R1 hasn't landed for that dimension, regardless of what the version banner says. It also made a broader, harder-to-dismiss point: across four rounds of "verify this shipped," the one item checkable against fresh data (R1) turned out unimplemented for these two dimensions both times, and nothing in the test suite would have caught it, because nothing in the test suite called these functions.
Rather than trust the claim or dismiss it, it was checked directly against the repository's own code and its own frozen regression fixture (eval/fixtures/regression/getvari_vari-mcp_*.json -- a real server whose initialize genuinely 404s):
task_success_score: registry says 0.0 -- actual (pre-fix): 8.33 / 10
installability_score: registry says 0.0 -- actual (pre-fix): 7.35 / 10
Both exceeded their own documented zero point by roughly 7-8x. SCORE_COMPONENT_ZERO_POINT (the registry R1 built and publishes on /methodology) already claimed both metrics "naturally floor at 0" -- that claim was false, and had been false since it was written.
Root causes (four distinct bugs, not one)
**1. score_task_success -- a tautological criterion.** One of six pass/fail criteria was (initialize_status != "auth_required") or (oauth check ok). Whenever initialize fails outright (any status other than auth_required), the left side of that or is trivially true -- the criterion always "passed" for a server that never spoke MCP at all. Combined with tools_list occasionally succeeding via a fallback probe despite initialize failing (a real, observed case in the frozen fixture), this let a non-functional server pass most criteria.
**2. score_protocol_conformance (feeds installability_score) -- vacuous-truthiness criteria.** extract_check_payload defaults to {} on any missing/failed check. Several criteria checked things like isinstance(initialize_payload, dict) and "error" not in initialize_payload -- both trivially true for an empty dict, since an empty dict is a dict and has no "error" key to be absent of an error, not because a check actually confirmed anything. A tenth criterion, bool(server.remote_url), credited merely having a URL configured. 7 of 10 criteria passed on a server with zero real evidence, scoring 8.0/10.
**3. score_execution_success_bayes (also feeds installability_score) -- an optimistic Bayesian prior.** alpha_prior = 2.0 baked in 2 phantom prior successes before any real evidence. A server on its first-ever validation run, which failed, scored (2+0)/(2+1+1) = 0.5 -> 5.0/10 purely from the prior, with zero observed successes.
4. The core scoring pipeline's own scale coercion -- the one that mattered most. compute_algorithmic_score_components's return line used coerce_component_score_points(), a heuristic that only rescales values above 4 and passes values in [0, 4] through unchanged, on the (false) assumption they might already be point-scale. Every value flowing into that line is guaranteed fresh raw 0-10 output straight from a score_*() call -- never a legacy or mixed-scale value -- so the heuristic was always wrong here. Any raw score that happened to land in (0, 4] -- which R1 made more common, since it made more functions correctly return low raw values for broken servers instead of generous 6-8.5 defaults -- was stored as if it were already on the 0-4 public point scale instead of being divided by 2.5. A legitimate raw floor of 1.5/10 was displayed as 1.5/4 points (37.5%) instead of 0.6/4 (15%). This is the exact same bug class already found and fixed in scripts/backfill_score_and_capability_corrections.py during R6 (see that script's own comment) -- except this was the live scoring path every validation run has gone through, not a one-off backfill script. Fixed by switching to normalize_component_score_to_points(), which unconditionally treats input as raw 0-10, matching what every score_* function actually returns.
What a fully re-audited "everything is broken" server now looks like
Built a test (test_r1_zero_point_registry_is_an_enforced_ceiling_not_just_documentation) that runs the real validator against a mock transport where every HTTP request fails, then checks every one of the 59 score components against its registered zero point. First run (before fix #4 above) surfaced far more than the two originally reported: 25 components exceeded their documented floor. Most of that was fix #4's blast radius; after fixing the coercion bug, 11 remained. Of those, several were artifacts of the test's own construction (a synthetic server needs some title/registry_source to exercise the check-dependent metrics meaningfully, and several metrics -- discovery_metadata_score, maintenance_signal_score, adoption_signal_score, safety_transparency_score, registry_consistency_score -- have zero-point registry text that already, explicitly hedges non-zero credit from static registry metadata independent of live reachability; asserting exactly 0.0 for those on a server that legitimately has some registry metadata isn't a real claim their own documentation makes). Excluded those five from the strict assertion with that reasoning spelled out in the test.
Three more were genuine bugs, found only because the test drives the real validator (not hand-guessed check statuses) end to end:
- **
build_provenance_divergence_probe's "nothing to compare" guard never actually fired.** The conditionofficial_probe is Nonewas meant to catch "no registry evidence at all," butbuild_official_registry_probealways returns a realCheckResult(status"missing"when there's no match, neverNone) -- so the guard was dead code. A server with zero registry payload and zero server-card data fell through to "no drift fields found" ->status="ok"->score_provenance_divergence's maximum credit (9.5/10). Comparing nothing to nothing isn't confirmed consistency. Fixed the guard to check for genuine absence of evidence on either side, not aNonethat never occurs. - **
score_action_safety's core-success gate was unreachable.** The function'sprobe is Nonebranch already correctly gated onis_core_success_from_check_results-- butbuild_action_safety_probeis called unconditionally at every validation run, sochecks["action_safety_probe"]is never actuallyNone. With an empty tool inventory, the probe still computed "0 high-risk tools" and reportedstatus="ok", which read as 8.5/10 confirmed-safe credit -- the exact "no exec tools flagged" default this whole remediation plan targeted, reached through a probe that technically ran instead of one that's technically absent. Moved the gate onto the"ok"branch itself (not"warning"/"error", which are genuine negative signals regardless of confirmation status -- an existing regression test already pins"error"as unaffected by R1, and that test still passes unchanged). - **
score_session_resumedidn't recognize"skipped"as "no evidence."** This one is a regression introduced by this session's own R14 work, not a pre-existing R1 gap: R14 added an explicit"skipped"status for probes deliberately never attempted afterinitializewas confirmed unreachable.score_session_resume's existing, already-correct gate checked specifically forprobe.status == "missing", written before"skipped"existed as a status value, so a genuinely-skipped probe fell through to a flat2.0instead of the 0.0 floor. Checked every other.status == "missing"comparison in the file for the same interaction (score_installability's determinism-probe check, most others) -- only this one was actually affected, since it's the only one of R14's newly-skippable checks that had a pre-existing status-string-specific gate; the others either weren't made skippable by R14 or already gate on!= "ok"rather than== "missing".
Two items were resolved as documentation corrections, not code changes, after confirming the code's actual behavior is already consistent with sibling functions and already conservative:
least_privilege_scope_score: has an unconditional zero-tool floor of 1.5, identical in shape and reasoning todestructive_operation_safety_score/egress_ssrf_resilience_score/execution_sandbox_safety_score/secret_handling_hygiene_score's already-documented "zero-tool floor already fixed in a prior round." Its own registry entry just never got updated to say so.transport_fidelity_score: its+2.0/+1.5declared/inferred-transport credit comes fromserver.transport_type(static configuration) and URL-shape inference -- not a live check -- same reasoning asdiscovery_metadata_score's documented "independent of live reachability" exception. The registry entry previously claimed "confirmed evidence only," which the code never actually enforced for this component.
Net effect
The frozen fixture and a from-scratch fully-broken test server both now score their composite at (or very near) 0 when every live check fails, matching what /methodology has claimed since R1 shipped. Verified end to end through the real validator, not by calling individual functions in isolation:
summary_status: failing
overall score: 0.0
Every one of the 49 non-hedged score components is exactly 0.0; the hedged, registry-metadata-derived ones (discovery/maintenance/adoption/safety-transparency/registry-consistency/official-registry-presence/transport-fidelity's declared-transport credit) show small, now-honestly-documented non-zero values consistent with their own stated design.
Scope note: this affects live production scores, corpus-wide
Bug #4 (the coercion fix) is not scoped to failing servers -- it affects the stored current_score_components for every server whose next validation run produces any raw score in (0, 4] on any dimension, which is common across the corpus, not just broken servers. This is the same class of corpus-wide impact R6's backfill addressed for the R1 zero-anchoring fixes. A corpus recompute was not run as part of this pass -- consistent with this session's established practice, a real, semi-irreversible, corpus-scale production write requires an explicit decision from someone who can see the actual before/after numbers, not a default baked into a code change. New validations get the corrected code automatically as the scheduler revalidates each server; a backfill (mirroring scripts/backfill_score_and_capability_corrections.py's pattern) would be needed to correct scores for servers that won't be revalidated again soon.
Verification
PYTHONPATH=verify/src:verify/tests pytest verify/tests -- 298 passed (up from 293 at the start of this round: 4 targeted regression tests for the four root-cause bugs, plus the systemic zero-point-registry enforcement test). One pre-existing test (test_r1_action_safety_unaffected_when_probe_errored_not_missing) initially broke from an over-broad first version of the score_action_safety fix (gating the whole function instead of just the "ok" branch) -- caught by the existing test suite exactly as it's supposed to, and narrowed before landing.