MCP Verify — Audit Round 16 Architecture & Development Plan
Source: "MCP Verify — remediation requirements (audit round 16)" (11 Aug 2026 audit, target build v1.0.546). This plan is grounded in direct reads of the current code (file:line citations throughout) plus four parallel research passes over validation/service.py, insights.py, main.py, db/repository.py, and the registry ingestion client — not the audit's assumptions alone. Several places below where the grounding changed the picture from what the audit assumed are called out explicitly, same as prior rounds.
Note on naming: this is audit round 16, the auditor's own numbering. It is unrelated to this repo's docs/mcp-verify-round-16-implementation.md (a different, already-shipped round in this project's own sequential doc numbering). Filed under a distinct name to avoid collision.
---
0. A13 resolved directly — version audit result
The audit's headline framing ("methodology page was serving v1.0.503") is not what's happening. Live-checked at v1.0.546 across 15 pages (/, /search, /methodology, /docs/scoring-specification, /methodology/changelog, /pricing, /trust-index, /about, /security, /changelog, /status, /rankings, /badges, /demo, /ecosystem, /github-action, /research/mcp-security-findings, /compare): every page's footer reports v1.0.546, from one source of truth (VERIFY_SITE_VERSION, main.py:1221), applied uniformly by inject_site_footer() (main.py:21557-21564) inside _html_response() (main.py:2736-2745, specifically the normalize_public_site_version_markup() call at 2745/21558) on every response. /methodology itself (main.py:8249-8251) renders fresh per-request — it isn't behind _get_cached_route_value at all, so there's no cache-staleness vector here either.
What's actually there: two sentences of rendered body copy, both on /docs/scoring-specification (which /methodology cross-links, and which the auditor evidently read as part of the "methodology" page family):
main.py:28109(insiderender_scoring_specification_page's "Subscore zero points" section): "Per the post-1.0.503 remediation plan, that value must be 0 unless..."main.py:1991(aSCORE_COMPONENT_ZERO_POINT"Why" entry, rendered viarender_methodology_zero_point_rows()into the same page's table): "...the post-1.0.503 review named as the clearest floor contributor..."
Both are naming a historical planning effort by the version it started at, not reporting the page's own version. It reads exactly like a stale version stamp to any scanner — human or automated — that doesn't know "post-1.0.503 remediation plan" is a proper noun.
Required fix (A13):
- Reword both strings to drop the version-number-shaped token, e.g. "Per the mid-2026 scoring remediation review, ..." / "...the mid-2026 scoring review named as the clearest floor contributor...". Pick one canonical name for the effort and use it in both places (this file's own audit doc calls it the same thing three times as "the round 15 standing constraints" heritage — consistency matters here for the same reason A9-A12's fixes matter below: one fact, one label).
- Add a regression test asserting no rendered page body (excluding
/changelogand/methodology/changelog, which legitimately list historical version numbers in a dated table, and excluding the site-footer's own current-version span) contains a token matching\bv?1\.0\.\d+\bother than the currentVERIFY_SITE_VERSION. This is a cheap, durable guard against exactly this class of finding recurring — a naivegrep-style scan is precisely how a future audit (human or automated) will re-check this. - Grep sweep confirms these are the only two live occurrences (
grep -c 'post-1.0.503'on rendered/docs/scoring-specificationoutput = 2; no other route contains any1.0.5xx-shaped token besides the current version). No further page-by-page audit needed beyond the regression test in (2).
Effort: trivial, ungated, do any time — bundle with A1 or ship standalone in the next release. Not sequenced below since it doesn't touch scoring/verdict logic.
---
1. Standing constraints (§1)
I could not locate a verbatim, checked-in copy of round 15 §8. I grepped every docs/mcp-verify-round-*.md file in the repo for the rule text and for a numbered "Rule 1/2/3..." block; several docs reference "the standing constraint" (e.g. docs/mcp-verify-round-9-implementation.md:33, docs/mcp-verify-round-13-implementation.md:211) but none contain the authoritative list itself — it appears to have lived only in an earlier conversation, not as a repo artifact. Per the audit doc's own instruction, this must be pasted in verbatim by the document's owner before this plan is treated as final. Until then, this plan proceeds on the six rules the audit doc itself reconstructed (marked there as "indicative only"), because four of them are directly load-bearing for the fixes below:
- Unknown is never a pass and never a conflict. (governs A4-A8, A9)
- One question, one computation. (governs A9, A10, A11, A12 — see §2 below, this is the recurring theme)
- A signal read in place of another signal is not a fix. (governs A1's output-schema branch, A3's alias/probe conflation)
- A validation artifact is not evidence about the server under test. (governs A2)
- Fixing A must not un-fix B. (governs sequencing — A4-A8 as one workstream, A9 before A10)
- Never publish an adverse claim against a party the evidence does not implicate. (governs A1, A2, A10a)
The four rules the audit doc didn't reconstruct are unknown to me and could change scope on any item below. Recommend blocking sign-off on this plan until the verbatim §8 is supplied, even though I've proceeded far enough to ground every fix in real code so no time is lost either way.
---
2. Cross-cutting architectural theme (found across the research, not called out as one thing in the audit)
A1, A9, A10a, A10b, A11, and A12 are not six unrelated bugs. They're six instances of the same shape: two renderers compute "the same" fact from overlapping-but-different inputs, and nobody enforces they agree.
| Item | Renderer A | Renderer B | Divergent input | |---|---|---|---| | A9 | Client-compat status badge | Client-compat reason/checklist | profile["compatibility"] (narrow criteria) vs. _full_client_blocker_list (criteria ∪ policy gates) | | A10a | Remediation row title (publisher-toned) | Remediation row playbook (registry-toned) | no addressee field exists at all — title is generic, playbook prose is honest | | A10b | critical_alerts counter | Remediation row severity | active_alerts list vs. three independent severity sources (static REMEDIATION_RULES table, hardcoded "medium", and alert["severity"]) | | A11 | Main score chip | Validation timeline / incident feed | public_display_score() (suppressed) vs. raw compute_current_score() (unsuppressed) | | A12 | Ranking pool population | Homepage coverage stats | fresh per-row public_display_score() recompute + 4-condition filter vs. raw possibly-stale Server.current_score/current_status columns + no filter | | A1 | Divergence count | Divergence wording | len(tool_divergences) (accurate) vs. always-"parameters" wording (ignores which of 8 compared_dimensions fired) |
This repo has already fixed this exact shape once (per insights.py:1503-1524, R77.1: unified two blocker-consuming reason surfaces) — the fix pattern is proven, it just wasn't applied to the status computation alongside the reason computation in that same pass, which is exactly A9. I'd treat "does a second renderer exist that reads a narrower or differently-filtered version of the same input" as a standing question for every fix below, not just this round's six.
---
3. Reproduction fixtures
Same two as the audit: awesome-malinoto/tracepass-mcp-server (healthy path) and ryanatindago/playwright-mcp-example (failing/unreachable path). Freeze both (a SELECT snapshot of Server + latest ValidationRun rows, or a pytest fixture seeded from their current field values) before starting, since P0/P1 acceptance criteria are diffs against these two pages' current rendered output.
---
P0 — Adverse claims unsupported by evidence
A1 — schema_divergence_probe
Correction to the audit's framing: card-shape detection already exists (validation/service.py:3504-3516 — dict entries with a name key go to card_tools, bare strings go to card_tool_names_only). The parameter-dimension logic already correctly routes name-only-card tools to tools_with_omitted_schema (service.py:3595-3606), not to tool_divergences. This is a narrower bug than "the probe has no shape detection" — it's one unguarded branch.
Root cause. The output-schema check (service.py:3633-3637) doesn't reuse the shape guard the parameter checks already have:
card_output_declared = bool((card_tool or {}).get("outputSchema")) # card_tool is None for name-only cards -> always False
Since card_tool is None for every tool on a name-array card, this unconditionally produces output_schema_card_declared: False for every tool, and any tool whose live schema declares an output schema shows as "diverging" — indistinguishable from a real card that declares no output schema on purpose.
Fix:
service.py:3633-3637— gate the output-schema comparison the same way the parameter dimensions already are: ifcard_tool is None(card carries no per-tool schema object at all), route totools_with_omitted_schema/ markoutput_schema_presencenot-assessed for that tool, don't populatetool_divergenceson this dimension.- Alert wording,
insights.py:3637("Server card and live tool schemas disagree") +insights.py:3652-3656(the count/list message) — the message always says "diverge on parameters" regardless of which of the 8compared_dimensions(service.py:3522-3531:server_name,server_version,declared_vs_observed_auth,tool_membership,parameter_names,required_parameters,parameter_types,output_schema_presence) actually fired. Build the wording from the set of dimensions that actually diverged for the affected tools (e.g. "diverge on required parameters" vs. "diverge on parameter types" vs. a generic "diverge on schema" only when genuinely mixed) — reserve "invalid-call" language (rule per audit: reserved for input-schema divergence) for whenparameter_names/required_parameters/parameter_typesare among the fired dimensions. - Minor:
insights.py:3652-3656's truncation concatenates"..."directly onto the last name then a period ("....") — insert a separating comma/space, e.g.", …"before the trailing period, when truncated. - Remediation text,
insights.py:4267(REMEDIATION_RULES["schema_divergence_probe"], "...match exactly") — same wording-selection fix as (2) applies here since it's a fixed template regardless of dimension.
Tests: a synthetic fixture with a name-only card + live output-schema tools must show output_schema_presence as not-assessed, zero tool_divergences from it, and no alert. A second synthetic fixture with a genuine object-shaped card (has outputSchema fields) and a real live mismatch must still produce the alert, correctly worded to the dimension that diverged, with the full tool list (no silent count/list disagreement).
Acceptance / catalog-wide retraction (open question 4). Run the corrected probe across the catalog once merged; report how many servers lose the false alert. Given card-shape detection already exists and the bug is isolated to the output-schema branch, this is likely every catalog server using a name-array card (probably a meaningful fraction — worth the count either way). Recommend yes on a changelog entry + affected-publisher notification once the number is known, consistent with rule 6 — this fixed a real published false claim, and the audit doc's own instinct here is right. Add the count and disposition to the CHANGELOG entry for this fix, same pattern used for the 1.0.533/1.0.534 registry-endpoint fixes.
---
A2 + A3 — one root cause (ingestion), plus two independent downstream bugs (status propagation, and rendering-layer conflation)
Confirmed: A2 and A3 are the "same bug, two symptom sites" the audit suspected them to be, but there are actually three separable fixes here, not one. Doing only the ingestion fix would silently look like it solved A2 (the specific reproduction case), while the status-propagation gap and the rendering conflation remain live for the next server that hits either by a different path.
3a. Ingestion: Glama listing URLs get persisted as remote_url
GlamaRegistryClient.normalize_directory_summary → extract_first_mcp_endpoint_from_text → _extract_mcp_endpoint_urls (ingest/registry_client.py:1424-1631). The URL-extraction regex (_MCP_ENDPOINT_URL_RE, 1517-1520) matches any URL containing /mcp or /sse; the scorer excludes /badge/mcp/ and penalizes lobehub.com/mcpservers.org/server.smithery.ai (1620-1623) but has **no exclusion for glama.ai itself or the listing-page path shape** /mcp/servers/<id>. Confirmed live: https://glama.ai/mcp/servers/<id> matches the regex and scores with no penalty, i.e. can be persisted as remote_url/an alias endpoint even though it's a directory listing page, not a transport endpoint.
Fix: add glama.ai + /mcp/servers/ path shape to the penalty/exclusion list at 1620-1623, same mechanism already used for the other three known-non-endpoint hosts. This is a one-line addition to an existing, proven pattern — low risk.
3b. Status propagation: the registry-vs-server distinction only changes alert text, not status or verdict
The existing 1.0.533/1.0.534 fix (insights.py:3529-3553, _non_mcp_endpoint_content_type() at 3467-3492) only swaps which alert object gets appended (medium/registry-directed vs. critical/publisher-directed text) — confirmed it never touches ValidationRun.summary_status, Server.current_status, or the verdict. compute_summary_status (validation/service.py:2222-2276) has no branch distinguishing "never reached the server" from any other initialize failure — a non-MCP endpoint and a genuinely-broken server both hit the same "failing" branch at 2250-2251.
Also confirmed: the two verdict implementations the audit worried about duplicating are already unified — build_production_readiness (insights.py:3200-3205) explicitly calls infer_server_verdict (db/repository.py:1456-1486) and says so in its own comment. (This also confirms the stale steady-wibbling-river.md plan file surfaced in this session's context — an old "Round 3" plan whose TASK-31 proposed exactly this consolidation — is obsolete; it already shipped. No action needed there.) But infer_server_verdict checks current_status == "failing" (1478-1479) before tool_count <= 0 (1480-1481), so for this exact case (failing status, zero tools) it returns "failing", never reaching the methodology's own stated "zero tools → metadata_only regardless" rule. That rule is enforced for never-validated servers but not for validated-and-failed-with-zero-tools servers — confirmed deliberate per an existing R71.1 comment (repository.py:1464-1475) that intentionally stopped conflating "validated and failed" with "metadata_only" in general. That general decision is correct (a server that fails on its own endpoint should read Failing, not the softer Metadata only) — it's specifically the "never reached the server at all" sub-case that needs to be pulled out from under it.
Fan-out — which of the six downstream surfaces would auto-inherit a fix, and which need explicit changes:
| Surface | Reads canonical verdict/status? | Citation | |---|---|---| | Profile headline/status chip | Yes — render_production_readiness_badge(production["label"]) | main.py:26052 | | /report (server + production_readiness) | Yes, for headline; No for embedded timeline/incident-feed (see A11) | main.py:14604 | | Client-compat verdicts (ChatGPT/Claude) | No — derives from checks via client_profiles, not current_status/verdict | insights.py:1577-1616 | | Benchmark task labels | No — reads checks["initialize"]/checks["tools_list"] directly | insights.py:1357-1409 | | Incident feed text | No — literal current.summary_status | insights.py:3369-3376 | | Badge SVG | No — badge_server.current_status directly | main.py:12774-12784 | | /policy | Reads detail.current_score/status via the detail model — needs direct check, not yet traced to the same depth as the others | main.py:11234+ | | /trust-summary (both paths) | No — compare_service.py:141-198 and main.py:10929-10948 both check server.current_status directly | see citations |
Decision required from the owner (open question 1, now grounded): two real options, with a real cost difference now that the fan-out is known.
- Option A — fold into Metadata only. Cheapest at the verdict layer (
infer_server_verdictgets a reordered check: "never reached own endpoint" routes tometadata_onlybefore thefailingcheck fires), but does not fix the 5 surfaces that key offcurrent_status/rawchecksdirectly (client-compat, benchmarks, incident feed, badge, both trust-summary paths) — those would still read "failing" and render adversely unlesscompute_summary_statusitself also changes. So Option A still requires acompute_summary_statuschange; it isn't actually cheaper once you count that. - Option B — new distinct state. Same
compute_summary_statuschange is required either way (see above), so the marginal cost of Option B over Option A is small: introduce acurrent_statusvalue (e.g."unreachable"or reuse"unverified", which already exists perservice.py:2245/2273for a different case — check whether that's semantically close enough to reuse rather than inventing a fifth value) instead of routing straight to a verdict label. This is more honest per rule 6 ("never publish an adverse claim... against a party the evidence does not implicate" — Metadata only still implies "we looked and there wasn't much," when the truth is "we never looked at all").
Recommendation: Option B, reusing the existing "unverified" status value if its current semantics ("tool_count <= 0" fallback at service.py:2273-2274) are compatible, rather than inventing a fifth status — but this is exactly the kind of call the audit doc flags as owner-only, and the fan-out table above should settle it rather than my preference. Either way, the fix touches: compute_summary_status (new/reused branch keyed on _non_mcp_endpoint_content_type(), checked before the generic initialize-failure branch), infer_server_verdict (route this status to metadata_only, and reorder so it's checked before current_status == "failing"), and the 5 independently-re-deriving surfaces above (each needs one added condition checking for the new/reused status, since none of them will auto-inherit a verdict-layer fix).
3c. Rendering layer: provenance-divergence probe status gets overwritten by unrelated alias data
build_provenance_divergence_detail (insights.py:575-629) takes the raw probe's drift_fields (empty on tracepass) and **unions in alias_disagreements** from build_alias_consolidation's source_disagreements (insights.py:2221-2276, cross-alias diff — a genuinely different signal), escalating status to error/warning even when the raw probe itself was not_assessed with drift_fields: [] (insights.py:611-617). Two other rendering sites — score-breakdown "not assessed" list (main.py:25298-25309) and the check-table ?Not_Assessed badge (main.py:31734-31760) — both iterate validation.checks directly and never see the alias data, so they correctly show not-assessed. This is why the same probe renders three different states on one page: two sites read the raw probe, two read a merged structure that silently overrides it.
Fix: stop merging alias disagreements into the provenance-divergence probe's own status. Render them as a clearly separate, independently-labeled block — e.g. "Registry alias disagreements" as its own subsection with its own status line, next to (not inside) "Provenance & registry divergence." The remediation "Resolve conflicting endpoints across registry aliases" (insights.py:3856-3866) already reads source_disagreements["remote_url"] directly (not the probe status) — leave that gating as-is once the alias data itself is clean per 3a; the fix here is presentational/structural (don't let one signal impersonate another — rule 3), not a remediation-gating change.
Sequencing within A2+A3: 3a (ingestion) first — it makes the specific reproduction case disappear and is the safest, most isolated change. 3b (status propagation) and 3c (rendering conflation) are independent of 3a and of each other; do 3c anytime (small, presentational), but do 3b as its own reviewed change given it touches compute_summary_status, infer_server_verdict, and five independent rendering call sites — this is the largest single fix in the whole round and deserves its own PR.
Tests (open question 5). One live fixture isn't enough — 3a/3b/3c are three different layers. Add: a unit test on _extract_mcp_endpoint_urls/normalize_directory_summary directly (registry_client tests) asserting a Glama listing-page URL is excluded/penalized regardless of what live corpus data currently looks like; a compute_summary_status/infer_server_verdict unit test for the new/reused status with a synthetic "initialize raised NonMCPContentTypeError, zero tools" run; and the 5 fan-out surfaces each get a direct assertion (not just an end-to-end page-render check) that they read the new status, so a future refactor of any one doesn't silently regress the others (rule 5).
---
P1 — Unknown resolving to pass
Correction to the audit's framing: these are not one shared root cause. Research found three distinct bug shapes:
- An explicit three-state check whose accepted set wrongly includes
"missing"(A4, A5). - A raw, unfiltered count with no status check at all (A6, contributes to A8).
- An existing "was this run's tool surface actually observed" filter (
validation_has_observed_tool_surface, already defined and used atinsights.py:3013/3039) that was applied to one function but not its sibling (A7, A8).
Recommend introducing one module-level constant, e.g. PASSING_PROBE_STATUSES = {"ok", "warning"}, and auditing every inline {"ok","warning","missing"}-shaped literal to use it — this doesn't fix any of the five by itself, but it's the guard that stops #1 from recurring at a sixth site next round.
A4 — Client-compat fixtures show "Passes" on missing
insights.py:1230-1231, build_compatibility_fixtures:
request_association: explicit set check includes"missing"in the passing set — drop it.frozen_tool_snapshot_refresh: worse — never readsconnector_replay["status"]at all, only the derived booleanwould_break_after_refresh(insights.py:645/668), which defaultsFalsewhen the probe is missing. Gate onconnector_replay.get("status") == "ok"in addition to the boolean.
A5 — connector_publishability_probe criteria
validation/service.py:3716-3722, build_connector_publishability_probe. Four of six criteria (session_ready, step_up_ready, connector_replay_ready, request_association_ready) use probe is None or probe.status in {"ok","warning","missing"} — both the is None short-circuit and the explicit "missing" need removing. transport_ready (3719) is inconsistent with the rest (keeps the is None loophole but not "missing") — fix all six uniformly to probe is not None and probe.status in {"ok","warning"}.
A6 — "Live checks captured: N"
insights.py:3104, build_evidence_confidence: live_check_count = len(checks) — a bare len(dict), no status filter of any kind. Fix: sum(1 for c in checks.values() if isinstance(c, dict) and c.get("status") in {"ok","warning"}).
A7 — 7d/30d success ratios
The fix pattern already exists and already shipped for one function (build_history_summary, insights.py:3013, per CHANGELOG "healthy-ratio and evidence-confidence history now exclude unobserved/unverified tool-surface runs") but was not applied to its sibling, build_public_server_reputation (insights.py:1265-1281) — last_7d/last_30d are windowed by timestamp only, no validation_has_observed_tool_surface filter, before computing success_7d/success_30d (1280-1281) → validation_success_ratio_7d/_30d (1315-1316). Fix: apply the same filter (already defined at 3039) to 1270-1279. This is the cheapest, highest-confidence fix in the whole round — literally finishing a fix that's already proven correct elsewhere in the same file.
A8 — recent-validations counter vs. history table; 0 validations still "Medium"
Two separate fixes:
main.py:24822-24823(history-table row loop) iterates rawvalidationswith no filter, while the counter it should agree with (insights.py:3122-3123,history_depth = min(len(observed_validations), 20)whereobserved_validationsis filtered) does filter. Either filter the table loop to match, or keep all rows visible but label unfiltered ones distinctly (e.g. "not fetched") so the counted number and the displayed number reconcile by design rather than by coincidence.insights.py:3125-3146,build_evidence_confidence: the confidence score is purely additive (age +history_depth*1.25+live_check_count*2.0+ status bonus) with no gate onhistory_depth == 0. A run with zero observed-tool-surface history can still sum into the 50-74 "medium" band. Add a hard cap/floor:history_depth == 0forceslabel == "low"(or a distinct "not assessed" label, consistent with A4-A8's shared theme) regardless of the other additive terms.
---
P2 — Verdict contradicting its own rationale
A9 — Client-compat verdict ignores its own blocker list
Root cause (see §2's cross-cutting table): status (insights.py:1610-1616, drives the "✓ Client-compatible" badge) is computed only from profile["compatibility"], itself from a narrow criteria checklist (_make_profile, insights.py:4465-4503 — transport/initialize/tools_list/oauth/dcr only). The reason text and remediation checklist (insights.py:1495-1574, _full_client_blocker_list) are computed from a broader union: narrow criteria ∪ policy gates (write_actions_present, admin_refresh_required, safe_for_company_knowledge, safe_for_messages_api_remote_mcp, request_association, transport_compliance — insights.py:1801). A server can pass the narrow set (badge: compatible) while failing policy gates (reason text: real blockers) — exactly the audit's reproduction.
Fix: derive status from the same broad blocker list the reason/checklist already use, not from profile["compatibility"] alone — status = "blocked" if _full_client_blocker_list(...) else "ready" (with a "partial" tier if the codebase wants to distinguish narrow-only failures from broad failures; check whether that granularity is used anywhere downstream before adding it). This is the same fix shape as R77.1 (insights.py:1503-1524, which unified two reason surfaces but not the status computation alongside them) — finish that pass.
Test: the assertion the audit suggests is the right one — "a page can never show a non-empty blocker list alongside a compatible verdict or an n-of-n score" — write it as a property test over both fixtures' rendered data, not just a fixture-specific check.
A10a + A10b — must ship together (per the audit's own note; confirmed necessary by the research)
A10a. No addressee field exists anywhere in the remediation schema today (build_remediations, insights.py:3771-3876 — items carry code, severity, title, why, action, playbook, maintainer_context, nothing else). The only place responsibility is expressed is prose buried in the playbook body (insights.py:4321/4347, guarded by a comment at 4310-4317 that already explicitly calls out "the registry entry is the party responsible, not the server operator" — the intent is already in the code, it just never became structured data). Add an addressee: Literal["publisher", "registry", "verify"] field, populate it per-row (check-based rows default to publisher except registry-detection-derived ones like the non-MCP-endpoint case, which should be registry; alert-based rows should inherit the alert's own addressee — natural to add alongside A2's registry_endpoint_invalid alert, which is conceptually already registry-addressed). Filter the publisher-facing "Recommended actions" summary and the score-facing remediation table to exclude non-publisher rows (or render them in a visually separate section).
A10b. Three independent severity sources feed the remediation table (static per-check REMEDIATION_RULES severity, insights.py:138+; hardcoded "medium" for score-decomposition rows, insights.py:3825; alert["severity"] for alert-based rows, insights.py:3839), while the page's own "High/critical-severity alerts: N" counter (main.py:27302-27328, insights.py:3230) counts only active_alerts. A Critical/High remediation row can exist with zero matching entry in the counted alert list.
Recommendation (open question 2): make active_alerts the single source of severity truth, per rule 2 and rule 6 (adverse-severity claims should always be counted, not hidden in a table nobody's summing). Concretely: every check-based and score-decomposition remediation row that would warrant Critical/High severity should either (a) read its severity from a matching alert if one exists, or (b) if none exists, that's itself the bug to fix — synthesize the alert from the same check-status data the remediation already keys off, so the two collapse into one computation rather than maintaining a second, disconnected severity table. This is more work than "cap remediation by alert severity" but is the direction that doesn't quietly downgrade real findings to make the counter agree.
Must ship together: reclassifying a registry-addressed row (A10a) changes what severity is even appropriate for it (a registry-directed medium finding shouldn't inherit a publisher-directed critical severity computed for the wrong party) — confirmed by tracing both through the same rows (insights.py:3771-3876) in this research pass, the audit's warning here holds.
A11 — Score suppressed on one surface, published on two others
Canonical suppression is public_display_score() (validation/service.py:2499-2538, "the one score any public surface may show" per its own docstring). Two functions bypass it and call compute_current_score() directly:
build_validation_timeline(insights.py:2172-2218, score computed at2203-2207).build_server_incident_feed(insights.py:3352-3377, score computed at3365-3368, literally interpolated intof"Score {score:.1f} with status {current.summary_status}."at3375).
Both feed ServerDetailResponse.validation_timeline/.incident_feed (main.py:1280, 1282), which flow unsuppressed into /v1/.../report's embedded "server": detail.model_dump() (main.py:14585) — so the JSON leak is a direct consequence of these two functions, not a separate endpoint-level bug. /policy (main.py:18314) and /badge (main.py:12538, explicit public_display_score() call) are already correct. /trust-summary doesn't expose a raw numeric score field at all (confirmed via compare_service.py:286-359) — not a bug surface.
Fix: both build_validation_timeline and build_server_incident_feed must run their computed score through public_display_score's suppression logic (either call it directly with the right inputs, or accept the already-suppressed value from the caller and thread it through) before rendering. This is a two-function fix; /report's JSON self-corrects once these land, no separate endpoint change needed.
Test: for a status=="failing" fixture, assert validation_timeline entries and incident_feed text never contain a numeric score — not just that the main chip says "n/a."
A12 — Ranking pool exceeds its own stated population
Ranking pool (db/repository.py:322-405, feeds "#N of M scored public servers") filters on current_score is not None, tool_count > 0, current_status != "failing", last_validated_at is not None (347-353), then recomputes public_display_score() per candidate row and drops None results (381-383) — which also excludes opted-out servers. Homepage coverage stats (build_public_coverage_stats, main.py:19713-19792) instead count raw Server.current_score/current_status columns directly, with only a smithery-blocklist filter (db/repository.py:826-841) — no tool_count, status, or opted-out filtering, and no fresh score recompute, so a stale current_score column value still counts.
Recommendation (open question 3): the ranking pool's definition is the more rigorous one (fresh recompute + 4-condition filter matching its own stated cohort text at db/repository.py:401-404). Recommend the homepage's scored_servers/handshake_and_tools_ok/healthy_servers counts adopt the same filter + recompute discipline — likely by extracting the ranking pool's filter+recompute as a shared helper both call, rather than loosening the ranking pool to match a looser, possibly-stale homepage count. Flagged as owner sign-off per the audit's own framing; this is a recommendation, not a unilateral call.
---
P3 — Classifier under-call (blocked on R11)
No fixes proposed here, per the audit's explicit scope guard ("stop and escalate" if a fix would change classifier conclusions). Current-state grounding only, so the eventual R11 work has real starting citations instead of needing to re-discover them:
A14. Risk/capability classification operates on tool-level shape and never descends into a tool's action enum values (confirmed: no code path in the classifier reads enum members of an action-typed parameter). The tracepass_passports/tracepass_products archive/archive_by_serial case is a real instance of this blind spot, structurally common to any action-dispatch-pattern tool. Separately-worth-filing: the server's own initialize instructions reference a suspend_passport tool that doesn't exist — a genuine agent-facing defect on the server, unrelated to Verify's classifier, worth its own low-effort finding (no R11 dependency — this one could be surfaced today as a simple "tool referenced in initialize instructions but not present in tools/list" check, if the team wants a quick independent win while A14 proper waits on R11).
A15. annotation_conflict_tools (validation/service.py:6611-6644, populated by detect_tool_annotation_conflicts, 6497-6523) only fires when readOnlyHint is True and schema evidence contradicts it (6511-6522) — one cell of the audit's 3-way split, conflating "server misrepresents itself" with "detector defect" under one label. No path exists for annotation-false-but-detection-true, and no taxonomy-scope-gap case. The field is computed but never read anywhere else (zero references in main.py/insights.py) — currently dead data, not even surfaced. Post-R11 work here is a genuine 3-way classification (data model + UI), not a small patch.
---
4. Sequencing (updated from the audit's proposal with what the research changed)
- A13 — trivial, ungated, do whenever (already fully specified above; not a blocker for anything else, contrary to the audit's assumption that it might invalidate other acceptance checks — the site version infrastructure is already correct).
- A2 §3a (ingestion fix) first, independently — smallest, safest, eliminates the concrete reproduction case immediately.
- A1 — fully self-contained, one unguarded branch, no dependency on anything else. Ship early given it's currently emitting a false High against a real fraction of the catalog (rule 6 urgency).
- A2 §3b (status propagation) — the largest single change in the round (touches
compute_summary_status,infer_server_verdict, 5 independent rendering surfaces). Needs the owner decision (open question 1) before starting. Do this as its own reviewed unit, not bundled with §3a/§3c. - A2 §3c (rendering conflation) — independent, small, do anytime after §3a (so the "conflicting endpoints" example in the fixture is already clean when this ships, making the before/after easier to verify).
- A4–A8 as one workstream, per the audit's own note, but now with three known sub-shapes rather than one — do A7 first (cheapest, proven pattern, zero design risk), then A4/A5 (same shape as each other), then A6, then A8 (depends conceptually on A6's fix being in place, since A8's confidence-score gate interacts with
live_check_count). - A9, then A10a + A10b together (confirmed necessary, not just cautious).
- A11, A12 — independent of everything above, can run in parallel with 6-7 if capacity allows; A12 needs owner sign-off (open question 3) before the homepage-side change, but the two
insights.pyfunction fixes in A11 don't need any decision and can ship immediately. - A14, A15 — blocked on R11, no earlier action beyond the low-effort "tool referenced in instructions but absent from tools/list" check called out under A14.
5. Out of scope (unchanged from the audit)
Same two guards apply: nothing here proceeds without the R11 labeled set beyond what's explicitly noted for A14/A15, and no fix in this plan should change what the classifier concludes — every item above is status-handling, rendering, addressee/severity plumbing, or ingestion filtering, not a decision-rule change. If implementation surfaces a case where a proposed fix would change a classifier's conclusion rather than how an existing conclusion is read/displayed/routed, stop and escalate before proceeding, per the audit's own instruction.
6. Open questions requiring owner sign-off before implementation starts
- A2 — fold "unreachable registry endpoint" into Metadata only, or a new/reused distinct status? Both require the same
compute_summary_statuschange and the same 5-surface fan-out fix regardless of which label wins — recommend reusing the existing"unverified"status if its semantics fit, over inventing a fifth value, but this is the owner's call. - A10b — severity ceiling from alerts, or alerts raised to match remediation? Recommend making
active_alertsthe single source (synthesize missing alerts rather than maintaining a second severity table), per rule 6. - A12 — which population is correct? Recommend the ranking pool's stricter, freshly-recomputed definition; homepage stats should adopt it.
- A1 — catalog-wide retraction count + changelog/notification once the fix lands and the probe reruns. Recommend yes.
- Standing constraints §1 — verbatim round-15 §8 text could not be located in the repo; needs to be supplied before this plan is final, since 4 of the 10 rules are unknown to me and could affect scope on any item above.
---
7. Test plan summary
PYTHONPATH=verify/src pytest verify/tests -qafter every item; new/extended files:test_validation_service.py(A1, A2 status branch, A4-A8 probe-status handling), a new or extended registry-ingestion test module (A2 §3a),test_api.py(A9 property test, A11 suppression-on-timeline/incident-feed, A13 version-token regression test),insights-focused tests wherever they currently live (A3 rendering separation, A10 addressee/severity, A12 population parity).python3 scripts/export_openapi.py --checkif any response schema changes (A10a's newaddresseefield is the most likely trigger).- Live verification per this session's established discipline: don't trust a green test suite alone for anything touching cached routes or prewarmed pages — curl the two frozen fixtures post-deploy and diff against the frozen pre-fix snapshot for each P0/P1 item's specific acceptance criterion.