Sentinel Signal

v1.0.642 — Admission-shed metric + detail-page fast-miss rollout

Source: docs/mcp-verify-v1.0.642-admission-shed-counter-and-fast-miss-rollout.md

Document Content

v1.0.642 — Admission-shed metric + detail-page fast-miss rollout

Context

Follow-up to v1.0.640 (docs/mcp-verify-v1.0.639-server-detail-fast-fallback.md), which shipped fast-fallback caching for /servers/{namespace}/{name} but left the detail-page fast-miss flag off and the Tuesday DB-tuning regression un-root-caused. This round root-caused the regression and turned the flag on.

What the v1.0.639 doc got wrong

That doc's leading hypothesis — idle_in_transaction_session_timeout=60000ms (new in the ecdd0cd deploy, web-only) causing new 5xx — is disproven. Loki retains logs across container restarts (the doc's "log durability" concern was itself mistaken; we hadn't actually checked Loki). Direct queries found:

  • Zero idle in transaction kill lines anywhere from just before the
  • original incident through the time of this investigation.

  • Zero real application exceptions/tracebacks in verify-web since the
  • hot_unclaimed_servers index fix (v1.0.637, 2026-08-15 17:10 UTC) — the canceling statement due to statement timeout bursts that dominated logs before that fix (18 in one 10-minute window) dropped to ~18 total across the following 35 hours.

What's actually happening

/compare, /v1/compare, and /servers/{namespace}/{name} return explicit 503 Service Unavailable ("server_trust_capacity_busy") from server_detail_admission_middleware (main.py:2677-2733) — a pre-existing backpressure mechanism, not a crash. It gates every request to those paths through one of two asyncio.Semaphores (human/machine, split by traffic class) and sheds with 503 if no slot frees up within server_detail_admission_timeout_seconds. Confirmed live: concurrency is 1 human slot, 1 machine slot, site-wide, timeout 10s (MCP_VERIFY_SERVER_DETAIL_HUMAN_CONCURRENCY / _MACHINE_CONCURRENCY / _ADMISSION_TIMEOUT_SECONDS, all at their code defaults — not overridden in the live prod.env, which predates these settings existing). Loki: zero /compare 503s in the 44 hours before the ecdd0cd deploy, 50 in a recent 6-hour sample, still recurring at last check.

Code inspection ruled out /compare//v1/compare themselves as the direct cost driver: both already have always-on (no settings flag) miss_fallback_builders (build_fast_compare_page / build_fast_compare_payload, main.py:13932-13966, 13609-13637) that take over on any page-level cache miss, so a cold /compare request rarely pays the full per-server _get_cached_server_detail_response cost. The actual mechanism is collateral contention: at concurrency=1 per class, one slow request — e.g. a real detail-page cold build, still 4-30s before the fast-miss flag was on — holds the only slot in its class for its entire duration, and every other concurrent request in that class (including otherwise-cheap /compare lookups) queues behind it and times out at 10s. The new recurring DATABASE_MAINTENANCE_JOB (same deploy, unconditionally enqueued every 300s regardless of the retention flag, scanning/updating servers and analytics_events) is the leading candidate for what tipped this from rare to routine — combined with the newly-enforced 30s statement timeout (previously unlimited), background contention now more often cancels a web query outright instead of just slowing it down.

What shipped

  1. **mcp_verify_server_detail_capacity_shed_total{traffic_class}** — new
  2. Prometheus counter, incremented at the admission middleware's except TimeoutError: branch (main.py, next to the counter definitions near slow_requests_counter). This closes the biggest visibility gap this investigation hit: nothing previously distinguished "503 from admission shedding" from any other 5xx in metrics — every check this round required grepping raw Loki log lines instead.

  3. **MCP_VERIFY_SERVER_DETAIL_FAST_MISS_ENABLED=true** in
  4. deploy/ionos/prod.env.example and the live prod.env, alongside the three sibling flags that were already on. This is the direct fix for the collateral-contention mechanism above: it makes a cold detail-page miss a fast-fallback build instead of a full build_server_insights pass, so the shared admission slot is rarely held for more than about a second by any detail-page request — starving /compare and other detail-page requests far less.

Not done this round

  • Extending fast-fallback into /compare//v1/compare's per-server
  • builds specifically — investigated and found largely redundant with the always-on miss_fallback_builders those routes already have (see above). Not implemented.

  • Raising MCP_VERIFY_SERVER_DETAIL_HUMAN_CONCURRENCY /
  • _MACHINE_CONCURRENCY above 1. The plan explicitly treats this as a monitored follow-up, not a first move — the fast-miss flag needs to prove out in production first, since the admission control exists as deliberate DB protection.

  • Confirming the DATABASE_MAINTENANCE_JOB correlation directly (e.g. a
  • 5-minute rhythm in the shed counter). The counter shipped this round; correlating it against the job's schedule is the next diagnostic step, now possible for the first time.

Verification

  • PYTHONPATH=verify/src:verify/tests python -m pytest verify/tests -q — 821
  • passed. Extended test_server_detail_admission_prevents_crawler_starvation to assert the new counter appears in /metrics with traffic_class="crawler" after a shed event, rather than adding a new test harness.

  • python -m pytest -m unit -q.
  • Live, after deploy: confirm mcp_verify_server_detail_capacity_shed_total
  • appears in /metrics; watch it, the 5xx ratio, and the slow-request ratio for at least an hour against the pre-deploy baseline; spot-check a previously-slow, cold server page for a fast-fallback response and confirm the real content replaces it shortly after.