Sentinel Signal

Verify: observed-attention cache stampede + export back-off

Source: docs/mcp-verify-observed-attention-stampede-fix.md

Document Content

Verify: observed-attention cache stampede + export back-off

Targeted production reliability fix for the 2026-09-28 06:02–06:04 UTC stall. Not a feature release.

Incident

  • The nightly Trust Data Feed snapshot export (sentinel-signal-data-snapshot.timer,
  • 03:15 UTC) fetches /v1/intelligence/servers/{ns}/{name} for ~97k servers, 10 threads, 10 s client timeout, through the public URL.

  • Every Intelligence call builds a policy payload → build_observed_attention →
  • _cached_session_sets, a per-process cache of three raw 30-day analytics_events rollups (MCP_VERIFY_OBSERVED_ATTENTION_SESSION_CACHE_SECONDS=1800).

  • The cache was check-then-compute with no single-flight. At each expiry every
  • in-flight request recomputed the rollups concurrently → Postgres temp spill → verify-web pool (20 + 15 overflow) exhausted → every DB-backed route stalled (policy 41 s, compare 12 s, /mcp 13 s, 500/503s), not just Intelligence. It recurred at ~30-minute boundaries during every export; 06:02 was the worst.

  • The exporter amplified it: a 10 s client timeout abandoned requests that kept
  • running server-side and immediately sent the next one. On 09-28 the exporter counted 637 fetch errors while verify-web logged only 52 non-200 responses — ~585 were abandoned requests that later completed 200.

Change

  1. observed_attention._cached_session_sets: single-flight + stale-while-refresh.
  • fresh entry → returned;
  • expired entry → exactly one caller recomputes; concurrent callers get the stale value;
  • failed refresh → previous value kept and served; retry after
  • MCP_VERIFY_OBSERVED_ATTENTION_SESSION_REFRESH_RETRY_SECONDS (default 60);

  • cold start → one caller computes; others wait at most
  • MCP_VERIFY_OBSERVED_ATTENTION_SESSION_COLD_WAIT_SECONDS (default 60), then ObservedAttentionSessionSetsUnavailable. Attention results are unchanged; only who computes them and when.

  1. Exporter (intelligence/data/export.py):
  • DATA_FEED_EXPORT_CONCURRENCY default 10 → 4;
  • DATA_FEED_EXPORT_TIMEOUT_SECONDS 60 (was a hard-coded 10);
  • bounded retries (DATA_FEED_EXPORT_MAX_ATTEMPTS=3) on timeout / transport
  • error / 429 / 5xx only; 4xx and bad JSON fail immediately;

  • ExportFetchBackoffGate: a retryable failure pauses all export threads
  • for base * 2**(n-1) s (base 5, cap 120), so the exporter stops adding load to a struggling server instead of refilling its queue;

  • result/progress JSON gains fetch_backoff_pauses.

Known limits

  • The single-flight lock is process-local. N web processes can still mean N
  • concurrent recomputes per expiry. Production verify-web runs one uvicorn process, so this collapses to one. If verify-web is ever scaled out, move the session sets to a maintenance-worker-precomputed shared store.

  • The refreshing request still pays the full recompute latency itself.
  • Not changed here (separate decisions): skipping observed attention for the
  • internal export / a lighter export endpoint; Loki health.

  • Lower concurrency may lengthen the export (baseline 2h32m–2h53m at 10); watch
  • run duration alongside fetch_errors.

Validation

  • verify/tests/test_observed_attention_single_flight.py (12 concurrent callers at
  • expiry, cold start, bounded wait, failed refresh, retry window, cold-start failure handoff, unchanged results) — 6 of 8 fail on the pre-fix code.

  • verify/tests/test_intelligence_data_export_backoff.py (defaults, exponential +
  • capped backoff, bounded retries, non-retryable errors, shared pause across threads).

  • Production: no DB-wide latency spike at the ≥2 30-minute expiry boundaries
  • after deploy, and during the next 03:15 export; compare export fetch_errors (baseline 469–767/run) and duration.