Files
jbod-monitor/services/hwgate.py
adam 7878e17e55 Harden poller from Codex review: close residual storm/robustness gaps
Adversarial review (gpt-5.5) of the new architecture surfaced real gaps:

Hardware-safety / storm prevention:
- Restart-storm guard: on startup, defer the first sweep if a sweep ran
  < POLL_SWEEP_INTERVAL ago (persisted in Redis), so deploy-cycling can no
  longer trigger back-to-back full hardware sweeps (poller.py).
- Centralize pacing in the hardware gate: POLL_DRIVE_GAP is now held after
  EVERY hardware op (SMART, SES, host, MegaRAID, ledctl), not just the
  enclosure-drive loop (hwgate.py); removed the now-redundant per-loop sleeps.
- Single-instance lock made atomic (Lua compare-and-set / compare-and-expire)
  and the refresher is now fatal on any error + has a done-callback, so a
  dropped lock can never leave two pollers sweeping concurrently (store.py,
  poller.py).
- Warn loudly when POLL_CONCURRENCY > 1.

Correctness:
- Heartbeat now actually writes: cache_set omits EX when ttl<=0 (Redis
  rejects EX 0), so poller meta/last_sweep_ts persists — this also enables
  the restart-storm guard and the poller_fresh health flag (cache.py).
- Web image runs a single uvicorn worker so the MQTT publisher is a true
  singleton (two workers shared a client_id and flapped the broker) (Dockerfile).
- LED enqueue catches Redis errors -> API returns 503, not 500 (store.py).

Deploy independence:
- build.sh takes a target (web|poller|all) so web iterations never rebuild
  or repush the poller image.

Nits: drop dead SMART_CACHE_TTL constant.
2026-06-16 15:58:43 +00:00

50 lines
1.6 KiB
Python

"""Global hardware-access gate + pacing.
Every command that touches the SAS bus — smartctl, sg_ses, zpool, lsblk,
ledctl — must run inside this gate. It does two things:
1. Serializes hardware I/O (semaphore, default concurrency 1) so the
poller can never fire a thundering herd at fragile SAS expanders.
2. Paces it: after every hardware op it holds the gate for POLL_DRIVE_GAP
seconds, so back-to-back ops (gathered SMART, MegaRAID scans, SES,
ledctl) are always spaced — not just the per-drive loop.
This is a per-process gate; only the poller process calls hardware, so
per-process serialization is sufficient. Created lazily so it binds to the
running event loop.
"""
import asyncio
import contextlib
import logging
import os
logger = logging.getLogger(__name__)
_sem: asyncio.Semaphore | None = None
def _semaphore() -> asyncio.Semaphore:
global _sem
if _sem is None:
n = max(1, int(os.environ.get("POLL_CONCURRENCY", "1")))
if n > 1:
logger.warning(
"POLL_CONCURRENCY=%d (>1) — hardware ops will run concurrently; "
"only do this on backplanes known to tolerate it", n,
)
_sem = asyncio.Semaphore(n)
return _sem
@contextlib.asynccontextmanager
async def gate():
"""Acquire the hardware gate for one op, then hold it for the pacing gap."""
sem = _semaphore()
async with sem:
try:
yield
finally:
gap = float(os.environ.get("POLL_DRIVE_GAP", "0.5"))
if gap > 0:
await asyncio.sleep(gap)