We use essential cookies to keep you signed in, and — with your OK — product analytics and session recordings to fix bugs and improve PassTraq. Recordings never include the trading workbench and mask financial figures. Privacy Policy
Status & post-mortems
PassTraq does not publish an uptime percentage. Our monitors stayed green through several of the incidents below — a number built on those monitors would be marketing, not information. What we publish instead is every incident that affected users: how long it really lasted, what actually caused it, and what structurally changed so the same failure cannot repeat quietly.
An incident appears here once its fix is deployed and verified. Newest first.
Volume profile froze partially loaded
Duration: Intermittent over several days for workbench users · Not caught by monitoring at the time
Impact
The volume profile study could render a smaller, clean-looking profile built from a subset of sessions and never fill in the rest — with no error and no retry.
Root cause
Two-layer caching defect: the market-data gateway memoized days a short upstream replay never reached as "empty", and the browser then cached those empty days permanently. Every refresh concluded nothing was missing.
What changed
The gateway no longer records an unproven empty day; the browser never caches an empty trading day; every network wait on the path is now bounded; and the profile displays an explicit completeness verdict ("7/10 sessions — retrying") instead of rendering partial data silently.
Chart history unavailable (upstream outage)
Duration: ~4 hours, overnight · Not caught by monitoring at the time
Impact
Historical candle loads and volume profile failed while the upstream history service refused TLS connections. Live streaming data was unaffected.
Root cause
Upstream (Rithmic History Plant) outage. Our gateway re-dialed the dead endpoint for every request — hundreds of dials against a service that was down.
What changed
A circuit breaker per upstream credential (stop dialing a dead endpoint, retry on a schedule), readable error responses instead of opaque 5xxs, and market-phase-aware messaging so an outage is not mislabeled as a quiet market.
Own-broker chart history silently served from the shared feed
Duration: 5 days (Aug 17–22) · Not caught by monitoring at the time
Impact
Chart history for connected-broker accounts was answered from the operator’s shared data source instead of the trader’s own broker connection. Charts looked normal; the data source was wrong.
Root cause
Three environment secrets were missing after a deployment, which disabled the per-credential history path. Nothing surfaced it: the fallback worked, so there was no symptom.
What changed
Every history answer now carries a data-source header; surfaces that must only ever use the trader’s own feed fail closed instead of falling back; the boot log line is checked after every gateway restart.
Connection cap at ~996 concurrent chart sessions
Duration: Load-test finding (pre-launch), would have hit at scale · Not caught by monitoring at the time
Impact
The market-data gateway could serve at most ~996 concurrent WebSocket subscribers; the 997th chart simply failed to connect with nothing in the logs.
Root cause
The systemd default file-descriptor limit (1024) — one descriptor per subscriber, minus the process’s own.
What changed
Raised the limit explicitly and documented it in the capacity baseline.
Frozen charts during regular trading hours
Duration: ~2 hours degraded · Not caught by monitoring at the time
Impact
Live charts intermittently froze for workbench users while the rest of the app stayed healthy.
Root cause
A cache-thrash loop inside the market-data gateway: a prefetcher and an 8-entry cache starved the layer that serves live data.
What changed
Sealed-day disk cache so historical reads stop competing with live serving; the prefetcher stays off.
Signup dead for ~16 weeks
Duration: ~16 weeks (April → August 11) · Not caught by monitoring at the time
Impact
Every signup attempt failed at the confirmation-email step and was rolled back. Roughly 40 real people tried to sign up during the window and could not.
Root cause
The email-provider API key stored as the auth service’s SMTP password had been revoked in April. Every auth email failed; nothing monitored that path. Found via a session recording of a real user hitting the error.
What changed
A daily canary that performs a real SMTP login with the exact credential the auth service uses, a 5-minute external monitor on it, and a support notification on every new signup.
Order-history writes failing for 35+ minutes
Duration: ~35 minutes · Not caught by monitoring at the time
Impact
The order-management service could not persist order events to the database. Trading itself continued; the record of it lagged.
Root cause
Database-provider degradation; the service had no bounded retry, only log spam (~30k warning lines).
What changed
A bounded retry queue with backoff, full persist-outcome counters, and a dedicated persist-health endpoint with its own external monitor.
Market data frozen for 2h46m during trading hours
Duration: 2 hours 46 minutes · Not caught by monitoring at the time
Impact
No live ticks reached any chart during regular trading hours.
Root cause
The data provider moved its gateway IPs; our connection was pinned to an IP instead of the hostname.
What changed
Connect by hostname, never IP; and a new freshness monitor that asks the only question that matters — is market data actually arriving right now? — instead of whether processes are up.
Live health endpoints (public, no auth): /api/health, /api/health/feed, /api/health/deep.