Better Stack incident

Delayed data processing in ...

Minor Resolved View vendor source →

Better Stack experienced a minor incident on September 4, 2026 affecting Telemetry (Telemetry), lasting 2d 5h. The incident has been resolved; the full update timeline is below.

Started
Sep 04, 2026, 01:50 PM UTC
Resolved
Sep 06, 2026, 07:13 PM UTC
Duration
2d 5h
Detected by Pingoru
Sep 04, 2026, 01:50 PM UTC

Affected components

Telemetry (Telemetry)

Update timeline

  1. investigating Sep 04, 2026, 01:50 PM UTC

    We're seeing delayed data processing for some sources located in the Europe region again. The team is looking into resolving as soon as possible. There has been no data loss, all incoming data is ingested and will become visible as the system catches up. If you have any questions, please reach out to [email protected]. Thank you for your patience!

  2. resolved Sep 06, 2026, 07:13 PM UTC

    This incident is now resolved. All sources in the Europe region are processing in real time and the delayed backlog has been processed in full. What happened: On September 3, we consolidated several EU Telemetry clusters onto new, higher-capacity infrastructure. The cutover multiplied concurrent write streams on the cluster, and the volume of concurrent inserts saturated per-host write throughput on I/O contention. Ingestion fell behind, and affected sources saw delays in Live tail, dashboards, alerts, and queries. Our initial fixes, tuning the write path and adding capacity, restored real-time ingestion within a day, but draining the backlog alongside live traffic re-created the contention. We then rebuilt the recovery pipeline and held live data with strict priority while separately replaying the backlog at a controlled rate. Replaying several days of data was the slowest part of the recovery by design, as we rate-limited the replay so that real-time ingestion wasn't put at risk. We know teams rely on this data to operate their own systems, and losing that visibility for days is serious. We are sorry, both for the disruption and for turning what we originally announced as 90-minute maintenance into a multi-day incident. The fixes we made here are permanent, and this kind of contention can't happen again. Thank you very much for your patience. Please reach out to [email protected] if you have any questions.