LiveKit incident
Investigating reports of elevated egress & connector API errors
LiveKit experienced a minor incident on August 20, 2026 affecting Global Egress, lasting 37m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- investigating Aug 20, 2026, 09:10 PM UTC
We are currently investigating reports of elevated API errors across the Egress service. We will post another update in 5 minutes.
- investigating Aug 20, 2026, 09:24 PM UTC
We had a spike of errors between 20:52 and 21:01 UTC, and the errors have subsided to baseline as of 21:01 UTC. We are actively monitoring and will share more updates in 30 minutes.
- resolved Aug 20, 2026, 09:47 PM UTC
This incident has been resolved. Between 20:52 and 21:01 UTC, we observed an increase in failed API requests. Errors have returned to baseline as of 21:01 UTC. We will follow up with a postmortem with further details.
- postmortem Sep 10, 2026, 07:08 PM UTC
On August 20, 2026, customers in our US East region experienced approximately 10 minutes of 503 errors on the LiveKit Ingress API and failures launching media track egress, between 20:52 and 21:01 UTC. All other regions were unaffected. **Root Cause** An internal routing service lost one of its two instances, and a concurrent infrastructure issue with our cloud provider was preventing new compute nodes from provisioning in US East, so the replacement couldn't start. At 20:51 UTC, an automated maintenance process removed that last instance, making the service fully unavailable. It recovered at 21:01 UTC once the evicted instance freed capacity for a replacement. **Timeline \(UTC\)** * Aug 20, 20:52 - Automated maintenance removes last instance in the pool; customer-facing errors begin * Aug 20, 21:01 - Service restored * Aug 21, 03:00 - Cloud provider resolves the underlying infrastructure issue; region fully healthy **Mitigations & Follow-ups** * The service is restored with two instances guaranteed to run on separate physical nodes. * We updated the automated maintenance process to avoid evicting the last healthy instance of a service. * We added disruption protection for the affected service and are continuing broader resilience improvements. * Our cloud provider resolved the infrastructure issue that had prevented new nodes from scaling up.