- Detected by Pingoru
- Sep 15, 2026, 10:10 PM UTC
- Resolved
- Sep 15, 2026, 11:15 PM UTC
- Duration
- 1h 4m
Affected: Global EgressGlobal IngressGlobal SIP
Timeline · 3 updates
-
monitoring Sep 15, 2026, 10:10 PM UTC
From 21:00 to 21:25 UTC we observed an occurrence of the same incident experienced earlier today impacting Global SIP, Ingress, and Egress due to the same underlying issue. We are currently operating normally and are continuing to monitor.
-
resolved Sep 15, 2026, 11:15 PM UTC
We have resolved the underlying issue and have not observed any further occurrences of this incident. We will follow up with a detailed postmortem in the original incident.
-
postmortem Sep 17, 2026, 01:53 PM UTC
Please see a full postmortem [here](https://status.livekit.io/incidents/dbqb4jhxcg9h).
Read the full incident report →
- Detected by Pingoru
- Sep 15, 2026, 05:04 PM UTC
- Resolved
- Sep 15, 2026, 06:16 PM UTC
- Duration
- 1h 11m
Affected: Global EgressGlobal IngressGlobal SIP
Timeline · 5 updates
Read the full incident report →
- Detected by Pingoru
- Sep 15, 2026, 12:31 PM UTC
- Resolved
- Sep 15, 2026, 03:30 PM UTC
- Duration
- 2h 58m
Affected: Cloud Dashboard (cloud.livekit.io)US West - Analytics Ingestion
Timeline · 6 updates
-
investigating Sep 15, 2026, 12:08 PM UTC
We are investigating an issue where users may be unexpectedly signed out of the LiveKit Cloud dashboard and unable to sign back in.
-
investigating Sep 15, 2026, 12:31 PM UTC
We are continuing to investigate this issue.
-
investigating Sep 15, 2026, 12:49 PM UTC
We are continuing to investigate this issue.
-
identified Sep 15, 2026, 01:04 PM UTC
We have identified the cause of the issue affecting the Cloud Dashboard and Analytics Services APIs, and recovery work is underway. Customers may continue to see errors accessing the Cloud Dashboard and Analytics Services APIs during this time. Room APIs, RTC APIs, and all other services remain healthy and unaffected. We have no indication of data loss from this incident. We will post another update as recovery progresses.
-
monitoring Sep 15, 2026, 02:29 PM UTC
A fix has been deployed as of 14:23 UTC, and access to the Cloud Dashboard and Analytics Services APIs is returning to normal. We are continuing to monitor before marking this resolved. Recovery may take a few additional minutes for some clients as the change fully propagates; retrying or reconnecting will restore access.
-
resolved Sep 15, 2026, 03:30 PM UTC
This incident has been resolved. Between 12:08 and 14:23 UTC, customers were unable to access the Cloud Dashboard and the Analytics Services APIs, caused by a loss of network reachability to the infrastructure serving them. We routed traffic to a healthy endpoint, and access returned to normal by 14:23 UTC. Room APIs, RTC APIs, and all other services were unaffected, and we have no indication of data loss. We will follow up with a postmortem.
Read the full incident report →
- Detected by Pingoru
- Sep 08, 2026, 08:45 AM UTC
- Resolved
- Sep 04, 2026, 10:00 AM UTC
- Duration
- —
Timeline · 1 update
-
resolved Sep 08, 2026, 08:45 AM UTC
A subset of room sessions that ended around 4 September 09:57 UTC were not recorded with an end time. These sessions continue to appear as active in the Cloud dashboard and in session records, and no `room_ended` end time is reflected for them. This was a reporting issue only. Rooms ended normally for participants, live traffic was not affected, and there is no impact on usage or billing. The underlying cause has been identified and fixed. Affected sessions from this window may continue to display as active; they can be safely disregarded.
Read the full incident report →
- Detected by Pingoru
- Sep 05, 2026, 11:36 PM UTC
- Resolved
- Sep 06, 2026, 01:12 AM UTC
- Duration
- 1h 35m
Affected: Europe Central - Cloud Agents
Timeline · 4 updates
Read the full incident report →
- Detected by Pingoru
- Sep 02, 2026, 09:25 PM UTC
- Resolved
- Sep 02, 2026, 09:45 PM UTC
- Duration
- 19m
Affected: Global SIP
Timeline · 3 updates
-
investigating Sep 02, 2026, 09:03 PM UTC
We are investigating elevated issues in calls connected to LiveKit Phone numbers where callee audio is not being heard by the caller. This is currently only experienced for numbers issues by LiveKit, and no other SIP provider connectivity should be affected. Our team is actively investigating this.
-
monitoring Sep 02, 2026, 09:25 PM UTC
An upstream carrier reverted a codec change that was causing one-way audio issues on calls to LiveKit-issued phone numbers. As of 21:10 UTC calls are connecting normally and we are no longer seeing the issue. The upstream codec change has impacted calls from T-Mobile and AT&T carriers, and no other SIP provider connectivity was affected. We are monitoring before marking this resolved.
-
resolved Sep 02, 2026, 09:45 PM UTC
This incident has been resolved. We will follow up with a detailed postmortem as soon as possible.
Read the full incident report →
- Detected by Pingoru
- Sep 02, 2026, 08:48 AM UTC
- Resolved
- Sep 02, 2026, 12:03 PM UTC
- Duration
- 3h 14m
Affected: Global SIP
Timeline · 1 update
-
monitoring Sep 02, 2026, 08:48 AM UTC
Since ~06:30 UTC, outbound calls routed through Twilio are failing at an elevated rate. Twilio is reporting degraded voice connectivity (https://status.twilio.com/), affecting North America, Latin America, Europe, and the Middle East & Africa.
Read the full incident report →
- Detected by Pingoru
- Sep 02, 2026, 08:15 AM UTC
- Resolved
- Sep 02, 2026, 11:57 AM UTC
- Duration
- 3h 41m
Affected: Cloud Dashboard (cloud.livekit.io)
Timeline · 4 updates
-
investigating Sep 02, 2026, 08:15 AM UTC
We are currently investigating reports of our dashboard failing to load across multiple regions, beginning at 0700 UTC. Other services remain unaffected and the impact is limited to the dashboard.
-
identified Sep 02, 2026, 09:01 AM UTC
We have identified the cause: a problematic node in US-East and are deploying a fix. Impact continues to be limited to the dashboard.
-
monitoring Sep 02, 2026, 10:55 AM UTC
A fix has been deployed as of 1055 UTC, and the dashboard is now loading successfully. We are continuing to monitor before marking this resolved.
-
resolved Sep 02, 2026, 11:57 AM UTC
This incident has been resolved and dashboard load times returned to normal by 1055 UTC
Read the full incident report →
- Detected by Pingoru
- Aug 26, 2026, 10:44 AM UTC
- Resolved
- Aug 26, 2026, 12:22 PM UTC
- Duration
- 1h 38m
Timeline · 2 updates
-
investigating Aug 26, 2026, 10:44 AM UTC
We are investigating a period of elevated error rates and timeouts (08:47:00-08:53:35 UTC) for LiveKit Inference requests in our US East region. Service has since recovered and inference requests are completing normally. We are continuing to investigate the cause and confirm the full scope and duration of impact. A further update will follow.
-
resolved Aug 26, 2026, 12:22 PM UTC
Between 08:47 and 08:54 UTC, the service that routes LLM, STT and TTS requests through LiveKit Inference in our US East region was unavailable, causing inference requests from agents in that region to fail or time out. Sessions depending on those responses may have degraded or ended prematurely. Replacement capacity came online at 08:54 UTC and request volumes returned fully to normal levels by 09:02 UTC. We have confirmed no further impact since then.
Read the full incident report →
- Detected by Pingoru
- Aug 25, 2026, 10:33 AM UTC
- Resolved
- Aug 25, 2026, 11:47 AM UTC
- Duration
- 1h 13m
Affected: India - Real Time Communication
Timeline · 1 update
-
investigating Aug 25, 2026, 10:33 AM UTC
Between approximately 09:00 and 10:05 UTC, we observed intermittent elevated error rates on room and participant management APIs in the Mumbai region. Error rates have since returned to normal, and other regions were not affected. We are continuing to investigate the root cause and are monitoring the region.
Read the full incident report →
- Detected by Pingoru
- Aug 24, 2026, 10:11 AM UTC
- Resolved
- Aug 24, 2026, 02:14 PM UTC
- Duration
- 4h 3m
Affected: India - Analytics IngestionEurope Central - Analytics Ingestion
Timeline · 3 updates
-
monitoring Aug 24, 2026, 10:11 AM UTC
We identified an issue where agent session analytics in the Cloud dashboard displayed zero concurrent sessions, beginning late on August 21 (UTC). This was a reporting issue only — agent sessions, calls, and all real-time services in these regions operated normally throughout. Ingestion has been restored and current session data is now reporting correctly. We are backfilling historical analytics for the affected window; some charts may show incomplete history until this completes. We are monitoring while the backfill finishes.
-
monitoring Aug 24, 2026, 01:55 PM UTC
We are continuing to monitor while the historical data backfill completes, and will resolve this incident once all analytics for the affected window are fully restored.
-
resolved Aug 24, 2026, 02:14 PM UTC
Agent session analytics in the Cloud dashboard have recovered and current data is reporting correctly for all affected projects. Some gaps may remain in historical data from the affected window (August 21–24 UTC); these will be backfilled.
Read the full incident report →
- Detected by Pingoru
- Aug 23, 2026, 09:00 PM UTC
- Resolved
- Aug 23, 2026, 09:00 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved Aug 23, 2026, 08:28 PM UTC
From 09:18 UTC on 21 August, 0.008% of egress recordings processed in our US East region were terminated before completion. Affected recordings were marked EGRESS_FAILED ("egress timed out"). For file outputs, media was not written to the configured destination; for HLS and streaming destinations, media was delivered up to the point of failure. The cause was node scale-down evicting egress workers with active recordings. A fix preventing this has been deployed on 23 August 09:00 UTC.
Read the full incident report →
- Detected by Pingoru
- Aug 20, 2026, 09:10 PM UTC
- Resolved
- Aug 20, 2026, 09:47 PM UTC
- Duration
- 37m
Affected: Global Egress
Timeline · 4 updates
Read the full incident report →
- Detected by Pingoru
- Aug 19, 2026, 07:33 PM UTC
- Resolved
- Aug 20, 2026, 12:30 AM UTC
- Duration
- 4h 57m
Affected: Cloud Dashboard (cloud.livekit.io)
Timeline · 6 updates
-
investigating Aug 19, 2026, 07:33 PM UTC
We are currently investigating longer than usual dashboard loading times. No services appear to be impacted and we don't expect any data to be lost.
-
identified Aug 19, 2026, 09:22 PM UTC
Sessions data ingest remains delayed for about an hour. It's not impacting other dashboards. we've identified the source of the slow ingest and are working on a mitigation. will update again in 30 mins
-
identified Aug 19, 2026, 09:25 PM UTC
We are continuing to work on a fix for this issue.
-
identified Aug 19, 2026, 10:07 PM UTC
the amount of data we are ingesting in the sessions view is outpacing our ability to ingest, thus causing a delay in ingestion. we are pursuing a few mitigations right now including moving to a larger database instance to relieve the processing bottleneck.
-
monitoring Aug 20, 2026, 12:01 AM UTC
mitigations have been deployed and we are now catching up to realtime.
-
resolved Aug 20, 2026, 12:30 AM UTC
This incident has been resolved. Between 19:20 and 00:10 UTC, session data and observability ingestion was delayed by 45-60 min. The root cause has been addressed and data ingestion returned to normal by 00:10 UTC. No data was lost during this window.
Read the full incident report →
- Detected by Pingoru
- Aug 17, 2026, 06:10 PM UTC
- Resolved
- Aug 17, 2026, 07:22 PM UTC
- Duration
- 1h 11m
Affected: US East - SIP
Timeline · 4 updates
Read the full incident report →
- Detected by Pingoru
- Aug 14, 2026, 09:10 PM UTC
- Resolved
- Aug 14, 2026, 11:01 PM UTC
- Duration
- 1h 51m
Affected: US East - Cloud Agents
Timeline · 5 updates
Read the full incident report →
- Detected by Pingoru
- Aug 13, 2026, 08:42 PM UTC
- Resolved
- Aug 14, 2026, 12:15 AM UTC
- Duration
- 3h 33m
Affected: Cloud Dashboard (cloud.livekit.io)
Timeline · 3 updates
-
investigating Aug 13, 2026, 08:42 PM UTC
We are currently investigating longer than usual dashboard loading times. No services appear to be impacted and we don't expect any data to be lost.
-
monitoring Aug 13, 2026, 09:34 PM UTC
We are seeing a 5-10 minute delay in loading some sessions, but we are confident that no data is being lost. We are currently monitoring and will update again once ingestion times have returned to baseline.
-
resolved Aug 14, 2026, 12:15 AM UTC
We have resolved the incident as session load times have returned to baseline.
Read the full incident report →
- Detected by Pingoru
- Aug 12, 2026, 06:34 PM UTC
- Resolved
- Aug 12, 2026, 08:09 PM UTC
- Duration
- 1h 34m
Timeline · 3 updates
-
investigating Aug 12, 2026, 06:34 PM UTC
We are currently investigating alerts for elevated server error rates on LiveKit RoomService APIs (GetParticipant, ListParticipants, RemoveParticipant) across multiple regions. The impact appears to be intermittent and limited to a subset of room-management API requests. We will follow with more details as soon possible.
-
monitoring Aug 12, 2026, 07:20 PM UTC
Error rates on ListParticipants API requests returned to normal as of 18:58 UTC. We will follow up with more details as soon possible.
-
resolved Aug 12, 2026, 08:09 PM UTC
Error rates on ListParticipants API requests recovered at 18:58 UTC and have remained stable since.
Read the full incident report →
- Detected by Pingoru
- Aug 07, 2026, 09:22 PM UTC
- Resolved
- Aug 07, 2026, 10:27 PM UTC
- Duration
- 1h 5m
Timeline · 3 updates
-
investigating Aug 07, 2026, 09:22 PM UTC
We are currently investigating automated alerting which triggered for degraded Room service performance in the US West and US Central regions beginning around 19:45 UTC. We are working to determine if there is any customer-facing impact. If there is impact, we don't currently have reason to believe that it is widespread. We will update again within 30 minutes.
-
monitoring Aug 07, 2026, 09:57 PM UTC
We have applied a mitigation to the affected regions and are now monitoring to ensure that performance signals return to baseline. We are still working on determining impact (including how users can determine if they were affected) and will share those details as soon as possible.
-
resolved Aug 07, 2026, 10:27 PM UTC
This incident has been resolved. Between 19:45 and 20:00 UTC, US West and US Central experienced intermittent degradation in Room Service and SIP performance. We rerouted traffic and performance has returned to baseline as of 22:00 UTC. We will follow up with a postmortem, which will include further details regarding scale and scope.
Read the full incident report →
- Detected by Pingoru
- Aug 06, 2026, 07:51 PM UTC
- Resolved
- Aug 07, 2026, 02:51 AM UTC
- Duration
- 7h
Affected: Global Inference
Timeline · 4 updates
-
identified Aug 06, 2026, 07:47 PM UTC
Between approximately 18:45 and 19:15 UTC today, a subset of LiveKit Inference requests using the google/gemma-4-31b-it model returned errors. Affected agent sessions would have seen an LLM request error on those requests. Other models and other LiveKit services were not affected. No action is needed on your part. If you continue to see errors, please reach out to support. We apologize for the disruption.
-
identified Aug 06, 2026, 07:51 PM UTC
We are actively investigating elevated error rate on LiveKit Inference (google/gemma-4-31b-it)
-
monitoring Aug 06, 2026, 09:08 PM UTC
A burst of requests has increased the failure rate. Both primary and secondary model providers failed to fulfill the requests at the time; we are root-causing the issue. We are continuing to monitor.
-
resolved Aug 07, 2026, 02:51 AM UTC
Today there was an issue with one of our GPU providers that caused it run at about 50% capacity. Requests above capacity return 503s, which we route to another, fallback provider. This fallback provider was not able to rise to the occasion and also returned 503s. In the very near future we will be onboarding additional providers to avoid these kinds of issues.
Read the full incident report →
- Detected by Pingoru
- Aug 04, 2026, 08:45 PM UTC
- Resolved
- Aug 04, 2026, 10:11 PM UTC
- Duration
- 1h 25m
Affected: Global Inference
Timeline · 2 updates
-
monitoring Aug 04, 2026, 08:45 PM UTC
We are monitoring elevated timeouts that affected LiveKit Inference between 14:05 and approximately 15:00 UTC today. During that window, roughly 3% of LLM requests using the google/gemma-4-31b-it model timed out. Impacted customers would have seen LLM request timeouts in their agent sessions (for example, APITimeoutError in agent logs). Other models were not affected. We have identified the cause. During a period of increased traffic, response times for this model slowed, and a defect in our health-monitoring logic incorrectly marked a backup deployment as unhealthy and removed it from rotation, preventing requests from failing over as designed. Error rates returned to baseline by approximately 15:00 UTC, and a fix for the health-monitoring defect is in progress. We are continuing to monitor before marking this resolved. No action is needed on your part. If you continue to see timeouts, please reach out to support. We apologize for the disruption.
-
resolved Aug 04, 2026, 10:11 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 03, 2026, 03:30 PM UTC
- Resolved
- Aug 03, 2026, 07:48 PM UTC
- Duration
- 4h 17m
Affected: Cloud Dashboard (cloud.livekit.io)
Timeline · 3 updates
-
investigating Aug 03, 2026, 03:30 PM UTC
We are currently investigating an issue where the sessions view does not populate in the Cloud Dashboard. Real-time connectivity is not affected and no data loss is observed.
-
investigating Aug 03, 2026, 04:51 PM UTC
We have traced the underlying issue to a backlog in our data pipeline, and we are continuing to investigate the root cause while working to clear the backlog. This data ingest delay results in the Sessions view returning no results for any time range ending at the current time, including all Quick Ranges options in the dashboard. As a temporary workaround, select "Custom range" and set the end time at least 30 minutes in the past to load your sessions. Real-time connectivity is not affected and no data has been lost.
-
resolved Aug 03, 2026, 07:48 PM UTC
We have fixed the data pipeline delay and all sessions should be loading now in the Cloud Dashboard.
Read the full incident report →
- Detected by Pingoru
- Jul 30, 2026, 10:55 PM UTC
- Resolved
- Jul 30, 2026, 10:55 PM UTC
- Duration
- —
Timeline · 2 updates
Read the full incident report →
- Detected by Pingoru
- Jul 28, 2026, 02:46 PM UTC
- Resolved
- Jul 28, 2026, 02:23 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved Jul 28, 2026, 02:46 PM UTC
Between approximately 10:30 and 11:55 UTC on July 28, customers with telephony workloads in our London region experienced elevated error rates on SIP APIs. Small number of call transfers were affected. We mitigated by removing the affected infrastructure from service at 11:48 UTC, and error rates returned to normal by 11:55 UTC.
Read the full incident report →
- Detected by Pingoru
- Jul 27, 2026, 05:00 PM UTC
- Resolved
- Jul 27, 2026, 05:00 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved Jul 27, 2026, 05:25 PM UTC
On July 27, our San Jose region experienced a minor networking incident beginning at 16:14 UTC. Users connected to this region may have experienced reconnects and increased session start times. Traffic was routed away from this region at 16:39 UTC and errors have since returned to baseline.
Read the full incident report →