LiveKit Outage History

LiveKit is up right now

LiveKit had 87 outages in the last 2 years totaling 48h 36m of downtime — averaging 3.6 incidents per month.

There were 87 LiveKit outages since September 22, 2025 totaling 48h 36m of downtime. Each is summarised below — incident details, duration, and resolution information.

Source: https://status.livekit.io

Minor September 15, 2026

Monitoring reports of intermittently increased API latency impacting Egress, Ingress, and SIP

Detected by Pingoru
Sep 15, 2026, 10:10 PM UTC
Resolved
Sep 15, 2026, 11:15 PM UTC
Duration
1h 4m
Affected: Global EgressGlobal IngressGlobal SIP
Timeline · 3 updates
  1. monitoring Sep 15, 2026, 10:10 PM UTC

    From 21:00 to 21:25 UTC we observed an occurrence of the same incident experienced earlier today impacting Global SIP, Ingress, and Egress due to the same underlying issue. We are currently operating normally and are continuing to monitor.

  2. resolved Sep 15, 2026, 11:15 PM UTC

    We have resolved the underlying issue and have not observed any further occurrences of this incident. We will follow up with a detailed postmortem in the original incident.

  3. postmortem Sep 17, 2026, 01:53 PM UTC

    Please see a full postmortem [here](https://status.livekit.io/incidents/dbqb4jhxcg9h).

Read the full incident report →

Major September 15, 2026

Cloud Dashboard, Analytics Services, and Billing API failures

Detected by Pingoru
Sep 15, 2026, 12:31 PM UTC
Resolved
Sep 15, 2026, 03:30 PM UTC
Duration
2h 58m
Affected: Cloud Dashboard (cloud.livekit.io)US West - Analytics Ingestion
Timeline · 6 updates
  1. investigating Sep 15, 2026, 12:08 PM UTC

    We are investigating an issue where users may be unexpectedly signed out of the LiveKit Cloud dashboard and unable to sign back in.

  2. investigating Sep 15, 2026, 12:31 PM UTC

    We are continuing to investigate this issue.

  3. investigating Sep 15, 2026, 12:49 PM UTC

    We are continuing to investigate this issue.

  4. identified Sep 15, 2026, 01:04 PM UTC

    We have identified the cause of the issue affecting the Cloud Dashboard and Analytics Services APIs, and recovery work is underway. Customers may continue to see errors accessing the Cloud Dashboard and Analytics Services APIs during this time. Room APIs, RTC APIs, and all other services remain healthy and unaffected. We have no indication of data loss from this incident. We will post another update as recovery progresses.

  5. monitoring Sep 15, 2026, 02:29 PM UTC

    A fix has been deployed as of 14:23 UTC, and access to the Cloud Dashboard and Analytics Services APIs is returning to normal. We are continuing to monitor before marking this resolved. Recovery may take a few additional minutes for some clients as the change fully propagates; retrying or reconnecting will restore access.

  6. resolved Sep 15, 2026, 03:30 PM UTC

    This incident has been resolved. Between 12:08 and 14:23 UTC, customers were unable to access the Cloud Dashboard and the Analytics Services APIs, caused by a loss of network reachability to the infrastructure serving them. We routed traffic to a healthy endpoint, and access returned to normal by 14:23 UTC. Room APIs, RTC APIs, and all other services were unaffected, and we have no indication of data loss. We will follow up with a postmortem.

Read the full incident report →

Notice September 8, 2026

Some sessions incorrectly shown as active

Detected by Pingoru
Sep 08, 2026, 08:45 AM UTC
Resolved
Sep 04, 2026, 10:00 AM UTC
Duration
Timeline · 1 update
  1. resolved Sep 08, 2026, 08:45 AM UTC

    A subset of room sessions that ended around 4 September 09:57 UTC were not recorded with an end time. These sessions continue to appear as active in the Cloud dashboard and in session records, and no `room_ended` end time is reflected for them. This was a reporting issue only. Rooms ended normally for participants, live traffic was not affected, and there is no impact on usage or billing. The underlying cause has been identified and fixed. Affected sessions from this window may continue to display as active; they can be safely disregarded.

Read the full incident report →

Minor September 2, 2026

Investigating reports of one-way audio issues in LiveKit Phone Numbers

Detected by Pingoru
Sep 02, 2026, 09:25 PM UTC
Resolved
Sep 02, 2026, 09:45 PM UTC
Duration
19m
Affected: Global SIP
Timeline · 3 updates
  1. investigating Sep 02, 2026, 09:03 PM UTC

    We are investigating elevated issues in calls connected to LiveKit Phone numbers where callee audio is not being heard by the caller. This is currently only experienced for numbers issues by LiveKit, and no other SIP provider connectivity should be affected. Our team is actively investigating this.

  2. monitoring Sep 02, 2026, 09:25 PM UTC

    An upstream carrier reverted a codec change that was causing one-way audio issues on calls to LiveKit-issued phone numbers. As of 21:10 UTC calls are connecting normally and we are no longer seeing the issue. The upstream codec change has impacted calls from T-Mobile and AT&T carriers, and no other SIP provider connectivity was affected. We are monitoring before marking this resolved.

  3. resolved Sep 02, 2026, 09:45 PM UTC

    This incident has been resolved. We will follow up with a detailed postmortem as soon as possible.

Read the full incident report →

Minor September 2, 2026

Twilio Voice outage

Detected by Pingoru
Sep 02, 2026, 08:48 AM UTC
Resolved
Sep 02, 2026, 12:03 PM UTC
Duration
3h 14m
Affected: Global SIP
Timeline · 1 update
  1. monitoring Sep 02, 2026, 08:48 AM UTC

    Since ~06:30 UTC, outbound calls routed through Twilio are failing at an elevated rate. Twilio is reporting degraded voice connectivity (https://status.twilio.com/), affecting North America, Latin America, Europe, and the Middle East & Africa.

Read the full incident report →

Major September 2, 2026

LiveKit dashboard failing to load

Detected by Pingoru
Sep 02, 2026, 08:15 AM UTC
Resolved
Sep 02, 2026, 11:57 AM UTC
Duration
3h 41m
Affected: Cloud Dashboard (cloud.livekit.io)
Timeline · 4 updates
  1. investigating Sep 02, 2026, 08:15 AM UTC

    We are currently investigating reports of our dashboard failing to load across multiple regions, beginning at 0700 UTC. Other services remain unaffected and the impact is limited to the dashboard.

  2. identified Sep 02, 2026, 09:01 AM UTC

    We have identified the cause: a problematic node in US-East and are deploying a fix. Impact continues to be limited to the dashboard.

  3. monitoring Sep 02, 2026, 10:55 AM UTC

    A fix has been deployed as of 1055 UTC, and the dashboard is now loading successfully. We are continuing to monitor before marking this resolved.

  4. resolved Sep 02, 2026, 11:57 AM UTC

    This incident has been resolved and dashboard load times returned to normal by 1055 UTC

Read the full incident report →

Notice August 26, 2026

Elevated inference errors — US East

Detected by Pingoru
Aug 26, 2026, 10:44 AM UTC
Resolved
Aug 26, 2026, 12:22 PM UTC
Duration
1h 38m
Timeline · 2 updates
  1. investigating Aug 26, 2026, 10:44 AM UTC

    We are investigating a period of elevated error rates and timeouts (08:47:00-08:53:35 UTC) for LiveKit Inference requests in our US East region. Service has since recovered and inference requests are completing normally. We are continuing to investigate the cause and confirm the full scope and duration of impact. A further update will follow.

  2. resolved Aug 26, 2026, 12:22 PM UTC

    Between 08:47 and 08:54 UTC, the service that routes LLM, STT and TTS requests through LiveKit Inference in our US East region was unavailable, causing inference requests from agents in that region to fail or time out. Sessions depending on those responses may have degraded or ended prematurely. Replacement capacity came online at 08:54 UTC and request volumes returned fully to normal levels by 09:02 UTC. We have confirmed no further impact since then.

Read the full incident report →

Minor August 25, 2026

Investigating elevated API errors in India

Detected by Pingoru
Aug 25, 2026, 10:33 AM UTC
Resolved
Aug 25, 2026, 11:47 AM UTC
Duration
1h 13m
Affected: India - Real Time Communication
Timeline · 1 update
  1. investigating Aug 25, 2026, 10:33 AM UTC

    Between approximately 09:00 and 10:05 UTC, we observed intermittent elevated error rates on room and participant management APIs in the Mumbai region. Error rates have since returned to normal, and other regions were not affected. We are continuing to investigate the root cause and are monitoring the region.

Read the full incident report →

Minor August 24, 2026

Agent session analytics displaying incorrect values in EU Central and India regions

Detected by Pingoru
Aug 24, 2026, 10:11 AM UTC
Resolved
Aug 24, 2026, 02:14 PM UTC
Duration
4h 3m
Affected: India - Analytics IngestionEurope Central - Analytics Ingestion
Timeline · 3 updates
  1. monitoring Aug 24, 2026, 10:11 AM UTC

    We identified an issue where agent session analytics in the Cloud dashboard displayed zero concurrent sessions, beginning late on August 21 (UTC). This was a reporting issue only — agent sessions, calls, and all real-time services in these regions operated normally throughout. Ingestion has been restored and current session data is now reporting correctly. We are backfilling historical analytics for the affected window; some charts may show incomplete history until this completes. We are monitoring while the backfill finishes.

  2. monitoring Aug 24, 2026, 01:55 PM UTC

    We are continuing to monitor while the historical data backfill completes, and will resolve this incident once all analytics for the affected window are fully restored.

  3. resolved Aug 24, 2026, 02:14 PM UTC

    Agent session analytics in the Cloud dashboard have recovered and current data is reporting correctly for all affected projects. Some gaps may remain in historical data from the affected window (August 21–24 UTC); these will be backfilled.

Read the full incident report →

Notice August 23, 2026

Egress recordings terminated early in US East

Detected by Pingoru
Aug 23, 2026, 09:00 PM UTC
Resolved
Aug 23, 2026, 09:00 PM UTC
Duration
Timeline · 1 update
  1. resolved Aug 23, 2026, 08:28 PM UTC

    From 09:18 UTC on 21 August, 0.008% of egress recordings processed in our US East region were terminated before completion. Affected recordings were marked EGRESS_FAILED ("egress timed out"). For file outputs, media was not written to the configured destination; for HLS and streaming destinations, media was delivered up to the point of failure. The cause was node scale-down evicting egress workers with active recordings. A fix preventing this has been deployed on 23 August 09:00 UTC.

Read the full incident report →

Minor August 19, 2026

Investigating delayed session data ingestion on Cloud Dashboard

Detected by Pingoru
Aug 19, 2026, 07:33 PM UTC
Resolved
Aug 20, 2026, 12:30 AM UTC
Duration
4h 57m
Affected: Cloud Dashboard (cloud.livekit.io)
Timeline · 6 updates
  1. investigating Aug 19, 2026, 07:33 PM UTC

    We are currently investigating longer than usual dashboard loading times. No services appear to be impacted and we don't expect any data to be lost.

  2. identified Aug 19, 2026, 09:22 PM UTC

    Sessions data ingest remains delayed for about an hour. It's not impacting other dashboards. we've identified the source of the slow ingest and are working on a mitigation. will update again in 30 mins

  3. identified Aug 19, 2026, 09:25 PM UTC

    We are continuing to work on a fix for this issue.

  4. identified Aug 19, 2026, 10:07 PM UTC

    the amount of data we are ingesting in the sessions view is outpacing our ability to ingest, thus causing a delay in ingestion. we are pursuing a few mitigations right now including moving to a larger database instance to relieve the processing bottleneck.

  5. monitoring Aug 20, 2026, 12:01 AM UTC

    mitigations have been deployed and we are now catching up to realtime.

  6. resolved Aug 20, 2026, 12:30 AM UTC

    This incident has been resolved. Between 19:20 and 00:10 UTC, session data and observability ingestion was delayed by 45-60 min. The root cause has been addressed and data ingestion returned to normal by 00:10 UTC. No data was lost during this window.

Read the full incident report →

Minor August 13, 2026

Investigating longer than usual dashboard loading times

Detected by Pingoru
Aug 13, 2026, 08:42 PM UTC
Resolved
Aug 14, 2026, 12:15 AM UTC
Duration
3h 33m
Affected: Cloud Dashboard (cloud.livekit.io)
Timeline · 3 updates
  1. investigating Aug 13, 2026, 08:42 PM UTC

    We are currently investigating longer than usual dashboard loading times. No services appear to be impacted and we don't expect any data to be lost.

  2. monitoring Aug 13, 2026, 09:34 PM UTC

    We are seeing a 5-10 minute delay in loading some sessions, but we are confident that no data is being lost. We are currently monitoring and will update again once ingestion times have returned to baseline.

  3. resolved Aug 14, 2026, 12:15 AM UTC

    We have resolved the incident as session load times have returned to baseline.

Read the full incident report →

Notice August 12, 2026

Investigating alerts for increased timeout error rates on ListParticipants APIs

Detected by Pingoru
Aug 12, 2026, 06:34 PM UTC
Resolved
Aug 12, 2026, 08:09 PM UTC
Duration
1h 34m
Timeline · 3 updates
  1. investigating Aug 12, 2026, 06:34 PM UTC

    We are currently investigating alerts for elevated server error rates on LiveKit RoomService APIs (GetParticipant, ListParticipants, RemoveParticipant) across multiple regions. The impact appears to be intermittent and limited to a subset of room-management API requests. We will follow with more details as soon possible.

  2. monitoring Aug 12, 2026, 07:20 PM UTC

    Error rates on ListParticipants API requests returned to normal as of 18:58 UTC. We will follow up with more details as soon possible.

  3. resolved Aug 12, 2026, 08:09 PM UTC

    Error rates on ListParticipants API requests recovered at 18:58 UTC and have remained stable since.

Read the full incident report →

Minor August 7, 2026

A percentage of cross-region calls between US West and US Central not completing

Detected by Pingoru
Aug 07, 2026, 09:22 PM UTC
Resolved
Aug 07, 2026, 10:27 PM UTC
Duration
1h 5m
Timeline · 3 updates
  1. investigating Aug 07, 2026, 09:22 PM UTC

    We are currently investigating automated alerting which triggered for degraded Room service performance in the US West and US Central regions beginning around 19:45 UTC. We are working to determine if there is any customer-facing impact. If there is impact, we don't currently have reason to believe that it is widespread. We will update again within 30 minutes.

  2. monitoring Aug 07, 2026, 09:57 PM UTC

    We have applied a mitigation to the affected regions and are now monitoring to ensure that performance signals return to baseline. We are still working on determining impact (including how users can determine if they were affected) and will share those details as soon as possible.

  3. resolved Aug 07, 2026, 10:27 PM UTC

    This incident has been resolved. Between 19:45 and 20:00 UTC, US West and US Central experienced intermittent degradation in Room Service and SIP performance. We rerouted traffic and performance has returned to baseline as of 22:00 UTC. We will follow up with a postmortem, which will include further details regarding scale and scope.

Read the full incident report →

Minor August 6, 2026

Elevated error rate on LiveKit Inference (google/gemma-4-31b-it)

Detected by Pingoru
Aug 06, 2026, 07:51 PM UTC
Resolved
Aug 07, 2026, 02:51 AM UTC
Duration
7h
Affected: Global Inference
Timeline · 4 updates
  1. identified Aug 06, 2026, 07:47 PM UTC

    Between approximately 18:45 and 19:15 UTC today, a subset of LiveKit Inference requests using the google/gemma-4-31b-it model returned errors. Affected agent sessions would have seen an LLM request error on those requests. Other models and other LiveKit services were not affected. No action is needed on your part. If you continue to see errors, please reach out to support. We apologize for the disruption.

  2. identified Aug 06, 2026, 07:51 PM UTC

    We are actively investigating elevated error rate on LiveKit Inference (google/gemma-4-31b-it)

  3. monitoring Aug 06, 2026, 09:08 PM UTC

    A burst of requests has increased the failure rate. Both primary and secondary model providers failed to fulfill the requests at the time; we are root-causing the issue. We are continuing to monitor.

  4. resolved Aug 07, 2026, 02:51 AM UTC

    Today there was an issue with one of our GPU providers that caused it run at about 50% capacity. Requests above capacity return 503s, which we route to another, fallback provider. This fallback provider was not able to rise to the occasion and also returned 503s. In the very near future we will be onboarding additional providers to avoid these kinds of issues.

Read the full incident report →

Minor August 4, 2026

Elevated timeouts on LiveKit Inference (google/gemma-4-31b-it)

Detected by Pingoru
Aug 04, 2026, 08:45 PM UTC
Resolved
Aug 04, 2026, 10:11 PM UTC
Duration
1h 25m
Affected: Global Inference
Timeline · 2 updates
  1. monitoring Aug 04, 2026, 08:45 PM UTC

    We are monitoring elevated timeouts that affected LiveKit Inference between 14:05 and approximately 15:00 UTC today. During that window, roughly 3% of LLM requests using the google/gemma-4-31b-it model timed out. Impacted customers would have seen LLM request timeouts in their agent sessions (for example, APITimeoutError in agent logs). Other models were not affected. We have identified the cause. During a period of increased traffic, response times for this model slowed, and a defect in our health-monitoring logic incorrectly marked a backup deployment as unhealthy and removed it from rotation, preventing requests from failing over as designed. Error rates returned to baseline by approximately 15:00 UTC, and a fix for the health-monitoring defect is in progress. We are continuing to monitor before marking this resolved. No action is needed on your part. If you continue to see timeouts, please reach out to support. We apologize for the disruption.

  2. resolved Aug 04, 2026, 10:11 PM UTC

    This incident has been resolved.

Read the full incident report →

Minor August 3, 2026

Sessions view not populating in the Cloud Dashboard

Detected by Pingoru
Aug 03, 2026, 03:30 PM UTC
Resolved
Aug 03, 2026, 07:48 PM UTC
Duration
4h 17m
Affected: Cloud Dashboard (cloud.livekit.io)
Timeline · 3 updates
  1. investigating Aug 03, 2026, 03:30 PM UTC

    We are currently investigating an issue where the sessions view does not populate in the Cloud Dashboard. Real-time connectivity is not affected and no data loss is observed.

  2. investigating Aug 03, 2026, 04:51 PM UTC

    We have traced the underlying issue to a backlog in our data pipeline, and we are continuing to investigate the root cause while working to clear the backlog. This data ingest delay results in the Sessions view returning no results for any time range ending at the current time, including all Quick Ranges options in the dashboard. As a temporary workaround, select "Custom range" and set the end time at least 30 minutes in the past to load your sessions. Real-time connectivity is not affected and no data has been lost.

  3. resolved Aug 03, 2026, 07:48 PM UTC

    We have fixed the data pipeline delay and all sessions should be loading now in the Cloud Dashboard.

Read the full incident report →

Notice July 28, 2026

Elevated SIP API error rates in the London region

Detected by Pingoru
Jul 28, 2026, 02:46 PM UTC
Resolved
Jul 28, 2026, 02:23 PM UTC
Duration
Timeline · 1 update
  1. resolved Jul 28, 2026, 02:46 PM UTC

    Between approximately 10:30 and 11:55 UTC on July 28, customers with telephony workloads in our London region experienced elevated error rates on SIP APIs. Small number of call transfers were affected. We mitigated by removing the affected infrastructure from service at 11:48 UTC, and error rates returned to normal by 11:55 UTC.

Read the full incident report →

Notice July 27, 2026

Elevated connectivity issues in US West

Detected by Pingoru
Jul 27, 2026, 05:00 PM UTC
Resolved
Jul 27, 2026, 05:00 PM UTC
Duration
Timeline · 1 update
  1. resolved Jul 27, 2026, 05:25 PM UTC

    On July 27, our San Jose region experienced a minor networking incident beginning at 16:14 UTC. Users connected to this region may have experienced reconnects and increased session start times. Traffic was routed away from this region at 16:39 UTC and errors have since returned to baseline.

Read the full incident report →