Is LiveKit down?

Last checked 11m ago
Current status
LiveKit is up

No incidents right now.

Official status page: https://status.livekit.io · Polled every 5 minutes · 116 components tracked

LiveKit is operational right now. Last checked 11m ago; the most recent incident resolved 2d ago.

Real-time LiveKit status, recent outages, and incident history — pulled directly from LiveKit's official status page at https://status.livekit.io every 5 minutes. Pingoru tracks 116 LiveKit services and has captured 30 incidents in the last 90 days (99.76% uptime). Get email, Slack, Discord, or webhook alerts the moment LiveKit reports a new incident — free for 5 monitors, no credit card.

Users who monitor LiveKit also follow these Development services: Bitbucket OpenAI Datadog US1 Sentry NPM Travis CI Postman Pusher Pendo SonarCloud View all 6,000+ providers
LiveKit uptime 99.76% uptime · past 90 days
Mon Wed Fri
JunJulAugSep
Less More

Recent outages & incidents

Past 90 days
  1. Resolved 1h 4m
    Started Sep 15, 2026, 10:10 PM UTC · Resolved Sep 15, 2026, 11:15 PM UTC
    Global EgressGlobal IngressGlobal SIP
    Timeline · 3 updates
    • monitoring · Sep 15, 2026, 10:10 PM UTC

      From 21:00 to 21:25 UTC we observed an occurrence of the same incident experienced earlier today impacting Global SIP, Ingress, and Egress due to the same underlying issue. We are currently operating normally and are continuing to monitor.

    • resolved · Sep 15, 2026, 11:15 PM UTC

      We have resolved the underlying issue and have not observed any further occurrences of this incident. We will follow up with a detailed postmortem in the original incident.

    • postmortem · Sep 17, 2026, 01:53 PM UTC

      Please see a full postmortem [here](https://status.livekit.io/incidents/dbqb4jhxcg9h).

    Latest: Please see a full postmortem [here](https://status.livekit.io/incidents/dbqb4jhxcg9h).

  2. Resolved 1h 11m
    Started Sep 15, 2026, 05:04 PM UTC · Resolved Sep 15, 2026, 06:16 PM UTC
    Global EgressGlobal IngressGlobal SIP
    Timeline · 5 updates
    • investigating · Sep 15, 2026, 05:04 PM UTC

      We are investigating reports of issues across multiple services stemming from increased latency starting at 13:30 UTC.

    • monitoring · Sep 15, 2026, 05:37 PM UTC

      We are still investigating, but believe that customer impact should be mostly mitigated - although some spikes may still be occurring. We are investigating reports of aborted egresses which we believe were also related to this issue. We will post another update as soon as possible.

    • monitoring · Sep 15, 2026, 06:03 PM UTC

      We are continuing to monitor, but want to keep users up to date with how we believe impact may have materialized: - Two impact windows 3 minutes long centered around 13:40 and 16:11 UTC - Impacted services appear to be concentrated around SIP, Egress, and Ingress

    • resolved · Sep 15, 2026, 06:16 PM UTC

      We are resolving this incident as the underlying issue has been resolved. Thank you for your patience and apologies for the disruption. We will follow up with a postmortem as soon as possible.

    • postmortem · Sep 17, 2026, 02:26 PM UTC

      ### Summary On 15 September, LiveKit Cloud experienced three periods of degraded performance between 13:34–13:41, 16:06–16:11 and 21:10–21:22 UTC. Ingress creation was the most affected: more than half of `CreateIngress` requests failed at the worst point. A number of in-progress recordings ended prematurely, and a small percentage of inbound SIP calls failed during a one-minute period \(during each impact window\). Realtime connections were not affected and media already flowing continued normally. The cause was a dropped connection to our primary metadata database that neither the client nor server side detected, which left a transaction open and holding row locks for up to 17 minutes. Requests queued behind those locks, and the sudden release of that backlog — not the wait itself — is what briefly degraded the wider platform. ### Root Cause Our services share a distributed global metadata database that stores ingress and egress state. At the time of the incident, a connection between one of our services and that database was dropped in a way neither end observed: the client treated the connection as closed and returned an error, while the database continued to consider the session live. Because the client had a transaction open at the time, the database kept that transaction's row locks held. The impact then came in two distinct phases. **While the lock was held**, only requests that needed the same rows were affected. Ingress creation failed, and recordings that could not report their status ended early. Most other APIs were unaffected, because they never touched the locked rows. **When the lock was released**, several hundred transactions that had queued behind it — some waiting more than seventeen minutes — all executed within a fraction of a second. That burst briefly saturated the database and degraded it for every service using it, not just those touching the original rows. This is why the broadest impact, including room management APIs and SIP participant creation, appears at the very end of each window rather than during it. The database only reclaimed the abandoned session when the operating system's TCP timeout expired, roughly 17 minutes after the connection was dropped. Two design choices amplified a single stuck row into regional impact: * **Recording workers treated "cannot report status" as "cannot accept work."** Every worker in a region reports through one shared service, so when that service became slow, all workers in the region stopped accepting new work at the same time and our router saw no available capacity. * **Several internal queries scanned the affected table without a narrowing filter.** That meant one locked row could block reads that were otherwise unrelated to it, which is what allowed the backlog to grow large enough to be disruptive on release. ### Scope of Impact Ingress creation was the most affected API: at the worst point in each window, 50.3%, 62.4% and 58.1% of `CreateIngress` requests returned server errors. Many failed requests are retried automatically, so the share of ingresses that ultimately could not be created is lower than those figures suggest. Starting a new egress was largely unaffected, staying under 2% throughout. SIP was affected at the end of each window, with `CreateSIPParticipant` reaching 7.50% and between 1.3% and 2.2% of attempted inbound calls failing during call setup; calls already connected were unaffected. Room management APIs stayed below 0.5%. Realtime connections were not affected, and media already flowing continued normally. The more consequential impact was to recordings already in progress. Up to 16.5% of egresses started during an affected window ended prematurely — 0.39% of all egresses that day — and did so without surfacing an error to indicate the recording had failed. ### Mitigations and Follow-ups Completed: * We have set a database-side idle transaction timeout so an abandoned transaction can no longer hold locks for more than 10 seconds. This caps both the wait and the size of any backlog that can accumulate behind it. We have verified this against a reproduction of the original failure. Underway: * We are removing an unnecessary transaction wrapper around single-statement writes, which shrinks the window in which a dropped connection can leave locks held. * We are adding database-side statement and lock timeouts so that no query can wait indefinitely on a lock. * We are changing recording workers so that a single status-reporting failure no longer removes an entire regional fleet from service. * We are making egress startup retry transient database errors rather than aborting the job. We appreciate your understanding and are committed to continuously improving our platform's reliability. If you have any questions, please reach out to our support team.

    Latest: ### Summary On 15 September, LiveKit Cloud experienced three periods of degraded performance between 13:34–13:41, 16:06–16:11 and 21:10–21:22 UTC. Ingress creation was the most aff…

  3. Resolved 2h 58m
    Started Sep 15, 2026, 12:31 PM UTC · Resolved Sep 15, 2026, 03:30 PM UTC
    Cloud Dashboard (cloud.livekit.io)US West - Analytics Ingestion
    Timeline · 6 updates
    • investigating · Sep 15, 2026, 12:08 PM UTC

      We are investigating an issue where users may be unexpectedly signed out of the LiveKit Cloud dashboard and unable to sign back in.

    • investigating · Sep 15, 2026, 12:31 PM UTC

      We are continuing to investigate this issue.

    • investigating · Sep 15, 2026, 12:49 PM UTC

      We are continuing to investigate this issue.

    • identified · Sep 15, 2026, 01:04 PM UTC

      We have identified the cause of the issue affecting the Cloud Dashboard and Analytics Services APIs, and recovery work is underway. Customers may continue to see errors accessing the Cloud Dashboard and Analytics Services APIs during this time. Room APIs, RTC APIs, and all other services remain healthy and unaffected. We have no indication of data loss from this incident. We will post another update as recovery progresses.

    • monitoring · Sep 15, 2026, 02:29 PM UTC

      A fix has been deployed as of 14:23 UTC, and access to the Cloud Dashboard and Analytics Services APIs is returning to normal. We are continuing to monitor before marking this resolved. Recovery may take a few additional minutes for some clients as the change fully propagates; retrying or reconnecting will restore access.

    • resolved · Sep 15, 2026, 03:30 PM UTC

      This incident has been resolved. Between 12:08 and 14:23 UTC, customers were unable to access the Cloud Dashboard and the Analytics Services APIs, caused by a loss of network reachability to the infrastructure serving them. We routed traffic to a healthy endpoint, and access returned to normal by 14:23 UTC. Room APIs, RTC APIs, and all other services were unaffected, and we have no indication of data loss. We will follow up with a postmortem.

    Latest: This incident has been resolved. Between 12:08 and 14:23 UTC, customers were unable to access the Cloud Dashboard and the Analytics Services APIs, caused by a loss of network reach…

  4. Resolved
    Started Sep 08, 2026, 08:45 AM UTC · Resolved Sep 04, 2026, 10:00 AM UTC
    Timeline · 1 update
    • resolved · Sep 08, 2026, 08:45 AM UTC

      A subset of room sessions that ended around 4 September 09:57 UTC were not recorded with an end time. These sessions continue to appear as active in the Cloud dashboard and in session records, and no `room_ended` end time is reflected for them. This was a reporting issue only. Rooms ended normally for participants, live traffic was not affected, and there is no impact on usage or billing. The underlying cause has been identified and fixed. Affected sessions from this window may continue to display as active; they can be safely disregarded.

    Latest: A subset of room sessions that ended around 4 September 09:57 UTC were not recorded with an end time. These sessions continue to appear as active in the Cloud dashboard and in sess…

  5. Resolved 1h 35m
    Started Sep 05, 2026, 11:36 PM UTC · Resolved Sep 06, 2026, 01:12 AM UTC
    Europe Central - Cloud Agents
    Timeline · 4 updates
    • identified · Sep 05, 2026, 11:36 PM UTC

      We have identified an issue with LiveKit Cloud hosted agents in eu-central. Existing agents in eu-central are facing scaling issues resulting in increased join latencies. Existing agents are still functional and can serve sessions normally, but they may be slower to join new sessions. Additionally, new agents cannot be created in this region. We are actively working on a mitigation to alleviate this issue. We will provide another update as soon as possible.

    • monitoring · Sep 06, 2026, 12:24 AM UTC

      We have applied a mitigation and believe that symptoms should be improving. We will update again with confirmation as soon as possible.

    • resolved · Sep 06, 2026, 01:12 AM UTC

      The mitigation was successful and users shouldn't notice any issues with join latencies or creating new agents. We will follow up as soon as possible with a postmortem.

    • postmortem · Sep 08, 2026, 09:59 PM UTC

      **Root Cause** A configuration change introduced a formatting error in the startup settings for the server pools that run hosted agents in our eu-central region. Newly provisioned servers failed to start, so the region could not add capacity. Agents that were already running were unaffected, but new agent deployments, version rollouts, and automatic scale-ups were delayed or stuck pending until the change was reverted. **Timeline \(UTC\)** `2026-09-04 23:56` - Configuration change deployed. The error only affected newly provisioned servers, so there was no immediate impact. `2026-09-05 18:07` - New servers began failing to start. Impact begins, intermittent at first. `2026-09-05 22:23` - Our monitoring alerted us to agent deployments stuck pending and we began investigating. `2026-09-06 00:50` - We identified the malformed startup configuration and reverted the change. `2026-09-06 01:17` - New capacity provisioned successfully, all pending deployments recovered, and we validated the fix. **Scope of Impact** Limited to hosted agents in our eu-central region. During the impact window, new agent deployments and rollouts were delayed or stuck pending, and automatic scale-up was blocked. Agents already running continued to serve sessions normally, and we estimate that, at the peak, roughly 1% of agent instances in the region were affected. No other regions or products were affected. **Mitigations and Follow-ups** * The faulty configuration change was reverted and provisioning has been stable since. * We are adding alerting for new servers that fail to start via this failure mode, which today fails silently. This lets us detect this class of failure directly and validate future changes to server provisioning quickly. We appreciate your understanding and are committed to continuously improving our platform's reliability. If you have any questions, please reach out to our support team.

    Latest: **Root Cause** A configuration change introduced a formatting error in the startup settings for the server pools that run hosted agents in our eu-central region. Newly provisioned …

See the full LiveKit outage history

25 more incidents in the last 90 days, plus the full multi-year archive of per-service events and update timelines.

Browse LiveKit outage history →

Or sign up free to get alerts when LiveKit breaks · 10 free monitors · No credit card

Outage history

Past 90 days · 30 incidents View full outage history →