Cloudli incident

Cloudli Service API and Quoting Tool | API de service Cloudli et outil de devis

Major Resolved View vendor source →

Cloudli experienced a major incident on September 3, 2026 affecting Other, lasting 4h 39m. The incident has been resolved; the full update timeline is below.

Started
Sep 03, 2026, 04:20 PM UTC
Resolved
Sep 03, 2026, 09:00 PM UTC
Duration
4h 39m
Detected by Pingoru
Sep 03, 2026, 04:20 PM UTC

Affected components

Other

Update timeline

  1. investigating Sep 03, 2026, 04:20 PM UTC

    Dear clients, We are currently investigating an issue with our Cloudli service as well as the impact it may have on some of our customers. We will update you upon further discovery. Services Impacted: Quoting tool & Customer API *** Chers clients, Nous enquêtons actuellement sur un problème affectant notre service Cloudli et son impact potentiel sur certains de nos clients. Nous vous tiendrons informés dès que nous aurons plus d'informations. Services concernés : Outil de devis et API client

  2. monitoring Sep 03, 2026, 04:26 PM UTC

    A fix has been implemented and we are monitoring the results.

  3. monitoring Sep 03, 2026, 04:26 PM UTC

    A fix has been implemented and we are monitoring the results. *** Un correctif a été mis en œuvre et nous surveillons les résultats.

  4. resolved Sep 03, 2026, 09:00 PM UTC

    This incident has been resolved. *** Cet incident a été résolu.

  5. postmortem Sep 04, 2026, 09:06 PM UTC

    ## Summary of Events: On September 3rd, 2026, at 11:52 AM EDT, Cloudli Communications identified an issue with the WorkerNode Auto-scaler incorrectly shutting down multiple instances hosting micro-services, including those supporting the VS API and Cloudli Quoting Tool, resulting in temporary service disruption. The affected services were restored at 12:11 PM EDT. The WorkerNode auto-scaler determines the compute resources required to run micro-services in our hosting facility and scales instances accordingly. The process relies on resource-usage metrics. During this incident, the WorkerNode auto-scaler did not receive all required metric information and scaled the environment down, draining some micro-services nodes. Because the configured minimum node count was set too low, the environment was allowed to scale below the capacity required to re-allocate all affected micro-services to alternate instances. ## Incident Analysis and Mitigation Measures: Post-incident analysis identified incomplete metrics as the condition that caused the WorkerNode auto-scaler to reduce capacity beyond a safe operating threshold. The minimum auto-scaler node setting has been increased to a safe level to prevent the environment from scaling below the capacity required to maintain service availability. Engineering has also identified the metrics-reading issue and is working on permanent correction. These safeguards are intended to prevent a similar metrics issue from causing an unsafe reduction in WorkerNode micro-services capacity in the future. ## Final Remarks: While the incident was short in duration, we take any interruption of service very seriously and are continuously evaluating new processes and mitigation measures that can be proactively implemented to ensure service continuity. We thank you for your continued support. Please feel free to reach out if you would like to discuss the particulars of this incident report further.