Google Cloud incident

Incident affecting AlloyDB for PostgreSQL, Apigee Edge Public Cloud and Apigee X

Critical Resolved

Google Cloud experienced a critical incident on August 20, 2026 affecting AlloyDB for PostgreSQL and Apigee Edge Public Cloud and 1 more component, lasting 3h 40m. The incident has been resolved; the full update timeline is below.

Started
Aug 20, 2026, 03:40 PM UTC
Resolved
Aug 20, 2026, 07:20 PM UTC
Duration
3h 40m
Detected by Pingoru
Aug 20, 2026, 03:40 PM UTC

Affected components

AlloyDB for PostgreSQLApigee Edge Public CloudApigee XArtifact RegistryBigQuery Data Transfer ServiceCloud BuildCloud Data FusionCloud FilestoreCloud Key Management ServiceCloud Monitoring

Update timeline

  1. investigating Aug 20, 2026, 04:44 PM UTC

    Description: We are experiencing an issue with multiple products, beginning on Thursday, 2026-08-20 08:40 PDT. Our engineering team continues to investigate the issue. We will provide an update by Thursday, 2026-08-20 10:30 PDT with details. We apologize to all who are affected by the disruption. Diagnosis / customer symptoms: Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. Workaround: None at this time.

  2. investigating Aug 20, 2026, 05:13 PM UTC

    Description: Mitigation actions are implemented by our engineering teams, recovery trends have been observed across infrastructure layers. Active efforts remain underway to bring impacted cloud services back to full operation. We will provide an update by Thursday, 2026-08-20 10:45 PDT with details. Diagnosis / customer symptoms: Mitigation actions are implemented by our engineering teams, recovery trends have been observed across infrastructure layers. Active efforts remain underway to bring impacted cloud services back to full operation. We will provide an update by Thursday, 2026-08-20 10:45 PDT with details. Workaround: We recommend customers to failover to other regions where feasible.

  3. investigating Aug 20, 2026, 05:32 PM UTC

    Description: Mitigation actions are implemented by our engineering teams, recovery trends have been observed across infrastructure layers. Active efforts remain underway to bring impacted cloud services back to full operation. We will provide an update by Thursday, 2026-08-20 11:00 PDT with details. Diagnosis / customer symptoms: Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. Workaround: We recommend customers to failover to other regions where feasible.

  4. investigating Aug 20, 2026, 06:03 PM UTC

    Description: Mitigation actions have been completed by our engineering teams and we are seeing recovery from multiple products. We are continuing to work to recover the remaining products. We will provide an update by Thursday, 2026-08-20 11:45 PDT with details. Diagnosis / customer symptoms: Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. Workaround: No workarounds needed at this time.

  5. investigating Aug 20, 2026, 06:53 PM UTC

    Description: Mitigation actions have been completed by our engineering teams and we are seeing recovery from multiple products. We are continuing to work to recover the remaining products. We will provide an update by Thursday, 2026-08-20 12:30 PDT with details. Diagnosis / customer symptoms: Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. Workaround: No workarounds needed at this time.

  6. investigating Aug 20, 2026, 07:21 PM UTC

    Description: Mitigation actions have been completed by our engineering teams and we are seeing recovery from multiple products. We are continuing to work to recover the remaining products. We will provide an update by Thursday, 2026-08-20 13:30 PDT with details. Diagnosis / customer symptoms: Customers in us-west1 may experience timeouts, service degradations, errors, and elevated latencies across multiple products. Workaround: No workarounds needed at this time.

  7. resolved Aug 20, 2026, 07:37 PM UTC

    Description: We have mitigated the issue impacting multiple products in our us-west1 region as of Thursday, 2026-08-20 10:22 PDT. Our engineering teams have restored capacity on a planned optical maintenance that caused unexpected congestion in Dalles, Oregon metro /us-west1 region. Our systems stabilized and services have recovered once capacity was restored. Our teams are continuing to monitor for any residual impact. We will publish an analysis of this incident once we have completed our internal investigation. We thank you for your patience while we worked on resolving the issue. Diagnosis / customer symptoms: Customers in us-west1 may have experienced timeouts, service degradations, errors, and elevated latencies across multiple products. Workaround: This issue is now mitigated.

  8. resolved Aug 25, 2026, 06:20 AM UTC

    Affected services and features: The following services experienced elevated latencies and/or increased error rates, across their respective data planes and control planes: * AlloyDB * Apache Kafka * Apigee Edge Public Cloud * Apigee X * Artifact Registry * BigQuery & BigQuery Data Transfer Service * Cloud Build * Cloud Data Fusion * Cloud Dataflow * Cloud Filestore * Cloud Key Management Service (KMS) * Cloud Monitoring * Cloud Run * Cloud SQL * Contact Center AI Platform * Dataproc Metastore * Google App Engine * Google Cloud Bigtable * Google Cloud Pub/Sub * Google Cloud Storage (GCS) * Google Compute Engine (GCE) * Google Kubernetes Engine (GKE) * Identity and Access Management (IAM) * Managed Airflow (Cloud Composer) *Managed Service for Apache Spark (Dataproc) Persistent Disk

  9. resolved Aug 27, 2026, 09:45 PM UTC

    Affected services and features: The following services experienced elevated latencies and/or increased error rates, across their respective data planes and control planes: * AlloyDB * Apache Kafka * Apigee Edge Public Cloud * Apigee X * Artifact Registry * BigQuery & BigQuery Data Transfer Service * Cloud Build * Cloud Data Fusion * Cloud Dataflow * Cloud Filestore * Cloud Key Management Service (KMS) * Cloud Monitoring * Cloud Run * Cloud SQL * Contact Center AI Platform * Dataproc Metastore * Google App Engine * Google Cloud Bigtable * Google Cloud Pub/Sub * Google Cloud Storage (GCS) * Google Compute Engine (GCE) * Google Kubernetes Engine (GKE) * Identity and Access Management (IAM) * Managed Airflow (Cloud Composer) * Managed Service for Apache Spark (Dataproc) * Persistent Disk The regional incident in us-west1 followed a structured three-phase recovery dictated by platform dependency layers: While core platform connectivity and live request serving were restored by 10:22 US/Pacific, certain services took additional time to fully recover due to asynchronous backlog processing and localized control-plane state reconciliation for a very small set of customers. For event-driven and pipeline services, inbound error rates dropped to 0 immediately, but some operations experienced elevated latency while workers processed through backlogs accumulated during the outage, clearing later for Cloud Pub/Sub, Cloud Build & Deploy, Cloud Dataflow, and Cloud Storage lifecycle deletions. Concurrently, while primary read/write traffic was healthy across the region, specific long-running lifecycle workflows required extra time to clear locks and reconcile distributed state machines. Compute Engine VM provisioning in zone us-west1-c and Cloud Filestore control plane instance allocation locks and resource validation checks took longer to normalize across regional storage backends.