Blacksmith Outage History

Blacksmith is up right now

Blacksmith had 96 outages in the last 2 years totaling 83h 38m of downtime — averaging 3.9 incidents per month.

There were 96 Blacksmith outages since January 30, 2026 totaling 83h 38m of downtime. Each is summarised below — incident details, duration, and resolution information.

Source: https://status.blacksmith.sh

Minor September 29, 2026

Backend Degradation causing job slowness and caching failures

Detected by Pingoru
Sep 29, 2026, 09:10 PM UTC
Resolved
Sep 29, 2026, 11:04 PM UTC
Duration
1h 54m
Affected: EU Central (eu-central ARM)EU Central (eu-central x86)US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)US Central (us-central MacOS)Incremental Docker Builders (eu-central Storage Cluster)Incremental Docker Builders (us-west Storage Cluster)Docker Container Cache (eu-central Storage Cluster)Docker Container Cache (us-west Storage Cluster)Sticky Disks (eu-central Storage Cluster)Sticky Disks (us-west Storage Cluster)Sticky Disks (eu-west Storage Cluster)Incremental Docker Builders (eu-west Storage Cluster)Docker Container Cache (eu-west Storage Cluster)EU West (eu-west ARM)US East (us-east x86)Actions Cache (US West Cache)Actions Cache (US East Cache)Actions Cache (EU West Cache)Actions Cache (EU Central Cache)Incremental Docker Builders (us-east Storage Cluster)Docker Container Cache (us-east Storage Cluster)Sticky Disks (us-east Storage Cluster)Actions Cache (US Central Cache)
Timeline · 6 updates
  1. investigating Sep 29, 2026, 09:10 PM UTC

    Jobs across all regions are taking longer to start and cache operations are failing. We are continuing to investigate the root cause.

  2. identified Sep 29, 2026, 09:28 PM UTC

    We have applied a fix and are monitoring its effect. Customers may still see delayed job starts and failing cache operations while recovery completes. We will provide an update within the next 30 minutes.

  3. monitoring Sep 29, 2026, 09:43 PM UTC

    Our mitigation has taken effect and job start times and cache operations are improving. Customers may still see some delayed job starts and cache failures while recovery completes. We will provide an update within the next 30 minutes.

  4. monitoring Sep 29, 2026, 10:20 PM UTC

    Caching has recovered, and job queue times across EU regions have recovered. Queue times for larger jobs (16 and 32 vcpu jobs) in US west and US east continue to remain elevated. We are continuing to monitor for any regressions.

  5. monitoring Sep 29, 2026, 10:42 PM UTC

    This incident is resolved. Job starts and cache operations returned to normal in all regions, and we are continuing to monitor. Jobs that failed during the incident can be re-run.

  6. resolved Sep 29, 2026, 11:04 PM UTC

    This incident has been resolved.

Read the full incident report →

Minor September 29, 2026

US west cache failure

Detected by Pingoru
Sep 29, 2026, 05:00 PM UTC
Resolved
Sep 29, 2026, 06:32 PM UTC
Duration
1h 32m
Affected: Actions Cache (US West Cache)
Timeline · 5 updates
  1. investigating Sep 29, 2026, 05:00 PM UTC

    We are investigating failing GitHub Actions cache operations for jobs running in us-west since approximately 16:10 UTC. Affected jobs may see cache restores and saves fail and fall back to a full install, which can make those jobs run longer or fail, while jobs in other regions are not affected.

  2. identified Sep 29, 2026, 05:09 PM UTC

    We have identified the cause of the failures, and we are implementing a fix to restore cache service. Customers with jobs in us-west may still see cache restores and saves fail and fall back to a full install, which can make those jobs run longer or fail, while jobs in other regions are not affected.

  3. monitoring Sep 29, 2026, 05:35 PM UTC

    We have restored the cache infrastructure in us-west, and cache failures have decreased but are not yet back to normal. Jobs in us-west may still see some cache restores and saves fail and fall back to a full install, while the earlier delayed job starts have cleared and other regions are unaffected. We are bringing additional cache capacity online to fully restore service.

  4. monitoring Sep 29, 2026, 06:06 PM UTC

    Cache errors for jobs in us-west have largely subsided. As we complete recovery, some jobs may see a one-time cache miss on their next run. We are monitoring and will provide an update within the next hour.

  5. resolved Sep 29, 2026, 06:32 PM UTC

    This incident is resolved. GitHub Actions cache operations for jobs in us-west have returned to normal, and any jobs that failed during the incident can be re-run.

Read the full incident report →

Minor September 28, 2026

Job failures in us-west

Detected by Pingoru
Sep 28, 2026, 08:52 PM UTC
Resolved
Sep 28, 2026, 10:14 PM UTC
Duration
1h 22m
Affected: US West (us-west ARM)US West (us-west x86)
Timeline · 3 updates
  1. identified Sep 28, 2026, 08:52 PM UTC

    A rack-level power failure at our us-west data center caused some jobs to fail with runner communication errors. The issue has been communicated with our upstream provider and we are identifying the affected machines to mitigate the impact to our customers.

  2. monitoring Sep 28, 2026, 09:53 PM UTC

    The affected rack remains offline while our upstream provider works to restore its power, and jobs in us-west are running normally on the rest of the fleet. Any jobs that failed with runner communication errors can be safely re-run.

  3. resolved Sep 28, 2026, 10:14 PM UTC

    This incident has been resolved.

Read the full incident report →

Minor September 24, 2026

GitHub connectivity degraded in us-west

Detected by Pingoru
Sep 24, 2026, 07:16 PM UTC
Resolved
Sep 24, 2026, 07:50 PM UTC
Duration
33m
Affected: US West (us-west ARM)US West (us-west x86)
Timeline · 3 updates
  1. investigating Sep 24, 2026, 07:16 PM UTC

    We are investigating degraded network connectivity between our us-west region and GitHub. Customers running jobs in us-west may see slower-than-normal git checkouts, while jobs in other regions are not affected.

  2. monitoring Sep 24, 2026, 07:31 PM UTC

    After rerouting traffic, congestion on the upstream network link has subsided, and git checkout performance in us-west has returned to normal. We are monitoring and will resolve this incident once performance has continued to stay stable.

  3. resolved Sep 24, 2026, 07:50 PM UTC

    Git checkout performance in us-west has remained stable since we rerouted traffic away from the congested upstream network link, and this incident is now resolved.

Read the full incident report →

Minor September 21, 2026

Elevated error rates in Blacksmith APIs

Detected by Pingoru
Sep 21, 2026, 06:37 PM UTC
Resolved
Sep 21, 2026, 06:47 PM UTC
Duration
9m
Affected: EU Central (eu-central ARM)EU Central (eu-central x86)US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)US Central (us-central MacOS)APIIncremental Docker Builders (eu-central Storage Cluster)Incremental Docker Builders (us-west Storage Cluster)Docker Container Cache (eu-central Storage Cluster)Docker Container Cache (us-west Storage Cluster)Sticky Disks (eu-central Storage Cluster)Sticky Disks (us-west Storage Cluster)Sticky Disks (eu-west Storage Cluster)Incremental Docker Builders (eu-west Storage Cluster)Docker Container Cache (eu-west Storage Cluster)EU West (eu-west ARM)US East (us-east x86)Actions Cache (US West Cache)Actions Cache (US East Cache)Actions Cache (EU West Cache)Actions Cache (EU Central Cache)Runtime Build Caching (US West Runtime Build Cache)Runtime Build Caching (EU Central Runtime Build Cache)Runtime Build Caching (EU West Runtime Build Cache)
Timeline · 2 updates
  1. identified Sep 21, 2026, 06:37 PM UTC

    We are seeing elevated errors in job adoption and other control plane APIs. This is yielding higher latencies for your job to be adopted. You may also see certain actions fail, such as actions caching and sticky disks. We have identified the issue and are rolling out a mitigation.

  2. resolved Sep 21, 2026, 06:47 PM UTC

    This issue has been resolved, and all services have operating normally. If any of your jobs failed during this time, please re-run them.

Read the full incident report →

Minor September 13, 2026

Degraded US East AWS connectivity

Detected by Pingoru
Sep 13, 2026, 01:11 AM UTC
Resolved
Sep 13, 2026, 01:58 AM UTC
Duration
46m
Affected: US East (us-east x86)
Timeline · 3 updates
  1. investigating Sep 13, 2026, 01:11 AM UTC

    We are currently observing network stalling in our US East region with runner connectivity against AWS's us-east-1 region. Uploads and downloads to and from ECR and S3 may experience occasional stalling. We are investigating.

  2. monitoring Sep 13, 2026, 01:40 AM UTC

    We have identified a small subset of runners that are exhibiting this network stalling and have isolated them. We're continuing to monitor recovery.

  3. resolved Sep 13, 2026, 01:58 AM UTC

    This incident has been resolved.

Read the full incident report →

Minor September 11, 2026

Upstream Ubuntu package mirror outage affecting apt

Detected by Pingoru
Sep 11, 2026, 05:30 AM UTC
Resolved
Sep 11, 2026, 02:59 PM UTC
Duration
9h 29m
Timeline · 3 updates
  1. monitoring Sep 11, 2026, 05:30 AM UTC

    Since 06:30 UTC Canonical has had three outages affecting [archive.ubuntu.com](http://archive.ubuntu.com), [us.archive.ubuntu.com](http://us.archive.ubuntu.com) and [security.ubuntu.com](http://security.ubuntu.com). Although they have marked these as resolved, we are still seeing intermittent connection failures to those mirrors, which can cause apt-get steps to fail or hang. In the meantime, pointing apt at [azure.archive.ubuntu.com](http://azure.archive.ubuntu.com) instead of [archive.ubuntu.com](http://archive.ubuntu.com) and [security.ubuntu.com](http://security.ubuntu.com) will unblock affected jobs.

  2. monitoring Sep 11, 2026, 02:21 PM UTC

    We are deploying a mitigation on our side that routes apt around the affected Canonical mirrors. For the time being pointing apt at [azure.archive.ubuntu.com](http://azure.archive.ubuntu.com) in place of [archive.ubuntu.com](http://archive.ubuntu.com) and [security.ubuntu.com](http://security.ubuntu.com) will unblock affected jobs.

  3. resolved Sep 11, 2026, 02:59 PM UTC

    This incident has been resolved - Canonical's Ubuntu apt mirrors have recovered and apt installs are succeeding again. We are also rolling out an image change to reduce the impact of upstream mirror outages in future.

Read the full incident report →

Major September 9, 2026

Major outage with our caching infrastructure and our website

Detected by Pingoru
Sep 09, 2026, 04:46 PM UTC
Resolved
Sep 09, 2026, 04:46 PM UTC
Duration
—
Timeline · 1 update
  1. resolved Sep 09, 2026, 04:46 PM UTC

    Type: Incident Duration: 1 hour and 20 minutes Affected Components: eu-west Storage Cluster, us-west Storage Cluster, EU Central Cache, EU West Cache, eu-west Storage Cluster, eu-west Storage Cluster, eu-central Storage Cluster, US East Cache, EU West Runtime Build Cache, Dashboard, eu-central Storage Cluster, eu-central Storage Cluster, EU Central Runtime Build Cache, US West Runtime Build Cache, US West Cache, us-west Storage Cluster, us-west Storage Cluster Sep 9, 16:46:00 GMT+0 - Investigating - We are currently experiencing a major outage with both our caching infrastructure and our dashboards. These components are affected across all regions. We are investigating the issue and will provide an update soon. Sep 9, 17:21:48 GMT+0 - Monitoring - We are seeing recovery of our caching infrastructure and our dashboard. We will continue to monitor this situation closely. Sep 9, 18:06:25 GMT+0 - Resolved - This incident has been resolved.

Read the full incident report →

Minor September 8, 2026

Job failures in EU West

Detected by Pingoru
Sep 08, 2026, 03:00 PM UTC
Resolved
Sep 08, 2026, 04:44 PM UTC
Duration
1h 43m
Affected: Github → API Requests
Timeline · 3 updates
  1. investigating Sep 08, 2026, 03:00 PM UTC

    Since approximately 12:20 UTC, some jobs in our eu-west region have been failing with the GitHub error "The self-hosted runner lost communication with the server"; other regions are unaffected. Affected jobs can be safely re-run while we continue to investigate the underlying cause.

  2. monitoring Sep 08, 2026, 03:31 PM UTC

    Between approximately 12:20 and 15:15 UTC, some jobs in our eu-west region failed with the GitHub error "The self-hosted runner lost communication with the server"; other regions were unaffected. Job failures in eu-west returned to normal levels as of 15:15 UTC, and any affected jobs can be safely re-run while we continue to investigate the underlying cause.

  3. resolved Sep 08, 2026, 04:44 PM UTC

    This incident is resolved. Job failures in our eu-west region returned to normal levels at 15:15 UTC and have remained there since; any jobs that failed between approximately 12:20 and 15:15 UTC with the error "The self-hosted runner lost communication with the server" can be safely re-run.

Read the full incident report →

Minor September 2, 2026

US West storage cluster degradation

Detected by Pingoru
Sep 02, 2026, 04:13 PM UTC
Resolved
Sep 02, 2026, 07:23 PM UTC
Duration
3h 9m
Affected: Incremental Docker Builders (us-west Storage Cluster)Docker Container Cache (us-west Storage Cluster)Sticky Disks (us-west Storage Cluster)
Timeline · 5 updates
  1. investigating Sep 02, 2026, 04:13 PM UTC

    We are investigating reports of higher latency on sticky disk operations in our us-west region. Customers running jobs in us-west may see slower incremental Docker builds, Git caching, and container caching, so affected jobs can take longer than usual to complete.

  2. investigating Sep 02, 2026, 04:48 PM UTC

    Sticky disk storage in our us-west region is experiencing higher latency; our other regions are not affected. Customers running jobs in us-west may see slower Git-cached checkouts, incremental Docker builds, and container caching. We are continuing to investigate the underlying cause and will provide another update within the next 30 minutes.

  3. investigating Sep 02, 2026, 05:22 PM UTC

    We are still working to resolve elevated latency on sticky disk storage in our us-west region; other regions are not affected. Customers running jobs in us-west may continue to see slower Git-cached checkouts, incremental Docker builds, and container caching. Our investigation into the underlying cause is ongoing and we will provide another update within the next 30 minutes.

  4. monitoring Sep 02, 2026, 06:10 PM UTC

    Our us-west region storage cluster's latency and error rates have returned to baseline as of approximately 17:15 UTC, following mitigations that reduce load on the affected storage. Git-cached checkouts, incremental Docker builds, and container caching in us-west are back to normal, and we are monitoring to confirm the recovery holds while we continue to investigate the underlying cause. We will provide a final update within the next hour.

  5. resolved Sep 02, 2026, 07:23 PM UTC

    This incident is resolved. Sticky disk storage in our us-west region came under more read load than it could serve at normal latency, which slowed and produced higher error rates for Git-cached checkouts, incremental Docker builds, and container caching. We reduced and redistributed that load, and performance has been normal since approximately 17:15 UTC. Jobs that failed during the incident can be safely re-run.

Read the full incident report →

Minor September 2, 2026

Service degradation in upstream GHCR registries

Detected by Pingoru
Sep 02, 2026, 02:09 PM UTC
Resolved
Sep 02, 2026, 04:30 PM UTC
Duration
2h 21m
Affected: Github → API Requests
Timeline · 4 updates
  1. investigating Sep 02, 2026, 02:09 PM UTC

    We are currently observing higher rates of errors for the upstream GHCR registries which may indicate an undeclared GitHub incident. We are currently monitoring.

  2. investigating Sep 02, 2026, 02:20 PM UTC

    We are seeing evidence that Git Checkouts are also affected by this. We are seeing high TCP retransmit rates into the EU GitHub loadbalancer. Checkouts in the EU West, EU Central, US East regions are affected.

  3. monitoring Sep 02, 2026, 03:43 PM UTC

    Connections to [github.com](http://github.com) and [ghcr.io](http://ghcr.io) from our eu-west, eu-central, and us-east regions have returned to normal as of approximately 15:00 UTC, so checkouts are completing at normal speed and login failures to GitHub Container Registry have dropped to baseline levels. We are continuing to monitor, and re-running any jobs that failed during this period should succeed.

  4. resolved Sep 02, 2026, 04:30 PM UTC

    This incident has been resolved. GitHub connectivity from our EU regions was degraded from \~12:45 to 16:15 UTC, affecting checkouts, [ghcr.io](http://ghcr.io), and some eu-west job starts. Normal since 16:15\. Affected jobs can be re-run.

Read the full incident report →

Notice August 18, 2026

Requests to GitHub failing due to upstream incident

Detected by Pingoru
Aug 18, 2026, 08:45 PM UTC
Resolved
Aug 18, 2026, 11:14 PM UTC
Duration
2h 29m
Timeline · 2 updates
  1. monitoring Aug 18, 2026, 08:45 PM UTC

    Jobs failure rate at an elevated rate caused by an upstream GitHub being rejected by 429 rate-limit errors. We are seeing a single-digit percentage increase in job failure rate across all organizations and are seeing similar failures on non-Blacksmith infrastructure.

  2. resolved Aug 18, 2026, 11:14 PM UTC

    This incident has been resolved.

Read the full incident report →

Minor August 17, 2026

Job adoption delays due to missed webhooks

Detected by Pingoru
Aug 17, 2026, 06:58 PM UTC
Resolved
Aug 17, 2026, 07:00 PM UTC
Duration
1m
Affected: EU Central (eu-central ARM)EU Central (eu-central x86)US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)US Central (us-central MacOS)EU West (eu-west ARM)US East (us-east x86)
Timeline · 2 updates
  1. investigating Aug 17, 2026, 06:58 PM UTC

    We are seeing an increased error rate processing webhook events which will lead to delays in adoption jobs. We are actively investigating.

  2. resolved Aug 17, 2026, 07:00 PM UTC

    This incident has been resolved.

Read the full incident report →

Major August 17, 2026

GitHub outage affecting job failures and dashboard errors

Detected by Pingoru
Aug 17, 2026, 01:43 PM UTC
Resolved
Aug 17, 2026, 08:40 PM UTC
Duration
6h 57m
Affected: Github → API RequestsGithub → WebhooksDashboard
Timeline · 6 updates
  1. identified Aug 17, 2026, 01:43 PM UTC

    The Blacksmith Dashboard is unable to load. We've identified the root cause to upstream 503s being returned from GitHub on permission-check requests.

  2. identified Aug 17, 2026, 01:46 PM UTC

    Upstream incident has been declared: We are also seeing elevated error rates in jobs as they hit upstream GitHub errors. Job adoption times are also affected and are delayed. We are monitoring and are looking at potential mitigations.

  3. monitoring Aug 17, 2026, 04:56 PM UTC

    We're seeing signs of GitHub recovery. The dashboard is now loading and jobs should be running again. We are monitoring the recovery.

  4. monitoring Aug 17, 2026, 05:52 PM UTC

    We are still seeing intermitting GitHub API errors at a low rate.

  5. monitoring Aug 17, 2026, 07:32 PM UTC

    We are no longer seeing upstream errors, and are continuing to monitor impact of the upstream outage.

  6. resolved Aug 17, 2026, 08:40 PM UTC

    This incident has been resolved.

Read the full incident report →

Minor August 13, 2026

Storage degradation in us-west

Detected by Pingoru
Aug 13, 2026, 11:40 AM UTC
Resolved
Aug 14, 2026, 06:14 PM UTC
Duration
1d 6h
Affected: EU Central (eu-central ARM)EU Central (eu-central x86)US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)Actions CacheGithub → WebhooksIncremental Docker Builders (us-west Storage Cluster)Docker Container Cache (us-west Storage Cluster)Sticky Disks (us-west Storage Cluster)EU West (eu-west ARM)US East (us-east x86)Actions Cache (US West Cache)Actions Cache (US East Cache)Actions Cache (EU West Cache)Actions Cache (EU Central Cache)
Timeline · 28 updates

Read the full incident report →

Minor August 12, 2026

Elevated failures downloading GitHub release assets

Detected by Pingoru
Aug 12, 2026, 06:41 PM UTC
Resolved
Aug 13, 2026, 01:20 AM UTC
Duration
6h 38m
Affected: EU Central (eu-central ARM)EU Central (eu-central x86)US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)US Central (us-central MacOS)Github → ActionsEU West (eu-west ARM)
Timeline · 10 updates

Read the full incident report →

Minor August 12, 2026

Hanging apt package installs on us-west runners due to an upstream mirror issue

Detected by Pingoru
Aug 12, 2026, 03:30 AM UTC
Resolved
Aug 12, 2026, 03:30 AM UTC
Duration
—
Timeline · 1 update
  1. resolved Aug 12, 2026, 03:30 AM UTC

    Type: Incident Duration: 8 hours and 21 minutes Affected Components: us-west x86 Aug 12, 03:30:00 GMT+0 - Investigating - We are currently investigating this incident. Aug 12, 09:10:10 GMT+0 - Identified - We have identified the affected mirror and are implementing a fix Aug 12, 10:33:01 GMT+0 - Identified - We are currently deploying a mitigation Aug 12, 11:26:10 GMT+0 - Monitoring - We implemented a fix and are seeing improvements. We are continuing to monitor the result. Aug 12, 11:50:47 GMT+0 - Resolved - This incident has been resolved.

Read the full incident report →

Minor August 11, 2026

Increased action cache miss rate for certain customers

Detected by Pingoru
Aug 11, 2026, 07:00 PM UTC
Resolved
Aug 11, 2026, 08:00 PM UTC
Duration
59m
Affected: Actions Cache
Timeline · 2 updates
  1. investigating Aug 11, 2026, 07:00 PM UTC

    We are observing certain customers experiencing higher than baseline cache miss rates.

  2. resolved Aug 11, 2026, 08:00 PM UTC

    We resolved the configuration error leading to the increased rate of cache misses. Cache hit rates are now at the baseline.

Read the full incident report →

Minor August 10, 2026

Degraded job performance, metrics, and log ingestion in us-west

Detected by Pingoru
Aug 10, 2026, 07:17 PM UTC
Resolved
Aug 10, 2026, 08:44 PM UTC
Duration
1h 26m
Affected: US West (us-west ARM)US West (us-west x86)Dashboard
Timeline · 5 updates
  1. investigating Aug 10, 2026, 07:17 PM UTC

    We are experiencing degradation in our metrics and log ingestion services in our us-west region. We are actively investigating the issue.

  2. investigating Aug 10, 2026, 07:44 PM UTC

    We are continuing to investigate degraded metrics and log ingestion in our us-west region. Customers may still see metrics and logs for their jobs appear missing or delayed in the Blacksmith dashboard, while other regions remain unaffected. We will provide another update within the next 30 minutes.

  3. investigating Aug 10, 2026, 07:45 PM UTC

    We are experiencing degraded network performance in our us-west region, affecting metrics and log ingestion as well as job performance. Jobs in us-west that upload artifacts or transfer large amounts of data may run slower than normal and in some cases hit their configured timeouts and fail. Other regions are not affected, and we are actively investigating the issue.

  4. monitoring Aug 10, 2026, 08:18 PM UTC

    Metrics and log ingestion in our us-west region is recovering, and job performance in the region has returned to normal. We are monitoring to confirm the recovery holds and are continuing to investigate the underlying cause. We will provide an update shortly.

  5. resolved Aug 10, 2026, 08:44 PM UTC

    This incident is resolved, with job performance and the ingestion of metrics and logs in our us-west region stable for the past 30 minutes. Jobs that failed or timed out during the incident can be safely re-run.

Read the full incident report →

Major August 6, 2026

Github → Actions experiencing degraded performance

Detected by Pingoru
Aug 06, 2026, 03:30 PM UTC
Resolved
Aug 07, 2026, 12:57 AM UTC
Duration
9h 26m
Affected: Github → ActionsGithub → Webhooks
Timeline · 2 updates
  1. monitoring Aug 06, 2026, 03:30 PM UTC

    Github has reported degraded performance for Actions, jobs may take a moment to be adopted. We are monitoring this incident.

  2. resolved Aug 07, 2026, 12:57 AM UTC

    GitHub job success and adoption rates are now back at normal levels. Jobs that were not adopted during the incident are being requeued by our team and should be picked up shortly.

Read the full incident report →

Major August 6, 2026

Jobs not getting picked up

Detected by Pingoru
Aug 06, 2026, 01:18 AM UTC
Resolved
Aug 06, 2026, 02:53 AM UTC
Duration
1h 34m
Affected: EU Central (eu-central ARM)EU Central (eu-central x86)US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)Actions CacheUS Central (us-central MacOS)WebsiteIncremental Docker Builders (eu-central Storage Cluster)Incremental Docker Builders (us-west Storage Cluster)Docker Container Cache (eu-central Storage Cluster)Docker Container Cache (us-west Storage Cluster)Sticky Disks (eu-central Storage Cluster)Sticky Disks (us-west Storage Cluster)Sticky Disks (eu-west Storage Cluster)Incremental Docker Builders (eu-west Storage Cluster)Docker Container Cache (eu-west Storage Cluster)Website (https://blacksmith.sh)CodesmithEU West (eu-west ARM)Runtime Build Caching
Timeline · 6 updates
  1. investigating Aug 06, 2026, 01:18 AM UTC

    We're currently experiencing an outage of our control plane. GitHub job adoption and execution are affected, as well as dashboard access.

  2. identified Aug 06, 2026, 01:36 AM UTC

    We've mitigated the root cause and jobs are resuming to run. Some jobs may still be delayed to start as we catch up with the job backlog.

  3. monitoring Aug 06, 2026, 01:51 AM UTC

    Jobs adoption has recovered and are operational.

  4. monitoring Aug 06, 2026, 02:11 AM UTC

    There is still a delay in job adoption times as we continue to recover.

  5. monitoring Aug 06, 2026, 02:45 AM UTC

    Job adoption is now fully operational.

  6. resolved Aug 06, 2026, 02:53 AM UTC

    This incident has been resolved.

Read the full incident report →

Minor August 5, 2026

Elevated error rates in Sticky Disk and Actions Cache requests

Detected by Pingoru
Aug 05, 2026, 10:45 PM UTC
Resolved
Aug 05, 2026, 10:45 PM UTC
Duration
—
Timeline · 1 update
  1. resolved Aug 05, 2026, 10:45 PM UTC

    Type: Incident Duration: 53 minutes Affected Components: eu-west Storage Cluster, us-west Storage Cluster, , eu-west Storage Cluster, eu-west Storage Cluster, , eu-central Storage Cluster, eu-central Storage Cluster, eu-central Storage Cluster, us-west Storage Cluster, us-west Storage Cluster, Runtime Build Caching → Actions Cache → Aug 5, 22:45:00 GMT+0 - Investigating - We are seeing elevated error rates for Sticky Disk and Actions Cache requests. We are investigating. Aug 5, 22:55:00 GMT+0 - Monitoring - We implemented a fix and are currently monitoring the result. Aug 5, 23:30:42 GMT+0 - Monitoring - We are still monitoring recovery. Aug 5, 23:37:56 GMT+0 - Resolved - This incident has been resolved.

Read the full incident report →

Minor August 3, 2026

Delay in Job Adoption

Detected by Pingoru
Aug 03, 2026, 03:37 PM UTC
Resolved
Aug 03, 2026, 03:42 PM UTC
Duration
5m
Affected: EU Central (eu-central ARM)EU Central (eu-central x86)US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)US Central (us-central MacOS)EU West (eu-west ARM)
Timeline · 2 updates
  1. investigating Aug 03, 2026, 03:37 PM UTC

    We're seeing a number of cases where jobs may not be adopted promptly. We're currently investigating.

  2. resolved Aug 03, 2026, 03:42 PM UTC

    This incident has been resolved.

Read the full incident report →

Minor August 2, 2026

US-West storage cluster degraded

Detected by Pingoru
Aug 02, 2026, 06:16 AM UTC
Resolved
Aug 02, 2026, 06:37 AM UTC
Duration
20m
Affected: Incremental Docker Builders (us-west Storage Cluster)Docker Container Cache (us-west Storage Cluster)Sticky Disks (us-west Storage Cluster)
Timeline · 3 updates
  1. investigating Aug 02, 2026, 06:16 AM UTC

    We are seeing some failures with sticky disk availability and write latency with the US-West storage cluster.

  2. monitoring Aug 02, 2026, 06:25 AM UTC

    We've applied a fix and are seeing failure rates and latency starting to come down.

  3. resolved Aug 02, 2026, 06:37 AM UTC

    This incident has been resolved.

Read the full incident report →