Is Sardine down?
Last checked 6m agoNo incidents right now.
Sardine is operational right now. Last checked 6m ago; the most recent incident resolved 16h ago.
Real-time Sardine status, recent outages, and incident history — pulled directly from Sardine's official status page at https://status.sardine.ai every 5 minutes. Pingoru tracks 9 Sardine services and has captured 9 incidents in the last 90 days (97.78% uptime). Get email, Slack, Discord, or webhook alerts the moment Sardine reports a new incident — free for 3 monitors, no credit card.
Recent outages & incidents
Past 90 days- Device APIsCustomer APIsBusiness APIFeedback API
Timeline · 5 updates
- investigating · Oct 01, 2026, 03:36 PM UTC
We are investigating elevated latency and increased response times affecting Customer API, Feedback API and Business API requests in our PROD-EU environment. Our engineering team is actively investigating, and we will provide an update as soon as we have more information.
- investigating · Oct 01, 2026, 03:39 PM UTC
We are continuing to investigate this issue.
- identified · Oct 01, 2026, 04:39 PM UTC
A mitigation has been applied and latency is improving. We are continuing to monitor and will post another update later.
- monitoring · Oct 01, 2026, 05:03 PM UTC
A fix has been implemented and we are monitoring the results.
- resolved · Oct 01, 2026, 09:18 PM UTC
The elevated latency and increased response times affecting the APIs in PROD-EU have been resolved. Response times have returned to normal levels and have remained stable since the fix was applied. We will continue to monitor the environment closely. A root cause analysis (RCA) will be shared once it is complete. We apologize for the disruption.
Latest: The elevated latency and increased response times affecting the APIs in PROD-EU have been resolved. Response times have returned to normal levels and have remained stable since the…
-
- Issuing APICustomer APIs
Timeline · 4 updates
- investigating · Sep 27, 2026, 09:56 PM UTC
Australia instance (api.au.sardine.ai) is degraded for v1/customers API and v1/issuing/risks API. Team is actively looking into it.
- monitoring · Sep 27, 2026, 10:21 PM UTC
A fix has been implemented and we are monitoring the results.
- resolved · Sep 27, 2026, 10:36 PM UTC
This incident has been resolved. We'll follow up with post mortem.
- postmortem · Sep 29, 2026, 09:43 PM UTC
## Summary On September 27–28, 2026, Sardine experienced two periods of elevated latency and request timeouts in our **Australian production region**. Requests to `/v1/issuing/risks` , `/v1/customers`, and `/v1/feedbacks` were affected from 12:16–22:00 UTC on September 27 and again from 23:20 UTC on September 27 until 00:23 UTC on September 28. A separate high-volume data-processing workload overloaded a shared regional database. An unbounded historical aggregation initially consumed most of the database capacity. As records accumulated, stale database statistics caused another frequently used transaction lookup to select an inefficient query plan. Query timeouts then recycled database connections, repeatedly reintroducing that plan and amplifying the load. The incident affected API availability, transaction persistence, and the risk level evaluation. ## Timeline _All times UTC._ * **September 27, 12:16** — A high-volume data-processing workload begins. Database CPU reaches capacity and `/v1/issuing/risks` timeouts start. * **14:20** — We receive the first customer report. * **14:29** — The primary support on-call acknowledges a page, but the issue is not escalated effectively. * **18:02** — We send the first substantive customer impact update. * **18:05** — The need for engineering escalation is identified. * **21:06** — Sardine declares an incident. * **21:38** — Active engineering investigation and remediation begin. * **22:03** — A database resize and primary switchover restore customer-facing service. * **22:12** — The high-volume workload is paused and database utilization returns to normal. * **22:40** — The first event is marked resolved. * **22:49** — Database CPU monitors are added for all production regions. * **23:20** — The workload resumes and triggers a second degradation. * **23:41** — The new database monitor pages the on-call team and is acknowledged within 13 seconds. * **September 28 00:14** — A second database resize and primary switchover resolve the inefficient transaction lookup. * **00:23** — Optional aggregation work is disabled and the remaining customer-facing timeouts end. * **00:26 onward** — Database statistics are refreshed and an optimized composite index is deployed. ## Root Cause The incident resulted from the interaction of four conditions: 1. **Unbounded aggregation:** A synchronous query scanned the full transaction history for given condition. This consumed most database capacity when the high-volume workload began. 2. **Inefficient transaction lookup:** As new records accumulated, database statistics did not refresh quickly enough. The query planner selected an index that scanned a customer’s growing transaction history rather than performing a targeted lookup. 3. **Connection recycling:** Timed-out queries caused database connections to close. Replacement connections repeatedly began with the same inefficient query plan, creating a feedback loop of expensive scans, timeouts, and reconnects. ## Contributing Factors * Our API monitoring excluded the relevant client-timeout responses, and we did not alert on query cancellations or sustained database-connection churn. Database CPU monitor was only configured for non-pagerduty alert * The customer detected the first event before our monitoring did. Delayed engineering escalation extended the response time. * Our first recovery addressed the immediate pressure but did not prevent the workload from triggering a second event when it resumed. ## Resolution and Recovery For the first event, we resized the database and switched to the new primary, then paused the high-volume workload. For the second event, we again resized and switched the primary, then disabled optional aggregation to remove the remaining load. After service recovery, we: * Refreshed database statistics. * Deployed a composite index that gives both new and established connections an efficient transaction-lookup path. * Added database CPU monitors for every production region. * Removed optional, unused aggregation from the `/v1/customers` path * Kept the Australian production database at a higher baseline capacity while we complete the permanent query fixes and capacity review. ## Corrective Actions We are taking the following actions: ### Database and application resilience * Replace the affected transaction lookup with an atomic, consistently indexed operation. * Skip unnecessary historical aggregation and bound any required fallback query. * Evaluate a read pool so issuing-risk reads do not compete with customer-event writes on the primary. * Review steady-state database capacity after the query fixes are complete. ### Monitoring and response * Expand monitoring for client timeouts and database metrics * Improve incident routing so customer reports, support tickets, and on-call escalation converge quickly.
Latest: ## Summary On September 27–28, 2026, Sardine experienced two periods of elevated latency and request timeouts in our **Australian production region**. Requests to `/v1/issuing/risk…
-
- Dashboard
Timeline · 2 updates
- identified · Sep 17, 2026, 08:12 AM UTC
Severity: Minor Status: Identified / Investigating Affected Component: Connection Graph (Production - EU) Impact: Users may be unable to view Connection Graph data for the past three days. Graphs generated during this window are not currently visible. No other functionality is affected. Description: We have identified a delay in the Connection Graph pipeline in our EU production environment, which began approximately three days ago. As a result, graph data from this period is not currently reflected. We are actively investigating the root cause and working on a resolution. We will post updates here as more information becomes available.
- resolved · Sep 17, 2026, 11:25 AM UTC
This incident has been resolved.
Latest: This incident has been resolved.
-
- Customer APIsDashboard
Timeline · 3 updates
- investigating · Sep 16, 2026, 12:08 PM UTC
The Sandbox EU environment is experiencing degraded performance. We are working on resolution.
- investigating · Sep 16, 2026, 12:08 PM UTC
We are continuing to investigate this issue.
- resolved · Sep 16, 2026, 03:19 PM UTC
This incident has been resolved.
Latest: This incident has been resolved.
-
- Customer APIs
Timeline · 7 updates
- investigating · Sep 04, 2026, 10:19 PM UTC
We are currently investigating an issue affecting Advanced Aggregations, where some customers may experience inaccurate results. We will post updates as we learn more.
- investigating · Sep 04, 2026, 11:11 PM UTC
We are continuing to investigate this issue
- identified · Sep 05, 2026, 12:38 AM UTC
We have identified the root cause of the issue affecting Advanced Aggregations and our engineering team is actively working on a fix. We will provide a further update once the fix has been deployed and verified.
- identified · Sep 05, 2026, 01:49 AM UTC
team is actively looking into it. Issue is partially mitigated, and we're still working on fixing the root cause.
- monitoring · Sep 05, 2026, 03:15 AM UTC
fix is being validated and being rolled out
- resolved · Sep 05, 2026, 03:50 AM UTC
incident is resolved. We'll provide post mortem.
- postmortem · Sep 08, 2026, 10:48 PM UTC
### Summary On September 4–5, 2026 \(UTC\), we experienced an issue affecting **Advanced Aggregations** **in which some customers saw inaccurate aggregation results because recent card transaction data may have been missing**. We fixed **the issue, and backfilled the data**. Advanced Aggregations results are expected to be correct again after **2026-09-05 03:50 UTC**. ### What happened A subset of card transaction records failed to be written into one of our tables because the column used to store the payment method id could not support the size of newly generated card IDs. As a result, those specific transactions were missing from the aggregated dataset, which could lead to incorrect results for customers ingesting this type of data. An impact assessment is available upon request. No data was permanently lost: failed writes were routed to a durable deadletter queue for later reprocessing. After mitigation was deployed, we replayed the queued events and completed a backfill so that aggregations reflected the correct underlying transaction data. Incident started at 2026-09-04 01:02 UTC and resolved at 2026-09-05 03:50 UTC. ### Why it happened The root cause was a **data type mismatch**: the `cards` table `id` values grew beyond the maximum value supported by a **4-byte integer**. `id` field has 8-byte integer so there was no error storing cards. However, `card_id` field in the Advanced Aggregations schema was 4-byte, causing inserts to fail for transactions related to newer cards. To remediate safely at scale, we: * Added new **8-byte integer \(**`bigint`\) columns for the affected identifiers and updated write logic to use them. * Introduced a **compatibility read path** so reads return the correct identifier whether data is stored in the old or new columns, along with new indexes to preserve query performance. ### What we are doing about it * We’re reviewing improvements to monitor schema limit risks earlier and prevent similar issues in the future. * We’ll enhance monitoring on deadletter queue so we can detect issue like this sooner.
Latest: ### Summary On September 4–5, 2026 \(UTC\), we experienced an issue affecting **Advanced Aggregations** **in which some customers saw inaccurate aggregation results because recent …
-
See the full Sardine outage history
3 more incidents in the last 90 days, plus the full multi-year archive of per-service events and update timelines.
Browse Sardine outage history →Or sign up free to get alerts when Sardine breaks · 10 free monitors · No credit card
- Started Oct 01, 2026, 02:49 PM UTC · Resolved Oct 01, 2026, 07:58 PM UTC · 5h 8m
- Prod Australia instance performance degradation (v1/customers API and v1/issuing/risks API) ResolvedStarted Sep 27, 2026, 09:56 PM UTC · Resolved Sep 27, 2026, 10:36 PM UTC · 40m
- Started Sep 17, 2026, 08:12 AM UTC · Resolved Sep 17, 2026, 11:25 AM UTC · 3h 13m
- Sandbox EU degradation ResolvedStarted Sep 16, 2026, 12:08 PM UTC · Resolved Sep 16, 2026, 03:19 PM UTC · 3h 10m
- Started Sep 04, 2026, 10:19 PM UTC · Resolved Sep 05, 2026, 03:50 AM UTC · 5h 30m
- Started Jul 24, 2026, 12:42 AM UTC · Resolved Jul 24, 2026, 02:07 AM UTC · 1h 25m
- Started Jul 16, 2026, 12:40 PM UTC · Resolved Jul 16, 2026, 02:25 PM UTC · 1h 45m
- Performance degradation on APIs ResolvedStarted Jul 15, 2026, 07:56 PM UTC · Resolved Jul 15, 2026, 07:56 PM UTC · —