Back to overview
Resolved

SnowBee is experiencing degraded performance and partial downtime

Oct 2, 2026 at 1:03pm UTC
Affected services
Public API
Internal API

Resolved
Oct 2, 2026 at 2:22pm UTC

At 15:00:02 Oslo time, a faulty deployment of our webapp frontend caused an extremely elevated error rate and partial unavailability. At 15:32:50 Oslo time, the issue was fully resolved.

The root cause was a bad version of the webapp frontend that caused our API to become overloaded. A software bug caused a much larger number of API calls to happen than what is usual. Endpoints that typically sits at 30-40 or so request per minute now received almost 10 000 per minute, a rough 270x increase in traffic.

We became aware of the issue within just a few seonds after the incident started, when error logs started flooding our monitoring channels.

At 15:14 Oslo time, we deployed a new version of the webapp frontend that fixed the issue. However, the problem continued affecting our systems, as all webapp frontend clients had to be refreshed in order to stop sending the flood of API calls.

At 15:32:50, we deployed a firewall rule that blocked the specific version of the webapp frontend that caused the issue. That instantly stopped the flood of API calls to reach our backend, and restored operations to normal.

In terms of the quantity of successful retail sales compared to normal operations, the system experienced a 60% reduction between 15:00 and 15:10, and 90% reduction between 15:20 and 15:30, up until operations were restored 15:32:50.

Here are the core issues that caused this prolonged period of degraded performance and partial unavailability, and what we'll do to mitigate issues like this in the future:

We need faster emergency code updates in production. The initial fix was deployed to production 14 minutes after the incident started, but the fix was implemented at 15:03 Oslo time, only 3 minutes after the incident started. Our deployment process includes a handful of automated testing that takes extra time. On top of that, another build was already in the build pipeline that we cancelled, but the cancellation process is asynchronous and caused the actual build to 15:06 Oslo time, a delay of 3 extra minutes. We have documented and tested the procedure to circumvent this build process so we have the ability to emergency deploy and replace known poison pill versions. That would have caused the new version to be out at an estimate of 15:05, instead of 15:14.

We need a streamlined process to block poison pill versions at the firewall level. After the poison pill version was replaced with a deployed fix, It still took a little more than 15 minutes until we had the firewall rule installed that blocked the poison pill. That's because we simply had no plan in place, and had to figure out the process to implement a firewall rule properly without blocking the entire system. For monitoring causes, we've already set up the system so that each API call that we make precisely identifies exactly which version of the software that performed the call, so we could rely on that for the firewall rule. But we were initially not sure at which layer in the networking stack the rule should be installed, and exactly how to operationally implement the rule in a safe way. We have documented and automated the procedure to block poison pill versions.

In short, if an issue like this were to happen again, we'll now be able to deploy an update to replace the poison pill within 5 minutes, and when the update is deployed we can instantly block the poison pill version, causing an immediate full recovery. If we had that in place for this incident, the issue would only have lasted an estimated 5 minutes before fully recovering.

Updated
Oct 2, 2026 at 1:41pm UTC

Internal API recovered.

Updated
Oct 2, 2026 at 1:38pm UTC

Public API recovered.

Updated
Oct 2, 2026 at 1:06pm UTC

Internal API went down.

Updated
Oct 2, 2026 at 1:05pm UTC

Public API went down.

Updated
Oct 2, 2026 at 1:04pm UTC

Internal API recovered.

Created
Oct 2, 2026 at 1:03pm UTC

Internal API went down.