Website and services partially unavailable due to high database load

Incident Report for Expo

Postmortem

The primary database was under high CPU load from 1:11PM-1:29PM PT. The impact of this is that our website and API were partially unavailable (more precisely, at 80-90% availability) during that time.

Scope

What was affected: This affected most services including the website, builds, submissions, CI/CD jobs, and publishing updates.

What wasn’t affected: Serving updates to end users maintained 100.000% availability. App performance measurements sent to Observe were also unaffected.

Root cause analysis

An internal orchestrator service used to update the state of CI/CD jobs, including build jobs and store submission jobs, was updated to retry upon application-level failures. We are investigating more thoroughly and currently believe this led to a cascade of retries that caused more failures once the database was under more load than it could handle.

Remediation

The commit that changed retry behavior was rolled back and we are clearing the queue of retries.

For a more robust solution, the internal orchestrator must gracefully handle backpressure from the application server and database, and to apply more backoff.

Posted Oct 05, 2026 - 13:53 PDT

Resolved

Due to high database load, the Expo website and several services including submitting CI/CD jobs, build jobs, and store submissions were partially unavailable. The issue has subsided as of 1:30PM PT and we are continuing to diagnose the root cause and monitor the health of the services.
Posted Oct 05, 2026 - 13:11 PDT