The primary database was under high CPU load from 1:11PM-1:29PM PT. The impact of this is that our website and API were partially unavailable (more precisely, at 80-90% availability) during that time.
What was affected: This affected most services including the website, builds, submissions, CI/CD jobs, and publishing updates.
What wasn’t affected: Serving updates to end users maintained 100.000% availability. App performance measurements sent to Observe were also unaffected.
An internal orchestrator service used to update the state of CI/CD jobs, including build jobs and store submission jobs, was updated to retry upon application-level failures. We are investigating more thoroughly and currently believe this led to a cascade of retries that caused more failures once the database was under more load than it could handle.
The commit that changed retry behavior was rolled back and we are clearing the queue of retries.
For a more robust solution, the internal orchestrator must gracefully handle backpressure from the application server and database, and to apply more backoff.