On September 18, between 22:40–23:03 UTC, WorkOS experienced a partial outage affecting authentication and other API requests.
A failover of the cache used for feature flag evaluation left app connections stalled, and requests waiting on flag checks held database connections open, causing queues to build and other requests to fail.
Many of you experienced failures, including failed sign-ins and long delays. We're sorry for the disruption this caused. We're taking this event seriously and have already fixed the root cause, with work underway on additional layers of protection.
We will publish a full RCA next week explaining in detail what failed, why recovery was not automatic (should have been ~seconds), and what we're changing to prevent this issue from happening again.