Cloud Atlas replaces per-instance rate limits with fleet-wide protection
Cloud Atlas is the globally distributed identity data store behind a large share of Salesforce logins and token validations, run to five nines of availability. In this interview, the team’s engineering lead explains why overload at that layer is dangerous: slow responses cause upstream services to retry, the retries add traffic to a system already under pressure, and a local bottleneck spreads. Agent-driven and automated workloads have made traffic burstier and harder to predict than human logins.
The old protection was per-instance rate limiting. That held up while infrastructure was static, but with autoscaling each server’s limits went stale whenever instances were added or removed, and each server knew only its own load rather than a customer’s total usage across the fleet. Busy servers rejected requests while capacity sat idle elsewhere, and one tenant’s spike could still starve others.
The redesign protects the service as a whole instead of individual servers. It combines global quota management that needs no central coordinator with load shedding that acts before retries begin to amplify the overload, so a single customer’s surge is isolated rather than shared.