How Uber Protects Against Retry Storms
iscmt
78 points
33 comments
September 17, 2026
Related Discussions
Found 5 related stories in 109.4ms across 7,105 title embeddings via pgvector HNSW
- Dutch regulator fines Uber €825M for letting AI deactivate driver accounts biglyburrito · 21 pts · August 22, 2026 · 48% similar
- Uber is lobbying to keep robotaxi rides 85% human in New Jersey logickkk1 · 15 pts · July 12, 2026 · 47% similar
- Uber fined $966M in NL for automating driver suspensions mattashii · 20 pts · August 21, 2026 · 45% similar
- Uber ordered to pay $40M over death of woman left on Southern California freeway hentrep · 22 pts · September 18, 2026 · 44% similar
- Uber Faces €825M Dutch Fine over Driver Suspensions tcp_handshaker · 12 pts · August 21, 2026 · 44% similar
Discussion Highlights (7 comments)
aftbit
I'd be interested to hear other strategies in this space. I've done the naive thing of allowing retries everywhere, and gotten into retry storms. When I was next presented with the problem, I tried the other naive thing of only allowing retries from the very top level service, which led me to redoing absolutely tons of work for each failure. What's a nice middle path that doesn't add too much complexity?
maxchisto
I'm suspicious of load shedding not mentioned in the article. Combine that with exp backoff in the caller and you got yourself a pretty robust starting point
whoevercares
Token bucket is all you need
prologic
So, effectively if A → B → C → D and D is failing, C may retry D, but B and A are discouraged from retrying the whole chain. This is quite slever. I also really like the concept of an "Error Budget", inspired by SRE and SLO(s) no doubt :)
Scoundreller
Meanwhile Google keeps giving me “please wait, do not reload page” walls, so I ctrl-r as rapidly as possible. Or is that the human test and response?
UltraSane
This feels like trying to reinvent Fibre Channel's flow control mechanism.
whatever1
Easy. Take a larger cut from the driver for each retry.