Umbrella: the tool that saved Amazon $250M
When a storm is coming, Amazon extends delivery promises before the network slips. Umbrella is the platform that decides where those pads land, prices what each one costs, and explains a model's decision to the person who gets asked about it.
Accuracy is not speed
- Amazon's network is measured on speed, cost and Delivery Estimate Accuracy, and they pull against each other
- Accuracy: did the package arrive when the customer was told it would. Told Thursday, arrives Friday: a miss. Moved to Friday on Wednesday: accurate.
- The promise can move. It just has to move before it expires.
- The one thing Amazon can't automate is the weather
The promise pad lifecycle
Extend the promise before it expires, and it still holds when the weather hits.
A promise pad is a purchase: you spend speed to buy accuracy, and someone's targets pay for it.
The same storm was getting three different answers
Advanced machine learning system
Tramontane placed pads automatically from a machine learning model. Fast, hard to argue with, and it only handled snow.
Retrospective based system
A separate system with its own severity thresholds, built independently, and retrospective.
No system at all
Every pad was worked out and placed by hand, event after event.
None of them could see the others, and nobody could say what was happening across the network without assembling it by hand.
How might we automate a decision this consequential, without taking away the control people need to defend it?
Three groups wanted opposite things from the same product
- A week-long workshop in Seattle with operations teams from both regions, with engineers in the room from day one
- That's how the snow-only constraint surfaced in week one, not in a sprint review
- Configurators wanted depth. Leaders wanted breadth. Mid-event operators wanted speed, and nothing else, all competing for the same screen.
Three personas came out of the workshop, wanting opposite things
- Configurators write the rules before the season. They need depth: every parameter, every threshold.
- Senior leaders have one question: is my network about to be hit. They need breadth, in one glance.
- Manual overriders place pads mid-event, with a storm tracking in. They need speed, and nothing else.
We scored the three tools before touching anything
Tramontane · North America
The strongest of the estate: operators genuinely liked seeing pads placed on a map of the US. What held it back was snow-only coverage, and reasoning nobody could explain upward.
Rulebook system · Europe
Retrospective by design: it told teams what they should have done after the storm had been and gone.
Manual process · Canada
Nothing to put in front of a user: every pad was worked out by hand in spreadsheets, so there was no tool to score at all.
The industry average SUS is 68. Neither tool came close, and one region had nothing to measure in the first place.
The scores had voices behind them
The map is the one thing I'd keep: you can see the pads land across the US. But it only speaks snow, and when a wind event hits, or leadership asks why a pad exists, I'm on my own.
By the time the system speaks, the storm has been and gone. We are always reading last week's answer.
Every event is a night of copy and paste. One typo in a postcode and the wrong city gets padded.
Four findings came back. Every decision traces to one.
Needs pulled in opposite directions
Deep control, one answered question, or pure speed, depending on who you asked.
Nobody wanted full automation
Handle the routine 80%, but leave real control over the 20% that needs judgement.
Reporting was eating evenings
Leaders assembled weather impact reports by hand for every event, so the picture was always late.
The key trade-off was invisible
Teams chose between protecting the customer and protecting speed constantly, without ever seeing the price of the choice.
Cover every weather type from day one, and let models earn their place
| Weather type | Placement at launch | Automation status |
|---|---|---|
| Snow | Tramontane ML model | Model live |
| High wind | Retrospective rules | Model in training |
| Heavy rain & flooding | Retrospective rules | Model planned |
| Ice & freezing rain | Retrospective rules | Rules only |
| Extreme heat | Retrospective rules | Rules only |
| Hurricanes & named storms | Rules + manual judgement | Rules only |
The requirement that fell out of this: the product cannot care how a pad was placed. Model, rule or human, every pad lands in the same list with the same controls, so a weather type graduating to a model changes nothing for the operator.
The tenets of Umbrella
You only need one Umbrella
One product for every region, every weather type, and every way a pad gets placed. No more three answers to the same storm.
Umbrella should be low touch
You set a strategy before the season, and Umbrella does the rest. The routine 80% should never need a human hand.
Umbrella should report
What was protected and what it cost, compiled by nobody. The picture leaders used to assemble by hand, always there.
Umbrella should allow manual override
Automation people cannot overrule is worse than none. A person can always step in, and their pad gets the same standing as the model's.
Every decision that follows was argued against these four. One of them I knowingly broke, and the deck says where.
One model of the work, before any screens
- Three regions had three different definitions of what a “pad” even was
- Before any UI: one object model, the objects, the states, the language, agreed with engineering
- Once that existed the interface followed. Engineers built one thing, not a nicer version of three.
Four decisions, and what each one cost
For every one, the option I did not take, and what choosing the other one cost us. That second part is usually the bit that goes missing.
One product layered by depth, not three tailored views
- Separate views tested beautifully with each group alone
- They fell apart when a leader asked an operator why is that pad there: two products, nothing in common to point at
- One structure, three depths: cockpit → protections → configuration
The override belongs next to the result, not in a settings screen
- Automation people cannot overrule is worse than none
- System pads and human pads: one list, same controls on both
- Accuracy and speed impact side by side on every row; the explanation one click away
Show the real reasoning, and translate it in place
- Operators answer for postcodes; the model thinks in its own weather regions. People were defending a decision they could not see.
- The model's figures, unedited, with a plain-language glossary directly underneath
- A friendly summary would have gone out of date the first time the model was retrained
Price the pad before it counts, not after
- The accuracy–speed trade was happening constantly, invisibly, in people's heads
- Every pad shows its estimated effect on both, before it is submitted
- Then it goes to a named approver, one at a time or hundreds at once
Three ways to place a pad, one mental model
- A form for precision, a file for volume, a shape drawn on the map for speed
- Mid-event, people think in the shape of the storm; the system translates it back into delivery areas
- All three land on the same review step. Speed changes the input, never the safeguards.
One design system: Meridian
- A unified solution deserved a unified language: Umbrella is built on Meridian, the design system built for Amazon's transportation org
- Most of the product is out-of-the-box Meridian: tables, forms, navigation, approvals. Deliberately unexciting, so nothing needs relearning mid-storm.
- The gap was the map. Meridian had no map component, so we designed a new one and contributed it back, for every transportation team after us
The network, seen in place
The protections map view: active pads across the network, running on the map component we added to Meridian.
The report that used to eat evenings, compiled by nobody
Performance, per pad: what it protected and what it spent: the picture leaders used to assemble by hand.
attributed to the weather padding programme that Umbrella made operable
Umbrella: one entry point for weather contingency across Amazon's global network.
Umbrella's score with the same operators who rated the old tools 58 and 47. Anything above 80 is an A grade; the industry average is 68.
The programme became usable
An automated model, a separate rulebook and a manual process, replaced by one product.
Research, information architecture, every screen and the interaction specs, across two time zones.
A design function that hadn't existed at the start, with other designers hired into it and a product pipeline behind it.
Umbrella was the interface layer, not the model. What I own is that the programme became usable: a pad could be placed quickly, priced before it counted, and explained by name afterwards.