← All projects

Umbrella: The Tool That Saved Amazon $250M

When a storm is coming, Amazon extends delivery promises before the network slips, using what it calls a promise pad. Umbrella is the platform that decides where those pads land, prices what each one costs, and has to explain a model's decision to the person who gets asked about it.

Company Amazon
Role Solo Design Lead
Timeline 2022 - 2024
Scope North America & Europe
Impact $250M+ saved
The basics

What a promise pad actually costs

Amazon measures Delivery Estimate Accuracy as the percentage of packages that arrive when the customer was told they would. Accuracy is not speed. If I tell you Thursday and it arrives Friday, that is a miss. If I tell you on Wednesday that it is now Friday, that is accurate. The promise can move. It just has to move before it expires.

When a storm is coming, the way to protect accuracy is to add time to delivery promises before the network slips. That is a promise pad, and it is a purchase: you spend speed to buy accuracy, and someone's targets pay for it.

Diagram of the promise pad lifecycle, showing how a delivery promise is extended before it expires so it still holds when weather delays the shipment
Context

The same storm was getting three different answers

Across North America and Europe, weather padding had grown up locally. Each operation had its own tools, its own thresholds and its own definition of a bad day. None of them could see the others, and nobody could say what was happening across the network without assembling it by hand.

North America: the forecast feeds a model that places many pads automatically
North America

A tool called Tramontane used a forecasting model to place pads automatically. Fast, and almost impossible for the people affected to argue with.

Europe: the forecast is read against a separate rulebook, placing fewer pads with less automation
Europe

A separate system with its own rulebook and its own severity thresholds, built independently, with far less automation.

Canada: the forecast is worked out by a person, who places each pad by hand
Canada

Within North America, Canada had no system at all. Every pad was worked out and placed by hand, event after event.

Design opportunity

How might we automate a decision this consequential, without taking away the control people need to defend it?

Research

Three groups wanted opposite things from the same product

I ran research across the three groups who would use Umbrella, starting with a week-long workshop in Seattle with operations teams from both regions, mapping how each one actually handled an event.

Four findings came back. Every decision below traces to one of them.

Photographs and artefacts from the Seattle workshop, with walls of notes mapping the vision, personas and feature priorities across regions

A week in Seattle mapping how each region actually handled an event.

The three Umbrella personas: configurators who need depth, senior leaders who need breadth, and manual overriders who need speed

Depth, breadth and speed, competing for the same screen.

The work

Four decisions, and what each one cost

Each of these came out of one of the findings above. For every one I have tried to name the option I did not take, and what choosing the other one cost us, because that second part is usually the bit that goes missing.

01 Decision

One product layered by depth, not three tailored views

Giving each group its own view tested beautifully with each group alone. It fell apart in the situation that actually matters: a leader asking an operator why a pad is there, with two different products in front of them and nothing in common to point at.

So everyone gets the same structure and moves down a level when they need more. The cockpit answers the leader's question, protections sits under it for the people doing the work, configuration under that for the people writing rules.

I chose one shared product over three tailored ones. It cost the leaders simplicity: their view carries a route into depth they will never personally use, and I would not hide it, because the moment they ask why, they need to be able to walk down there with someone.
The Umbrella V1 information architecture, mapping every page and flow across the three persona types

One structure at three depths, rather than three products.

The Umbrella cockpit dashboard, showing weather event status and active protections across regions at a glance

The cockpit: the leader's question, answered without anyone compiling it.

02 Decision

The override belongs next to the result, not in a settings screen

The second finding was the sharpest: automation people cannot overrule is worse than no automation. The obvious build is to let the system place pads and notify you, with the rules tucked away somewhere in settings. That is the calmest design on an ordinary day and the worst one during a storm.

So pads the system placed and pads people placed live in one list, with the same controls on both.

The protections table, listing pads placed automatically by the system alongside pads placed by named people, with accuracy and speed impact on each row
01

Automatic and manual, one list

A single column tells you whether the system or a person placed a pad. Everything else about the row is identical, so nobody has to learn two mental models.

02

Both sides of the trade, on every row

Accuracy and speed impact sit side by side. A pad that reports what it protects but not what it spends hides half the decision.

03

Why, without leaving the row

The explanation is an action on the pad itself rather than a separate audit screen, so it is reachable at the moment somebody asks.

I chose to keep the override next to the result. It cost safety: there is a Remove rule button inside a panel people open in the middle of a storm, and I have never been comfortable with how easy I made that.
03 Decision

Show the real reasoning, and translate it in place

People were being asked to defend a decision they did not make and could not see. Underneath that sits a real mismatch: operators are accountable for postcodes and delivery areas, while the forecasting model works in its own weather regions, drawn for meteorology, which cut straight across the areas people manage.

A friendly summary would have been kinder to read and would have quietly gone out of date the first time the model was retrained.

Two maps that do not line up What the operator manages, postcodes and delivery areas, the station that serves them, and the targets they are judged on. What the model watches, the weather forecast, its own weather regions, and a probability crossing a line. The two meet only at a single delivery area, where the pad lands and someone has to explain it. WHAT THE OPERATOR MANAGES WHAT THE MODEL WATCHES Postcodes and delivery areas The station that serves them The targets they answer for The weather forecast Its own weather regions A probability crossing a line One delivery area where the pad lands, and someone has to explain it Two maps that do not line up What the operator manages and what the model watches meet only at a single delivery area, where the pad lands. WHAT THE OPERATOR MANAGES Postcodes and delivery areas The station that serves them The targets they answer for WHAT THE MODEL WATCHES The weather forecast Its own weather regions A probability crossing a line One delivery area where the pad lands, and someone explains it

The two maps meet at one point. People answer for the left one and can only watch the right one.

The model deepdive panel, showing the forecasting model's own figures with a plain language glossary of each term beneath, alongside a snowfall chart and a regional map
01

The model's figures, unedited

Shown as the model produces them, so the panel cannot drift out of step with the thing it is describing.

02

A glossary directly underneath

Every term explained in plain language in the same view. The translation happens where the reader is, not in a document they will never open.

03

The forecast beside the numbers

The chart and map are there so somebody who does not want to read a table can still see the weather that caused this.

The same panel showing every rule that contributed to a pad, listed in order, each with its own parameters, author and edit controls

Where several rules combine into one pad, the same panel lists all of them in order, and each stays editable from here.

I chose honesty over friendliness. It cost first-time readability: the panel needs its glossary to be understood, which makes it hardest on exactly the person opening it for the first time, mid-event, with someone waiting.
04 Decision

Price the pad before it counts, not after

The fourth finding was that the trade between accuracy and speed was happening constantly, in people's heads, with no feedback. Reporting it afterwards in a performance view keeps it exactly as invisible as the research found it.

So every pad, placed one at a time or hundreds at once, shows its estimated effect on both before it is submitted, and then goes to a named approver.

The bulk upload review step, showing each pad with its location, type, timeframe and simulated impact on accuracy and speed before submission for approval
01

A review step that is hard to skip

Bulk uploads expand into individual pads before submission, so hundreds of decisions cannot pass as one.

02

The estimated cost, per pad

Both sides of the trade shown at the moment somebody can still change their mind about it.

03

Submit, rather than apply

The button says what actually happens next. A named person approves it, which is the part that cost us speed.

I chose to put a review step and a human approval into the fastest task in the product. It cost speed at exactly the moment speed matters, and it cost me my own tenet that Umbrella should be low touch. I would still do it, because the alternative is one person spending another region's speed with nobody agreeing to it.
Interaction design

Three ways to place a pad, one mental model

A pad is one object, but the urgency around it varies enormously. Setting up a rule before a season is considered work. Loading a file someone prepared is administrative. Placing forty pads because a storm is tracking towards the Midlands is neither.

Three entry points, one product: a form for precision, a file for volume, a shape drawn onto the map for speed. All three land on the same review step and enter the same queue.

The draw pad location modal, where an operator draws a shape directly onto a map of the delivery network, with toggles to show delivery stations, sort centres and buildings inside it
01

Draw the weather, not the postcodes

Mid-event, people think in the shape of the storm. The system resolves that shape back into the areas it understands, which moves the translation off the person.

02

See what is inside before confirming

Toggles for delivery stations and sort centres, so the consequence of a shape is visible while it is still a shape.

03

Confirm, then join the same flow

The drawn area lands on the same review and approval steps as every other route in. Speed changes the input, never the safeguards.

Craft

The structure was settled in low fidelity

The wireframes carried the arguments about columns, density and what belongs on a row, and the shipped screens are recognisably the same tables. Two things changed, and both came out of the research: accuracy and speed became two columns instead of one, and the explanation became an action on the row rather than a separate screen.

Annotated low fidelity wireframe of the pads table, showing a single impact column

The wireframe: one impact column.

The shipped protections table, showing separate accuracy and speed impact columns and a why was this pad placed action on the row

The shipped table: both sides of the trade, explanation one click away.

Impact

$250M+
attributed to the weather padding programme that Umbrella made operable
3 → 1
an automated model, a separate rulebook and a manual process, replaced by one product
1 designer
research, information architecture, every screen and the interaction specs, across two time zones

Umbrella was the interface layer, not the model. What I own is that the programme became usable: a pad could be placed quickly, priced before it counted, and explained by name afterwards.

← All projects Next: Liverpool FC →