How to read your results: median, percentiles, and what "success" means
A retirement stress test hands you a few numbers: a percentage, a median, a "10th percentile" and a "90th percentile." Each means something precise, and each is easy to misread. This guide explains what Runway's Monte Carlo and historical backtest results actually say, using one couple at three spending levels so you can see how the numbers move together.
What does "chance your money lasts" actually count?
Runway simulates your plan year by year, from today to your plan-through age, in each of thousands of randomly generated market histories (how that works). In each simulated future it tries to pay your spending target, plus taxes and healthcare, every year. If a year comes up short by more than a dollar, that future is counted as a failure; otherwise it is a success. The success rate is the successes divided by the total. So 84.9% means 1,698 of 2,000 futures paid every bill through 95, and 302 did not.
That rule is strict and simple, and it leaves out two things the percentage can't tell you: when the shortfall happened and how big it was. Take the failures at three spending levels:
| Annual spending | Lasts to 95 | Futures that ran short (of 2,000) | Median age money first ran short | Share of those short before 85 |
|---|---|---|---|---|
| $95,000 | 99.4% | 13 | n/a (too few) | n/a |
| $110,000 | 84.9% | 302 | 90 | 10% |
| $125,000 | 51.4% | 972 | 87 | 30% |
At $110,000, even the futures that fail mostly fail late: half of them first ran short at 90 or later, and the earliest was 77. At $125,000 the failures arrive sooner: a typical one at 87, and 30% before 85. Two plans with similar success rates can therefore carry different risks, which is why the historical backtest labels every failing starting year with the age it ran short, not only a percentage.
Why does Runway report the median and not the average?
The median is the middle result: half of the futures ended above it and half below. The average adds up all the endings and divides, which means a handful of extremely lucky futures can drag it far from what is typical. In the $95,000 run the average ending balance is $1,660,728 and the median is $1,483,604: the average is $177,125 higher, only because of the lucky tail. The median is the better answer to "what is a typical outcome?"
What are the 10th and 90th percentiles?
A percentile says what share of outcomes fall below a level. The 10th percentile is the balance that 1 in 10 futures ended below, a bad but not a worst case. The 90th percentile is the level 9 in 10 stayed below, a good case. Together with the median they show the spread, not just the middle:
| Annual spending | 10th percentile ending | Median ending | 90th percentile ending | Average ending |
|---|---|---|---|---|
| $95,000 | $662,952 | $1,483,604 | $2,924,316 | $1,660,728 |
| $110,000 | $0 | $1,070,647 | $2,595,709 | $1,199,599 |
| $125,000 | $0 | $36,907 | $1,966,097 | $605,096 |
A 10th percentile of $0 means at least 1 in 10 futures ended with nothing left; at $110,000 it was 15.1%. Watch how the percentiles open up over time at that spending level, with Pat's age along the top (all in today's dollars):
| Pat's age | 10th percentile | Median | 90th percentile |
|---|---|---|---|
| 70 | $919,252 | $1,265,505 | $1,696,928 |
| 75 | $749,273 | $1,261,249 | $2,003,025 |
| 80 | $577,592 | $1,258,552 | $2,252,188 |
| 85 | $371,506 | $1,219,480 | $2,401,824 |
| 90 | $124,944 | $1,159,814 | $2,578,115 |
| 95 | $0 | $1,070,647 | $2,595,709 |
The typical future stays near $1.1 to $1.3 million the whole way, while the bad case drains to zero and the good case climbs to about $2.6 million. That widening fan is what "a range of outcomes instead of one average" means in practice. The single line most retirement tools show is just the middle of it.
Is a higher success rate always better?
Not automatically. At $95,000 the success rate is 99.4%, and in 51.2% of those futures the couple ends with more than the $1,460,000 they started with, in today's dollars. A number that high can mean a plan that leaves a lot unspent. Runway's Your Best Plan page uses a 90% chance as its yardstick: the most you could spend each year with a 90% chance your money lasts. That is a choice, not a law of nature. The right target depends on how much you value a cushion against how much you want to spend, and on whether you could cut spending if things went badly.
How is the historical backtest different?
The backtest doesn't draw random markets. It replays the plan from every real starting year in the data, 125 of them (1872 through 1996, each running 30 years). For the same couple at $95,000, all 125 starting years lasted (100.0%), with a median ending balance of $1,379,861 and 10th and 90th percentiles of $724,222 and $2,728,179. At $110,000, 76.0% of starting years lasted, against 84.9% in the Monte Carlo. The two methods can disagree, and the asset allocation guide shows why. One caution about the backtest percentages: the 125 starting years overlap heavily, so they are far fewer independent experiences than the number suggests.
What should I do with all this?
- Read the percentage as a stress score. It says how a plan holds up, not what will happen. Don't read 99.7% as safer than 99.4%.
- Look at the bad case. Ask what you'd do if you ended up at the 10th percentile, such as cutting spending, working a bit, or using guardrails.
- Check the age of the first shortfall. A plan that runs short at 91 is a different problem from one that runs short at 80.
- Remember it all rests on assumptions. Change the return and volatility numbers and every figure here moves.
See these numbers for your own plan. The Stress Test and Historical Backtest tabs under What-If Scenarios show the success rate and the 10th, median and 90th percentile results for your own accounts.
Try the plannerHow were these numbers computed?
All figures are Runway's engine on the sample couple: married filing jointly, Social Security claimed at 69 and 68, plan through 95, 60% stocks and 40% bonds with stocks at 7.5% and bonds at 2.93% a year after inflation and volatilities of 16% and 6%, long-term-care costs left out. Monte Carlo results use 2,000 simulated markets with seed 2026. A future counts as a success if the spending need is met in every year (a shortfall of more than $1 is a failure). The historical backtest uses annual real returns for U.S. stocks and 10-year bonds from Robert Shiller's public dataset. All amounts are in today's dollars.