How much of a win is luck?
Contents
- A strip of paper
- What the curve looks like
- Where skill comes from
- So how much is luck?
- Does the game agree?
- What it doesn't say
- Checked by machine
- What this might mean for learning agents
- One game says less than it seems
- The most useful opponent isn't an equal one
- The small version of the game might be the worst place to start
- Share the dice across a group
- Score agents by θ, and by β
- Back to the question
- Further reading
- Papers
Suppose you are the strongest player at a table, and everyone else picks their moves at random. How often should you win?
"Every time" is wrong as soon as there are dice. "Your fair share" is wrong as soon as there are choices. The truth sits somewhere in between, and this post is about one equation that says where. It comes with a picture that makes the equation feel obvious, so we'll start with the picture.
The setting is a family of dice games where each player controls some number of pieces, and every roll lets you choose which piece to move. Two knobs matter: how many players sit at the table, and how many pieces each one controls.
A strip of paper
Start with a table of players where nobody has an edge. By symmetry, each wins with the same chance, their fair share .
Here is a way to picture that. Take a strip of paper and cut it into slices of equal width, one per player. Drop a point on the strip, uniformly at random. Whoever owns the slice it lands in, wins. Equal slices, equal chances.
Now give one player an edge. The natural move is to make their slice wider. Say your slice has width , while each of the other players keeps a slice of width 1. The chance the point lands on you is your width over the total width:
That's the whole model. Play with it:
Your slice is eθ wide, with θ = 0.751 × (4 − 1) = 2.25. Each of the 3 other players gets a slice of width 1.
- Chance you win
- 25.0%
- Fair share 1/4
- 25.0%
- Luck share L
- 100.0%
A slice's width is only a promise about the long run. To see what one game looks like, drop the points yourself. Each drop is one game.
- Games played
- 0
- You won
- –
- The formula says
- 76.0%
1 gamegames played (log scale)10
The dashed line is the formula. The shaded funnel is where the running share should stay 95% of the time: wide after a few games, narrow after many.
Ten drops can say almost anything. At the starting setting, a player who should win 76% of the time wins six or fewer of ten about one time in five, and anywhere from five to all ten is ordinary. A thousand drops crowd into the funnel around the formula. That gap between one game and many is what luck feels like from the inside.
This rule, where everyone wins in proportion to a strength, is Luce's choice rule from 1959. With two players it's the Bradley–Terry model from 1952, which is what most rating systems are built on.
Why write the width as rather than just ? Because skill compounds. Raising the skill index by one multiplies your slice by , whatever it was before. And means a slice of width 1, the same as everyone else, which is exactly "no skill".
What the curve looks like
Hold the table fixed and turn up . At two players the formula collapses to the logistic curve, . With more players the curve keeps its S-shape but starts lower and climbs later.
- 2 players
- 3 players
- 4 players
Every curve starts at the fair share 1/n when θ = 0 and bends towards a sure win. More opponents pull the whole curve down.
Show data
| Series | Skill index θ | Win chance |
|---|---|---|
| 2 players | 0.0 | 50% |
| 2 players | 0.3 | 59% |
| 2 players | 0.8 | 68% |
| 2 players | 1.1 | 75% |
| 2 players | 1.4 | 81% |
| 2 players | 1.8 | 86% |
| 2 players | 2.2 | 90% |
| 2 players | 2.5 | 93% |
| 2 players | 2.9 | 95% |
| 2 players | 3.3 | 96% |
| 2 players | 3.6 | 97% |
| 2 players | 4.0 | 98% |
| 3 players | 0.0 | 33% |
| 3 players | 0.3 | 42% |
| 3 players | 0.8 | 51% |
| 3 players | 1.1 | 60% |
| 3 players | 1.4 | 68% |
| 3 players | 1.8 | 75% |
| 3 players | 2.2 | 82% |
| 3 players | 2.5 | 86% |
| 3 players | 2.9 | 90% |
| 3 players | 3.3 | 93% |
| 3 players | 3.6 | 95% |
| 3 players | 4.0 | 96% |
| 4 players | 0.0 | 25% |
| 4 players | 0.3 | 32% |
| 4 players | 0.8 | 41% |
| 4 players | 1.1 | 50% |
| 4 players | 1.4 | 59% |
| 4 players | 1.8 | 67% |
| 4 players | 2.2 | 75% |
| 4 players | 2.5 | 81% |
| 4 players | 2.9 | 86% |
| 4 players | 3.3 | 90% |
| 4 players | 3.6 | 93% |
| 4 players | 4.0 | 95% |
The formula also runs backwards. Rearranging gives
so any measured win rate can be turned into a skill index. That turns out to be the useful direction, because it lets you ask what actually is in a real game.
Where skill comes from
Skill needs choices. If you control only one piece, every roll moves that piece and there is nothing to decide. A player who thinks hard and a player who flips coins make identical moves, so the skilled player wins exactly the fair share, , and . That's not a measurement. It's forced.
Each extra piece adds options, and so adds room for skill. The simplest guess is that each extra piece adds the same amount:
The isn't a fitting choice. It's what the one-piece case demands. In odds, this says every extra piece multiplies your odds by the same factor , at any table size.
Multiplying is the key word. Watch what happens to your slice as pieces are added, first on an ordinary ruler and then on a ruler that counts in multiples:
On an ordinary ruler
Each piece multiplies the width, so the gaps keep growing.
On a log ruler
Take logs and every step is the same size: that's θ = β(k − 1), a straight line.
Measured against a random opponent, a player following a short list of rules of thumb lands close to that straight line:
- Rules of thumb, β = 0.751
- Shallow search, β = 0.411
- Rules of thumb, measured at 2 players
Both lines start at zero, because one piece leaves nothing to decide. Their slopes differ: β is a property of the player, not of the game.
Show data
| Series | Pieces each player controls, k | Skill index θ |
|---|---|---|
| Rules of thumb, β = 0.751 | 1 | 0.0 |
| Rules of thumb, β = 0.751 | 4 | 2.3 |
| Shallow search, β = 0.411 | 1 | 0.0 |
| Shallow search, β = 0.411 | 4 | 1.2 |
| Rules of thumb, measured at 2 players | 1 | 0.0 |
| Rules of thumb, measured at 2 players | 2 | 0.7 |
| Rules of thumb, measured at 2 players | 3 | 1.5 |
| Rules of thumb, measured at 2 players | 4 | 2.2 |
Getting from measured win rates to that line takes a few lines of code. Turn each rate into with the backwards formula, then fit a slope that is forced through zero at one piece:
import math
# Measured win rate against one random player, by pieces each
measured = {2: 0.667, 3: 0.813, 4: 0.900}
def theta_from(p, n=2):
"""Skill index implied by a win rate p at a table of n."""
return math.log((n - 1) * p / (1 - p))
thetas = {k: theta_from(p) for k, p in measured.items()}
# Least squares through the origin: one piece is exactly zero skill
num = sum((k - 1) * t for k, t in thetas.items())
den = sum((k - 1) ** 2 for k in thetas)
beta = num / den
print({k: round(t, 2) for k, t in thetas.items()})
# {2: 0.69, 3: 1.47, 4: 2.2}
print(round(beta, 2)) # 0.73
Two players alone give 0.73. Pooling the measurements from two, three and four players gives the 0.751 used in the rest of this post.
Here is the twist. Swap in a different skilled player, a shallow look-ahead search, and the line keeps its shape but changes its slope: instead of . On the same dice, the simple rules beat the search. So doesn't belong to the game. It belongs to the player, and describes how far that player's edge reaches against no skill at all.
For the rules-of-thumb player, . Every piece roughly doubles their odds.
So how much is luck?
Put your win chance on a line. At the left end sits the fair share : what you'd get in a game of pure luck. At the right end sits a certain win: pure skill. Your lands somewhere between. The luck share is the fraction of that distance skill hasn't covered:
It equals 1 when skill is worth nothing and falls towards 0 as skill takes over. The bar under the strip above draws exactly this.
From the first form to the second
Start from the strip: , the share of the strip that isn't yours.
The fair-share gap is .
Divide one by the other. The cancels:
Now read the second form. Your skill appears once, as in the denominator, so every piece pushes luck down. The number of players appears on top and underneath, and each extra opponent adds one more unit-width slice of paper that your slice has to beat. More pieces, less luck. More players, more luck.
- 2 players
- 3 players
- 4 players
One piece is pure luck at any table size. Pieces pull the share down; every extra opponent holds it up.
Show data
| Series | Pieces each player controls, k | Luck share L |
|---|---|---|
| 2 players | 1 | 100% |
| 2 players | 1 | 89% |
| 2 players | 2 | 81% |
| 2 players | 2 | 71% |
| 2 players | 2 | 61% |
| 2 players | 2 | 52% |
| 2 players | 3 | 46% |
| 2 players | 3 | 39% |
| 2 players | 3 | 32% |
| 2 players | 4 | 27% |
| 2 players | 4 | 23% |
| 2 players | 4 | 19% |
| 3 players | 1 | 100% |
| 3 players | 1 | 92% |
| 3 players | 2 | 87% |
| 3 players | 2 | 78% |
| 3 players | 2 | 70% |
| 3 players | 2 | 62% |
| 3 players | 3 | 56% |
| 3 players | 3 | 49% |
| 3 players | 3 | 42% |
| 3 players | 4 | 35% |
| 3 players | 4 | 31% |
| 3 players | 4 | 26% |
| 4 players | 1 | 100% |
| 4 players | 1 | 94% |
| 4 players | 2 | 90% |
| 4 players | 2 | 83% |
| 4 players | 2 | 76% |
| 4 players | 2 | 68% |
| 4 players | 3 | 63% |
| 4 players | 3 | 56% |
| 4 players | 3 | 49% |
| 4 players | 4 | 42% |
| 4 players | 4 | 38% |
| 4 players | 4 | 32% |
The same numbers as a grid make the two directions easy to see at once:
| Players | 1 piece | 2 pieces | 3 pieces | 4 pieces |
|---|---|---|---|---|
| 2 players | 100% | 64% | 36% | 19% |
| 3 players | 100% | 73% | 46% | 26% |
| 4 players | 100% | 78% | 53% | 32% |
Predicted from the equation with β = 0.751. Read across a row and luck falls; read down a column and it rises.
At the two ends of the measured range, for the rules-of-thumb player with four pieces each: two players leave a luck share of about 19%, and four players about 32%. Against the shallow search, the same four-player game is 62% luck.
The whole model fits in two functions. Here is the four-piece column of the grid above:
import math
def win_chance(n, theta):
"""One skilled player against n - 1 who choose at random."""
return math.exp(theta) / (math.exp(theta) + n - 1)
def luck_share(n, theta):
"""1 for a game of pure luck, 0 for a game of pure skill."""
return n / (n - 1 + math.exp(theta))
theta = 0.751 * (4 - 1) # four pieces each
for n in (2, 3, 4):
p, L = win_chance(n, theta), luck_share(n, theta)
print(n, round(p, 3), round(L, 3))
# 2 0.905 0.19
# 3 0.826 0.261
# 4 0.76 0.32
Does the game agree?
At two players the measurements sit almost on the curve. The pooled , with a 95% interval of 0.735 to 0.769, predicts 81.8% at three pieces and 90.5% at four. The measured rates are 81.3% and 90.0%. At two pieces the prediction of 67.9% sits just past the edge of the measured interval.
- 2 players
- 3 players (predicted)
- 4 players (predicted)
- 2 players, measured
Lines use the pooled β = 0.751. The dots and the shaded band are measured win rates with 95% intervals at two players; the dashed curves are predictions.
Show data
| Series | Pieces each player controls, k | Win chance |
|---|---|---|
| 2 players | 1 | 50% |
| 2 players | 1 | 56% |
| 2 players | 2 | 59% |
| 2 players | 2 | 65% |
| 2 players | 2 | 70% |
| 2 players | 2 | 74% |
| 2 players | 3 | 77% |
| 2 players | 3 | 81% |
| 2 players | 3 | 84% |
| 2 players | 4 | 87% |
| 2 players | 4 | 88% |
| 2 players | 4 | 90% |
| 3 players (predicted) | 1 | 33% |
| 3 players (predicted) | 1 | 39% |
| 3 players (predicted) | 2 | 42% |
| 3 players (predicted) | 2 | 48% |
| 3 players (predicted) | 2 | 53% |
| 3 players (predicted) | 2 | 59% |
| 3 players (predicted) | 3 | 62% |
| 3 players (predicted) | 3 | 68% |
| 3 players (predicted) | 3 | 72% |
| 3 players (predicted) | 4 | 77% |
| 3 players (predicted) | 4 | 79% |
| 3 players (predicted) | 4 | 83% |
| 4 players (predicted) | 1 | 25% |
| 4 players (predicted) | 1 | 29% |
| 4 players (predicted) | 2 | 33% |
| 4 players (predicted) | 2 | 38% |
| 4 players (predicted) | 2 | 43% |
| 4 players (predicted) | 2 | 49% |
| 4 players (predicted) | 3 | 53% |
| 4 players (predicted) | 3 | 58% |
| 4 players (predicted) | 3 | 63% |
| 4 players (predicted) | 4 | 69% |
| 4 players (predicted) | 4 | 72% |
| 4 players (predicted) | 4 | 76% |
| 2 players, measured | 1 | 50% |
| 2 players, measured | 2 | 67% |
| 2 players, measured | 3 | 81% |
| 2 players, measured | 4 | 90% |
But one number for every table size is an approximation, and the measurements say so in two ways.
- The slope drifts with table size. Going from two players to four raises by about 0.06, for both players, with intervals that exclude zero.
- Luce's rule bends a little. The rule assumes your chance depends only on the strengths at the table. That gives a second way to read : seat two skilled players against the rest, and both readings should agree. Across twelve such checks for the two players, two disagreed, where chance alone would explain about 0.6. Both were the rules-of-thumb player with four pieces, so it's suggestive rather than settled.
Here is what that second reading looks like on the strip. Two players each own a wide slice, and the strip says each of them should win
- Each skilled player wins
- 45.2%
- Each other player wins
- 4.8%
If Luce's rule holds, measuring this table and solving for θ gives the same θ as one skilled player against three. That agreement is the test.
So the equation is a good description, and not a law. That's a fine thing for an equation to be, as long as it says so.
What it doesn't say
- It isn't the most skill can do. Each describes one particular player against random ones. A stronger player would have a steeper line.
- It isn't about evenly matched players. Luck between two strong players is a different question.
- It's only checked inside its range: two to four players and one to four pieces. Three points of slope make a pattern, not a law of nature.
Checked by machine
The algebra in this post, though not the fit, is proven in the Lean theorem prover. That covers the following statements:
- Luce's rule reduces to the formula for , and the shares sum to one.
- With one piece, and , exactly.
- always lies between 0 and 1 (given ), falls strictly as grows, and tends to zero.
- For any , adding a player strictly increases .
- Each extra piece multiplies the odds by exactly .
What this might mean for learning agents
A reinforcement-learning agent in a game like this usually learns from one number per game: did it win? The equation lets us ask how much that number can tell it.
One game says less than it seems
Differentiate with respect to and something tidy happens, at every table size:
A win is a coin flip with bias , and the information one flip carries about works out to that same . So the number of games it takes to notice a change in skill grows like
To see θ move by 0.1, at two standard errors, takes about 1,600 games when you win half the time. It takes about 4,400 when you already win 90% of them. It's the funnel from dropping points, read the other way: how many drops before the funnel is narrow enough to tell two slices apart.
- 2 players
- 3 players
- 4 players
Each curve peaks at 0.25, where the skilled player wins exactly half the time. For two players that's an even match; for four it's θ = ln 3 ≈ 1.1.
Show data
| Series | Skill index θ | Information per game, P(1 − P) |
|---|---|---|
| 2 players | 0.0 | 0.25 |
| 2 players | 0.3 | 0.24 |
| 2 players | 0.8 | 0.22 |
| 2 players | 1.1 | 0.19 |
| 2 players | 1.4 | 0.15 |
| 2 players | 1.8 | 0.12 |
| 2 players | 2.2 | 0.09 |
| 2 players | 2.5 | 0.07 |
| 2 players | 2.9 | 0.05 |
| 2 players | 3.3 | 0.04 |
| 2 players | 3.6 | 0.02 |
| 2 players | 4.0 | 0.02 |
| 3 players | 0.0 | 0.22 |
| 3 players | 0.3 | 0.24 |
| 3 players | 0.8 | 0.25 |
| 3 players | 1.1 | 0.24 |
| 3 players | 1.4 | 0.22 |
| 3 players | 1.8 | 0.19 |
| 3 players | 2.2 | 0.15 |
| 3 players | 2.5 | 0.12 |
| 3 players | 2.9 | 0.09 |
| 3 players | 3.3 | 0.07 |
| 3 players | 3.6 | 0.05 |
| 3 players | 4.0 | 0.03 |
| 4 players | 0.0 | 0.19 |
| 4 players | 0.3 | 0.22 |
| 4 players | 0.8 | 0.24 |
| 4 players | 1.1 | 0.25 |
| 4 players | 1.4 | 0.24 |
| 4 players | 1.8 | 0.22 |
| 4 players | 2.2 | 0.19 |
| 4 players | 2.5 | 0.15 |
| 4 players | 2.9 | 0.12 |
| 4 players | 3.3 | 0.09 |
| 4 players | 3.6 | 0.07 |
| 4 players | 4.0 | 0.05 |
If a learner's progress is limited by the same signal, outcome-only rewards in luck-heavy games could need many more episodes than the size of the game suggests. Those are the pass/fail rewards that a lot of recent RL relies on.
The most useful opponent isn't an equal one
Every curve above peaks where the learner wins exactly half its games. At two players that means an even match. At four players it means being clearly stronger than each opponent: . Self-play among four equals sits at a quarter of the wins and gets 25% less information per game than the peak. A trainer that picks opponents to keep the learner near 50% wins, rather than near its fair share, might learn faster at bigger tables.
The small version of the game might be the worst place to start
Curricula usually start small. Here, with one piece, the reward doesn't depend on the policy at all, so any gradient from it is pure noise. With two pieces the luck share is still 64% to 78%, depending on the table. Starting small could mean starting where the outcome says least about the choices.
Share the dice across a group
Group-relative methods, such as GRPO, compare several attempts from the same starting point. If the attempts in a group also shared their dice, the comparison would cancel much of the luck before it reached the gradient. In evaluation, reusing the same dice across seatings made confidence intervals about a fifth narrower. Whether training gains anything similar is untested, and it needs an environment whose randomness can be replayed.
Score agents by θ, and by β
Win rate against random players depends heavily on the table size. moves much less, because the per-piece slope drifts by only about 0.06 between two and four players. And condenses how an agent's edge grows as the game gives it more choices into a single number. A learned agent could be placed on the same chart as the rules-of-thumb player and the shallow search above, so its progress has a scale.
Back to the question
Add an opponent and the game gets luckier. Add a piece and it gets more skilful. Both answers fall out of a strip of paper.
Further reading
Six books, from the general to the technical:
- The Success Equation
Where any activity sits between pure luck and pure skill, and how to tell. The closest book to this post's question.
- Thinking in Bets
Judging a decision by how it was made rather than how it turned out: the everyday version of telling a good choice from a lucky one.
- Individual Choice Behavior
Where the strip of paper comes from: the choice rule, and the independence assumption the whole equation rests on.
- Discrete Choice Methods with Simulation
The logit model in full, including what goes wrong when independence fails. Free to read from the author.
- An Introduction to the Bootstrap
Where the confidence intervals here come from: resample the games you played instead of trusting a formula.
- Reinforcement Learning: An Introduction
The standard text on learning from reward, for following up the untested ideas above. Free to read from the authors.
Papers
On measuring skill against luck:
- Borm & van der Genugten, On a relative measure of skill for games with chance elements, TOP 9 (2001). Places a game on a 0–1 scale by comparing a beginner, an optimal player, and an imaginary player who knows every roll in advance.
- Duersch, Lambrecht & Oechssler, Measuring skill and chance in games, European Economic Review 127 (2020). The wider the spread of a game's fitted Elo ratings, the more skill decides it, measured against a deliberately half-random version of chess.
- Getty, Li, Yano, Gao & Hosoi, Luck and the Law: Quantifying Chance in Fantasy Sports and Other Contests, SIAM Review 60 (2018). A skill–luck measure built from players' records, aimed at the legal question of whether a contest is gambling.
- Jerdee & Newman, Luck, skill, and depth of competition in games and social hierarchies, Science Advances 10 (2024). A paired-comparison model with separate parameters for upsets and for depth of competition, fitted across sports, games and animal hierarchies.
On the models and methods used here:
- Bradley & Terry, Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons, Biometrika 39 (1952). The two-player case of the strip of paper, and the root of most rating systems.
- Efron, Bootstrap Methods: Another Look at the Jackknife, Annals of Statistics 7 (1979). The resampling idea behind every interval in this post.
- Burch, Schmid, Moravčík & Bowling, AIVAT: A New Variance Reduction Technique for Agent Evaluation in Imperfect Information Games (2016). Subtracting out the luck of the deal to compare agents with fewer games, a close cousin of sharing the dice.
- Shao et al., DeepSeekMath (2024). Introduces GRPO, the group-relative training method in the shared-dice idea above.