How much of a win is luck?

Contents
  1. A strip of paper
  2. What the curve looks like
  3. Where skill comes from
  4. So how much is luck?
  5. Does the game agree?
  6. What it doesn't say
  7. Checked by machine
  8. What this might mean for learning agents
  9. One game says less than it seems
  10. The most useful opponent isn't an equal one
  11. The small version of the game might be the worst place to start
  12. Share the dice across a group
  13. Score agents by θ, and by β
  14. Back to the question
  15. Further reading
  16. Papers

Suppose you are the strongest player at a table, and everyone else picks their moves at random. How often should you win?

"Every time" is wrong as soon as there are dice. "Your fair share" is wrong as soon as there are choices. The truth sits somewhere in between, and this post is about one equation that says where. It comes with a picture that makes the equation feel obvious, so we'll start with the picture.

The setting is a family of dice games where each player controls some number of pieces, and every roll lets you choose which piece to move. Two knobs matter: how many players sit at the table, and how many pieces each one controls.

A strip of paper

Start with a table of nn players where nobody has an edge. By symmetry, each wins with the same chance, their fair share 1/n\cY{1/n}.

Here is a way to picture that. Take a strip of paper and cut it into nn slices of equal width, one per player. Drop a point on the strip, uniformly at random. Whoever owns the slice it lands in, wins. Equal slices, equal chances.

Now give one player an edge. The natural move is to make their slice wider. Say your slice has width eθ\cB{e^{\theta}}, while each of the n−1n-1 other players keeps a slice of width 1. The chance the point lands on you is your width over the total width:

P  =  eθeθ+(n−1)P \;=\; \frac{\cB{e^{\theta}}}{\cB{e^{\theta}} + \cR{(n-1)}}

That's the whole model. Play with it:

One strip, one random point

Your slice is eθ wide, with θ = 0.751 × (4 − 1) = 2.25. Each of the 3 other players gets a slice of width 1.

Chance you win
25.0%
Fair share 1/4
25.0%
Luck share L
100.0%

A slice's width is only a promise about the long run. To see what one game looks like, drop the points yourself. Each drop is one game.

Drop points on the strip
Games played
0
You won
–
The formula says
76.0%
100%0%

1 gamegames played (log scale)10

The dashed line is the formula. The shaded funnel is where the running share should stay 95% of the time: wide after a few games, narrow after many.

Ten drops can say almost anything. At the starting setting, a player who should win 76% of the time wins six or fewer of ten about one time in five, and anywhere from five to all ten is ordinary. A thousand drops crowd into the funnel around the formula. That gap between one game and many is what luck feels like from the inside.

This rule, where everyone wins in proportion to a strength, is Luce's choice rule from 1959. With two players it's the Bradley–Terry model from 1952, which is what most rating systems are built on.

Why write the width as eθ\cB{e^{\theta}} rather than just ww? Because skill compounds. Raising the skill index θ\cB{\theta} by one multiplies your slice by ee, whatever it was before. And θ=0\cB{\theta} = 0 means a slice of width 1, the same as everyone else, which is exactly "no skill".

What the curve looks like

Hold the table fixed and turn up θ\cB{\theta}. At two players the formula collapses to the logistic curve, P=1/(1+e−θ)P = 1/(1+e^{-\cB{\theta}}). With more players the curve keeps its S-shape but starts lower and climbs later.

Chance the skilled player wins, by skill index θ
  • 2 players
  • 3 players
  • 4 players
0%25%50%75%100%0.01.02.03.04.0Skill index θWin chance2 players3 players4 players

Every curve starts at the fair share 1/n when θ = 0 and bends towards a sure win. More opponents pull the whole curve down.

Show data
SeriesSkill index θWin chance
2 players0.050%
2 players0.359%
2 players0.868%
2 players1.175%
2 players1.481%
2 players1.886%
2 players2.290%
2 players2.593%
2 players2.995%
2 players3.396%
2 players3.697%
2 players4.098%
3 players0.033%
3 players0.342%
3 players0.851%
3 players1.160%
3 players1.468%
3 players1.875%
3 players2.282%
3 players2.586%
3 players2.990%
3 players3.393%
3 players3.695%
3 players4.096%
4 players0.025%
4 players0.332%
4 players0.841%
4 players1.150%
4 players1.459%
4 players1.867%
4 players2.275%
4 players2.581%
4 players2.986%
4 players3.390%
4 players3.693%
4 players4.095%

The formula also runs backwards. Rearranging gives

θ  =  ln⁡(n−1) P1−P,\cB{\theta} \;=\; \ln \frac{\cR{(n-1)}\,P}{1-P},

so any measured win rate can be turned into a skill index. That turns out to be the useful direction, because it lets you ask what θ\cB{\theta} actually is in a real game.

Where skill comes from

Skill needs choices. If you control only one piece, every roll moves that piece and there is nothing to decide. A player who thinks hard and a player who flips coins make identical moves, so the skilled player wins exactly the fair share, 1/n\cY{1/n}, and θ=0\cB{\theta} = 0. That's not a measurement. It's forced.

Each extra piece adds options, and so adds room for skill. The simplest guess is that each extra piece adds the same amount:

θ  =  β (k−1)\cB{\theta} \;=\; \cB{\beta}\,(\cG{k} - 1)

The k−1\cG{k}-1 isn't a fitting choice. It's what the one-piece case demands. In odds, this says every extra piece multiplies your odds by the same factor eβe^{\cB{\beta}}, at any table size.

Multiplying is the key word. Watch what happens to your slice as pieces are added, first on an ordinary ruler and then on a ruler that counts in multiples:

Your slice's width, eθ, as pieces are added

On an ordinary ruler

k = 11.0
k = 22.1
k = 34.5
k = 49.5
×2.1×2.1×2.1010

Each piece multiplies the width, so the gaps keep growing.

On a log ruler

k = 11.0
k = 22.1
k = 34.5
k = 49.5
+0.751+0.751+0.751110

Take logs and every step is the same size: that's θ = β(k − 1), a straight line.

Measured against a random opponent, a player following a short list of rules of thumb lands close to that straight line:

Skill index θ against pieces, for two different players
  • Rules of thumb, β = 0.751
  • Shallow search, β = 0.411
  • Rules of thumb, measured at 2 players
0.00.51.01.52.02.51234Pieces each player controls, kSkill index θRules of thumbShallow search

Both lines start at zero, because one piece leaves nothing to decide. Their slopes differ: β is a property of the player, not of the game.

Show data
SeriesPieces each player controls, kSkill index θ
Rules of thumb, β = 0.75110.0
Rules of thumb, β = 0.75142.3
Shallow search, β = 0.41110.0
Shallow search, β = 0.41141.2
Rules of thumb, measured at 2 players10.0
Rules of thumb, measured at 2 players20.7
Rules of thumb, measured at 2 players31.5
Rules of thumb, measured at 2 players42.2

Getting from measured win rates to that line takes a few lines of code. Turn each rate into θ\cB{\theta} with the backwards formula, then fit a slope that is forced through zero at one piece:

Python
import math

# Measured win rate against one random player, by pieces each
measured = {2: 0.667, 3: 0.813, 4: 0.900}

def theta_from(p, n=2):
    """Skill index implied by a win rate p at a table of n."""
    return math.log((n - 1) * p / (1 - p))

thetas = {k: theta_from(p) for k, p in measured.items()}

# Least squares through the origin: one piece is exactly zero skill
num = sum((k - 1) * t for k, t in thetas.items())
den = sum((k - 1) ** 2 for k in thetas)
beta = num / den

print({k: round(t, 2) for k, t in thetas.items()})
# {2: 0.69, 3: 1.47, 4: 2.2}
print(round(beta, 2))  # 0.73

Two players alone give 0.73. Pooling the measurements from two, three and four players gives the 0.751 used in the rest of this post.

Here is the twist. Swap in a different skilled player, a shallow look-ahead search, and the line keeps its shape but changes its slope: β≈0.41\cB{\beta} \approx 0.41 instead of 0.750.75. On the same dice, the simple rules beat the search. So β\cB{\beta} doesn't belong to the game. It belongs to the player, and describes how far that player's edge reaches against no skill at all.

For the rules-of-thumb player, e0.751≈2.1e^{0.751} \approx 2.1. Every piece roughly doubles their odds.

So how much is luck?

Put your win chance on a line. At the left end sits the fair share 1/n\cY{1/n}: what you'd get in a game of pure luck. At the right end sits a certain win: pure skill. Your PP lands somewhere between. The luck share is the fraction of that distance skill hasn't covered:

L  =  1−P1−1/n  =  nn−1+eθ\cR{L} \;=\; \frac{1-P}{1-\cY{1/n}} \;=\; \frac{\cR{n}}{\cR{n} - 1 + \cB{e^{\theta}}}

It equals 1 when skill is worth nothing and falls towards 0 as skill takes over. The bar under the strip above draws exactly this.

From the first form to the second

Start from the strip: 1−P=n−1eθ+n−11 - P = \dfrac{n-1}{e^{\theta} + n - 1}, the share of the strip that isn't yours.

The fair-share gap is 1−1n=n−1n1 - \dfrac{1}{n} = \dfrac{n-1}{n}.

Divide one by the other. The n−1n-1 cancels:

L=n−1eθ+n−1⋅nn−1=nn−1+eθ.L = \frac{n-1}{e^{\theta}+n-1}\cdot\frac{n}{n-1} = \frac{n}{n-1+e^{\theta}}.

Now read the second form. Your skill appears once, as eθ\cB{e^{\theta}} in the denominator, so every piece pushes luck down. The number of players appears on top and underneath, and each extra opponent adds one more unit-width slice of paper that your slice has to beat. More pieces, less luck. More players, more luck.

Luck share L as pieces are added (rules-of-thumb player)
  • 2 players
  • 3 players
  • 4 players
0%25%50%75%100%1234Pieces each player controls, kLuck share L4 players3 players2 players

One piece is pure luck at any table size. Pieces pull the share down; every extra opponent holds it up.

Show data
SeriesPieces each player controls, kLuck share L
2 players1100%
2 players189%
2 players281%
2 players271%
2 players261%
2 players252%
2 players346%
2 players339%
2 players332%
2 players427%
2 players423%
2 players419%
3 players1100%
3 players192%
3 players287%
3 players278%
3 players270%
3 players262%
3 players356%
3 players349%
3 players342%
3 players435%
3 players431%
3 players426%
4 players1100%
4 players194%
4 players290%
4 players283%
4 players276%
4 players268%
4 players363%
4 players356%
4 players349%
4 players442%
4 players438%
4 players432%

The same numbers as a grid make the two directions easy to see at once:

Luck share L, by players and pieces (rules-of-thumb player)
Players1 piece2 pieces3 pieces4 pieces
2 players100%64%36%19%
3 players100%73%46%26%
4 players100%78%53%32%

Predicted from the equation with β = 0.751. Read across a row and luck falls; read down a column and it rises.

At the two ends of the measured range, for the rules-of-thumb player with four pieces each: two players leave a luck share of about 19%, and four players about 32%. Against the shallow search, the same four-player game is 62% luck.

The whole model fits in two functions. Here is the four-piece column of the grid above:

Python
import math

def win_chance(n, theta):
    """One skilled player against n - 1 who choose at random."""
    return math.exp(theta) / (math.exp(theta) + n - 1)

def luck_share(n, theta):
    """1 for a game of pure luck, 0 for a game of pure skill."""
    return n / (n - 1 + math.exp(theta))

theta = 0.751 * (4 - 1)   # four pieces each

for n in (2, 3, 4):
    p, L = win_chance(n, theta), luck_share(n, theta)
    print(n, round(p, 3), round(L, 3))
# 2 0.905 0.19
# 3 0.826 0.261
# 4 0.76 0.32

Does the game agree?

At two players the measurements sit almost on the curve. The pooled β=0.751\cB{\beta} = 0.751, with a 95% interval of 0.735 to 0.769, predicts 81.8% at three pieces and 90.5% at four. The measured rates are 81.3% and 90.0%. At two pieces the prediction of 67.9% sits just past the edge of the measured interval.

Win chance of the rules-of-thumb player against random opponents
  • 2 players
  • 3 players (predicted)
  • 4 players (predicted)
  • 2 players, measured
25%50%75%100%1234Pieces each player controls, kWin chance2 players3 players4 players

Lines use the pooled β = 0.751. The dots and the shaded band are measured win rates with 95% intervals at two players; the dashed curves are predictions.

Show data
SeriesPieces each player controls, kWin chance
2 players150%
2 players156%
2 players259%
2 players265%
2 players270%
2 players274%
2 players377%
2 players381%
2 players384%
2 players487%
2 players488%
2 players490%
3 players (predicted)133%
3 players (predicted)139%
3 players (predicted)242%
3 players (predicted)248%
3 players (predicted)253%
3 players (predicted)259%
3 players (predicted)362%
3 players (predicted)368%
3 players (predicted)372%
3 players (predicted)477%
3 players (predicted)479%
3 players (predicted)483%
4 players (predicted)125%
4 players (predicted)129%
4 players (predicted)233%
4 players (predicted)238%
4 players (predicted)243%
4 players (predicted)249%
4 players (predicted)353%
4 players (predicted)358%
4 players (predicted)363%
4 players (predicted)469%
4 players (predicted)472%
4 players (predicted)476%
2 players, measured150%
2 players, measured267%
2 players, measured381%
2 players, measured490%

But one number for every table size is an approximation, and the measurements say so in two ways.

  • The slope drifts with table size. Going from two players to four raises β\cB{\beta} by about 0.06, for both players, with intervals that exclude zero.
  • Luce's rule bends a little. The rule assumes your chance depends only on the strengths at the table. That gives a second way to read θ\cB{\theta}: seat two skilled players against the rest, and both readings should agree. Across twelve such checks for the two players, two disagreed, where chance alone would explain about 0.6. Both were the rules-of-thumb player with four pieces, so it's suggestive rather than settled.

Here is what that second reading looks like on the strip. Two players each own a wide slice, and the strip says each of them should win

P  =  eθ2 eθ+(n−2).P \;=\; \frac{\cB{e^{\theta}}}{2\,\cB{e^{\theta}} + \cR{(n-2)}}.
Two skilled players at a table of four, four pieces each
Each skilled player wins
45.2%
Each other player wins
4.8%

If Luce's rule holds, measuring this table and solving for θ gives the same θ as one skilled player against three. That agreement is the test.

So the equation is a good description, and not a law. That's a fine thing for an equation to be, as long as it says so.

What it doesn't say

  • It isn't the most skill can do. Each β\cB{\beta} describes one particular player against random ones. A stronger player would have a steeper line.
  • It isn't about evenly matched players. Luck between two strong players is a different question.
  • It's only checked inside its range: two to four players and one to four pieces. Three points of slope make a pattern, not a law of nature.

Checked by machine

The algebra in this post, though not the fit, is proven in the Lean theorem prover. That covers the following statements:

  • Luce's rule reduces to the formula for PP, and the shares sum to one.
  • With one piece, P=1/nP = 1/n and L=1L = 1, exactly.
  • LL always lies between 0 and 1 (given θ≥0\theta \ge 0), falls strictly as θ\theta grows, and tends to zero.
  • For any θ>0\theta > 0, adding a player strictly increases LL.
  • Each extra piece multiplies the odds by exactly eβe^{\beta}.

What this might mean for learning agents

A reinforcement-learning agent in a game like this usually learns from one number per game: did it win? The equation lets us ask how much that number can tell it.

One game says less than it seems

Differentiate PP with respect to θ\cB{\theta} and something tidy happens, at every table size:

dPdθ=P (1−P)\frac{dP}{d\cB{\theta}} = P\,(1-P)

A win is a coin flip with bias PP, and the information one flip carries about θ\cB{\theta} works out to that same P(1−P)P(1-P). So the number of games it takes to notice a change Δθ\Delta\cB{\theta} in skill grows like

games  ∝  1P (1−P) (Δθ)2\text{games} \;\propto\; \frac{1}{P\,(1-P)\,(\Delta\cB{\theta})^{2}}

To see θ move by 0.1, at two standard errors, takes about 1,600 games when you win half the time. It takes about 4,400 when you already win 90% of them. It's the funnel from dropping points, read the other way: how many drops before the funnel is narrow enough to tell two slices apart.

How much one game tells you about θ
  • 2 players
  • 3 players
  • 4 players
0.000.100.200.300.01.02.03.04.0Skill index θInformation per game, P(1 − P)peak, θ = ln 2peak, θ = ln 34 players3 players2 players

Each curve peaks at 0.25, where the skilled player wins exactly half the time. For two players that's an even match; for four it's θ = ln 3 ≈ 1.1.

Show data
SeriesSkill index θInformation per game, P(1 − P)
2 players0.00.25
2 players0.30.24
2 players0.80.22
2 players1.10.19
2 players1.40.15
2 players1.80.12
2 players2.20.09
2 players2.50.07
2 players2.90.05
2 players3.30.04
2 players3.60.02
2 players4.00.02
3 players0.00.22
3 players0.30.24
3 players0.80.25
3 players1.10.24
3 players1.40.22
3 players1.80.19
3 players2.20.15
3 players2.50.12
3 players2.90.09
3 players3.30.07
3 players3.60.05
3 players4.00.03
4 players0.00.19
4 players0.30.22
4 players0.80.24
4 players1.10.25
4 players1.40.24
4 players1.80.22
4 players2.20.19
4 players2.50.15
4 players2.90.12
4 players3.30.09
4 players3.60.07
4 players4.00.05

If a learner's progress is limited by the same signal, outcome-only rewards in luck-heavy games could need many more episodes than the size of the game suggests. Those are the pass/fail rewards that a lot of recent RL relies on.

The most useful opponent isn't an equal one

Every curve above peaks where the learner wins exactly half its games. At two players that means an even match. At four players it means being clearly stronger than each opponent: θ=ln⁡3≈1.1\cB{\theta} = \ln 3 \approx 1.1. Self-play among four equals sits at a quarter of the wins and gets 25% less information per game than the peak. A trainer that picks opponents to keep the learner near 50% wins, rather than near its fair share, might learn faster at bigger tables.

The small version of the game might be the worst place to start

Curricula usually start small. Here, with one piece, the reward doesn't depend on the policy at all, so any gradient from it is pure noise. With two pieces the luck share is still 64% to 78%, depending on the table. Starting small could mean starting where the outcome says least about the choices.

Share the dice across a group

Group-relative methods, such as GRPO, compare several attempts from the same starting point. If the attempts in a group also shared their dice, the comparison would cancel much of the luck before it reached the gradient. In evaluation, reusing the same dice across seatings made confidence intervals about a fifth narrower. Whether training gains anything similar is untested, and it needs an environment whose randomness can be replayed.

Score agents by θ, and by β

Win rate against random players depends heavily on the table size. θ\cB{\theta} moves much less, because the per-piece slope β\cB{\beta} drifts by only about 0.06 between two and four players. And β\cB{\beta} condenses how an agent's edge grows as the game gives it more choices into a single number. A learned agent could be placed on the same chart as the rules-of-thumb player and the shallow search above, so its progress has a scale.

Back to the question

Add an opponent and the game gets luckier. Add a piece and it gets more skilful. Both answers fall out of a strip of paper.

Further reading

Six books, from the general to the technical:

  • The Success Equation

    Michael J. Mauboussin. Harvard Business Review Press, 2012.

    Where any activity sits between pure luck and pure skill, and how to tell. The closest book to this post's question.

  • Thinking in Bets

    Annie Duke. Portfolio, 2018.

    Judging a decision by how it was made rather than how it turned out: the everyday version of telling a good choice from a lucky one.

  • Individual Choice Behavior

    R. Duncan Luce. Wiley, 1959; Dover reprint, 2005.

    Where the strip of paper comes from: the choice rule, and the independence assumption the whole equation rests on.

  • Discrete Choice Methods with Simulation

    Kenneth E. Train. Cambridge University Press, 2nd ed., 2009. Free online.

    The logit model in full, including what goes wrong when independence fails. Free to read from the author.

  • An Introduction to the Bootstrap

    Bradley Efron & Robert J. Tibshirani. Chapman & Hall, 1993.

    Where the confidence intervals here come from: resample the games you played instead of trusting a formula.

  • Reinforcement Learning: An Introduction

    Richard S. Sutton & Andrew G. Barto. MIT Press, 2nd ed., 2018. Free online.

    The standard text on learning from reward, for following up the untested ideas above. Free to read from the authors.

Papers

On measuring skill against luck:

On the models and methods used here: