Choices No Utility Function Represents

What you will be able to do

The learner can compute the expected value of a prospect, derive the inequality a stated preference imposes on a utility function, demonstrate that a common-consequence pair of preferences admits no solution, and distinguish a violation of the theory's axioms from mere risk aversion or from a taste the theory permits.

Orientation

What would refute a representation theorem

Expected utility theory is often introduced as a claim about how people decide, and then dismissed on the grounds that nobody computes utilities. Both halves of that exchange miss what the theory says.

It is a representation theorem. If a person's preferences satisfy a handful of axioms, then a function u exists such that their choices coincide with maximising ∑ i p i u ( x i ) . The function is constructed from the preferences rather than assumed in advance, so it accommodates any attitude to risk. Refusing a gamble worth 1,390,000 in expectation in favour of a certain 1,000,000 violates nothing: it says the person's u is concave, which is what risk aversion means here.

So a single choice can never contradict the theory. What would is a pattern of choices from which no such function can be built.

That is what this unit exhibits. Two decision problems, differing only in a component common to both options in each, produce a modal preference pair whose two halves demand u ( 1 M ) > 10 / 11 and u ( 1 M ) < 10 / 11 . The same threshold, opposite directions, no solution. The demonstration is algebra, not a survey result, and it holds for every utility function rather than for most of them.

The unit then covers what a descriptive model adds to accommodate the pattern, and, separately, because the two are often run together, what it does not thereby establish about whether the pattern is a mistake.

Definition

The representation theorem and its axioms

The canonical statements above give the representation, the independence axiom and the two prospect-theory components. What follows is the logical status of each, which decides what could refute it.

The direction of the theorem matters. It does not say people maximise utility. It says: if the axioms hold, then a utility function exists reproducing the choices. So evidence about mental process, that nobody multiplies probabilities by utilities, bears on nothing. The theory is falsified only by preferences no function can reproduce.

The utility function is determined only up to a positive affine transformation. Replacing u by a u + b with a > 0 leaves every comparison unchanged, which is why two outcomes can be fixed arbitrarily. Setting u ( 0 ) = 0 and u ( 5 M ) = 1 costs nothing and leaves exactly one unknown.

AxiomRequiresRuled out
completenessevery pair is comparablerefusing to rank
transitivity P ≻ Q ≻ R ⇒ P ≻ R cyclic preference
continuitysmall probability changes move preference smoothlylexicographic safety-first rules
independencea common component cancelsthe Allais pattern

The Allais pair violates independence and none of the others. The preferences are complete, transitive and continuous. This matters because "irrational" usually connotes a cycle or an inability to choose, and neither occurs here.

Why the common component cancels under independence. Problem 1's options both contain a 0.89 chance of 1 M ; problem 2's both contain a 0.89 chance of nothing. Independence says the shared part contributes identically to each side and therefore cannot determine the ranking. Removing it leaves the identical residual comparison in both problems, 1 M with probability 0.11 against 5 M with probability 0.10 , so the two problems must receive the same answer.

Prospect theory breaks the cancellation deliberately. If probabilities enter through a weighting function w , the common component contributes w ( 0.89 ) u ( 1 M ) in one problem and w ( 0.89 ) u ( 0 ) in the other, and these are not the same quantity. The model therefore permits the two problems to be answered differently without contradiction.

Weighting is not belief. A decision-maker may know a probability is 0.01 exactly and still act on a weight of 0.0673 . The distortion is in how a known probability enters the decision, not in what the person thinks the probability is, which is why offering better information does not remove it.

Intuition

Why the contradiction is exact rather than statistical

Most empirical claims about behaviour are statistical: a tendency, a proportion, an effect size with an interval. The Allais result is not one of them. It is a statement about algebra, and it would hold if exactly one person made the two choices.

The mechanism. Each preference is an inequality between two weighted sums. Write u ( 0 ) = 0 , u ( 5 M ) = 1 , and let u stand for u ( 1 M ) .

Preferring the certain million:

u > 0.10 ( 1 ) + 0.89 u + 0.01 ( 0 ) ⟹ 0.11 u > 0.10 ⟹ u > 10 11 .

Preferring the larger gamble in the second problem:

0.10 ( 1 ) > 0.11 u ⟹ u < 10 11 .

The two thresholds are the same number, 10 / 11 = 0.909091 , approached from opposite sides. There is no parameter value, however extreme, that satisfies both, so the pair is not merely unusual, it is unrepresentable.

What survives the demonstration. Risk aversion survives entirely: taking 1,000,000 certain over a prospect worth 1,390,000 requires only u > 10 / 11 , which any sufficiently concave function provides. The theory is comfortable with that. What it cannot absorb is holding that preference and the second one.

Why the certainty matters. The two problems are built so that the difference between them is a component shared by both options within each problem. Under independence a shared component cannot matter, so the problems should be answered alike. The empirical finding is that they are not, and the salient difference is that in problem 1 one option is certain, while in problem 2 nothing is. Moving from 0.89 to 0.90 feels ordinary; moving from 0.99 to 1.00 does not.

How weighting reproduces it. Under the Prelec form with parameter 0.65 , a probability of 0.01 enters as 0.0673 and one of 0.90 enters as 0.7933 : small chances loom larger than they are, large ones smaller. The function crosses the diagonal at exactly 1 / e ≈ 0.367879 , below which probabilities are overweighted and above which they are underweighted. Because w ( 0.89 ) ≠ 0.89 , the common component no longer cancels, and a single coherent model can produce both observed choices.

A caution about what has been shown. That a model predicts the pattern is one claim; that the pattern is an error to be corrected is another, requiring an argument that independence ought to govern. Allais proposed the example precisely to dispute that, and the dispute is not settled by the arithmetic above.

Example

Five choice patterns and what each one shows

Declining a favourable gamble. Offered 1,000,000 certain against a prospect worth 1,390,000 , most people take the certainty. This violates nothing: it requires only u ( 1 M ) > 10 / 11 , which any sufficiently concave utility function supplies. Risk aversion is a taste the theory represents, not an anomaly.

Preferring a cycle. Someone who prefers P to Q , Q to R and R to P violates transitivity, and no utility function represents that either, but for a different reason, and one most people accept as a genuine error, since the cycle can be exploited by an intermediary who trades around it repeatedly. The Allais pattern has no such exploit, which is part of why it is contested.

The Allais pair. Both preferences are individually unremarkable and jointly unrepresentable, with the two demands meeting at exactly 10 / 11 . Completeness, transitivity and continuity all survive; independence does not.

The same prospect described two ways. A treatment presented as saving 200 of 600 lives, against one presented as leaving 400 of 600 to die, produces different choices though the outcomes are identical. Expected utility is defined over outcomes and cannot distinguish the two descriptions, so this is a violation of a different kind: the theory does not merely predict the wrong preference, it has no place for the distinction that drives it. A reference-dependent value function does.

Buying insurance and a lottery ticket at once. Paying a premium above expected loss indicates risk aversion; buying a ticket worth less than its price indicates risk seeking. One concave utility function cannot do both, but a weighting function that overweights small probabilities can, since w ( 0.01 ) = 0.0673 against a true 0.01 makes an unlikely jackpot loom larger than it is while an unlikely catastrophe does the same.

---

The first is permitted, the second is a violation nearly everyone regards as an error, the third is a violation many defend, the fourth is outside what the theory can express at all, and the fifth is a pattern that one model forbids and another explains. "Violates expected utility" covers all four of the latter and means something different in each.

Procedure

Testing a pattern of choices against the theory

To test whether a set of preferences is representable.

  1. Write each prospect as a distribution over outcomes, with probabilities summing to one. Include zero-probability outcomes explicitly if it helps the comparison line up.
  2. Count the distinct outcomes across all problems. With k outcomes, the normalisation fixes two and leaves k − 2 unknowns.
  3. Normalise at the extremes: u ( worst ) = 0 and u ( best ) = 1 . This is free, because expected utility is invariant to positive affine transformations.
  4. Convert each stated preference into an inequality in the remaining unknowns, by writing both expected utilities and keeping the direction of the preference.
  5. Ask whether the system has a solution. With one unknown this is immediate: collect the inequalities and see whether they leave an interval. An empty interval means no utility function represents the pattern.
  6. Report the threshold, not merely the emptiness. Two constraints meeting at a single point is a sharper finding than two that miss by a wide margin, because it shows the impossibility is structural rather than parametric.

To locate the axiom that fails.

  1. Look for a component common to both options within a problem. The same outcome at the same probability on each side.
  2. Strip it from both options. If the residual comparisons in two problems are identical, independence requires the same answer to both.
  3. If the answers differ, independence is the casualty. Check transitivity and completeness separately before naming anything else: the Allais pattern violates neither, and saying so precisely matters.

To model the pattern rather than diagnose it.

  1. Choose what to relax. Non-linear probability weighting breaks the cancellation; a reference-dependent value function handles patterns involving gains against losses.
  2. Fit or assume a weighting function, and verify it reproduces the observed choices rather than assuming it does.
  3. State what the model establishes: it predicts the pattern. It does not establish that the pattern is correct, and a descriptive model has no normative authority on its own.

Checks. Confirm each prospect's probabilities sum to one. Confirm the normalisation is at the extreme outcomes, since normalising in the middle can hide a sign error. And before describing any choice as a mistake, verify that the pattern is genuinely unrepresentable. A single choice never is, and departure from expected value never is.

Worked example

Two problems, four prospects, one impossible utility

The problems. Amounts in millions.

Problem 1.

Outcome and probability
A 1 M with certainty
B 5 M with 0.10 ; 1 M with 0.89 ; 0 with 0.01

Problem 2.

Outcome and probability
C 1 M with 0.11 ; 0 with 0.89
D 5 M with 0.10 ; 0 with 0.90

Step 1: expected values.

E [ A ] = 1,000,000 E [ B ] = 0.10 ( 5,000,000 ) + 0.89 ( 1,000,000 ) = 1,390,000 E [ C ] = 0.11 ( 1,000,000 ) = 110,000 E [ D ] = 0.10 ( 5,000,000 ) = 500,000

The modal choices are A over B , and D over C .

Step 2: note what is not yet wrong. Choosing A means taking 1,000,000 over a prospect worth 1,390,000 . That is risk aversion, which expected utility represents with a concave u . Nothing has been violated.

Step 3: normalise. Expected utility is invariant to positive affine transformations, so two values may be chosen freely. Set

u ( 0 ) = 0 , u ( 5 M ) = 1 , u ( 1 M ) = u  (unknown) .

Step 4: convert each preference into an inequality.

A ≻ B gives

u > 0.10 ( 1 ) + 0.89 u + 0.01 ( 0 ) = 0.10 + 0.89 u ,

so 0.11 u > 0.10 and therefore

u > 10 11 = 0.909091 .

D ≻ C gives

0.10 ( 1 ) + 0.90 ( 0 ) > 0.11 u + 0.89 ( 0 ) ,

so 0.10 > 0.11 u and therefore

u < 10 11 = 0.909091 .

Step 5: read the result. The two requirements are u > 10 / 11 and u < 10 / 11 . They meet at exactly the same value and point in opposite directions, so no utility function whatever represents both preferences. This is not a matter of the utility being implausible; there is no candidate at all.

Step 6: locate the axiom. Both options in problem 1 contain a 0.89 chance of 1 M ; both in problem 2 contain a 0.89 chance of 0 . Strip the common component from each pair and the identical residual remains in both problems:

1 M with probability  0.11 against 5 M with probability  0.10 .

Independence says a common component cannot affect a ranking, so the two problems are the same problem and must be answered alike. The observed pattern answers them differently, which is precisely the violation.

Step 7: what a weighting function does to it. Suppose probabilities enter as w ( p ) rather than p , with the Prelec form w ( p ) = exp ⁡ ( − ( − ln ⁡ p ) 0.65 ) :

p w ( p )
0.01 0.0673 overweighted
0.10 0.1791 overweighted
0.367879 0.367879 fixed point, exactly 1 / e
0.50 0.4547 underweighted
0.90 0.7933 underweighted

The common component now contributes w ( 0.89 ) u ( 1 M ) in problem 1 and w ( 0.89 ) u ( 0 ) = 0 in problem 2. These differ, so the cancellation fails and both observed choices can come from one coherent model.

What step 7 does not do. It explains the pattern; it does not show the pattern is correct. Whether independence ought to govern choice is a separate argument, and it is the argument Allais was making when he constructed the example.

Contrast

Pairs that differ in one respect

A single choice against a pattern.

taking 1 M over a 1.39 M prospectthe Allais pair
requires u > 10 / 11 u > 10 / 11 and u < 10 / 11
representableyes, by any concave u by no u whatever
what it showsa risk attitudea failure of the axioms

No single choice can contradict expected utility, because a utility function can always be built to accommodate it. Only a pattern can.

Problem 1 against problem 2.

They differ by exactly one substitution: a common 0.89 chance of 1 M becomes a common 0.89 chance of nothing. Strip the common part from both options in each problem and the identical residual remains, 1 M at 0.11 against 5 M at 0.10 . Independence therefore requires the same answer twice.

Independence against transitivity.

violated byexploitable
transitivitya preference cycleyes, by repeated trades
independencethe Allais patternno known mechanism

Both make preferences unrepresentable. The first is almost universally regarded as an error because it can be turned into a money pump; the second is defended by some on the grounds that certainty is a legitimate thing to value. The distinction matters when deciding whether to correct behaviour or model it.

Probability against decision weight.

p w ( p )
0.01 0.0673
0.10 0.1791
0.367879 0.367879
0.90 0.7933

A decision-maker may know the probability exactly and still act on the weight. That is why the distortion is not a belief error and is not removed by better information, and why the fixed point sits at 1 / e rather than at 0.5 , which a symmetric distortion would give.

A descriptive model against a normative one.

Prospect theory predicts the Allais pattern; expected utility says the pattern cannot be represented. Neither settles whether somebody choosing that way has erred. That question needs an argument about whether independence ought to govern, and the arithmetic is silent on it.

Warning

Claims the demonstration does not support

That people are irrational. What the Allais pair establishes is that no utility function represents both preferences, so the axioms fail for that decision-maker. Calling the result irrationality imports a judgement the algebra does not contain, and the axiom at issue has been disputed on its own terms since the example was constructed. Contrast a preference cycle, which can be turned into a money pump. There is no comparable mechanism here.

That expected utility is refuted by risk aversion. Taking 1,000,000 over a prospect worth 1,390,000 is represented by any concave utility function. A theory that accommodates every risk attitude is not embarrassed by one, and citing such a choice as evidence against it misreads what the theory claims.

That the theory predicts how people think. It is a representation theorem: if the axioms hold, a function exists reproducing the choices. Evidence that nobody computes expected utilities bears on nothing, because the theory never asserted they do.

That prospect theory shows the choices are correct. It shows they are describable by a coherent model. A descriptive model has no normative authority, and predicting a pattern is not endorsing it. The same model predicts choices most people would want to revise on reflection.

---

Two errors in applying the framework.

Treating a decision weight as a belief. Under the Prelec form a probability of 0.01 enters as 0.0673 . This is not a mistaken estimate of the chance; the decision-maker may know it exactly. Supplying better information therefore does not remove the effect, which is why interventions built on that assumption tend to fail.

Fitting a weighting function and declaring the matter settled. A flexible enough function reproduces most patterns, so reproduction is weak evidence. What makes prospect theory substantive is that it was specified in advance and predicts particular reversals, not that it can be made to fit.

---

And one about scope. Everything above concerns choices among stated prospects with known probabilities. Decisions under ambiguity, where probabilities are not given, raise separate failures that this apparatus does not address, and extending the conclusions there is unwarranted.

Application

Where the departure from expected utility has consequences

Insurance pricing. People buy cover against low-probability losses at premiums well above expected loss, and simultaneously buy lottery tickets. One concave utility function cannot produce both; overweighting small probabilities produces both at once. Insurers price against observed demand rather than against expected loss for this reason, and the gap is largest exactly where probabilities are smallest.

Retirement saving defaults. Enrolment rates differ sharply between opt-in and opt-out schemes offering identical terms. Under expected utility the default is irrelevant, since it changes no outcome and no probability. Under a reference-dependent account the default sets the reference point, so leaving it is a loss and inertia has a cost. Whether the change should be made is a policy argument the model does not settle.

Medical risk communication. The same treatment described by survival rate and by mortality rate produces different decisions from the same patients. The framing carries no information, which is what makes it a violation rather than a preference: the theory defines choice over outcomes, and these descriptions share their outcomes exactly.

Regulatory cost-benefit analysis. Expected-value reasoning weights a one-in-a-million fatality risk linearly; public willingness to pay for its reduction is consistently higher. Treating that as error implies overriding stated preferences, and treating it as taste implies spending against the linear calculus. Agencies handle this explicitly rather than by assumption, because the arithmetic does not decide it.

Financial product design. Structured products offering capital protection with limited upside sell well against their expected value. The certainty of the protected floor is exactly the feature the Allais pattern shows people pay for beyond what linear probability weighting justifies.

---

The recurring shape. In each case a model that fits behaviour better is available, and in none does the better fit decide what to do. The design question, nudge, override, inform, or leave alone, turns on whether the pattern is judged an error or a preference, and that judgement is argued rather than computed.

Next step

Practice Choices No Utility Function Represents

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.