Module 1 of 3 · Lesson 1 of 4

Random Variables, Expectation and Variance

What you will be able to do

The learner can compute the expectation and variance of a discrete random variable by either route, and apply the linearity and scaling rules to a linear transformation of it.

Orientation

What expectation and variance each report

Two numbers summarise a distribution, and they answer different questions.

A random variable is a function, not a variable. X attaches a number to each possible outcome, roll a die, and X reports the pips. The randomness lies in which outcome occurs; X itself is a fixed rule. That distinction is what makes expressions like E [ X ] and Var ⁡ ( 2 X + 3 ) meaningful.

Expectation is a weighted average. For a fair die,

E [ X ] = 1 6 ( 1 + 2 + 3 + 4 + 5 + 6 ) = 7 2 = 3.5 ,

a number no face shows. Expectation is where the distribution balances, not a prediction of any single roll. The first thing to unlearn about the name.

Variance measures spread. For the same die, Var ⁡ ( X ) = 35 12 ≈ 2.916667 , and σ = 35 / 12 ≈ 1.707825 in pips. Two distributions can share an expectation and differ in variance by a factor of ten, so a mean reported alone does not describe a distribution.

Both behave predictably under a linear transformation, and differently from each other. Adding a constant moves the mean and leaves the variance untouched. Multiplying by a doubles neither: it scales the mean by a and the variance by a 2 .

What this unit does not settle is what happens when two random variables are combined. E [ X + Y ] = E [ X ] + E [ Y ] always, but whether Var ⁡ ( X + Y ) is Var ⁡ ( X ) + Var ⁡ ( Y ) depends on how the two move together, which is the subject of the next unit.

Definition

Reading the definitions: why each rule takes the form it does

Why variance squares the deviations. The average deviation E [ X − E [ X ] ] is always zero, by linearity it equals E [ X ] − E [ X ] . The mean is by construction the point about which deviations cancel, so any measure that lets them cancel measures nothing. Squaring removes the signs. For the die: deviations − 2.5 , − 1.5 , − 0.5 , 0.5 , 1.5 , 2.5 sum to zero, while their squares average to 35 12 ≈ 2.916667 .

The cost is units. Squared pips are not pips, which is why σ = 35 / 12 ≈ 1.707825 is usually reported instead.

Why the two variance formulas agree. Expanding ( X − μ ) 2 = X 2 − 2 μ X + μ 2 and taking expectations gives E [ X 2 ] − 2 μ E [ X ] + μ 2 = E [ X 2 ] − μ 2 . Both routes give 35 12 for the die, verified exactly in rational arithmetic. The second is usually easier to compute, the first easier to interpret.

Why a is squared in Var ⁡ ( a X + b ) and b vanishes. Shifting every value by b moves the mean by b too, leaving all deviations unchanged, so a constant shift cannot alter spread. Scaling by a multiplies every deviation by a and hence every squared deviation by a 2 . For Y = 2 X + 3 on the die: E [ Y ] = 10 = 2 ( 3.5 ) + 3 and Var ⁡ ( Y ) = 35 3 = 4 ⋅ 35 12 , both exact.

Why E [ g ( X ) ] ≠ g ( E [ X ] ) for non-linear g . Linearity works because expectation is a sum over outcomes and a linear function passes through a sum. A non-linear function does not. With the die and g ( x ) = x 2 , E [ X 2 ] = 91 6 ≈ 15.166667 against ( E [ X ] ) 2 = 12.25 ; the difference is 35 12 , which is the variance, since Var ⁡ ( X ) = E [ X 2 ] − E [ X ] 2 . The size of that gap is what the variance measures, and it is zero exactly when X is constant.

Intuition

Balance points and moment arms

Place the probabilities as weights along a number line. Two pictures follow, and most of what is counterintuitive about expectation and variance dissolves in them.

Expectation is where the line balances. For a fair die the weights are equal and the balance point sits midway between 1 and 6, at 3.5, which is why no face shows it. A balance point need not coincide with any weight, and nothing requires it to be attainable. This is also why the expectation of a count can be 1.8 children per household: the balance point of a distribution over integers is rarely an integer.

Variance is the weighted average of squared distances from that point. Squaring is what stops the deviations cancelling, since the balance point is exactly where they sum to zero. It also means distant values count disproportionately: a value twice as far contributes four times as much. That is why a single extreme outcome can dominate a variance while barely moving a mean, and why a variance is sensitive to the tail of a distribution in a way an average is not.

The transformation rules read straight off the picture. Sliding every weight along the line by b moves the balance point by b and changes no distance, so the variance is untouched. Stretching the line by a factor a multiplies every distance by a , and every squared distance by a 2 . Neither rule requires anything of the distribution's shape.

Where the picture stops. It holds one variable. The moment two variables are combined, the question becomes whether their deviations reinforce or cancel, which no single number line can show. That is the subject of the unit on covariance.

Example

A die, and a linear transformation

One distribution, worked completely, then put through the rules.

---

1. A fair die. X takes 1 , … , 6 each with probability 1 6 .

Expectation.

E [ X ] = 1 6 ( 1 + 2 + 3 + 4 + 5 + 6 ) = 21 6 = 7 2 = 3.5 .

Variance, both ways. First by the definition, averaging squared deviations from 3.5:

1 6 [ ( − 2.5 ) 2 + ( − 1.5 ) 2 + ( − 0.5 ) 2 + ( 0.5 ) 2 + ( 1.5 ) 2 + ( 2.5 ) 2 ] = 1 6 [ 6.25 + 2.25 + 0.25 + 0.25 + 2.25 + 6.25 ] = 17.5 6 = 35 12 .

Then by the shortcut, with E [ X 2 ] = 1 6 ( 1 + 4 + 9 + 16 + 25 + 36 ) = 91 6 :

E [ X 2 ] − E [ X ] 2 = 91 6 − 49 4 = 182 − 147 12 = 35 12 .

Both give 35 12 ≈ 2.916667 , verified exactly in rational arithmetic. The standard deviation is 35 / 12 ≈ 1.707825 , in pips.

Note the deviations themselves sum to zero: − 2.5 − 1.5 − 0.5 + 0.5 + 1.5 + 2.5 = 0 . Squaring is what prevents that cancellation.

---

2. A linear transformation. Let Y = 2 X + 3 , doubling the pips and adding three.

By the rules: E [ Y ] = 2 ( 3.5 ) + 3 = 10 and Var ⁡ ( Y ) = 4 ⋅ 35 12 = 35 3 .

Verified directly by computing over Y 's values 5 , 7 , 9 , 11 , 13 , 15 : E [ Y ] = 10 and Var ⁡ ( Y ) = 35 3 ≈ 11.666667 , both exact.

The + 3 left the variance untouched while moving the mean by 3; the × 2 doubled the mean and quadrupled the variance. In standard deviations the scaling is gentler: σ Y = 2 σ X = 3.415650 , since the square root undoes the squaring.

---

3. Where this is heading. The same two rules applied to an average, X ¯ = 1 n ∑ X i , give E [ X ¯ ] = μ directly, since expectation passes through the sum and the 1 n . The variance needs one more ingredient, a condition on how the observations relate to one another, and that is taken up in the next unit.

Procedure

Computing an expectation and a variance

To compute an expectation.

Step 1 — List the values and their probabilities, and check they sum to 1. A mass function failing this is misread or mis-specified, and everything downstream inherits the error.

Step 2 — Multiply and add: E [ X ] = ∑ x x p ( x ) .

Step 3 — For a transformed variable, prefer the rules. E [ a X + b ] = a E [ X ] + b avoids re-tabulating. For a non-linear transformation there is no such shortcut: E [ g ( X ) ] = ∑ x g ( x ) p ( x ) must be computed term by term, and E [ g ( X ) ] ≠ g ( E [ X ] ) in general.

To compute a variance.

Step 4 — Use E [ X 2 ] − E [ X ] 2 , computing E [ X 2 ] = ∑ x x 2 p ( x ) . It is usually less work than averaging squared deviations, and the two agree by construction.

Step 5 — Sanity-check the result. Variance is never negative; if it comes out so, the subtraction was done in the wrong order or E [ X 2 ] is wrong. It should also be roughly the square of a plausible spread: for values in [ 1 , 6 ] , a variance near 3 is sensible and one near 300 is not.

Step 6 — For transformations, square the multiplier and ignore the shift: Var ⁡ ( a X + b ) = a 2 Var ⁡ ( X ) .

Where it goes wrong.

  • Treating E [ X ] as a typical value. A fair die averages 3.5 and shows no 3.5.
  • Forgetting the square in Var ⁡ ( a X + b ) , giving a instead of a 2 .
  • Adding b to the variance. A shift cannot change a spread.
  • Applying E [ g ( X ) ] = g ( E [ X ] ) , which fails for every non-linear g .
  • Stopping at E [ X 2 ] without subtracting E [ X ] 2 .
  • Reporting a variance where a standard deviation was wanted, leaving the answer in squared units.

Warning

Interpreting an expected value

An expected value is a probability-weighted average of the values a random variable takes. It is not a prediction of any single outcome and need not be a value the variable can take.

A fair die has E [ X ] = 3.5 , and no face shows 3.5 . The expectation has probability zero of occurring.

Other cases behave the same way. A variable taking 0 or 1 with equal probability has E [ X ] = 0.5 ; a count of children per household can have expectation 1.8 . Both numbers are correct and neither is an attainable value.

Expectation does not determine spread. These two distributions share E [ X ] = 3.5 :

DistributionValues Var
Fair die 1 , … , 6 equally likely 35 12 ≈ 2.92
Two values 3 or 4 , equally likely 0.25

The variances differ by a factor of about 11.7 . A mean reported without a measure of spread does not describe the distribution.

Expectation is a property of the distribution, not of a sample. E [ X ] = 3.5 holds for the die regardless of any sequence of rolls. A sample mean of 3.2 from twenty rolls is an estimate of it. The sample mean varies across samples; E [ X ] does not.

This distinction also separates two different quantities. Var ⁡ ( X ) describes the spread of a single observation of X . The uncertainty of an estimator, such as X ¯ , is the spread of the estimator across repeated samples, Var ⁡ ( X ¯ ) , whose square root is the standard error. For n independent observations with common variance σ 2 these are related by Var ⁡ ( X ¯ ) = σ 2 / n , but they answer different questions and are reported for different purposes.

E [ g ( X ) ] ≠ g ( E [ X ] ) for non-linear g . Linearity gives E [ 2 X + 3 ] = 2 E [ X ] + 3 , and this does not extend to non-linear transformations. With the die and g ( x ) = x 2 :

E [ X 2 ] = 91 6 ≈ 15.166667 , ( E [ X ] ) 2 = 3.5 2 = 12.25 .

The difference is 35 12 , which is the variance, since Var ⁡ ( X ) = E [ X 2 ] − E [ X ] 2 . The size of the gap is what the variance measures.

Consequences: the average of squares is at least as large as the square of the average, with equality exactly when X is constant, since the gap is Var ⁡ ( X ) ≥ 0 ; the expected return of a volatile investment differs from the return computed at its expected value; and an average of transformed data is not the transform of the average.

In reporting, "the long-run average is 3.5 " states what the expectation is. "Expect 3.5 " suggests a prediction for a single roll, which the expectation does not provide.

Next step

Practice Random Variables, Expectation and Variance

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to Covariance, Independence and the Variance of a Sum

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.