Random Variables, Expectation and Variance

A random variable as a numeric function of an uncertain outcome, its expectation as the probability-weighted balance point of the distribution, and its variance as the weighted average squared deviation from that point, together with the linearity and scaling rules that govern how both behave under a linear transformation.

Definition

Outcomes and events. An experiment has a sample space Ω of possible outcomes; an event is a subset of Ω . A probability assigns each event a number in [ 0 , 1 ] with P ( Ω ) = 1 and P ( A ∪ B ) = P ( A ) + P ( B ) for disjoint A , B .

Random variables. A random variable X is a function from Ω to the reals: it attaches a number to each outcome. It is neither random nor a variable in the algebraic sense: the randomness is in which outcome occurs, and X merely reports a number once one does. A discrete X takes countably many values with probability mass function p ( x ) = P ( X = x ) .

Expectation. For discrete X ,

E [ X ] = ∑ x x p ( x ) ,

the average of the values weighted by their probabilities. It is a property of the distribution, not of any sample, and need not be an attainable value: a fair die has E [ X ] = 7 / 2 .

Variance. The spread about the mean,

Var ⁡ ( X ) = E [ ( X − E [ X ] ) 2 ] = E [ X 2 ] − E [ X ] 2 ,

with standard deviation σ = Var ⁡ ( X ) , in the units of X itself.

Linear transformations. For constants a , b :

RuleHolds
E [ a X + b ] = a E [ X ] + b always
E [ X + Y ] = E [ X ] + E [ Y ] always, even if dependent
Var ⁡ ( a X + b ) = a 2 Var ⁡ ( X ) always

Expectation passes through a linear transformation unchanged. Variance ignores the shift b , because shifting every value moves the mean equally and alters no deviation, and picks up a 2 rather than a , because scaling multiplies every deviation by a and hence every squared deviation by a 2 .

For a non-linear g there is no corresponding shortcut: E [ g ( X ) ] = ∑ x g ( x ) p ( x ) must be computed term by term, and E [ g ( X ) ] ≠ g ( E [ X ] ) in general. The gap in the quadratic case is exactly the variance, since E [ X 2 ] − E [ X ] 2 = Var ⁡ ( X ) .

Whether Var ⁡ ( X + Y ) decomposes into Var ⁡ ( X ) + Var ⁡ ( Y ) is a separate question, governed by a condition on the joint behaviour of X and Y rather than by either distribution alone.

Assumptions and scope

  • E [ X ] need not be a value X can attain, and is a property of the distribution rather than of any observed sample.

  • E [ g ( X ) ] = g ( E [ X ] ) holds for linear g only. For any non-linear g the two differ in general, and in the quadratic case the difference is exactly Var ⁡ ( X ) .

  • Var ⁡ ( X ) describes the spread of a single observation of X . The spread of an estimator computed from a sample is a different quantity, answering a different question, and is derived in the units on sampling distributions.

  • Whether the variance of a sum is the sum of the variances is not settled by this unit: it depends on how the two variables move together.

Worked material

Example

A die, and a linear transformation

One distribution, worked completely, then put through the rules.

---

1. A fair die. X takes 1 , … , 6 each with probability 1 6 .

Expectation.

E [ X ] = 1 6 ( 1 + 2 + 3 + 4 + 5 + 6 ) = 21 6 = 7 2 = 3.5 .

Variance, both ways. First by the definition, averaging squared deviations from 3.5:

1 6 [ ( − 2.5 ) 2 + ( − 1.5 ) 2 + ( − 0.5 ) 2 + ( 0.5 ) 2 + ( 1.5 ) 2 + ( 2.5 ) 2 ] = 1 6 [ 6.25 + 2.25 + 0.25 + 0.25 + 2.25 + 6.25 ] = 17.5 6 = 35 12 .

Then by the shortcut, with E [ X 2 ] = 1 6 ( 1 + 4 + 9 + 16 + 25 + 36 ) = 91 6 :

E [ X 2 ] − E [ X ] 2 = 91 6 − 49 4 = 182 − 147 12 = 35 12 .

Both give 35 12 ≈ 2.916667 , verified exactly in rational arithmetic. The standard deviation is 35 / 12 ≈ 1.707825 , in pips.

Note the deviations themselves sum to zero: − 2.5 − 1.5 − 0.5 + 0.5 + 1.5 + 2.5 = 0 . Squaring is what prevents that cancellation.

---

2. A linear transformation. Let Y = 2 X + 3 , doubling the pips and adding three.

By the rules: E [ Y ] = 2 ( 3.5 ) + 3 = 10 and Var ⁡ ( Y ) = 4 ⋅ 35 12 = 35 3 .

Verified directly by computing over Y 's values 5 , 7 , 9 , 11 , 13 , 15 : E [ Y ] = 10 and Var ⁡ ( Y ) = 35 3 ≈ 11.666667 , both exact.

The + 3 left the variance untouched while moving the mean by 3; the × 2 doubled the mean and quadrupled the variance. In standard deviations the scaling is gentler: σ Y = 2 σ X = 3.415650 , since the square root undoes the squaring.

---

3. Where this is heading. The same two rules applied to an average, X ¯ = 1 n ∑ X i , give E [ X ¯ ] = μ directly, since expectation passes through the sum and the 1 n . The variance needs one more ingredient, a condition on how the observations relate to one another, and that is taken up in the next unit.

Common errors

Common misconception

The expected value is the value to expect, a typical or most likely outcome, so E [ X ] should be something X can actually take.

Related units

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.