Covariance, Independence and the Variance of a Sum

What you will be able to do

The learner can determine whether the variance of a sum decomposes, justifying the decision by the covariance term, and can distinguish independence from zero covariance by exhibiting or interpreting a pair that is uncorrelated yet dependent.

Orientation

The one rule that carries a condition

Expectations add. Variances usually do not.

E [ X + Y ] = E [ X ] + E [ Y ]

holds for any two random variables, however they are related. It needs no hypothesis and admits no exception.

Var ⁡ ( X + Y ) = Var ⁡ ( X ) + Var ⁡ ( Y )

holds only when the two are uncorrelated. In general there is a third term, and this unit is about what that term does.

The asymmetry is not a technicality to note and move past. It is the reason a standard error depends on how a sample was drawn: Var ⁡ ( X ¯ ) = σ 2 / n is the addition rule applied n times, so it inherits the condition. A clustered sample, a repeated-measures design, a time series, each breaks the condition rather than merely straining it, and each makes the familiar formula report an uncertainty smaller than the true one.

The second idea in this unit is a distinction that looks pedantic and is not. Uncorrelated and independent are different conditions. Independence is strictly stronger. Variance addition needs only the weaker one, which is worth knowing precisely because it means a dependent pair can still satisfy it, and because a reported correlation of zero is much weaker evidence than it is usually read as.

Definition

Why the cross term appears, and what it detects

Where the third term comes from. Variance involves squares, and ( X + Y ) 2 contains the cross term 2 X Y . Taking expectations,

Var ⁡ ( X + Y ) = Var ⁡ ( X ) + Var ⁡ ( Y ) + 2 ( E [ X Y ] − E [ X ] E [ Y ] ) ,

and the bracket is the covariance. It vanishes exactly when E [ X Y ] = E [ X ] E [ Y ] , which is the definition of uncorrelated. Expectation, by contrast, is a sum over outcomes with no squaring anywhere, and sums rearrange freely, which is why it adds unconditionally.

What the sign means. Cov ⁡ ( X , Y ) = E [ ( X − E [ X ] ) ( Y − E [ Y ] ) ] averages the product of the two deviations. When X and Y are usually on the same side of their means, the products are usually positive and the sum spreads further than the parts. When they are usually on opposite sides, the products are negative and the sum is tamer. For two independent dice, E [ X Y ] = 49 4 = E [ X ] E [ Y ] , the covariance is zero, and Var ⁡ ( X + Y ) = 35 6 = 2 ⋅ 35 12 , verified over all 36 outcomes.

The perfectly dependent case. If Y = X , then X + Y = 2 X and the addition rule does not apply at all; the scaling rule gives Var ⁡ ( 2 X ) = 4 Var ⁡ ( X ) , twice what independent addition would have given. Expectation is 2 E [ X ] either way and cannot see the difference.

Covariance is not scale-free. Multiplying X by 100 multiplies the covariance by 100 without changing the relationship, which is why the correlation ρ = Cov ⁡ ( X , Y ) / ( σ X σ Y ) is reported when magnitudes are to be compared. Correlation lies in [ − 1 , 1 ] and is zero exactly when the covariance is.

Independence, and why it is more. X and Y are independent when the joint distribution factors at every pair of values. That implies E [ X Y ] = E [ X ] E [ Y ] and hence zero covariance. The converse fails, because covariance averages a product of deviations and a symmetric relationship can make that average zero while the variables determine each other completely.

Example

A pair with a covariance that does not vanish

Every pair examined so far had zero covariance. Here is one that does not, worked end to end, so the cross term can be seen doing its work.

---

The setup. Draw one card from four, labelled 1 , 2 , 3 , 4 , each equally likely. Let

X = the number on the card , Z = { 1 if the number is  3  or  4 0 otherwise.

Z is an indicator for a high card. It is a function of X , so knowing the card fixes it.

The marginals.

E [ X ] = 1 4 ( 1 + 2 + 3 + 4 ) = 2.5 , Var ⁡ ( X ) = 1 4 ( 1 + 4 + 9 + 16 ) − 2.5 2 = 7.5 − 6.25 = 1.25 .

Z is 1 with probability 1 2 , so E [ Z ] = 0.5 and Var ⁡ ( Z ) = E [ Z 2 ] − E [ Z ] 2 = 0.5 − 0.25 = 0.25 , using Z 2 = Z for an indicator.

The covariance. X Z takes the value X when the card is high and 0 otherwise, so

E [ X Z ] = 1 4 ( 0 + 0 + 3 + 4 ) = 1.75 , Cov ⁡ ( X , Z ) = 1.75 − ( 2.5 ) ( 0.5 ) = 1.75 − 1.25 = 0.5 .

Positive, as expected: high cards are exactly when Z is large, so the two deviations share a sign more often than not.

What that does to the sum. Let W = X + Z . Adding the variances alone would give 1.25 + 0.25 = 1.5 . The true value includes the cross term:

Var ⁡ ( W ) = 1.25 + 0.25 + 2 ( 0.5 ) = 2.5 .

Verified directly. W takes 1 , 2 , 4 , 5 each with probability 1 4 , so E [ W ] = 3 and E [ W 2 ] = 1 4 ( 1 + 4 + 16 + 25 ) = 11.5 , giving Var ⁡ ( W ) = 11.5 − 9 = 2.5 .

Adding the variances would have understated the spread by 40%. The expectation, meanwhile, adds correctly without any condition: E [ W ] = 2.5 + 0.5 = 3 .

---

The same calculation with the sign flipped. Let Z ′ = 1 − Z , an indicator for a low card. Then E [ Z ′ ] = 0.5 and Var ⁡ ( Z ′ ) = 0.25 , unchanged, but

Cov ⁡ ( X , Z ′ ) = − Cov ⁡ ( X , Z ) = − 0.5 ,

so Var ⁡ ( X + Z ′ ) = 1.25 + 0.25 − 1 = 0.5 , less than either variance added naively would suggest. Checking: X + Z ′ takes 2 , 3 , 3 , 4 , with mean 3 and variance 1 4 ( 1 + 0 + 0 + 1 ) = 0.5 .

What the two cases show. The cross term is not a correction for sloppiness; it carries real information about the pair. When the variables move together the sum ranges further than the parts; when they move oppositely, the sum is tamer than either. Negative covariance is why a hedged portfolio can be steadier than either holding alone, and why a difference of two positively correlated measurements is less variable than a difference of independent ones.

Scale dependence, seen once. Measure the card in tens, X ∗ = 10 X . Then Cov ⁡ ( X ∗ , Z ) = 5 , ten times larger, though nothing about the relationship changed. The correlation is unmoved: ρ = 0.5 / 1.25 × 0.25 ≈ 0.894 before and after. That invariance is why correlation is what gets reported when magnitudes are to be compared.

Contrast

Uncorrelated does not mean independent

Zero covariance is often treated as a synonym for independence. It is strictly weaker, and the gap matters wherever a claim of "no relationship" is made.

---

The pair. Let X take the values − 1 , 0 , 1 each with probability 1 3 , and let Y = X 2 .

Y is a deterministic function of X , knowing X determines Y exactly. No two variables could be more dependent.

Yet the covariance is zero.

E [ X ] = 1 3 ( − 1 + 0 + 1 ) = 0 , E [ Y ] = 1 3 ( 1 + 0 + 1 ) = 2 3 ,
E [ X Y ] = E [ X 3 ] = 1 3 ( ( − 1 ) 3 + 0 + 1 3 ) = 0 ,

so

Cov ⁡ ( X , Y ) = E [ X Y ] − E [ X ] E [ Y ] = 0 − 0 ⋅ 2 3 = 0 .

All four values verified exactly in rational arithmetic.

But independence fails plainly. Independence would require P ( X = 0 , Y = 0 ) = P ( X = 0 ) P ( Y = 0 ) . The left side is 1 3 , since X = 0 forces Y = 0 . The right side is 1 3 × 1 3 = 1 9 . Not equal, so X and Y are dependent.

Equivalently: P ( Y = 0 ) = 1 3 unconditionally, but P ( Y = 0 ∣ X = 0 ) = 1 . Learning X changes everything about Y .

---

Why covariance missed it. Covariance measures linear association, whether Y tends to rise as X rises. Here Y falls as X goes from − 1 to 0 and rises as X goes from 0 to 1 . The relationship is perfect and symmetric, so the rising and falling halves cancel exactly, leaving zero.

Covariance does not report "no relationship". It reports no net linear tendency, which a symmetric non-linear relationship produces just as readily as no relationship at all.

---

The comparison.

X and Y = X 2 Two independent dice
Cov 0 0
E [ X Y ] = E [ X ] E [ Y ] yes ( 0 = 0 )yes ( 49 4 both)
Var ⁡ ( X + Y ) = Var ⁡ X + Var ⁡ Y yesyes
P ( joint ) = P ⋅ P no ( 1 3 vs 1 9 )yes
Independentnoyes

The variance-addition rule holds in both columns, which is the practically important point. Uncorrelatedness is exactly what that rule needs; independence is more than required.

So the logical structure is:

independent ⟹ Cov = 0 ⟹ Var ⁡ ( X + Y ) = Var ⁡ ( X ) + Var ⁡ ( Y ) ,

with neither implication reversible.

---

Where this bites. A reported correlation of zero between two variables licenses "no linear relationship", not "no relationship" and certainly not "independent". A U-shaped dose–response, harmful at low and high doses, beneficial in the middle, can show a correlation near zero while the dose determines the outcome entirely.

The converse error is rarer but worse: assuming independence from an observed zero correlation, then using it to justify something genuinely stronger, such as treating observations as independent draws when computing a standard error. Variance addition survives; most other consequences of independence do not.

Next step

Practice Covariance, Independence and the Variance of a Sum

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.