Module 1 of 2 · Lesson 1 of 1

Estimators and How They Are Judged

Bias, variance and consistency are properties of the rule, not of the number it returned.

What you will be able to do

The learner can derive an estimator by the method of moments and by maximum likelihood, decompose its mean squared error into bias and variance, and judge competing estimators by unbiasedness, consistency and efficiency rather than by their value on one sample.

Orientation

The rule, not the number

A study reports that the mean is 2.14 . Asking whether that number is unbiased is a category error: a single number has no bias, no variance and no limiting behaviour. What can have those properties is the rule that produced it, which would have produced something else from a different sample.

So the object of study here is the estimator as a random variable, and the questions are about its distribution. Where is it centred, relative to the truth? How widely does it scatter? Does it improve as data accumulate? Those three questions are bias, variance and consistency, and they are independent enough that an estimator can do well on one and badly on another.

This unit covers the two standard ways of constructing an estimator, matching moments, and maximising likelihood, and the criteria for choosing between the results. The central finding is one that sounds wrong at first: an unbiased estimator is not automatically the better one, and a worked comparison later in the unit exhibits a biased estimator whose mean squared error is smaller by a factor of 1.875 .

Definition

What each property quantifies over

The canonical statements above give the definitions. What distinguishes them is the quantifier each carries, and that is where the properties come apart.

PropertyQuantifies overHolds at
unbiasednessevery θ in the parameter spaceeach fixed n
efficiencycompeting estimators of the same θ each fixed n
consistencya sequence indexed by n the limit only

Unbiasedness is a claim about every parameter value. An estimator that happens to be centred correctly when θ = 3 and is off elsewhere is not unbiased. This matters because θ is unknown, so a property holding at one value is unusable.

Consistency is a claim about no particular n . It says the sequence θ ^ 1 , θ ^ 2 , … concentrates on θ , and it is compatible with every member of that sequence being biased. It also promises nothing about the sample actually in hand.

Mean squared error is the one that combines them. The decomposition

MSE ( θ ^ ) = bias ⁡ ( θ ^ ) 2 + Var ⁡ ( θ ^ )

follows from adding and subtracting E [ θ ^ ] inside the square and observing that the cross term vanishes. Both components are in squared units of θ , which is what makes them addable and what makes two estimators comparable on one scale.

Why the two construction methods can disagree. Method of moments uses a few summaries of the sample; maximum likelihood uses the whole assumed density. When the density carries information the moments do not, the likelihood estimator is typically more efficient. When the model is wrong, the likelihood estimator is maximising the wrong function, while a moment estimator may still target something interpretable. A robustness difference, not a precision one.

A boundary case. For a uniform family on ( 0 , θ ) the likelihood is zero for any θ below the sample maximum and decreasing above it, so it is maximised at the boundary and the score equation has no solution at all. Solving ℓ ′ ( θ ) = 0 is the usual route, not the definition, and a derivation that only differentiates will miss this case entirely.

Intuition

Two ways to be wrong, and why they add

Picture the estimator's distribution over repeated samples as a cloud of possible answers, with the true θ marked. Two things can be wrong with the cloud: it can sit off to one side, and it can be wide. Bias measures the first, variance the second.

Neither alone settles anything. A cloud centred exactly on θ but spread across half the plausible range delivers a poor answer most of the time. A tight cloud sitting slightly off-target delivers a near-miss every time, which is usually preferable.

Mean squared error prices them together. Squaring the bias puts it in the same units as the variance, so the two add, and the sum is the expected squared distance from the truth. That single number lets two estimators be compared without deciding in advance which kind of wrongness matters more.

Why insisting on unbiasedness costs something. Restricting attention to unbiased estimators is a constraint on the search. Constraints do not improve an optimum; they can only leave it unchanged or make it worse. So the best unbiased estimator is, in mean squared error, never better than the best estimator overall. In the uniform example later the gap is a factor of 1.875 .

Consistency is about a different thing entirely. It describes what happens as the sample grows: does the cloud tighten onto θ ? An estimator can be off-centre at every sample size with the offset shrinking toward zero, which is consistent and never unbiased. The variance estimator with divisor n behaves exactly this way, its expectation running 0.5000 , 0.8000 , 0.9000 , 0.9800 , 0.9990 times σ 2 as n goes 2 , 5 , 10 , 50 , 1000 .

The converse also occurs. Take the first observation and ignore the rest. It is unbiased for the mean at every sample size, and its variance is σ 2 no matter how much data arrives, so the cloud never tightens. Unbiasedness has been achieved while the data are being thrown away.

Method of moments asks: for which θ would the population average match what I observed? Maximum likelihood asks: for which θ was this exact sample most probable? The second uses more of the model, which usually gives lower variance when the model is right, and larger error when it is wrong.

Example

Four estimators and what each property says about them

Each case fixes a model and an estimator, then asks the three questions separately.

The sample mean for a population mean. Unbiased at every n , since E [ X ¯ ] = μ whatever the distribution. Variance σ 2 / n , which falls with n , so it is also consistent. This is the case where the properties agree, and it is why the sample mean is the default, not because it is obvious, but because it satisfies every criterion at once.

The sample variance with divisor n − 1 . Unbiased for σ 2 ; the n − 1 exists precisely to make it so, compensating for estimating μ by X ¯ . Consistent as well. Its square root is not unbiased for σ : an unbiased estimator of a quantity does not give an unbiased estimator of a nonlinear function of it, and s understates σ on average.

The sample variance with divisor n . This is the maximum likelihood estimator under normality, and it is biased at every sample size, with E [ s n 2 ] = n − 1 n σ 2 :

n E [ s n 2 ]
2 0.5000 σ 2
5 0.8000 σ 2
10 0.9000 σ 2
50 0.9800 σ 2
1000 0.9990 σ 2

The bias never reaches zero at any finite n , and it goes to zero in the limit. Biased at every sample size, consistent nonetheless.

The first observation, as an estimator of the mean. E [ X 1 ] = μ , so it is unbiased at every n , perfectly, not approximately. Its variance is σ 2 regardless of how many observations were collected, so the estimator never improves and is not consistent. It is unbiased and useless, which is the sharpest available demonstration that unbiasedness alone is not a recommendation.

---

The first satisfies everything. The second and third differ only in a divisor and land on opposite sides of unbiasedness while both being consistent. The fourth is unbiased and inconsistent. So among these four, unbiasedness and consistency appear in all four combinations that matter, which is what it means for the two properties to be logically independent rather than one implying the other.

Procedure

Deriving an estimator, and checking what it is worth

To derive a method-of-moments estimator.

  1. Express the population moment in the parameter. For one unknown, write E θ [ X ] as a function of θ .
  2. Set it equal to the sample moment X ¯ .
  3. Solve for θ . The solution is the estimator.
  4. For k unknown parameters, use k equations, adding E [ X 2 ] and higher moments as needed.

To derive a maximum likelihood estimator.

  1. Write the likelihood L ( θ ) = ∏ i f ( x i ; θ ) , treating the data as fixed and θ as the variable. Include the support: a density that is zero outside a range makes L zero for parameter values inconsistent with the data.
  2. Take logarithms. Products become sums and the maximiser is unchanged, since log is increasing.
  3. Differentiate and solve ℓ ′ ( θ ) = 0 .
  4. Confirm it is a maximum by the sign of ℓ ″ , or by evaluating ℓ on either side.
  5. Check the boundary. If the support depends on θ , the maximum may lie at an endpoint where the derivative does not vanish. Skipping this step loses the answer entirely for uniform and similar families.

To evaluate an estimator.

  1. Compute E [ θ ^ ] and subtract θ for the bias.
  2. Compute Var ⁡ ( θ ^ ) .
  3. Assemble MSE = bias 2 + Var , expressed in units of θ 2 so competitors are comparable.
  4. Check consistency by asking whether both bias and variance go to zero as n grows. Either one failing to vanish defeats it.

To compare two estimators. Put both mean squared errors over the same θ 2 and take the ratio. State the criterion explicitly: mean squared error is a choice, and an application penalising overestimates more than underestimates calls for a different one.

If an estimator is biased and otherwise attractive, try debiasing it. When E [ θ ^ ] = c ( n ) θ for a known c ( n ) , the rescaled θ ^ / c ( n ) is unbiased and inherits the shape of the original. For the uniform maximum this is n + 1 n M , which improves mean squared error rather than merely removing the bias.

Checks. Confirm the estimate lies in the parameter space. A variance estimate must be non-negative, a probability inside [ 0 , 1 ] , a uniform endpoint at least the sample maximum. An estimator that can return impossible values has been derived without attention to its support.

Worked example

The unbiased estimator that loses

The model. X 1 , … , X n independent and uniform on ( 0 , θ ) , with θ unknown. Eight observations:

3.2 ,   7.9 ,   5.1 ,   9.4 ,   2.6 ,   8.8 ,   6.3 ,   4.7 .

Sample mean X ¯ = 6.0 , sample maximum M = 9.4 .

Step 1: method of moments. For this family E [ X ] = θ / 2 . Setting X ¯ = θ / 2 gives

θ ^ MoM = 2 X ¯ = 12.0000 .

Step 2: maximum likelihood. The density is 1 / θ on ( 0 , θ ) and zero elsewhere, so

L ( θ ) = θ − n for  θ ≥ M , L ( θ ) = 0 for  θ < M .

Above M the likelihood is strictly decreasing, so it is largest at the smallest permitted value:

θ ^ MLE = M = 9.4000 .

Note what did not happen. There is no score equation here: ℓ ′ ( θ ) = − n / θ never vanishes. The maximum sits at a boundary, and a derivation that only differentiates finds nothing.

Step 3: bias. E [ 2 X ¯ ] = 2 ⋅ θ / 2 = θ , so the moment estimator is unbiased.

For the maximum, E [ M ] = n n + 1 θ , which at n = 8 is 8 9 θ = 0.8889 θ . Its bias is − θ / 9 . The sample maximum can never exceed θ , so it understates at every sample size. The estimator is biased by construction, not by accident.

On this sample the contrast is visible: 12.0000 against 9.4000 , and the true θ must be at least 9.4 .

Step 4: mean squared error. In units of θ 2 :

Estimatorbias²varianceMSE
2 X ¯ 0 1 3 n = 0.041667 0.041667
M 1 ( n + 1 ) 2 = 0.012346 n ( n + 1 ) 2 ( n + 2 ) = 0.009877 0.022222
MSE ( 2 X ¯ ) MSE ( M ) = 0.041667 0.022222 = 1.875 .

The biased estimator is better by a factor of 1.875 . Its bias contributes 0.012346 , and it more than repays that with a variance of 0.009877 against the moment estimator's 0.041667 . Unbiasedness optimised the wrong quantity.

Step 5: taking the best of both. Since E [ M ] = n n + 1 θ , the estimator n + 1 n M is unbiased, and its mean squared error is

1 n ( n + 2 ) = 0.012500 θ 2 ,

better than both. On this sample it gives 9 8 ( 9.4 ) = 10.5750 .

So the ranking is 0.0125 < 0.0222 < 0.0417 : debiased maximum, then raw maximum, then the moment estimator. Unbiasedness is neither necessary nor sufficient for the win. The debiased maximum is unbiased and best, and the moment estimator is unbiased and worst.

---

A second case, where the two methods agree. Twelve Poisson counts:

2 ,   0 ,   3 ,   1 ,   4 ,   2 ,   1 ,   3 ,   0 ,   2 ,   5 ,   1 , ∑ x i = 24 .

The log-likelihood is ℓ ( λ ) = − n λ + ( ∑ i x i ) log ⁡ λ − ∑ i log ⁡ ( x i ! ) . The score is − n + ( ∑ i x i ) / λ , vanishing at

λ ^ = 1 n ∑ i x i = 24 12 = 2.000000 .

The second derivative − ( ∑ i x i ) / λ 2 is negative, confirming a maximum. Checking the value directly: ℓ ( 2.0 ) = − 20.992974 , against ℓ ( 1.5 ) = − 21.897343 and ℓ ( 2.5 ) = − 21.637528 , lower by 0.904 and 0.645 .

Here the method of moments gives the same answer, because E [ X ] = λ makes the moment equation and the score equation coincide. The two methods agreeing is the common case; the uniform example is the instructive one.

Contrast

Pairs that differ in one respect

Unbiased against minimum mean squared error, on the uniform family.

Estimatorbias²varianceMSE (units of θ 2 )
2 X ¯ 0 0.041667 0.041667
M (sample maximum) 0.012346 0.009877 0.022222
n + 1 n M 0 0.012500 0.012500

The first and third are unbiased and their mean squared errors differ by a factor of 3.3 . So unbiasedness does not determine the ranking in either direction: the best and the worst of these three are both unbiased, and the biased one sits between them.

Divisor n − 1 against divisor n in the sample variance.

One change to a denominator moves the estimator across the unbiasedness line. Both are consistent, both converge to the same thing, and the one that is biased is the maximum likelihood estimator. The choice is conventional rather than forced, and the convention exists because unbiasedness is easy to state, not because it is the better criterion here.

Method of moments against maximum likelihood.

momentslikelihood
usesa few sample summariesthe whole assumed density
closed formusuallysometimes
efficiency when the model is rightoften lowerusually higher
behaviour when the model is wrongmay still target something meaningfulmaximises the wrong function

For the Poisson family they coincide exactly, both giving X ¯ , because E [ X ] = λ makes the moment equation and the score equation the same equation. Agreement is common; the uniform case, where they give 12.0000 and 9.4000 , is the one that teaches the difference.

Unbiased for σ 2 against unbiased for σ .

s 2 is unbiased for the variance and s is not unbiased for the standard deviation. Unbiasedness is not preserved by nonlinear transformation, so "an unbiased estimate of the variance" and "an unbiased estimate of the spread" are different claims and only the first is true.

Consistency against unbiasedness, as claims.

The first says the estimator eventually concentrates on the truth and promises nothing about the data in hand. The second says the estimator is centred correctly for the data in hand and promises nothing about improvement. A study with n = 30 is not helped by a limiting guarantee, and a study accumulating data indefinitely is not helped by correct centring around an unchanging spread.

Warning

Derivations that end one step early

Solving the score equation is not the definition of maximum likelihood. The definition is the maximiser of the likelihood. Differentiating is the usual route to it and fails whenever the support depends on the parameter. For the uniform family on ( 0 , θ ) , ℓ ′ ( θ ) = − n / θ never vanishes, and the maximum sits at the boundary θ = M . A derivation that sets the derivative to zero and finds no solution has not shown the estimator does not exist; it has used the wrong method.

A stationary point is not automatically a maximum. The score equation locates where the slope is zero, which includes minima and inflection points. Confirm with the sign of ℓ ″ , or by evaluating ℓ on either side. In the Poisson case, ℓ ( 2.0 ) = − 20.992974 against ℓ ( 1.5 ) = − 21.897343 and ℓ ( 2.5 ) = − 21.637528 settles it directly, and the second derivative − ( ∑ i x i ) / λ 2 < 0 settles it in general.

Unbiasedness does not survive a nonlinear transformation. s 2 unbiased for σ 2 does not make s unbiased for σ , nor 1 / X ¯ unbiased for 1 / μ . Whenever an estimator is transformed, its bias has to be recomputed rather than inherited.

A property proved at one parameter value is not the property. Unbiasedness quantifies over every θ . Since θ is unknown, an estimator centred correctly at one value and off elsewhere gives no usable guarantee.

---

Two errors about what the properties describe.

They are not checkable against the sample in hand. Bias, variance and consistency are features of a distribution over samples that were not drawn. No computation on one dataset verifies them; they are derived from the model or they are not known. An estimate close to a value believed to be true is not evidence of unbiasedness, and a bad estimate is not evidence against it.

A consistency guarantee says nothing about n = 30 . It describes a limit. An estimator can be badly behaved at every sample size a study could afford and still be consistent, so quoting consistency as reassurance about a small sample uses a theorem outside its scope.

---

Finally, check that the estimate is possible. A variance estimate must be non-negative, a probability must lie in [ 0 , 1 ] , and an estimate of a uniform endpoint must be at least the largest observation. The moment estimator 2 X ¯ = 12.0000 satisfies the last of these on the worked sample, but it need not in general: a sample whose mean is small relative to its maximum yields 2 X ¯ < M , an estimate the model rules out. An estimator that can return impossible values was derived without attention to the support, and that is a defect in the derivation rather than bad luck with the data.

Application

Where the choice of estimator is the decision

Estimating a population total from a survey. The design supplies unequal selection probabilities, and the standard estimator weights each response by the inverse of its probability. It is unbiased by construction, and when a few units carry very large weights its variance is severe. Survey practice therefore trims or smooths extreme weights, knowingly introducing bias to reduce mean squared error. The trade of this unit, made as routine methodology.

Variance components in quality control. Maximum likelihood estimates of variance components are biased downward, which matters when the estimate feeds a tolerance. Restricted maximum likelihood exists to correct that bias, and the choice between them is the choice between the more efficient estimator and the better centred one.

Shrinkage in small-area estimation. Estimating unemployment for a district with few sampled households, the direct estimate is unbiased and unstable. Shrinking it toward a regional average introduces bias deliberately and reduces mean squared error, often substantially. Published figures are shrunk, and the documentation says so, because an unbiased and wildly variable district estimate is worse for every use it has.

Estimating a maximum from observed data. Reconstructing a total from serial numbers, or a peak load from observed peaks, is the uniform-endpoint problem. The sample maximum understates by construction, and the correction factor n + 1 n is the standard fix. Here debiasing is not a refinement: the raw estimate is known to be too small every time, and reporting it as the total would be a systematic error with a known direction.

Machine learning regularisation. Ridge and lasso shrink coefficients toward zero, making them biased for the underlying parameters while reducing prediction variance. The justification is the decomposition of this unit applied to prediction error rather than to a parameter estimate, which is why the bias-variance argument recurs there in the same algebraic form.

---

The common decision. In each case an unbiased estimator was available and something else was preferred, for a reason that was stated. That is what the criteria are for: not to establish that unbiasedness is undesirable, but to make the trade explicit so it can be defended. A method chosen because it is unbiased, with no comparison made, has not been chosen at all.

Next step

Practice Estimators and How They Are Judged

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to Testing Counts Against a Claim

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.