Estimators and How They Are Judged

An estimator as a random variable rather than a number, the two standard ways of constructing one, and the properties that decide between competitors: bias, variance, their combination as mean squared error, consistency, and why an unbiased estimator is not automatically the better choice.

Definition

An estimator θ ^ = T ( X 1 , … , X n ) is a function of the sample, so it is a random variable with a distribution of its own. An estimate is the value it takes on one realised sample. Properties below describe the estimator; none of them is a property of the estimate.

Bias is bias ⁡ ( θ ^ ) = E [ θ ^ ] − θ , and θ ^ is unbiased when this is zero for every θ .

Mean squared error combines the two ways an estimator can be wrong:

MSE ( θ ^ ) = E [ ( θ ^ − θ ) 2 ] = bias ⁡ ( θ ^ ) 2 + Var ⁡ ( θ ^ ) .

Consistency is a statement about the limit: θ ^ n → θ in probability as n → ∞ . It concerns a sequence of estimators, one per sample size, and is logically independent of unbiasedness at any fixed n .

Efficiency compares variances among estimators of the same class, the more efficient being the one with smaller variance.

The method of moments equates sample moments to population moments and solves. With one unknown parameter, set X ¯ = E θ [ X ] and solve for θ .

Maximum likelihood takes the likelihood L ( θ ) = ∏ i f ( x i ; θ ) , the probability of the observed sample read as a function of θ , and returns the maximising value. In practice one maximises ℓ ( θ ) = log ⁡ L ( θ ) , solving the score equation ℓ ′ ( θ ) = 0 and confirming the stationary point is a maximum.

Assumptions and scope

  • Unbiasedness must hold for every value of the parameter, not merely for the one that generated a particular dataset. An estimator unbiased only at a specific θ is not unbiased.

  • Maximum likelihood is not in general unbiased. The maximum likelihood estimator of a uniform upper endpoint is the sample maximum, which can never exceed the truth and so underestimates it at every sample size.

  • Solving the score equation locates a stationary point. Confirming it is a maximum, and checking the boundary of the parameter space, are part of the derivation rather than formalities; for the uniform family the likelihood is maximised at a boundary where the derivative does not vanish at all.

  • Consistency and unbiasedness are logically independent. The variance estimator with divisor n is biased at every n and consistent; the first observation alone is unbiased for the mean at every n and not consistent.

  • Comparing estimators by mean squared error requires a common target. Two estimators of different quantities cannot be ranked by it, however similar their formulas look.

Worked material

Example

Four estimators and what each property says about them

Each case fixes a model and an estimator, then asks the three questions separately.

The sample mean for a population mean. Unbiased at every n , since E [ X ¯ ] = μ whatever the distribution. Variance σ 2 / n , which falls with n , so it is also consistent. This is the case where the properties agree, and it is why the sample mean is the default, not because it is obvious, but because it satisfies every criterion at once.

The sample variance with divisor n − 1 . Unbiased for σ 2 ; the n − 1 exists precisely to make it so, compensating for estimating μ by X ¯ . Consistent as well. Its square root is not unbiased for σ : an unbiased estimator of a quantity does not give an unbiased estimator of a nonlinear function of it, and s understates σ on average.

The sample variance with divisor n . This is the maximum likelihood estimator under normality, and it is biased at every sample size, with E [ s n 2 ] = n − 1 n σ 2 :

n E [ s n 2 ]
2 0.5000 σ 2
5 0.8000 σ 2
10 0.9000 σ 2
50 0.9800 σ 2
1000 0.9990 σ 2

The bias never reaches zero at any finite n , and it goes to zero in the limit. Biased at every sample size, consistent nonetheless.

The first observation, as an estimator of the mean. E [ X 1 ] = μ , so it is unbiased at every n , perfectly, not approximately. Its variance is σ 2 regardless of how many observations were collected, so the estimator never improves and is not consistent. It is unbiased and useless, which is the sharpest available demonstration that unbiasedness alone is not a recommendation.

---

The first satisfies everything. The second and third differ only in a divisor and land on opposite sides of unbiasedness while both being consistent. The fourth is unbiased and inconsistent. So among these four, unbiasedness and consistency appear in all four combinations that matter, which is what it means for the two properties to be logically independent rather than one implying the other.

Contrast

Pairs that differ in one respect

Unbiased against minimum mean squared error, on the uniform family.

Estimatorbias²varianceMSE (units of θ 2 )
2 X ¯ 0 0.041667 0.041667
M (sample maximum) 0.012346 0.009877 0.022222
n + 1 n M 0 0.012500 0.012500

The first and third are unbiased and their mean squared errors differ by a factor of 3.3 . So unbiasedness does not determine the ranking in either direction: the best and the worst of these three are both unbiased, and the biased one sits between them.

Divisor n − 1 against divisor n in the sample variance.

One change to a denominator moves the estimator across the unbiasedness line. Both are consistent, both converge to the same thing, and the one that is biased is the maximum likelihood estimator. The choice is conventional rather than forced, and the convention exists because unbiasedness is easy to state, not because it is the better criterion here.

Method of moments against maximum likelihood.

momentslikelihood
usesa few sample summariesthe whole assumed density
closed formusuallysometimes
efficiency when the model is rightoften lowerusually higher
behaviour when the model is wrongmay still target something meaningfulmaximises the wrong function

For the Poisson family they coincide exactly, both giving X ¯ , because E [ X ] = λ makes the moment equation and the score equation the same equation. Agreement is common; the uniform case, where they give 12.0000 and 9.4000 , is the one that teaches the difference.

Unbiased for σ 2 against unbiased for σ .

s 2 is unbiased for the variance and s is not unbiased for the standard deviation. Unbiasedness is not preserved by nonlinear transformation, so "an unbiased estimate of the variance" and "an unbiased estimate of the spread" are different claims and only the first is true.

Consistency against unbiasedness, as claims.

The first says the estimator eventually concentrates on the truth and promises nothing about the data in hand. The second says the estimator is centred correctly for the data in hand and promises nothing about improvement. A study with n = 30 is not helped by a limiting guarantee, and a study accumulating data indefinitely is not helped by correct centring around an unchanging spread.

Common errors

Common misconception

That an unbiased estimator is always preferable to a biased one, since being centred on the truth is the point of estimating. Unbiasedness constrains where the sampling distribution is centred and says nothing about how far it spreads, and mean squared error, which combines both as bias 2 + Var , can favour the biased competitor. For a uniform distribution on ( 0 , θ ) with eight observations, the unbiased method-of-moments estimator 2 X ¯ has mean squared error 0.0417 θ 2 while the biased maximum likelihood estimator, the sample maximum, has 0.0222 θ 2 , smaller by a factor of 1.875 despite understating θ at every sample size. Unbiasedness is one criterion among several, and imposing it is a restriction that has a cost.

Common misconception

That bias, variance and consistency describe the number a study reported. They describe the rule that produced it. An estimator is a function of the sample and therefore a random variable with a distribution; an estimate is one realised value of it, and a single number has no variance and no bias. Saying that a reported figure of 2.14 is unbiased is a category error: what may be unbiased is the procedure that would have produced a different figure from a different sample. The confusion matters because it makes the properties sound checkable against the data in hand, when every one of them is a statement about samples that were not drawn.

Related units

Requires

Connected

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.