Module 1 of 2 · Lesson 1 of 1
Estimators and How They Are Judged
Bias, variance and consistency are properties of the rule, not of the number it returned.
What you will be able to do
The learner can derive an estimator by the method of moments and by maximum likelihood, decompose its mean squared error into bias and variance, and judge competing estimators by unbiasedness, consistency and efficiency rather than by their value on one sample.
Orientation
The rule, not the number
A study reports that the mean is
So the object of study here is the estimator as a random variable, and the questions are about its distribution. Where is it centred, relative to the truth? How widely does it scatter? Does it improve as data accumulate? Those three questions are bias, variance and consistency, and they are independent enough that an estimator can do well on one and badly on another.
This unit covers the two standard ways of constructing an estimator, matching moments, and maximising likelihood, and the criteria for choosing between the results. The central finding is one that sounds wrong at first: an unbiased estimator is not automatically the better one, and a worked comparison later in the unit exhibits a biased estimator whose mean squared error is smaller by a factor of
Definition
What each property quantifies over
The canonical statements above give the definitions. What distinguishes them is the quantifier each carries, and that is where the properties come apart.
| Property | Quantifies over | Holds at |
|---|---|---|
| unbiasedness | every | each fixed |
| efficiency | competing estimators of the same | each fixed |
| consistency | a sequence indexed by | the limit only |
Unbiasedness is a claim about every parameter value. An estimator that happens to be centred correctly when
Consistency is a claim about no particular
Mean squared error is the one that combines them. The decomposition
follows from adding and subtracting
Why the two construction methods can disagree. Method of moments uses a few summaries of the sample; maximum likelihood uses the whole assumed density. When the density carries information the moments do not, the likelihood estimator is typically more efficient. When the model is wrong, the likelihood estimator is maximising the wrong function, while a moment estimator may still target something interpretable. A robustness difference, not a precision one.
A boundary case. For a uniform family on
Intuition
Two ways to be wrong, and why they add
Picture the estimator's distribution over repeated samples as a cloud of possible answers, with the true
Neither alone settles anything. A cloud centred exactly on
Mean squared error prices them together. Squaring the bias puts it in the same units as the variance, so the two add, and the sum is the expected squared distance from the truth. That single number lets two estimators be compared without deciding in advance which kind of wrongness matters more.
Why insisting on unbiasedness costs something. Restricting attention to unbiased estimators is a constraint on the search. Constraints do not improve an optimum; they can only leave it unchanged or make it worse. So the best unbiased estimator is, in mean squared error, never better than the best estimator overall. In the uniform example later the gap is a factor of
Consistency is about a different thing entirely. It describes what happens as the sample grows: does the cloud tighten onto
The converse also occurs. Take the first observation and ignore the rest. It is unbiased for the mean at every sample size, and its variance is
Method of moments asks: for which
Example
Four estimators and what each property says about them
Each case fixes a model and an estimator, then asks the three questions separately.
The sample mean for a population mean. Unbiased at every
The sample variance with divisor
The sample variance with divisor
| 2 | |
| 5 | |
| 10 | |
| 50 | |
| 1000 |
The bias never reaches zero at any finite
The first observation, as an estimator of the mean.
---
The first satisfies everything. The second and third differ only in a divisor and land on opposite sides of unbiasedness while both being consistent. The fourth is unbiased and inconsistent. So among these four, unbiasedness and consistency appear in all four combinations that matter, which is what it means for the two properties to be logically independent rather than one implying the other.
Procedure
Deriving an estimator, and checking what it is worth
To derive a method-of-moments estimator.
- Express the population moment in the parameter. For one unknown, write
as a function of . - Set it equal to the sample moment
. - Solve for
. The solution is the estimator. - For
unknown parameters, use equations, adding and higher moments as needed.
To derive a maximum likelihood estimator.
- Write the likelihood
, treating the data as fixed andas the variable. Include the support: a density that is zero outside a range makes zero for parameter values inconsistent with the data. - Take logarithms. Products become sums and the maximiser is unchanged, since
is increasing. - Differentiate and solve
. - Confirm it is a maximum by the sign of
, or by evaluating on either side. - Check the boundary. If the support depends on
, the maximum may lie at an endpoint where the derivative does not vanish. Skipping this step loses the answer entirely for uniform and similar families.
To evaluate an estimator.
- Compute
and subtract for the bias. - Compute
. - Assemble
, expressed in units ofso competitors are comparable. - Check consistency by asking whether both bias and variance go to zero as
grows. Either one failing to vanish defeats it.
To compare two estimators. Put both mean squared errors over the same
If an estimator is biased and otherwise attractive, try debiasing it. When
Checks. Confirm the estimate lies in the parameter space. A variance estimate must be non-negative, a probability inside
Worked example
The unbiased estimator that loses
The model.
Sample mean
Step 1: method of moments. For this family
Step 2: maximum likelihood. The density is
Above
Note what did not happen. There is no score equation here:
Step 3: bias.
For the maximum,
On this sample the contrast is visible:
Step 4: mean squared error. In units of
| Estimator | bias² | variance | MSE |
|---|---|---|---|
The biased estimator is better by a factor of
Step 5: taking the best of both. Since
better than both. On this sample it gives
So the ranking is
---
A second case, where the two methods agree. Twelve Poisson counts:
The log-likelihood is
The second derivative
Here the method of moments gives the same answer, because
Contrast
Pairs that differ in one respect
Unbiased against minimum mean squared error, on the uniform family.
| Estimator | bias² | variance | MSE (units of |
|---|---|---|---|
The first and third are unbiased and their mean squared errors differ by a factor of
Divisor
One change to a denominator moves the estimator across the unbiasedness line. Both are consistent, both converge to the same thing, and the one that is biased is the maximum likelihood estimator. The choice is conventional rather than forced, and the convention exists because unbiasedness is easy to state, not because it is the better criterion here.
Method of moments against maximum likelihood.
| moments | likelihood | |
|---|---|---|
| uses | a few sample summaries | the whole assumed density |
| closed form | usually | sometimes |
| efficiency when the model is right | often lower | usually higher |
| behaviour when the model is wrong | may still target something meaningful | maximises the wrong function |
For the Poisson family they coincide exactly, both giving
Unbiased for
Consistency against unbiasedness, as claims.
The first says the estimator eventually concentrates on the truth and promises nothing about the data in hand. The second says the estimator is centred correctly for the data in hand and promises nothing about improvement. A study with
Warning
Derivations that end one step early
Solving the score equation is not the definition of maximum likelihood. The definition is the maximiser of the likelihood. Differentiating is the usual route to it and fails whenever the support depends on the parameter. For the uniform family on
A stationary point is not automatically a maximum. The score equation locates where the slope is zero, which includes minima and inflection points. Confirm with the sign of
Unbiasedness does not survive a nonlinear transformation.
A property proved at one parameter value is not the property. Unbiasedness quantifies over every
---
Two errors about what the properties describe.
They are not checkable against the sample in hand. Bias, variance and consistency are features of a distribution over samples that were not drawn. No computation on one dataset verifies them; they are derived from the model or they are not known. An estimate close to a value believed to be true is not evidence of unbiasedness, and a bad estimate is not evidence against it.
A consistency guarantee says nothing about
---
Finally, check that the estimate is possible. A variance estimate must be non-negative, a probability must lie in
Application
Where the choice of estimator is the decision
Estimating a population total from a survey. The design supplies unequal selection probabilities, and the standard estimator weights each response by the inverse of its probability. It is unbiased by construction, and when a few units carry very large weights its variance is severe. Survey practice therefore trims or smooths extreme weights, knowingly introducing bias to reduce mean squared error. The trade of this unit, made as routine methodology.
Variance components in quality control. Maximum likelihood estimates of variance components are biased downward, which matters when the estimate feeds a tolerance. Restricted maximum likelihood exists to correct that bias, and the choice between them is the choice between the more efficient estimator and the better centred one.
Shrinkage in small-area estimation. Estimating unemployment for a district with few sampled households, the direct estimate is unbiased and unstable. Shrinking it toward a regional average introduces bias deliberately and reduces mean squared error, often substantially. Published figures are shrunk, and the documentation says so, because an unbiased and wildly variable district estimate is worse for every use it has.
Estimating a maximum from observed data. Reconstructing a total from serial numbers, or a peak load from observed peaks, is the uniform-endpoint problem. The sample maximum understates by construction, and the correction factor
Machine learning regularisation. Ridge and lasso shrink coefficients toward zero, making them biased for the underlying parameters while reducing prediction variance. The justification is the decomposition of this unit applied to prediction error rather than to a parameter estimate, which is why the bias-variance argument recurs there in the same algebraic form.
---
The common decision. In each case an unbiased estimator was available and something else was preferred, for a reason that was stated. That is what the criteria are for: not to establish that unbiasedness is undesirable, but to make the trade explicit so it can be defended. A method chosen because it is unbiased, with no comparison made, has not been chosen at all.