Week 12 — Information bounds, efficiency, and asymptotic normality
Where this week starts
Week 11 closed with a comparison. Rao-Blackwell improves any unbiased estimator by conditioning on a sufficient statistic, and when that statistic is complete, Lehmann-Scheffe promotes the improved version to uniformly minimum variance among unbiased estimators. That is a genuine optimality claim, but a relative one: it says your estimator beats the others in a set, not how good the best member of that set is. It gives you nothing to say when someone reports a variance you find implausible.
This week supplies the missing absolute standard. Before any estimator exists — before you have data, in fact — the model itself pins down a number below which the variance of an unbiased estimator cannot go. That number is built from the Fisher information of Week 7, where it measured the sharpness of a log-likelihood peak and supplied a standard error by convention. Here it becomes the denominator of a floor, , and that convention finally gets its theorem.
Two questions organize the page, each with an honest complication. Is there a floor? Yes, the Cramer-Rao bound, easy to prove once you notice that the score is mean-zero and has a fixed covariance with every unbiased estimator. But it speaks about unbiased estimators in regular models, and both qualifications do real work. Does maximum likelihood reach the floor? Asymptotically yes: has a normal limit whose variance is the reciprocal of the information. At finite , though, “asymptotically” is a promise about a limit rather than a description of your estimate, and this page makes that difference numerical rather than rhetorical.
Then the uniform, one last time. Since Week 3 this course has kept a family whose support moves with the parameter, and since Week 7 you have known the score identities fail for it. Here the failure pays off spectacularly: the bound cannot be written down at all, and the natural estimator of the endpoint has standard deviation of order rather than — doing what a careless reading of the theorem calls impossible. By Thursday you should be able to decide before computing whether a floor applies, write it correctly when it does, and say how far a moderate sample sits from the limit theorem you quote.
Why this matters downstream
Here is the concrete stake. A laboratory reports an unbiased estimate of a sensor’s full-scale voltage with a standard error of volts from twelve readings. A reviewer objects: the model has information per observation, so no unbiased estimator can have standard deviation below volts, and the reported figure is impossible. But the reviewer has computed an information for a model whose support moves with the parameter, where the derivative-under-the-integral step behind the bound is illegal, and used it to reject the best unbiased estimator available. Both are confident, one is wrong, and telling them apart takes the conditions, not just the inequality.
The structural stake is larger. Week 13’s Wald interval is this week’s limit theorem with a normal quantile attached, so every caveat here becomes a coverage failure there. Sample-size planning uses to decide how many observations buy a target precision, and the standard errors printed by regression and generalized-linear-model software are the multiparameter version of the same formula with the information matrix inverted. The posterior has a matching limit, which is why Bayesian and frequentist intervals nearly coincide in large regular samples.
What you will be able to do
- Compute for a one-parameter family by both routes, and write the Cramer-Rao floor for and for a differentiable function .
- Check the regularity conditions before quoting the bound, and say which one fails when it does.
- Decide whether a given unbiased estimator attains the floor, using the equality condition rather than by comparing numbers.
- Report the efficiency of an estimator and the relative efficiency of two, and translate an efficiency of into a statement about sample size.
- Sketch the Taylor-expansion proof of and name which theorem does each step.
- Attach a standard error to a maximum likelihood estimate from observed information, and quantify how far a moderate sample sits from the limit.
Words worth owning
| Term | What it means in this course |
|---|---|
| Cramer-Rao bound | The inequality for any unbiased of , valid only under the regularity conditions of Week 7. |
| Information floor | The right-hand side of that inequality, read as a property of the model at a particular . It exists before any estimator does. |
| Efficiency | The floor divided by an estimator’s variance, a number in . An efficiency of means the estimator does the work of observations. |
| Relative efficiency | The ratio of two estimators’ variances, always stated in the order that makes the better estimator’s efficiency exceed one. |
| Attainment | Equality in the bound. It happens only when the score is a multiple of the centred estimator, which forces a one-parameter exponential family. |
| Asymptotic variance | The variance of the limit of , divided by when quoted on the estimator’s own scale. |
| Asymptotically efficient | Having asymptotic variance equal to : reaching the floor in the limit, whatever happens at finite . |
| Superefficiency | Beating asymptotically at isolated parameter values, at the cost of terrible behaviour nearby. It is why “the MLE is optimal” needs care. |
The Cramer-Rao floor on the variance of an unbiased estimator
Let be independent draws from with in an open interval , and write for the joint density of the sample. Let be any statistic with finite variance and for every , with differentiable. The word every carries the argument: unbiasedness is not a fact at one parameter value but an identity in , and an identity can be differentiated.
Assume the Week 7 regularity conditions: the support does not depend on ; is differentiable in on that support; differentiation may be moved inside the integral both for and for ; and . Under these,
A covariance inequality, and where the work sits
The proof is three lines and one idea, and the idea is that the score has exactly the same covariance with every unbiased estimator. Write for the sample score. From Week 7, and , both under the same conditions. Now compute the covariance between the estimator and the score:
The first equality uses ; the second is the definition of the score; the third is cancellation; the fourth is the interchange, and is the only step that can fail; the last is unbiasedness differentiated. Then Cauchy-Schwarz, in its covariance form , gives
and dividing by , which is strictly positive by assumption, finishes it. Setting recovers the familiar form .
Read the structure rather than the algebra. Every unbiased estimator of has the same covariance with the score, namely , however clever or crude it is. The correlation is not the same: it is , which changes from estimator to estimator, and that is what makes the argument bite: a fixed covariance against a free variance forces the variance up, since shrinking shrinks the covariance with it. Drop unbiasedness and is some other function whose derivative you do not control, so the constraint evaporates.
Audit habit — the floor is a property of the model, not of your data. The right-hand side involves , and ; no observation appears in it. You can compute it before collecting anything, which makes it a design tool — and a floor evaluated at rather than the unknown is itself an estimate, with its own uncertainty.
Efficiency, relative efficiency, and when the floor is reached
Define the efficiency of an unbiased as the floor divided by its variance, a number in . Efficiency has a concrete reading: since variance falls like , a half-efficient estimator needs roughly twice the sample to match an efficient one, throwing away half your budget. The relative efficiency of against , both unbiased, is ; values above one favour . Both depend on , so an estimator can be nearly efficient in one region and poor in another.
When is the bound attained? Cauchy-Schwarz is an equality precisely when the two variables are proportional, so equality here requires
for some function free of the data. This is a strong demand: must be affine in the score, with slope and intercept free of even though the score is not. Solving that constraint back through and integrating in produces a density of the form . Attainment therefore happens only in a one-parameter exponential family, and only for the one that the natural sufficient statistic is unbiased for. Week 3’s exponential-family structure is doing the work again.
The figure makes the geometry concrete for the exponential model worked below, in units of so the picture holds at every true rate. No unbiased estimator enters the shaded region beneath the green floor. The blue curve is the exact variance of the unbiased rate estimator , always above the floor and closing on it. The orange curve is the exact mean squared error of , and it sits higher for two reasons rather than one: rescaling by inflates the variance by and introduces the bias . Do not credit the whole excess to the bias. At , , and : of the between the curves, squared bias supplies about a third and inflated variance the rest; at the bias share is only a quarter. Rescaling by removes both defects at once. The floor constrains neither part of the orange curve — a biased estimator is outside the theorem’s scope, and drawing one here is a deliberate reminder of that.
For a case where efficiency is catastrophic rather than merely imperfect, take the exponential model with mean and estimate by rescaling the sample minimum. Since , the minimum is itself exponential with rate , so and . The floor is , so this estimator has efficiency — at , two and a half percent. Its variance is that of a single observation: the rescaled minimum is worth one draw no matter how many you collect.
Before leaving the regular case, look at what a floor depends on. The left panel plots against the rate: information is not a number attached to a family but a function on the parameter space, falling steeply here as the rate grows. The right panel turns that into the smallest standard deviation an unbiased estimator of could have at , namely , rising linearly from to and passing through at . The readings are not in tension: a large rate is harder to pin down in absolute units, while in relative terms they cancel, the floor over being at every rate — a cancellation special to this scale family, and the sort of check worth running whenever a precision claim looks surprising.
Where the floor does not exist at all
Now the standing counterexample. Let for , with . On the support , so and . Three routes that agree in a regular model disagree here: the mean square of the score gives ; the variance of the score gives , a constant having no variance; and minus the expected curvature gives , a negative number where a variance should be. The three coincide only because the score is centred, and the score is centred only because the interchange is licensed — and it is not licensed here.
So the situation is not that the bound is loose, nor that it is hard to attain: there is no bound. Substituting — the only positive candidate of the three, and so the one a careless calculation reaches for — produces the number , which is a floor on nothing. The second worked example puts the best unbiased estimator comfortably below it.
Asymptotic normality of the maximum likelihood estimator
The bound says what is possible; the second theorem says likelihood achieves it in the limit. Let be the true parameter value, interior to , with the model identifiable and the Week 7 conditions in force, and add one more: is three times differentiable in near with for some with . Let be a consistent root of the score equation. Then
Read the statement carefully. It concerns the standardized quantity , not ; the limit is a fixed distribution, not an approximation with an error term attached; and it presumes a consistent root has been selected, which in a multimodal likelihood is an assumption about your optimizer, not the model.
The Taylor expansion of the score, term by term
The proof is a single expansion of the score around the truth, worth carrying because each piece is a theorem you already have. Write . Since is a root, . Expanding around with a second-order remainder at some between and ,
Solve for the difference and multiply by , dividing numerator and denominator by :
Now take the three pieces in turn. The numerator is times a sum of independent terms with mean zero and variance — the Week 7 centring and information identities — so the central limit theorem of Week 5 gives . The first denominator term is an average of independent copies of , so the weak law sends it in probability to . The remainder is bounded by ; the first factor goes to zero in probability by consistency and the second converges to a finite mean, so the product vanishes. Slutsky’s theorem then divides a normal limit by a constant, giving , which is the claim.
Three observations are worth keeping. First, consistency was assumed, not proved; establishing it needs identifiability plus a uniform law of large numbers, and that is the part of the theory this sketch skips. Second, the third-derivative envelope is what makes the remainder negligible, a smoothness condition on the model rather than on the data. Third, Week 7’s failure modes reappear here: a boundary maximizer means is false, and the expansion never starts.
What a standard error from observed information really claims
Since in probability and is continuous, replacing by or by changes nothing in the limit, again by Slutsky. That licenses the formula Week 7 used on credit,
and it says the maximum likelihood estimator is asymptotically efficient: its asymptotic variance equals the Cramer-Rao floor. Practitioners often prefer the observed , because it responds to the sample in hand rather than to an average over samples you did not draw.
Check the special case — is asymptotic efficiency an optimality theorem? Not on its own. Take from and define when and otherwise. At any the threshold is eventually irrelevant and ; at it equals zero with probability tending to one, so the limit has variance zero, beating the floor. Hodges’ construction shows that pointwise asymptotic variance is too weak a criterion to crown a winner. The repairs — restricting to regular estimators, or comparing worst-case risk over shrinking neighbourhoods — are Mathematical Statistics II material, and the honest summary for now is that the MLE is efficient among estimators that behave stably in .
The figure shows how slowly “asymptotically” can arrive. Each curve is the exact density of , available in closed form because and the total is gamma. At the density is conspicuously skewed right and its mean sits at , more than half a standard deviation from the limit’s zero; at that displacement is , at it is . Because is built into the standardization, those numbers are the bias already divided by the asymptotic standard error — the same that step six of the worked example below reports, with no further division to perform. The bias itself, , dies like , an order faster than the standard error’s , so their ratio improves like : the bias does become negligible against the noise it sits in, but only at that same slow rate, which is why the curve is still visibly shifted. The theorem is not wrong; it describes a limit, and the distance to that limit at your is a separate quantity you can and should compute.
Worked example — the exponential rate, its mean, and a standard error at forty observations
Setting. The reliability laboratory of Week 7 extends its study: forty nominally identical components are run to failure and the lifetimes total hours, so hours. Model them as independent exponential with rate , density on . The support does not move with and is smooth in , so the regularity conditions hold and the bound is available.
Step one — the information, computed twice. From Week 7, and . The curvature route gives ; the score-variance route gives . They agree, which is the free check that regularity holds.
Step two — the floor for the rate. With and , any unbiased estimator of the rate has
Step three — the floor for the mean. The laboratory reports mean lifetime , so and the floor becomes
Reparameterizing the model directly in gives the same thing: , so , , and . Two routes, one floor.
Step four — who attains it. The sample mean is unbiased for with , which is the floor exactly. The equality condition confirms it rather than accidentally agreeing: , a constant multiple of the centred estimator, with . For the rate the picture is different. Attainment would force with and free of , whence ; no two constants make that equal for all . No unbiased estimator of the rate attains the floor.
Step five — numbers. The maximum likelihood estimates are hours and, by invariance, per hour. Observed information for the rate is , so
Step six — how far from the limit. Because is gamma with shape and rate , the reciprocal moments are exact: and, for the bias-corrected , . So the best unbiased rate estimator has efficiency ; its exact standard deviation is against the asymptotic , an understatement of percent; and the maximum likelihood estimator’s bias is , which is of a standard error. None of that is alarming at ; all of it is invisible if you quote only the limit theorem.
Step seven — confirm by simulation. The exact sampling distribution can be generated straight from the gamma total, so no data are needed to check the algebra.
n <- 40; lambda <- 1 / 17.5; sims <- 200000
tot <- rgamma(sims, shape = n, rate = lambda) # the sufficient total
lam_hat <- n / tot # the MLE
lam_tilde <- (n - 1) / tot # the unbiased version
mean(lam_hat) - lambda * n / (n - 1) # near zero: matches the exact bias
var(lam_tilde) - lambda^2 / (n - 2) # near zero: matches the exact variance
lambda^2 / n # the floor, 8.2e-05
var(lam_tilde) # about 5 percent above the floor
A simulation cannot prove a derivation, as Week 8 insisted, but a disagreement would be decisive, and agreement across a symbolic route, an exact moment formula and a Monte Carlo estimate is the strongest evidence short of a proof.
What this assumed. A constant hazard, complete observation of all forty failures, and independence. Censoring changes the score and the information; a hazard rising with age makes the exponential family wrong, and Week 14 explains what converges to then. The floor is a property of the model, so a wrong model gives a confidently computed floor for a quantity nobody wanted.
The same reasoning, transferred
Run the same steps on a Poisson model, where counts replace waiting times. Let be independent Poisson(), so . Then and : the curvature route gives and the score-variance route gives . The floor for unbiased estimators of is , and is unbiased with variance , so here the mean does attain the floor for the rate itself: has the required proportionality, with .
Now push to a function. Suppose the target is , the probability of a zero count. Then and the floor is . Week 11 produced the uniformly minimum variance unbiased estimator for this quantity by Lehmann-Scheffe, namely with . Since is Poisson() and , taking gives , so
the inequality holding because for . With and the floor is and the variance , an efficiency of ; as grows the ratio tends to one. What stayed the same: compute the information two ways, write the floor for the parameter and for the function, test attainment through the proportionality condition, quote an efficiency. What changed: the sufficient statistic is a count, the floor for is now attained exactly, and the best possible estimator of a nonlinear function of falls short at every finite while closing on it — the typical pattern, and why “UMVU” and “efficient” are different words.
Second worked example — the uniform ceiling, where no floor exists
Setting. A sensor’s output is modelled as uniform on , with the unknown full-scale value in volts. Twelve independent readings, sorted:
Here , the total is so , and .
Step one — what a careless calculation reports. Taking from the mean-square-of-the-score route gives the apparent floor , so an apparent smallest standard deviation of — the reviewer’s number from the opening of this page, about volts at .
Step two — the actual best unbiased estimator. From Week 3, has density on , with and , so
Rescaling to remove the bias, is unbiased with
By Week 11 this is the uniformly minimum variance unbiased estimator, being complete and sufficient here. Its standard deviation is ; numerically volts, with estimated standard deviation about volts.
Step three — compare. The best unbiased estimator has standard deviation times below the apparent floor. Nothing has gone wrong with Cramer-Rao. The support moves with , the derivative cannot be taken inside the integral, the score is not centred, and the covariance identity — the single step everything rested on — is false. The inequality was never in force, so nothing was violated.
Step four — the rate, not just the constant. What matters is not the factor but the exponent. A regular model gives variance of order ; here is of order , so quadrupling the sample divides the standard deviation by four rather than two. For contrast, the method-of-moments estimator is also unbiased, with variance , here — larger by at , and by a factor growing without bound. On this data it reports volts. A wrong rate loses to a right one eventually, and “eventually” arrives quickly.
Step five — the limit is not normal. For , using at ,
so converges in distribution to an Exp(1) variable. Differentiating the exact expression, the density of that rescaled quantity is on , decreasing from its maximum at zero for every : no interior mode at any sample size, so no rescaling produces a bell.
The left panel stacks the exact densities at , and on the limiting exponential; they are nearly indistinguishable by , and every one is one-sided and skewed. Convergence is fast, and it is to the wrong shape for a normal-based interval. The right panel plots the unbiased estimator’s standard deviation against the line a regular model would force: at the two are and , and the gap widens with every observation.
What this buys, and what it costs. You get an estimator that converges faster than anything a regular model permits. You lose the apparatus this week built: no information, no floor, no normal limit, no symmetric standard error, and so no Wald interval in Week 13. The interval you can build comes from the exact distribution of , available in closed form and more accurate than any approximation.
The misreading to avoid
“No estimator can beat the Cramer-Rao bound.” Students write this after seeing the proof, and the proof is correct, so the error is in the quantifier. The theorem constrains unbiased estimators in regular models, and there are two exits.
The first exit is bias. The bound says nothing about mean squared error, and an estimator willing to be biased can have smaller MSE than any unbiased one. Take with , and shrink it: with has . At and that is , below the floor of that binds every unbiased estimator. The gain is real but local: at the same estimator has MSE and is now worse. This is the Week 6 bias-variance trade-off with a sharper edge, and in three or more dimensions Stein’s phenomenon makes the shrinkage improvement uniform rather than local — a result Mathematical Statistics II treats properly.
The second exit is regularity, and the uniform above walked through it. The unbiased is not merely below the apparent floor, it is below it by a factor that grows with . When a variance claim looks impossible, check the support before checking the arithmetic.
“The maximum likelihood estimator is efficient.” Add “asymptotically” or the sentence is false. At in the exponential model the maximum likelihood estimator of the rate is biased upward by about a sixth of a standard error, and the best unbiased competitor has efficiency , not . The word describes the limit of a sequence of sampling distributions, not the estimate on your screen.
“Efficiency is a property of the estimator.” It is a property of the estimator, the parameterization, the parameter value, and the sample size together. In the exponential model attains the floor for the mean exactly, while nothing attains it for the rate: same experiment, same data, different verdict depending on which label you gave the unknown. What survives reparameterization is asymptotic efficiency, because the delta method carries both the estimator’s asymptotic variance and the floor through the same factor .
Practice on your own
These are for self-checking, not submission. Work them with a pencil before running anything.
- Bernoulli, and a floor that vanishes. For independent Bernoulli() observations, confirm and show attains the floor for . Then write the floor for and evaluate it at . Explain in words what a floor of zero is claiming, and whether any unbiased estimator of actually has zero variance there.
- Normal variance, with and without a known mean. With known, show and verify that has variance , attaining the floor. With unknown, has variance . Compute its efficiency and say in one sentence what the missing observation’s worth of precision was spent on.
- A counterexample hunt. For the shifted exponential on , identify the failing condition, find the maximum likelihood estimator, obtain the exact distribution of , and build an unbiased estimator. Compare its variance with the floor the careless calculation would report, and with the uniform’s behaviour above.
- A simulation to describe. Describe a study that draws exponential samples at each of and , forms the interval , and records how often it covers the true rate. State in advance which sample size you expect to undercover and in which direction the misses fall, then say what the figures predict about the shortfall.
- Audit a plausible argument. A colleague writes: “The maximum likelihood estimator is asymptotically efficient, so at it has the smallest mean squared error of any estimator of .” Identify every distinct error in that sentence, produce one concrete estimator that refutes it, and say which of the two exits above your counterexample uses.
Where to read more
- MIT OpenCourseWare 18.655, Mathematical Statistics — lecture notes on information bounds, efficiency and the asymptotics of maximum likelihood, with the regularity conditions stated formally.
- MIT OpenCourseWare 18.650, Statistics for Applications — the asymptotic normality lectures, with more worked models and a gentler pace.
- Penn State STAT 414, Introduction to Probability Theory — the distributional facts used here: gamma moments including reciprocal moments, the Poisson probability generating function, and the distribution of a sample maximum.
- The R Project — home of the software above;
rgamma,optimizeandnumericDerivare all in the base distribution. - Hogg, McKean and Craig treat the Rao-Cramer inequality and the asymptotics of maximum likelihood in consecutive chapters, in this page’s order. It is optional.
- Course pages: the resources index and the schedule.
Where this goes next
Week 13 turns these two theorems into intervals. The Wald interval is this week’s normal limit with a quantile attached, and it inherits every weakness the figures displayed: a skewed sampling distribution at moderate becomes asymmetric coverage error, and a boundary parameter becomes an interval running outside the parameter space. The likelihood-ratio interval, a horizontal cut through the log-likelihood rather than its curvature at one point, repairs much of that by respecting the shape of . Read Week 13 with this week’s floor in mind, and revisit Week 11 if the claim that is best unbiased for the uniform ceiling needs refreshing.
Two threads carry further. The Bayesian counterpart of asymptotic normality is that the posterior becomes approximately under the same conditions, which is why the Week 9 credible interval and the Week 4 confidence interval nearly coincide in large regular samples — and part company for the uniform, where neither the floor nor the normal limit survives. Week 14 asks what this means when the model is wrong: information from a misspecified family is the curvature of the wrong log-likelihood, standard errors from it are miscalibrated, and the sandwich correction exists because and stop agreeing once the model no longer contains the truth. The notes index has the full sequence.