Week 5 — Convergence, the laws of large numbers, and the delta method
Where this week starts
Week 4 handed you exact statements: under a normal model the sample mean is exactly normal, the scaled sample variance exactly chi-square, the two exactly independent, and the studentized mean exactly at every sample size, including . Those are facts about the normal family, not about statistics in general; they came from an orthogonal transformation no other family hands you.
Step one inch outside that family and the exact sampling distribution of even a mild statistic disappears. What is the exact density of for an exponential sample? Of for a Bernoulli sample? Of a ratio of two sample means? Occasionally you can grind one out; usually you cannot.
So this week asks: when the exact distribution is out of reach, what survives as grows, and how wrong is the approximation at the sample size you actually have? The first half needs a vocabulary of limits, together with the two theorems — continuous mapping and Slutsky — that build complicated limits from simple ones. The second half needs the delta method and an honest audit of what it costs at finite .
That second half matters most: “asymptotically normal” describes a limit no data set ever reaches. By Thursday you should be able to state each convergence mode precisely, assemble a limit from parts, run the delta method on a transformation of your choosing, and put a number on the error you accept.
Why this matters downstream
Here is the stake. A reliability laboratory measures time to failure on a strongly right-skewed population, collects thirty units, and reports a one-sided lower 95 percent bound for the mean from a normal approximation — a guaranteed minimum average life — believing the true mean falls below it five percent of the time. Track which tail that failure lives in: the bound overshoots the truth exactly when the sample mean runs high, that is when the standardized mean exceeds . With skewness near six at that tail carries about seven percent, so the guarantee fails nearly forty percent more often than advertised, while the matching upper bound fails through the opposite tail and almost never fails at all. A two-sided summary would have concealed both errors.
The machinery also carries the rest of the course. Every standard error you meet from Week 12 onward, and every one printed by a regression routine, is a delta-method calculation on an asymptotic variance; the asymptotic normality of the maximum likelihood estimator is a Taylor expansion of the score finished by Slutsky; the Wald interval of Week 13 is the delta method plus a normal quantile. Leave the vocabulary vague and all of that becomes recipe.
What you will be able to do
- State convergence in probability and in distribution precisely, and exhibit a sequence that converges in distribution but not in probability.
- Prove the weak law under a finite variance with Chebyshev’s inequality, and name the hypothesis that proof uses which the theorem does not need.
- Use continuous mapping and Slutsky to obtain the limiting distribution of a ratio, and explain why Slutsky needs one of its limits to be a constant.
- Derive the first-order delta-method variance for a differentiable , and recognise when forces a second-order argument with a non-normal limit.
- Build a variance-stabilizing transformation by solving , and carry an interval back to the original scale.
- Quantify how far a normal approximation sits from the exact result at a stated , and name the population feature that governs the gap.
Words worth owning
| Term | What it means in this course |
|---|---|
| Convergence in probability | The chance that misses its target by any fixed tolerance goes to zero. About closeness. |
| Convergence in distribution | The distribution functions converge at every continuity point of the limit. About shape only. |
| Weak law of large numbers | For an i.i.d. sample with a finite mean, the sample mean converges in probability to that mean. |
| Central limit theorem | For an i.i.d. sample with finite non-zero variance, the standardized sample mean converges in distribution to a standard normal. |
| Slutsky’s theorem | A limit in distribution combines with a limit in probability to a constant — by sum, product or quotient — and the constant comes along. |
| Delta method | A smooth transformation multiplies the asymptotic standard deviation by the absolute derivative, the variance by its square. |
| Asymptotic variance | The variance of the limiting normal law after scaling by . A property of the limit, not the limit of the variances. |
| Variance-stabilizing transformation | A transformation chosen so the asymptotic variance stops depending on the parameter. |
Modes of convergence and the tools that move between them
Throughout, are independent draws from a common distribution with mean and variance , both finite unless a passage says otherwise. Write and .
A sequence of random variables is not a sequence of numbers, so “converges” needs a definition. Two carry this course.
where and are the distribution functions of and . The first says the random variables get close to each other; the second says only that the pictures get close, and permits and to live on different probability spaces. That gap is why two definitions are needed.
Convergence in probability implies convergence in distribution, and the reverse fails. The counterexample is one line: let be standard normal and set for every . By symmetry has exactly the standard normal law, so and ; but and is a fixed positive number that never shrinks. The definitions ask different questions.
One partial converse holds and gets used constantly: if for a constant , then , since , both being continuity points of the degenerate limit. With a constant target there is no shape left to match.
Convergence in probability and the weak law
The weak law of large numbers says that if are i.i.d. with and mean , then . Assuming also a finite variance, the proof is one line of Chebyshev’s inequality, since and :
Notice what happened. The proof used a finite variance; the theorem needs only a finite mean. The conditions of a proof and of a theorem are different lists, and confusing them makes a result look more fragile than it is.
The finite mean, though, is genuinely needed. Let be standard Cauchy, density , with no mean at all. Its characteristic function is , so that of is : the average of Cauchy observations has exactly the distribution of one, for every . Averaging a thousand buys nothing.
Convergence in distribution and the central limit theorem
The central limit theorem, in the form used here, says that if are i.i.d. with mean and variance satisfying , then
Read that carefully, because it is routinely misquoted. The observations do not become normal; their distribution never changes. Nor does the sample mean become normal. The theorem is about after centring at and inflating by ; without that inflation collapses onto and the limit is degenerate. The factor holds the spread still while the centre stabilizes, and finding such a rate is a move you will repeat all term.
It is also an statement, silent about any particular . The figure below makes that concrete for an exponential population, where a sum of exponentials is gamma and the exact density of is available in closed form.
At the exact density is obviously not normal, and its left tail simply ends, because cannot be negative and so cannot fall below . By the curves are hard to separate by eye in the middle, yet the printed tail probabilities show the approximation still off where it matters: the region above , to which the normal gives probability , carries at and still carries at . Agreement in the middle arrives long before agreement in the tails, and inference lives in the tails.
Continuous mapping and Slutsky as the working tools
Two theorems turn these limits into a calculus. The continuous mapping theorem says that if and is continuous on a set carrying all of the probability of , then ; the same holds with throughout. Slutsky’s theorem says that if and for a constant , then
The word “constant” is load-bearing and students skip it. Convergence in distribution of two separate sequences tells you nothing about their sum: put and for standard normal, and both converge to while is identically zero. Marginal limits do not determine joint behaviour; Slutsky escapes this because a sequence converging to a constant has no room left to be dependent on anything.
The standard use is the studentized mean:
For the second factor, the identity gives . The weak law applied to gives , finite precisely because is; continuous mapping gives ; and . So , and , continuous at , gives . Slutsky assembles the pieces, with no normality assumption anywhere.
The figure separates the factors. On the left the density of is broad at and has collapsed to a spike at by — convergence in probability to a constant, drawn. On the right the ratio starts heavy-tailed and is already on the standard normal at . For a normal population that ratio is exactly at every , which is why the curve is fat-tailed rather than noisy; the asymptotic argument reaches the same limit while assuming far less.
The delta method and the price of a smooth transformation
Estimators rarely arrive in the form you want to report: you estimate a mean and want a log, a rate and want a half-life, a proportion and want the odds. The delta method settles the resulting question: given the approximate distribution of , what is that of ?
The first-order expansion and its variance
Suppose and is differentiable at with . Then
The proof is a Taylor expansion plus the two tools above. Differentiability at lets us write, for near ,
Substituting and multiplying by ,
Because converges in distribution and , Slutsky gives , and continuous mapping gives . The bracket converges in probability to the constant , the second factor in distribution to , and Slutsky multiplies them.
The picture is that argument without symbols. Over the window where lands, a few multiples of around , the curve is nearly its own tangent, and a straight line of slope stretches a spread by ; variances, being squared, pick up . The curvature the tangent discards is not nothing: it produces a bias of order , negligible beside a spread of order as grows, and not negligible at .
Two cautions. This is a statement about a limiting law, not about moments: a limit with variance does not imply is even finite, as the transfer example shows. And everything hinges on being differentiable at the true , with non-zero derivative there.
When the derivative vanishes, and the second-order version
If the first-order statement is still true and useless: it says , a point mass, meaning only that approaches faster than . Expand one more term. If is twice differentiable at with , then , so multiplying by rather than and applying continuous mapping to the square,
The rate improved from to and the limit stopped being normal. Take a Bernoulli sample and : here vanishes at , , and the sample proportion has there, so
Audit that before believing it. The limit sits on the negative half-line, and indeed always, so the sign is right. Its mean is , and the exact calculation gives , so exactly at every . Two independent checks agree, the pattern to imitate whenever a limit theorem produces something surprising.
Variance stabilizing as the practical payoff
Often depends on the parameter: for a proportion it is , for a Poisson mean , for an exponential mean . That is awkward, because the width of an approximate interval then varies with the unknown quantity. Write for the dependence and demand a that removes it:
Integrate the reciprocal of the asymptotic standard deviation: that is the whole recipe. For an exponential mean and the integral is ; for a Poisson mean it is ; for a proportion it is the arcsine transformation derived below. Three of the commonest transformations in applied statistics fall out of one integral.
Worked example — the logarithm of an exponential sample mean
Setting. A laboratory records times to failure , modelled as independent draws from for , with rate , mean and variance . Failure times span orders of magnitude, so the reported quantity is . Take .
Step one — the base limit. The exponential has finite non-zero variance, so
Step two — the transformation. With we have and , non-zero for every admissible , so
Step three — read it. The asymptotic variance is , free of , so the logarithm is variance-stabilizing for an exponential mean, exactly as the integral recipe predicted from . The approximate standard error of is whatever the truth, so at it is and the interval width can be written down before seeing data.
Step four — audit it. A claim this clean deserves a check, and this model supplies an exact one. Since is standard exponential, is gamma with shape and rate , so the law of does not involve at all. That makes an exact pivot, parameter-free at every and not merely in the limit, so the delta method’s promise here is no accident of the approximation.
The exact moments follow. With gamma of shape and rate , , whose mean is the digamma function of less and whose variance is the trigamma function of . At the mean is , the standard deviation against the delta-method value , and the skewness . So the approximation is right to about one percent in spread but misses a downward shift of a tenth of a standard error and a mild left skew — exactly what a first-order expansion is entitled to miss, a bias of order and a skewness of order .
The same reasoning, transferred
Keep the model, change the target. The rate is often what a reliability report wants, and the natural estimator is . The argument does not move: same base limit, same theorem, different . With and ,
What stayed the same: the base central limit theorem, the differentiability requirement, the squared-derivative rule. What changed: the variance of the limiting law is , no longer free of the parameter, so the approximate variance of at sample size is and this transformation stabilizes nothing. Keep those two quantities verbally apart, as the glossary above does: an asymptotic variance is a property of the limit and carries no , while the you would put in a report is the finite- approximation got by undoing the scaling. The estimator returns in Week 7 as the maximum likelihood estimator of , and in Week 12 as the inverse Fisher information.
Now the caution. The map blows up at zero, so the argument needs . Here with probability one and the limit is safe, but finite-sample behaviour still differs more than you might expect. Because is gamma with shape , and for . At the bias is , about four percent upward, and the exact variance exceeds by the factor . Take the square root before calling that a standard-error discrepancy. The exact standard deviation exceeds the asymptotic by , so the printed standard error is about eight percent too small, and the printed variance about fifteen percent, since . Neither is a rounding error, and one gap has now been quoted three ways. Whenever a discrepancy arrives as a percentage, say which scale it lives on and which way round the comparison runs.
Change the model and the caution turns fatal. Estimate from a normal sample whose mean is near zero: the limit still holds for any fixed , but puts positive density at zero, so has no finite mean and no finite variance at any . The asymptotic variance exists, the actual variance does not, and only the first is what the delta method ever claimed.
Second worked example — stabilizing the variance of a sample proportion
Setting. A survey of households records whether each has a particular utility connection. Model the responses as independent Bernoulli variables with success probability , let , and suppose respond yes, so . The base limit is
whose asymptotic variance depends on . That dependence is the problem to fix.
Step one — integrate for the stabilizing map. With , substitute for , so that and :
Step two — verify the derivative directly, since a substitution is worth re-deriving. With the chain rule gives , so and
with asymptotic standard error at every in the open unit interval.
Step three — see the stabilization in numbers. At the raw standard error runs from at to at , a factor of . On the arcsine scale the asymptotic value is throughout, and the exact standard deviation of , summed over the binomial distribution, is at , at , at and at — constant to within seven percent across a tenfold change in . Push closer to the boundary and the flatness gives way: at , where the expected count is only two and a half, the exact standard deviation is , twenty percent above the asymptotic value. Stabilization is a statement about a limit, and near the boundary that limit needs a larger .
Step four — carry an interval back. With we have , so an approximate 95 percent interval on the transformed scale is , and inverting through gives . The ordinary interval built directly on is . The widths are nearly equal, against , but the arcsine interval is shifted upward and asymmetric about : the transformation has noticed that is near the boundary, where the sampling distribution is right-skewed.
Step five — a competitor, and what it buys instead. The logit, , has and asymptotic variance : not constant, so it stabilizes nothing. What it does is map the unit interval onto the whole line, so an interval computed there can never be transported back outside . Here the logit is with standard error , giving and, back-transformed, . Both respect the boundary and only one flattens the variance.
They part company at the edge of the sample space, and it is worth being exact about how. If comes out exactly , the logit is with an infinite estimated standard error, so it returns nothing until you adjust the counts, typically by adding a half to each cell. The arcsine has no such trouble: and are finite, and the standard error never consulted at all. At and it gives on the transformed scale, which truncated to the range of runs from to and inverts to .
So the arcsine is the better behaved of the two here, in the narrow sense that it still returns an interval — not in the sense of returning a good one. An exact calculation from zero successes in one hundred trials puts the one-sided 95 percent upper limit at , three times as far out, so the arcsine interval is far too short and its coverage near zero is correspondingly poor. Stabilization is itself the reason: a standard error that refuses to depend on cannot register that is the outcome saying least about how small might be. Week 13 returns to this comparison properly.
A simulation check you can describe and run
Everything above is a claim about a limit, and the honest way to interrogate one is to look at a finite . Here is the idiom for the first worked example; it needs no packages and runs as it stands in the R Project environment.
set.seed(75063)
n <- 25
lam <- 3
reps <- 200000
xbar <- replicate(reps, mean(rexp(n, rate = lam)))
zdel <- sqrt(n) * (log(xbar) - log(1 / lam))
c(centre = mean(zdel),
spread = sd(zdel),
skewness = mean((zdel - mean(zdel))^3) / sd(zdel)^3)
The delta method predicts a centre of , a spread of and a skewness of ; the exact gamma calculation predicts , and . The simulation will side with the second, which is the point of running it. Change lam to any other positive number and the summaries will not move, confirming that is a genuine pivot. Change n to and the centre should move to about and the skewness to about , each shrinking at the advertised rate.
Notice the discipline: the simulation is not evidence that the derivation is right, only a check that derivation and sampling behaviour agree at one parameter value, up to Monte Carlo error of order . Week 8 makes that distinction its main subject.
The misreading to avoid
The sentence to dismantle is the one every applied course leaves behind: “the central limit theorem makes enough.” Repeated often enough, it acquires the feel of a theorem. No threshold on alone can be right: the claim never mentions the population, and the population is what sets the rate.
The leading correction to the normal approximation for a standardized mean is proportional to , where is the population skewness. The error is small when is large relative to , not when is large in the abstract: a symmetric population can be fine at , while a population with has , so has not yet bought one unit of the quantity that matters.
The figure makes this exact rather than heuristic, using two gamma populations, where a sum of observations is again gamma and the tail probabilities compute to full precision.
Read the numbers off. For the strongly skewed population at , the upper-tail region a normal approximation assigns probability actually carries , and the lower-tail region carries : wrong by a factor of a hundred on one side and nearly forty percent on the other, with both errors still visible at . Now the trap inside the trap. The errors have opposite signs, so the two-sided event assigned probability — falling more than standard errors out on either side — actually has probability at , which looks perfectly respectable. A two-sided check can pass while both halves are badly wrong.
A further lesson sits in the arithmetic. A gamma population with shape has skewness , and a sample of size totals to a gamma with shape , so the standardized sampling distribution depends on and the population only through , that is through . The exponential population at and the skewness- population at therefore have identical standardized sampling distributions. Make the rule of thumb “ large compared with ”, and treat any fixed number as the start of a question rather than the end of one.
Practice on your own
These are for your own checking, not for submission. Pencil first, computer second.
A limit that is not a limit. Let be uniform on and set for odd and for even . Decide whether converges in distribution, whether it converges in probability, and reconcile the verdicts with the implication between them.
Where Chebyshev is not enough. Construct an i.i.d. sequence with a finite mean and infinite variance for which the weak law still holds, and explain why the Chebyshev proof cannot reach it. Then modify it so the mean fails to exist, and say which conclusion is lost.
A second-order delta method of your own. For a Poisson sample with mean , take and find the limiting distribution of at , with the correct scaling. Confirm the sign of the limit against the shape of , then check its mean against an exact computation of .
A simulation to describe. Write out, in ten lines of R, a study estimating the true coverage of the ordinary 95 percent interval for a proportion at and , and of the arcsine interval on the same samples. Predict, before running it, which dips further below and why the boundary is responsible.
Audit a plausible argument. A colleague writes: “since and is continuous, ; therefore , so is asymptotically unbiased.” Identify the step that does not follow, name the extra condition that would rescue it, and give a counterexample using on a normal sample.
Where to read more
- MIT OpenCourseWare 18.655, Mathematical Statistics — the large-sample lectures cover modes of convergence, Slutsky, and the delta method, with the measure theory this course sets aside.
- MIT OpenCourseWare 18.650, Statistics for Applications — the central limit theorem and delta-method lectures, with worked applications.
- Penn State STAT 414, Introduction to Probability Theory — useful if Chebyshev’s inequality or the gamma family need refreshing first.
- Hogg, McKean and Craig, Introduction to Mathematical Statistics, chapter 5, treats convergence, Slutsky, and the delta method in the order used here; an optional reference, and nothing here depends on owning it.
- The course pages: the notes index, the schedule, and the resources page.
Where this goes next
Week 6 starts building estimators and comparing them, and leans on this week immediately. Consistency is convergence in probability under another name, while the mean squared error decomposition needs the finite-sample variance, not the asymptotic one — keeping those apart is the discipline the transfer example rehearsed. The comparison of with a bias-corrected sample maximum for the uniform model is a case where asymptotic vocabulary does not settle the matter and finite-sample arithmetic does. Read Week 6 with that tension in mind, and look back at Week 4 for what an exact result looks like.
Two threads run forward. Computationally, every simulation from now on probes an asymptotic claim at finite , and predicting the summary before running the code is the habit worth keeping. On the Bayesian side, the posteriors of Week 9 have a large-sample story of their own: under the same regularity conditions that make the maximum likelihood estimator asymptotically normal, the posterior concentrates and becomes approximately normal around it, which is why a credible interval and a confidence interval so often nearly coincide at large and diverge at small . Week 13 asks what each actually claims.