Week 3 — Uniformly most powerful tests and monotone likelihood ratios

Where this week starts

Week 2 finished with a complete result and an incomplete procedure. The Neyman-Pearson lemma says exactly which test is most powerful when the null and the alternative are both single distributions: threshold the likelihood ratio, randomizing on the boundary if the model is discrete. But the test it hands you was built from one particular alternative θ1\theta_1, and changing θ1\theta_1 changes the critical value and in principle the shape of the rejection region. Real alternatives are rarely single points: a laboratory monitoring a defect rate cares about every rate above the standard.

So this week asks a sharper question: when does one test stay most powerful against every alternative in a set at once? Such a test is called uniformly most powerful, and the word to hold on to is uniformly — the optimality runs over the whole alternative set, and holds within a stated class of competitors, never in the abstract. The property that delivers it is monotone likelihood ratio, the theorem that converts the property into a test is the Karlin-Rubin theorem, and the home of both is the one-parameter exponential family from Week 0.

Two things should feel different by the end. First, “which test is best” is not a well-formed question until you have said best against what and best among which competitors. Second, the two-sided problem should feel like a genuine impossibility rather than an oversight: no cleverer statistic produces a uniformly most powerful test of a normal mean against a two-sided alternative, and the repair narrows the competitors by demanding unbiasedness instead.

Why this matters beyond the theorem

Here is the claim someone gets wrong without this material. An analyst calibrating an instrument runs the one-sided test because a colleague called it the most powerful test available, sees a sample mean well below the reference value, and reports a discrepancy. In the setting worked below that test has power 0.0001340.000134 at a true mean four units below the reference: it is nearly blind in that direction, and the blindness is not a defect but what it paid for optimality on the other side. Worse, if the direction is chosen after seeing which way the data fell, the procedure’s actual size is 0.100.10 rather than 0.050.05, so the premise that made it optimal has quietly been withdrawn.

The second stake appears in every discrete model, where a test whose stated level is 0.050.05 may have size 0.0320.032 because no cutoff lands on the nominal figure. That gap is not free: the conservative test is beaten at every alternative by one that spends the whole allowance.

What you will be able to do

  • State what makes a test uniformly most powerful at level α\alpha, naming both the alternative set the optimality runs over and the class of competitors it beats.
  • Verify or refute monotone likelihood ratio for a stated family in a stated statistic, and give the direction the ratio moves.
  • Write a one-parameter exponential family in natural form and read off the sufficient statistic and the sign that fixes that direction.
  • Apply the Karlin-Rubin theorem to build the uniformly most powerful one-sided test, including the randomization exact size needs in a discrete model, and derive its power function.
  • Show that no uniformly most powerful test exists against a two-sided alternative by exhibiting two level-α\alpha tests neither of which dominates.
  • Define an unbiased test, diagnose unbiasedness from a power function, and state the two side conditions identifying the best unbiased test in a one-parameter exponential family.

Terms and notation worth fixing

Three of these are places where loose usage does real damage.

Symbol or term Meaning as used in this course
Θ0\Theta_0, Θ1\Theta_1 the null and alternative parameter sets; this week both are intervals of the real line
φ(x)\varphi(x) the test function, the probability of rejecting when the data are xx; a non-randomized test takes only the values 00 and 11
β(θ)=Eθφ(X)\beta(\theta) = E_\theta \varphi(X) the power function, defined on all of Θ\Theta; on Θ0\Theta_0 it is the probability of rejecting a true null
size and level the size is supθΘ0β(θ)\sup_{\theta \in \Theta_0} \beta(\theta), one number; a level is any α\alpha with size α\le \alpha, so a test of size 0.0320.032 is a level-0.050.05 test
T(X)T(X) the statistic the likelihood ratio is monotone in, named every time because the property is relative to it
monotone likelihood ratio in TT for every θ1<θ2\theta_1 < \theta_2, the ratio f(x;θ2)/f(x;θ1)f(x; \theta_2)/f(x; \theta_1) is a nondecreasing function of T(x)T(x)
uniformly most powerful β(θ)βψ(θ)\beta(\theta) \ge \beta_{\psi}(\theta) for all θΘ1\theta \in \Theta_1 and every competitor ψ\psi in the stated class
γ\gamma the probability of rejecting at the one boundary value of TT, which buys exact size in a discrete model

Monotone likelihood ratio as a family property

The Neyman-Pearson region for a simple pair is {x:f(x;θ1)/f(x;θ0)>k}\{x : f(x; \theta_1)/f(x; \theta_0) > k\}, so whether one test can serve every alternative comes down to a question about that ratio. Does the ordering it puts on the sample space depend on which alternative you picked? Monotone likelihood ratio is the statement that it does not.

The definition and the statistic it names

Let {f(x;θ):θΘ}\{f(x; \theta) : \theta \in \Theta\} be a family of densities or mass functions with Θ\Theta \subseteq \mathbb{R} an interval, and let TT be a real-valued statistic. The family has monotone likelihood ratio in TT if for every pair θ1<θ2\theta_1 < \theta_2 the ratio

Λθ1,θ2(x)=f(x;θ2)f(x;θ1) \Lambda_{\theta_1, \theta_2}(x) \;=\; \frac{f(x; \theta_2)}{f(x; \theta_1)}

is a nondecreasing function of T(x)T(x), on the set where at least one of the two densities is positive, reading the ratio as ++\infty where the denominator vanishes and the numerator does not.

Three features carry the week. It is a property of a pair, a family together with a named statistic, so “this family has monotone likelihood ratio” is an incomplete sentence. Direction is part of the claim: a family whose ratio is nonincreasing in TT has the property in T-T, which reflects every region that follows. And the requirement runs over all pairs, which is why it belongs to the family rather than to one testing problem. What it buys is immediate: if Λ\Lambda is nondecreasing in TT then {x:Λ(x)>k}\{x : \Lambda(x) > k\} is an upper set of TT, so some cc has {T>c}{Λ>k}{Tc}\{T > c\} \subseteq \{\Lambda > k\} \subseteq \{T \ge c\}, and only the cutoff, never the shape, depends on the alternative.

Why one-parameter exponential families have it

Write the family in natural form, f(x;θ)=h(x)exp{η(θ)T(x)A(η(θ))}f(x; \theta) = h(x) \exp\{\eta(\theta) T(x) - A(\eta(\theta))\}, with η\eta the natural parameter and AA the cumulant function. For θ1<θ2\theta_1 < \theta_2,

f(x;θ2)f(x;θ1)=exp{[η(θ2)η(θ1)]T(x)}exp{A(η(θ1))A(η(θ2))}, \frac{f(x; \theta_2)}{f(x; \theta_1)} \;=\; \exp\bigl\{ [\eta(\theta_2) - \eta(\theta_1)] \, T(x) \bigr\} \cdot \exp\bigl\{ A(\eta(\theta_1)) - A(\eta(\theta_2)) \bigr\} ,

and the second factor is free of xx. So the ratio is strictly increasing in T(x)T(x) exactly when η(θ2)>η(θ1)\eta(\theta_2) > \eta(\theta_1): a one-parameter exponential family whose natural parameter is strictly increasing in θ\theta has monotone likelihood ratio in its natural sufficient statistic, and one whose η\eta is strictly decreasing has it in T-T. The carrier h(x)h(x) never enters.

A five-column table of six families giving the natural parameter, the sufficient statistic, whether the likelihood ratio increases or decreases in it, and the one-sided rejection region, with the exponential rate row decreasing.

Standard families with their natural parameter, sufficient statistic, direction of the ratio, and the one-sided rejection region that follows.

The fourth row of the table is the warning to memorize. For exponential lifetimes with rate θ\theta the density is θeθx\theta e^{-\theta x}, so η(θ)=θ\eta(\theta) = -\theta decreases as the rate grows: the family has the property in iXi-\sum_i X_i, and the one-sided test for a large rate rejects when the total time on test is small. The last row points the other way. The uniform family on (0,θ)(0, \theta) is not an exponential family, its support moving with θ\theta, yet the ratio (θ1/θ2)n(\theta_1/\theta_2)^n below θ1\theta_1, jumping to ++\infty above it, is still nondecreasing in the sample maximum. Exponential structure is sufficient here and necessary for nothing.

A family that fails the property

Let XX be one Cauchy observation with location θ\theta, density f(x;θ)={π(1+(xθ)2)}1f(x; \theta) = \{\pi(1 + (x - \theta)^2)\}^{-1}. Against θ1=0\theta_1 = 0 and θ2=θ\theta_2 = \theta the ratio is

Λ(x)=1+x21+(xθ)2, \Lambda(x) \;=\; \frac{1 + x^2}{1 + (x - \theta)^2} ,

which tends to 11 as x±x \to \pm\infty and so cannot be monotone unless it is constant. Setting its derivative to zero gives x(xθ)=1x(x - \theta) = 1, so the maximizer is x=(θ+θ2+4)/2x = \bigl(\theta + \sqrt{\theta^2 + 4}\bigr)/2. At θ=2\theta = 2 that is 1+22.4141 + \sqrt{2} \approx 2.414, where the ratio attains (1+2)25.828\bigl(1 + \sqrt{2}\bigr)^2 \approx 5.828, while Λ(10)=101/651.554\Lambda(10) = 101/65 \approx 1.554: an observation at 1010, far out in the direction of the alternative, is weaker evidence for it than one at 2.42.4. Solving Λ(x)=5\Lambda(x) = 5 gives x25x+6=0x^2 - 5x + 6 = 0, so the most powerful region at that cutoff is the bounded interval (2,3)(2, 3), of null probability (arctan3arctan2)/π=0.0452(\arctan 3 - \arctan 2)/\pi = 0.0452.

That interval is the whole problem. Repeat the construction against θ2=4\theta_2 = 4 at the same size and the region is (3.20,6.16)(3.20, 6.16), a different set. Since a most powerful test at a given level is essentially unique, no single test is most powerful against both, so no uniformly most powerful test exists even for the one-sided Cauchy problem. Heavy-tailed location families generally fail the property, which is one reason Weeks 12 and 13 build procedures that do not need it.

Karlin-Rubin and the reach of one-sided optimality

Theorem (Karlin-Rubin). Let Θ\Theta \subseteq \mathbb{R} be an interval, let the family have monotone likelihood ratio in TT, and fix θ0\theta_0 in Θ\Theta. For testing H0:θθ0H_0: \theta \le \theta_0 against H1:θ>θ0H_1: \theta > \theta_0, define

φ(x)={1,T(x)>c,γ,T(x)=c,0,T(x)<c, \varphi(x) = \begin{cases} 1, & T(x) > c, \\ \gamma, & T(x) = c, \\ 0, & T(x) < c, \end{cases}

with cc and γ[0,1]\gamma \in [0, 1] chosen so that Eθ0φ(X)=αE_{\theta_0} \varphi(X) = \alpha. Then φ\varphi is uniformly most powerful at level α\alpha, and its power function β(θ)=Eθφ(X)\beta(\theta) = E_\theta \varphi(X) is nondecreasing.

Every hypothesis is used: the parameter must be one-dimensional and ordered, or “one-sided” means nothing, and the property must hold in the same TT for all pairs, since the argument quantifies over alternatives.

Why the power function is nondecreasing

This step is what makes the composite null free. Take θ<θ\theta' < \theta'' and consider the artificial problem of testing θ\theta' against θ\theta'' at level β(θ)=Eθφ\beta(\theta') = E_{\theta'}\varphi. By monotone likelihood ratio φ\varphi is the Neyman-Pearson test for that pair, hence most powerful at that level. Compare it with the constant test ψβ(θ)\psi \equiv \beta(\theta'), which ignores the data, has the same level, and has power exactly β(θ)\beta(\theta'). Most powerful means φ\varphi cannot do worse, so β(θ)β(θ)\beta(\theta'') \ge \beta(\theta'). Hence the supremum of β\beta over Θ0={θθ0}\Theta_0 = \{\theta \le \theta_0\} is attained at the boundary, and calibrating at the single value θ0\theta_0 makes the test level α\alpha over the whole composite null.

The proof in three moves

Move one, reduce to a simple pair. Fix any θ1>θ0\theta_1 > \theta_0. Neyman-Pearson says a most powerful level-α\alpha test of θ0\theta_0 against θ1\theta_1 rejects wherever Λ(x)=f(x;θ1)/f(x;θ0)>k\Lambda(x) = f(x; \theta_1)/f(x; \theta_0) > k, does not reject wherever Λ(x)<k\Lambda(x) < k, and may be assigned freely where Λ(x)=k\Lambda(x) = k so long as the size comes out right. Choose cc and γ\gamma to give size α\alpha at θ0\theta_0 and let kk be the value the ratio takes at T=cT = c. Monotone likelihood ratio then puts φ\varphi in exactly that form, since nondecreasing in TT means T>cT > c forces Λk\Lambda \ge k and T<cT < c forces Λk\Lambda \le k, so φ\varphi never rejects below the threshold and never withholds rejection above it.

Say why that covers the flat case, because the definition gives only nondecreasing. The ratio may sit at height exactly kk across a whole stretch of values of TT rather than meeting kk at one boundary value, and then {Λ>k}\{\Lambda > k\} is strictly smaller than {T>c}\{T > c\}: the sharp cut at cc rejects outright at some points of the stretch and not at others. The conclusion survives untouched, because the sufficiency inequality of Week 2 turns on the factor f1kf0f_1 - k f_0, which vanishes identically there, so the pointwise inequality holds however φ\varphi is assigned on the stretch. Where the ratio is strictly increasing in TT, as in every exponential family above, the stretch collapses to a single value and the usual picture returns.

Move two, notice what is missing. The pair (c,γ)(c, \gamma) was determined by α\alpha and the null distribution alone. The alternative entered only through the direction of the inequality, which the property fixed once for all pairs, so the same φ\varphi is most powerful against every θ1>θ0\theta_1 > \theta_0 at once.

Move three, enlarge the null. Every level-α\alpha test for the composite null θθ0\theta \le \theta_0 is in particular a level-α\alpha test for the simple null θ=θ0\theta = \theta_0, so the competitors to beat form a subset of the class already handled, and the previous subsection put φ\varphi itself in that subset. A rule best in a class and belonging to a subclass is best in the subclass.

Drop the property and move one fails: a most powerful region still exists for each alternative, but its shape depends on which one, as the Cauchy interval showed. Drop one-sidedness and move two fails, since the direction of the inequality is no longer fixed.

Where uniform optimality runs out

Suppose the family has monotone likelihood ratio in TT, and test H0:θ=θ0H_0: \theta = \theta_0 against H1:θθ0H_1: \theta \ne \theta_0. Against an alternative above θ0\theta_0, move one gives an upper-tail region; against one below, the same argument with the inequality reversed gives a lower-tail region. The necessity half of the Neyman-Pearson lemma pins a most powerful test down only off the tie set where the likelihood ratio equals its threshold, so quote it that way rather than as uniqueness up to probability zero: on a family whose ratio is merely nondecreasing in TT the tie set can carry positive probability. When the ratio is continuously distributed, as it is in the normal model of the second worked example, that tie set is null and the pinning is complete, and then a single test would have to agree almost everywhere with two tests that disagree on a set of positive probability. No uniformly most powerful level-α\alpha test exists, and the second worked example does this with numbers.

Be precise about what failed. Reverse the roles of null and alternative and optimality returns: in a one-parameter exponential family, testing H0:θθ1H_0: \theta \le \theta_1 or θθ2\theta \ge \theta_2 against θ1<θ<θ2\theta_1 < \theta < \theta_2 does have a uniformly most powerful test, rejecting when c1<T<c2c_1 < T < c_2. What breaks uniform optimality is a two-sided alternative, asking one region to point two ways at once.

Unbiased tests as the repair

A level-α\alpha test is unbiased when β(θ)α\beta(\theta) \le \alpha for every θΘ0\theta \in \Theta_0 and β(θ)α\beta(\theta) \ge \alpha for every θΘ1\theta \in \Theta_1: it is never less likely to reject when the null is false than when it is true. That is a coherence demand rather than an optimality criterion, and it is what the one-sided tests fail in a two-sided problem, since the test rejecting for large sample means has power far below α\alpha on the low side.

Two consequences follow when the power function is smooth, as it is in an exponential family, where β\beta may be differentiated under the integral sign in the interior of the natural parameter space. If θ0\theta_0 is a boundary point between Θ0\Theta_0 and Θ1\Theta_1 and β\beta is continuous, unbiasedness forces β(θ0)=α\beta(\theta_0) = \alpha, since β\beta is at most α\alpha from one side and at least α\alpha from the other; and if β\beta has an interior minimum at θ0\theta_0 then β(θ0)=0\beta'(\theta_0) = 0. Those two equalities identify the uniformly most powerful unbiased test in a one-parameter exponential family against a two-sided alternative: it rejects when Tc1T \le c_1 or Tc2T \ge c_2, with the constants fixed by

Eθ0φ(X)=αandEθ0[T(X)φ(X)]=αEθ0T(X), E_{\theta_0} \varphi(X) = \alpha \qquad \text{and} \qquad E_{\theta_0}\bigl[ T(X)\, \varphi(X) \bigr] = \alpha \, E_{\theta_0} T(X) ,

the second being β(θ0)=0\beta'(\theta_0) = 0 written out, since differentiating β\beta with respect to the natural parameter gives Eθ0[Tφ]A(η0)Eθ0[φ]E_{\theta_0}[T\varphi] - A'(\eta_0) E_{\theta_0}[\varphi]. For a normal mean with known variance, symmetry of the null distribution makes the familiar equal-tailed test satisfy both at once. Do not generalize that shortcut: for the normal variance with known mean, equal tails fail the second condition.

Worked example — the one-sided binomial test at level 0.05

The model and the question. A quality audit inspects n=20n = 20 items independently, each defective with probability θ\theta. The process standard is a defect rate of 0.20.2 and the audit exists to detect deterioration, so test H0:θ0.2H_0: \theta \le 0.2 against H1:θ>0.2H_1: \theta > 0.2 at level 0.050.05.

Step 1: exponential family form. Writing the Bernoulli mass function as f(x;θ)=(1θ)exp{xlog[θ/(1θ)]}f(x; \theta) = (1 - \theta)\exp\{x \log[\theta/(1 - \theta)]\} for x{0,1}x \in \{0, 1\} identifies the natural parameter η(θ)=log[θ/(1θ)]\eta(\theta) = \log[\theta/(1-\theta)], the log odds, strictly increasing on (0,1)(0, 1). The natural sufficient statistic is T=i=120XiT = \sum_{i=1}^{20} X_i, with TBinomial(20,θ)T \sim \text{Binomial}(20, \theta).

Step 2: verify the property in that statistic. For θ1<θ2\theta_1 < \theta_2 the ratio of the mass functions of TT is

f(t;θ2)f(t;θ1)=(1θ21θ1)20(θ2(1θ1)θ1(1θ2))t, \frac{f(t; \theta_2)}{f(t; \theta_1)} \;=\; \left( \frac{1 - \theta_2}{1 - \theta_1} \right)^{\!20} \left( \frac{\theta_2 (1 - \theta_1)}{\theta_1 (1 - \theta_2)} \right)^{\!t} ,

the binomial coefficient having cancelled. The bracket in the second factor is an odds ratio and exceeds 11 whenever θ2>θ1\theta_2 > \theta_1, so the ratio is strictly increasing in tt. With θ1=0.2\theta_1 = 0.2 and θ2=0.4\theta_2 = 0.4 it is (0.4×0.8)/(0.2×0.6)=8/32.667(0.4 \times 0.8)/(0.2 \times 0.6) = 8/3 \approx 2.667, so the log ratio is a straight line in tt of slope log(8/3)=0.981\log(8/3) = 0.981.

Three straight lines rising with the number of successes, one for each alternative, with the region from eight upward shaded; the lines have different slopes but all increase, so each names the same upper-tail rejection region.

Log likelihood ratio against the null value, for three alternatives, as a function of the number of defectives.

The figure is move one drawn: three alternatives give three lines with three cutoffs, and because all three increase, all three name an upper tail of TT.

Step 3: apply the theorem and find the cutoff. The parameter space (0,1)(0, 1) is an interval, θ0=0.2\theta_0 = 0.2 lies inside it, and the property holds in TT, so the uniformly most powerful level-0.050.05 test rejects for large TT with randomization at one boundary value. Under θ0=0.2\theta_0 = 0.2 the exact tails are P(T8)=0.03214P(T \ge 8) = 0.03214 and P(T7)=0.08669P(T \ge 7) = 0.08669, so no non-randomized upper-tail region has size 0.050.05. In the notation of the theorem take c=7c = 7: reject outright when T>7T > 7, and reject with probability γ\gamma when T=7T = 7, where

γ=0.05P0.2(T8)P0.2(T=7)=0.050.032140.05455=0.3274. \gamma \;=\; \frac{0.05 - P_{0.2}(T \ge 8)}{P_{0.2}(T = 7)} \;=\; \frac{0.05 - 0.03214}{0.05455} \;=\; 0.3274 .

Step 4: the power function. By construction β(θ)=Pθ(T8)+0.3274Pθ(T=7)\beta(\theta) = P_\theta(T \ge 8) + 0.3274 \, P_\theta(T = 7), which is 0.05000.0500 at θ=0.2\theta = 0.2 and 0.28150.2815, 0.63840.6384, 0.89260.8926 at θ=0.3\theta = 0.3, 0.40.4, 0.50.5. Dropping the randomization leaves a test of size 0.032140.03214 whose power at 0.40.4 is 0.58410.5841.

n <- 20; theta0 <- 0.2; alpha <- 0.05
cstar  <- qbinom(1 - alpha, n, theta0) + 1               # 8
p_tail <- 1 - pbinom(cstar - 1, n, theta0)               # 0.032143
gam    <- (alpha - p_tail) / dbinom(cstar - 1, n, theta0) # 0.327358
power <- function(th) 1 - pbinom(cstar - 1, n, th) + gam * dbinom(cstar - 1, n, th)
power(c(0.2, 0.3, 0.4, 0.5))       # 0.0500 0.2815 0.6384 0.8926
1 - pbinom(cstar - 1, n, 0.4)      # 0.5841, the conservative test

lr <- dbinom(0:20, n, 0.4) / dbinom(0:20, n, theta0)
all(diff(lr) > 0)                  # TRUE: the ratio increases in the count

Two rising power curves against the success probability, the randomized test above the non-randomized one at every value, with a dashed level line at 0.05 and the null set shaded to the left of 0.2.

Power of the randomized and non-randomized one-sided tests against the defect rate, with the level marked.

What this licenses and what it does not. Both curves rise, which is the monotonicity proved above, so each test has its largest null rejection probability at θ=0.2\theta = 0.2 and is genuinely level 0.050.05 over all of θ0.2\theta \le 0.2; at θ=0.1\theta = 0.1 the randomized test rejects with probability 0.001060.00106. That test is uniformly most powerful among level-0.050.05 tests of this one-sided pair, and the figure shows it above the conservative test everywhere. The claim dies as soon as the setup moves: it says nothing about an audit that would also act on an unexpectedly low rate. And the conservative test most laboratories would actually run is uniformly most powerful at level 0.0320.032, not at 0.050.05 — a weaker claim, bought by refusing to let a random number settle a borderline case.

The same reasoning, transferred

Now put n=12n = 12 components on a burn-in rack with independent exponential lifetimes of rate θ\theta per hour, and test H0:θ0.5H_0: \theta \le 0.5 against H1:θ>0.5H_1: \theta > 0.5 at level 0.050.05. The natural parameter η(θ)=θ\eta(\theta) = -\theta is strictly decreasing, so the family has the property in iXi-\sum_i X_i and Karlin-Rubin rejects for small total time on test. The exact null distribution is available because 2θiXiχ2422\theta \sum_i X_i \sim \chi^2_{24}, so with the chi-square 0.050.05 quantile at twenty-four degrees of freedom equal to 13.84813.848 the test rejects when iXi13.848\sum_i X_i \le 13.848 hours, that is when the mean lifetime is at most 1.1541.154 hours. Its power is β(θ)=P(χ24227.696θ)\beta(\theta) = P\bigl(\chi^2_{24} \le 27.696\,\theta\bigr), which returns 0.050.05 at θ=0.5\theta = 0.5 and 0.7270.727 at θ=1\theta = 1.

What stayed the same: an exponential family, a sufficient statistic, the property verified in that statistic, and a size attained at the boundary. What changed: the natural parameter decreases in θ\theta, so the region is a lower tail, and the statistic is continuous, so the cutoff sits exactly at level 0.050.05. Discreteness belongs to the binomial model, not to the theorem.

Second worked example — the two-sided normal mean with no best test

The model and the question. An instrument is checked against a reference material whose stated value is 100100. Twenty-five independent measurements are modelled as N(θ,σ2)N(\theta, \sigma^2) with σ=10\sigma = 10 known from calibration history, so XN(θ,4)\bar{X} \sim N(\theta, 4) with standard error 22. Test H0:θ=100H_0: \theta = 100 against H1:θ100H_1: \theta \ne 100 at level 0.050.05; miscalibration either way matters, which makes the alternative two-sided.

Step 1: two one-sided competitors. Let φ+\varphi_+ reject when X100+1.645×2=103.290\bar{X} \ge 100 + 1.645 \times 2 = 103.290 and φ\varphi_- reject when X96.710\bar{X} \le 96.710. Each has size exactly 0.050.05, so each is a level-0.050.05 test of this problem, and each is uniformly most powerful for its own one-sided alternative, the normal family with known variance having monotone likelihood ratio in X\bar{X}.

Step 2: compare their power. At θ=104\theta = 104,

β+(104)=1Φ(103.2901042)=1Φ(0.355)=0.6388,β(104)=Φ(96.7101042)=0.0001. \beta_+(104) = 1 - \Phi\!\left( \frac{103.290 - 104}{2} \right) = 1 - \Phi(-0.355) = 0.6388 , \qquad \beta_-(104) = \Phi\!\left( \frac{96.710 - 104}{2} \right) = 0.0001 .

By symmetry the two exchange at θ=96\theta = 96, where β(96)=0.6388\beta_-(96) = 0.6388 and β+(96)=0.000134\beta_+(96) = 0.000134.

Step 3: the equal-tailed test. It rejects when |X100|1.96×2=3.920|\bar{X} - 100| \ge 1.96 \times 2 = 3.920, that is when X96.080\bar{X} \le 96.080 or X103.920\bar{X} \ge 103.920. Its power at 104104 is Φ(3.960)+1Φ(0.040)=0.5160\Phi(-3.960) + 1 - \Phi(-0.040) = 0.5160, and the same at 9696.

Three power curves over the mean: one rising to the right, one rising to the left, and a symmetric two-sided curve between them, all equal to 0.05 at the null value of 100, so no curve lies above the others everywhere.

Power functions of two one-sided tests and the equal-tailed two-sided test of the same mean.

Step 4: no uniformly most powerful test exists. Suppose some level-0.050.05 test φ*\varphi^* were one. At θ=104\theta = 104 it would need β*(104)0.6388\beta^*(104) \ge 0.6388; but φ+\varphi_+ is the most powerful level-0.050.05 test of 100100 against 104104, and by the necessity half of the Neyman-Pearson lemma any test attaining that power agrees with φ+\varphi_+ almost everywhere off the tie set where the likelihood ratio equals its threshold. Here the ratio is a strictly increasing function of the continuously distributed X\bar{X}, so the tie set is the single point X=103.290\bar{X} = 103.290 and has probability zero, leaving φ*=φ+\varphi^* = \varphi_+ almost everywhere. The same argument at θ=96\theta = 96 forces φ*=φ\varphi^* = \varphi_- almost everywhere. The two differ on {X103.290}\{\bar{X} \ge 103.290\}, an event of null probability 0.050.05, so no such φ*\varphi^* exists. The figure draws the same fact: every curve is beaten somewhere.

Step 5: the repair and its price. Since β+(96)=0.000134\beta_+(96) = 0.000134 sits far below 0.050.05 at an alternative, φ+\varphi_+ is not unbiased here, and restricting to unbiased tests eliminates both one-sided tests at once. The equal-tailed test meets the two side conditions, its size being 0.050.05 and its power function having a stationary minimum at 100100 by symmetry, so it is uniformly most powerful unbiased. Its power of 0.51600.5160 at θ=104\theta = 104 against 0.63880.6388 is what the insurance cost.

se    <- 10 / sqrt(25)                       # 2
cut_u <- 100 + qnorm(0.95) * se              # 103.2897
cut_2 <- 100 + qnorm(0.975) * se             # 103.9199
b_up  <- function(th) 1 - pnorm(cut_u, th, se)
b_two <- function(th) pnorm(200 - cut_2, th, se) + 1 - pnorm(cut_2, th, se)
c(b_up(104), b_two(104))                     # 0.6388 0.5160
c(b_up(96),  b_two(96))                      # 0.0001 0.5160

Size 0.050.05 alone does not deliver unbiasedness. Split the tails unequally, rejecting when X100+2.3263×2=104.653\bar{X} \ge 100 + 2.3263 \times 2 = 104.653 or X1001.7507×2=96.499\bar{X} \le 100 - 1.7507 \times 2 = 96.499, and the size is still 0.01+0.04=0.050.01 + 0.04 = 0.05. That power function is least where the normal densities at the two cutoffs balance, at the midpoint θ=100.576\theta = 100.576, where it takes the value 0.04150.0415.

Two panels of power curves for two size 0.05 tests; the zoomed right panel shows the unequal-tailed curve dipping to 0.0415 just above the null while the equal-tailed curve has its minimum of 0.05 exactly at the null.

Two size 0.05 two-sided tests, one unbiased and one not, shown full scale and zoomed near the null value.

What this licenses and what it does not. At a true mean of 100.576100.576 the unequal-tailed test rejects less often than if the null were exactly true, which is bias, and is why the class had to be narrowed by a condition rather than by taste. What it does not license is calling the equal-tailed test best in any wider sense: it is best among unbiased tests, and beaten by a one-sided test on either side.

The misreading to avoid

The misreading sounds like this: “The one-sided test is the most powerful test at level 0.050.05, so I should use it and stop giving power away to the two-sided test.” Two things are wrong, and only the second is arithmetic. First, “most powerful” is not a property a test carries on its own; it is a relation to an alternative set and a class of competitors. The one-sided test above is uniformly most powerful against θ>100\theta > 100 among level-0.050.05 tests, and against θ<100\theta < 100 it has power 0.0001340.000134 at a mean of 9696, worse than ignoring the data and rejecting with probability 0.050.05. Second, the power it seems to give away is bought with a commitment made before the data. Choose the direction after seeing which way the sample mean fell and you are running the union of two size-0.050.05 regions, a procedure of size 0.100.10, so the premise that made it optimal is gone.

A second misreading travels with it: “this family has monotone likelihood ratio, so the test rejects for large values of the statistic.” The property relates a family to a named statistic, and its direction is part of the claim. The exponential rate row of the table figure is the trap: since η(θ)=θ\eta(\theta) = -\theta, the family has the property in minus the total time on test, and the test for a large rate rejects when that total is small.

The third is the reflex that “no uniformly most powerful test exists” means nothing is optimal, so any reasonable test will do. Optimality did not disappear; it changed classes. Narrow the competitors to unbiased tests and a best test returns, pinned down by the two side conditions; change the criterion to average or worst-case risk and best rules return in the forms Week 14 develops. Nonexistence of a uniform optimum says only that a partial ordering has no greatest element, the fact Week 1 met when two risk curves crossed.

Practice on your own

These are for self-checking as you read, not for submission. Work each with a pencil before touching R.

  1. Verify the property, then break it. For X1,,XnX_1, \dots, X_n independent N(θ,σ2)N(\theta, \sigma^2) with σ2\sigma^2 known, write the likelihood ratio at θ1<θ2\theta_1 < \theta_2, confirm it increases in X\bar{X}, and state the uniformly most powerful test of H0:θθ0H_0: \theta \le \theta_0. Then for N(θ,θ2)N(\theta, \theta^2) with θ>0\theta > 0 show the ratio involves both iXi\sum_i X_i and iXi2\sum_i X_i^2, in a combination that changes with the pair, so neither alone can serve as TT.
  2. A family outside the exponential class. For X1,,XnX_1, \dots, X_n uniform on (0,θ)(0, \theta), show the ratio is nondecreasing in X(n)X_{(n)}, derive the uniformly most powerful level-α\alpha test of H0:θθ0H_0: \theta \le \theta_0, and check that its cutoff is θ0(1α)1/n\theta_0 (1 - \alpha)^{1/n}. Say which Week 0 regularity conditions this never needed.
  3. The discrete cutoff again. For n=10n = 10 independent Poisson counts with mean θ\theta, test H0:θ2H_0: \theta \le 2 against H1:θ>2H_1: \theta > 2 at level 0.050.05: find the smallest non-randomized upper-tail region and its exact size, compute the randomization probability that restores exact size, and compare both power functions at θ=3\theta = 3.
  4. The Cauchy interval. Confirm the maximizer (θ+θ2+4)/2\bigl(\theta + \sqrt{\theta^2 + 4}\bigr)/2, that the region where the ratio exceeds 55 is (2,3)(2, 3) when θ=2\theta = 2, and that it has null probability 0.04520.0452. Then say why a bounded region that moves with the alternative rules out uniform optimality.
  5. A simulation that shows the crossing. Simulate the two-sided normal problem at n=25n = 25 and σ=10\sigma = 10 across true means from 9494 to 106106, estimate the power of φ+\varphi_+, φ\varphi_-, and the equal-tailed test at each, and plot the three curves with Monte Carlo standard errors.

Where to read more

Where this goes next

Week 4 turns every test on this page into an interval. The device is duality: a family of level-α\alpha acceptance regions, one for each candidate parameter value, becomes a 1α1 - \alpha confidence set by collecting the values whose region contains the observed data. The one-sided test built here inverts into a one-sided bound and the two-sided test into the familiar interval, so this week’s power question reappears as a question about what an interval is asked to say. The discreteness that forced a randomization probability of 0.32740.3274 reappears too, as the conservatism of an exact interval for a proportion.

Read Week 4 — Confidence-test duality, pivots, and exact procedures next, and return to the notes index for the full sequence. If move one felt like a quotation rather than a step you could reproduce, reread Week 2 — Tests as decision rules and the Neyman-Pearson lemma first: every optimality claim from here to Week 6 is built on it.