Skip to content
CompStats PlaygroundMAST90083 · 2023 S2

Assignment 3 · Question 2 · bootstrap bias

Resample, estimate the bias, subtract it.

Two independent exponential samples have means and , and the target is . The plug-in estimator is biased upwards. The submission derived that bias by hand, proposed estimating it with the bootstrap, and subtracted it.

The original answer was pure algebra with no code. This page turns each step into a simulation you can rerun.

2.00
1.00
20
2,000

[Q2.4–2.5] theta_star <- replicate(B, (mean(sample(x, replace = TRUE)) - mean(sample(y, replace = TRUE)))^2)

One pair of samples, B bootstrap replicates

Resample each sample with replacement, recompute and watch the replicates accumulate. Their average minus estimates the bias , and subtracting it gives the bias-corrected .
X samplemean 2.22Y samplemean 0.650 … 7.3

  • = 1.000 (truth)
  • = 2.469
  • = 2.636
  • = 2.303

2,000 / 2,000 replicates

Plug-in estimate
2.4694
truth = 1.0000
Bootstrap bias estimate
0.1665
theory = 0.2500
Bias-corrected
2.3029

95% intervals for from these 2,000 replicates (seed 2023). In a simulation the truth is known, so you can see whether this one interval caught it.

Methodloweruppercovers θ₀ = 1.000?
Percentile2.5% and 97.5% quantiles of θ̂*0.6325.983yes
BCaz₀ = 0.029, a = 0.0208; quantiles at 3.4% and 98.2%0.6856.243yes

[Q2.3, Q2.6–2.7] E(theta_hat - theta0) and E(theta_1 - theta0) by simulation

Is the corrected estimator unbiased?

Repeat the whole procedure on many fresh pairs of samples (each with its own B = 200 bootstrap) and average. This checks the derived b₁ and the claim b₂ = 0.
Starting the worker…

A refinement of the submitted answer

The submission derived correctly and concluded . The bootstrap actually estimates with the plug-in variance (divisor n), so and . That is not zero, but it is one order of n smaller, which is the real gain from bias correction. At small n (try n = 5 with 5,000 repetitions) the simulation separates the two values. The chart below shows both rates across n.

sapply(c(5, 10, 20, 40, 80), function(n) ...)

Bias against sample size

b₁ decays like 1/n, b₂ like 1/n². Each point averages 1,000 repetitions with B = 100.
Run five sample sizes, about 30 million resampled draws in total, in a background worker.

[rigour] coverage of percentile and BCa intervals for θ₀, by simulation

Do the bootstrap intervals cover the truth?

The 2023 answer stopped at the bias. A natural next step is an interval for θ₀. Simulate many datasets from the model above, build a nominal 95% percentile and BCa interval from each dataset's own bootstrap, and count how often the interval contains the known θ₀.
μx = 2, μy = 1, n = 20 · θ₀ = 1.000500 datasets × B = 999 · seed 2026computed at build time

Coverage with Wilson 95% intervals

The dashed line is the nominal 95%. Coverage is a binomial proportion over the simulated datasets.

Percentile: 90.2% (87.3% to 92.5%); BCa: 85.6% (82.3% to 88.4%)
Intervalcoveragemiss lowmiss high
Percentile90.2%87.3%–92.5%6.6%3.2%
BCa85.6%82.3%–88.4%9.8%4.6%

“Miss low”: the whole interval lies below θ₀; “miss high”: above. Mean widths: percentile 4.12, BCa 4.43. Bias check from the same runs: mean b̂₁ = 0.2317 ± 0.0060, against an exact E(b̂₁) = 0.2375 and a true b₁ = 0.2500.

Reading the result

The percentile interval under-covers: its Wilson interval (87.3%–92.5%) lies below 95%. BCa under-covers: its Wilson interval (82.3%–88.4%) lies below 95%.

Paired on the same 500 datasets, only the percentile interval covers θ₀ in 23 and only BCa in 0 (exact McNemar p < 0.001); the coverage difference, percentile minus BCa, is +4.6 points (Newcombe paired 95% interval +2.8 to +6.7). BCa is not automatically the better interval here. It misses more often on both sides (BCa low 9.8%, high 4.6%; percentile low 6.6%, high 3.2%).

With 500 datasets, each coverage estimate has a Monte Carlo SE of about 1.0 points near 95%. Because both intervals are built on the same datasets, compare them with the paired test above rather than by eye. Set μx = μy to see both intervals fail at the boundary θ₀ = 0, where every replicate is above the truth.

Explain this simulationoptional · your own key

Add your own Anthropic or OpenAI key in AI settings to enable. Nothing is sent without one.