Assignment 3 · Question 2 · bootstrap bias
Resample, estimate the bias, subtract it.
Two independent exponential samples have means and , and the target is . The plug-in estimator is biased upwards. The submission derived that bias by hand, proposed estimating it with the bootstrap, and subtracted it.
The original answer was pure algebra with no code. This page turns each step into a simulation you can rerun.
[Q2.4–2.5] theta_star <- replicate(B, (mean(sample(x, replace = TRUE)) - mean(sample(y, replace = TRUE)))^2)
One pair of samples, B bootstrap replicates
- = 1.000 (truth)
- = 2.469
- = 2.636
- = 2.303
2,000 / 2,000 replicates
95% intervals for from these 2,000 replicates (seed 2023). In a simulation the truth is known, so you can see whether this one interval caught it.
| Method | lower | upper | width | covers θ₀ = 1.000? |
|---|---|---|---|---|
| Percentile2.5% and 97.5% quantiles of θ̂* | 0.632 | 5.983 | 5.351 | yes |
| BCaz₀ = 0.029, a = 0.0208; quantiles at 3.4% and 98.2% | 0.685 | 6.243 | 5.558 | yes |
[Q2.3, Q2.6–2.7] E(theta_hat - theta0) and E(theta_1 - theta0) by simulation
Is the corrected estimator unbiased?
A refinement of the submitted answer
The submission derived correctly and concluded . The bootstrap actually estimates with the plug-in variance (divisor n), so and . That is not zero, but it is one order of n smaller, which is the real gain from bias correction. At small n (try n = 5 with 5,000 repetitions) the simulation separates the two values. The chart below shows both rates across n.
sapply(c(5, 10, 20, 40, 80), function(n) ...)
Bias against sample size
[rigour] coverage of percentile and BCa intervals for θ₀, by simulation
Do the bootstrap intervals cover the truth?
Coverage with Wilson 95% intervals
The dashed line is the nominal 95%. Coverage is a binomial proportion over the simulated datasets.
| Interval | coverage | miss low | miss high | mean width |
|---|---|---|---|---|
| Percentile | 90.2%87.3%–92.5% | 6.6% | 3.2% | 4.12± 0.13 |
| BCa | 85.6%82.3%–88.4% | 9.8% | 4.6% | 4.43± 0.16 |
“Miss low”: the whole interval lies below θ₀; “miss high”: above. Width ± its Monte Carlo SE.Mean widths: percentile 4.12, BCa 4.43. Bias check from the same runs: mean b̂₁ = 0.2317 ± 0.0060, against an exact E(b̂₁) = 0.2375 and a true b₁ = 0.2500.
Reading the result
The percentile interval under-covers: its Wilson interval (87.3%–92.5%) lies below 95%. BCa under-covers: its Wilson interval (82.3%–88.4%) lies below 95%.
Paired on the same 500 datasets, only the percentile interval covers θ₀ in 23 and only BCa in 0 (exact McNemar p < 0.001); the coverage difference, percentile minus BCa, is +4.6 points (Newcombe paired 95% interval +2.8 to +6.7). BCa is not automatically the better interval here. It misses more often on both sides (BCa low 9.8%, high 4.6%; percentile low 6.6%, high 3.2%).
With 500 datasets, each coverage estimate has a Monte Carlo SE of about 1.0 points near 95%. Because both intervals are built on the same datasets, compare them with the paired test above rather than by eye. Set μx = μy to see both intervals fail at the boundary θ₀ = 0, where every replicate is above the truth.
Explain this simulationoptional · your own key
Add your own Anthropic or OpenAI key in AI settings to enable. Nothing is sent without one.