Keep the Q2.4 bug reproducible and show the correction beside it
Decision: the model-selection page reproduces the submitted Q2.4 samples by default, bug included, and offers a clearly labelled "Corrected" toggle next to them. Nothing in the 2023 output is silently replaced.
Scope: Assignment 1, Questions 2.4 and 2.5 (web/src/lib/ar/model-selection.ts, web/src/components/model-selection/single-sample.tsx)
Context
Question 2.4 asked for 100 observations from two autoregressive models, M1 = AR(5) and M2 = AR(2), so that Question 2.5 could compare three information criteria on one sample of each. The submitted code built the series from y, a name that still held the Hitters salary vector from Question 1, instead of from the series being generated:
y1[6:100] <- sapply(6:100, function(t) y[(t-1):(t-length(phi.m1)) %*% phi.m1 + rnorm(1)])
y2[3:100] <- sapply(3:100, function(t) y[(t-1):(t-length(phi.m2))] %*% phi.m2 + rnorm(1))
For M1 a noisy weighted sum of time indices became an index into the salaries. For M2 each value is a weighted sum of two neighbouring salaries plus noise. The Q2.5 plots therefore sit on a log-salary scale, which is why they needed ylim = c(11.5, 14). The Monte Carlo function used from Q2.6 onwards generated its own series correctly, so the 1,000-replicate tables are unaffected.
Decision
Reproduce the submitted generator exactly, as the default view, and put a "Corrected" toggle beside it that uses the same recursion the submission itself used in Q2.6. Both are labelled. The bug is described next to the chart, in coursework/README.md and in the README's "Known issues" list.
Options considered
- Fix it silently. Show only the corrected series and the criteria they produce.
- Show the submitted version only, with a footnote.
- Show both, submitted by default, with a toggle (chosen).
- Drop Q2.4 and Q2.5 from the site and start the page at the Monte Carlo study.
Why
The project's rule is that original coursework results stay faithful: upgrades add analysis around them and never change them quietly. Option 1 would have rewritten history and made the parity tests meaningless for this question. Option 2 is honest but leaves a reader with numbers that answer the wrong question. Option 4 hides the most instructive mistake in the submission. Option 3 keeps the 2023 numbers checkable (seed 10 reproduces the printed Q2.5 tables to five decimals, which model-selection.test.ts asserts) and shows what the question intended.
The bug is also a useful lesson about R sessions: state from one question leaked into the next because every chunk shared one global environment.
What happened
- With seed 10 the submitted samples reproduce the printed tables. The criteria choose orders 3, 3 and 2 for "M1" and 7, 3 and 3 for "M2", neither of which means anything, because the data are salaries.
- For 2 of the first 1,000 seeds the noisy M1 index falls below 1. R would then return a zero-length vector,
sapplywould return a list andIC()would stop with an error. The page says so instead of inventing a result. - With the corrected generator and seed 10, all three criteria choose p = 2 for both models. That is right for M2 and wrong for M1. Over the first 1,000 seeds, single corrected samples recover M1's true order only 6.1% (IC₁, Wilson 95% 4.8%–7.8%), 6.0% (IC₂, 4.7%–7.6%) and 1.2% (IC₃, 0.7%–2.1%) of the time, against 57.5% (54.4%–60.5%), 63.1% (60.1%–66.0%) and 77.5% (74.8%–80.0%) for M2. The correction does not make single-sample selection reliable. M1's trailing coefficients (0.108 and 0.087) are too small to detect at n = 100, which the Monte Carlo study and its Wilson intervals now show directly.
What I'd change
- Never reuse a generic name such as
yacross questions. Run each question in its own environment (local({ ... })) or clear the workspace between sections. - Add a sanity check to every simulator: the sample variance of a simulated AR series should be close to its theoretical value. A salary-scale series would have failed that check at once.
- Plot the raw series before computing anything on it. The 2023 report plotted only the criteria, which hid the scale.