Skip to content
CompStats PlaygroundMAST90083 · 2023 S2

Methods · model card · AI use · decision records

How the numbers were made, and where they stop.

The 2023 coursework reported point estimates. This page sets out where the data came from, how each result is evaluated, what it assumes, and how much it would move under resampling. The decisions behind the revival are recorded too, including the ones whose numbers turned out weak.

Every document here is also in the repository's docs/ folder. The model card's metrics are generated by the site's own code and checked by the test suite.

4 decision recordsWilson, bootstrap and BCa helpers tested against SciPy and RAI: optional, your own key, logged

How the numbers on this site are produced, how they are evaluated, and where they stop being trustworthy. The 2023 results are reproduced unchanged; everything marked "rigour" was added in 2026 around them.

Data provenance

Data Source Size Used for
Hitters ISLR R package (James, Witten, Hastie and Tibshirani), 1986 season statistics and 1987 salaries in $000s 322 players; 59 with no salary removed, leaving 263; 19 predictors after model.matrix (three factors become dummies) Assignment 1, Q1: ridge, lasso and least squares
AR series Simulated: M1 with φ = (0.434, 0.217, 0.145, 0.108, 0.087), M2 with φ = (0.682, 0.346), N(0, 1) innovations, zero start, no burn-in n = 100 and n = 15; 1,000 replicates per setting Assignment 1, Q2: order selection
Gaussian classes Simulated: three classes with means (0, 0), (0, 3) and (3, 0), unit variance 300 training points (seed 50), 300 test points (seed 100) Assignment 3, Q1: SVMs
Exponential samples Simulated: X ~ Exp(mean μx), Y ~ Exp(mean μy), equal sizes n any μx, μy and n from the page controls Assignment 3, Q2: bootstrap bias

web/src/lib/data/hitters.json is exported from the ISLR package by scripts/export_artefacts.R. No other data is shipped; everything else is simulated in the browser or at build time with R's random number generator, so every seed reproduces exactly.

Method

  • Ridge and lasso: a TypeScript port of glmnet's covariance-update coordinate descent, including its standardisation and default λ path. λ is chosen by 10-fold cv.glmnet on a 131-player training split; test error is measured on the other 132 players.
  • Order selection: AR(p) for p = 1…10 fitted by least squares; three information criteria as the assignment defined them (see DR-002).
  • SVMs: a port of LIBSVM's SMO solver with e1071's defaults and tune()'s 10-fold cross-validation.
  • Bootstrap: the plug-in estimator θ̂ = (x̄ − ȳ)², its bootstrap bias estimate and the bias-corrected estimator, as derived in the submission.
  • Rigour helpers (2026): Wilson score intervals and Monte Carlo standard errors for simulated proportions, type-7 quantiles, percentile and BCa bootstrap intervals, paired bootstrap comparisons, and, for two rates measured on the same units, an exact McNemar test with Newcombe's paired interval for the difference, in web/src/lib/stats/. Each is unit-tested against values from SciPy (scripts/stats_reference.py) or R (base R and the contingencytables package).

Evaluation design

Question Design Sample size Uncertainty shown
Which of ridge, lasso and OLS predicts salaries best? One 131/132 split (seed 10); paired bootstrap over test players, the same resamples for every model 132 test players, B = 2,000 (seed 2026) 95% percentile intervals for each MSE, each paired difference and each MSE ratio; share of players where one model wins, with a Wilson interval
How much does λmin depend on the fold assignment? Same training split, folds redrawn with set.seed(s) for s = 1…200 200 fold draws Middle 95% of λmin and of test MSE; Wilson intervals on the share of draws that equal or fall below the 2023 value
How often does each criterion pick the true order? Monte Carlo: simulate, fit p = 1…10, count the arg-min; all three criteria scored on the same series 1,000 replicates per setting (seed 10 reproduces 2023) Wilson 95% intervals, Monte Carlo SE, the running estimate with its band; paired McNemar tests and Newcombe intervals between criteria
Is the corrected bootstrap estimator unbiased? Repeat the full bootstrap on fresh samples 500–5,000 repetitions, B = 200 ±2 Monte Carlo SE
Do bootstrap intervals for θ₀ cover the truth? Simulate datasets with known θ₀; build percentile and BCa intervals from each dataset's own bootstrap 500 datasets × B = 999 (seed 2026 by default) Wilson interval on coverage; Monte Carlo SE of mean width; paired McNemar test between the two methods on the same datasets
Are the tuned SVMs different? Same 300 test points for both models 300 Wilson intervals on error rates; exact McNemar test on the discordant points

All seeds are fixed and shown next to the result they produced.

Assumptions

  • The 132 held-out players are a fair sample of the population the model would be used on. They are not: the split is one random draw from one season.
  • Players are exchangeable, so resampling them with replacement approximates sampling variability. Teams and positions are ignored.
  • Simulated AR series start at zero with no burn-in, as in the submission, so early observations are not from the stationary distribution.
  • Exponential samples are independent and the two samples have equal size.
  • Monte Carlo replicates are independent draws from one long Mersenne-Twister stream.

Limitations

  • One split. The regression comparison holds the training split fixed. Changing the split seed on the regularisation page moves all three test MSEs, so the bootstrap intervals understate the total uncertainty of "which method is better".
  • Heavy tails. Salaries are skewed. The five largest OLS squared errors make up about 43% of its test error, so a handful of players drive every comparison and the intervals are wide.
  • No detectable differences. On the submitted split, every paired difference between ridge, lasso and OLS has a 95% interval that includes zero. The 2023 ranking by point estimate is within resampling noise.
  • Order selection at n = 15 is degenerate: with eight or more lags the fit has no residual degrees of freedom, and the counts depend on floating-point round-off.
  • BCa is not always better. In the default coverage study BCa covers θ₀ less often than the percentile interval, and the paired comparison on the same datasets (exact McNemar test) shows the gap is not Monte Carlo noise. Both intervals fail near the boundary θ₀ = 0.
  • Monte Carlo error. A rate from 1,000 replicates carries a standard error of up to 1.6 percentage points. That describes each rate on its own. The three criteria are scored on the same replicates, so their differences are judged with a paired test on the replicates where they disagree, not by whether two separate intervals overlap: overlapping intervals can hide a decisive difference (for M2 at n = 100, IC₁ and IC₂ disagree on 62 replicates and IC₂ is the one that is right on 61 of them).
  • AI explanations are optional, can be wrong, and are not evaluated against reference answers. See the AI use statement.

What I'd change

  • Replace the single split with repeated random splits or nested cross-validation, and report the spread of test MSE across splits next to the bootstrap intervals.
  • Model log salary, which tames the heavy tail and makes the comparison less dependent on a few stars.
  • Fit every AR order on a common estimation sample and add a burn-in to the simulator.
  • Use the 1-SE rule or repeated CV to choose λ, since λmin is sensitive to the fold draw.
  • Report intervals from the first draft. Several 2023 conclusions were drawn from point estimates that the intervals now show cannot support them.

docs/model-card.md

Model card

This card covers the three fitted models on the site: penalised regression of 1987 baseball salaries, three-class support vector machines on simulated data, and information-criterion order selection for AR models. All three are coursework models, rebuilt in 2026 for an interactive notebook. None is a product.

Intended use

  • Teaching and demonstration of computational statistics: shrinkage, cross-validation, model selection, margins and the bootstrap.
  • Showing how much a reported number would move under resampling, and how to report it with uncertainty.

Out of scope: any decision about a real person's pay, contract or worth; any use of these models on data other than the coursework data; any claim about baseball beyond 1986–87.

Training data provenance

  • Salary regression: ISLR Hitters (public R package): 263 Major League Baseball players with a recorded 1987 salary, 19 predictors from the 1986 season and careers to date. 131 players train the models (set.seed(10); sample(263, 131)); 132 are held out.
  • SVMs: 300 simulated points in three Gaussian classes (seed 50); a fresh 300-point test set (seed 100).
  • Order selection: series simulated from two known AR models (seed 10).

Evaluation

Every figure below is computed by the site's own code and checked against this file by the test suite.

Salary regression (A1 Q1): test MSE on the 132 held-out players

Model λ (CV) Test MSE 95% bootstrap interval
Ridge 300.8959 143,261 79,414–235,366
Lasso 4.468208 141,633 79,751–234,205
OLS none 145,023 86,146–232,357
Paired difference Δ MSE 95% bootstrap interval MSE ratio (95%)
Ridge − OLS −1,762 −23,361 to +21,493 0.988 (0.809–1.149)
Lasso − OLS −3,390 −16,249 to +10,374 0.977 (0.862–1.063)
Ridge − Lasso +1,629 −9,533 to +13,714 1.011 (0.924–1.108)

Paired percentile bootstrap over players, B = 2,000, seed 2026, split seed 10. The five largest OLS squared errors are 42.6% of its test error.

λmin across 200 CV fold seeds (same split): ridge middle 95% 188.5–436.5 (2023: 300.9, 40.5% of seeds lower); lasso middle 95% 4.06–23.85 (2023: 4.47, 6.0% of seeds lower, median 16.44). The lasso draws split into 2 separate modes on the λ grid; the 2023 value is in one that 14.5% of seeds pick (Wilson 95% 10.3%–20.0%).

Three-class SVM (A3 Q1): 300 fresh test points

Model Tuned parameters Test errors Error rate (Wilson 95%)
Linear cost 1 28 / 300 9.3% (6.5%–13.2%)
Radial cost 100, γ 2 31 / 300 10.3% (7.4%–14.3%)

Paired on the same points: only linear correct on 10, only radial correct on 7; exact McNemar p = 0.63. Accuracy difference (linear − radial) +1.0 points, Newcombe paired 95% interval −1.9 to +4.0.

AR order selection (A1 Q2): P(selects the true order), 1,000 replicates each

True model n IC₁ (AIC-like) IC₂ (AICc-like) IC₃ (BIC-like)
M1 · AR(5) 100 7.8% (6.3%–9.6%) 7.8% (6.3%–9.6%) 1.4% (0.8%–2.3%)
M2 · AR(2) 100 57.7% (54.6%–60.7%) 63.7% (60.7%–66.6%) 79.8% (77.2%–82.2%)
M1 · AR(5) 15 0.0% (0.0%–0.4%) 0.0% (0.0%–0.4%) 0.0% (0.0%–0.4%)
M2 · AR(2) 15 0.0% (0.0%–0.4%) 0.0% (0.0%–0.4%) 0.0% (0.0%–0.4%)

Wilson 95% intervals on the 2023 counts (seed 10), which the port reproduces exactly at n = 100.

Paired comparisons at n = 100 (the criteria are scored on the same 1,000 replicates):

True model Pair Only first right Only second right Difference, points (Newcombe 95%) Exact McNemar p
M1 · AR(5) IC₁ vs IC₂ 9 9 0.0 (−0.9 to +0.9) 1.00
M1 · AR(5) IC₂ vs IC₃ 66 2 +6.4 (+4.9 to +8.1) 1.6 × 10⁻¹⁷
M1 · AR(5) IC₁ vs IC₃ 67 3 +6.4 (+4.9 to +8.1) 9.7 × 10⁻¹⁷
M2 · AR(2) IC₁ vs IC₂ 1 61 −6.0 (−7.5 to −4.5) 2.7 × 10⁻¹⁷
M2 · AR(2) IC₂ vs IC₃ 73 234 −16.1 (−19.4 to −12.8) 8.0 × 10⁻²¹
M2 · AR(2) IC₁ vs IC₃ 64 285 −22.1 (−25.5 to −18.7) 2.1 × 10⁻³⁴

At n = 15 no criterion ever picks the true order, so there are no discordant replicates to compare.

Known failure modes

  • Test-set dependence. All salary comparisons use one held-out split of 132 players. A different split changes every test MSE, often by more than the differences between methods.
  • Heavy-tailed errors. A few high earners dominate squared error. Mean squared error on the dollar scale is a fragile yardstick for this data.
  • λ instability. The cross-validated λ moves with the fold assignment. For the lasso the fold draws split into two separate modes on the λ grid, because the CV curve has two competing minima. The 2023 λmin falls in the smaller, low-λ mode, so the 2023 lasso kept more predictors than a typical fold draw would.
  • Small trailing coefficients. For M1 the order-selection criteria almost never find the true order at n = 100, and none of them works at n = 15.
  • Simulated classes are easy. The SVM data are well separated Gaussians; error rates say nothing about harder or real-world classification.

Ethical considerations

  • The regression predicts what named people were paid. It reflects 1987 pay practices, including any bias in them, and it must not be used to judge anyone's value.
  • Player names appear only because they are part of the public dataset; nothing is inferred about individuals beyond the coursework.
  • Optional AI explanations are labelled, human-reviewed and logged; see the AI use statement.

docs/ai-use.md

AI use statement

This site uses generative AI for one optional feature: "Explain this simulation", a short plain-language reading of a result that is already on the page. It is informed by the transparency and human-oversight principles in the Australian Government's policy for the responsible use of AI in government (DTA), the EU AI Act's transparency provisions and the NIST AI Risk Management Framework. It is not certified against, and makes no claim of compliance with, any of them.

What the AI does

  • Reads a small JSON summary of the result in front of you (rates, intervals, sample sizes, seeds) and returns a two-to-four sentence summary, a few key points and a few caveats.
  • Is told to use only the numbers it was given and to say when the numbers cannot answer something.

What the AI never does

  • It never computes, changes or replaces a number on the site. Every statistic, chart and table is produced by tested code without AI.
  • It never runs without your action, and never runs without a key you supplied.
  • It never sees your key in the prompt, and the prompt contains nothing from your browser except the summary shown under "Prompt sent" on every explanation.

Data sent to the provider

  • The prompt: the facts summary and fixed instructions, which you can read in full under "Prompt sent" on every explanation and in the audit log.
  • The provider also receives the output schema and model settings (model, maximum output tokens and, for Claude Sonnet, the effort setting), plus standard request metadata: your IP address, browser user agent and this site's origin. It links the request to the account that owns the key.
  • The request goes from your browser directly to the provider you chose (api.anthropic.com or api.openai.com) with your own key. This site has no server and receives nothing. The provider's own terms and data retention apply to that request.
  • No personal data from this site is put in the prompt. The facts are summaries of public or simulated data.

Your key

  • Stored in sessionStorage by default, which the browser clears when the tab closes. If you tick "Remember on this device", it is stored in localStorage instead.
  • "Forget key" removes it from both. It is never logged, never written to the audit log and never committed anywhere.
  • The site's Content-Security-Policy restricts fetches and resource loads (images, styles, fonts, workers) to this site, plus the two provider APIs for the AI calls. It cannot stop a script already running on the page from navigating the page away, so it reduces, rather than eliminates, the risk of a key being sent elsewhere.

Transparency and human oversight

  • Every explanation is labelled AI-generated, with the model, latency and token usage the provider reported.
  • A grounding check lists any number in the explanation that cannot be traced to the facts that were sent. It catches invented numbers, not flawed reasoning.
  • You decide what happens next: Accept, Edit (your version is kept next to the original) or Reject.
  • Every call, including failed ones, is written to an audit log in your browser's IndexedDB: id, timestamp, feature, provider, model, request settings, input, output, latency, token usage, the grounding result and your decision, with every earlier decision kept in an append-only history. When a reply fails validation and is not shown, its raw text is still logged (with anything key-like redacted, truncated to 8,000 characters). View the log at /ai-log and export it as JSON or CSV.

Models

  • Anthropic: Claude Haiku 4.5 by default (claude-haiku-4-5), with Claude Sonnet 5.5 (claude-sonnet-5-5) as the stronger option.
  • OpenAI: any chat model that supports JSON-schema output; the id is editable and defaults to gpt-5-mini.

Open this browser's AI audit log

docs/decisions/

Decision records

Each record states the decision first, then the context, the options, why, what actually happened (weak numbers included) and what I would do differently. Past records are never edited; a later record supersedes them.

DR-001accepted2026-10-06

Keep the Q2.4 bug reproducible and show the correction beside it

Decision: the model-selection page reproduces the submitted Q2.4 samples by default, bug included, and offers a clearly labelled "Corrected" toggle next to them. Nothing in the 2023 output is silently replaced.

Scope: Assignment 1, Questions 2.4 and 2.5 (web/src/lib/ar/model-selection.ts, web/src/components/model-selection/single-sample.tsx)

Context

Question 2.4 asked for 100 observations from two autoregressive models, M1 = AR(5) and M2 = AR(2), so that Question 2.5 could compare three information criteria on one sample of each. The submitted code built the series from y, a name that still held the Hitters salary vector from Question 1, instead of from the series being generated:

y1[6:100] <- sapply(6:100, function(t) y[(t-1):(t-length(phi.m1)) %*% phi.m1 + rnorm(1)])
y2[3:100] <- sapply(3:100, function(t) y[(t-1):(t-length(phi.m2))] %*% phi.m2 + rnorm(1))

For M1 a noisy weighted sum of time indices became an index into the salaries. For M2 each value is a weighted sum of two neighbouring salaries plus noise. The Q2.5 plots therefore sit on a log-salary scale, which is why they needed ylim = c(11.5, 14). The Monte Carlo function used from Q2.6 onwards generated its own series correctly, so the 1,000-replicate tables are unaffected.

Decision

Reproduce the submitted generator exactly, as the default view, and put a "Corrected" toggle beside it that uses the same recursion the submission itself used in Q2.6. Both are labelled. The bug is described next to the chart, in coursework/README.md and in the README's "Known issues" list.

Options considered

  1. Fix it silently. Show only the corrected series and the criteria they produce.
  2. Show the submitted version only, with a footnote.
  3. Show both, submitted by default, with a toggle (chosen).
  4. Drop Q2.4 and Q2.5 from the site and start the page at the Monte Carlo study.

Why

The project's rule is that original coursework results stay faithful: upgrades add analysis around them and never change them quietly. Option 1 would have rewritten history and made the parity tests meaningless for this question. Option 2 is honest but leaves a reader with numbers that answer the wrong question. Option 4 hides the most instructive mistake in the submission. Option 3 keeps the 2023 numbers checkable (seed 10 reproduces the printed Q2.5 tables to five decimals, which model-selection.test.ts asserts) and shows what the question intended.

The bug is also a useful lesson about R sessions: state from one question leaked into the next because every chunk shared one global environment.

What happened

  • With seed 10 the submitted samples reproduce the printed tables. The criteria choose orders 3, 3 and 2 for "M1" and 7, 3 and 3 for "M2", neither of which means anything, because the data are salaries.
  • For 2 of the first 1,000 seeds the noisy M1 index falls below 1. R would then return a zero-length vector, sapply would return a list and IC() would stop with an error. The page says so instead of inventing a result.
  • With the corrected generator and seed 10, all three criteria choose p = 2 for both models. That is right for M2 and wrong for M1. Over the first 1,000 seeds, single corrected samples recover M1's true order only 6.1% (IC₁, Wilson 95% 4.8%–7.8%), 6.0% (IC₂, 4.7%–7.6%) and 1.2% (IC₃, 0.7%–2.1%) of the time, against 57.5% (54.4%–60.5%), 63.1% (60.1%–66.0%) and 77.5% (74.8%–80.0%) for M2. The correction does not make single-sample selection reliable. M1's trailing coefficients (0.108 and 0.087) are too small to detect at n = 100, which the Monte Carlo study and its Wilson intervals now show directly.

What I'd change

  • Never reuse a generic name such as y across questions. Run each question in its own environment (local({ ... })) or clear the workspace between sections.
  • Add a sanity check to every simulator: the sample variance of a simulated AR series should be close to its theoretical value. A salary-scale series would have failed that check at once.
  • Plot the raw series before computing anything on it. The 2023 report plotted only the criteria, which hid the scale.
DR-002accepted2026-10-06

Use the assignment's information-criterion definitions, labelled as "AIC-like", "AICc-like" and "BIC-like"

Decision: every reproduced number uses the three criteria exactly as the assignment defined and the submission coded them, including T = n − p and the 1e-10 ridge term. The site calls them IC₁, IC₂ and IC₃ ("AIC-like", "AICc-like", "BIC-like"), never AIC or BIC, and documents the textbook alternatives instead of switching to them.

Scope: Assignment 1, Question 2 (web/src/lib/ar/model-selection.ts, web/src/components/model-selection/*)

Context

The assignment defined, for an AR(p) fit by least squares with T = n − p usable observations and σ̂²ₚ = RSSₚ / T:

  • IC₁ = log σ̂²ₚ + 2(p + 1) / T
  • IC₂ = log σ̂²ₚ + (T + p) / (T − p − 2)
  • IC₃ = log σ̂²ₚ + p log(T) / T

These are per-observation versions of AIC, AICc and BIC, but they are not the textbook forms (−2 log L + penalty with an explicit parameter count). Two further details matter:

  1. Each order p is fitted on a different set of rows (T = n − p), so the criteria compare models estimated on slightly different data.
  2. The submitted Monte Carlo function simulate_and_evaluate_IC added 1e-10 * diag(p) to XᵀX and wrote IC₂ as (n − p + p) / (n − p − p − 2), which is the same quantity.

Decision

Keep the assignment's definitions everywhere a 2023 number is reproduced, port the 1e-10 ridge term, and label the criteria by their family rather than by the textbook names.

Options considered

  1. The assignment's definitions, as submitted (chosen).
  2. Textbook AIC and BIC from the Gaussian log-likelihood with k = p + 1 (or p + 2 counting the variance).
  3. A common estimation sample: condition every order on the first pmax = 10 observations so that all fits use T = n − 10 rows, as most software does.
  4. A toggle between definitions on the page.

Why

The questions, the submitted counts and the Q2.10 overfitting derivations are all written in terms of these forms. Options 2 and 3 would change the Monte Carlo counts, which would break parity with the PDF and blur what the coursework actually found. Option 4 doubles the interface for a point that a short note makes just as well. Labelling them "AIC-like" keeps the family resemblance without claiming they are something they are not.

What happened

With Wilson 95% intervals added to the 2023 counts (1,000 replicates each):

  • M2 = AR(2), n = 100: IC₃ selects p = 2 in 79.8% of replicates (77.2%–82.2%), IC₂ in 63.7% (60.7%–66.6%) and IC₁ in 57.7% (54.6%–60.7%). The three criteria are scored on the same 1,000 series, so they are compared pairwise with an exact McNemar test on the replicates where they disagree (replaying set.seed(10) recovers each replicate's choices). IC₃ is right where IC₂ is wrong on 234 replicates and the reverse happens on 73 (p = 8.0 × 10⁻²¹; difference 16.1 points, Newcombe paired 95% interval 12.8 to 19.4). IC₂ beats IC₁ on 61 replicates against 1 (p = 2.7 × 10⁻¹⁷; 6.0 points, 4.5 to 7.5). The separate Wilson intervals of IC₁ and IC₂ overlap slightly, so an overlap rule would have called that difference noise. The pairing shows it is real, and the familiar AIC versus BIC contrast holds here.
  • M1 = AR(5), n = 100: no criterion recovers p = 5 often. IC₁ and IC₂ manage 7.8% (6.3%–9.6%) and IC₃ only 1.4% (0.8%–2.3%). Paired, IC₁ and IC₂ disagree on 18 replicates, 9 each way (p = 1), and IC₁ beats IC₃ on 67 against 3 (p = 9.7 × 10⁻¹⁷). M1's last two coefficients are small, and the BIC-like penalty prefers orders 2 or 3. This is a weak result and the page now says so instead of letting the bar chart speak for itself.
  • n = 15: every criterion selects p ≥ 7 in every replicate, so the share choosing the true order is 0.0% (0.0%–0.4%) for both models. With eight or more coefficients only five to seven rows remain and RSS collapses towards zero. The counts there also depend on floating-point round-off in near-singular solves, which is why the n = 15 tables match only "within round-off".
  • The Q2.11 closed forms (lower tail instead of upper, an inverted scale factor and a mis-simplified IC₂ penalty) are reproduced exactly and set beside a simulation, now with Wilson intervals on the simulated probabilities.

What I'd change

  • Fit every order on a common sample (drop the first pmax observations) so that the criteria compare like with like.
  • Write the criteria as −2 log L + penalty with the parameter count stated, and cross-check against stats::ar(method = "ols") and AIC() in R.
  • Derive the overfitting probabilities with the upper tail and verify them by simulation before writing them up.
  • Report selection rates with intervals from the start, and compare criteria with a paired test because they share replicates. Raw counts invited comparisons the data could not support, and unpaired intervals would have understated the ones it could.
DR-003accepted2026-10-06 (records a decision taken during the 2026 revival)

Port the algorithms to TypeScript instead of shipping precomputed R grids

Decision: the site recomputes everything in the browser with TypeScript ports of R's random number generator, glmnet, LIBSVM (as wrapped by e1071) and the submitted functions. R stays the reference: scripts/export_artefacts.R exports fixtures and the Vitest parity suite checks the ports against them.

Scope: web/src/lib/rng, web/src/lib/glmnet, web/src/lib/svm, web/src/lib/ar, web/src/lib/bootstrap

Context

The revival turned two R Markdown assignments into an interactive, statically hosted site. The interactive controls are continuous or open-ended: λ anywhere from 10⁻² to 10¹⁰, any seed, any SVM cost and γ, any bootstrap sample size. The numbers had to stay faithful to the 2023 PDFs, and there is no budget for a server.

Decision

Port the algorithms line by line, replay R's random streams bit for bit, and test the ports against R output.

Options considered

  1. Precompute in R over a grid of control values and ship the results as JSON.
  2. WebR, R compiled to WebAssembly and run in the browser.
  3. A server running R (Plumber or Shiny) behind the site.
  4. TypeScript ports with R parity tests (chosen).
  5. Fresh TypeScript implementations without matching R's random numbers.

Why

  • A grid (option 1) is either coarse or enormous once seeds are free: one ridge path per seed per λ, one tuning table per seed. It also cannot answer questions nobody precomputed.
  • WebR (option 2) would run the original code unchanged, but it downloads tens of megabytes and starts slowly, and glmnet and e1071 would need to be built for WebAssembly.
  • A server (option 3) costs money and adds an attack surface to a portfolio site.
  • Without RNG parity (option 5) the site could only be "statistically similar" to 2023, which would not show that the ports are right. Replaying set.seed(10) means the browser draws the same 131 training players, the same folds, the same 4,000 series and the same SVM clouds, and matching to the last printed digit is strong evidence that every step is correct.

What happened

  • Parity is tight: ridge paths agree with R to about 1e-10 relative, CV curves to 1e-15, SVM tuning errors and support vectors exactly, and the n = 100 Monte Carlo tables count for count. The n = 15 tables match only within round-off, because their fits leave no residual degrees of freedom and R itself differs between versions there.
  • The cost was real engineering: Mersenne-Twister seeding, inversion normals with R's 2²⁷ trick, rejection sampling in sample(), glmnet's standardised-y ridge quirk and LIBSVM's float32 kernel cache all had to match.
  • The ports paid for the 2026 rigour work. Repeated cross-validation over 200 fold seeds takes about three seconds, and a 500-dataset coverage study with 999 bootstrap resamples each takes about two on a laptop, so both run at build time and again in a Web Worker on demand. With precomputed grids neither would exist.
  • The risk is silent divergence on inputs the fixtures do not cover. The tests reduce it but cannot remove it.

What I'd change

  • Record R's sessionInfo() and package versions inside every fixture file, not only in the README.
  • Add randomised comparisons against R (random seeds, random λ, random SVM parameters) as a scheduled job, instead of relying on fixed fixtures alone.
  • Consider WebR for an "open in R" view of the original code, alongside the ports rather than instead of them.
DR-004accepted2026-10-06

Optional AI explanations with the visitor's own key, called from the browser and audit-logged locally

Decision: the only AI feature is an optional "Explain this simulation" panel. It uses the visitor's own Anthropic or OpenAI key, calls the provider directly from the browser, sends only the numbers already on screen, labels every output as AI-generated, and records every call (without the key) in an audit log in the visitor's IndexedDB. No AI output ever feeds a computed result.

Scope: web/src/lib/ai, web/src/components/ai, /ai-log, the "AI use statement" on /methods

Context

The upgrade had to show working GenAI fluency with governance, on a static site with no server and no budget for AI. The governance principles it is informed by (the Australian DTA policy for the responsible use of AI in government, the EU AI Act's transparency principles and the NIST AI RMF) all point the same way: tell people when content is AI-generated, keep a human in the loop, keep records, and be clear about what data goes where. The site makes no claim of compliance with any of them.

Decision

Bring-your-own-key only. Keys are stored in sessionStorage by default, or localStorage if the visitor opts in, and can be forgotten with one button. Requests go straight to api.anthropic.com (with the anthropic-dangerous-direct-browser-access header) or api.openai.com. Replies are constrained with structured outputs, validated with zod, checked for numbers that cannot be traced to the inputs, and shown with Accept, Edit and Reject controls whose outcome is written to the audit log.

Options considered

  1. No AI at all.
  2. A server-side proxy with a site-owned key. It needs a budget, rate limiting and abuse handling.
  3. Bring your own key through a proxy. The visitor's key would pass through infrastructure they do not control.
  4. Bring your own key, called directly from the browser (chosen).
  5. An on-device model (WebLLM or similar). It needs gigabyte downloads and a capable GPU.

Why

Option 4 costs nothing to run, keeps the key off any infrastructure the site controls, and keeps the audit trail with the person who made the calls. The feature is deliberately small: it explains results and never produces them, so a wrong explanation cannot change a number on the page. Plain fetch adapters rather than the provider SDKs keep the static bundle small and make the exact request auditable. The panel shows the prompt that was sent and the other request settings.

What happened

  • The adapters, key storage, audit store, CSV and JSON export, grounding check and error mapping are covered by 35 unit tests with fetch mocked. They have not been exercised against the live APIs in CI, by design: there is no key to use. The request shapes follow the providers' documentation and are checked against it only through mocks, which is a real limitation.
  • The grounding check catches invented numbers but not wrong reasoning built on correct numbers, so it is presented as a warning rather than a guarantee.
  • Direct browser access means any script running on the page could read the key. The site loads no third-party scripts, keeps the key in sessionStorage by default and sends a Content-Security-Policy that keeps fetches and resource loads (images, styles, fonts, workers) on this site, with connect-src adding only the two provider APIs. That reduces the ways an injected script could send a key elsewhere, but it cannot stop a script from navigating the page away, so it is a mitigation rather than a guarantee. A first version set only connect-src; a review showed an image request could still reach another host, and the policy now sets default-src 'self' and explicit image, font and worker sources.
  • Server-side refusal fallbacks are not enabled. For Claude Sonnet 5.5 they cover refusal categories that cannot plausibly apply to explaining a statistics result, and the extra beta header could fail for an organisation that has not enabled it. Refusals are reported to the visitor instead.
  • OpenAI 401 messages echo a masked fragment of the key. Error text is redacted before it is shown or logged, and a record that contains the key is refused outright.
  • The audit record keeps the request settings (max tokens, output schema, effort), the grounding result shown to the reader, an append-only decision history, and, when a reply fails validation, the provider's raw text (redacted and truncated to 8,000 characters). Failed calls are marked "not applicable" rather than waiting for a decision.

What I'd change

  • Add an evaluation harness if the feature grows: a fixed set of results with reference readings, scored for factual agreement with intervals across repeated runs and models.
  • Tighten the Content-Security-Policy further with nonce-based script-src.
  • Let visitors export and import the audit log across browsers, and add a retention setting.