Assignment 1 · Question 1 · Hitters salaries
Shrink, select, validate.
The task: predict 263 Major League Baseball players' 1987 salaries from 19 career statistics, and compare ridge regression with the lasso. Both minimise squared error plus a penalty whose strength is set by . Ridge uses and the lasso uses , which can zero coefficients out.
Everything below is recomputed live by a line-by-line TypeScript port of glmnet's coordinate descent, with R's random number generator replayed so set.seed(10) picks the same training rows and folds as the 2023 submission.
[Q1.2] glmnet(x, y, alpha, lambda = 10^seq(10, -2, length = 100))
Ridge coefficient paths
[Q1.3] sqrt(sum(beta^2)) vs log(lambda)
ℓ₂ norm of the coefficients
[Q1.4] cv.glmnet(X_train, y_train, nfolds = 10, alpha = 0)
Cross-validation curve
sample(), replayed bit-for-bit.[Q1.4–1.5] mean((predict(fit, X_test) - y_test)^2)
Test MSE on the 132 held-out players
[Q1.4–1.5] cbind(coef(ridge), coef(lm), coef(lasso))
Coefficients on the full data
ridge and lasso at λ = 301
| Predictor | Ridge | Lasso | OLSLeast squares |
|---|---|---|---|
| (Intercept) | 13.522 | 535.93 | 163.10 |
| 0.0705 | · | −1.980 | |
| 0.8839 | · | 7.501 | |
| 0.5225 | · | 4.331 | |
| 1.074 | · | −2.376 | |
| 0.8794 | · | −1.045 | |
| 1.657 | · | 6.231 | |
| 1.164 | · | −3.489 | |
| 0.0114 | · | −0.1713 | |
| 0.0588 | · | 0.1340 | |
| 0.4147 | · | −0.1729 | |
| 0.1169 | · | 1.454 | |
| 0.1238 | · | 0.8077 | |
| 0.0494 | · | −0.8116 | |
| 23.007 | · | 62.599 | |
| −81.473 | · | −116.85 | |
| 0.1710 | · | 0.2819 | |
| 0.0315 | · | 0.3711 | |
| −1.440 | · | −3.361 | |
| 8.912 | · | −24.762 |
What the submission concluded
At the CV-chosen λ (300.9) every ridge coefficient is still present but smaller than its least-squares counterpart. At λ = 4.468 the lasso drops HmRun, Runs, RBI, CAtBat, CHits and NewLeagueN entirely. Press λmin with each method selected to reproduce the table in the PDF.
[rigour] paired bootstrap over the 132 held-out players
Is any model actually better on the test set?
Test MSE on n = 132 players, with 95% bootstrap intervals
| Model | MSE | 95% interval |
|---|---|---|
| Ridgeλ = 300.9 | 143,261 | 79,414–235,366 |
| Lassoλ = 4.468 | 141,633 | 79,751–234,205 |
| OLSno penalty | 145,023 | 86,146–232,357 |
Paired differences (first minus second; negative favours the first)
| Comparison | Δ MSE [95%] | ratio | first wins |
|---|---|---|---|
| Ridge − OLS | −1,762−23,361 to 21,493 | 0.9880.81–1.15 | 61.4%52.8%–69.2% |
| Lasso − OLS | −3,390−16,249 to 10,374 | 0.9770.86–1.06 | 54.5%46.0%–62.8% |
| Ridge − Lasso | +1,629−9,533 to 13,714 | 1.0110.92–1.11 | 58.3%49.8%–66.4% |
Paired MSE differences with 95% intervals
An interval that crosses the dashed zero line is not a detectable difference.
What this shows
All three intervals include zero: on 132 test players, ridge, lasso and least squares cannot be told apart by mean squared error. The ranking by point estimate (Lasso < Ridge < OLS) is within resampling noise.
Ridge has the smaller error for 61.4% of players (Wilson 52.8%–69.2%), so it is usually a little closer than OLS, yet the mean difference is not detectable: a few large errors decide the average.
Salaries are heavy-tailed: the five largest OLS squared errors make up 42.6% of its test error, which is why the intervals are wide. These intervals hold the training split fixed; a new split (try another seed above) moves all three MSEs as well.
Explain this simulationoptional · your own key
Add your own Anthropic or OpenAI key in AI settings to enable. Nothing is sent without one.
[rigour] cv.glmnet(X_train, y_train, nfolds = 10, foldid = <set.seed(s)>), s = 1, 2, …
How stable is λmin? Repeated cross-validation
Ridge: log λmin over 200 fold draws
log λmin
- 2023 draw: λ = 300.9
- median: λ = 300.9
- λmin, middle 95% of draws
- 188.5 to 436.5
- draws on the 2023 grid point (index 76)
- 14.0% [9.9%, 19.5%]
- draws with a smaller λmin than 2023
- 40.5% [33.9%, 47.4%]
- test MSE at λmin, middle 95%
- 142,104 to 144,261
Lasso: log λmin over 200 fold draws
log λmin
- 2023 draw: λ = 4.468
- median: λ = 16.44
- λmin, middle 95% of draws
- 4.062 to 23.85
- draws on the 2023 grid point (index 47)
- 5.5% [3.1%, 9.6%]
- draws with a smaller λmin than 2023
- 6.0% [3.5%, 10.2%]
- non-zero coefficients (training fit): 2023 vs median
- 10 vs 8
- test MSE at λmin, middle 95%
- 141,408 to 142,793
What this changes about the 2023 reading
Ridge's 2023 λmin of 300.9 is a typical draw: 40.5% of fold seeds give a smaller value and the median draw is 300.9. The lasso's 2023 λmin of 4.468 falls in a secondary mode that 14.5% of fold seeds pick (the CV curve has competing minima): 6.0% of fold seeds give a smaller value and the median draw is 16.44. With a smaller penalty than usual, the 2023 lasso kept 10 predictors on the training fit where the median fold draw keeps 8. Across the middle 95% of draws the lasso's test MSE moves by about 1,385, small next to the bootstrap uncertainty of the test MSE itself.
The submitted values are unchanged and still reproduced exactly; this cell only shows how much they depend on one partition. Averaging the CV curve over several fold draws, or using λ1se, would be the more stable choice.
Shares in brackets are Wilson 95% intervals over the 200 fold seeds. λmin is always one of the 100 (ridge) or 94 (lasso) values on glmnet's default grid for this training set, so the histograms are lumpy by construction. Grid index 76 is R's 1-based position.
Explain this simulationoptional · your own key
Add your own Anthropic or OpenAI key in AI settings to enable. Nothing is sent without one.