Skip to content
CompStats PlaygroundMAST90083 · 2023 S2

Assignment 3 · Question 1 · multiclass SVM

Three clouds, one margin at a time.

Three hundred points are drawn from standard normals and shifted to class means (0, 0), (0, 3) and (3, 0). The task was to fit support vector classifiers with R's e1071 package, tune the cost (and the radial kernel's γ) by ten-fold cross-validation, and score the best models on a fresh test set. Here a port of LIBSVM's SMO solver retrains as you move the sliders. It finds the same support vectors and the same tuning tables as R.

[Q1.2–1.5] svm(y ~ ., data = tdata, kernel = "linear", cost = 10)

Decision regions of a three-class support vector machine

One-vs-one C-classification on standardised inputs, exactly as e1071 fits it. Shading shows the predicted class; outlined points are support vectors. Shapes repeat the class colours.
class 1class 2class 3support vector
C = 10

Small C tolerates margin violations (many support vectors); large C fits the training points hard.

means (0,0), (0,3), (3,0)

The assignment used 3. Same seeds, so only the shift changes.

Support vectors
…
training
Test misclassified
…
 

Test confusion table

Why one class has more than 100 correct predictions (Q1.4)

The test labels are drawn with sample(1:3, 300, replace = TRUE), so the classes are not balanced: class 3 has 114 test points. The linear model misclassifies 28 of 300 and the tuned radial model 31, so a curved boundary buys nothing here. Three shifted Gaussians are close to linearly separable.

2026: with uncertainty. Error rates 9.3% (Wilson 95% 6.5–13.2%) and 10.3% (7.4–14.3%). Both models are scored on the same 300 points: only the linear model is right on 10 and only the radial on 7, exact McNemar p = 0.63. The accuracy difference is +1.0 points (Newcombe paired 95% interval −1.9 to +4.0), so the 28 vs 31 gap is not detectable: the data cannot tell the two boundaries apart.

[Q1.3, Q1.5] set.seed(50); tune(svm, y ~ ., data = tdata, kernel, ranges = list(...))

Ten-fold cross-validation over the tuning grid

e1071's tune() draws one permutation with sample(300), cuts it into ten folds of 30, and records the misclassification rate and its spread for every parameter combination. The best model is refitted on all 300 points.

Linear kernel

Run the linear tune to reproduce Q1.3: best cost 1 with 8.0% CV error and 81 support vectors.

Radial kernel

Run the 25-combination radial tune to reproduce Q1.5: best cost 100, γ = 2 with 7.3% CV error.