A4.3.3 HL · evaluating models
A model that fits its training data perfectly can still do badly on new data. The first demo shows how making a model more complex changes its error on training data and on test data. The second shows how moving the classification threshold changes the accuracy, precision, recall and F1 score of the same model.
training error
0.003
test error
0.007
This looks like a good fit.
Filled dots are training data. Orange rings are test data. A low polynomial degree may miss the pattern (underfitting). A high degree may fit training noise and perform poorly on new data (overfitting). Compare degrees using the test error here. In a real project, tune on a validation set and reserve the test set for final evaluation.
12
TP · caught
1
FN · missed
0
FP · false alarm
13
TN · correctly cleared
Slide the threshold to the right and the model predicts fewer positives. Precision can go up or down. Recall goes down each time an actual positive drops below the threshold. F1 score combines precision and recall into one number.