A classifier does more than return a label: it makes correct predictions and different kinds of mistakes. Classification metrics tell us which mistakes a model makes. This matters when, for example, a spam filter must catch unwanted mail without hiding important messages.
Start with a confusion matrix
Suppose a filter checks 100 messages. Twenty are actually spam and 80 are good mail. It blocks 16 spam messages, misses 4, and incorrectly blocks 8 good messages.
| Result | Meaning | Count |
|---|---|---|
| True positive (TP) | Spam correctly blocked | 16 |
| False positive (FP) | Good mail incorrectly blocked | 8 |
| False negative (FN) | Spam incorrectly delivered | 4 |
| True negative (TN) | Good mail correctly delivered | 72 |
Here, positive means “predicted spam.” A false positive is not simply a wrong answer; it is specifically a good message blocked as spam. Always define the positive class before reading a confusion matrix.
Accuracy: how often is the model correct?
Accuracy = (TP + TN) / all examples. In this example, (16 + 72) / 100 = 88%. Accuracy is useful when the classes and the costs of their errors are reasonably balanced.
Accuracy alone can hide a weak model. A filter that labels every message “not spam” would be right on 80 of these 100 messages, yet it would catch no spam. When one class is uncommon, examine the other metrics too.
Precision: can we trust a positive prediction?
Precision = TP / (TP + FP). The filter blocked 24 messages in total; 16 were truly spam. Its precision is 16 / 24 = 66.7%. Put another way, about one in three blocked messages was good mail.
Precision matters when acting on a false alarm is costly. For a spam filter, a low-precision rule may hide a receipt or a job offer. A system could raise its decision threshold to block fewer uncertain messages, but it might then miss more spam.
Recall: how many real positives did we find?
Recall = TP / (TP + FN). There were 20 spam messages; the filter caught 16. Its recall is 16 / 20 = 80%. The remaining four spam messages were false negatives.
Recall matters when missing a real case is costly. Lowering a classifier's threshold often catches more positives, but it can also create more false alarms. The right balance depends on the task, not on a universal “best” score.
F1 score: combine precision and recall
F1 = 2 × precision × recall / (precision + recall). For this filter, F1 is about 72.7%. F1 is a compact way to compare precision and recall when both matter. It does not include true negatives, so keep the confusion matrix and the task's error costs in view.
Run the numbers in Python
The following example calculates all four metrics from the same counts. Edit the values to see how one kind of error changes the scores.
true_positive = 16
false_positive = 8
false_negative = 4
true_negative = 72
total = true_positive + false_positive + false_negative + true_negative
accuracy = (true_positive + true_negative) / total
precision = true_positive / (true_positive + false_positive)
recall = true_positive / (true_positive + false_negative)
f1 = 2 * precision * recall / (precision + recall)
print(f"Accuracy: {accuracy:.1%}")
print(f"Precision: {precision:.1%}")
print(f"Recall: {recall:.1%}")
print(f"F1 score: {f1:.1%}")
Output: Accuracy: 88.0%; Precision: 66.7%; Recall: 80.0%; F1 score: 72.7%.
If a denominator is zero—for example, the model never predicts the positive class—you must define how your evaluation handles that case rather than dividing by zero.
Which metric should you choose?
- Use precision when false alarms are especially harmful.
- Use recall when missed positives are especially harmful.
- Use F1 when you want one summary of precision and recall.
- Always inspect the confusion matrix to understand the actual error counts.
Calculate metrics on examples the model did not train on. Comparing only training predictions can give an overly optimistic result. Start by naming the positive class and asking which mistake matters more in the real application.