Classification Metrics: Accuracy, Precision, Recall, and F1

Artificial Intelligence Tutorials


A classifier does more than return a label: it makes correct predictions and different kinds of mistakes. Classification metrics tell us which mistakes a model makes. This matters when, for example, a spam filter must catch unwanted mail without hiding important messages.

Start with a confusion matrix

Suppose a filter checks 100 messages. Twenty are actually spam and 80 are good mail. It blocks 16 spam messages, misses 4, and incorrectly blocks 8 good messages.

A spam-filter confusion matrix separates correct predictions from false positives and false negatives.
A spam-filter confusion matrix separates correct predictions from false positives and false negatives.
Result Meaning Count
True positive (TP) Spam correctly blocked 16
False positive (FP) Good mail incorrectly blocked 8
False negative (FN) Spam incorrectly delivered 4
True negative (TN) Good mail correctly delivered 72

Here, positive means “predicted spam.” A false positive is not simply a wrong answer; it is specifically a good message blocked as spam. Always define the positive class before reading a confusion matrix.

Accuracy: how often is the model correct?

Accuracy = (TP + TN) / all examples. In this example, (16 + 72) / 100 = 88%. Accuracy is useful when the classes and the costs of their errors are reasonably balanced.

Accuracy alone can hide a weak model. A filter that labels every message “not spam” would be right on 80 of these 100 messages, yet it would catch no spam. When one class is uncommon, examine the other metrics too.

Precision: can we trust a positive prediction?

Precision = TP / (TP + FP). The filter blocked 24 messages in total; 16 were truly spam. Its precision is 16 / 24 = 66.7%. Put another way, about one in three blocked messages was good mail.

Precision matters when acting on a false alarm is costly. For a spam filter, a low-precision rule may hide a receipt or a job offer. A system could raise its decision threshold to block fewer uncertain messages, but it might then miss more spam.

Recall: how many real positives did we find?

Recall = TP / (TP + FN). There were 20 spam messages; the filter caught 16. Its recall is 16 / 20 = 80%. The remaining four spam messages were false negatives.

Recall matters when missing a real case is costly. Lowering a classifier's threshold often catches more positives, but it can also create more false alarms. The right balance depends on the task, not on a universal “best” score.

F1 score: combine precision and recall

F1 = 2 × precision × recall / (precision + recall). For this filter, F1 is about 72.7%. F1 is a compact way to compare precision and recall when both matter. It does not include true negatives, so keep the confusion matrix and the task's error costs in view.

Run the numbers in Python

The following example calculates all four metrics from the same counts. Edit the values to see how one kind of error changes the scores.

true_positive = 16
false_positive = 8
false_negative = 4
true_negative = 72

total = true_positive + false_positive + false_negative + true_negative
accuracy = (true_positive + true_negative) / total
precision = true_positive / (true_positive + false_positive)
recall = true_positive / (true_positive + false_negative)
f1 = 2 * precision * recall / (precision + recall)

print(f"Accuracy: {accuracy:.1%}")
print(f"Precision: {precision:.1%}")
print(f"Recall: {recall:.1%}")
print(f"F1 score: {f1:.1%}")

Output: Accuracy: 88.0%; Precision: 66.7%; Recall: 80.0%; F1 score: 72.7%.

If a denominator is zero—for example, the model never predicts the positive class—you must define how your evaluation handles that case rather than dividing by zero.

Which metric should you choose?

  • Use precision when false alarms are especially harmful.
  • Use recall when missed positives are especially harmful.
  • Use F1 when you want one summary of precision and recall.
  • Always inspect the confusion matrix to understand the actual error counts.

Calculate metrics on examples the model did not train on. Comparing only training predictions can give an overly optimistic result. Start by naming the positive class and asking which mistake matters more in the real application.



Found This Page Useful? Share It!
Get the Latest Tutorials and Updates
Join us on Telegram