Supervised Learning (Classification) & Model Evaluation
From the Machine Learning curriculum
TL;DR
Supervised learning for classification trains models to predict categories from labeled data. You split data into training and testing sets to evaluate how well your model generalizes. Key evaluation metrics like accuracy, precision, recall, and F1-score help you understand your model's performance beyond simple correctness.
1. The Mental Model
Imagine you're teaching a child to tell the difference between pictures of cats and dogs. You show them many labeled examples (cat or dog), and they learn to identify new pictures. Supervised classification is the same, but with algorithms and data.
2. The Core Material
In supervised classification, you're trying to predict a discrete label or category for a given input. This is different from regression, where you predict a continuous value. Think "spam or not spam," "disease or no disease," or "customer will churn or not churn."
The "supervised" part means your training data comes with the correct answers (labels). Your model learns the patterns that link the input features to these known labels.
2.1 The Classification Process

Photo by Tima Miroshnichenko on Pexels
The general steps for a classification task look like this:
graph TD
A["Gather & Prepare Labeled Data"] --> B["Split Data (Train/Test)"];
B --> C["Choose a Classification Model (e.g., Logistic Regression, Decision Tree)"];
C --> D["Train Model on Training Data"];
D --> E["Make Predictions on Test Data"];
E --> F["Evaluate Model Performance"];
F --> G["Fine-tune Model or Repeat"];
- Gather & Prepare Labeled Data: Collect your data and ensure each data point has the correct category assigned. This often involves cleaning and transforming the data.
- Split Data (Train/Test): It's crucial to split your dataset into at least two parts: a training set (what the model learns from) and a testing set (what you use to evaluate how well the model generalizes to new, unseen data). A common split is 70-80% for training and 20-30% for testing. Never train on your test set.
- Choose a Model: Select an algorithm suitable for classification. Examples include Logistic Regression, Decision Trees, Support Vector Machines (SVMs), and K-Nearest Neighbors (KNN).
- Train Model: The chosen algorithm "learns" the patterns from your training data.
- Make Predictions: Use your trained model to predict the categories for the data in your unseen test set.
- Evaluate Model Performance: Compare the model's predictions on the test set to the actual labels. This is where evaluation metrics come in.
2.2 Model Evaluation Metrics

Photo by RDNE Stock project on Pexels
Simply looking at accuracy (correct predictions / total predictions) can be misleading, especially with imbalanced datasets (e.g., 95% of emails are not spam, so predicting "not spam" all the time gives 95% accuracy but catches no spam).
For a deeper understanding, we use a confusion matrix and metrics derived from it. Let's assume a binary classification problem (e.g., "positive" or "negative" class).
| Predicted Positive | Predicted Negative | |
|---|---|---|
| Actual Positive | True Positive (TP) | False Negative (FN) |
| Actual Negative | False Positive (FP) | True Negative (TN) |
- TP (True Positive): Correctly predicted the positive class.
- FN (False Negative): Incorrectly predicted the negative class (missed a positive).
- FP (False Positive): Incorrectly predicted the positive class (a false alarm).
- TN (True Negative): Correctly predicted the negative class.
From these, we derive key metrics:
- Accuracy: (TP + TN) / (TP + TN + FP + FN)
- What it tells you: Overall correctness.
- Precision (Positive Predictive Value): TP / (TP + FP)
- What it tells you: Out of all predicted positives, how many were actually positive? High precision means fewer false alarms.
- Recall (Sensitivity, True Positive Rate): TP / (TP + FN)
- What it tells you: Out of all actual positives, how many did the model correctly identify? High recall means fewer missed positives.
- F1-Score: 2 * (Precision * Recall) / (Precision + Recall)
- What it tells you: The harmonic mean of precision and recall. Useful when you need a balance between precision and recall, especially with uneven class distribution.
- Specificity (True Negative Rate): TN / (TN + FP)
- What it tells you: Out of all actual negatives, how many did the model correctly identify?
The best metric to optimize depends on your problem.
* If missing a positive is very costly (e.g., medical diagnosis for a serious disease), you'll want high recall.
* If false positives are very costly (e.g., flagging innocent people as criminals), you'll want high precision.
3. Worked Example
Let's say you've trained a model to detect fraudulent transactions. You test it on 100 transactions, and here are the results:
- Actual Fraud (Positive): 10 transactions
- Actual Legitimate (Negative): 90 transactions
Your model's predictions:
- It correctly identified 8 fraudulent transactions (TP = 8).
- It missed 2 fraudulent transactions (FN = 2).
- It incorrectly flagged 5 legitimate transactions as fraud (FP = 5).
- It correctly identified 85 legitimate transactions (TN = 85).
Let's calculate the metrics:
- Accuracy: (8 + 85) / (8 + 2 + 5 + 85) = 93 / 100 = 0.93 (93%)
- Precision: 8 / (8 + 5) = 8 / 13 = 0.615 (61.5%)
- Interpretation: When your model says a transaction is fraud, it's correct about 61.5% of the time.
- Recall: 8 / (8 + 2) = 8 / 10 = 0.80 (80%)
- Interpretation: Your model caught 80% of all actual fraudulent transactions.
- F1-Score: 2 * (0.615 * 0.80) / (0.615 + 0.80) = 2 * 0.492 / 1.415 = 0.984 / 1.415 = 0.695
While the accuracy looks good at 93%, the precision of 61.5% means many legitimate transactions are being falsely flagged. This might be acceptable if catching fraud is paramount, but it could lead to customer frustration. The high recall (80%) is good for not missing much fraud.
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, confusion_matrix
# 1. Dummy Data (Replace with your actual data)
data = {
'feature_1': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20],
'feature_2': [2, 3, 4, 5, 6, 7, 8, 9, 10, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11],
'target': [0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 1] # 6 ones (fraud), 14 zeros (legit)
}
df = pd.DataFrame(data)
X = df[['feature_1', 'feature_2']]
y = df['target']
# 2. Split Data
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42, stratify=y)
# stratify=y ensures both train/test sets have similar proportions of target classes
# 3. Choose and Train Model
model = LogisticRegression(random_state=42)
model.fit(X_train, y_train)
# 4. Make Predictions
y_pred = model.predict(X_test)
# 5. Evaluate Model Performance
print("--- Model Evaluation ---")
print(f"Accuracy: {accuracy_score(y_test, y_pred):.2f}")
print(f"Precision: {precision_score(y_test, y_pred):.2f}")
print(f"Recall: {recall_score(y_test, y_pred):.2f}")
print(f"F1-Score: {f1_score(y_test, y_pred):.2f}")
print("\nConfusion Matrix:")
print(confusion_matrix(y_test, y_pred))
# Example output on this dummy data (may vary slightly due to random split):
# --- Model Evaluation ---
# Accuracy: 0.83
# Precision: 0.67
# Recall: 0.67
# F1-Score: 0.67
#
# Confusion Matrix:
# [[4 1]
# [1 2]]
#
# Interpretation of Confusion Matrix:
# Actual Negatives (row 0): 4 True Negatives, 1 False Positive
# Actual Positives (row 1): 1 False Negative, 2 True Positives
4. Key Takeaways
- Supervised classification predicts discrete categories using labeled historical data.
- Always split your data into training and testing sets to evaluate generalization.
- Accuracy alone isn't sufficient for evaluating classification models, especially with imbalanced datasets.
- Understand the meaning of True Positives, False Positives, True Negatives, and False Negatives from a confusion matrix.
- Choose evaluation metrics (precision, recall, F1-score) based on the specific costs of false positives vs. false negatives for your problem.
Common Mistakes to Avoid:
* **Training and testing on the
Frequently asked about Supervised Learning (Classification) & Model Evaluation
Study this next
Get the full Machine Learning curriculum
Clone the complete plan to your dashboard for unlimited AI-generated notes, practice quizzes, and a personalised revision schedule.
Create Free Account