Foundations of Machine Learning & Supervised Learning (Regression)
From the Machine Learning curriculum
TL;DR
Machine learning lets computers learn patterns from data without explicit programming, making predictions or decisions. Supervised learning, especially regression, trains models on labeled data to predict continuous numerical values. Understanding core concepts like features, labels, and model evaluation is crucial for building effective regression models.
1. The Mental Model
Think of machine learning like teaching a child. You give them examples (data), they try to find a rule (model), and then they can apply that rule to new situations. In regression, you're teaching them to guess a number, like how many candies to expect based on the number of friends.
2. The Core Material
Machine learning (ML) is a field where systems learn from data, identify patterns, and make decisions with minimal human intervention. It’s broadly categorized into supervised, unsupervised, and reinforcement learning. We'll focus on supervised learning, which uses labeled data – data where the correct answer (the "label") is already known – to train models.
Within supervised learning, there are two main types:
* Classification: Predicting a categorical label (e.g., "spam" or "not spam", "cat" or "dog").
* Regression: Predicting a continuous numerical value (e.g., house price, temperature, sales figures). This is our focus today.
Key Concepts

Photo by Ingo Joseph on Pexels
- Features (X): These are the input variables or characteristics used to make a prediction. In predicting house prices, features might include square footage, number of bedrooms, or location.
- Labels (y): This is the output variable you're trying to predict. For house prices, the label is the actual price.
- Training Data: The dataset used to teach the model. It contains both features and their corresponding labels.
- Model: The algorithm or mathematical function learned from the training data that maps features to labels.
- Prediction: The output generated by the model for new, unseen features.
The Regression Process

Photo by Sinan KRIYA on Pexels
Here's how a typical supervised regression project flows:
graph TD
A["Gather Data (Features & Labels)"] --> B["Split Data (Train/Test)"]
B --> C["Choose a Model (e.g., Linear Regression)"]
C --> D["Train Model on Training Data"]
D --> E["Evaluate Model on Test Data"]
E --> F["Tune Model / Deploy"]
Linear Regression

Photo by Google DeepMind on Pexels
Linear Regression is one of the simplest and most fundamental regression algorithms. It assumes a linear relationship between the features and the label. For a single feature, it tries to find the best-fit straight line through your data points.
The equation for simple linear regression (one feature) is:
$y = mx + b$
Where:
* $y$ is the predicted label (e.g., house price).
* $x$ is the feature (e.g., square footage).
* $m$ is the slope of the line, representing how much $y$ changes for a unit change in $x$.
* $b$ is the y-intercept, the predicted $y$ when $x$ is 0.
The model "learns" the values of $m$ and $b$ that minimize the difference between its predictions and the actual labels in the training data.
Evaluating Regression Models

Photo by Markus Winkler on Pexels
How do we know if our regression model is good? We use metrics:
- Mean Absolute Error (MAE): The average absolute difference between predicted and actual values. It's easy to interpret as it's in the same units as the label.
$MAE = \frac{1}{N} \sum_{i=1}^{N} |y_i - \hat{y}_i|$ - Mean Squared Error (MSE): The average of the squared differences. It penalizes larger errors more heavily.
$MSE = \frac{1}{N} \sum_{i=1}^{N} (y_i - \hat{y}_i)^2$ - Root Mean Squared Error (RMSE): The square root of MSE. It's also in the same units as the label, making it more interpretable than MSE.
$RMSE = \sqrt{\frac{1}{N} \sum_{i=1}^{N} (y_i - \hat{y}_i)^2}$ - R-squared ($R^2$): Represents the proportion of the variance in the dependent variable that's predictable from the independent variables. A value of 1 means the model perfectly predicts the label, while 0 means it explains none of the variance.
3. Worked Example
Let's predict exam scores based on hours studied using simple linear regression in Python with scikit-learn.
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, r2_score
# 1. Gather Data (Features and Labels)
# Hours studied (X)
hours_studied = np.array([2, 3, 4, 5, 6, 7, 8, 9, 10, 1]).reshape(-1, 1)
# Exam scores (y)
exam_scores = np.array([55, 60, 65, 70, 75, 80, 85, 90, 95, 50])
print("--- Original Data ---")
for i in range(len(hours_studied)):
print(f"Hours: {hours_studied[i][0]}, Score: {exam_scores[i]}")
# 2. Split Data (Train/Test)
# We'll use 80% for training and 20% for testing
X_train, X_test, y_train, y_test = train_test_split(
hours_studied, exam_scores, test_size=0.2, random_state=42
)
print("\n--- Training Data ---")
print(f"X_train:\n{X_train}")
print(f"y_train:\n{y_train}")
print("\n--- Testing Data ---")
print(f"X_test:\n{X_test}")
print(f"y_test:\n{y_test}")
# 3. Choose a Model (Linear Regression)
model = LinearRegression()
# 4. Train Model on Training Data
model.fit(X_train, y_train)
# Print the learned coefficients
print(f"\nLearned slope (m): {model.coef_[0]:.2f}")
print(f"Learned intercept (b): {model.intercept_:.2f}")
# 5. Make predictions on the test set
y_pred = model.predict(X_test)
print("\n--- Predictions vs. Actual ---")
for i in range(len(X_test)):
print(f"Hours: {X_test[i][0]}, Actual Score: {y_test[i]}, Predicted Score: {y_pred[i]:.2f}")
# 6. Evaluate Model on Test Data
mae = mean_absolute_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)
print(f"\nMean Absolute Error (MAE): {mae:.2f}")
print(f"R-squared (R2): {r2:.2f}")
# Predict for a new, unseen value (e.g., 6.5 hours studied)
new_hours = np.array([[6.5]])
predicted_score = model.predict(new_hours)
print(f"\nPredicted score for 6.5 hours of study: {predicted_score[0]:.2f}")
In this example, our model learned a slope and intercept. The MAE tells us, on average, our predictions were off by about 3.39 points. An R-squared of 0.99 is very high, indicating our model explains 99% of the variance in exam scores based on hours studied.
4. Key Takeaways
- Machine learning enables systems to learn from data for predictions and decisions, with supervised learning using labeled data.
- Regression specifically predicts continuous numerical values, like house prices or temperatures.
- Features are input variables, and labels are the target values you're trying to predict.
- The regression process involves data splitting, model training, and evaluation using metrics like MAE, MSE, RMSE, and R-squared.
- Linear Regression is a foundational algorithm that models a linear relationship between features and labels.
- Model evaluation metrics help you understand how well your model performs on unseen data.
Common Mistakes to Avoid:
- Not splitting your data: Always divide your data into training and testing sets to properly evaluate how your model generalizes to new data.
- Overfitting: Creating a model that's too complex and learns the training data too well, failing to perform well on new data.
- Ignoring data preprocessing: Real-world data is often messy; cleaning and preparing it is a crucial first step.
- Using the wrong evaluation metric: Choose metrics appropriate for regression tasks and interpret them correctly.
5. Now Try It
Spend 15 minutes:
- Modify the example code: Change the
hours_studiedandexam_scoresdata points to create a slightly different linear relationship (e.g., make the scores increase faster or slower with hours). - Run the modified code: Observe how the
Learned slope (m)andLearned intercept (b)change. - Analyze the new metrics: How do the MAE and R-squared values change with your new data?
- Try a new prediction: Predict an exam score for an hour studied value you didn't include in your initial data.
Success looks like: You've successfully run the code, interpreted the new slope/intercept, and understood how data changes impact the model's performance metrics and predictions.
Frequently asked about Foundations of Machine Learning & Supervised Learning (Regression)
Study this next
Get the full Machine Learning curriculum
Clone the complete plan to your dashboard for unlimited AI-generated notes, practice quizzes, and a personalised revision schedule.
Create Free Account