Stanford University CS229

Foundations of Machine Learning & Supervised Learning (Regression)

SA
StudyAI Editorial
Reviewed by StudyAI tutors
· Published Updated

From the Machine Learning curriculum

TL;DR

Machine learning lets computers learn patterns from data without explicit programming, making predictions or decisions. Supervised learning, especially regression, trains models on labeled data to predict continuous numerical values. Understanding core concepts like features, labels, and model evaluation is crucial for building effective regression models.

1. The Mental Model

Think of machine learning like teaching a child. You give them examples (data), they try to find a rule (model), and then they can apply that rule to new situations. In regression, you're teaching them to guess a number, like how many candies to expect based on the number of friends.

2. The Core Material

Machine learning (ML) is a field where systems learn from data, identify patterns, and make decisions with minimal human intervention. It’s broadly categorized into supervised, unsupervised, and reinforcement learning. We'll focus on supervised learning, which uses labeled data – data where the correct answer (the "label") is already known – to train models.

Within supervised learning, there are two main types:
* Classification: Predicting a categorical label (e.g., "spam" or "not spam", "cat" or "dog").
* Regression: Predicting a continuous numerical value (e.g., house price, temperature, sales figures). This is our focus today.

Key Concepts

A set of metal house keys on a wooden surface, ideal for real estate themes.
Photo by Ingo Joseph on Pexels

  • Features (X): These are the input variables or characteristics used to make a prediction. In predicting house prices, features might include square footage, number of bedrooms, or location.
  • Labels (y): This is the output variable you're trying to predict. For house prices, the label is the actual price.
  • Training Data: The dataset used to teach the model. It contains both features and their corresponding labels.
  • Model: The algorithm or mathematical function learned from the training data that maps features to labels.
  • Prediction: The output generated by the model for new, unseen features.

The Regression Process

Masked worker operating machinery in an industrial factory setting, processing materials.
Photo by Sinan KRIYA on Pexels

Here's how a typical supervised regression project flows:

graph TD
    A["Gather Data (Features & Labels)"] --> B["Split Data (Train/Test)"]
    B --> C["Choose a Model (e.g., Linear Regression)"]
    C --> D["Train Model on Training Data"]
    D --> E["Evaluate Model on Test Data"]
    E --> F["Tune Model / Deploy"]

Linear Regression

Visual representation of geometric calculations comparing bits and qubits in black and white.
Photo by Google DeepMind on Pexels

Linear Regression is one of the simplest and most fundamental regression algorithms. It assumes a linear relationship between the features and the label. For a single feature, it tries to find the best-fit straight line through your data points.

The equation for simple linear regression (one feature) is:

$y = mx + b$

Where:
* $y$ is the predicted label (e.g., house price).
* $x$ is the feature (e.g., square footage).
* $m$ is the slope of the line, representing how much $y$ changes for a unit change in $x$.
* $b$ is the y-intercept, the predicted $y$ when $x$ is 0.

The model "learns" the values of $m$ and $b$ that minimize the difference between its predictions and the actual labels in the training data.

Evaluating Regression Models

Scrabble tiles spelling out 'risk' scattered on a rustic wooden background, symbolizing uncertainty.
Photo by Markus Winkler on Pexels

How do we know if our regression model is good? We use metrics:

  • Mean Absolute Error (MAE): The average absolute difference between predicted and actual values. It's easy to interpret as it's in the same units as the label.
    $MAE = \frac{1}{N} \sum_{i=1}^{N} |y_i - \hat{y}_i|$
  • Mean Squared Error (MSE): The average of the squared differences. It penalizes larger errors more heavily.
    $MSE = \frac{1}{N} \sum_{i=1}^{N} (y_i - \hat{y}_i)^2$
  • Root Mean Squared Error (RMSE): The square root of MSE. It's also in the same units as the label, making it more interpretable than MSE.
    $RMSE = \sqrt{\frac{1}{N} \sum_{i=1}^{N} (y_i - \hat{y}_i)^2}$
  • R-squared ($R^2$): Represents the proportion of the variance in the dependent variable that's predictable from the independent variables. A value of 1 means the model perfectly predicts the label, while 0 means it explains none of the variance.

3. Worked Example

Let's predict exam scores based on hours studied using simple linear regression in Python with scikit-learn.

import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, r2_score

# 1. Gather Data (Features and Labels)
# Hours studied (X)
hours_studied = np.array([2, 3, 4, 5, 6, 7, 8, 9, 10, 1]).reshape(-1, 1)
# Exam scores (y)
exam_scores = np.array([55, 60, 65, 70, 75, 80, 85, 90, 95, 50])

print("--- Original Data ---")
for i in range(len(hours_studied)):
    print(f"Hours: {hours_studied[i][0]}, Score: {exam_scores[i]}")

# 2. Split Data (Train/Test)
# We'll use 80% for training and 20% for testing
X_train, X_test, y_train, y_test = train_test_split(
    hours_studied, exam_scores, test_size=0.2, random_state=42
)

print("\n--- Training Data ---")
print(f"X_train:\n{X_train}")
print(f"y_train:\n{y_train}")
print("\n--- Testing Data ---")
print(f"X_test:\n{X_test}")
print(f"y_test:\n{y_test}")

# 3. Choose a Model (Linear Regression)
model = LinearRegression()

# 4. Train Model on Training Data
model.fit(X_train, y_train)

# Print the learned coefficients
print(f"\nLearned slope (m): {model.coef_[0]:.2f}")
print(f"Learned intercept (b): {model.intercept_:.2f}")

# 5. Make predictions on the test set
y_pred = model.predict(X_test)

print("\n--- Predictions vs. Actual ---")
for i in range(len(X_test)):
    print(f"Hours: {X_test[i][0]}, Actual Score: {y_test[i]}, Predicted Score: {y_pred[i]:.2f}")

# 6. Evaluate Model on Test Data
mae = mean_absolute_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)

print(f"\nMean Absolute Error (MAE): {mae:.2f}")
print(f"R-squared (R2): {r2:.2f}")

# Predict for a new, unseen value (e.g., 6.5 hours studied)
new_hours = np.array([[6.5]])
predicted_score = model.predict(new_hours)
print(f"\nPredicted score for 6.5 hours of study: {predicted_score[0]:.2f}")

In this example, our model learned a slope and intercept. The MAE tells us, on average, our predictions were off by about 3.39 points. An R-squared of 0.99 is very high, indicating our model explains 99% of the variance in exam scores based on hours studied.

4. Key Takeaways

  • Machine learning enables systems to learn from data for predictions and decisions, with supervised learning using labeled data.
  • Regression specifically predicts continuous numerical values, like house prices or temperatures.
  • Features are input variables, and labels are the target values you're trying to predict.
  • The regression process involves data splitting, model training, and evaluation using metrics like MAE, MSE, RMSE, and R-squared.
  • Linear Regression is a foundational algorithm that models a linear relationship between features and labels.
  • Model evaluation metrics help you understand how well your model performs on unseen data.

Common Mistakes to Avoid:

  • Not splitting your data: Always divide your data into training and testing sets to properly evaluate how your model generalizes to new data.
  • Overfitting: Creating a model that's too complex and learns the training data too well, failing to perform well on new data.
  • Ignoring data preprocessing: Real-world data is often messy; cleaning and preparing it is a crucial first step.
  • Using the wrong evaluation metric: Choose metrics appropriate for regression tasks and interpret them correctly.

5. Now Try It

Spend 15 minutes:

  1. Modify the example code: Change the hours_studied and exam_scores data points to create a slightly different linear relationship (e.g., make the scores increase faster or slower with hours).
  2. Run the modified code: Observe how the Learned slope (m) and Learned intercept (b) change.
  3. Analyze the new metrics: How do the MAE and R-squared values change with your new data?
  4. Try a new prediction: Predict an exam score for an hour studied value you didn't include in your initial data.

Success looks like: You've successfully run the code, interpreted the new slope/intercept, and understood how data changes impact the model's performance metrics and predictions.

Frequently asked about Foundations of Machine Learning & Supervised Learning (Regression)

Machine learning lets computers learn patterns from data without explicit programming, making predictions or decisions. Supervised learning, especially regression, trains models on labeled data to predict continuous numerical values. Read the full notes above for the details.

Foundations of Machine Learning & Supervised Learning (Regression) is a core topic in Machine Learning. Most exam papers test it via a mix of definitions, worked examples, and applied problems. The notes above cover the high-yield sub-topics, common pitfalls, and the kind of questions examiners typically set.

Yes — every note in the StudyAI Campus Hub is free to read in full, right here on this page, with no account needed. If you clone the plan into your own dashboard, the free plan shows a preview of each note there; Basic and above unlock the full notes in your dashboard, along with practice quizzes, flashcards and offline study. You can always come back here to read the complete note for free.
Continue with
Supervised Learning (Classification) & Model Evaluation

Study this next


Get the full Machine Learning curriculum

Clone the complete plan to your dashboard for unlimited AI-generated notes, practice quizzes, and a personalised revision schedule.

Create Free Account