Deep Dive
From the Introduction to AI for Students curriculum
TL;DR
You're going to explore how AI learns from data, focusing on neural networks and how they adjust to get better. We'll cover the basic setup of these networks and the math behind how they improve. You'll also learn how to apply these ideas in a simple, practical way.
1. The Mental Model
Imagine you're teaching a child to recognize a cat. At first, they might guess wrong often. You correct them, explaining what makes a cat a cat, and over time, they get better. Deep learning models learn in a very similar, iterative way.
2. The Core Material
Deep learning is a part of machine learning that uses structures called neural networks. These networks are inspired by the human brain and are designed to recognize patterns in data.
What's a Neural Network?

Photo by Google DeepMind on Pexels
A neural network is made of layers of interconnected "neurons" (or nodes). Each neuron takes some inputs, performs a simple calculation, and then passes its output to other neurons.
- Input Layer: This is where your data enters the network. Each neuron here represents a feature of your input (e.g., a pixel in an image, a word in a sentence).
- Hidden Layers: These are the "thinking" layers between the input and output. A network can have one or many hidden layers – "deep" refers to having many.
- Output Layer: This layer gives you the network's final prediction or classification.
Each connection between neurons has a weight, which determines how important that connection is. Each neuron also has a bias, which shifts its output. Learning in a neural network is all about adjusting these weights and biases.
How Does a Neural Network Learn?

Photo by Google DeepMind on Pexels
The learning process typically involves these steps:
- Forward Pass: Input data goes through the network, layer by layer, until it produces an output.
- Loss Calculation: This output is compared to the correct answer (the "ground truth"). The difference between the predicted output and the actual output is called the loss. A high loss means the network is making big mistakes.
- Backward Pass (Backpropagation): The network figures out how much each weight and bias contributed to the loss. It then adjusts these weights and biases slightly to reduce the loss for the next round. This is the gradient descent algorithm in action – it's like finding the bottom of a valley by taking small steps downhill.
- Iteration: Steps 1-3 are repeated many times, with different batches of data, until the network's loss is minimized and it makes accurate predictions.
Activation Functions

Photo by Ann H on Pexels
Neurons don't just pass numbers directly; they often apply an activation function to their sum of inputs. This function introduces non-linearity, which is crucial for the network to learn complex patterns. Without activation functions, a deep network would just be a linear model, no matter how many layers it had. Common ones include:
- ReLU (Rectified Linear Unit):
f(x) = max(0, x)– simply outputs the input if positive, otherwise zero. Very popular. - Sigmoid:
f(x) = 1 / (1 + e^-x)– squashes values between 0 and 1, useful for probabilities. - Softmax: Used in the output layer for multi-class classification, converting values into probabilities that sum to 1.
graph TD
A[Input Features] --> B{Weighted Sum + Bias}
B --> C{Activation Function}
C --> D[Output to Next Layer]
style A fill:#f9f,stroke:#333,stroke-width:2px
style D fill:#f9f,stroke:#333,stroke-width:2px
style B fill:#bbf,stroke:#333,stroke-width:2px
style C fill:#ccf,stroke:#333,stroke-width:2px
Optimizers
Optimizers are algorithms that update the weights and biases based on the calculated gradients during backpropagation. They determine how the network "learns."
- Stochastic Gradient Descent (SGD): A basic optimizer that updates weights after processing each small batch of data.
- Adam: A more advanced optimizer that generally converges faster and often performs better. It adapts the learning rate for each weight individually.
Hyperparameters
These are settings you choose before training the network. They aren't learned by the network itself.
- Learning Rate: How big of a step the optimizer takes when adjusting weights. Too high, and you might overshoot the optimal solution; too low, and training takes forever.
- Number of Layers/Neurons: The architecture of your network.
- Batch Size: How many data samples are processed before weights are updated.
- Epochs: How many times the entire dataset is passed through the network during training.
3. Worked Example
Let's imagine a tiny network with one input, one neuron, and one output, trying to learn to multiply by 2.
Suppose your data is:
- Input x = 3, Correct output y = 6
- Input x = 5, Correct output y = 10
Let's start with a random weight w = 0.5 and bias b = 0.
Our neuron's calculation: output = (x * w) + b. For simplicity, we'll skip an activation function here.
Training with x = 3, y = 6:
- Forward Pass:
predicted_output = (3 * 0.5) + 0 = 1.5 - Loss Calculation (using simple squared error):
loss = (predicted_output - actual_output)^2 = (1.5 - 6)^2 = (-4.5)^2 = 20.25
This is a high loss, meaning we're far off! - Backward Pass (adjusting
wto reduce loss):
The goal is to changewso thatpredicted_outputgets closer to6. Ifpredicted_outputis too low (1.5 vs 6),wneeds to increase.
Let's say our optimizer tells us to increasewby0.1(this is a simplified gradient step).
new_w = 0.5 + 0.1 = 0.6
Training with x = 3, y = 6 again (after update):
- Forward Pass:
predicted_output = (3 * 0.6) + 0 = 1.8 - Loss Calculation:
loss = (1.8 - 6)^2 = (-4.2)^2 = 17.64
The loss is now17.64, which is less than20.25. We're getting closer! If we keep doing this with many more examples and small adjustments,wwould eventually get close to2.
4. Key Takeaways
- Deep learning uses neural networks with interconnected layers of neurons to learn patterns.
- Learning involves adjusting weights and biases in the network through a process of prediction, error calculation, and adjustment.
- The forward pass makes a prediction, and the backward pass (backpropagation) corrects the network's parameters.
- Activation functions are essential for neural networks to learn complex, non-linear relationships.
- Optimizers guide how weights and biases are updated during training, with Adam being a popular choice.
- Hyperparameters like learning rate and network size are set before training and greatly influence performance.
Common Mistakes to Avoid:
- Ignoring the learning rate: Setting it too high can prevent the model from learning; too low makes training painfully slow.
- Not using activation functions (or using only linear ones): This limits the network's ability to learn complex patterns.
- Overfitting: When the model learns the training data too well, including its noise, and performs poorly on new, unseen data.
- Underfitting: When the model is too simple or hasn't trained enough to capture the underlying patterns in the data.
5. Now Try It
Pick a simple dataset, like the Iris dataset (classifying flowers based on measurements), or a basic regression problem. Using a library like Keras or TensorFlow, build a small neural network with 1-2 hidden layers. Experiment with changing the learning rate, the number of neurons in your hidden layers, and different activation functions (e.g., ReLU vs. Sigmoid) to see how they impact your model's performance (e.g., accuracy for classification or mean squared error for regression). What does a learning rate that's too high do to your loss during training?
Frequently asked about Deep Dive
More from Introduction to AI for Students
Get the full Introduction to AI for Students curriculum
Clone the complete plan to your dashboard for unlimited AI-generated notes, practice quizzes, and a personalised revision schedule.
Create Free Account