Introduction to Reinforcement Learning & Q-Learning Fundamentals
From the https://www.youtube.com/watch?v=-Bqx2BuFjik curriculum
TL;DR
Reinforcement Learning (RL) teaches an "agent" to make decisions in an environment by trial and error, aiming to maximize a cumulative reward. Q-Learning is a specific RL algorithm that helps the agent learn the "value" of taking an action in a given state. This value, or Q-value, guides the agent to choose optimal actions over time.
1. The Mental Model
Imagine you're training a dog: it tries different behaviors, and you give it a treat for good ones and nothing for bad ones. Over time, the dog learns which behaviors lead to treats. Reinforcement Learning works similarly, where an "agent" learns to make good decisions by getting "rewards" or "penalties" from its "environment."
2. The Core Material
Reinforcement Learning (RL) is a machine learning paradigm where an agent learns how to behave in an environment by performing actions and receiving rewards or penalties. The agent's goal is to learn a policy—a strategy that tells it what action to take in any given state—to maximize its total cumulative reward.
Here are the key components:
* Agent: The learner or decision-maker.
* Environment: The world the agent interacts with.
* State (S): A snapshot of the environment at a specific time.
* Action (A): What the agent can do in a given state.
* Reward (R): A numerical feedback from the environment after an action. Positive for good, negative for bad.
* Policy ($\pi$): A mapping from states to actions, telling the agent what to do.
* Q-Value (Q(S, A)): Represents the expected future reward for taking action A in state S, and then following an optimal policy thereafter.
The RL Loop

Photo by Nothing Ahead on Pexels
The process unfolds as a continuous loop:
1. The agent observes the current state (S) of the environment.
2. Based on its policy, the agent chooses an action (A).
3. The environment transitions to a new state (S') and gives the agent a reward (R).
4. The agent updates its policy or value estimates based on this experience.
5. Repeat from step 1.
graph TD
Agent["Agent"] -->|Chooses Action (A)| Environment["Environment"]
Environment -->|New State (S') & Reward (R)| Agent
Agent -->|Updates Policy/Values| Agent
Q-Learning
Q-Learning is a model-free RL algorithm. "Model-free" means the agent doesn't need to know how the environment works (e.g., what state an action will lead to, or what reward it will give) beforehand. It learns by interacting with the environment.
The core idea of Q-Learning is to learn the Q-values for all possible (state, action) pairs. These Q-values are typically stored in a table called the Q-table.
The Q-value update rule is crucial:
$Q(S, A) \leftarrow Q(S, A) + \alpha [R + \gamma \max_{A'} Q(S', A') - Q(S, A)]$
Let's break down this equation:
* $Q(S, A)$: The current estimated Q-value for taking action A in state S.
* $\alpha$ (alpha): The learning rate (0 to 1). How much new information overrides old information. A high $\alpha$ means the agent learns quickly but might be unstable; a low $\alpha$ means slower, more stable learning.
* $R$: The immediate reward received after taking action A in state S and transitioning to state S'.
* $\gamma$ (gamma): The discount factor (0 to 1). How much future rewards are valued. A $\gamma$ close to 1 means the agent cares a lot about long-term rewards; a $\gamma$ close to 0 means it focuses more on immediate rewards.
* $\max_{A'} Q(S', A')$: The maximum Q-value for the next state (S'). This represents the optimal expected future reward from the next state.
* $[R + \gamma \max_{A'} Q(S', A') - Q(S, A)]$: This is the temporal difference (TD) error. It's the difference between the new, more informed estimate of the Q-value (the "target") and the current estimate. The agent uses this error to update its knowledge.
Exploration vs. Exploitation

Photo by Volker Braun on Pexels
For the agent to learn effectively, it needs to balance:
* Exploration: Trying new actions to discover potentially better rewards.
* Exploitation: Choosing actions that have yielded the highest rewards in the past.
A common strategy is $\epsilon$-greedy (epsilon-greedy). With probability $\epsilon$ (epsilon), the agent chooses a random action (exploration). With probability $1 - \epsilon$, it chooses the action with the highest Q-value for the current state (exploitation). $\epsilon$ usually starts high and decreases over time, allowing for more exploration early on and more exploitation later.
3. Worked Example
Let's consider a simple 2x2 grid world. The agent starts at (0,0), wants to reach (1,1) for a reward of +10, and avoids (0,1) which gives -10. All other moves give -1.
States: (0,0), (0,1), (1,0), (1,1)
Actions: Up, Down, Left, Right (if legal, otherwise stay put)
Initial Q-table (all zeros):
| State | Action: Up | Action: Down | Action: Left | Action: Right |
|---|---|---|---|---|
| (0,0) | 0 | 0 | 0 | 0 |
| (0,1) | 0 | 0 | 0 | 0 |
| (1,0) | 0 | 0 | 0 | 0 |
| (1,1) | 0 | 0 | 0 | 0 |
Let's trace one learning step.
Assume: $\alpha = 0.1$, $\gamma = 0.9$
Episode 1, Step 1:
* Current state: S = (0,0)
* Agent chooses action: Right (due to exploration, let's say)
* Environment: Agent moves to (0,1), receives Reward R = -10 (penalty!).
* New state: S' = (0,1)
Now, we update $Q((0,0), \text{Right})$:
1. Find $\max_{A'} Q(S', A')$, which is $\max_{A'} Q((0,1), A')$. Since all Q-values are 0 initially, this is 0.
2. Apply the update rule:
$Q((0,0), \text{Right}) \leftarrow Q((0,0), \text{Right}) + \alpha [R + \gamma \max_{A'} Q((0,1), A') - Q((0,0), \text{Right})]$
$Q((0,0), \text{Right}) \leftarrow 0 + 0.1 [-10 + 0.9 * 0 - 0]$
$Q((0,0), \text{Right}) \leftarrow 0.1 * (-10)$
$Q((0,0), \text{Right}) \leftarrow -1$
Now the Q-table entry for $Q((0,0), \text{Right})$ is -1. This means the agent has learned that going right from (0,0) is a bad idea. Over many episodes, the agent will fill out this table, learning that moves towards (1,1) are good, and moves towards (0,1) are bad.
4. Key Takeaways
- Reinforcement Learning involves an agent learning optimal actions through trial and error in an environment to maximize cumulative reward.
- The agent observes a state, takes an action, receives a reward, and transitions to a new state.
- Q-Learning is a model-free algorithm that learns the value (Q-value) of taking an action in a particular state.
- The Q-value update rule uses immediate rewards and discounted future rewards to refine the agent's understanding.
- The learning rate ($\alpha$) controls how quickly new information updates old Q-values.
- The discount factor ($\gamma$) determines how much the agent values future rewards versus immediate rewards.
- Balancing exploration (trying new things) and exploitation (using known good actions) is crucial for effective learning, often managed by an $\epsilon$-greedy strategy.
Common Mistakes to Avoid

Photo by KATRIN BOLOVTSOVA on Pexels
- Not setting appropriate learning rates ($\alpha$) or discount factors ($\gamma$), leading to unstable or slow learning.
- Insufficient exploration, which can cause the agent to settle on sub-optimal strategies without discovering better paths.
- Not handling terminal states or rewards correctly, which can confuse the agent about episode boundaries.
- Forgetting that Q-values are expected future rewards, so a single bad outcome doesn't necessarily mean a Q-value will plummet if other paths are good.
5. Now Try It
Think of a simple game like Tic-Tac-Toe.
1. Define the states: How would you represent the board for an RL agent? (e.g., a 9-element tuple or string).
2. Define the actions: What actions can the agent take in any given state? (e.g., placing a mark in an empty cell).
3. Define the rewards: What reward would you give for winning, losing, or drawing? What about illegal moves or just playing a turn?
4. Imagine the Q-table: If you were to start a Q-table for this game, what would its dimensions be? (Don't build it, just think about the size and structure).
Success looks like you being able to clearly articulate how each of these components would be defined for Tic-Tac-Toe, and understanding why each is necessary for an RL agent to learn to play.
Frequently asked about Introduction to Reinforcement Learning & Q-Learning Fundamentals
Study this next
Get the full https://www.youtube.com/watch?v=-Bqx2BuFjik curriculum
Clone the complete plan to your dashboard for unlimited AI-generated notes, practice quizzes, and a personalised revision schedule.
Save this course free