Unsupervised Learning & Dimensionality Reduction
From the Machine Learning curriculum
TL;DR
Unsupervised learning finds hidden patterns in data without pre-labeled answers, like grouping similar customers. Dimensionality reduction simplifies complex datasets by reducing the number of features, making them easier to analyze and visualize. Together, these techniques help uncover structure and prepare data for further analysis.
1. The Mental Model
Imagine you have a giant pile of LEGO bricks without instructions; unsupervised learning is like sorting them into similar color piles or finding out which pieces often connect. Dimensionality reduction is like taking a 3D model and making a 2D drawing of it, keeping the important features but making it simpler to look at.
2. The Core Material
In unsupervised learning, you don't have target variables or labels. The goal is to discover inherent structures, patterns, or groupings within the data itself. This is different from supervised learning, where you train a model on input-output pairs.
Common unsupervised tasks include:
* Clustering: Grouping similar data points together. Think of customer segmentation, where you group customers with similar purchasing habits.
* Association: Finding rules that describe large portions of your data, like "people who buy bread also tend to buy milk."
* Dimensionality Reduction: Reducing the number of features (variables) in your dataset while preserving as much important information as possible. This is crucial for visualization, improving model performance, and handling high-dimensional data.
Clustering: K-Means

Photo by Pixabay on Pexels
K-Means is a popular clustering algorithm. Here's how it generally works:
- Choose
k: Decide how many clusters you want (e.g., 3 customer segments). - Initialize centroids: Randomly place
k"centroids" (center points) in your data space. - Assign points: Assign each data point to the closest centroid. This forms
kinitial clusters. - Update centroids: Recalculate the position of each centroid to be the mean (average) of all points currently assigned to its cluster.
- Repeat: Go back to step 3 and re-assign points based on the new centroids. Keep repeating steps 3 and 4 until the centroids no longer move significantly or a maximum number of iterations is reached.
Dimensionality Reduction: Principal Component Analysis (PCA)

Photo by Tima Miroshnichenko on Pexels
PCA is a powerful technique to reduce the number of features in a dataset. It transforms the original features into a new set of uncorrelated features called Principal Components (PCs).
Here's the idea:
* The first principal component (PC1) captures the most variance (spread) in the data.
* The second principal component (PC2) captures the second most variance, and it's orthogonal (at a right angle) to PC1, meaning it's independent.
* This continues for subsequent PCs.
By selecting the top few principal components, you can represent most of the information in your dataset with fewer features, without losing too much important detail. This is especially useful for visualizing high-dimensional data (e.g., reducing 50 features to 2 or 3 for a scatter plot).
Here's a simple visualization of the relationship between these concepts:
graph TD
A["Raw Data (No Labels)"] --> B["Unsupervised Learning"]
B --> C["Clustering (e.g., K-Means)"]
B --> D["Dimensionality Reduction (e.g., PCA)"]
C --> E["Groups / Segments Discovered"]
D --> F["Simplified Feature Set"]
F --> G["Visualization / Further Analysis"]
E --> G
3. Worked Example
Let's imagine you have a dataset of customer purchase habits with 10 features (e.g., "avg_items_per_purchase", "total_spend_last_month", "visits_per_week", etc.). You want to reduce these to 2 principal components for visualization and then cluster them.
Here's a simplified Python example using scikit-learn:
import numpy as np
import matplotlib.pyplot as plt
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.cluster import KMeans
# 1. Generate some dummy data (imagine these are your 10 customer features)
# In a real scenario, you'd load your actual data.
np.random.seed(42)
X = np.random.rand(100, 10) * 100 # 100 customers, 10 features
# 2. Standardize the data (important for PCA and K-Means)
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
# 3. Apply PCA to reduce to 2 components
pca = PCA(n_components=2)
X_pca = pca.fit_transform(X_scaled)
print("Original data shape:", X.shape)
print("Reduced data shape (after PCA):", X_pca.shape)
print("Variance explained by 2 components:", sum(pca.explained_variance_ratio_))
# 4. Apply K-Means clustering on the reduced data
kmeans = KMeans(n_clusters=3, random_state=42, n_init=10) # We want 3 customer segments
clusters = kmeans.fit_predict(X_pca)
# 5. Visualize the clustered data in 2D
plt.figure(figsize=(8, 6))
plt.scatter(X_pca[:, 0], X_pca[:, 1], c=clusters, cmap='viridis', s=50, alpha=0.8)
plt.scatter(kmeans.cluster_centers_[:, 0], kmeans.cluster_centers_[:, 1],
s=200, marker='X', c='red', label='Centroids')
plt.title("Customer Segments after PCA and K-Means")
plt.xlabel("Principal Component 1")
plt.ylabel("Principal Component 2")
plt.legend()
plt.grid(True)
plt.show()
# Now 'clusters' contains the segment ID for each customer.
# You can analyze what each cluster represents by looking at the original features.
This code first scales the data, then uses PCA to transform the 10 features into 2 principal components. Finally, it applies K-Means to these 2 components, identifying 3 distinct customer segments and visualizing them.
4. Key Takeaways
- Unsupervised learning finds inherent patterns in data without relying on pre-labeled outcomes.
- Clustering groups similar data points, like customers, into distinct segments.
- Dimensionality reduction simplifies datasets by reducing the number of features.
- PCA is a key dimensionality reduction technique that creates new, uncorrelated features called principal components.
- Principal components capture variance in descending order, with the first PC explaining the most data spread.
- Reducing dimensions makes data easier to visualize and can improve the performance of other machine learning models.
- Standardizing your data is often crucial before applying PCA or K-Means.
Common mistakes to avoid:
- Not scaling your data before PCA or K-Means, which can lead to features with larger scales dominating the results.
- Choosing an inappropriate number of clusters (k) for K-Means without proper evaluation methods (like the elbow method).
- Interpreting principal components as easily understandable original features; they are often abstract combinations.
- Applying dimensionality reduction unnecessarily when the original feature space is already small and interpretable.
5. Now Try It
Take a small dataset (e.g., the Iris dataset, which is often used for examples but ignore its labels for this exercise, or a dataset you create yourself with 4-5 numeric features). Apply StandardScaler to it, then use PCA to reduce it to 2 components. Visualize these 2 components in a scatter plot. Then, apply KMeans with 3 clusters to the PCA-reduced data and plot the results, coloring points by their assigned cluster. Success looks like a scatter plot showing your data points colored by cluster, demonstrating how PCA helps prepare data for clustering and visualization.
Frequently asked about Unsupervised Learning & Dimensionality Reduction
More from Machine Learning
Get the full Machine Learning curriculum
Clone the complete plan to your dashboard for unlimited AI-generated notes, practice quizzes, and a personalised revision schedule.
Create Free Account