Phase 2: Data Collection and Preparation
From the AI TERM 1 curriculum
TL;DR
This phase is all about getting the right data and making it ready for your AI model. You'll gather data from various sources, clean it up, and transform it into a usable format. High-quality, well-prepared data is crucial for building effective AI.
1. The Mental Model
Think of data collection and preparation like gathering ingredients and prepping them for a complex recipe. You need to find the right ingredients (data), wash and chop them (clean and transform), and measure them accurately (format) before you can start cooking (training your AI).
2. The Core Material
Data collection and preparation is often the most time-consuming part of an AI project, but it's where success begins. Poor data leads to poor models, regardless of how sophisticated your algorithms are.
2.1 Data Collection: Finding Your Ingredients

Photo by Markus Winkler on Pexels
Data collection involves identifying and acquiring the data you'll use to train and test your AI model.
- Internal Data: Data you already own, like customer databases, transaction logs, or sensor readings.
- External Data: Data from outside your organization, such as public datasets (e.g., Kaggle, UCI Machine Learning Repository), web scraping, or third-party data providers.
- Synthetic Data: Data artificially generated to mimic real-world data, often used when real data is scarce or sensitive.
When collecting, consider:
* Relevance: Does the data actually help solve your problem?
* Volume: Do you have enough data?
* Variety: Does it cover different scenarios?
* Velocity: How fast is new data generated (for real-time systems)?
* Veracity: How trustworthy and accurate is the data?
2.2 Data Understanding: Knowing Your Ingredients

Photo by Ron Lach on Pexels
Once you have data, it's vital to understand what you're working with. This involves:
- Exploratory Data Analysis (EDA): Using statistical techniques and visualizations to understand data patterns, identify anomalies, and summarize main characteristics.
- Feature Identification: Determining which columns (features) are most relevant to your prediction or task.
2.3 Data Cleaning: Washing and Chopping

Photo by Gustavo Fring on Pexels
Raw data is rarely perfect. Cleaning addresses issues that can skew your model.
- Handling Missing Values:
- Deletion: Remove rows or columns with too many missing values (use with caution to avoid losing valuable data).
- Imputation: Fill in missing values using strategies like mean, median, mode, or more advanced methods (e.g., regression imputation).
- Dealing with Duplicates: Identify and remove redundant entries that can bias your model.
- Correcting Errors: Fix typos, inconsistencies (e.g., "NY" vs. "New York"), or incorrect data types (e.g., numbers stored as text).
- Outlier Detection and Treatment: Identify data points that are significantly different from others. You might remove them, transform them, or cap them, depending on their nature and impact.
2.4 Data Transformation: Shaping Your Ingredients

Photo by Google DeepMind on Pexels
After cleaning, you often need to transform data into a format suitable for your chosen AI algorithm.
- Feature Scaling:
- Normalization: Scales values to a fixed range, usually 0 to 1 (useful for algorithms sensitive to feature magnitudes, like neural networks).
X_scaled = (X - X_min) / (X_max - X_min) - Standardization: Scales data to have a mean of 0 and a standard deviation of 1 (useful for algorithms assuming normally distributed data, like SVMs or linear regression).
X_scaled = (X - mean) / std_dev
- Normalization: Scales values to a fixed range, usually 0 to 1 (useful for algorithms sensitive to feature magnitudes, like neural networks).
- Encoding Categorical Data: Many AI models require numerical input.
- One-Hot Encoding: Creates new binary columns for each category (e.g., "Red", "Blue", "Green" becomes
[1,0,0],[0,1,0],[0,0,1]). - Label Encoding: Assigns a unique integer to each category (e.g., "Red" = 0, "Blue" = 1, "Green" = 2). Use carefully, as it implies an ordinal relationship that might not exist.
- One-Hot Encoding: Creates new binary columns for each category (e.g., "Red", "Blue", "Green" becomes
- Feature Engineering: Creating new features from existing ones to improve model performance. For example, combining "day" and "month" into "season," or calculating "age" from "date of birth."
Here's a simplified process flow for data collection and preparation:
graph TD
A["Define Problem & Data Needs"] --> B("Identify Data Sources");
B --> C["Collect Raw Data"];
C --> D("Exploratory Data Analysis (EDA)");
D --> E{"Is Data Clean?"};
E -- No --> F("Clean Data: Handle Missing, Duplicates, Errors");
E -- Yes --> G{"Is Data Transformed?"};
F --> G;
G -- No --> H("Transform Data: Scale, Encode, Engineer Features");
G -- Yes --> I("Ready for Model Training");
H --> I;
3. Worked Example
Let's say you're building a model to predict house prices, and you've collected a small dataset.
Raw Data Snippet:
| ID | SqFt | Bedrooms | Bathrooms | Neighborhood | Price |
|----|------|----------|-----------|--------------|-------|
| 1 | 1500 | 3 | 2 | Downtown | 300000|
| 2 | 2000 | 4 | 2.5 | Suburb | 450000|
| 3 | NULL | 2 | 1 | Downtown | 250000|
| 4 | 1800 | 3 | 2 | NULL | 380000|
| 5 | 1500 | 3 | 2 | Downtown | 300000|
| 6 | 2500 | 5 | 30 | Rural | 600000|
Preparation Steps:
- Missing Values:
ID 3:SqFtisNULL. Let's impute with the medianSqFtof available houses (e.g., 1800).ID 4:NeighborhoodisNULL. Let's impute with the mode (most frequent) neighborhood, which is "Downtown".
- Duplicates:
ID 1andID 5are identical rows. RemoveID 5.
- Correcting Errors/Outliers:
ID 6:Bathroomsis30. This is likely a typo. Based on typical house sizes,3.0is more plausible. We'll correct it to3.0(or decide it's an outlier and remove the row if it drastically skews the data).
- Encoding Categorical Data:
Neighborhoodis categorical (Downtown,Suburb,Rural). We'll use One-Hot Encoding.Downtownbecomes[1,0,0]Suburbbecomes[0,1,0]Ruralbecomes[0,0,1]
- Feature Scaling:
SqFtandPricevalues vary widely. We'd standardize these columns to have a mean of 0 and std dev of 1.
Cleaned and Transformed Snippet (Conceptual):
| ID | SqFt_scaled | Bedrooms | Bathrooms | Neighborhood_Downtown | Neighborhood_Suburb | Neighborhood_Rural | Price_scaled |
|----|-------------|----------|-----------|-----------------------|---------------------|--------------------|--------------|
| 1 | -0.5 | 3 | 2.0 | 1 | 0 | 0 | -0.8 |
| 2 | 0.5 | 4 | 2.5 | 0 | 1 | 0 | 0.5 |
| 3 | 0.0 | 2 | 1.0 | 1 | 0 | 0 | -1.2 |
| 4 | 0.0 | 3 | 2.0 | 1 | 0 | 0 | 0.0 |
| 6 | 1.5 | 5 | 3.0 | 0 | 0 | 1 | 1.5 |
(Note: _scaled values are illustrative, actual numbers would depend on the dataset's statistics.)
4. Key Takeaways
- Data quality directly impacts model performance. Bad data in means bad predictions out.
- Data collection should be purposeful, aiming for relevant, diverse, and sufficient data for your problem.
- Exploratory Data Analysis (EDA) is crucial for understanding your data's characteristics and potential issues.
- Data cleaning addresses common imperfections like missing values, duplicates, and errors.
- Data transformation makes data compatible with algorithms through scaling, encoding, and feature engineering.
- Spend sufficient time in this phase, as it sets the foundation for your entire AI project.
Common mistakes to avoid:
- Rushing data preparation: This often leads to subtle errors that are hard to diagnose later.
- Ignoring domain knowledge: Always consult experts about your data to understand context and potential issues.
- Over-imputing missing values: Imputing too many values can introduce bias or artificial patterns.
- Not splitting data before scaling/encoding: Apply transformations to your training data then use the same transformations (e.g., mean/std from training) on your test data to prevent data leakage.
- Using inappropriate encoding for categorical data: Label encoding for non-ordinal features can confuse models.
5. Now Try It
Take a small, publicly available dataset (like the Iris dataset or a simple CSV from Kaggle).
1. Load the data into a Pandas DataFrame.
2. Identify if there are any missing values, duplicates, or obvious errors.
3. Choose one numerical column and one categorical column.
4. Apply a suitable cleaning technique to any issues you found (e.g., impute a missing value, remove duplicates).
5. Apply either standardization or normalization to your chosen numerical column.
6. Apply one-hot encoding to your chosen categorical column.
What success looks like: You should have a DataFrame where the selected columns are cleaned and transformed, and you can explain why you chose those specific techniques.
Frequently asked about Phase 2: Data Collection and Preparation
Study this next
Get the full AI TERM 1 curriculum
Clone the complete plan to your dashboard for unlimited AI-generated notes, practice quizzes, and a personalised revision schedule.
Create Free Account