Phase 2: Data Collection and Preparation

SA
StudyAI Editorial
Reviewed by StudyAI tutors
· Published Updated

From the AI TERM 1 curriculum

TL;DR

This phase is all about getting the right data and making it ready for your AI model. You'll gather data from various sources, clean it up, and transform it into a usable format. High-quality, well-prepared data is crucial for building effective AI.

1. The Mental Model

Think of data collection and preparation like gathering ingredients and prepping them for a complex recipe. You need to find the right ingredients (data), wash and chop them (clean and transform), and measure them accurately (format) before you can start cooking (training your AI).

2. The Core Material

Data collection and preparation is often the most time-consuming part of an AI project, but it's where success begins. Poor data leads to poor models, regardless of how sophisticated your algorithms are.

2.1 Data Collection: Finding Your Ingredients

Wooden letter tiles spelling 'DATA' on a wood textured surface, symbolizing data concepts.
Photo by Markus Winkler on Pexels

Data collection involves identifying and acquiring the data you'll use to train and test your AI model.

  • Internal Data: Data you already own, like customer databases, transaction logs, or sensor readings.
  • External Data: Data from outside your organization, such as public datasets (e.g., Kaggle, UCI Machine Learning Repository), web scraping, or third-party data providers.
  • Synthetic Data: Data artificially generated to mimic real-world data, often used when real data is scarce or sensitive.

When collecting, consider:
* Relevance: Does the data actually help solve your problem?
* Volume: Do you have enough data?
* Variety: Does it cover different scenarios?
* Velocity: How fast is new data generated (for real-time systems)?
* Veracity: How trustworthy and accurate is the data?

2.2 Data Understanding: Knowing Your Ingredients

Top view of petri dishes with samples in a scientific laboratory setting.
Photo by Ron Lach on Pexels

Once you have data, it's vital to understand what you're working with. This involves:

  • Exploratory Data Analysis (EDA): Using statistical techniques and visualizations to understand data patterns, identify anomalies, and summarize main characteristics.
  • Feature Identification: Determining which columns (features) are most relevant to your prediction or task.

2.3 Data Cleaning: Washing and Chopping

Two adults wash and slice fresh vegetables on a kitchen counter.
Photo by Gustavo Fring on Pexels

Raw data is rarely perfect. Cleaning addresses issues that can skew your model.

  • Handling Missing Values:
    • Deletion: Remove rows or columns with too many missing values (use with caution to avoid losing valuable data).
    • Imputation: Fill in missing values using strategies like mean, median, mode, or more advanced methods (e.g., regression imputation).
  • Dealing with Duplicates: Identify and remove redundant entries that can bias your model.
  • Correcting Errors: Fix typos, inconsistencies (e.g., "NY" vs. "New York"), or incorrect data types (e.g., numbers stored as text).
  • Outlier Detection and Treatment: Identify data points that are significantly different from others. You might remove them, transform them, or cap them, depending on their nature and impact.

2.4 Data Transformation: Shaping Your Ingredients

A conceptual image blending technology and nature, symbolizing AI's role in sustainable energy.
Photo by Google DeepMind on Pexels

After cleaning, you often need to transform data into a format suitable for your chosen AI algorithm.

  • Feature Scaling:
    • Normalization: Scales values to a fixed range, usually 0 to 1 (useful for algorithms sensitive to feature magnitudes, like neural networks).
      X_scaled = (X - X_min) / (X_max - X_min)
    • Standardization: Scales data to have a mean of 0 and a standard deviation of 1 (useful for algorithms assuming normally distributed data, like SVMs or linear regression).
      X_scaled = (X - mean) / std_dev
  • Encoding Categorical Data: Many AI models require numerical input.
    • One-Hot Encoding: Creates new binary columns for each category (e.g., "Red", "Blue", "Green" becomes [1,0,0], [0,1,0], [0,0,1]).
    • Label Encoding: Assigns a unique integer to each category (e.g., "Red" = 0, "Blue" = 1, "Green" = 2). Use carefully, as it implies an ordinal relationship that might not exist.
  • Feature Engineering: Creating new features from existing ones to improve model performance. For example, combining "day" and "month" into "season," or calculating "age" from "date of birth."

Here's a simplified process flow for data collection and preparation:

graph TD
    A["Define Problem & Data Needs"] --> B("Identify Data Sources");
    B --> C["Collect Raw Data"];
    C --> D("Exploratory Data Analysis (EDA)");
    D --> E{"Is Data Clean?"};
    E -- No --> F("Clean Data: Handle Missing, Duplicates, Errors");
    E -- Yes --> G{"Is Data Transformed?"};
    F --> G;
    G -- No --> H("Transform Data: Scale, Encode, Engineer Features");
    G -- Yes --> I("Ready for Model Training");
    H --> I;

3. Worked Example

Let's say you're building a model to predict house prices, and you've collected a small dataset.

Raw Data Snippet:

| ID | SqFt | Bedrooms | Bathrooms | Neighborhood | Price |
|----|------|----------|-----------|--------------|-------|
| 1  | 1500 | 3        | 2         | Downtown     | 300000|
| 2  | 2000 | 4        | 2.5       | Suburb       | 450000|
| 3  | NULL | 2        | 1         | Downtown     | 250000|
| 4  | 1800 | 3        | 2         | NULL         | 380000|
| 5  | 1500 | 3        | 2         | Downtown     | 300000|
| 6  | 2500 | 5        | 30        | Rural        | 600000|

Preparation Steps:

  1. Missing Values:
    • ID 3: SqFt is NULL. Let's impute with the median SqFt of available houses (e.g., 1800).
    • ID 4: Neighborhood is NULL. Let's impute with the mode (most frequent) neighborhood, which is "Downtown".
  2. Duplicates:
    • ID 1 and ID 5 are identical rows. Remove ID 5.
  3. Correcting Errors/Outliers:
    • ID 6: Bathrooms is 30. This is likely a typo. Based on typical house sizes, 3.0 is more plausible. We'll correct it to 3.0 (or decide it's an outlier and remove the row if it drastically skews the data).
  4. Encoding Categorical Data:
    • Neighborhood is categorical (Downtown, Suburb, Rural). We'll use One-Hot Encoding.
      • Downtown becomes [1,0,0]
      • Suburb becomes [0,1,0]
      • Rural becomes [0,0,1]
  5. Feature Scaling:
    • SqFt and Price values vary widely. We'd standardize these columns to have a mean of 0 and std dev of 1.

Cleaned and Transformed Snippet (Conceptual):

| ID | SqFt_scaled | Bedrooms | Bathrooms | Neighborhood_Downtown | Neighborhood_Suburb | Neighborhood_Rural | Price_scaled |
|----|-------------|----------|-----------|-----------------------|---------------------|--------------------|--------------|
| 1  | -0.5        | 3        | 2.0       | 1                     | 0                   | 0                  | -0.8         |
| 2  | 0.5         | 4        | 2.5       | 0                     | 1                   | 0                  | 0.5          |
| 3  | 0.0         | 2        | 1.0       | 1                     | 0                   | 0                  | -1.2         |
| 4  | 0.0         | 3        | 2.0       | 1                     | 0                   | 0                  | 0.0          |
| 6  | 1.5         | 5        | 3.0       | 0                     | 0                   | 1                  | 1.5          |

(Note: _scaled values are illustrative, actual numbers would depend on the dataset's statistics.)

4. Key Takeaways

  • Data quality directly impacts model performance. Bad data in means bad predictions out.
  • Data collection should be purposeful, aiming for relevant, diverse, and sufficient data for your problem.
  • Exploratory Data Analysis (EDA) is crucial for understanding your data's characteristics and potential issues.
  • Data cleaning addresses common imperfections like missing values, duplicates, and errors.
  • Data transformation makes data compatible with algorithms through scaling, encoding, and feature engineering.
  • Spend sufficient time in this phase, as it sets the foundation for your entire AI project.

Common mistakes to avoid:
- Rushing data preparation: This often leads to subtle errors that are hard to diagnose later.
- Ignoring domain knowledge: Always consult experts about your data to understand context and potential issues.
- Over-imputing missing values: Imputing too many values can introduce bias or artificial patterns.
- Not splitting data before scaling/encoding: Apply transformations to your training data then use the same transformations (e.g., mean/std from training) on your test data to prevent data leakage.
- Using inappropriate encoding for categorical data: Label encoding for non-ordinal features can confuse models.

5. Now Try It

Take a small, publicly available dataset (like the Iris dataset or a simple CSV from Kaggle).
1. Load the data into a Pandas DataFrame.
2. Identify if there are any missing values, duplicates, or obvious errors.
3. Choose one numerical column and one categorical column.
4. Apply a suitable cleaning technique to any issues you found (e.g., impute a missing value, remove duplicates).
5. Apply either standardization or normalization to your chosen numerical column.
6. Apply one-hot encoding to your chosen categorical column.

What success looks like: You should have a DataFrame where the selected columns are cleaned and transformed, and you can explain why you chose those specific techniques.

Frequently asked about Phase 2: Data Collection and Preparation

This phase is all about getting the right data and making it ready for your AI model. You'll gather data from various sources, clean it up, and transform it into a usable format. High-quality, well-prepared data is crucial for building effective AI. Read the full notes above for the details.

Phase 2: Data Collection and Preparation is a core topic in AI TERM 1. Most exam papers test it via a mix of definitions, worked examples, and applied problems. The notes above cover the high-yield sub-topics, common pitfalls, and the kind of questions examiners typically set.

Yes — every note in the StudyAI Campus Hub is free to read in full, right here on this page, with no account needed. If you clone the plan into your own dashboard, the free plan shows a preview of each note there; Basic and above unlock the full notes in your dashboard, along with practice quizzes, flashcards and offline study. You can always come back here to read the complete note for free.

Study this next


Get the full AI TERM 1 curriculum

Clone the complete plan to your dashboard for unlimited AI-generated notes, practice quizzes, and a personalised revision schedule.

Create Free Account