Foundations of Biostatistics: Data and Variables

SA
StudyAI Editorial
Reviewed by StudyAI tutors
· Published Updated

From the bio stat curriculum

Foundations of Biostatistics: Data and Variables

TL;DR

Biostatistics is about making sense of health-related data. Understanding different data types is crucial because it dictates what statistical methods you can use. You'll mostly work with numbers or categories that describe something about your subjects.

1. The Mental Model

Think of data as all the information you collect about people, animals, or cells in a study. Variables are the specific characteristics you measure, like someone's age or whether they have a certain disease. Knowing what kind of variable you have helps you pick the right tools for analysis.

2. The Core Material

When you're doing biostatistics, you're constantly dealing with data. Data is just a collection of observations. For example, if you measure the height of 10 people, those 10 measurements are your data. A variable is a characteristic that can take on different values. In our height example, "height" is the variable.

Understanding the type of variable you have is super important because it determines what you can do with it statistically. There are two main categories: quantitative and qualitative.

Quantitative Variables

Hand analyzing business graphs on a wooden desk, focusing on data results and growth analysis.
Photo by Lukas Blazek on Pexels

These are variables that represent counts or measurements. They are inherently numerical.

  • Discrete: These variables can only take specific, separate values, often whole numbers, and usually result from counting. Think of the number of children in a family, or the number of heart attacks a patient has had. You can't have 2.5 children or 1.3 heart attacks.
  • Continuous: These variables can take any value within a given range. They usually result from measuring. Examples include height, weight, blood pressure, or drug dosage. You can have a height of 170.5 cm or a weight of 65.32 kg.

Qualitative (or Categorical) Variables

Hand analyzing business graphs on a wooden desk, focusing on data results and growth analysis.
Photo by Lukas Blazek on Pexels

These variables represent characteristics or categories. They don't have a numerical meaning that you can add or subtract.

  • Nominal: These are categories without any natural order or ranking. Examples include blood type (A, B, AB, O), gender (Male, Female), or disease presence (Yes, No). You can't say "A" blood type is "better" or "more" than "B".
  • Ordinal: These are categories with a meaningful order or ranking, but the difference between categories isn't necessarily equal or quantifiable. Think of pain levels (Mild, Moderate, Severe), cancer stages (Stage I, Stage II, Stage III), or educational attainment (High School, Bachelor's, Master's, PhD). You know "Severe" is more pain than "Moderate," but the jump from Mild to Moderate isn't necessarily the same "amount" of pain as from Moderate to Severe.

It's common for researchers to assign numbers to qualitative variables (e.g., 1 for Male, 2 for Female). However, these numbers are just labels; you shouldn't perform mathematical operations on them. Averaging "gender" (1+2)/2 = 1.5 doesn't make sense.

Here's a diagram to help visualize the variable types:

graph TD
    A["Variable Types"] --> B["Quantitative (Numerical)"]
    A --> C["Qualitative (Categorical)"]

    B --> D["Discrete (Counts)"]
    B --> E["Continuous (Measurements)"]

    C --> F["Nominal (No Order)"]
    C --> G["Ordinal (Ordered Categories)"]

3. Worked Example

Let's say you're designing a study to look at factors affecting recovery from a certain illness. You decide to collect the following information from each patient:

  1. Patient ID: A unique number assigned to each patient (e.g., P001, P002).
  2. Age: Patient's age in years.
  3. Gender: Male or Female.
  4. Severity of Illness: Rated on admission as Mild, Moderate, or Severe.
  5. Number of Hospital Days: How many full days the patient stayed in the hospital.
  6. Blood Pressure (Systolic): Measured in mmHg.

Here's how you'd classify each variable:

  • Patient ID: This is just a label. While it's a number, it doesn't have numerical meaning you can calculate with. It's a nominal variable, serving as an identifier.
  • Age: You can have 25 years, 25.5 years, etc. This is a measurement. It's a quantitative continuous variable.
  • Gender: Male or Female. There's no order. This is a qualitative nominal variable.
  • Severity of Illness: Mild, Moderate, Severe. There's a clear order. This is a qualitative ordinal variable.
  • Number of Hospital Days: You count full days (1, 2, 3...). You can't have 2.7 days. This is a quantitative discrete variable.
  • Blood Pressure (Systolic): Measured in mmHg (e.g., 120, 122.5). This is a measurement. It's a quantitative continuous variable.

4. Key Takeaways

  • Data are the raw observations you collect, and variables are the specific characteristics you measure.
  • Quantitative variables represent numerical measurements or counts.
  • Qualitative variables represent categories or labels.
  • Discrete variables are counts (whole numbers), while continuous variables are measurements that can take any value within a range.
  • Nominal variables are categories without order; ordinal variables are categories with a meaningful order.
  • Correctly identifying variable types is fundamental for choosing appropriate statistical tests.

Common Mistakes to Avoid:

Flat lay of a spiral notebook and eraser on a pastel pink background with crossed out words.
Photo by KATRIN BOLOVTSOVA on Pexels

  • Confusing nominal numerical identifiers (like patient ID numbers) with true quantitative variables.
  • Treating ordinal variables as if the difference between categories is equal (e.g., assuming the difference between "Mild" and "Moderate" pain is the same as "Moderate" and "Severe").
  • Incorrectly categorizing a continuous variable as discrete, or vice-versa, which can lead to inappropriate analyses.
  • Forgetting that qualitative variables, even if coded with numbers, should not be used in arithmetic calculations.

5. Now Try It

For 15 minutes, imagine you're planning a small survey about people's diet and exercise habits. List 5 different variables you would collect (e.g., "favorite fruit"). For each variable, determine if it's quantitative or qualitative, and then specify if it's discrete, continuous, nominal, or ordinal.

What success looks like: You should have a clear list of 5 variables, each correctly classified down to the specific type (e.g., "Daily steps" -> Quantitative, Discrete).

Frequently asked about Foundations of Biostatistics: Data and Variables

Biostatistics is about making sense of health-related data. Understanding different data types is crucial because it dictates what statistical methods you can use. You'll mostly work with numbers or categories that describe something about your subjects. Read the full notes above for the details.

Foundations of Biostatistics: Data and Variables is a core topic in bio stat. Most exam papers test it via a mix of definitions, worked examples, and applied problems. The notes above cover the high-yield sub-topics, common pitfalls, and the kind of questions examiners typically set.

Yes. Every note in the StudyAI Campus Hub is free to read. Create a free account if you want to clone the full plan, generate your own notes from your textbook, or get AI-powered practice quizzes and flashcards.

Get the full bio stat curriculum

Clone the complete plan to your dashboard for unlimited AI-generated notes, practice quizzes, and a personalised revision schedule.

Create Free Account