Good decisions depend on how data was collected.
Where data comes from shapes what it can honestly tell you.
Data collection is the process of gathering information from sources such as sensors, surveys, forms, apps, websites, and transactions. The way data is collected affects its quality, its potential bias, and the privacy of the people involved.
This topic is part of Big Idea 2: Data. You will learn the difference between structured and unstructured data and why responsible, consent-based collection matters.
Why this matters: Computing innovations often rely on collected data. Understanding bias and privacy helps you evaluate them fairly and build responsibly.
Data moves through a pipeline before it informs a decision.
Source, collection, storage, cleaning, analysis, decision.
Raw data rarely answers a question on its own. It travels through a pipeline: a source produces it, a collection method captures it, it is stored, then cleaned to fix errors and gaps, then analyzed, and finally used to support a decision.
Data also comes in two broad forms. Structured data is organized neatly, such as rows in a table. Unstructured data is messier, such as free-text comments, photos, or audio.
At every stage, watch for problems: biased samples, missing values, poor data quality, and privacy concerns. A clean pipeline with biased input still produces a biased conclusion.
Data flows from source through collection, storage, cleaning, and analysis to a decision.
Trace the arrows from left to right and notice the example sources feeding the start. The note at the bottom is the key idea: privacy, bias, and quality must be considered at every stage, not just at the end. On the exam, weak data early in the pipeline weakens the final decision.
Words for thinking about data quality.
These terms appear in many exam scenarios.
- Data source: where data originates, such as a sensor, survey, or app.
- Structured data: organized data that fits neatly into tables or fields.
- Unstructured data: data without a fixed format, like comments or images.
- Bias: a systematic error that makes data unrepresentative of reality.
- Consent: permission from people for their data to be collected and used.
- Privacy: a person's right to control information about themselves.
Memory hook: Structured data slots into a spreadsheet. Unstructured data does not, like a paragraph of free text.
How data collection is tested.
Bias, privacy, and quality are central.
Be ready to:
- Identify likely sources of data for a scenario.
- Spot sample bias that would make a conclusion unreliable.
- Explain privacy and consent concerns in data collection.
- Distinguish structured from unstructured data.
AP CSP places strong emphasis on the impact of computing on people. Many questions test whether you can reason about responsible data use, not just technical correctness.
Exam tip: If a sample only includes one type of person or situation, suspect bias. Ask "Who or what is missing from this data?"
Improving cafeteria lunch choices.
Find the sources, the privacy concerns, and the bias risks.
A school wants data to improve cafeteria lunches. Let's analyze the collection.
Possible data sources: a student survey, point-of-sale records of what was bought, leftover-waste measurements, and a suggestion form.
Privacy concerns: linking purchases to individual student accounts could reveal personal eating habits. The school should collect only what is needed and protect identities.
Bias risks: if only students who stay for lunch respond to the survey, the data ignores those who bring food from home or skip lunch. A voluntary online survey may also over-represent students who feel strongly. These gaps could push the school toward choices that do not reflect everyone.
A responsible plan uses multiple sources, seeks broad participation to reduce bias, and limits personal data to protect privacy.
Data collection pitfalls.
Avoid these to reason clearly about data.
- Assuming data is always accurate. Collected data can contain errors and gaps.
- Ignoring sample bias. A narrow sample leads to misleading conclusions.
- Collecting more than needed. Extra data raises privacy risk without adding value.
- Forgetting consent and privacy. People should know and agree to how their data is used.
- Treating missing data as meaningless. Gaps can themselves reveal a problem with collection.
Reframe: Before trusting a conclusion, ask how the data was collected, who is represented, and who is left out.
Data sources and forms compared.
Know what each source typically provides.
| Type | What it is | Example |
|---|---|---|
| Sensor data | Measurements from devices | Temperature readings |
| Survey data | Answers people provide | Lunch preference poll |
| Transaction data | Records of purchases or actions | Items bought at checkout |
| App log data | Automatic records of app use | Buttons tapped, time spent |
| Structured data | Organized into fields | A spreadsheet table |
| Unstructured data | No fixed format | Free-text comments |
Run the "who is missing?" check.
It exposes bias fast in any scenario.
Whenever you evaluate a dataset, ask three quick questions: Where did this come from? Who is represented? Who is left out? This routine reveals sample bias and privacy issues, exactly the reasoning AP CSP rewards in data-impact questions.
Try it: A fitness app draws conclusions from users who own smartwatches. Who might be missing, and how could that bias the results?
Practice — attempt these now.
AP-style assessments aligned to this lesson. Time them.
AP CSP Big Idea 2 Topic 4: Data Collection — Set 1
AP-style topic practice assessment
AP CSP Big Idea 2 Topic 4: Data Collection — Set 2
AP-style topic practice assessment
AP CSP Big Idea 2 Topic 4: Data Collection — Set 3
AP-style topic practice assessment
AP CSP Big Idea 2 Topic 4: Data Collection — Set 4
AP-style topic practice assessment