Data Collection Methods
Hamilton College · Year 12 · 2026
For C02 you must collect data using at least three different methods and explain why you chose each one. This page covers the four methods VCE Software Development recognises — what they are, when each works best, and how to use them in your SAT project.
The C2-1 rubric rewards specific, justified method choice at the 9–10 band. Picking interviews "because everyone uses interviews" will not score. Picking interviews "because I needed depth on three users' workflows that a survey couldn't capture" will.
The big picture

Researchers split data collection into two families:
- Primary — data you collect, first-hand: surveys, interviews, observations, experiments.
- Secondary — data someone else already collected, repurposed for your question: literature reviews, government databases, commercial databases, web data.
For your VCE SD project we use four of these — surveys, interviews, observations, and the secondary group bundled together as reports / existing data. (Experiments are not a standard SAT method — your project is software design, not scientific research.)
The four methods at a glance
| Method | Best for | Data type | Time cost | Save raw to |
|---|---|---|---|---|
| Interview | Depth, motivations, follow-up questions | Mostly qualitative | High (per person) | C02/raw/interview-<name>-<date>.md |
| Survey | Breadth, countable patterns across many people | Mostly quantitative + some qualitative | Low (per response) | C02/raw/survey-results.csv |
| Observation | What people actually do (vs what they say) | Mostly qualitative | Medium | C02/raw/observation-<context>-<date>.md |
| Reports / existing data | Context, constraints, technical environment | Mixed | Low | C02/raw/report-<topic>.md |
A strong project usually combines: 1 method for depth (interview), 1 for breadth (survey), and 1 for grounding (observation or reports).
1. Interview
What it is — A structured or semi-structured conversation between you and one person (your client, a target user, an expert). You ask predetermined questions but follow up on interesting answers.
When to use it. When you need depth — understanding why someone does something, what frustrates them, or what they would value most. Best for the first 1–3 stakeholders, before you know enough to design a survey.
Strengths
- Lets you ask follow-up questions and clarify confusing answers
- Captures rich detail and personal context a survey can't reach
- Adapts on the fly — you can probe an unexpected response
Weaknesses
- Slow — a serious interview is 20–45 minutes per person, plus prep and writing it up
- Your wording, body language, and tone affect what people say (interviewer bias)
- Hard to keep two interviews directly comparable
In your project
- Capture key points and direct quotes, not full transcripts. A short bulleted note file is enough.
- After each interview, write a 2-line summary of "biggest insight" and "thing I didn't expect" — these often become your strongest poster content.
- Plan 2–4 interviews, not 10. The C2-1 rubric rewards depth of analysis, not interview count.
Real-world example
In 2024, university researchers interviewed teachers about how they use ChatGPT in lesson planning. The interviews revealed something a survey would have missed: most teachers used the tool to draft rubrics and generate question variants, not to write lesson content directly. They distrusted AI for content but trusted it for structure — a nuance that only emerged when interviewers asked follow-up questions like "can you walk me through the last time you used it?"
2. Survey
What it is — A short questionnaire (Google Forms, Microsoft Forms, etc.) sent to many people. Mostly closed-ended questions (multiple choice, scales) so you can count results; one or two open-ended for texture.
When to use it. When you need breadth — to test whether something you heard in an interview is widely true, or to count preferences across a target user group.
Strengths
- Cheap and fast per response
- Reaches many people who would never agree to an interview
- Quantitative results are easy to summarise on a poster ("78% of Year 9 students said…")
Weaknesses
- People misread or skip questions you thought were clear (ambiguity = bias)
- You only get answers to the questions you asked — no follow-up
- Self-report; what people say they do is often not what they actually do
In your project
- Keep it short. 5–10 questions, max 5 minutes to complete, or response rate collapses.
- Mix closed (countable) with 1–2 open (texture) questions.
- Pilot it with 1–2 people before sending widely — you will catch ambiguous wording every time.
- Save the raw CSV export, not just a summary screenshot.
Real-world example
Pew Research's 2024 survey of US adults on AI tools reached over 11,000 respondents in a few weeks. They found roughly 1 in 4 adults had used ChatGPT, with usage skewed strongly to younger and more-educated demographics. That kind of demographic spread can only be established by surveys — interviewing 11,000 people would take years.
3. Observation
What it is — Watching real people do the task your software will eventually replace or support. You note what works, what frustrates them, and what they do that contradicts what they told you in interview.
When to use it. When you want to know what people do, not just what they say they do — these are usually different. Especially powerful for workflows people perform so often they no longer notice the friction.
Strengths
- Captures real behaviour in context
- Catches workarounds and unspoken steps people forget to mention
- Cuts through self-report bias — what you observe happened
Weaknesses
- Time-consuming; you usually need multiple sessions to spot patterns
- People behave differently when watched (observer effect)
- You can only see external behaviour, not motivation
In your project
- Participant observation: you join in (e.g. shadow a teacher running the canteen). Non-participant: you watch from the side (e.g. observe a library queue at lunch).
- Note specific moments, not generalities. "At 12:47, a student gave up looking for the book and asked the librarian instead" beats "it took a while to find books".
- Cross-check observations against interview claims. Contradictions are gold for your poster's "explaining why" section.
Real-world example
GitHub's 2023 controlled study observed developers completing the same coding task with and without GitHub Copilot (an AI coding assistant). Developers using Copilot finished about 55% faster on average — a number self-reports would have understated, because developers using AI assistants often feel slower (the tool interrupts their flow) even when measurement shows they're producing more working code.
4. Reports / existing data
What it is — Information that already exists, collected by someone else for some other purpose. Includes school policy documents, government statistics, industry reports, your school's IT device list, academic studies, and existing software documentation.
When to use it. When you need to ground your project in real constraints — the technical environment, the regulatory environment, or population statistics that would take you months to collect yourself.
Strengths
- Free or cheap; you didn't pay to collect it
- Often covers populations or time spans you could never access yourself
- Excellent for the rubric's constraints and technical environment categories — areas students often have nothing to say about
Weaknesses
- Was collected for someone else's question, not yours — may not exactly fit
- Quality varies; you must judge the source
- Can be out of date (especially anything pre-2023)
In your project
- For technical environment: your school's IT device list, browser stats, OS versions, network policy.
- For constraints: school policy on student data, age-appropriate design code (UK ICO), accessibility guidelines (WCAG).
- For user characteristics: ABS census data, government education statistics, your school's enrolment breakdown.
- Always save the source link plus your own one-paragraph summary — examiners want to see you read it, not just cited it.
Sub-types worth knowing
- Literature review — academic papers, textbooks, OER like this wiki
- Government databases — ABS, Department of Education, ACMA
- Industry reports — Statista, Gartner, vendor white papers
- Web data — public APIs, open datasets (data.gov.au)
Choosing your mix
For C02 you must use three or more methods. A balanced mix:
Recommended pattern
- 1 interview with your client or primary user → depth, motivation, must-have features
- 1 survey of target users → breadth, validation of interview claims
- 1 observation OR report set → grounding in real behaviour or constraints
Avoid these traps
- Three interviews and nothing else — no breadth, no constraints, no observed behaviour. Caps at 6–7 on C2-1.
- A survey of one class with no interviews — no depth, no rationale beyond "it was easy".
- Reports only — you've researched the area but haven't engaged with real users; the rubric's user characteristics section will be weak.
The 9–10 descriptor expects you to explain why each method was chosen for its specific job — what data it would yield that the others wouldn't. Practise that explanation out loud before your C2-1 station defence.
Check Your Understanding
Answer in your head first, then click the spoiler to check.
1. Which of these is a secondary data collection method? (a) interview · (b) survey · (c) observation · (d) report
(d) report. Secondary methods reuse data someone else already collected — that's what reports / existing data covers. Interviews, surveys, and observations are all primary — you collect the data first-hand.
2. You want to know whether your survey finding "most students prefer dark mode" is real or just a fluke. Which method would best triangulate that claim? (a) another survey with the same question · (b) an interview asking why · (c) observation of students using the system in their usual environment · (d) a report on UI trends
(c) observation. Triangulation = checking the claim against a different type of evidence. Re-asking in another survey reproduces the same self-report bias. Observation tests whether stated preference matches actual behaviour — the strongest cross-check.
3. A classmate's data collection plan is: "three interviews with my client". According to this page, what's the biggest risk? (a) too few people for breadth · (b) interviewer bias · (c) qualitative-only data · (d) all of the above
(d) all of the above. Three interviews with one stakeholder gives no breadth (only one person's perspective), introduces interviewer bias unchecked, and produces only qualitative data. The rubric explicitly looks for a justified mix; this caps at 6–7 on C2-1.
4. You collect 50 survey responses to the question "how often do you forget to bring your laptop?" with options Never / Rarely / Sometimes / Often / Always. What kind of data is this? (a) qualitative · (b) quantitative · (c) both · (d) neither
(b) quantitative. Even though the topic feels qualitative ("how often" is a subjective judgement), the data type is quantitative — closed-ended responses you can count and compare. Data type is determined by what you do with it, not what it's about.
5. For your project's technical environment section of the SRS, which method is usually most efficient? (a) interview the IT coordinator · (b) survey students about their devices · (c) observation in the computer lab · (d) request the school's existing IT device list (a report)
(d) request the school's IT device list. That data already exists — your IT staff maintain it. Surveying or interviewing recreates work that's already done. (a) is a useful follow-up if the list is unclear, but the report is your starting point.
See also
- Essential Terms — definitions of qualitative, quantitative, open-ended, close-ended, and the four method names
- Explaining vs Describing Your Data Collection — the verb ladder the C2-1 rubric uses to separate 7 from 9–10
- Context Diagram — once you have data, the entities you discovered go on this diagram
- What Is an Entity — your data sources may themselves be entities (e.g. a payment gateway, an LMS)
Adapted for VCE SD students from EDUCBA's Data Collection Methods overview, with project-specific framing for C02.
