Blame
|
1 | # Data Collection Methods |
||||||
| 2 | ||||||||
| 3 | Hamilton College · Year 12 · 2026 |
|||||||
| 4 | ||||||||
| 5 | For C02 you must collect data using **at least three different methods** and explain *why* you chose each one. This page covers the four methods VCE Software Development recognises — what they are, when each works best, and how to use them in your SAT project. |
|||||||
| 6 | ||||||||
| 7 | The C2-1 rubric rewards **specific, justified method choice** at the 9–10 band. Picking interviews "because everyone uses interviews" will not score. Picking interviews "because I needed depth on three users' workflows that a survey couldn't capture" will. |
|||||||
| 8 | ||||||||
| 9 | --- |
|||||||
| 10 | ||||||||
|
11 | ## The big picture |
||||||
| 12 | ||||||||
| 13 |  |
|||||||
| 14 | ||||||||
| 15 | Researchers split data collection into two families: |
|||||||
| 16 | ||||||||
| 17 | - **Primary** — data *you* collect, first-hand: surveys, interviews, observations, experiments. |
|||||||
| 18 | - **Secondary** — data *someone else already collected*, repurposed for your question: literature reviews, government databases, commercial databases, web data. |
|||||||
| 19 | ||||||||
| 20 | For your VCE SD project we use **four** of these — surveys, interviews, observations, and the secondary group bundled together as **reports / existing data**. (Experiments are not a standard SAT method — your project is software design, not scientific research.) |
|||||||
| 21 | ||||||||
| 22 | --- |
|||||||
| 23 | ||||||||
|
24 | ## The four methods at a glance |
||||||
| 25 | ||||||||
| 26 | | Method | Best for | Data type | Time cost | Save raw to | |
|||||||
| 27 | |---|---|---|---|---| |
|||||||
| 28 | | **Interview** | Depth, motivations, follow-up questions | Mostly qualitative | High (per person) | `C02/raw/interview-<name>-<date>.md` | |
|||||||
| 29 | | **Survey** | Breadth, countable patterns across many people | Mostly quantitative + some qualitative | Low (per response) | `C02/raw/survey-results.csv` | |
|||||||
| 30 | | **Observation** | What people *actually* do (vs what they say) | Mostly qualitative | Medium | `C02/raw/observation-<context>-<date>.md` | |
|||||||
| 31 | | **Reports / existing data** | Context, constraints, technical environment | Mixed | Low | `C02/raw/report-<topic>.md` | |
|||||||
| 32 | ||||||||
| 33 | A strong project usually combines: **1 method for depth** (interview), **1 for breadth** (survey), and **1 for grounding** (observation or reports). |
|||||||
| 34 | ||||||||
| 35 | --- |
|||||||
| 36 | ||||||||
| 37 | ## 1. Interview |
|||||||
| 38 | ||||||||
| 39 | ::: info |
|||||||
| 40 | **What it is** — A structured or semi-structured conversation between you and one person (your client, a target user, an expert). You ask predetermined questions but follow up on interesting answers. |
|||||||
| 41 | ::: |
|||||||
| 42 | ||||||||
| 43 | **When to use it.** When you need *depth* — understanding why someone does something, what frustrates them, or what they would value most. Best for the first 1–3 stakeholders, before you know enough to design a survey. |
|||||||
| 44 | ||||||||
| 45 | **Strengths** |
|||||||
| 46 | ||||||||
| 47 | - Lets you ask follow-up questions and clarify confusing answers |
|||||||
| 48 | - Captures rich detail and personal context a survey can't reach |
|||||||
| 49 | - Adapts on the fly — you can probe an unexpected response |
|||||||
| 50 | ||||||||
| 51 | **Weaknesses** |
|||||||
| 52 | ||||||||
| 53 | - Slow — a serious interview is 20–45 minutes per person, plus prep and writing it up |
|||||||
| 54 | - Your wording, body language, and tone affect what people say (interviewer bias) |
|||||||
| 55 | - Hard to keep two interviews directly comparable |
|||||||
| 56 | ||||||||
| 57 | **In your project** |
|||||||
| 58 | ||||||||
| 59 | - Capture **key points and direct quotes**, not full transcripts. A short bulleted note file is enough. |
|||||||
| 60 | - After each interview, write a 2-line summary of "biggest insight" and "thing I didn't expect" — these often become your strongest poster content. |
|||||||
| 61 | - Plan **2–4 interviews**, not 10. The C2-1 rubric rewards depth of analysis, not interview count. |
|||||||
| 62 | ||||||||
| 63 | >| ### Real-world example |
|||||||
|
64 | >| In 2024, university researchers interviewed teachers about how they use ChatGPT in lesson planning. The interviews revealed something a survey would have missed: most teachers used the tool to **draft rubrics and generate question variants**, not to write lesson content directly. They distrusted AI for content but trusted it for structure — a nuance that only emerged when interviewers asked follow-up questions like *"can you walk me through the last time you used it?"* |
||||||
|
65 | |||||||
| 66 | --- |
|||||||
| 67 | ||||||||
| 68 | ## 2. Survey |
|||||||
| 69 | ||||||||
| 70 | ::: info |
|||||||
| 71 | **What it is** — A short questionnaire (Google Forms, Microsoft Forms, etc.) sent to many people. Mostly **closed-ended** questions (multiple choice, scales) so you can count results; one or two **open-ended** for texture. |
|||||||
| 72 | ::: |
|||||||
| 73 | ||||||||
| 74 | **When to use it.** When you need *breadth* — to test whether something you heard in an interview is widely true, or to count preferences across a target user group. |
|||||||
| 75 | ||||||||
| 76 | **Strengths** |
|||||||
| 77 | ||||||||
| 78 | - Cheap and fast per response |
|||||||
| 79 | - Reaches many people who would never agree to an interview |
|||||||
| 80 | - Quantitative results are easy to summarise on a poster ("78% of Year 9 students said…") |
|||||||
| 81 | ||||||||
| 82 | **Weaknesses** |
|||||||
| 83 | ||||||||
| 84 | - People misread or skip questions you thought were clear (ambiguity = bias) |
|||||||
| 85 | - You only get answers to the questions you asked — no follow-up |
|||||||
| 86 | - Self-report; what people *say* they do is often not what they *actually* do |
|||||||
| 87 | ||||||||
| 88 | **In your project** |
|||||||
| 89 | ||||||||
| 90 | - Keep it short. **5–10 questions**, max 5 minutes to complete, or response rate collapses. |
|||||||
| 91 | - Mix **closed** (countable) with **1–2 open** (texture) questions. |
|||||||
| 92 | - Pilot it with 1–2 people before sending widely — you will catch ambiguous wording every time. |
|||||||
| 93 | - Save the **raw CSV export**, not just a summary screenshot. |
|||||||
| 94 | ||||||||
| 95 | >| ### Real-world example |
|||||||
|
96 | >| Pew Research's 2024 survey of US adults on AI tools reached over 11,000 respondents in a few weeks. They found roughly **1 in 4 adults had used ChatGPT**, with usage skewed strongly to younger and more-educated demographics. That kind of demographic spread can only be established by surveys — interviewing 11,000 people would take years. |
||||||
|
97 | |||||||
| 98 | --- |
|||||||
| 99 | ||||||||
| 100 | ## 3. Observation |
|||||||
| 101 | ||||||||
| 102 | ::: info |
|||||||
| 103 | **What it is** — Watching real people do the task your software will eventually replace or support. You note what works, what frustrates them, and what they do that contradicts what they told you in interview. |
|||||||
| 104 | ::: |
|||||||
| 105 | ||||||||
| 106 | **When to use it.** When you want to know what people *do*, not just what they *say* they do — these are usually different. Especially powerful for workflows people perform so often they no longer notice the friction. |
|||||||
| 107 | ||||||||
| 108 | **Strengths** |
|||||||
| 109 | ||||||||
| 110 | - Captures real behaviour in context |
|||||||
| 111 | - Catches workarounds and unspoken steps people forget to mention |
|||||||
| 112 | - Cuts through self-report bias — what you observe happened |
|||||||
| 113 | ||||||||
| 114 | **Weaknesses** |
|||||||
| 115 | ||||||||
| 116 | - Time-consuming; you usually need multiple sessions to spot patterns |
|||||||
| 117 | - People behave differently when watched (observer effect) |
|||||||
| 118 | - You can only see external behaviour, not motivation |
|||||||
| 119 | ||||||||
| 120 | **In your project** |
|||||||
| 121 | ||||||||
| 122 | - **Participant** observation: you join in (e.g. shadow a teacher running the canteen). **Non-participant**: you watch from the side (e.g. observe a library queue at lunch). |
|||||||
| 123 | - Note **specific moments**, not generalities. "*At 12:47, a student gave up looking for the book and asked the librarian instead*" beats "*it took a while to find books*". |
|||||||
| 124 | - Cross-check observations against interview claims. Contradictions are gold for your poster's *"explaining why"* section. |
|||||||
| 125 | ||||||||
| 126 | >| ### Real-world example |
|||||||
|
127 | >| GitHub's 2023 controlled study observed developers completing the same coding task with and without GitHub Copilot (an AI coding assistant). Developers using Copilot finished **about 55% faster** on average — a number self-reports would have understated, because developers using AI assistants often *feel* slower (the tool interrupts their flow) even when measurement shows they're producing more working code. |
||||||
|
128 | |||||||
| 129 | --- |
|||||||
| 130 | ||||||||
| 131 | ## 4. Reports / existing data |
|||||||
| 132 | ||||||||
| 133 | ::: info |
|||||||
| 134 | **What it is** — Information that already exists, collected by someone else for some other purpose. Includes school policy documents, government statistics, industry reports, your school's IT device list, academic studies, and existing software documentation. |
|||||||
| 135 | ::: |
|||||||
| 136 | ||||||||
| 137 | **When to use it.** When you need to ground your project in **real constraints** — the technical environment, the regulatory environment, or population statistics that would take you months to collect yourself. |
|||||||
| 138 | ||||||||
| 139 | **Strengths** |
|||||||
| 140 | ||||||||
| 141 | - Free or cheap; you didn't pay to collect it |
|||||||
| 142 | - Often covers populations or time spans you could never access yourself |
|||||||
| 143 | - Excellent for the rubric's *constraints* and *technical environment* categories — areas students often have nothing to say about |
|||||||
| 144 | ||||||||
| 145 | **Weaknesses** |
|||||||
| 146 | ||||||||
| 147 | - Was collected for someone else's question, not yours — may not exactly fit |
|||||||
| 148 | - Quality varies; you must judge the source |
|||||||
| 149 | - Can be out of date (especially anything pre-2023) |
|||||||
| 150 | ||||||||
| 151 | **In your project** |
|||||||
| 152 | ||||||||
| 153 | - For *technical environment*: your school's IT device list, browser stats, OS versions, network policy. |
|||||||
| 154 | - For *constraints*: school policy on student data, age-appropriate design code (UK ICO), accessibility guidelines (WCAG). |
|||||||
| 155 | - For *user characteristics*: ABS census data, government education statistics, your school's enrolment breakdown. |
|||||||
| 156 | - Always save **the source link plus your own one-paragraph summary** — examiners want to see you read it, not just cited it. |
|||||||
| 157 | ||||||||
| 158 | **Sub-types worth knowing** |
|||||||
| 159 | ||||||||
| 160 | - **Literature review** — academic papers, textbooks, OER like this wiki |
|||||||
| 161 | - **Government databases** — ABS, Department of Education, ACMA |
|||||||
| 162 | - **Industry reports** — Statista, Gartner, vendor white papers |
|||||||
| 163 | - **Web data** — public APIs, open datasets (data.gov.au) |
|||||||
| 164 | ||||||||
| 165 | --- |
|||||||
| 166 | ||||||||
|
167 | ## Qualitative or quantitative? |
||||||
| 168 | ||||||||
| 169 | Every data point you collect is either **qualitative** (descriptive — words, themes, observations) or **quantitative** (countable — numbers, frequencies, ratings). Knowing which is which matters for your poster's labelling section and for explaining why you chose each method. |
|||||||
| 170 | ||||||||
| 171 | | Same project, two data points | Type | How you'd present it | |
|||||||
| 172 | |---|---|---| |
|||||||
| 173 | | *"Year 9 students said the canteen queue is too slow at lunch."* | **Qualitative** | Direct quote, theme tag (*pain points*), source attribution | |
|||||||
| 174 | | *"On average, students spend 6.5 minutes in the canteen queue (n=30 observations)."* | **Quantitative** | Number, sample size, summary statistic | |
|||||||
| 175 | ||||||||
| 176 | ::: info |
|||||||
| 177 | **The decision rule:** the *data type* is determined by what you can do with the result, not by how it was collected. Surveys can produce both: a closed-ended scale gives quantitative data; an open-ended response gives qualitative. The same goes for interviews and observations. |
|||||||
| 178 | ::: |
|||||||
| 179 | ||||||||
| 180 | --- |
|||||||
| 181 | ||||||||
|
182 | ## Choosing your mix |
||||||
| 183 | ||||||||
| 184 | For C02 you must use **three or more** methods. A balanced mix: |
|||||||
| 185 | ||||||||
| 186 | ::: success |
|||||||
| 187 | **Recommended pattern** |
|||||||
| 188 | ||||||||
| 189 | - **1 interview** with your client or primary user → depth, motivation, must-have features |
|||||||
| 190 | - **1 survey** of target users → breadth, validation of interview claims |
|||||||
| 191 | - **1 observation OR report set** → grounding in real behaviour or constraints |
|||||||
| 192 | ::: |
|||||||
| 193 | ||||||||
| 194 | ::: warning |
|||||||
| 195 | **Avoid these traps** |
|||||||
| 196 | ||||||||
| 197 | - **Three interviews and nothing else** — no breadth, no constraints, no observed behaviour. Caps at 6–7 on C2-1. |
|||||||
| 198 | - **A survey of one class with no interviews** — no depth, no rationale beyond "*it was easy*". |
|||||||
| 199 | - **Reports only** — you've researched the area but haven't engaged with real users; the rubric's *user characteristics* section will be weak. |
|||||||
| 200 | ::: |
|||||||
| 201 | ||||||||
| 202 | The 9–10 descriptor expects you to **explain why each method was chosen for its specific job** — what data it would yield that the others wouldn't. Practise that explanation out loud before your C2-1 station defence. |
|||||||
| 203 | ||||||||
| 204 | --- |
|||||||
| 205 | ||||||||
|
206 | ## Check Your Understanding |
||||||
| 207 | ||||||||
| 208 | Answer in your head first, then click the spoiler to check. |
|||||||
| 209 | ||||||||
| 210 | **1.** Which of these is a **secondary** data collection method? |
|||||||
|
211 | *(a) interview · (b) survey · (c) observation · (d) report* |
||||||
|
212 | |||||||
|
213 | >! **(d) report.** Secondary methods reuse data someone else already collected — that's what *reports / existing data* covers. Interviews, surveys, and observations are all *primary* — you collect the data first-hand. |
||||||
|
214 | |||||||
| 215 | **2.** You want to know whether your survey finding *"most students prefer dark mode"* is real or just a fluke. Which method would best **triangulate** that claim? |
|||||||
| 216 | *(a) another survey with the same question · (b) an interview asking why · (c) observation of students using the system in their usual environment · (d) a report on UI trends* |
|||||||
| 217 | ||||||||
| 218 | >! **(c) observation.** Triangulation = checking the claim against a *different* type of evidence. Re-asking in another survey reproduces the same self-report bias. Observation tests whether stated preference matches actual behaviour — the strongest cross-check. |
|||||||
| 219 | ||||||||
| 220 | **3.** A classmate's data collection plan is: *"three interviews with my client"*. According to this page, what's the biggest risk? |
|||||||
| 221 | *(a) too few people for breadth · (b) interviewer bias · (c) qualitative-only data · (d) all of the above* |
|||||||
| 222 | ||||||||
| 223 | >! **(d) all of the above.** Three interviews with one stakeholder gives no breadth (only one person's perspective), introduces interviewer bias unchecked, and produces only qualitative data. The rubric explicitly looks for a *justified mix*; this caps at 6–7 on C2-1. |
|||||||
| 224 | ||||||||
| 225 | **4.** You collect 50 survey responses to the question *"how often do you forget to bring your laptop?"* with options *Never / Rarely / Sometimes / Often / Always*. What kind of data is this? |
|||||||
| 226 | *(a) qualitative · (b) quantitative · (c) both · (d) neither* |
|||||||
| 227 | ||||||||
| 228 | >! **(b) quantitative.** Even though the *topic* feels qualitative ("how often" is a subjective judgement), the *data type* is quantitative — closed-ended responses you can count and compare. Data type is determined by what you do with it, not what it's about. |
|||||||
| 229 | ||||||||
| 230 | **5.** For your project's **technical environment** section of the SRS, which method is usually most efficient? |
|||||||
| 231 | *(a) interview the IT coordinator · (b) survey students about their devices · (c) observation in the computer lab · (d) request the school's existing IT device list (a report)* |
|||||||
| 232 | ||||||||
| 233 | >! **(d) request the school's IT device list.** That data already exists — your IT staff maintain it. Surveying or interviewing recreates work that's already done. (a) is a useful follow-up if the list is unclear, but the report is your starting point. |
|||||||
| 234 | ||||||||
| 235 | --- |
|||||||
| 236 | ||||||||
|
237 | ## See also |
||||||
| 238 | ||||||||
| 239 | - [Essential Terms](Essential%20Terms.md) — definitions of *qualitative*, *quantitative*, *open-ended*, *close-ended*, and the four method names |
|||||||
| 240 | - [Explaining vs Describing Your Data Collection](Explaining%20vs%20Describing%20Your%20Data%20Collection.md) — the verb ladder the C2-1 rubric uses to separate 7 from 9–10 |
|||||||
| 241 | - [Context Diagram](Context%20Diagram.md) — once you have data, the entities you discovered go on this diagram |
|||||||
| 242 | - [What Is an Entity](What%20Is%20an%20Entity.md) — your data sources may themselves be entities (e.g. a payment gateway, an LMS) |
|||||||
| 243 | ||||||||
| 244 | --- |
|||||||
| 245 | ||||||||
| 246 | *Adapted for VCE SD students from EDUCBA's [Data Collection Methods](https://www.educba.com/data-collection-methods/) overview, with project-specific framing for C02.* |
|||||||
