> **DRAFT** — under teacher review. # Evaluation vs Testing — What's the Difference? The Hamilton and Alexandra College · Year 12 · 2026 One of the most persistent mistakes in C4-2 is writing about *testing* when the rubric is asking about *evaluation*. These two processes are defined differently, use different evidence, and answer different questions. You can pass every test and still fail an evaluation criterion. --- ## The one-sentence rule | Process | One-sentence definition | Question it answers | |---|---|---| | **Testing** | Checking whether the software works correctly | "Does it **do** what it is supposed to do?" | | **Evaluation** | Judging how well the software meets the criteria set at design time | "Does it meet the **standard** we set?" | Testing is binary: the system either passes or fails a test case. Evaluation is a judgement against a criterion — and the criterion was derived from an SRS requirement. --- ## A concrete example Imagine a student builds an event-booking app. Here are three test results: | Test | Result | |---|---| | User can search for events | Pass | | Booking confirmation email is sent | Pass | | All records saved correctly to file | Pass | All tests pass. Now look at the evaluation matrix: | Criterion | SRS Requirement | Score (1–5) | Result | |---|---|---|---| | A first-time user shall complete a booking in under 4 minutes without assistance | NFR2 — Usability | 2 | **Fails criterion** | | The event list shall load in under 2 seconds on school Wi-Fi | NFR1 — Performance | 4 | Meets criterion | The app *works* (tests pass) but *does not meet the usability standard* (evaluation fails). Testing cannot catch this — only evaluation can. This is the distinction the rubric language at 7–8 and 9–10 is checking: > *"Uses the evaluation criteria to explain which elements of the design ideas should be further developed…"* — 7–8 band > *"Uses the evaluation criteria to **justify** which elements…"* — 9–10 band The rubric is asking you to reason from criterion scores, not from test results. --- ## The most common mistakes ### Mistake 1: Treating testing as evaluation > ~~"I evaluated the app by running three test cases. All tests passed, so the design is good."~~ Test pass rates are not evaluation evidence. You need criterion scores linked to SRS requirements. ### Mistake 2: Writing evaluation criteria that are just test cases > ~~"Criterion: The search function returns results."~~ This is a test case disguised as a criterion — it is binary (pass/fail) and measures whether a feature exists, not how well it works. A genuine evaluation criterion must be measurable against a standard, for example: *"The search function returns results in under 1.5 seconds on a standard device."* ### Mistake 3: Inventing criteria that float free of the SRS > ~~"Criterion: The interface is attractive."~~ Attractiveness is a valid effectiveness factor, but only if your SRS included a requirement about it. If it is not in the SRS, you cannot evaluate against it — you would be judging the solution against a goal you never set. --- ## Why VCAA cares At **5–6**, the rubric requires you to *develop and apply evaluation criteria*. This means you need actual criteria with scoring, not a description of test results. At **7–8**, you must *use the evaluation criteria to explain* decisions. This means referencing specific scores: "Design Idea A scored 4/5 for usability because..." At **9–10**, you must *use the evaluation criteria to justify* decisions. This means arguing that the scores demonstrate a design element is worth developing further, with explicit links back to SRS requirements. --- ## See also - C04 Resources: [external reading and resources](/sd/Resources/C04-Resources) --- ← Back to [C04 Home](/sd/C04/C04-home) · [VCE Software Development Hub](/sd/VCE%20Software%20Development%20Hub)
