DRAFT — under teacher review.
Evaluation vs Testing — What's the Difference?
The Hamilton and Alexandra College · Year 12 · 2026
One of the most persistent mistakes in C4-2 is writing about testing when the rubric is asking about evaluation. These two processes are defined differently, use different evidence, and answer different questions. You can pass every test and still fail an evaluation criterion.
The one-sentence rule
| Process | One-sentence definition | Question it answers |
|---|---|---|
| Testing | Checking whether the software works correctly | "Does it do what it is supposed to do?" |
| Evaluation | Judging how well the software meets the criteria set at design time | "Does it meet the standard we set?" |
Testing is binary: the system either passes or fails a test case. Evaluation is a judgement against a criterion — and the criterion was derived from an SRS requirement.
A concrete example
Imagine a student builds an event-booking app. Here are three test results:
| Test | Result |
|---|---|
| User can search for events | Pass |
| Booking confirmation email is sent | Pass |
| All records saved correctly to file | Pass |
All tests pass. Now look at the evaluation matrix:
| Criterion | SRS Requirement | Score (1–5) | Result |
|---|---|---|---|
| A first-time user shall complete a booking in under 4 minutes without assistance | NFR2 — Usability | 2 | Fails criterion |
| The event list shall load in under 2 seconds on school Wi-Fi | NFR1 — Performance | 4 | Meets criterion |
The app works (tests pass) but does not meet the usability standard (evaluation fails). Testing cannot catch this — only evaluation can.
This is the distinction the rubric language at 7–8 and 9–10 is checking:
"Uses the evaluation criteria to explain which elements of the design ideas should be further developed…" — 7–8 band
"Uses the evaluation criteria to justify which elements…" — 9–10 band
The rubric is asking you to reason from criterion scores, not from test results.
The most common mistakes
Mistake 1: Treating testing as evaluation
"I evaluated the app by running three test cases. All tests passed, so the design is good."
Test pass rates are not evaluation evidence. You need criterion scores linked to SRS requirements.
Mistake 2: Writing evaluation criteria that are just test cases
"Criterion: The search function returns results."
This is a test case disguised as a criterion — it is binary (pass/fail) and measures whether a feature exists, not how well it works. A genuine evaluation criterion must be measurable against a standard, for example: "The search function returns results in under 1.5 seconds on a standard device."
Mistake 3: Inventing criteria that float free of the SRS
"Criterion: The interface is attractive."
Attractiveness is a valid effectiveness factor, but only if your SRS included a requirement about it. If it is not in the SRS, you cannot evaluate against it — you would be judging the solution against a goal you never set.
Why VCAA cares
At 5–6, the rubric requires you to develop and apply evaluation criteria. This means you need actual criteria with scoring, not a description of test results.
At 7–8, you must use the evaluation criteria to explain decisions. This means referencing specific scores: "Design Idea A scored 4/5 for usability because..."
At 9–10, you must use the evaluation criteria to justify decisions. This means arguing that the scores demonstrate a design element is worth developing further, with explicit links back to SRS requirements.
See also
- C04 Resources: external reading and resources
← Back to C04 Home · VCE Software Development Hub
