Blame
|
1 | > **DRAFT** — under teacher review. |
||||||
| 2 | ||||||||
| 3 | # Evaluation vs Testing — What's the Difference? |
|||||||
| 4 | ||||||||
|
5 | The Hamilton and Alexandra College · Year 12 · 2026 |
||||||
|
6 | |||||||
| 7 | One of the most persistent mistakes in C4-2 is writing about *testing* when the rubric is asking about *evaluation*. These two processes are defined differently, use different evidence, and answer different questions. You can pass every test and still fail an evaluation criterion. |
|||||||
| 8 | ||||||||
| 9 | --- |
|||||||
| 10 | ||||||||
| 11 | ## The one-sentence rule |
|||||||
| 12 | ||||||||
| 13 | | Process | One-sentence definition | Question it answers | |
|||||||
| 14 | |---|---|---| |
|||||||
| 15 | | **Testing** | Checking whether the software works correctly | "Does it **do** what it is supposed to do?" | |
|||||||
| 16 | | **Evaluation** | Judging how well the software meets the criteria set at design time | "Does it meet the **standard** we set?" | |
|||||||
| 17 | ||||||||
| 18 | Testing is binary: the system either passes or fails a test case. Evaluation is a judgement against a criterion — and the criterion was derived from an SRS requirement. |
|||||||
| 19 | ||||||||
| 20 | --- |
|||||||
| 21 | ||||||||
| 22 | ## A concrete example |
|||||||
| 23 | ||||||||
| 24 | Imagine a student builds an event-booking app. Here are three test results: |
|||||||
| 25 | ||||||||
| 26 | | Test | Result | |
|||||||
| 27 | |---|---| |
|||||||
| 28 | | User can search for events | Pass | |
|||||||
| 29 | | Booking confirmation email is sent | Pass | |
|||||||
| 30 | | All records saved correctly to file | Pass | |
|||||||
| 31 | ||||||||
| 32 | All tests pass. Now look at the evaluation matrix: |
|||||||
| 33 | ||||||||
| 34 | | Criterion | SRS Requirement | Score (1–5) | Result | |
|||||||
| 35 | |---|---|---|---| |
|||||||
| 36 | | A first-time user shall complete a booking in under 4 minutes without assistance | NFR2 — Usability | 2 | **Fails criterion** | |
|||||||
| 37 | | The event list shall load in under 2 seconds on school Wi-Fi | NFR1 — Performance | 4 | Meets criterion | |
|||||||
| 38 | ||||||||
| 39 | The app *works* (tests pass) but *does not meet the usability standard* (evaluation fails). Testing cannot catch this — only evaluation can. |
|||||||
| 40 | ||||||||
| 41 | This is the distinction the rubric language at 7–8 and 9–10 is checking: |
|||||||
| 42 | ||||||||
| 43 | > *"Uses the evaluation criteria to explain which elements of the design ideas should be further developed…"* — 7–8 band |
|||||||
| 44 | ||||||||
| 45 | > *"Uses the evaluation criteria to **justify** which elements…"* — 9–10 band |
|||||||
| 46 | ||||||||
| 47 | The rubric is asking you to reason from criterion scores, not from test results. |
|||||||
| 48 | ||||||||
| 49 | --- |
|||||||
| 50 | ||||||||
| 51 | ## The most common mistakes |
|||||||
| 52 | ||||||||
| 53 | ### Mistake 1: Treating testing as evaluation |
|||||||
| 54 | ||||||||
| 55 | > ~~"I evaluated the app by running three test cases. All tests passed, so the design is good."~~ |
|||||||
| 56 | ||||||||
| 57 | Test pass rates are not evaluation evidence. You need criterion scores linked to SRS requirements. |
|||||||
| 58 | ||||||||
| 59 | ### Mistake 2: Writing evaluation criteria that are just test cases |
|||||||
| 60 | ||||||||
| 61 | > ~~"Criterion: The search function returns results."~~ |
|||||||
| 62 | ||||||||
| 63 | This is a test case disguised as a criterion — it is binary (pass/fail) and measures whether a feature exists, not how well it works. A genuine evaluation criterion must be measurable against a standard, for example: *"The search function returns results in under 1.5 seconds on a standard device."* |
|||||||
| 64 | ||||||||
| 65 | ### Mistake 3: Inventing criteria that float free of the SRS |
|||||||
| 66 | ||||||||
| 67 | > ~~"Criterion: The interface is attractive."~~ |
|||||||
| 68 | ||||||||
| 69 | Attractiveness is a valid effectiveness factor, but only if your SRS included a requirement about it. If it is not in the SRS, you cannot evaluate against it — you would be judging the solution against a goal you never set. |
|||||||
| 70 | ||||||||
| 71 | --- |
|||||||
| 72 | ||||||||
| 73 | ## Why VCAA cares |
|||||||
| 74 | ||||||||
| 75 | At **5–6**, the rubric requires you to *develop and apply evaluation criteria*. This means you need actual criteria with scoring, not a description of test results. |
|||||||
| 76 | ||||||||
| 77 | At **7–8**, you must *use the evaluation criteria to explain* decisions. This means referencing specific scores: "Design Idea A scored 4/5 for usability because..." |
|||||||
| 78 | ||||||||
| 79 | At **9–10**, you must *use the evaluation criteria to justify* decisions. This means arguing that the scores demonstrate a design element is worth developing further, with explicit links back to SRS requirements. |
|||||||
| 80 | ||||||||
| 81 | --- |
|||||||
| 82 | ||||||||
| 83 | ## See also |
|||||||
| 84 | ||||||||
| 85 | - C04 Resources: [external reading and resources](/sd/Resources/C04-Resources) |
|||||||
| 86 | ||||||||
| 87 | --- |
|||||||
| 88 | ||||||||
| 89 | ← Back to [C04 Home](/sd/C04/C04-home) · [VCE Software Development Hub](/sd/VCE%20Software%20Development%20Hub) |
|||||||
