Blame

530bd6 lisa 2026-04-25 22:54:52
Add C04 wiki content (DRAFT) — landing page, Resources rows Light first pass: C04-home, 1 key-concept page marked DRAFT (Evaluation vs Testing), 4 external resource rows in C04-Resources.md (Isaac CS: testing/evaluation, design evaluation, wireframe annotation, evaluation criteria). Excalidraw template and teacher-recorded videos remain un-ticked in backlog. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
1
> **DRAFT** — under teacher review.
2
3
# Evaluation vs Testing — What's the Difference?
4
1abe81 lisa 2026-04-30 18:17:38
School name → 'The Hamilton and Alexandra College' across all wiki pages Sweep rename of the standard subtitle line on every wiki page (sd/, hub, resources). 25 files affected.
5
The Hamilton and Alexandra College · Year 12 · 2026
530bd6 lisa 2026-04-25 22:54:52
Add C04 wiki content (DRAFT) — landing page, Resources rows Light first pass: C04-home, 1 key-concept page marked DRAFT (Evaluation vs Testing), 4 external resource rows in C04-Resources.md (Isaac CS: testing/evaluation, design evaluation, wireframe annotation, evaluation criteria). Excalidraw template and teacher-recorded videos remain un-ticked in backlog. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
6
7
One of the most persistent mistakes in C4-2 is writing about *testing* when the rubric is asking about *evaluation*. These two processes are defined differently, use different evidence, and answer different questions. You can pass every test and still fail an evaluation criterion.
8
9
---
10
11
## The one-sentence rule
12
13
| Process | One-sentence definition | Question it answers |
14
|---|---|---|
15
| **Testing** | Checking whether the software works correctly | "Does it **do** what it is supposed to do?" |
16
| **Evaluation** | Judging how well the software meets the criteria set at design time | "Does it meet the **standard** we set?" |
17
18
Testing is binary: the system either passes or fails a test case. Evaluation is a judgement against a criterion — and the criterion was derived from an SRS requirement.
19
20
---
21
22
## A concrete example
23
24
Imagine a student builds an event-booking app. Here are three test results:
25
26
| Test | Result |
27
|---|---|
28
| User can search for events | Pass |
29
| Booking confirmation email is sent | Pass |
30
| All records saved correctly to file | Pass |
31
32
All tests pass. Now look at the evaluation matrix:
33
34
| Criterion | SRS Requirement | Score (1–5) | Result |
35
|---|---|---|---|
36
| A first-time user shall complete a booking in under 4 minutes without assistance | NFR2 — Usability | 2 | **Fails criterion** |
37
| The event list shall load in under 2 seconds on school Wi-Fi | NFR1 — Performance | 4 | Meets criterion |
38
39
The app *works* (tests pass) but *does not meet the usability standard* (evaluation fails). Testing cannot catch this — only evaluation can.
40
41
This is the distinction the rubric language at 7–8 and 9–10 is checking:
42
43
> *"Uses the evaluation criteria to explain which elements of the design ideas should be further developed…"* — 7–8 band
44
45
> *"Uses the evaluation criteria to **justify** which elements…"* — 9–10 band
46
47
The rubric is asking you to reason from criterion scores, not from test results.
48
49
---
50
51
## The most common mistakes
52
53
### Mistake 1: Treating testing as evaluation
54
55
> ~~"I evaluated the app by running three test cases. All tests passed, so the design is good."~~
56
57
Test pass rates are not evaluation evidence. You need criterion scores linked to SRS requirements.
58
59
### Mistake 2: Writing evaluation criteria that are just test cases
60
61
> ~~"Criterion: The search function returns results."~~
62
63
This is a test case disguised as a criterion — it is binary (pass/fail) and measures whether a feature exists, not how well it works. A genuine evaluation criterion must be measurable against a standard, for example: *"The search function returns results in under 1.5 seconds on a standard device."*
64
65
### Mistake 3: Inventing criteria that float free of the SRS
66
67
> ~~"Criterion: The interface is attractive."~~
68
69
Attractiveness is a valid effectiveness factor, but only if your SRS included a requirement about it. If it is not in the SRS, you cannot evaluate against it — you would be judging the solution against a goal you never set.
70
71
---
72
73
## Why VCAA cares
74
75
At **5–6**, the rubric requires you to *develop and apply evaluation criteria*. This means you need actual criteria with scoring, not a description of test results.
76
77
At **7–8**, you must *use the evaluation criteria to explain* decisions. This means referencing specific scores: "Design Idea A scored 4/5 for usability because..."
78
79
At **9–10**, you must *use the evaluation criteria to justify* decisions. This means arguing that the scores demonstrate a design element is worth developing further, with explicit links back to SRS requirements.
80
81
---
82
83
## See also
84
85
- C04 Resources: [external reading and resources](/sd/Resources/C04-Resources)
86
87
---
88
89
← Back to [C04 Home](/sd/C04/C04-home) · [VCE Software Development Hub](/sd/VCE%20Software%20Development%20Hub)