> **DRAFT** — under teacher review.

# Evaluation vs Testing — What's the Difference?

The Hamilton and Alexandra College · Year 12 · 2026

One of the most persistent mistakes in C4-2 is writing about *testing* when the rubric is asking about *evaluation*. These two processes are defined differently, use different evidence, and answer different questions. You can pass every test and still fail an evaluation criterion.

---

## The one-sentence rule

| Process | One-sentence definition | Question it answers |
|---|---|---|
| **Testing** | Checking whether the software works correctly | "Does it **do** what it is supposed to do?" |
| **Evaluation** | Judging how well the software meets the criteria set at design time | "Does it meet the **standard** we set?" |

Testing is binary: the system either passes or fails a test case. Evaluation is a judgement against a criterion — and the criterion was derived from an SRS requirement.

---

## A concrete example

Imagine a student builds an event-booking app. Here are three test results:

| Test | Result |
|---|---|
| User can search for events | Pass |
| Booking confirmation email is sent | Pass |
| All records saved correctly to file | Pass |

All tests pass. Now look at the evaluation matrix:

| Criterion | SRS Requirement | Score (1–5) | Result |
|---|---|---|---|
| A first-time user shall complete a booking in under 4 minutes without assistance | NFR2 — Usability | 2 | **Fails criterion** |
| The event list shall load in under 2 seconds on school Wi-Fi | NFR1 — Performance | 4 | Meets criterion |

The app *works* (tests pass) but *does not meet the usability standard* (evaluation fails). Testing cannot catch this — only evaluation can.

This is the distinction the rubric language at 7–8 and 9–10 is checking:

> *"Uses the evaluation criteria to explain which elements of the design ideas should be further developed…"* — 7–8 band

> *"Uses the evaluation criteria to **justify** which elements…"* — 9–10 band

The rubric is asking you to reason from criterion scores, not from test results.

---

## The most common mistakes

### Mistake 1: Treating testing as evaluation

> ~~"I evaluated the app by running three test cases. All tests passed, so the design is good."~~

Test pass rates are not evaluation evidence. You need criterion scores linked to SRS requirements.

### Mistake 2: Writing evaluation criteria that are just test cases

> ~~"Criterion: The search function returns results."~~

This is a test case disguised as a criterion — it is binary (pass/fail) and measures whether a feature exists, not how well it works. A genuine evaluation criterion must be measurable against a standard, for example: *"The search function returns results in under 1.5 seconds on a standard device."*

### Mistake 3: Inventing criteria that float free of the SRS

> ~~"Criterion: The interface is attractive."~~

Attractiveness is a valid effectiveness factor, but only if your SRS included a requirement about it. If it is not in the SRS, you cannot evaluate against it — you would be judging the solution against a goal you never set.

---

## Why VCAA cares

At **5–6**, the rubric requires you to *develop and apply evaluation criteria*. This means you need actual criteria with scoring, not a description of test results.

At **7–8**, you must *use the evaluation criteria to explain* decisions. This means referencing specific scores: "Design Idea A scored 4/5 for usability because..."

At **9–10**, you must *use the evaluation criteria to justify* decisions. This means arguing that the scores demonstrate a design element is worth developing further, with explicit links back to SRS requirements.

---

## See also

- C04 Resources: [external reading and resources](/sd/Resources/C04-Resources)

---

← Back to [C04 Home](/sd/C04/C04-home) · [VCE Software Development Hub](/sd/VCE%20Software%20Development%20Hub)
0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9