> **DRAFT** — under teacher review.

# Designing an Evaluation Matrix

The Hamilton and Alexandra College · Year 12 · 2026

Most students learn to *fill in* an evaluation matrix. Far fewer learn to *design* one — and that is the harder, more valuable skill. A good matrix is not a neutral measuring device that fell from the sky; it is a **series of justified design decisions**. Change who you are evaluating for, and a sensible matrix changes with it.

This matters directly for **Unit 4 Outcome 1**, where you evaluate your own solution against criteria *you* chose and *you* must justify. The single most common failure is writing criteria *after* the solution is built, to flatter what you happened to make. Designing the matrix first — and defending it — is how you avoid that.

---

## Start by taking a great matrix apart

Before you build one, reverse-engineer one that professionals already trust. **CHOICE magazine's May 2026 headphones test** ("Hi-fly way" — the best headphones *for travel*) is an excellent model. Look past the scores and ask *why the matrix is shaped the way it is*:

- It produces a single composite **CHOICE Expert Rating** from **six weighted criteria**: sound quality 45%, active noise cancellation 15%, ease of use 17.5%, battery 10%, comfort 7.5%, durability 5%.
- The weights are not equal, and they are not arbitrary — they encode a **point of view**. Sound at 45% says, plainly, that it matters nearly as much as everything else combined; ANC earns its own 15% precisely *because* this is a travel test (planes, trains, noisy hostels).
- Two things are measured and shown in the table but **deliberately excluded** from the rating: **microphone quality** and **cleaning**. CHOICE reports them because some readers care, but decided they should not move the headline score.
- Criteria use **different measurement methods**: lab measurement for sound and ANC (measured with ANC active at maximum reduction), a subjective panel for comfort, and claimed-versus-measured figures for battery.
- It sets a **recommendation cut-off** — CHOICE only recommends headphones scoring **75% or above**. A matrix can decide not just *ranking* but *what counts as good enough*.
- It handles **missing criteria** gracefully: for headphones with no ANC or no battery, those weights are **redistributed evenly** across the remaining criteria, so different product types are still judged fairly.
- The weights themselves were **revised** (new as of April 2025) — proof that even a trusted matrix is a set of *decisions* its owners revisit, not a law of nature.

Every one of those is a decision someone made and could defend. That is what you are learning to do.

Here is that same matrix laid out **as a set of design decisions** — not the product scores (those are CHOICE's data), but the *shape* you are reverse-engineering:

| Criterion | Weight | How it's measured | Counts toward the rating? |
|---|---|---|---|
| Sound quality | **45%** | Lab measurement | ✅ |
| Active noise cancellation | **15%** | Lab (ANC active, max reduction) | ✅ |
| Ease of use | **17.5%** | Tester rating | ✅ |
| Battery | **10%** | Claimed vs measured run-time | ✅ |
| Comfort | **7.5%** | Subjective panel | ✅ |
| Durability | **5%** | Tester rating | ✅ |
| Microphone | — | Lab measurement | ❌ shown, not scored |
| Cleaning | — | Tester rating | ❌ shown, not scored |

Read the table *downwards* and the design jumps out: six criteria carry the score, the weights tilt hard toward sound, and two whole columns are measured but **deliberately left out**. In this test the top travel picks landed at Sony WH-1000XM6 80%, JBL Tour One M3 79% and Bose QuietComfort Ultra 78% — but the lesson is the *shape* above, not the ranking.

*Structure and weights: CHOICE, May 2026 (choice.com.au). The per-product scores are CHOICE's own data — see the original.*

*Want to see which of these criteria are **efficiency** and which are **effectiveness** — and why most product matrices come out lopsided? See [CHOICE Headphones — Spotting Efficiency and Effectiveness](/sd/C04/CHOICE%20Headphones%20-%20Spotting%20Efficiency%20and%20Effectiveness).*

---

## The seven design decisions

Designing a matrix means making — and justifying — these seven choices, roughly in order:

1. **Define the user.** Criteria flow from *who* you are evaluating for. "Best headphones" is meaningless; "best headphones *for a daily train commuter*" tells you what to measure.
2. **Choose criteria that matter to that user.** Not everything that *can* be measured *should* be. Pick the handful that decide the outcome for this person.
3. **Decide how each criterion is measured.** Objective lab test, subjective rating, feature checklist, or spec sheet — and say which, because it changes how much to trust the number.
4. **Assign weights, and justify them.** Weights are where your point of view lives. Equal weights are themselves a claim (that everything matters the same) — usually a weak one.
5. **Choose a scale.** 0–5, 0–100, or a traffic-light system. Be consistent so scores are comparable across options.
6. **Decide what to exclude — on purpose.** Naming what you *leave out*, and why, is a mark of a considered matrix (CHOICE excluding microphone quality). Silent omissions look like oversights; stated exclusions look like judgement.
7. **Stress-test it.** Swap the user, try to game it, and hunt for the criterion you forgot. A matrix that survives this is one you can defend.

---

## Watch: CHOICE evaluates hot cross buns (1 min)

Same organisation, same craft as the headphones test — just tastier. Before CHOICE rates anything, an expert panel scores it against agreed criteria.

**🎯 Watch for:** the *criteria* the panel judges against — and that they're agreed *before* the tasting, never invented afterwards to fit a favourite. That's the whole discipline, applied to a hot cross bun.

{{Video|src=https://www.youtube.com/watch?v=k7m04xW0fzk}}

*Hot cross buns taste test — CHOICE Australia. The [written version](https://www.choice.com.au/food-and-drink/bread-cereal-and-grains/bread/articles/hot-cross-buns) shows the criteria and method in full.*

---

## Worked example — choosing a Python IDE

Let's make all seven decisions explicit, using three tools you might actually meet: **Thonny**, **VS Code**, and **IDLE**.

1. **User:** a Year 11 student in their *first* programming class, on a school laptop, who needs to get a program running quickly without being overwhelmed.
2. **Criteria that matter to them:** how fast you get started, how simple the interface is, how helpful it is for a beginner's errors, whether it runs on school hardware, and how much room there is to grow.
3. **Measurement method:** *getting started* = time to first running program (objective); *simple interface* and *beginner debugging* = a subjective rating against a checklist; *runs on school laptop* = does it install and run within the SOE; *room to grow* = a feature checklist.
4. **Weights, justified:** a first-week beginner is won or lost on getting started and not being overwhelmed, so those carry the most; "room to grow" barely matters *yet*, so it carries the least.
5. **Scale:** 0–5.
6. **Deliberate exclusion:** extension ecosystems and Git integration are reported but *not scored* — they matter enormously to a professional, and not at all to this student in week one.

| Criterion (weight) | Thonny | VS Code | IDLE |
|---|---|---|---|
| Getting started — time to first run (30%) | 5 | 2 | 4 |
| Simple, uncluttered interface (25%) | 5 | 2 | 4 |
| Helpful for beginner errors / step-debugging (25%) | 5 | 3 | 2 |
| Runs on the school laptop (12%) | 4 | 3 | 5 |
| Room to grow (8%) | 3 | 5 | 2 |
| **Weighted total / 5** | **4.7** | **2.6** | **3.5** |

*(Scores are illustrative judgements, not benchmarks.)* For *this* user, **Thonny** wins — which makes sense, because it was designed for exactly this person.

---

## Stress-test your matrix

The result above is only as good as the persona behind it. Pressure-test it three ways:

- **Swap the persona.** Re-run it for a Year 12 student mid-SAT, building a large program. Now "room to grow", debugging power and extensions matter far more, getting-started far less — re-weight, and **VS Code** climbs to the top. Same options, different user, opposite answer. That is the whole lesson.
- **Try to game it.** Could someone inflate one weight to force their favourite to win? If a single criterion at 60% decides everything, your matrix is really a one-criterion decision wearing a disguise.
- **Hunt the missing criterion.** What did you forget? Cost? Offline use? Accessibility? A matrix is only as honest as its least-considered gap.

---

## The quality factors are a menu, not a matrix

VCE names software quality factors — **effectiveness, efficiency, usability, reliability, maintainability, robustness, security**. Treat that list as a **menu to select from**, never a ready-made matrix. Your job is to *select* the few that matter to your user, *weight* them, and *justify* both choices. Tipping the whole menu into a table with equal weights is the opposite of evaluation. For the efficiency-versus-effectiveness distinction specifically, see [Efficiency vs Effectiveness](/sd/C04/Efficiency%20vs%20Effectiveness).

---

## Connection to your SAT

In **C4-2** you apply this exact skill to your *own* design ideas — building [your evaluation matrix](/sd/C04/C4-2%20Checkpoint%20-%20Evaluation%20Matrix) to decide which design proceeds. In **Unit 4 Outcome 1** you do it again to evaluate your finished solution. The discipline is the same every time: **decide and defend the criteria before you judge anything against them.** Criteria written to flatter a finished product fool no examiner.

---

## Practise

> A Year 12 student wants to choose note-taking software for VCE — **OneNote**, **Notion**, or **Obsidian**. Design an evaluation matrix for this user: define the persona precisely, choose and weight four or five criteria (with a measurement method for each), name **one** thing you deliberately exclude and why, then state which tool you would recommend (*explain*) and what you trade away by not choosing the runner-up (*justify*).

---

## See also

- [Activity — Design a Hot Cross Bun Matrix](/sd/C04/Activity%20-%20Hot%20Cross%20Bun%20Matrix) — practise the seven decisions on something edible before you do it for software
- [C4-2 Checkpoint — Your Evaluation Matrix](/sd/C04/C4-2%20Checkpoint%20-%20Evaluation%20Matrix) — build the matrix for your own design ideas
- [Evaluation Criteria and the SRS](/sd/C04/Evaluation%20Criteria%20and%20the%20SRS) · [Efficiency vs Effectiveness](/sd/C04/Efficiency%20vs%20Effectiveness)
- [Evaluation vs Testing](/sd/C04/Evaluation%20vs%20Testing)

---

← Back to [C04 Home](/sd/C04/C04-home) · [VCE Software Development Hub](/sd/VCE%20Software%20Development%20Hub)
0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9