> **DRAFT** — under teacher review. # Designing an Evaluation Matrix The Hamilton and Alexandra College · Year 12 · 2026 Most students learn to *fill in* an evaluation matrix. Far fewer learn to *design* one — and that is the harder, more valuable skill. A good matrix is not a neutral measuring device that fell from the sky; it is a **series of justified design decisions**. Change who you are evaluating for, and a sensible matrix changes with it. This matters directly for **Unit 4 Outcome 1**, where you evaluate your own solution against criteria *you* chose and *you* must justify. The single most common failure is writing criteria *after* the solution is built, to flatter what you happened to make. Designing the matrix first — and defending it — is how you avoid that. --- ## Start by taking a great matrix apart Before you build one, reverse-engineer one that professionals already trust. **CHOICE magazine's May 2026 headphones test** ("Hi-fly way" — the best headphones *for travel*) is an excellent model. Look past the scores and ask *why the matrix is shaped the way it is*: - It produces a single composite **CHOICE Expert Rating** from **six weighted criteria**: sound quality 45%, active noise cancellation 15%, ease of use 17.5%, battery 10%, comfort 7.5%, durability 5%. - The weights are not equal, and they are not arbitrary — they encode a **point of view**. Sound at 45% says, plainly, that it matters nearly as much as everything else combined; ANC earns its own 15% precisely *because* this is a travel test (planes, trains, noisy hostels). - Two things are measured and shown in the table but **deliberately excluded** from the rating: **microphone quality** and **cleaning**. CHOICE reports them because some readers care, but decided they should not move the headline score. - Criteria use **different measurement methods**: lab measurement for sound and ANC (measured with ANC active at maximum reduction), a subjective panel for comfort, and claimed-versus-measured figures for battery. - It sets a **recommendation cut-off** — CHOICE only recommends headphones scoring **75% or above**. A matrix can decide not just *ranking* but *what counts as good enough*. - It handles **missing criteria** gracefully: for headphones with no ANC or no battery, those weights are **redistributed evenly** across the remaining criteria, so different product types are still judged fairly. - The weights themselves were **revised** (new as of April 2025) — proof that even a trusted matrix is a set of *decisions* its owners revisit, not a law of nature. Every one of those is a decision someone made and could defend. That is what you are learning to do. Here is that same matrix laid out **as a set of design decisions** — not the product scores (those are CHOICE's data), but the *shape* you are reverse-engineering: | Criterion | Weight | How it's measured | Counts toward the rating? | |---|---|---|---| | Sound quality | **45%** | Lab measurement | ✅ | | Active noise cancellation | **15%** | Lab (ANC active, max reduction) | ✅ | | Ease of use | **17.5%** | Tester rating | ✅ | | Battery | **10%** | Claimed vs measured run-time | ✅ | | Comfort | **7.5%** | Subjective panel | ✅ | | Durability | **5%** | Tester rating | ✅ | | Microphone | — | Lab measurement | ❌ shown, not scored | | Cleaning | — | Tester rating | ❌ shown, not scored | Read the table *downwards* and the design jumps out: six criteria carry the score, the weights tilt hard toward sound, and two whole columns are measured but **deliberately left out**. In this test the top travel picks landed at Sony WH-1000XM6 80%, JBL Tour One M3 79% and Bose QuietComfort Ultra 78% — but the lesson is the *shape* above, not the ranking. *Structure and weights: CHOICE, May 2026 (choice.com.au). The per-product scores are CHOICE's own data — see the original.* *Want to see which of these criteria are **efficiency** and which are **effectiveness** — and why most product matrices come out lopsided? See [CHOICE Headphones — Spotting Efficiency and Effectiveness](/sd/C04/CHOICE%20Headphones%20-%20Spotting%20Efficiency%20and%20Effectiveness).* --- ## The seven design decisions Designing a matrix means making — and justifying — these seven choices, roughly in order: 1. **Define the user.** Criteria flow from *who* you are evaluating for. "Best headphones" is meaningless; "best headphones *for a daily train commuter*" tells you what to measure. 2. **Choose criteria that matter to that user.** Not everything that *can* be measured *should* be. Pick the handful that decide the outcome for this person. 3. **Decide how each criterion is measured.** Objective lab test, subjective rating, feature checklist, or spec sheet — and say which, because it changes how much to trust the number. 4. **Assign weights, and justify them.** Weights are where your point of view lives. Equal weights are themselves a claim (that everything matters the same) — usually a weak one. 5. **Choose a scale.** 0–5, 0–100, or a traffic-light system. Be consistent so scores are comparable across options. 6. **Decide what to exclude — on purpose.** Naming what you *leave out*, and why, is a mark of a considered matrix (CHOICE excluding microphone quality). Silent omissions look like oversights; stated exclusions look like judgement. 7. **Stress-test it.** Swap the user, try to game it, and hunt for the criterion you forgot. A matrix that survives this is one you can defend. --- ## Watch: CHOICE evaluates hot cross buns (1 min) Same organisation, same craft as the headphones test — just tastier. Before CHOICE rates anything, an expert panel scores it against agreed criteria. **🎯 Watch for:** the *criteria* the panel judges against — and that they're agreed *before* the tasting, never invented afterwards to fit a favourite. That's the whole discipline, applied to a hot cross bun. {{Video|src=https://www.youtube.com/watch?v=k7m04xW0fzk}} *Hot cross buns taste test — CHOICE Australia. The [written version](https://www.choice.com.au/food-and-drink/bread-cereal-and-grains/bread/articles/hot-cross-buns) shows the criteria and method in full.* --- ## Worked example — choosing a Python IDE Let's make all seven decisions explicit, using three tools you might actually meet: **Thonny**, **VS Code**, and **IDLE**. 1. **User:** a Year 11 student in their *first* programming class, on a school laptop, who needs to get a program running quickly without being overwhelmed. 2. **Criteria that matter to them:** how fast you get started, how simple the interface is, how helpful it is for a beginner's errors, whether it runs on school hardware, and how much room there is to grow. 3. **Measurement method:** *getting started* = time to first running program (objective); *simple interface* and *beginner debugging* = a subjective rating against a checklist; *runs on school laptop* = does it install and run within the SOE; *room to grow* = a feature checklist. 4. **Weights, justified:** a first-week beginner is won or lost on getting started and not being overwhelmed, so those carry the most; "room to grow" barely matters *yet*, so it carries the least. 5. **Scale:** 0–5. 6. **Deliberate exclusion:** extension ecosystems and Git integration are reported but *not scored* — they matter enormously to a professional, and not at all to this student in week one. | Criterion (weight) | Thonny | VS Code | IDLE | |---|---|---|---| | Getting started — time to first run (30%) | 5 | 2 | 4 | | Simple, uncluttered interface (25%) | 5 | 2 | 4 | | Helpful for beginner errors / step-debugging (25%) | 5 | 3 | 2 | | Runs on the school laptop (12%) | 4 | 3 | 5 | | Room to grow (8%) | 3 | 5 | 2 | | **Weighted total / 5** | **4.7** | **2.6** | **3.5** | *(Scores are illustrative judgements, not benchmarks.)* For *this* user, **Thonny** wins — which makes sense, because it was designed for exactly this person. --- ## Stress-test your matrix The result above is only as good as the persona behind it. Pressure-test it three ways: - **Swap the persona.** Re-run it for a Year 12 student mid-SAT, building a large program. Now "room to grow", debugging power and extensions matter far more, getting-started far less — re-weight, and **VS Code** climbs to the top. Same options, different user, opposite answer. That is the whole lesson. - **Try to game it.** Could someone inflate one weight to force their favourite to win? If a single criterion at 60% decides everything, your matrix is really a one-criterion decision wearing a disguise. - **Hunt the missing criterion.** What did you forget? Cost? Offline use? Accessibility? A matrix is only as honest as its least-considered gap. --- ## The quality factors are a menu, not a matrix VCE names software quality factors — **effectiveness, efficiency, usability, reliability, maintainability, robustness, security**. Treat that list as a **menu to select from**, never a ready-made matrix. Your job is to *select* the few that matter to your user, *weight* them, and *justify* both choices. Tipping the whole menu into a table with equal weights is the opposite of evaluation. For the efficiency-versus-effectiveness distinction specifically, see [Efficiency vs Effectiveness](/sd/C04/Efficiency%20vs%20Effectiveness). --- ## Connection to your SAT In **C4-2** you apply this exact skill to your *own* design ideas — building [your evaluation matrix](/sd/C04/C4-2%20Checkpoint%20-%20Evaluation%20Matrix) to decide which design proceeds. In **Unit 4 Outcome 1** you do it again to evaluate your finished solution. The discipline is the same every time: **decide and defend the criteria before you judge anything against them.** Criteria written to flatter a finished product fool no examiner. --- ## Practise > A Year 12 student wants to choose note-taking software for VCE — **OneNote**, **Notion**, or **Obsidian**. Design an evaluation matrix for this user: define the persona precisely, choose and weight four or five criteria (with a measurement method for each), name **one** thing you deliberately exclude and why, then state which tool you would recommend (*explain*) and what you trade away by not choosing the runner-up (*justify*). --- ## See also - [Activity — Design a Hot Cross Bun Matrix](/sd/C04/Activity%20-%20Hot%20Cross%20Bun%20Matrix) — practise the seven decisions on something edible before you do it for software - [C4-2 Checkpoint — Your Evaluation Matrix](/sd/C04/C4-2%20Checkpoint%20-%20Evaluation%20Matrix) — build the matrix for your own design ideas - [Evaluation Criteria and the SRS](/sd/C04/Evaluation%20Criteria%20and%20the%20SRS) · [Efficiency vs Effectiveness](/sd/C04/Efficiency%20vs%20Effectiveness) - [Evaluation vs Testing](/sd/C04/Evaluation%20vs%20Testing) --- ← Back to [C04 Home](/sd/C04/C04-home) · [VCE Software Development Hub](/sd/VCE%20Software%20Development%20Hub)
