Blame
|
1 | > **DRAFT** — under teacher review. |
||||||
| 2 | ||||||||
| 3 | # Designing an Evaluation Matrix |
|||||||
| 4 | ||||||||
| 5 | The Hamilton and Alexandra College · Year 12 · 2026 |
|||||||
| 6 | ||||||||
| 7 | Most students learn to *fill in* an evaluation matrix. Far fewer learn to *design* one — and that is the harder, more valuable skill. A good matrix is not a neutral measuring device that fell from the sky; it is a **series of justified design decisions**. Change who you are evaluating for, and a sensible matrix changes with it. |
|||||||
| 8 | ||||||||
| 9 | This matters directly for **Unit 4 Outcome 1**, where you evaluate your own solution against criteria *you* chose and *you* must justify. The single most common failure is writing criteria *after* the solution is built, to flatter what you happened to make. Designing the matrix first — and defending it — is how you avoid that. |
|||||||
| 10 | ||||||||
| 11 | --- |
|||||||
| 12 | ||||||||
| 13 | ## Start by taking a great matrix apart |
|||||||
| 14 | ||||||||
| 15 | Before you build one, reverse-engineer one that professionals already trust. **CHOICE magazine's May 2026 headphones test** is an excellent model. Look past the scores and ask *why the matrix is shaped the way it is*: |
|||||||
| 16 | ||||||||
| 17 | - It produces a single composite **CHOICE Expert Rating** from **six weighted criteria**: sound 45%, noise cancelling 15%, ease of use 17.5%, battery 10%, comfort 7.5%, durability 5%. |
|||||||
| 18 | - The weights are not equal, and they are not arbitrary — they encode a **point of view**. Sound at 45% says, plainly, "for this product, how it sounds matters more than everything else combined nearly does." |
|||||||
| 19 | - Two things are measured but **deliberately excluded** from the score: microphone quality and ease of cleaning. CHOICE reports them because some readers care, but decided they should not move the headline rating. |
|||||||
| 20 | - The criteria suit a **specific user context** (here, travel — which is why noise cancelling earns its own 15%). |
|||||||
| 21 | - Different criteria use **different measurement methods**: lab instruments for sound, a subjective panel for comfort, a spec sheet for battery. |
|||||||
| 22 | ||||||||
| 23 | Every one of those is a decision someone made and could defend. That is what you are learning to do. |
|||||||
| 24 | ||||||||
| 25 | --- |
|||||||
| 26 | ||||||||
| 27 | ## The seven design decisions |
|||||||
| 28 | ||||||||
| 29 | Designing a matrix means making — and justifying — these seven choices, roughly in order: |
|||||||
| 30 | ||||||||
| 31 | 1. **Define the user.** Criteria flow from *who* you are evaluating for. "Best headphones" is meaningless; "best headphones *for a daily train commuter*" tells you what to measure. |
|||||||
| 32 | 2. **Choose criteria that matter to that user.** Not everything that *can* be measured *should* be. Pick the handful that decide the outcome for this person. |
|||||||
| 33 | 3. **Decide how each criterion is measured.** Objective lab test, subjective rating, feature checklist, or spec sheet — and say which, because it changes how much to trust the number. |
|||||||
| 34 | 4. **Assign weights, and justify them.** Weights are where your point of view lives. Equal weights are themselves a claim (that everything matters the same) — usually a weak one. |
|||||||
| 35 | 5. **Choose a scale.** 0–5, 0–100, or a traffic-light system. Be consistent so scores are comparable across options. |
|||||||
| 36 | 6. **Decide what to exclude — on purpose.** Naming what you *leave out*, and why, is a mark of a considered matrix (CHOICE excluding microphone quality). Silent omissions look like oversights; stated exclusions look like judgement. |
|||||||
| 37 | 7. **Stress-test it.** Swap the user, try to game it, and hunt for the criterion you forgot. A matrix that survives this is one you can defend. |
|||||||
| 38 | ||||||||
| 39 | --- |
|||||||
| 40 | ||||||||
| 41 | ## Watch: how weighting works (5 min) |
|||||||
| 42 | ||||||||
| 43 | **🎯 Watch for:** how the *weights* — not the scores — are where "what matters most to this user" actually gets encoded. |
|||||||
| 44 | ||||||||
| 45 | {{Video|src=https://www.youtube.com/watch?v=hUlB8lkfw9Q}} |
|||||||
| 46 | ||||||||
| 47 | *The Weighted Scoring Decision Matrix — airfocus. (It's framed for product managers; the method is exactly the same for software evaluation.)* |
|||||||
| 48 | ||||||||
| 49 | --- |
|||||||
| 50 | ||||||||
| 51 | ## Worked example — choosing a Python IDE |
|||||||
| 52 | ||||||||
| 53 | Let's make all seven decisions explicit, using three tools you might actually meet: **Thonny**, **VS Code**, and **IDLE**. |
|||||||
| 54 | ||||||||
| 55 | 1. **User:** a Year 11 student in their *first* programming class, on a school laptop, who needs to get a program running quickly without being overwhelmed. |
|||||||
| 56 | 2. **Criteria that matter to them:** how fast you get started, how simple the interface is, how helpful it is for a beginner's errors, whether it runs on school hardware, and how much room there is to grow. |
|||||||
| 57 | 3. **Measurement method:** *getting started* = time to first running program (objective); *simple interface* and *beginner debugging* = a subjective rating against a checklist; *runs on school laptop* = does it install and run within the SOE; *room to grow* = a feature checklist. |
|||||||
| 58 | 4. **Weights, justified:** a first-week beginner is won or lost on getting started and not being overwhelmed, so those carry the most; "room to grow" barely matters *yet*, so it carries the least. |
|||||||
| 59 | 5. **Scale:** 0–5. |
|||||||
| 60 | 6. **Deliberate exclusion:** extension ecosystems and Git integration are reported but *not scored* — they matter enormously to a professional, and not at all to this student in week one. |
|||||||
| 61 | ||||||||
| 62 | | Criterion (weight) | Thonny | VS Code | IDLE | |
|||||||
| 63 | |---|---|---|---| |
|||||||
| 64 | | Getting started — time to first run (30%) | 5 | 2 | 4 | |
|||||||
| 65 | | Simple, uncluttered interface (25%) | 5 | 2 | 4 | |
|||||||
| 66 | | Helpful for beginner errors / step-debugging (25%) | 5 | 3 | 2 | |
|||||||
| 67 | | Runs on the school laptop (12%) | 4 | 3 | 5 | |
|||||||
| 68 | | Room to grow (8%) | 3 | 5 | 2 | |
|||||||
| 69 | | **Weighted total / 5** | **4.7** | **2.6** | **3.5** | |
|||||||
| 70 | ||||||||
| 71 | *(Scores are illustrative judgements, not benchmarks.)* For *this* user, **Thonny** wins — which makes sense, because it was designed for exactly this person. |
|||||||
| 72 | ||||||||
| 73 | --- |
|||||||
| 74 | ||||||||
| 75 | ## Stress-test your matrix |
|||||||
| 76 | ||||||||
| 77 | The result above is only as good as the persona behind it. Pressure-test it three ways: |
|||||||
| 78 | ||||||||
| 79 | - **Swap the persona.** Re-run it for a Year 12 student mid-SAT, building a large program. Now "room to grow", debugging power and extensions matter far more, getting-started far less — re-weight, and **VS Code** climbs to the top. Same options, different user, opposite answer. That is the whole lesson. |
|||||||
| 80 | - **Try to game it.** Could someone inflate one weight to force their favourite to win? If a single criterion at 60% decides everything, your matrix is really a one-criterion decision wearing a disguise. |
|||||||
| 81 | - **Hunt the missing criterion.** What did you forget? Cost? Offline use? Accessibility? A matrix is only as honest as its least-considered gap. |
|||||||
| 82 | ||||||||
| 83 | --- |
|||||||
| 84 | ||||||||
| 85 | ## The quality factors are a menu, not a matrix |
|||||||
| 86 | ||||||||
| 87 | VCE names software quality factors — **effectiveness, efficiency, usability, reliability, maintainability, robustness, security**. Treat that list as a **menu to select from**, never a ready-made matrix. Your job is to *select* the few that matter to your user, *weight* them, and *justify* both choices. Tipping the whole menu into a table with equal weights is the opposite of evaluation. For the efficiency-versus-effectiveness distinction specifically, see [Efficiency vs Effectiveness](/sd/C04/Efficiency%20vs%20Effectiveness). |
|||||||
| 88 | ||||||||
| 89 | --- |
|||||||
| 90 | ||||||||
| 91 | ## Connection to your SAT |
|||||||
| 92 | ||||||||
| 93 | In **C4-2** you apply this exact skill to your *own* design ideas — building [your evaluation matrix](/sd/C04/C4-2%20Checkpoint%20-%20Evaluation%20Matrix) to decide which design proceeds. In **Unit 4 Outcome 1** you do it again to evaluate your finished solution. The discipline is the same every time: **decide and defend the criteria before you judge anything against them.** Criteria written to flatter a finished product fool no examiner. |
|||||||
| 94 | ||||||||
| 95 | --- |
|||||||
| 96 | ||||||||
| 97 | ## Practise |
|||||||
| 98 | ||||||||
| 99 | > A Year 12 student wants to choose note-taking software for VCE — **OneNote**, **Notion**, or **Obsidian**. Design an evaluation matrix for this user: define the persona precisely, choose and weight four or five criteria (with a measurement method for each), name **one** thing you deliberately exclude and why, then state which tool you would recommend (*explain*) and what you trade away by not choosing the runner-up (*justify*). |
|||||||
| 100 | ||||||||
| 101 | --- |
|||||||
| 102 | ||||||||
| 103 | ## See also |
|||||||
| 104 | ||||||||
| 105 | - [C4-2 Checkpoint — Your Evaluation Matrix](/sd/C04/C4-2%20Checkpoint%20-%20Evaluation%20Matrix) — build the matrix for your own design ideas |
|||||||
| 106 | - [Evaluation Criteria and the SRS](/sd/C04/Evaluation%20Criteria%20and%20the%20SRS) · [Efficiency vs Effectiveness](/sd/C04/Efficiency%20vs%20Effectiveness) |
|||||||
| 107 | - [Evaluation vs Testing](/sd/C04/Evaluation%20vs%20Testing) |
|||||||
| 108 | ||||||||
| 109 | --- |
|||||||
| 110 | ||||||||
| 111 | ← Back to [C04 Home](/sd/C04/C04-home) · [VCE Software Development Hub](/sd/VCE%20Software%20Development%20Hub) |
|||||||
