Commit 408909
2026-05-25 22:25:13 lisa: SD C04: add 'Designing an Evaluation Matrix' guide (first-principles, CHOICE model) Teaches matrix DESIGN as seven justified decisions, reverse-engineering the CHOICE May-2026 headphones test, with a worked Thonny/VS Code/IDLE example (persona-driven, stress-tested), the quality-factors-as-menu point, a verified weighted-matrix video, and an exam-style practice prompt. Linked from C04-home and the C4-2 checkpoint. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>| sd/C04/C04-home.md .. | |
| @@ 25,6 25,7 @@ | |
| - [Evaluation vs Testing](/sd/C04/Evaluation%20vs%20Testing) — why a solution can pass all tests and still fail evaluation; the exact distinction the 7–10 band requires you to hold | |
| - [Efficiency vs Effectiveness](/sd/C04/Efficiency%20vs%20Effectiveness) — side-by-side scenario; how each term maps to separate evaluation matrix tables; why getting this wrong costs marks | |
| - [Evaluation Criteria and the SRS](/sd/C04/Evaluation%20Criteria%20and%20the%20SRS) — why every criterion must trace to a numbered SRS requirement, with worked counter-examples and a pre-submission checklist | |
| + | - [Designing an Evaluation Matrix](/sd/C04/Designing%20an%20Evaluation%20Matrix) — how to *design* a matrix from first principles (seven justified decisions), reverse-engineering a CHOICE magazine test; for C4-2 and U4 O1 |
| - [C04 → C05 — Choose, then Detail](/sd/C05/C04%20to%20C05%20-%20Choose%20then%20Detail) — how this criterion (choose *which* idea) differs from C05 (detail the chosen idea); breadth vs depth | |
| ## Resources | |
| sd/C04/C4-2 Checkpoint - Evaluation Matrix.md .. | |
| @@ 91,6 91,7 @@ | |
| ## See also | |
| + | - [Designing an Evaluation Matrix](/sd/C04/Designing%20an%20Evaluation%20Matrix) — design your criteria and weights from first principles (the *why* behind the matrix) |
| - [C4-1 Checkpoint — Your Design Pack](/sd/C04/C4-1%20Checkpoint%20-%20Design%20Pack) — the other half of C4 | |
| - [Efficiency vs Effectiveness](/sd/C04/Efficiency%20vs%20Effectiveness) · [Evaluation Criteria and the SRS](/sd/C04/Evaluation%20Criteria%20and%20the%20SRS) | |
| - [C4 Design Ideas and Evaluation Explained](/sd/C4%20Design%20Ideas%20and%20Evaluation%20Explained) — the walkthrough videos | |
| /dev/null .. sd/C04/Designing an Evaluation Matrix.md | |
| @@ 0,0 1,111 @@ | |
| + | > **DRAFT** — under teacher review. |
| + | |
| + | # Designing an Evaluation Matrix |
| + | |
| + | The Hamilton and Alexandra College · Year 12 · 2026 |
| + | |
| + | Most students learn to *fill in* an evaluation matrix. Far fewer learn to *design* one — and that is the harder, more valuable skill. A good matrix is not a neutral measuring device that fell from the sky; it is a **series of justified design decisions**. Change who you are evaluating for, and a sensible matrix changes with it. |
| + | |
| + | This matters directly for **Unit 4 Outcome 1**, where you evaluate your own solution against criteria *you* chose and *you* must justify. The single most common failure is writing criteria *after* the solution is built, to flatter what you happened to make. Designing the matrix first — and defending it — is how you avoid that. |
| + | |
| + | --- |
| + | |
| + | ## Start by taking a great matrix apart |
| + | |
| + | Before you build one, reverse-engineer one that professionals already trust. **CHOICE magazine's May 2026 headphones test** is an excellent model. Look past the scores and ask *why the matrix is shaped the way it is*: |
| + | |
| + | - It produces a single composite **CHOICE Expert Rating** from **six weighted criteria**: sound 45%, noise cancelling 15%, ease of use 17.5%, battery 10%, comfort 7.5%, durability 5%. |
| + | - The weights are not equal, and they are not arbitrary — they encode a **point of view**. Sound at 45% says, plainly, "for this product, how it sounds matters more than everything else combined nearly does." |
| + | - Two things are measured but **deliberately excluded** from the score: microphone quality and ease of cleaning. CHOICE reports them because some readers care, but decided they should not move the headline rating. |
| + | - The criteria suit a **specific user context** (here, travel — which is why noise cancelling earns its own 15%). |
| + | - Different criteria use **different measurement methods**: lab instruments for sound, a subjective panel for comfort, a spec sheet for battery. |
| + | |
| + | Every one of those is a decision someone made and could defend. That is what you are learning to do. |
| + | |
| + | --- |
| + | |
| + | ## The seven design decisions |
| + | |
| + | Designing a matrix means making — and justifying — these seven choices, roughly in order: |
| + | |
| + | 1. **Define the user.** Criteria flow from *who* you are evaluating for. "Best headphones" is meaningless; "best headphones *for a daily train commuter*" tells you what to measure. |
| + | 2. **Choose criteria that matter to that user.** Not everything that *can* be measured *should* be. Pick the handful that decide the outcome for this person. |
| + | 3. **Decide how each criterion is measured.** Objective lab test, subjective rating, feature checklist, or spec sheet — and say which, because it changes how much to trust the number. |
| + | 4. **Assign weights, and justify them.** Weights are where your point of view lives. Equal weights are themselves a claim (that everything matters the same) — usually a weak one. |
| + | 5. **Choose a scale.** 0–5, 0–100, or a traffic-light system. Be consistent so scores are comparable across options. |
| + | 6. **Decide what to exclude — on purpose.** Naming what you *leave out*, and why, is a mark of a considered matrix (CHOICE excluding microphone quality). Silent omissions look like oversights; stated exclusions look like judgement. |
| + | 7. **Stress-test it.** Swap the user, try to game it, and hunt for the criterion you forgot. A matrix that survives this is one you can defend. |
| + | |
| + | --- |
| + | |
| + | ## Watch: how weighting works (5 min) |
| + | |
| + | **🎯 Watch for:** how the *weights* — not the scores — are where "what matters most to this user" actually gets encoded. |
| + | |
| + | {{Video|src=https://www.youtube.com/watch?v=hUlB8lkfw9Q}} |
| + | |
| + | *The Weighted Scoring Decision Matrix — airfocus. (It's framed for product managers; the method is exactly the same for software evaluation.)* |
| + | |
| + | --- |
| + | |
| + | ## Worked example — choosing a Python IDE |
| + | |
| + | Let's make all seven decisions explicit, using three tools you might actually meet: **Thonny**, **VS Code**, and **IDLE**. |
| + | |
| + | 1. **User:** a Year 11 student in their *first* programming class, on a school laptop, who needs to get a program running quickly without being overwhelmed. |
| + | 2. **Criteria that matter to them:** how fast you get started, how simple the interface is, how helpful it is for a beginner's errors, whether it runs on school hardware, and how much room there is to grow. |
| + | 3. **Measurement method:** *getting started* = time to first running program (objective); *simple interface* and *beginner debugging* = a subjective rating against a checklist; *runs on school laptop* = does it install and run within the SOE; *room to grow* = a feature checklist. |
| + | 4. **Weights, justified:** a first-week beginner is won or lost on getting started and not being overwhelmed, so those carry the most; "room to grow" barely matters *yet*, so it carries the least. |
| + | 5. **Scale:** 0–5. |
| + | 6. **Deliberate exclusion:** extension ecosystems and Git integration are reported but *not scored* — they matter enormously to a professional, and not at all to this student in week one. |
| + | |
| + | | Criterion (weight) | Thonny | VS Code | IDLE | |
| + | |---|---|---|---| |
| + | | Getting started — time to first run (30%) | 5 | 2 | 4 | |
| + | | Simple, uncluttered interface (25%) | 5 | 2 | 4 | |
| + | | Helpful for beginner errors / step-debugging (25%) | 5 | 3 | 2 | |
| + | | Runs on the school laptop (12%) | 4 | 3 | 5 | |
| + | | Room to grow (8%) | 3 | 5 | 2 | |
| + | | **Weighted total / 5** | **4.7** | **2.6** | **3.5** | |
| + | |
| + | *(Scores are illustrative judgements, not benchmarks.)* For *this* user, **Thonny** wins — which makes sense, because it was designed for exactly this person. |
| + | |
| + | --- |
| + | |
| + | ## Stress-test your matrix |
| + | |
| + | The result above is only as good as the persona behind it. Pressure-test it three ways: |
| + | |
| + | - **Swap the persona.** Re-run it for a Year 12 student mid-SAT, building a large program. Now "room to grow", debugging power and extensions matter far more, getting-started far less — re-weight, and **VS Code** climbs to the top. Same options, different user, opposite answer. That is the whole lesson. |
| + | - **Try to game it.** Could someone inflate one weight to force their favourite to win? If a single criterion at 60% decides everything, your matrix is really a one-criterion decision wearing a disguise. |
| + | - **Hunt the missing criterion.** What did you forget? Cost? Offline use? Accessibility? A matrix is only as honest as its least-considered gap. |
| + | |
| + | --- |
| + | |
| + | ## The quality factors are a menu, not a matrix |
| + | |
| + | VCE names software quality factors — **effectiveness, efficiency, usability, reliability, maintainability, robustness, security**. Treat that list as a **menu to select from**, never a ready-made matrix. Your job is to *select* the few that matter to your user, *weight* them, and *justify* both choices. Tipping the whole menu into a table with equal weights is the opposite of evaluation. For the efficiency-versus-effectiveness distinction specifically, see [Efficiency vs Effectiveness](/sd/C04/Efficiency%20vs%20Effectiveness). |
| + | |
| + | --- |
| + | |
| + | ## Connection to your SAT |
| + | |
| + | In **C4-2** you apply this exact skill to your *own* design ideas — building [your evaluation matrix](/sd/C04/C4-2%20Checkpoint%20-%20Evaluation%20Matrix) to decide which design proceeds. In **Unit 4 Outcome 1** you do it again to evaluate your finished solution. The discipline is the same every time: **decide and defend the criteria before you judge anything against them.** Criteria written to flatter a finished product fool no examiner. |
| + | |
| + | --- |
| + | |
| + | ## Practise |
| + | |
| + | > A Year 12 student wants to choose note-taking software for VCE — **OneNote**, **Notion**, or **Obsidian**. Design an evaluation matrix for this user: define the persona precisely, choose and weight four or five criteria (with a measurement method for each), name **one** thing you deliberately exclude and why, then state which tool you would recommend (*explain*) and what you trade away by not choosing the runner-up (*justify*). |
| + | |
| + | --- |
| + | |
| + | ## See also |
| + | |
| + | - [C4-2 Checkpoint — Your Evaluation Matrix](/sd/C04/C4-2%20Checkpoint%20-%20Evaluation%20Matrix) — build the matrix for your own design ideas |
| + | - [Evaluation Criteria and the SRS](/sd/C04/Evaluation%20Criteria%20and%20the%20SRS) · [Efficiency vs Effectiveness](/sd/C04/Efficiency%20vs%20Effectiveness) |
| + | - [Evaluation vs Testing](/sd/C04/Evaluation%20vs%20Testing) |
| + | |
| + | --- |
| + | |
| + | ← Back to [C04 Home](/sd/C04/C04-home) · [VCE Software Development Hub](/sd/VCE%20Software%20Development%20Hub) |
