<!-- Generated from applied-computing-au vic/unit3-4/sat/C09-2026 by port-reference-godot-to-wiki.py โ€” do not hand-edit; re-run the port. -->
# 9-1 Beta Testing Plan โ€” Worked Example (Godot Tetris)

> **How to use this:** this is a *completed* plan at the scale your SAT expects โ€” 3 testers, 3+ collection methods, a testing window of half a week plus a weekend. Read it beside `C091-Beta-Testing-Plan-Template.md` (same sections, filled in), then write the same kind of plan for **your** solution against **your** SRS. The ๐Ÿ’ก boxes explain what each section is doing for the score โ€” leave them out of your own plan.

---

## Beta Testing Plan For:

**Software Solution Name:** Godot Tetris โ€” a modern falling-blocks game (solo + two-player LAN versus), built in Godot 4.7

**Target Users:** students aged 12โ€“18 who play casual games on school laptops; two players sharing a LAN for versus mode

---

## 1. Objective

**Primary Goal:** Confirm that the finished game is playable, understandable and stable in the hands of real players โ€” the interface reads clearly, the modern mechanics (hold, ghost piece, hard drop, wall kicks) behave as players expect, and a two-player LAN match runs start-to-finish without failure.

**Testing Focus:**

- [x] **Appearance** (interface design, visual consistency)
- [x] **Functionality** (core features working correctly)
- [x] **User Experience** (ease of use, workflow efficiency)
- [x] **Performance** (speed, reliability, compatibility)

> ๐Ÿ’ก *The objective names the actual solution and what "working" means for it โ€” not "test the game to find bugs". It targets appearance AND functionality AND requirements: that's the 3โ€“4, 5โ€“6 and 7โ€“8 band behaviours in one sentence each.*

---

## 2. Test Scenarios

### Scenario 1 โ€” Appearance Testing: "Can you read the board?"

*Targets SRS non-functional requirements: usability (visual clarity), consistency of the interface.*

Testers judge whether the play field, next-piece preview, hold slot, ghost piece, score and level display communicate without explanation.

**Detailed User Steps:**

- **Step 1 โ€” Preparation:** open the game on a school laptop; do not explain anything.
- **Step 2 โ€” Starting Activity:** tester starts a solo game from the menu unaided.
- **Step 3 โ€” During Activity:** after two minutes, tester points at each on-screen element and says what they think it shows (score, level, next piece, hold slot, ghost outline).
- **Step 4 โ€” Ending Activity:** tester plays until game over and describes what the game-over screen tells them.
- **Step 5 โ€” Post-Activity Review:** short survey items on visual clarity (1โ€“5 scales).

**Specific Feedback Points:** Was the ghost piece understood without being told? Is the hold slot's "once per piece" state visible? Does the score/level area draw attention when a level-up happens?

### Scenario 2 โ€” Functionality Testing: "Do the mechanics do what players expect?"

*Targets SRS functional requirements: rotation with wall kicks (FR-3), hold (FR-5), hard/soft drop and scoring (FR-7), line clears and levelling (FR-8).*

**Detailed User Steps:**

- **Step 1 โ€” Preparation:** hand the tester the one-page controls card (โ†/โ†’ move, โ†“ soft drop, โ†‘/X and Z rotate, Space hard drop, C/Shift hold).
- **Step 2 โ€” Starting Activity:** tester attempts each control once in their first game.
- **Step 3 โ€” During Activity:** set tasks โ€” rotate a piece flush against the wall (wall kick), hold a piece and retrieve it, clear two lines with one drop, reach level 2.
- **Step 4 โ€” Ending Activity:** observer checks the final score against the scoring table (100/300/500/800 ร— level, +1 per soft-drop cell, +2 per hard-drop cell).
- **Step 5 โ€” Post-Activity Review:** tester reports anything that "felt wrong" (a rotation that refused, a piece that locked too early, a hold that didn't respond).

**Specific Feedback Points:** Did any rotation near the wall surprise the tester? Did the 0.5 s lock delay feel fair or frustrating at speed? Did the displayed score match the events observed?

### Scenario 3 โ€” User Experience Testing: "A full versus match, cold"

*Targets UX characteristics: learnability, efficiency, error tolerance. Targets SRS reliability requirement: a LAN session survives a complete match.*

**Detailed User Steps:**

- **Step 1 โ€” Preparation:** two testers, two laptops, same LAN; neither has played versus mode.
- **Step 2 โ€” Starting Activity:** testers follow the lobby screen to connect to each other unaided (host + join).
- **Step 3 โ€” During Activity:** play one full match; observer logs any confusion, disconnection, or garbage-row event the players don't understand.
- **Step 4 โ€” Ending Activity:** the match ends; testers state who won and how they know.
- **Step 5 โ€” Post-Activity Review:** paired interview โ€” what nearly stopped you, what would you change?

**Specific Feedback Points:** Time from "open game" to "match running" without help; whether incoming garbage rows read as an attack; whether the win/lose screen is unambiguous.

> ๐Ÿ’ก *Scenarios are real tasks with steps a stranger could run, not "click each button". Each names the FR/NFRs it exercises โ€” that's the 7โ€“8 band โ€” and Scenario 3 targets UX characteristics by name, which is the 9โ€“10 band behaviour.*

---

## 3. Potential Users

| User | Who are they? | Why selected? | Available when? |
|------|---------------|---------------|-----------------|
| **User 1** | Year 8 student, plays mobile puzzle games, has never played Tetris | Reads the interface with fresh eyes โ€” the learnability test can't be faked with an experienced player | Lunchtimes this week |
| **User 2** | Year 11 student, experienced Tetris player (plays online guideline Tetris) | Knows how hold, ghost and wall kicks *should* behave, so deviations from expectations surface immediately | After school Thu/Fri |
| **User 3** | Parent, plays no games, uses a laptop daily for work | Extreme-novice check on menus, controls card and game-over flow; also my weekend tester | Saturday |

**Why these users represent your target audience:** the target users are casual players on school laptops โ€” Users 1 and 2 bracket that range (novice โ†’ expert), and User 3 tests whether the interface survives someone outside it. Users 1+2 together also form the LAN pair for Scenario 3.

> ๐Ÿ’ก *Each tester has a reason tied to what they reveal โ€” that's "explains why potential users have been selected" (5โ€“6 band). Role descriptions are fine; full names are not required.*

---

## 4. Methodology

### User Recruitment:

Ask in person this week; confirm each session time by message the day before. Consent forms handed out and **signed before any session is recorded** โ€” no consent form, no session video.

### Data Collection Methods:

- [x] **Direct observation** (watch users during testing)
- [x] **Interviews** (verbal feedback sessions)
- [x] **Surveys/questionnaires** (structured feedback forms)
- [x] **Error logging** (track problems encountered)

### How Results Will Be Collected:

**Method 1:** Observation sheet (printed, one per session) โ€” one row per event: time, what happened, tester reaction.

**Method 2:** Post-session survey (Google Form, 10 items: 1โ€“5 scales on clarity/controls/fun + two open questions) โ€” link sent as the session ends.

**Method 3:** Recorded interview (audio) using the six-question interview sheet; ~5 minutes per tester.

**Method 4:** Error log (spreadsheet) โ€” every crash, disconnect, or "that's wrong" moment with steps to reproduce.

### Data Validation Methods:

**Comparison Method:** survey answers cross-checked against what the observation sheet actually recorded โ€” a tester who rates controls 5/5 but fumbled hold for three pieces gets a follow-up question in the interview.

**Known Benchmarks:** final scores checked against the scoring table; gravity/level progression checked against the level formula (level = 1 + lines รท 10).

**External Validation:** User 2's expectations from mainstream guideline Tetris act as the reference for "standard" mechanic behaviour.

> ๐Ÿ’ก *Four methods, and every method names its instrument โ€” a survey that exists beats a "survey" that doesn't. Instruments must be ready before the first session: that's what the simulation checks.*

---

## 5. Timeline

| Phase | Duration | Activities |
|-------|----------|------------|
| **Setup** | ยฝ day | Print observation sheets + controls cards, build the survey form, collect signed consent forms, test the LAN pairing on two school laptops |
| **Active Testing** | 3 days (two weekdays + Saturday) | Users 1 & 2 individually (Scenarios 1โ€“2), Users 1+2 paired (Scenario 3), User 3 on Saturday (Scenarios 1โ€“2) |
| **Results & Feedback** | 1 day | File raw data into the evidence folder, complete the error log, write the deviation note |

> ๐Ÿ’ก *Total: half a week plus a weekend โ€” the scale the assessment expects. Three testers, four sessions. A 6-week 12-tester plan (like the EasyRetail commercial exemplar) would fail the "realistic for the window" check.*

---

## 6. Resources

### Hardware/Software Requirements:

Two school laptops on the same LAN (versus needs both); the exported game build installed on each; any keyboard works โ€” controls use physical key positions, not letters.

### Support Materials for Users:

One-page controls card; consent form; the tester never sees the code or the plan.

---

## 7. Success Criteria

### Primary Success Measures:

- Every tester starts a solo game unaided in under 1 minute (learnability).
- All set mechanic tasks (wall-kick rotation, hold, double line clear, reach level 2) completed by Users 1 and 2.
- One full LAN versus match completes with no disconnection and an unambiguous result.
- Displayed scores match the scoring table in every observed session.

### User Satisfaction Targets:

- Visual clarity and controls both average โ‰ฅ 4/5 on the survey.
- No tester abandons a session.

### Decision Framework:

**What results would indicate success?** All primary measures met and the error log holds only minor items โ†’ proceed to recommendations with priorities from the survey's open questions.

**What results would require major changes?** Any crash or LAN failure, a mechanic that confused both novice testers, or scores that don't match the table โ†’ these become the top recommended modifications in the C9-3/C9-4 report.

> ๐Ÿ’ก *Success criteria are countable afterwards โ€” "โ‰ฅ 4/5", "under 1 minute", "no disconnection" โ€” so the report can say* whether *the test passed, not just how it felt.*

---

## Quality Checklist

**C9-1 Assessment Criteria:**

- [x] **C9-1-1:** Components clearly identified for testing
- [x] **C9-1-3:** Plan targets software appearance AND outlines potential users
- [x] **C9-1-5:** Plan targets functionality AND explains why users were selected
- [x] **C9-1-7:** Plan targets functional AND non-functional requirements AND documents how results will be collected
- [x] **C9-1-9:** Test scenarios target user experience characteristics AND documentation is clear and concise

**Professional Standards:**

- [x] All user selections include clear rationale (**"why selected"**)
- [x] Multiple data collection methods documented (**"how collected"**)
- [x] Test scenarios focus on relevant characteristics of your software solution
- [x] Timeline is realistic for user coordination and testing
- [x] All sections completed with clear, professional information
0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9