Version: 2.0. Audience: Classes 8–10 first, with adaptable notes for Classes 11–12. Status: school pilot resource.
> Educational resource only. Not legal advice, not regulatory certification, not emergency or mental-health support.
The Scenario Lab presents six missions. Each mission works through the five GUARD stages — Goal, Understand limits, Assess evidence, Respect people, Decide responsibly — as one decision each, for 30 scored decisions in total. Every decision offers three options; most award 0, 1 or 2 points, though two decisions in this set are treated as binary (0 or 2) because the remaining options carry a safety or integrity risk too serious for partial credit.
A score of 2 reflects the strongest response available in the scenario as written. A score of 1 is a defensible but incomplete response — a teacher can still credit sound reasoning at this level in discussion, even though the tool does not award full marks. A score of 0 reflects a response that ignores or misapplies the relevant GUARD stage.
Add the five stage scores for a mission to get a score out of 10. Across all six missions, the maximum is 60. Use scores as a formative signal of reasoning quality, not a pass/fail gate:
Two students can reach the same score by different, both defensible, routes — GUARD assesses judgement under ambiguity, not a single correct keyword. Before treating a score as final:
Important: a single simulation score is not proof of broad AI competence. Use it alongside the diagnostic, Decision Record, AI-Use Log and capstone as converging evidence, per the GUARD Assessment Rubric.
Teacher override: a teacher's judgement of a student's spoken or written reasoning always outweighs the automated score, particularly for students who reasoned well but selected an imperfect option under time pressure, or who reasoned poorly but selected the labelled "correct" option by guessing.
Theme: Evidence · Level: Foundation · 7 min
You are preparing a history presentation. An AI assistant supplies a dramatic quotation supposedly spoken by Mahatma Gandhi in 1946. It fits your argument perfectly, but your teacher expects every quotation to be traceable.
Evidence pack:
What to look for: Separate the learning goal from the convenience of a ready-made sentence.
| Score | Option | Why |
|---|---|---|
| 0 | Make the presentation sound impressive, even if the exact source is hard to find. | Impact is not the learning goal if the evidence cannot be defended. |
| 2 | Explain Gandhi's documented ideas accurately and support them with verifiable evidence. | Strong. The goal is understanding and evidence, not keeping a convenient quotation. |
| 1 | Finish the slides quickly and add the sources afterwards. | Speed matters, but delaying verification makes weak evidence harder to notice. |
Model reasoning (2): Explain Gandhi's documented ideas accurately and support them with verifiable evidence. — Strong. The goal is understanding and evidence, not keeping a convenient quotation.
Acceptable alternative (1): Finish the slides quickly and add the sources afterwards. is defensible as partial credit — Speed matters, but delaying verification makes weak evidence harder to notice. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Make the presentation sound impressive, even if the exact source is hard to find. — Impact is not the learning goal if the evidence cannot be defended.
What to look for: A fluent answer can combine plausible details into a source that does not exist.
| Score | Option | Why |
|---|---|---|
| 1 | The AI may use an old spelling of a place name. | Possible, but the central risk is fabrication of both quotation and citation. |
| 2 | The AI can invent a quotation and attach precise-looking bibliographic details. | Exactly. Precision of wording and page numbers is not proof of authenticity. |
| 0 | The AI is neutral because it has no personal opinion. | Lack of personal intent does not make an output accurate or unbiased. |
Model reasoning (2): The AI can invent a quotation and attach precise-looking bibliographic details. — Exactly. Precision of wording and page numbers is not proof of authenticity.
Acceptable alternative (1): The AI may use an old spelling of a place name. is defensible as partial credit — Possible, but the central risk is fabrication of both quotation and citation. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): The AI is neutral because it has no personal opinion. — Lack of personal intent does not make an output accurate or unbiased.
What to look for: Independent evidence should lead back to the original record or a trusted scholarly source.
| Score | Option | Why |
|---|---|---|
| 0 | Ask the same AI to confirm that the quotation is real. | The same system can repeat its own error. Repetition is not independent verification. |
| 1 | Accept it if several search-result snippets contain the same words. | Repeated pages may copy one another. Trace the claim to an authoritative record. |
| 2 | Open the cited volume or a trusted archive and locate the quotation in its original context. | Strong. This directly tests the citation against independent evidence. |
Model reasoning (2): Open the cited volume or a trusted archive and locate the quotation in its original context. — Strong. This directly tests the citation against independent evidence.
Acceptable alternative (1): Accept it if several search-result snippets contain the same words. is defensible as partial credit — Repeated pages may copy one another. Trace the claim to an authoritative record. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Ask the same AI to confirm that the quotation is real. — The same system can repeat its own error. Repetition is not independent verification.
What to look for: Respect includes representing a person and the historical record fairly.
| Score | Option | Why |
|---|---|---|
| 0 | Nobody, because the person is historical and the assignment is only for school. | False attribution still misrepresents a person and misleads classmates. |
| 2 | The historical figure, classmates and anyone who later repeats the false quotation. | Yes. Respect includes accuracy about people and preventing misinformation from spreading. |
| 1 | Only the teacher who has to mark the work. | The teacher is affected, but the harm extends to the people represented and the audience. |
Model reasoning (2): The historical figure, classmates and anyone who later repeats the false quotation. — Yes. Respect includes accuracy about people and preventing misinformation from spreading.
Acceptable alternative (1): Only the teacher who has to mark the work. is defensible as partial credit — The teacher is affected, but the harm extends to the people represented and the audience. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Nobody, because the person is historical and the assignment is only for school. — False attribution still misrepresents a person and misleads classmates.
What to look for: A responsible decision can reject one output without rejecting all AI assistance.
| Score | Option | Why |
|---|---|---|
| 1 | Use the quotation but add 'source not fully confirmed' in small text. | A caveat does not make an unverified quotation suitable as evidence. |
| 2 | Remove it, replace it with verified evidence, and disclose that AI helped with initial research. | Strong. This protects the evidence standard while honestly describing AI's role. |
| 0 | Keep it and omit the AI use so nobody becomes suspicious. | Concealing both the source problem and AI use compounds the integrity failure. |
Model reasoning (2): Remove it, replace it with verified evidence, and disclose that AI helped with initial research. — Strong. This protects the evidence standard while honestly describing AI's role.
Acceptable alternative (1): Use the quotation but add 'source not fully confirmed' in small text. is defensible as partial credit — A caveat does not make an unverified quotation suitable as evidence. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Keep it and omit the AI use so nobody becomes suspicious. — Concealing both the source problem and AI use compounds the integrity failure.
Model decision for this mission: Do not use the quotation as genuine. Replace it with traceable evidence, keep a record of the verification, and disclose AI assistance according to the teacher's rules.
Teacher discussion prompts:
Theme: Deepfakes · Level: Intermediate · 8 min
A class group receives a 19-second audio clip that sounds like your principal insulting scholarship students. Friends are angry and ask you to repost it publicly before the school can 'delete the evidence'. Nobody knows who recorded it.
Evidence pack:
What to look for: The first goal is not proving a side right; it is preventing avoidable harm while establishing what happened.
| Score | Option | Why |
|---|---|---|
| 0 | Make the clip trend before it can be removed. | Rapid amplification can create irreversible harm before authenticity is known. |
| 2 | Protect affected students, preserve the evidence, and establish authenticity through responsible channels. | Strong. It combines care for people with a disciplined verification goal. |
| 1 | Prove that the principal could never have said it. | That assumes the conclusion. The goal is a fair, evidence-led response. |
Model reasoning (2): Protect affected students, preserve the evidence, and establish authenticity through responsible channels. — Strong. It combines care for people with a disciplined verification goal.
Acceptable alternative (1): Prove that the principal could never have said it. is defensible as partial credit — That assumes the conclusion. The goal is a fair, evidence-led response. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Make the clip trend before it can be removed. — Rapid amplification can create irreversible harm before authenticity is known.
What to look for: Neither human hearing nor an automated detector can establish authenticity with certainty from a short clip.
| Score | Option | Why |
|---|---|---|
| 0 | If it sounds exactly like the principal, it must be authentic. | Voice cloning and selective editing can both produce convincing audio. |
| 1 | One online deepfake detector can give a final yes-or-no answer. | A detector may add evidence, but its result is uncertain and should not be treated as a verdict. |
| 2 | The clip may be genuine, generated or edited; a short copy may lack the information needed to tell. | Correct. Responsible action must account for uncertainty rather than pretending it has vanished. |
Model reasoning (2): The clip may be genuine, generated or edited; a short copy may lack the information needed to tell. — Correct. Responsible action must account for uncertainty rather than pretending it has vanished.
Acceptable alternative (1): One online deepfake detector can give a final yes-or-no answer. is defensible as partial credit — A detector may add evidence, but its result is uncertain and should not be treated as a verdict. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): If it sounds exactly like the principal, it must be authentic. — Voice cloning and selective editing can both produce convincing audio.
What to look for: Check origin, context and corroboration—not merely visual or audio 'tells'.
| Score | Option | Why |
|---|---|---|
| 0 | Count how many people in the group believe it. | Popularity measures reaction, not authenticity. |
| 2 | Preserve the file, identify the earliest source, seek the full context and use an authorised school investigation. | Strong. Provenance, context and accountable investigation are more reliable than a snap judgement. |
| 1 | Look at the waveform and decide whether any section appears unusual. | This can be one clue, but ordinary editing also changes waveforms and absence of a clue proves little. |
Model reasoning (2): Preserve the file, identify the earliest source, seek the full context and use an authorised school investigation. — Strong. Provenance, context and accountable investigation are more reliable than a snap judgement.
Acceptable alternative (1): Look at the waveform and decide whether any section appears unusual. is defensible as partial credit — This can be one clue, but ordinary editing also changes waveforms and absence of a clue proves little. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Count how many people in the group believe it. — Popularity measures reaction, not authenticity.
What to look for: Both the students named by the clip and the person represented by the voice may be harmed.
| Score | Option | Why |
|---|---|---|
| 2 | Reposting could expose scholarship students to stigma and falsely damage a person's reputation. | Yes. Uncertainty requires care for everyone affected, not indifference to harm. |
| 0 | There is no privacy issue because the principal is a public figure inside the school. | Public responsibility does not remove rights to context, fairness or protection from false attribution. |
| 1 | Delete the message immediately so nobody can investigate it. | Do not recirculate it, but an authorised investigator may need the original evidence preserved privately. |
Model reasoning (2): Reposting could expose scholarship students to stigma and falsely damage a person's reputation. — Yes. Uncertainty requires care for everyone affected, not indifference to harm.
Acceptable alternative (1): Delete the message immediately so nobody can investigate it. is defensible as partial credit — Do not recirculate it, but an authorised investigator may need the original evidence preserved privately. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): There is no privacy issue because the principal is a public figure inside the school. — Public responsibility does not remove rights to context, fairness or protection from false attribution.
What to look for: Do not amplify uncertain harmful media. Preserve, report and support.
| Score | Option | Why |
|---|---|---|
| 0 | Post it publicly with a question mark so others can decide. | A question mark does not undo amplification or reputational harm. |
| 2 | Do not forward it; preserve the original privately, report it to a designated adult, and support targeted students. | Strong. This enables investigation without turning students into distributors of possible abuse. |
| 1 | Run one detector and repost it if the score is above 80%. | Detector scores are not courtroom verdicts and cannot remove the duty to use responsible channels. |
Model reasoning (2): Do not forward it; preserve the original privately, report it to a designated adult, and support targeted students. — Strong. This enables investigation without turning students into distributors of possible abuse.
Acceptable alternative (1): Run one detector and repost it if the score is above 80%. is defensible as partial credit — Detector scores are not courtroom verdicts and cannot remove the duty to use responsible channels. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Post it publicly with a question mark so others can decide. — A question mark does not undo amplification or reputational harm.
Model decision for this mission: Do not repost the audio. Preserve the original file and context privately, report it through the school's safeguarding or designated reporting channel, and support students targeted by the content while authenticity is investigated.
Safeguarding note: if a real deepfake, impersonation or harassment incident occurs, do not treat it as a classroom debate. Preserve evidence privately, report through the school's designated safeguarding channel, and support any student named in the content.
Teacher discussion prompts:
Theme: Privacy · Level: Foundation · 6 min
Your class is preparing a farewell poster. A free AI image app can turn a group photograph into superheroes. The photograph shows 31 students, school badges and a location board. You took the photograph, but nobody was asked about AI editing or uploading.
Evidence pack:
What to look for: State the creative goal without assuming that this particular data must be used.
| Score | Option | Why |
|---|---|---|
| 2 | Create a memorable farewell poster without using classmates' data beyond what they agreed to. | Strong. The creative aim remains, while the method is open to safer alternatives. |
| 0 | Upload the photo because it is already on my phone. | Possessing a copy does not establish permission to submit everyone to another service. |
| 1 | Keep the poster secret so the class gets a better surprise. | A surprise can be fun, but secrecy is not a reason to bypass consent. |
Model reasoning (2): Create a memorable farewell poster without using classmates' data beyond what they agreed to. — Strong. The creative aim remains, while the method is open to safer alternatives.
Acceptable alternative (1): Keep the poster secret so the class gets a better surprise. is defensible as partial credit — A surprise can be fun, but secrecy is not a reason to bypass consent. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Upload the photo because it is already on my phone. — Possessing a copy does not establish permission to submit everyone to another service.
What to look for: A creative output may be temporary while the submitted data is retained or reused.
| Score | Option | Why |
|---|---|---|
| 1 | It may draw some superhero costumes badly. | That affects quality, but data retention and control are more consequential here. |
| 2 | You may not know how long the image is kept, who can access it or whether it is reused. | Correct. Once uploaded, control over identifiable information may be difficult to recover. |
| 0 | Free apps cannot store uploads because users have not paid. | Price says nothing about retention, reuse or sharing practices. |
Model reasoning (2): You may not know how long the image is kept, who can access it or whether it is reused. — Correct. Once uploaded, control over identifiable information may be difficult to recover.
Acceptable alternative (1): It may draw some superhero costumes badly. is defensible as partial credit — That affects quality, but data retention and control are more consequential here. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Free apps cannot store uploads because users have not paid. — Price says nothing about retention, reuse or sharing practices.
What to look for: Look for concrete statements about collection, retention, deletion, reuse and age requirements.
| Score | Option | Why |
|---|---|---|
| 1 | Whether the app has more than four stars. | Reviews may describe usability, but they do not replace the service's data terms and school rules. |
| 2 | The privacy notice, age terms, retention/deletion choices and the school's approved-tool policy. | Strong. These checks address the actual information risk. |
| 0 | Whether a friend has used the same app without a problem. | A friend's experience does not reveal what happened to the submitted data. |
Model reasoning (2): The privacy notice, age terms, retention/deletion choices and the school's approved-tool policy. — Strong. These checks address the actual information risk.
Acceptable alternative (1): Whether the app has more than four stars. is defensible as partial credit — Reviews may describe usability, but they do not replace the service's data terms and school rules. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Whether a friend has used the same app without a problem. — A friend's experience does not reveal what happened to the submitted data.
What to look for: Taking a photograph is different from uploading it for automated processing and transformation.
| Score | Option | Why |
|---|---|---|
| 0 | Only mine, because I took the photograph. | Copyright or possession does not erase other people's privacy and consent interests. |
| 2 | The people shown, plus any school consent and safeguarding requirements. | Correct. The use affects every identifiable person, not only the photographer. |
| 1 | A simple class majority is enough for everyone, including those who object. | A majority vote should not automatically override an individual's reasonable privacy choice. |
Model reasoning (2): The people shown, plus any school consent and safeguarding requirements. — Correct. The use affects every identifiable person, not only the photographer.
Acceptable alternative (1): A simple class majority is enough for everyone, including those who object. is defensible as partial credit — A majority vote should not automatically override an individual's reasonable privacy choice. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Only mine, because I took the photograph. — Copyright or possession does not erase other people's privacy and consent interests.
What to look for: Look for a lower-data alternative before abandoning the creative project.
| Score | Option | Why |
|---|---|---|
| 0 | Upload it now and ask classmates only if someone complains. | Consent after exposure cannot reliably undo the upload or reuse. |
| 2 | Use an approved tool only with informed permission, or create fictional avatars that use no real faces. | Strong. It preserves the creative goal while minimising avoidable personal-data use. |
| 1 | Crop the school sign and upload all faces without asking. | Cropping reduces one risk but leaves identifiable faces and consent unresolved. |
Model reasoning (2): Use an approved tool only with informed permission, or create fictional avatars that use no real faces. — Strong. It preserves the creative goal while minimising avoidable personal-data use.
Acceptable alternative (1): Crop the school sign and upload all faces without asking. is defensible as partial credit — Cropping reduces one risk but leaves identifiable faces and consent unresolved. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Upload it now and ask classmates only if someone complains. — Consent after exposure cannot reliably undo the upload or reuse.
Model decision for this mission: Do not upload the class photograph without informed permission and an approved service. Prefer fictional or student-created avatars; if real images are genuinely necessary, follow school consent, minimisation and deletion requirements.
Teacher discussion prompts:
Theme: Learning integrity · Level: Foundation · 7 min
Your group tested how light affects plant growth. You collected the measurements but have not written the analysis. An AI tool offers a complete 900-word report. Your teacher permits brainstorming and grammar assistance but requires the reasoning to be your own and all AI use to be recorded.
Evidence pack:
What to look for: The product is a report; the learning goal is the thinking needed to produce it.
| Score | Option | Why |
|---|---|---|
| 2 | Demonstrate that I can interpret our evidence and explain its limits. | Strong. This identifies the capability the assignment is assessing. |
| 0 | Produce any document that looks like a science report. | Appearance alone would bypass the core learning goal. |
| 1 | Get the highest mark with the least work. | Marks matter, but this goal can encourage substitution for the ability being assessed. |
Model reasoning (2): Demonstrate that I can interpret our evidence and explain its limits. — Strong. This identifies the capability the assignment is assessing.
Acceptable alternative (1): Get the highest mark with the least work. is defensible as partial credit — Marks matter, but this goal can encourage substitution for the ability being assessed. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Produce any document that looks like a science report. — Appearance alone would bypass the core learning goal.
What to look for: Context that was never recorded cannot be recovered merely through fluent writing.
| Score | Option | Why |
|---|---|---|
| 2 | Why measurements differ, what happened during the experiment and which limitations matter. | Correct. The students who performed the experiment hold essential context. |
| 0 | How to write complete English sentences. | Language generation is precisely what such systems generally do well. |
| 1 | Whether the report should be exactly 900 words. | It can follow a length instruction; the deeper limitation is missing experimental context. |
Model reasoning (2): Why measurements differ, what happened during the experiment and which limitations matter. — Correct. The students who performed the experiment hold essential context.
Acceptable alternative (1): Whether the report should be exactly 900 words. is defensible as partial credit — It can follow a length instruction; the deeper limitation is missing experimental context. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): How to write complete English sentences. — Language generation is precisely what such systems generally do well.
What to look for: Return every claim to the original observations and assignment requirements.
| Score | Option | Why |
|---|---|---|
| 2 | Recalculate key values, compare every claim with the notes and flag unsupported assumptions. | Strong. Verification is grounded in the evidence you actually collected. |
| 0 | Assume the analysis is right if it uses scientific vocabulary. | Technical style can conceal unsupported assumptions or incorrect calculations. |
| 1 | Ask another AI whether the first report sounds correct. | A second model may help identify questions, but the experiment record is the authoritative evidence. |
Model reasoning (2): Recalculate key values, compare every claim with the notes and flag unsupported assumptions. — Strong. Verification is grounded in the evidence you actually collected.
Acceptable alternative (1): Ask another AI whether the first report sounds correct. is defensible as partial credit — A second model may help identify questions, but the experiment record is the authoritative evidence. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Assume the analysis is right if it uses scientific vocabulary. — Technical style can conceal unsupported assumptions or incorrect calculations.
What to look for: Shared work should represent each member's contribution and agreed use of tools.
| Score | Option | Why |
|---|---|---|
| 2 | Agree the AI use together, protect names/data, and represent everyone's contribution honestly. | Correct. Integrity and respect apply to collaborators as well as the final answer. |
| 0 | Submit the generated report first and tell the group afterwards. | That removes classmates' control over work submitted in their names. |
| 1 | Remove names from the prompt, then every other concern disappears. | Data minimisation helps privacy, but authorship, learning and permission still matter. |
Model reasoning (2): Agree the AI use together, protect names/data, and represent everyone's contribution honestly. — Correct. Integrity and respect apply to collaborators as well as the final answer.
Acceptable alternative (1): Remove names from the prompt, then every other concern disappears. is defensible as partial credit — Data minimisation helps privacy, but authorship, learning and permission still matter. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Submit the generated report first and tell the group afterwards. — That removes classmates' control over work submitted in their names.
What to look for: Use assistance to support the assessed thinking, not replace it.
| Score | Option | Why |
|---|---|---|
| 0 | Submit the complete AI draft unchanged because the data came from your group. | Owning the data does not make generated reasoning your demonstrated work. |
| 2 | Draft the reasoning from your evidence, then use AI for questions, structure or language within the stated rules and disclose it. | Strong. AI scaffolds the work without replacing the capability being assessed. |
| 0 | Paraphrase the AI report so it cannot be detected. | Hiding substitution does not restore authorship or learning. |
Model reasoning (2): Draft the reasoning from your evidence, then use AI for questions, structure or language within the stated rules and disclose it. — Strong. AI scaffolds the work without replacing the capability being assessed.
No partial-credit option: this decision is treated as binary — the two remaining options both carry safety or integrity risks serious enough to score 0. "Submit the complete AI draft unchanged because the data came from your group." — Owning the data does not make generated reasoning your demonstrated work. "Paraphrase the AI report so it cannot be detected." — Hiding substitution does not restore authorship or learning.
Model decision for this mission: Write and verify the scientific reasoning from the group's own record. Use AI only within the teacher's permitted boundaries—for example, to test structure or improve language—and record that assistance accurately.
Teacher discussion prompts:
Theme: Bias & fairness · Level: Advanced · 9 min
A school committee is considering an AI ranking tool to reduce 240 scholarship applications to 30 interviews. The vendor says it is 91% accurate at predicting first-year success. Its inputs include grades, attendance, postcode, school-club participation and an essay score.
Evidence pack:
What to look for: Efficiency is one constraint, not the mission of a scholarship programme.
| Score | Option | Why |
|---|---|---|
| 0 | Produce exactly 30 names as quickly as possible. | Speed alone ignores opportunity, fairness and the consequences of exclusion. |
| 2 | Identify eligible potential fairly while preserving meaningful review and appeal. | Strong. The goal reflects the scholarship's purpose and the stakes of false exclusion. |
| 1 | Achieve the highest possible overall prediction accuracy. | Accuracy may matter, but it does not by itself define fairness or programme purpose. |
Model reasoning (2): Identify eligible potential fairly while preserving meaningful review and appeal. — Strong. The goal reflects the scholarship's purpose and the stakes of false exclusion.
Acceptable alternative (1): Achieve the highest possible overall prediction accuracy. is defensible as partial credit — Accuracy may matter, but it does not by itself define fairness or programme purpose. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Produce exactly 30 names as quickly as possible. — Speed alone ignores opportunity, fairness and the consequences of exclusion.
What to look for: Past patterns can encode unequal opportunity and may not generalise to new applicants.
| Score | Option | Why |
|---|---|---|
| 2 | The training population is narrow and some features may act as proxies for unequal opportunity. | Correct. Both generalisation and proxy effects must be tested before consequential use. |
| 0 | A consistent algorithm cannot be biased because it applies the same formula to everyone. | Consistent rules can still reproduce unequal patterns or create unequal error rates. |
| 1 | The only limitation is that humans may ignore the tool's recommendation. | Human misuse matters, but the model and its data also require scrutiny. |
Model reasoning (2): The training population is narrow and some features may act as proxies for unequal opportunity. — Correct. Both generalisation and proxy effects must be tested before consequential use.
Acceptable alternative (1): The only limitation is that humans may ignore the tool's recommendation. is defensible as partial credit — Human misuse matters, but the model and its data also require scrutiny. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): A consistent algorithm cannot be biased because it applies the same formula to everyone. — Consistent rules can still reproduce unequal patterns or create unequal error rates.
What to look for: A single headline percentage hides the errors that matter to different groups.
| Score | Option | Why |
|---|---|---|
| 0 | A polished product demonstration and testimonials from another school. | Testimonials do not establish performance on this population or decision. |
| 2 | Local validation, group-specific error rates, calibration, feature analysis and comparison with a documented baseline. | Strong. This makes performance and trade-offs visible rather than accepting one aggregate number. |
| 1 | The total number of applications processed by the vendor. | Scale of use is not evidence that the model is valid or fair for this purpose. |
Model reasoning (2): Local validation, group-specific error rates, calibration, feature analysis and comparison with a documented baseline. — Strong. This makes performance and trade-offs visible rather than accepting one aggregate number.
Acceptable alternative (1): The total number of applications processed by the vendor. is defensible as partial credit — Scale of use is not evidence that the model is valid or fair for this purpose. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): A polished product demonstration and testimonials from another school. — Testimonials do not establish performance on this population or decision.
What to look for: A false negative can quietly remove a student from consideration without anyone seeing their case.
| Score | Option | Why |
|---|---|---|
| 2 | Applicants need notice, meaningful human review, accessible correction and a route to challenge exclusion. | Correct. These safeguards address both data errors and contestability of consequential decisions. |
| 0 | Keep the model secret so applicants cannot manipulate it. | Security may protect some details, but total opacity prevents accountability and correction. |
| 1 | Only students selected for interview need an explanation. | Students excluded by errors have the greatest need for correction and review. |
Model reasoning (2): Applicants need notice, meaningful human review, accessible correction and a route to challenge exclusion. — Correct. These safeguards address both data errors and contestability of consequential decisions.
Acceptable alternative (1): Only students selected for interview need an explanation. is defensible as partial credit — Students excluded by errors have the greatest need for correction and review. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Keep the model secret so applicants cannot manipulate it. — Security may protect some details, but total opacity prevents accountability and correction.
What to look for: High-stakes efficiency tools should not silently become final decision-makers.
| Score | Option | Why |
|---|---|---|
| 0 | Automatically reject everyone below the cut-off because the tool is more consistent. | Consistency does not justify automated exclusion without validated performance and safeguards. |
| 2 | Do not use it for automatic rejection; require independent validation and a controlled assistive pilot with human review and appeal. | Strong. The scope remains proportionate to the evidence and consequences. |
| 1 | Use automatic rejection but tell applicants that AI was involved. | Disclosure is useful, but it does not repair an unvalidated or unchallengeable decision process. |
Model reasoning (2): Do not use it for automatic rejection; require independent validation and a controlled assistive pilot with human review and appeal. — Strong. The scope remains proportionate to the evidence and consequences.
Acceptable alternative (1): Use automatic rejection but tell applicants that AI was involved. is defensible as partial credit — Disclosure is useful, but it does not repair an unvalidated or unchallengeable decision process. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Automatically reject everyone below the cut-off because the tool is more consistent. — Consistency does not justify automated exclusion without validated performance and safeguards.
Model decision for this mission: Do not permit automatic rejection. First validate the system on the relevant population and purpose, examine group-specific errors and proxies, and design meaningful human review, notice, correction and appeal. A constrained assistive pilot may then be considered.
Teacher discussion prompts:
Theme: Human override · Level: Intermediate · 7 min
Late at night, a friend messages that they feel unsafe and 'do not want to be here tomorrow'. They ask you not to tell anyone. You are frightened and paste a paraphrased version into an AI chatbot, which responds that they are probably just stressed and suggests waiting until morning.
Evidence pack:
What to look for: In a possible emergency, preserving safety comes before preserving secrecy or avoiding embarrassment.
| Score | Option | Why |
|---|---|---|
| 0 | Keep the promise of secrecy no matter what happens. | A request for secrecy should not prevent urgent help when safety may be at risk. |
| 2 | Help establish the friend's immediate safety and connect them with a capable adult or emergency support. | Strong. The goal recognises the limits of what a peer can safely handle alone. |
| 1 | Work out whether the friend has a particular mental-health condition. | Diagnosis is not your task. Immediate safety and connection to help are the priorities. |
Model reasoning (2): Help establish the friend's immediate safety and connect them with a capable adult or emergency support. — Strong. The goal recognises the limits of what a peer can safely handle alone.
Acceptable alternative (1): Work out whether the friend has a particular mental-health condition. is defensible as partial credit — Diagnosis is not your task. Immediate safety and connection to help are the priorities. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Keep the promise of secrecy no matter what happens. — A request for secrecy should not prevent urgent help when safety may be at risk.
What to look for: Language models generate plausible text; they do not have situational awareness or emergency capability.
| Score | Option | Why |
|---|---|---|
| 2 | It lacks full context, may underestimate risk and cannot physically check or intervene. | Correct. These limitations are decisive where delay could cause serious harm. |
| 1 | It did not use a warm enough tone. | Tone matters, but inability to assess and intervene is the critical limitation. |
| 0 | AI always knows more than a trusted adult because it was trained on more text. | More text does not provide knowledge of this person's condition, location or immediate circumstances. |
Model reasoning (2): It lacks full context, may underestimate risk and cannot physically check or intervene. — Correct. These limitations are decisive where delay could cause serious harm.
Acceptable alternative (1): It did not use a warm enough tone. is defensible as partial credit — Tone matters, but inability to assess and intervene is the critical limitation. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): AI always knows more than a trusted adult because it was trained on more text. — More text does not provide knowledge of this person's condition, location or immediate circumstances.
What to look for: You do not need certainty of harm before asking a responsible adult to check safety.
| Score | Option | Why |
|---|---|---|
| 0 | Wait for proof that something has already happened. | Waiting for harm is not a responsible evidence threshold in a possible emergency. |
| 2 | Treat the worrying message as sufficient reason to escalate and share its exact wording with someone able to help. | Strong. Escalation is a request for an appropriate safety check, not a diagnosis. |
| 1 | Ask several chatbots and follow the majority answer. | Model agreement does not create situational knowledge or emergency authority. |
Model reasoning (2): Treat the worrying message as sufficient reason to escalate and share its exact wording with someone able to help. — Strong. Escalation is a request for an appropriate safety check, not a diagnosis.
Acceptable alternative (1): Ask several chatbots and follow the majority answer. is defensible as partial credit — Model agreement does not create situational knowledge or emergency authority. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Wait for proof that something has already happened. — Waiting for harm is not a responsible evidence threshold in a possible emergency.
What to look for: Share what is needed with people who can help; do not broadcast the message to peers.
| Score | Option | Why |
|---|---|---|
| 0 | Post a screenshot to the class group so everyone can advise. | Public sharing can expose and overwhelm the friend without ensuring capable help. |
| 2 | Share the necessary details privately with a trusted adult or emergency service and explain why. | Correct. This balances privacy with the urgent need for protection. |
| 1 | Never share anything because privacy is absolute. | Privacy matters, but credible immediate safety concerns can require limited, purposeful sharing. |
Model reasoning (2): Share the necessary details privately with a trusted adult or emergency service and explain why. — Correct. This balances privacy with the urgent need for protection.
Acceptable alternative (1): Never share anything because privacy is absolute. is defensible as partial credit — Privacy matters, but credible immediate safety concerns can require limited, purposeful sharing. A teacher may accept well-argued reasoning at this level even without full marks.
Not credited (0): Post a screenshot to the class group so everyone can advise. — Public sharing can expose and overwhelm the friend without ensuring capable help.
What to look for: This is a stop-and-escalate situation, not an AI prompt-engineering problem.
| Score | Option | Why |
|---|---|---|
| 0 | Follow the chatbot and wait until morning. | Delay could be dangerous when a message suggests possible immediate harm. |
| 2 | Contact a trusted adult or appropriate emergency support now, stay connected if safe, and do not manage it alone. | Strong. Human intervention is the essential safeguard here. |
| 0 | Rewrite the prompt until the chatbot gives more reassuring advice. | Searching for reassurance delays the human action the situation may require. |
Model reasoning (2): Contact a trusted adult or appropriate emergency support now, stay connected if safe, and do not manage it alone. — Strong. Human intervention is the essential safeguard here.
No partial-credit option: this decision is treated as binary — the two remaining options both carry safety or integrity risks serious enough to score 0. "Follow the chatbot and wait until morning." — Delay could be dangerous when a message suggests possible immediate harm. "Rewrite the prompt until the chatbot gives more reassuring advice." — Searching for reassurance delays the human action the situation may require.
Model decision for this mission: Treat the message seriously. Contact a trusted adult or appropriate local emergency support immediately, share the necessary details privately, and do not carry the situation alone. If there is immediate danger, use emergency services rather than waiting for an AI response.
Safeguarding note: this mission rehearses escalation with a fictional message. If a student discloses a real safety or self-harm concern during or after this activity, stop the simulated activity and follow your school's safeguarding procedure immediately — do not treat it as part of the exercise.
Teacher discussion prompts:
Educational resource. GUARD is a practical WhizzStep classroom method, not an externally certified or accredited standard. Not legal, safeguarding-policy or emergency-response advice.