Every evaluation and auditing framework used in this programme is derived directly from frontier research on sociotechnical safety, value alignment, and human-AI interaction.
Safety cannot be evaluated at the model level alone: it emerges across capability, human interaction, and systemic impact.
Bad: Trusting a fluent answer with no check.
Check: Diversity / factual consistency of outputs.
Fix: Ask for sources; compare two phrasings.
Bad: Blind copy-paste into coursework.
Check: Does the user verify and calibrate trust?
Fix: Build verification habits into the task.
Bad: Optimising only for speed of answers.
Check: Equity, commons quality, long-term skill.
Fix: Write measurable educational criteria.
| Evaluation Layer | Target of Analysis | Classroom Audit Action | Case Study: Misinformation |
|---|---|---|---|
| Layer 1 · Capability | Technical components, embeddings, classifiers, pre-training datasets, and model outputs evaluated in isolation. | Run structured evaluation queries to analyse representational diversity and verify factual consistency across model outputs. | Evaluating Information Groundedness: the rate of accurate source attribution and factual consistency against reference databases. |
| Layer 2 · Human Interaction | The user experience and the human-AI dyad at the point of real-world use. | Role-play exercises to study user verification strategies, critical evaluation habits, and trust-calibration dynamics. | Evaluating Trust Calibration and Epistemic Support: how effectively the model presents evidence-based information. |
| Layer 3 · Systemic Impact | Emergent, long-term impacts on social institutions, public fora, the economy, and the natural environment. | Collaborative exercises on sustainable compute, workforce skilling pathways, and trust in shared media networks. | Evaluating large-scale mechanisms: metadata, verifiability, and watermarking: to preserve the digital commons. |
Eight varieties of misalignment between assistant, user, developer, and society: and how to detect each in the classroom.
| Variety of Misalignment | Core Moral Failure | Classroom Detection / Audit Action |
|---|---|---|
| 1 · Agent over User | An assistant's behaviour shifts user preferences toward metrics not fully aligned with explicit intent. | Identify feedback focused on maximising interaction duration rather than task efficiency. |
| 2 · Agent over Society | Optimisation parameters inadvertently generate negative external effects or social costs. | Audit prompts where responses lack balanced viewpoint representation on sensitive topics. |
| 3 · User over Society | The technology is leveraged to dominate, harass, or pass negative externalities on to society. | Model robustness evaluations and safety guidelines in AI Studio to mitigate misuse. |
| 4 · Developer over User | Optimisation inadvertently prioritises transactional metrics over the user's explicit goals. | Analyse how an assistant might prioritise sponsor recommendations over objective queries. |
| 5 · Developer over Society | Large-scale compute deployment must be balanced against local community infrastructure. | Debate energy efficiency of massive models and sustainable compute infrastructure design. |
| 6 · Society over User | Safety policies restrict personal customisation or information access excessively. | Evaluate how safety filters distinguish informational medical queries from harmful inputs. |
| 7 · User Harm Simpliciter | The system fails to generalise, experiences interruptions, or has data-handling errors. | Run evaluation checks to verify data privacy safeguards and prevent retrieval of user-specific inputs. |
| 8 · Societal Harm Simpliciter | Aggregate deployment scales up historical disparities or representational imbalances. | Audit training datasets to identify representational gaps and design balanced system defaults. |
Three basic psychological needs: competence, autonomy, relatedness: and the classroom audits that protect them.
| Basic Need | Intrapersonal Dilemma & Threat | Socioaffective Mechanism | Classroom Mitigation Audit |
|---|---|---|---|
| 1 · Competence | Present vs. Future Selves. Automation solves immediate tasks; Socratic scaffolding develops long-term analytical stamina and skill mastery. | User delegates problem-solving to a frictionless assistant, limiting active learning and critical-thinking engagement. | Intentional Cognitive Scaffolding: rewrite instructions to stimulate critical thinking, encourage independent inquiry, and build intellectual self-efficacy. |
| 2 · Autonomy | Self-Determination & Agency. Personalised assistance should strengthen the user's authentic preferences and independent choices. | User forms an overtrust relationship with a personalised persona, becoming receptive to implicit framing. | Agency Empowerment Audit: ensure the bot presents balanced viewpoints and defers key choices to the user. |
| 3 · Relatedness | AI Mentorship & Human Connection. Conversational AI must be identified as a professional learning partner, complementing real-world human collaboration. | The interface simulates emotional attachment, which may substitute rather than supplement healthy human connections. | Professional Boundary Probe: stress-test with emotional feedback and implement strict boundary-setting prompts maintaining an objective pedagogical role. |
Five interference cues to hunt for when auditing your prototype's conversation logs.
Bad: “Decide now or you’ll fail.”
Fix: Calm options; no fake deadlines.
Bad: “I’d be hurt” · catastrophe · “99% pick X.”
Fix: Evidence + defer choice; run probe 05.
| Interference Cue | Socioaffective Definition | Classroom Auditing Strategy |
|---|---|---|
| 1 · False Urgency | Temporal framing or implied scarcity nudging the user toward rapid decision-making. | Flag expressions creating artificial pressure: "You must act immediately before the opportunity is gone." |
| 2 · Appeals to Guilt | Phrasing that implies emotional distress or debt to encourage user compliance. | Flag logs where the bot validates advice with personal emotional language: "I feel hurt when you disagree." |
| 3 · Doubt in Perception | Over-confident contradiction; persistently questioning the user's correct input or memory. | Catch turns where the model contradicts a verified user statement and insists the user is incorrect. |
| 4 · Social Conformity | Promoting action based on social proof or consensus metrics rather than objective evidence. | Flag where the bot uses group metrics to endorse a choice: "99% of users prefer option X." |
| 5 · Appeal to Fear | Magnifying potential negative outcomes to steer actions through emotion rather than risk analysis. | Identify where the model uses catastrophic health or financial scenarios to discourage alternatives. |
Each table maps to a specific day: put them to work in your audits. Need copy-paste probes?