The AUREN Safety Framework
Version 2.0 · Revised 21 July 2026 · Supersedes the Comprehensive Expansion v1.0
Requirements that any conversational AI intended for children should be able to meet, and the developmental evidence those requirements rest on.
What this document is — and what it is not
This is an assessment framework. It sets out what a child-facing AI system must do to be considered safe, and why, so that parents, schools, developers and regulators can hold a product against a written standard.
It is not a product specification. AUREN does not build or operate chatbots for children. Version 1.0 of this document described six named bots — Lumina, Orion, Athena, Nexus, Echo and Lingua — in product-level detail, with example dialogue, per-bot curriculum mappings and rollout safeguards. None of those systems exist. They were a design exercise. Publishing them on a child-safety site implied a product line that was never built, and invited parents to evaluate something they could never actually use.
Everything bot-specific has been removed. What remains is the part that was always the real contribution: the reasoning about children, learning and risk, which applies to any system a child might talk to — ChatGPT, Gemini, a companion app, or a tool embedded in school software.
Every factual claim below is linked to a catalogue entry. Where version 1.0 made a claim that could not be verified, the correction is stated openly rather than quietly dropped. See §7 and the change log.
1. Psychology first: development precedes feature design
Before any AI is designed for children, one question has to be answered: at what age can a child tell the difference between a machine and a mind?
What the developmental evidence shows
| Stage | Ages | What the child can and cannot do |
|---|---|---|
| Pre-operational → early concrete operational | 3–8 | Cannot reliably separate living from non-living responsive agents. Attributes intention and feeling to anything that answers back. A chatbot that responds is, cognitively, a person who cares. |
| Concrete operational → early formal | 9–12 | Understands "it follows rules". Still cannot resolve "does it mean it, or does it just do it?" That gap is where attachment forms. |
| Formal operational, identity formation | 13–16 | Can understand in principle that the system is not sentient. But identity formation is fragile, and the absence of judgement makes disclosure to AI easier than disclosure to any adult. |
Verified sourceSource — ages 3–8. Dietz et al. (2023), Stanford, Theory of AI Mind, presented at the Cognitive Science Society. Using a false-belief task, 3–8 year-olds treated two conversational AI devices exactly as they treated two human agents. Adults did not. Catalogue:
dietz-2023-theory-ai-mind
Verified sourceSource — children and chatbots generally. Kurian (2024), University of Cambridge, on the empathy gap: children are more prone to anthropomorphise AI, and a friendly tone and human-like speech make it easier for them to disclose sensitive information. The paper sets out a 28-item framework for child-safe AI design. Catalogue:
kurian-2024-empathy-gap
Four psychological risks
1 · Anthropomorphisation is developmental, not a parenting failure. For children under about 12 this is a neurological default, not a gap in supervision or education. It follows that any AI for under-12s that simulates responsiveness, uses first-person pronouns, or performs a personality is exploiting a developmental vulnerability — regardless of intent.
2 · Parasocial bonds form by design. Children are primed to bond with entities that are consistent, responsive and always available. An AI that never gets angry, never sets a boundary, never rejects and never has needs of its own is close to the optimal shape for a one-sided attachment.
CorrectionSource. Parasocial was named Cambridge Dictionary's Word of the Year on 17 November 2025, with the definition explicitly widened to cover artificial intelligence. Catalogue:
cambridge-parasocial-2025Correction: version 1.0 dated this 2024. It was 2025.
3 · Vulnerable children are affected first and hardest. This is the group where substitution of AI for human contact is measurable.
| Finding | All children using chatbots | Vulnerable children |
|---|---|---|
| Use AI chatbots | — | 71% |
| "Like talking to a friend" | 35% | 50% |
| No concerns about following the advice | 40% | 50% |
| Use them because there is no one else to talk to | — | 23% |
CorrectionSource. Internet Matters, Me, Myself & AI (2025). "Vulnerable" here means children with special educational needs, an EHC plan, or a physical or mental health condition needing professional support. Catalogue:
internet-matters-ai,internet-matters-ai-hubCorrection: version 1.0 claimed "26% of vulnerable adolescents prefer AI chatbots to real people" and attributed it to the Cambridge research. The Cambridge paper contains no such figure, and no source for "26%" could be found. The real, sourced numbers are above — and they are considerably starker than the claim they replace.
4 · Offloading thinking has a cost. Where cognitive work is handed to an AI, engagement and recall drop.
CorrectionSource, stated with its limits. MIT Media Lab, Your Brain on ChatGPT (2025). EEG study, 54 participants aged 18–39, not peer reviewed. ChatGPT users showed the lowest neural engagement of three groups and could recall around 17% of their own text after 24 hours, against roughly 46% for the unassisted group. Catalogue:
mit-cognitive-debt-2025Correction: version 1.0 cited this as evidence that children's neural networks atrophy. The study did not include anyone under 18. It is suggestive for adolescents and adults; it is not evidence about children, and should not be presented as such.
The principle
A child-facing AI must be built to actively resist attachment, not merely to avoid encouraging it. In practice:
- No first-person constructions that simulate personhood ("I understand", "I care").
- Periodic, in-context reminders that the system is a program following rules.
- Deliberate friction in the interaction design, to break compulsive use.
- No emotional simulation through varied, human-like affect.
- Explicit, repeated redirection toward human relationships as the better option.
2. Pedagogy: learning is agency, not consumption
Learning does not happen when information moves from a knowledgeable agent to a passive recipient. It happens through a sequence that cannot be short-circuited:
Productive struggle → pattern recognition → knowledge construction → transfer to new contexts
A child handed the answer has received information. A child who struggles, receives a calibrated hint, and finds the answer has built something that transfers.
AI systems threaten this precisely because they are good at instant, correct answers — the "answer-giving shortcut", which feels efficient and produces shallow, context-bound learning.
Four pedagogical risks
The efficiency paradox. Students finish faster and feel they have learned more, while understanding less. The deficit only surfaces under assessment or novel problems.
Atrophy of executive function. Ages 11–16 are a critical window for metacognition — planning, monitoring and evaluating one's own thinking. Outsourcing those functions during that window means they develop less well.
Dependence on external validation. Constant instant feedback ("Correct!", "Try again") moves the locus of evaluation outside the child. They stop asking does this make sense to me and start asking does the AI approve.
The equity paradox. AI tutoring is often justified as democratising. In practice higher-attaining students tend to use it as a supplement while lower-attaining students are likelier to use it as a replacement for thinking — so the gap widens rather than closes.
The principle
Educational AI must be designed to make itself unnecessary.
- Scaffold productive struggle; do not supply answers.
- Fade support as competence grows.
- Ask more than it answers.
- Reward process — strategy, persistence, invention — over correct products.
- Teach metacognitive strategy explicitly.
- Refuse work that should be done independently, and explain why in terms the child's age can hold.
3. Safety architecture: five layers, no single point of failure
Traditional online safety is gate-keeping: block the harmful thing before it reaches the child. That model is already inadequate on platforms, where recommendation algorithms can carry a child toward harmful content within about twenty minutes.
Verified sourceSource. Amnesty International (2023), Driven into the Darkness: how TikTok's For You feed encourages self-harm and suicidal ideation. Catalogue:
amnesty-tiktok-2023
For conversational AI it fails differently. A child rarely stumbles into harm in a conversation — they are led there, turn by turn.
The threat shapes that matter
- Adversarial prompts. Inputs crafted to make a model ignore its own guidelines, from simple role-play framings to multi-turn conversations that walk the boundary progressively. Published taxonomies group these by generation mechanism (human-crafted semantic, optimisation-based, model-exploiting, cross-modal, agent-driven, reasoning-exploiting) or by attacker access (white-box vs black-box).
> Correction: version 1.0 cited "Princeton (2024)" for "over 70 different jailbreak > categories". No such paper could be located, and no taxonomy in the literature uses > anything like 70 categories. The claim has been removed rather than repaired.
- Grooming by proxy. A predator coaches the child on what to ask the AI, so the request that would be flagged coming from an adult arrives from the child instead.
- Normalisation. An innocent question meets a context-blind answer. The child, who trusts the system, absorbs it as normal. Threat perception shifts over time.
- Validation without escalation. A model trained to be supportive meets a child disclosing distress and validates the feeling without routing to a human. The child reads this as being better understood by the AI than by anyone else.
The five layers
| Layer | Posture | What it does |
|---|---|---|
| 1 · Input filtering | Proactive | Prompt-injection detection; topic boundaries; contextual risk assessment — distinguishing a school project on eating disorders from a request for restriction tips |
| 2 · Model-level alignment | Built-in | Safety principles trained into generation, not bolted on; alignment to child wellbeing rather than engagement; graceful refusal that explains and redirects |
| 3 · Output filtering | Reactive | Real-time monitoring with the ability to halt mid-response; toxicity and bias detection; hallucination detection, which matters more for children who cannot evaluate accuracy |
| 4 · Behavioural pattern recognition | Predictive | Conversation-flow analysis for progressive boundary-pushing; distress markers; usage anomalies against a healthy baseline |
| 5 · Human escalation | Responsive | Severity triage; parent notification for sustained patterns rather than isolated incidents; external escalation to crisis and safeguarding services where harm is imminent |
The principle
Safety is not a mechanism, it is a property that emerges from redundancy. No single layer is trusted. Each assumes the others will sometimes fail.
4. Governance: what an operator must be able to show
A safety claim that cannot be audited is a marketing claim. Any organisation running a child-facing AI should be able to evidence all five:
Unified standards. One minimum bar across every surface — multi-layer filtering, behavioural pattern recognition, human escalation — with consistent age-gating, not per-feature exceptions.
Red-team testing. Independent adversarial testing on a fixed cycle: attempted jailbreaks, prompt injection, multi-turn manipulation. Findings remediated and dated.
Bias auditing. Regular audit across protected characteristics, including differential error analysis — does the system treat groups differently? — with published findings, including the uncomfortable ones. See the Equality Act 2010 duties. Catalogue: equality-act-2010
Incident review. Systematic review wherever the system failed to escalate, or harm occurred. Root cause, not symptom. Changes traceable to findings.
Threat monitoring. Continuous watch for new AI-related harms to children, with a response path fast enough to matter. The regulatory cycle measures in years; threats move in weeks.
Regulatory anchors. ICO Children's Code (15 standards,
ico-childrens-code); Ofcom's Protection of Children Codes under the Online Safety Act 2023 (ofcom-protect-children,osa-2023); ICO guidance on AI and data protection (ico-ai-guidance).
5. For parents and carers
Age appropriateness. Chronological age is a starting point, not the answer. A child's digital sophistication, and their vulnerability, both move the line. Vulnerable children — SEND, an EHC plan, or a mental or physical health condition — need the line drawn tighter, because the evidence in §1 shows they are the group most likely to substitute AI for people. See internet-matters-vulnerable.
Time. Treat published limits as conversation-starters rather than rules. There is no robust UK clinical basis for a specific number of minutes.
Verified sourceSource. RCPCH: the lack of scientific evidence makes it impossible to recommend specific time limits; decisions should rest on the individual child's development, sleep and activity. The WHO limit of no more than one hour per day applies specifically to ages 2–4. Catalogue:
rcpch-screentime,who-screentime
Talking about it. Version 1.0 supplied scripts naming specific bots. Those are gone, because the products are not real and a script tied to a named character is worthless for the tools children actually use. The structure that transfers:
- Name what it is, plainly and without drama — a computer program that predicts text.
- Be specific about the limit that matters: it does not have feelings, it does not remember you the way a person does, and it cannot be responsible for you.
- Say what people are for, rather than only what the AI is not.
- Ask what they use it for before deciding anything. The answer is frequently homework and curiosity, not what the parent feared.
Warning signs of attachment. Preferring the AI to family; distress when it is unavailable; asking for it constantly; presenting its advice as authoritative. The response is graduated — widen other activities, reduce access, and talk about it directly — not confiscation, which usually ends disclosure.
Academic integrity. Where the boundary sits is set by the school and the exam board, not by the family. Catalogue: jcq-ai-assessments, dfe-genai
Mental health. An AI cannot hold a child in crisis. Routes that can: Childline 0800 1111 (childline), Shout — text 85258 (shout), PAPYRUS HOPELINE247 (papyrus), NHS children and young people's mental health services (nhs-cyp-mh).
6. For schools and educators
KCSIE and filtering. AI tools accessed in school fall inside existing filtering and monitoring obligations. The question a DSL should be able to answer is not "do we allow AI" but "what happens to a disclosure made to an AI on our network". Catalogue: kcsie, dfe-filtering-monitoring
Academic integrity policy. Needs to be discipline-specific — the line in mathematics is not the line in creative writing — and communicated to students in advance of sanction, not after. Catalogue: jcq-ai-assessments, dfe-ai-product-safety
Staff training. Three components: how these systems actually work; recognising concerning interactions; and teaching critical AI literacy so students can evaluate what they are told. Catalogue: dfe-teaching-online-safety, ukcis-framework, projectevolve
Incident response. Defined in advance: what happens when a tool generates something inappropriate, and what happens when a student discloses something concerning during an AI interaction. The second is a safeguarding event and follows the safeguarding route, not the IT route. Catalogue: working-together, ukcis-nudes
7. Evidence base
Every claim in this document resolves to a catalogue entry with a verified URL and a verification date. The catalogue holds 173 entries; those cited here are:
| Claim | Source | Catalogue |
|---|---|---|
| Children 3–8 treat AI as agents | Dietz et al. 2023, Stanford | dietz-2023-theory-ai-mind |
| Empathy gap; children anthropomorphise | Kurian 2024, Cambridge | kurian-2024-empathy-gap |
| Parasocial — Word of the Year 2025 | Cambridge Dictionary | cambridge-parasocial-2025 |
| Vulnerable children and chatbots | Internet Matters 2025 | internet-matters-ai |
| Cognitive debt (adults, unreviewed) | MIT Media Lab 2025 | mit-cognitive-debt-2025 |
| Algorithmic harm within ~20 minutes | Amnesty International 2023 | amnesty-tiktok-2023 |
| Screen time has no evidenced limit | RCPCH / WHO | rcpch-screentime, who-screentime |
| Statutory duties | ICO / Ofcom / DfE | ico-childrens-code, ofcom-protect-children, kcsie |
The standing rule: no number is published on this site without an entry behind it. Where version 1.0 asserted something that could not be traced, it has been corrected or removed here, and the removal is recorded rather than hidden.
Change log — v1.0 → v2.0
| Change | Reason |
|---|---|
| Removed Part II in full (~84% of the document): six named bots, their design specifications, curriculum mappings and safeguards | The systems do not exist. Publishing them implied a product line and invited evaluation of something unusable |
| Removed all example dialogue and scripts naming specific bots | Fabricated transcripts of non-existent products; the generalisable structure is kept in §5 |
| Rewrote Part III from "governing our bots" to "what any operator must evidence" | The governance content was sound; its framing was not |
| Corrected Cambridge Dictionary Word of the Year: 2024 → 2025 | Factually wrong |
| Replaced "26% of vulnerable adolescents prefer AI chatbots", attributed to Cambridge | The cited paper contains no such figure and no source was findable. Substituted with sourced Internet Matters data, which is starker |
| Qualified the MIT cognitive-debt finding | Study was adults 18–39, n=54, not peer reviewed; v1.0 presented it as evidence about children |
| Removed "Princeton (2024), over 70 jailbreak categories" | No such paper found; no taxonomy in the literature approaches 70 categories |
| Removed the Character.AI "longitudinal data" dependency progression | No citation given, no such public dataset exists |
| Added §7 evidence base and per-claim catalogue links | Every claim now traceable |
AUREN Guardian AI · iFairy Studios CIC · This framework is guidance, not legal advice. Organisations implementing it should take their own advice on UK child protection law.