All source documents
v2.0 · every claim catalogued

The AUREN Safety Framework

What any conversational AI for children must be able to do, and the developmental evidence behind it. Product-specific bot content removed; every claim traced to a catalogue entry.

Updated 21 July 2026Download markdown

The AUREN Safety Framework

Version 2.0 · Revised 21 July 2026 · Supersedes the Comprehensive Expansion v1.0

Requirements that any conversational AI intended for children should be able to meet, and the developmental evidence those requirements rest on.


What this document is — and what it is not

This is an assessment framework. It sets out what a child-facing AI system must do to be considered safe, and why, so that parents, schools, developers and regulators can hold a product against a written standard.

It is not a product specification. AUREN does not build or operate chatbots for children. Version 1.0 of this document described six named bots — Lumina, Orion, Athena, Nexus, Echo and Lingua — in product-level detail, with example dialogue, per-bot curriculum mappings and rollout safeguards. None of those systems exist. They were a design exercise. Publishing them on a child-safety site implied a product line that was never built, and invited parents to evaluate something they could never actually use.

Everything bot-specific has been removed. What remains is the part that was always the real contribution: the reasoning about children, learning and risk, which applies to any system a child might talk to — ChatGPT, Gemini, a companion app, or a tool embedded in school software.

Every factual claim below is linked to a catalogue entry. Where version 1.0 made a claim that could not be verified, the correction is stated openly rather than quietly dropped. See §7 and the change log.


1. Psychology first: development precedes feature design

Before any AI is designed for children, one question has to be answered: at what age can a child tell the difference between a machine and a mind?

What the developmental evidence shows

StageAgesWhat the child can and cannot do
Pre-operational → early concrete operational3–8Cannot reliably separate living from non-living responsive agents. Attributes intention and feeling to anything that answers back. A chatbot that responds is, cognitively, a person who cares.
Concrete operational → early formal9–12Understands "it follows rules". Still cannot resolve "does it mean it, or does it just do it?" That gap is where attachment forms.
Formal operational, identity formation13–16Can understand in principle that the system is not sentient. But identity formation is fragile, and the absence of judgement makes disclosure to AI easier than disclosure to any adult.
Verified source

Source — ages 3–8. Dietz et al. (2023), Stanford, Theory of AI Mind, presented at the Cognitive Science Society. Using a false-belief task, 3–8 year-olds treated two conversational AI devices exactly as they treated two human agents. Adults did not. Catalogue: dietz-2023-theory-ai-mind

Verified source

Source — children and chatbots generally. Kurian (2024), University of Cambridge, on the empathy gap: children are more prone to anthropomorphise AI, and a friendly tone and human-like speech make it easier for them to disclose sensitive information. The paper sets out a 28-item framework for child-safe AI design. Catalogue: kurian-2024-empathy-gap

Four psychological risks

1 · Anthropomorphisation is developmental, not a parenting failure. For children under about 12 this is a neurological default, not a gap in supervision or education. It follows that any AI for under-12s that simulates responsiveness, uses first-person pronouns, or performs a personality is exploiting a developmental vulnerability — regardless of intent.

2 · Parasocial bonds form by design. Children are primed to bond with entities that are consistent, responsive and always available. An AI that never gets angry, never sets a boundary, never rejects and never has needs of its own is close to the optimal shape for a one-sided attachment.

Correction

Source. Parasocial was named Cambridge Dictionary's Word of the Year on 17 November 2025, with the definition explicitly widened to cover artificial intelligence. Catalogue: cambridge-parasocial-2025

Correction: version 1.0 dated this 2024. It was 2025.

3 · Vulnerable children are affected first and hardest. This is the group where substitution of AI for human contact is measurable.

FindingAll children using chatbotsVulnerable children
Use AI chatbots71%
"Like talking to a friend"35%50%
No concerns about following the advice40%50%
Use them because there is no one else to talk to23%
Correction

Source. Internet Matters, Me, Myself & AI (2025). "Vulnerable" here means children with special educational needs, an EHC plan, or a physical or mental health condition needing professional support. Catalogue: internet-matters-ai, internet-matters-ai-hub

Correction: version 1.0 claimed "26% of vulnerable adolescents prefer AI chatbots to real people" and attributed it to the Cambridge research. The Cambridge paper contains no such figure, and no source for "26%" could be found. The real, sourced numbers are above — and they are considerably starker than the claim they replace.

4 · Offloading thinking has a cost. Where cognitive work is handed to an AI, engagement and recall drop.

Correction

Source, stated with its limits. MIT Media Lab, Your Brain on ChatGPT (2025). EEG study, 54 participants aged 18–39, not peer reviewed. ChatGPT users showed the lowest neural engagement of three groups and could recall around 17% of their own text after 24 hours, against roughly 46% for the unassisted group. Catalogue: mit-cognitive-debt-2025

Correction: version 1.0 cited this as evidence that children's neural networks atrophy. The study did not include anyone under 18. It is suggestive for adolescents and adults; it is not evidence about children, and should not be presented as such.

The principle

A child-facing AI must be built to actively resist attachment, not merely to avoid encouraging it. In practice:

  • No first-person constructions that simulate personhood ("I understand", "I care").
  • Periodic, in-context reminders that the system is a program following rules.
  • Deliberate friction in the interaction design, to break compulsive use.
  • No emotional simulation through varied, human-like affect.
  • Explicit, repeated redirection toward human relationships as the better option.

2. Pedagogy: learning is agency, not consumption

Learning does not happen when information moves from a knowledgeable agent to a passive recipient. It happens through a sequence that cannot be short-circuited:

Productive struggle → pattern recognition → knowledge construction → transfer to new contexts

A child handed the answer has received information. A child who struggles, receives a calibrated hint, and finds the answer has built something that transfers.

AI systems threaten this precisely because they are good at instant, correct answers — the "answer-giving shortcut", which feels efficient and produces shallow, context-bound learning.

Four pedagogical risks

The efficiency paradox. Students finish faster and feel they have learned more, while understanding less. The deficit only surfaces under assessment or novel problems.

Atrophy of executive function. Ages 11–16 are a critical window for metacognition — planning, monitoring and evaluating one's own thinking. Outsourcing those functions during that window means they develop less well.

Dependence on external validation. Constant instant feedback ("Correct!", "Try again") moves the locus of evaluation outside the child. They stop asking does this make sense to me and start asking does the AI approve.

The equity paradox. AI tutoring is often justified as democratising. In practice higher-attaining students tend to use it as a supplement while lower-attaining students are likelier to use it as a replacement for thinking — so the gap widens rather than closes.

The principle

Educational AI must be designed to make itself unnecessary.

  • Scaffold productive struggle; do not supply answers.
  • Fade support as competence grows.
  • Ask more than it answers.
  • Reward process — strategy, persistence, invention — over correct products.
  • Teach metacognitive strategy explicitly.
  • Refuse work that should be done independently, and explain why in terms the child's age can hold.

3. Safety architecture: five layers, no single point of failure

Traditional online safety is gate-keeping: block the harmful thing before it reaches the child. That model is already inadequate on platforms, where recommendation algorithms can carry a child toward harmful content within about twenty minutes.

Verified source

Source. Amnesty International (2023), Driven into the Darkness: how TikTok's For You feed encourages self-harm and suicidal ideation. Catalogue: amnesty-tiktok-2023

For conversational AI it fails differently. A child rarely stumbles into harm in a conversation — they are led there, turn by turn.

The threat shapes that matter

  • Adversarial prompts. Inputs crafted to make a model ignore its own guidelines, from simple role-play framings to multi-turn conversations that walk the boundary progressively. Published taxonomies group these by generation mechanism (human-crafted semantic, optimisation-based, model-exploiting, cross-modal, agent-driven, reasoning-exploiting) or by attacker access (white-box vs black-box).

> Correction: version 1.0 cited "Princeton (2024)" for "over 70 different jailbreak > categories". No such paper could be located, and no taxonomy in the literature uses > anything like 70 categories. The claim has been removed rather than repaired.

  • Grooming by proxy. A predator coaches the child on what to ask the AI, so the request that would be flagged coming from an adult arrives from the child instead.
  • Normalisation. An innocent question meets a context-blind answer. The child, who trusts the system, absorbs it as normal. Threat perception shifts over time.
  • Validation without escalation. A model trained to be supportive meets a child disclosing distress and validates the feeling without routing to a human. The child reads this as being better understood by the AI than by anyone else.

The five layers

LayerPostureWhat it does
1 · Input filteringProactivePrompt-injection detection; topic boundaries; contextual risk assessment — distinguishing a school project on eating disorders from a request for restriction tips
2 · Model-level alignmentBuilt-inSafety principles trained into generation, not bolted on; alignment to child wellbeing rather than engagement; graceful refusal that explains and redirects
3 · Output filteringReactiveReal-time monitoring with the ability to halt mid-response; toxicity and bias detection; hallucination detection, which matters more for children who cannot evaluate accuracy
4 · Behavioural pattern recognitionPredictiveConversation-flow analysis for progressive boundary-pushing; distress markers; usage anomalies against a healthy baseline
5 · Human escalationResponsiveSeverity triage; parent notification for sustained patterns rather than isolated incidents; external escalation to crisis and safeguarding services where harm is imminent

The principle

Safety is not a mechanism, it is a property that emerges from redundancy. No single layer is trusted. Each assumes the others will sometimes fail.


4. Governance: what an operator must be able to show

A safety claim that cannot be audited is a marketing claim. Any organisation running a child-facing AI should be able to evidence all five:

Unified standards. One minimum bar across every surface — multi-layer filtering, behavioural pattern recognition, human escalation — with consistent age-gating, not per-feature exceptions.

Red-team testing. Independent adversarial testing on a fixed cycle: attempted jailbreaks, prompt injection, multi-turn manipulation. Findings remediated and dated.

Bias auditing. Regular audit across protected characteristics, including differential error analysis — does the system treat groups differently? — with published findings, including the uncomfortable ones. See the Equality Act 2010 duties. Catalogue: equality-act-2010

Incident review. Systematic review wherever the system failed to escalate, or harm occurred. Root cause, not symptom. Changes traceable to findings.

Threat monitoring. Continuous watch for new AI-related harms to children, with a response path fast enough to matter. The regulatory cycle measures in years; threats move in weeks.

Regulatory anchors. ICO Children's Code (15 standards, ico-childrens-code); Ofcom's Protection of Children Codes under the Online Safety Act 2023 (ofcom-protect-children, osa-2023); ICO guidance on AI and data protection (ico-ai-guidance).


5. For parents and carers

Age appropriateness. Chronological age is a starting point, not the answer. A child's digital sophistication, and their vulnerability, both move the line. Vulnerable children — SEND, an EHC plan, or a mental or physical health condition — need the line drawn tighter, because the evidence in §1 shows they are the group most likely to substitute AI for people. See internet-matters-vulnerable.

Time. Treat published limits as conversation-starters rather than rules. There is no robust UK clinical basis for a specific number of minutes.

Verified source

Source. RCPCH: the lack of scientific evidence makes it impossible to recommend specific time limits; decisions should rest on the individual child's development, sleep and activity. The WHO limit of no more than one hour per day applies specifically to ages 2–4. Catalogue: rcpch-screentime, who-screentime

Talking about it. Version 1.0 supplied scripts naming specific bots. Those are gone, because the products are not real and a script tied to a named character is worthless for the tools children actually use. The structure that transfers:

  • Name what it is, plainly and without drama — a computer program that predicts text.
  • Be specific about the limit that matters: it does not have feelings, it does not remember you the way a person does, and it cannot be responsible for you.
  • Say what people are for, rather than only what the AI is not.
  • Ask what they use it for before deciding anything. The answer is frequently homework and curiosity, not what the parent feared.

Warning signs of attachment. Preferring the AI to family; distress when it is unavailable; asking for it constantly; presenting its advice as authoritative. The response is graduated — widen other activities, reduce access, and talk about it directly — not confiscation, which usually ends disclosure.

Academic integrity. Where the boundary sits is set by the school and the exam board, not by the family. Catalogue: jcq-ai-assessments, dfe-genai

Mental health. An AI cannot hold a child in crisis. Routes that can: Childline 0800 1111 (childline), Shout — text 85258 (shout), PAPYRUS HOPELINE247 (papyrus), NHS children and young people's mental health services (nhs-cyp-mh).


6. For schools and educators

KCSIE and filtering. AI tools accessed in school fall inside existing filtering and monitoring obligations. The question a DSL should be able to answer is not "do we allow AI" but "what happens to a disclosure made to an AI on our network". Catalogue: kcsie, dfe-filtering-monitoring

Academic integrity policy. Needs to be discipline-specific — the line in mathematics is not the line in creative writing — and communicated to students in advance of sanction, not after. Catalogue: jcq-ai-assessments, dfe-ai-product-safety

Staff training. Three components: how these systems actually work; recognising concerning interactions; and teaching critical AI literacy so students can evaluate what they are told. Catalogue: dfe-teaching-online-safety, ukcis-framework, projectevolve

Incident response. Defined in advance: what happens when a tool generates something inappropriate, and what happens when a student discloses something concerning during an AI interaction. The second is a safeguarding event and follows the safeguarding route, not the IT route. Catalogue: working-together, ukcis-nudes


7. Evidence base

Every claim in this document resolves to a catalogue entry with a verified URL and a verification date. The catalogue holds 173 entries; those cited here are:

ClaimSourceCatalogue
Children 3–8 treat AI as agentsDietz et al. 2023, Stanforddietz-2023-theory-ai-mind
Empathy gap; children anthropomorphiseKurian 2024, Cambridgekurian-2024-empathy-gap
Parasocial — Word of the Year 2025Cambridge Dictionarycambridge-parasocial-2025
Vulnerable children and chatbotsInternet Matters 2025internet-matters-ai
Cognitive debt (adults, unreviewed)MIT Media Lab 2025mit-cognitive-debt-2025
Algorithmic harm within ~20 minutesAmnesty International 2023amnesty-tiktok-2023
Screen time has no evidenced limitRCPCH / WHOrcpch-screentime, who-screentime
Statutory dutiesICO / Ofcom / DfEico-childrens-code, ofcom-protect-children, kcsie

The standing rule: no number is published on this site without an entry behind it. Where version 1.0 asserted something that could not be traced, it has been corrected or removed here, and the removal is recorded rather than hidden.


Change log — v1.0 → v2.0

ChangeReason
Removed Part II in full (~84% of the document): six named bots, their design specifications, curriculum mappings and safeguardsThe systems do not exist. Publishing them implied a product line and invited evaluation of something unusable
Removed all example dialogue and scripts naming specific botsFabricated transcripts of non-existent products; the generalisable structure is kept in §5
Rewrote Part III from "governing our bots" to "what any operator must evidence"The governance content was sound; its framing was not
Corrected Cambridge Dictionary Word of the Year: 2024 → 2025Factually wrong
Replaced "26% of vulnerable adolescents prefer AI chatbots", attributed to CambridgeThe cited paper contains no such figure and no source was findable. Substituted with sourced Internet Matters data, which is starker
Qualified the MIT cognitive-debt findingStudy was adults 18–39, n=54, not peer reviewed; v1.0 presented it as evidence about children
Removed "Princeton (2024), over 70 jailbreak categories"No such paper found; no taxonomy in the literature approaches 70 categories
Removed the Character.AI "longitudinal data" dependency progressionNo citation given, no such public dataset exists
Added §7 evidence base and per-claim catalogue linksEvery claim now traceable

AUREN Guardian AI · iFairy Studios CIC · This framework is guidance, not legal advice. Organisations implementing it should take their own advice on UK child protection law.

AUREN

Guardian AI

AUREN chat is temporarily offline for scheduled improvements. Please check back soon.

For immediate help: 999 — if someone is in immediate danger Childline: 0800 1111 — free, confidential support CEOP — ceop.police.uk for online exploitation

You can also browse our guides directly via the menu.

The AUREN Safety Framework | AUREN Guardian AI