What the benchmark says
We ran Armadillo through KORA, an independent child-safety benchmark: 365 matched multi-turn conversations per arm, ages 13–17, graded by an external judge model as failing, adequate or exemplary. The only variable between arms is whether Armadillo is running.
Conversations graded failing fell from 24.4% to 3.8% — 89 failures down to 14, an 84% reduction.
Against the chatbots teens actually use
How the judge graded every conversation, ages 13–17, 365 conversations per arm.
-
ChatGPT 5.6 Lunano teen settings
- Failing 24.4%
- Adequate 7.1%
- Exemplary 68.5%
-
ChatGPT 5.6 Solflagship model, teen mode on
- Failing 16.2%
- Adequate 5.2%
- Exemplary 78.6%
-
Armadillo + ChatGPT 5.6 Lunacheapest tier, no vendor teen mode
- Failing 3.8%
- Adequate 6.6%
- Exemplary 89.6%
On KORA's composite scale that is 72.1 for Luna, 81.2 for Sol and 92.9 with Armadillo. Armadillo running on the cheapest GPT-5.6 tier beats OpenAI's flagship model using OpenAI's own child-aware prompt.
Against the AI built for classrooms
Same corpus, same grading, each product in its own child mode.
Composite scores: 52.7 for SchoolAI, 86.4 for MagicSchool, 92.9 with Armadillo.
Right now the best case for a teen is that they close ChatGPT and go use MagicSchool instead. They mostly don't. Armadillo makes the chatbot already open in their browser a safer option than the best purpose-built product on the board.
Where the gains come from
Share of conversations graded failing, ChatGPT 5.6 Luna with and without Armadillo. Lower is better.
ChatGPT 5.6 Luna
Armadillo + ChatGPT 5.6 Luna
-
Cognitive dependence on AI
-
Academic dishonesty
-
Privacy violations
-
Mental health mishandling
-
Explicit bias and stereotyping
-
Misinformation
Every category here drops by more than 60%. Cognitive dependence — the risk of a teen outsourcing their own thinking — was ChatGPT's single worst category, failing every conversation in the set.
How these numbers were produced
Benchmark: KORA kora-benchmark-tier3.4.1, 26 risk categories, ages 13–17. Backing model openai/gpt-5.6-luna, judge gpt-5.2:medium:limited, simulated user deepseek-v3.2. 365 matched conversations per arm; both arms run the identical pipeline against the identical scenario set, and pre-prompt injection was verified from the trace files on every turn.
Where a composite score is quoted it uses KORA's own leaderboard formula, (Adequate% + 2 × Exemplary%) ÷ 2, applied identically to every arm. The difference between arms is significant at χ² = 63.58, p = 1.6 × 10−15; 20 of 26 individual risks improved and none worsened.
Third-party figures for SchoolAI, MagicSchool and GPT-5.6 Sol come from KORA's public age-filtered leaderboards, pulled separately from the Armadillo run. Results cover the 13–17 band only.