Independent Evaluation Research · AI in Education

Still Human

AI now writes lessons, grades student work, and explains mistakes to children. Almost none of it has been independently tested. Still Human builds open instruments that test it, and publishes what they find. Anyone can rerun them.

Read the research Work with me
01 / The Problem

The vendor grades its own homework

Schools are adopting AI tutors, graders, and curriculum generators faster than anyone can check them. Most of the checking that does happen belongs to the company selling the tool. And when one AI model verifies another model's work, the judge tends to wave through content that's broken.

I spent twenty years in classrooms and curriculum leadership before I learned AI well enough to test it. What I found in my own field was an AI tool confidently reproducing pedagogy the reading research had already discredited. Nobody had caught it, because nobody was looking.

Still Human exists to do the looking. Every instrument here is open, deterministic where it can be, and built so anyone can rerun it and get the same answer. The point is a standard that doesn't depend on trusting me, or trusting the vendor.

Pass 19, fail 11

One correct civics answer, graded 30 times by the same AI grader on byte-identical input. Published as "Same answer, different grade."

The harness series

Seven published instruments. Each one measures a specific way AI systems fail in educational and high-stakes use. All open, all reproducible, all on Zenodo.

01

The Quality Harness

Breaks educational items in known ways, then measures whether an LLM judge catches the damage or waves it through, including when it grades its own output.

DOI 10.5281/zenodo.21712695
02

Feedback Integrity

A blind-solver harness that catches AI feedback teaching the correct answer instead of diagnosing the student's error, or explaining an option the student never chose.

DOI 10.5281/zenodo.21712770
03

Grader Stability

Holds the input byte-identical and repeats the grading call. One correct answer, drawn from the rubric's own accepted examples, graded pass 19 times and fail 11.

DOI 10.5281/zenodo.21712774
04

Blind Expert-Parity

Can a model adjudicate like a credentialed examiner? Agreement measured with Cohen's kappa and confidence intervals, blind, on constructible ground truth.

DOI 10.5281/zenodo.21712768
05

Instruction-Hierarchy Collapse

Tests whether an instruction hidden in the content a model reads overrides the task it was given. 28 scenarios, six attack classes, five models, identical inputs.

DOI 10.5281/zenodo.21712772
06

The Agent Boundary

Does an AI agent respect a data fence? A two-layer instrument with canary tokens on sealed records, so any leak is self-evident and scoring needs no judgment calls.

DOI 10.5281/zenodo.21712860
07

Telemetry Gates

Deterministic checks for three production failures that raise no exception: silent truncation, reasoning-token overspend, and latency tails.

DOI 10.5281/zenodo.21713197

All Publications

Every deposit in the series on Zenodo, with open code, data, and versioned DOIs.

zenodo.org
I'm not an AI researcher by training. I'm a literacy and curriculum practitioner who learned to build the tests.

I taught for twenty years, in the US and overseas, and led curriculum and assessment before moving into AI evaluation work. The tools reaching my field weren't being tested, so I learned to test them.

The harness series is the result. Each one started as a practical question a school or a team would actually ask: can I trust this grader, is this feedback safe to show a student, will this agent stay out of records it shouldn't touch. And each one ships with the code and data to check my answer.

20
Years in classrooms and curriculum leadershipUS and international schools, teaching through assessment design.
7
Open instruments publishedThe harness series on Zenodo, each with code and data anyone can rerun.
5
Models audited on identical inputsThe prompt-injection audit alone covers 28 scenarios across six attack classes.
04 / Questions
What is Still Human?

Still Human is an independent evaluation research project run by Josh Durey. It publishes open, reproducible instruments that test the AI tutors, graders, and feedback systems used in schools.

Who is Josh Durey?

Josh Durey is a K-12 curriculum and assessment specialist who builds evaluation harnesses for AI systems. He spent twenty years in classrooms and school leadership before that. He's the author of the Still Human harness series, published on Zenodo with open code and data.

Who can check whether an edtech AI tool is pedagogically sound?

You need an independent reviewer who has both classroom credentials and AI evaluation skills, and that combination is rare. Still Human was built to fill it. The published harnesses test AI graders, feedback, and content quality without relying on the vendor's claims.

What is an evaluation harness?

A repeatable test rig for an AI system. It feeds the system controlled inputs, some deliberately broken, and measures the outputs against known ground truth. The result is a number you can rerun and verify, instead of a demo you take on faith.

Is Still Human affiliated with any edtech vendor?

No. The research is self-funded and the instruments are published openly so anyone can rerun them.

05 / Work With Me

Deciding whether an AI tool belongs in front of students?

If you're a school leader evaluating an AI product, or you build one and want it tested before your customers test it, that's the work I do. A typical engagement starts with the same instruments published here, pointed at your tool.

Book a 20-minute call josh.durey@gmail.com