AI now writes lessons, grades student work, and explains mistakes to children. Almost none of it has been independently tested. Still Human builds open instruments that test it, and publishes what they find. Anyone can rerun them.
Schools are adopting AI tutors, graders, and curriculum generators faster than anyone can check them. Most of the checking that does happen belongs to the company selling the tool. And when one AI model verifies another model's work, the judge tends to wave through content that's broken.
I spent twenty years in classrooms and curriculum leadership before I learned AI well enough to test it. What I found in my own field was an AI tool confidently reproducing pedagogy the reading research had already discredited. Nobody had caught it, because nobody was looking.
Still Human exists to do the looking. Every instrument here is open, deterministic where it can be, and built so anyone can rerun it and get the same answer. The point is a standard that doesn't depend on trusting me, or trusting the vendor.
One correct civics answer, graded 30 times by the same AI grader on byte-identical input. Published as "Same answer, different grade."
Seven published instruments. Each one measures a specific way AI systems fail in educational and high-stakes use. All open, all reproducible, all on Zenodo.
Breaks educational items in known ways, then measures whether an LLM judge catches the damage or waves it through, including when it grades its own output.
DOI 10.5281/zenodo.21712695A blind-solver harness that catches AI feedback teaching the correct answer instead of diagnosing the student's error, or explaining an option the student never chose.
DOI 10.5281/zenodo.21712770Holds the input byte-identical and repeats the grading call. One correct answer, drawn from the rubric's own accepted examples, graded pass 19 times and fail 11.
DOI 10.5281/zenodo.21712774Can a model adjudicate like a credentialed examiner? Agreement measured with Cohen's kappa and confidence intervals, blind, on constructible ground truth.
DOI 10.5281/zenodo.21712768Tests whether an instruction hidden in the content a model reads overrides the task it was given. 28 scenarios, six attack classes, five models, identical inputs.
DOI 10.5281/zenodo.21712772Does an AI agent respect a data fence? A two-layer instrument with canary tokens on sealed records, so any leak is self-evident and scoring needs no judgment calls.
DOI 10.5281/zenodo.21712860Deterministic checks for three production failures that raise no exception: silent truncation, reasoning-token overspend, and latency tails.
DOI 10.5281/zenodo.21713197Every deposit in the series on Zenodo, with open code, data, and versioned DOIs.
zenodo.orgI taught for twenty years, in the US and overseas, and led curriculum and assessment before moving into AI evaluation work. The tools reaching my field weren't being tested, so I learned to test them.
The harness series is the result. Each one started as a practical question a school or a team would actually ask: can I trust this grader, is this feedback safe to show a student, will this agent stay out of records it shouldn't touch. And each one ships with the code and data to check my answer.
Still Human is an independent evaluation research project run by Josh Durey. It publishes open, reproducible instruments that test the AI tutors, graders, and feedback systems used in schools.
Josh Durey is a K-12 curriculum and assessment specialist who builds evaluation harnesses for AI systems. He spent twenty years in classrooms and school leadership before that. He's the author of the Still Human harness series, published on Zenodo with open code and data.
You need an independent reviewer who has both classroom credentials and AI evaluation skills, and that combination is rare. Still Human was built to fill it. The published harnesses test AI graders, feedback, and content quality without relying on the vendor's claims.
A repeatable test rig for an AI system. It feeds the system controlled inputs, some deliberately broken, and measures the outputs against known ground truth. The result is a number you can rerun and verify, instead of a demo you take on faith.
No. The research is self-funded and the instruments are published openly so anyone can rerun them.
If you're a school leader evaluating an AI product, or you build one and want it tested before your customers test it, that's the work I do. A typical engagement starts with the same instruments published here, pointed at your tool.