A score is a joint effort. The only part that belongs to the student is the part that survives once the help is taken away.
Every program wants the same piece of evidence: that the people coming out of it can actually do the work. So it measures performance. Students run the exercises, clear the labs, finish the assessments, and the numbers go up.
But a score is never produced by the student alone. It is produced by the student plus everything standing behind them: the structure of the lesson, the framing of the task, the tool that was just demonstrated, the hint at the bottom of the page, and now an AI assistant sitting one keystroke away.
Performance is the sum of all of it, but readiness is the part that remains when the help is taken away.
That distinction is not a quibble, and it is not new. In a 2015 review of decades of learning research, cognitive scientists Nicholas Soderstrom and Robert Bjork drew a hard line between performance, what a learner can do during training, and learning, what durably remains afterward. Their conclusion was uncomfortable. The two can move in opposite directions. Gains in performance can fail to produce gains in learning, and some of the conditions that boost one can actively suppress the other.
The pattern is simple. Heavy guidance makes people look better during training, but weaker once the guidance is gone. Easy, repetitive practice feels like progress, but often leaves less behind. The work that builds real skill usually feels harder, messier, and less impressive in the moment.
AI assistance has quietly made this harder to see, because an assistant is the most powerful scaffold a student has ever had and one of the least visible. It does not show up in the gradebook. A 2026 arXiv preprint on human-AI collaboration in a live capture-the-flag competition, written by researchers at William & Mary and IBM, named the risk in almost these terms. It warned of “performance without transfer,” where learners complete tasks more efficiently with AI guidance but fail to build skills that persist once the guidance is removed. It also pointed to education research where AI-generated hints improved students’ immediate success while the advantage disappeared when the support was withdrawn. The student looked more capable. The student was not more capable. The assistant was.
This is why a program can look like it is improving while the operators inside it stay exactly where they were. Completion rates climb, solve times drop, scores rise, and every one of those signals can be explained by the support structure getting better rather than the people getting better. You end up watching the scaffolding improve and crediting the student. The one measurement that would separate the two, what a person can do once the scaffolding is gone, is the one almost nobody collects, because collecting it means deliberately taking things away.
That is the move most programs skip. Readiness is not something you observe. It is something you isolate by engineering the subtraction. Change the conditions so the technique that worked last time is not the technique that works this time. Strip the framing so the task does not arrive pre-sorted into a category. Pull the tool, close the assistant, and put the person in front of something they have to reason through alone. What they do in that moment is the only number that was ever really about them.
There is a catch, and it is an honest one. Subtraction has to be continuous to mean anything. The first time you remove the supports, you learn something true about the operator. The second time, against the same scenario, you are back to measuring memory, because a scenario someone has already seen has quietly become its own kind of scaffold. To keep testing the residue, you need a steady supply of conditions the student has not encountered. That is the slow, expensive part to produce by hand. It is also why most programs build a single unsupported assessment and call the question answered.
We build the part that makes the subtraction sustainable: environments varied enough and disposable enough that the help can always be taken away again. The principle holds regardless of what you use to get there. The number on the board belongs to the student and everything assisting them, together. If you want to know what your program actually produced, stop adding support and start removing it, then watch what is still standing.
Sources
- Soderstrom, N. C., & Bjork, R. A. (2015). Learning versus performance: An integrative review. Perspectives on Psychological Science, 10(2), 176–199. https://doi.org/10.1177/1745691615569000
- Tang, T., Janis, N., Montague, K. A., Eykholt, K., Kirat, D., Park, Y., Jang, J., Nadkarni, A., & Xiao, Y. (2026). Understanding Human-AI Collaboration in Cybersecurity Competitions. arXiv preprint arXiv:2602.20446. https://arxiv.org/abs/2602.20446