The horse seemed to know mathematics. Someone asked a calculation and he answered by tapping his hoof. Spectators counted the taps; he stopped at the correct result. To the audience, the evidence was plain: question, answer, success.
Clever Hans became famous in the early twentieth century. The interesting part is not mocking those who believed him. Establishing that there was no deliberate trick did not explain what was happening. Hans could still succeed without his trainer, which seemed to strengthen the interpretation that he was calculating.
Oskar Pfungst’s controls changed the question. What if the person asking did not know the answer? What if Hans could not see them? Performance no longer held up. The horse had learned to respond to subtle cues from the observer rather than perform the attributed calculations.
A century later, that gap between what we observe and what we believe we have demonstrated is uncomfortable again when evaluating AI. A correct answer can be useful and still tell us less than we would like about the system’s capacity.
Melanie Mitchell discusses such cases in “Six principles for evaluating cognitive capabilities in AI models”. Her proposal is not to declare models incapable of reasoning. It asks for controls before turning good performance into broad claims about understanding, reasoning or intelligence.
The number that arrives before the question
“The model solved ninety out of a hundred problems” describes an observation. “The model understands the concept” proposes an explanation. Between them, we need evidence that the test measures that concept and that other strategies cannot explain the correct answers.
Evaluation calls this construct validity. A construct is the capacity under study, such as analogical reasoning. A test contains observable behavior intended to provide evidence about that capacity. Putting “reasoning” in a benchmark’s name does not establish the relationship.
Imagine a fictional sequence-completion test where the correct answer is always the longest option. A system might exploit that detail. If examples circulate online, it might recognize close versions. If a new question resembles a familiar one, an exact duplicate is not necessary for an alternative explanation.
The answer remains correct on the test. What changes is how far we can generalize. For an application builder, this matters: an assistant that works with familiar templates may fail when a user phrases the problem differently.
I find “under what conditions does it work?” more useful than attempting to settle whether a machine thinks like a person. The first question leads to an evaluation design. The second can occupy a long conversation without changing a single test.
Six principles at the workbench
Mitchell’s six principles can become review habits. They are not a new score or a guarantee that we will uncover the model’s internal mechanism.
Recognize anthropomorphic bias. Fluent explanations encourage us to attribute intention and understanding. Translate that impression into a testable hypothesis: which task should the system solve, with what information and under which variations?
Test alternative explanations. If a superficial cue could suffice, remove it. This was the decisive move with Hans: change a condition separating two explanations rather than repeat the impressive demonstration.
Create novel variations. Change names, symbols or context while preserving the relevant relationship. A drop in performance deserves investigation. Check that the variation has not introduced a genuinely new difficulty, too.
Investigate how success occurs. Response patterns and controlled interventions provide additional evidence. A verbal explanation is another behavior to study, not an infallible window onto the system’s internal processes.
Separate competence from performance. Failure can reflect the capacity we intend to measure, but also formatting, perception or execution. A model identifying a grid transformation but failing to draw it presents two tasks worth distinguishing.
Study errors and negative results. Unexpected responses can expose families of failures. Hiding them behind an average loses useful information and leaves us documenting only what worked.
A small experiment before a large conclusion
Consider a fictional assistant deciding whether a booking satisfies a cancellation policy. We provide the policy in context rather than expect memorized knowledge. We want to test whether it applies the conditions correctly.
I would first create simple, manually verifiable cases: inside and outside the cancellation window, an explicit exception and a missing required date. Then I would create paired variants: change the hotel name, reorder sentences or use different dates preserving the same temporal relationship.
Inventing a date in the incomplete case is a failure even if the final decision happens to match an expected answer. This changes the metric we need. Binary accuracy considering only “allow / deny” could overlook fabricated evidence.
I would repeat cases under a recorded configuration: model version, instructions, parameters and actual execution date. I would retain inputs and outputs, inspect aggregate results and compare variants. I would not publish numbers for this example without executing that protocol.
Failure only after changing names suggests a possible shortcut. Failure with one date format might suggest an interpretation difficulty. Different responses to identical inputs raise consistency questions. These are hypotheses, not definitive diagnoses. More controls are needed to separate them.
There is another practical benefit to paired examples: a single aggregate score no longer tells the whole story. We can see whether an irrelevant change flips a decision and examine exactly which input changed. For a user facing that decision, the individual flip matters even if the average looks satisfactory.
What I would write in the report
I would avoid “the model understands policies.” Instead, describe the supplied rules, tested case families, tolerated variations and observed failures. If it detects missing information but mishandles exceptions, that distinction belongs in the summary rather than an appendix.
I would also distinguish robustness, generalization and consistency. Changing an irrelevant word probes a type of robustness. Testing another policy family asks about transfer. Repeating an input examines stability. These questions are connected but not interchangeable.
This does not make benchmarks useless. They compare performance under known conditions. Trouble begins when conclusions move so far beyond those conditions that a score supports a story the test never examined.
Hans leaves a practical lesson: the most valuable control may be the one that spoils the demonstration. If removing an accidental cue breaks performance, we have learned something. That knowledge helps build a reliable application more than another brilliant, unexplained percentage.

Found this useful? If you would like to support this space, you can buy me a coffee.
Buy me a coffee Optional support through PayPal. You choose the amount.

