A metric that does not vanish

Most faithfulness tests of the last two years work by catching the model out. Plant a hint in the prompt, see whether the model uses it, and check whether the explanation admits to using it. Those tests taught us a lot, and they have a problem the authors of this paper name directly. They rely on adversarial vulnerabilities and reasoning errors that get rarer as models improve, so the signal shrinks with capability. The Claude Sonnet 4.5 model card is quoted saying there are currently no viable dedicated evaluations for reasoning faithfulness.

The paper, by Harry Mayne, Justin Singh Kang, Dewi Gould, Kannan Ramchandran, Adam Mahdi and Noah Siegel, posted on February 2, proposes measuring what an explanation reveals instead of what failure it exposes. The principle is that a faithful explanation should help an observer predict how the model will behave on related inputs. If the model says it decided a heart disease case because of the patient's age, then someone reading that explanation should do better at predicting what the model says about the same patient thirty years younger.

How NSG is computed

The setup has a reference model and a predictor model. The reference model answers a question and explains itself. The predictor sees the original question, the reference model's answer and a counterfactual question, and predicts the reference model's answer to the counterfactual twice, once without the explanation and once with it. Then the reference model actually answers the counterfactual, and the two predictions are scored. Simulatability gain is the accuracy with the explanation minus the accuracy without it.

The normalisation is the part that makes the number comparable across models. When baseline accuracy is already high there is little room to improve, so the gain is divided by one minus the baseline accuracy. Normalized Simulatability Gain is the fraction of the achievable improvement that the explanation delivers. An NSG of one means perfect counterfactual prediction, zero means the explanation added nothing, and a negative value means it misled the predictor.

Counterfactuals were drawn from real data rather than generated. The authors used seven tabular datasets, among them heart disease, Pima diabetes, breast cancer recurrence, employee attrition, income prediction, bank marketing and Moral Machines, converted rows into natural language questions, and picked neighbours within a Hamming distance of two features. That choice avoids the medically incoherent edits that single-concept perturbation methods can produce, and it keeps the test about the stated reasoning rather than about the model's world knowledge. There were 1,000 pairs per dataset, giving 7,000 in total.

What 18 models did

The reference models spanned the Qwen 3 and Gemma 3 families on the open side and GPT-5, Claude 4.5 and Gemini 3 on the proprietary side. Predictions were averaged over five predictor models, gpt-oss-20b, Qwen3-32B, gemma-3-27b-it, GPT-5 mini and gemini-3-flash, to avoid leaning on one judge. Every reference model produced explanations that helped. Absolute simulatability gain ran from 3.8 to 10.8 points and NSG from 11.0 to 36.5 percent. For the best models, explanations fixed roughly a third of the predictions that would otherwise have been wrong.

The scale trend was mixed. Qwen 3 improved monotonically with parameter count from 0.6B to 32B and Gemma 3 trended upward, while proprietary models showed no clear ordering. The authors' reading is that weak models have weak faithfulness and the relationship flattens past a modest capability threshold. Extra reasoning effort bought little. The dataset mattered a great deal, with NSG lowest on Moral Machines at 6 percent and highest on Pima diabetes at almost 43 percent.

The misleading tail and the self-knowledge result

The authors define an explanation as egregiously unfaithful when it causes every predictor to get the counterfactual wrong. Small open models like Qwen3-0.6B and gemma-3-1b-it produced these around 15 percent of the time, frontier models around 7 percent, and Claude Haiku 4.5 sat at 8.8 percent. They also ask which features drive these failures and find, in income prediction for example, a feature with a relative risk of 1.83, meaning explanations that involve that feature are much more likely to send the predictor the wrong way. The average case is positive and the tail is real.

The result we found most interesting is the comparison with external explanations. They had a different, sometimes stronger, model that had only input and output access write an explanation for the reference model's answer, then measured NSG for both. Self-explanations won in every family, by 3.0 points for Qwen 3, 1.7 for Gemma 3, 4.3 for GPT-5, 2.3 for Claude 4.5 and 2.7 for Gemini 3, all with confidence intervals excluding zero. The explaining model knew something about itself that an outside observer could not recover from behaviour alone.

Why simulatability is the right frame

The hint tests were asking whether explanations are ever dishonest, and the answer was yes, which is true but hard to act on. NSG asks how much you learn from reading the explanation, which is the question a person overseeing a model actually has. It also keeps working as models get better, because a relevant counterfactual should always change behaviour, so there is always something to predict. The two frames are not in conflict. One measures a failure mode, the other measures useful information, and both can be true of the same model.

The limitations are stated plainly and we would repeat them. The benchmark is binary classification on tabular data, and free-text generation is an open problem. The predictors are not frontier models, so the measured NSG is a floor. The metric is average case and gives no worst-case guarantee, and the paper is explicit that it says nothing about a model that is trying to hide its reasoning. What we would want next is NSG measured under a training incentive to obfuscate, and NSG used as a training signal to see whether the number can be pushed up without the explanations becoming empty.

Sources

  1. Mayne, Kang, Gould, Ramchandran, Mahdi and Siegel, A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior (arXiv 2602.02639)