Some of the most capable models are also the hardest to inspect. One common way to understand them is to build a simpler, interpretable model that mimics the complex one, a surrogate, and study the surrogate instead. Done well, this offers real insight into how a system behaves. Done carelessly, it offers a confident story that may not match what the real model actually does, which can be worse than no explanation at all.
Key Takeaways
- A surrogate approximates a complex model with a simpler, interpretable one.
- Surrogates can give genuine insight into behavior that is otherwise hard to inspect.
- An unfaithful surrogate produces explanations that mislead rather than clarify.
- The fidelity of a surrogate must be measured, not assumed, and its scope kept honest.
The ProblemBlack boxes resist inspection
A complex model may make excellent decisions while offering no legible account of why. For high-stakes use, that opacity is a problem: reviewers, regulators, and the people affected often need to understand the basis of a decision, not just its output. Surrogates are one response, approximate the complex model with something simple enough to read. But the approximation is exactly where the danger lies, because a simple model that usually agrees with a complex one can still diverge on the cases that matter, and its clean explanation can lend those divergences a false credibility.
Why It MattersAn explanation that misleads is worse than none
People act on explanations. If a surrogate suggests a model relies on reasonable factors when the real model actually relies on something else, the explanation does not just fail to help; it actively misleads, giving false confidence in a system that has not earned it. In high-stakes settings, that is a serious failure, because decisions and oversight are built on the belief that the explanation is true. The value of an interpretable approximation is entirely contingent on its faithfulness, and an explanation that is trusted but wrong can be more harmful than honest opacity.
The TeraSystemsAI PerspectiveUse surrogates, but measure their faithfulness
Our position is that surrogate models are a legitimate and useful tool, treated with the right skepticism. The central discipline is measuring fidelity: how closely, and under what conditions, the surrogate actually tracks the model it stands in for, and where it does not. An honest surrogate comes with its own limits stated plainly, this approximation holds here, and breaks down there, so that the insight it offers is not overextended. Interpretability is valuable, but only when the interpretable object faithfully represents the real one; otherwise it is a comforting fiction. The goal is understanding you can trust, not a story that merely sounds explanatory.
Practical ImplicationsFidelity, scope, and honest claims
In practice, using surrogates responsibly means quantifying how well the surrogate agrees with the underlying model, overall and especially on the consequential cases, rather than assuming agreement. It means being explicit about the surrogate's scope, the regions where it is faithful and the regions where it should not be trusted. It means resisting the temptation to present a clean surrogate explanation as the whole truth about a complex system. And it means pairing surrogate insight with other evidence about the model's behavior, so that understanding rests on more than a single approximation. Used this way, surrogates illuminate; used carelessly, they mislead.
Work with us on trustworthy AI
Join a community of researchers and engineers building accountable, evidence-grounded systems.
Join the Community