counterfactual explanation
/ kown-ter-FAK-choo-ul ek-spluh-NAY-shun /
A counterfactual explanation answers a refreshingly human question: "What would need to change for the model to decide differently?" Instead of trying to dissect the model's tangled reasoning, it gives you a concrete alternative: "Your loan was denied. Had your annual income been $5,000 higher, it would have been approved." It explains a decision not by opening the box, but by showing you the nearest world in which the answer flips.
The appeal is that this matches how people naturally reason about decisions — we understand things through small what-ifs. A good counterfactual is minimal (change as little as possible), realistic (you can't ask someone to change their age or birthplace), and actionable (it points to something the person could actually do). When all three hold, it becomes recourse: not just an explanation but practical advice on how to get a different outcome next time.
Counterfactuals matter because they are often the most useful and least misleading form of explanation for the person on the receiving end of a decision. They sidestep the trap of pretending to reveal the model's true inner logic — they only describe the model's behavior near one input, which is exactly what an affected person needs. The honest caveats: a model usually has many possible counterfactuals (which to show is a choice with consequences), and a counterfactual describes the model's quirks, not necessarily the real-world causes — following its advice changes the model's verdict, not always your underlying situation.
A job-screening model rejects an applicant. A counterfactual explanation says: "With one more year of relevant experience, this application would have been accepted." That's actionable and clear. Contrast a bad one: "Had you been five years younger, you'd have been accepted" — technically a counterfactual, but useless as advice and a red flag for age discrimination.
A useful counterfactual points to something you can actually change — and can expose bias.
A counterfactual tells you how to flip the model's decision, which is not the same as how to change your real-world outcome. If the model leans on a spurious feature, gaming that feature may flip the verdict without improving anything real. Counterfactuals are advice about the model, so they're only as trustworthy as the model behind them.