ResearcharXiv AI / CL
An Empirical Study of Counterfactual Self-Explanations in LLMs

Summary
Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior.
Original Article
Captured source content or English translation, normalized into this reading format.
This story does not yet have captured source text. Open the source link to read it.
Region
Global
Heat Score
76
Category
Research
Language
en
