Are Stated Reasoning Steps Causally Load-Bearing?

cs.AI updates on arXiv.org · 1h ago
Research Papers

arXiv:2609.27038v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring assumes that the reasoning a model writes reflects the computation that directly produces its answer. Previous faithfulness metrics have been predominantly behavioral, as they simply edit the reasoning text and observe the resulting answer. However, our methodology aims to measure faithfulness causally at the…

Read original article on cs.AI updates on arXiv.org →