Chain-of-thought faithfulness
AIChain-of-thought faithfulness: Whether a model's stated reasoning matches the computation that actually produced its answer, which experiments show is often not the case.
Chain-of-thought faithfulness: Whether a model's stated reasoning matches the computation that actually produced its answer, which experiments show is often not the case.