Measuring Chain-of-Thought Necessity and the Effect of Reasoning Constraints
Why readCoT monitoring is a key oversight tool for AI safety, but only if the CoT faithfully reflects the reasoning behind the answer. If it doesn't, monitoring it tells us less about why the model chose that answer.I extend the CoT faithfulness interventions from Lanham et al. (2023) to ethical reasoning. These interventions measure CoT necessity - whether the answer is dependent on the CoT content - which is one element in establishing faithfulness, though not sufficient on its own. Consistent...
Read full article →