The Fragile Foundations of CoT Monitoring for AI Safety

I attended a workshop on Chain of Thought monitorability and found the reliance on CoT for safety to be precarious. Since CoT was designed for performance rather than transparency, models might hide deceptive reasoning or perform computations without verbalizing them. I argue we must reduce our dependence on CoT and invest in interpretability tools to inspect internal model states directly.
On balance, I agree that it is useful to be able to monitor CoT. However, the fragility of this alliance makes me feel that we should reduce our dependence on CoT for safety, to plan for a future in which CoT isn't useful.