What is neuralese and why is it bad?
What is Nueralese and Why is it Bad
Neuralese is the idea of replacing a model's natural-language chain-of-thought with a stream of numbers, making its reasoning unreadable to humans. This piece explains why that's dangerous: it would eliminate chain-of-thought monitoring, one of the few safety tools we have for tracking model intent. Citing recent leaks about OpenAI's Astra model, the author argues that even hybrid approaches are a slippery slope toward full neuralese, and urges journalists and employees to push back.
So moving away from monitorable chain-of-thought to neuralese seems very bad, with dubious benefits.