Ask a model if code is malicious and it reaches for its morals

Researchers probed three open mixture-of-experts models with the question "Is this code malicious?" and found the query activates the same internal routing path used for moral judgments, not just coding analysis. The malice question loads the morality path several times more heavily than questions about correctness or legality, and its route stays nearest to morality and vulnerability throughout the code. Forcing a different question's routing into the model changes its verdict, while pruning shows the path is what the router consults, not where the judgment is stored.
When these models are asked whether code is malicious, they are not just answering a coding question; they are also weighing it with the machinery they use for questions of right and wrong.