The Softmax Function and Its Derivative: A Deep Dive

The Softmax Function and Its Derivative: A Deep Dive

Softmax turns any real vector into probabilities that sum to 1, making it a staple of multiclass classification. This post derives its Jacobian from first principles using vector calculus, explains why naive exponentiation causes NaNs and how shifting by the maximum fixes it, and shows how the derivative simplifies when combined with cross-entropy loss.

Intuitively, the softmax function is a "soft" version of the maximum function. Instead of just selecting one maximal element, softmax breaks the vector up into parts of a whole (1.0) with the maximal input element getting a proportionally larger chunk, but the other elements getting some of it as well.

More from this day

2026-10-03