by Haytham ElFadeel - [email protected] September 6, 2026

Every deep learning practitioner carries around a set tradition, today we open up one of them ‘Weight Norm and Weight Decay’. This essay explore what does weight decay actually do, and it relationship with learning rate, Grokking, effective learning rate, Muon, Spectral spheres and Maximal Update Parametrization.

1. Why Would Magnitude Matter?

The classical reasons are:

We rarely ask why. Almost every modern architecture is full of normalization layers (LayerNorm, RMSNorm, QK-Norm). A linear layer followed by a normalization is scale-invariant:

$f(\alpha W) = f(W) \qquad \forall, \alpha > 0$ .

Multiply the weights by 10 or by 0.1 and the network computes exactly the same function. So in what sense can "small weights" be better? For these layers, a weight-norm penalty can't regularize the function at all, because there is nothing to regularize. So why we care about Weight decay and weight magnitude. Van Laarhoven (2017) pointed this out years ago, and Zhang et al. (2019) gave the answer this blog post builds on. Weight decay still matters a lot in these networks, but through the optimization dynamics, not through the function class.

So here's the question behind the rest of this post:

If the norm is a free choice, why does controlling it matter so much?

Spoiler alert / TLDR: The magnitude of the weights doesn't set what the network computes. It sets how big a step the optimizer takes in function space. Without weight decay, the weights grow and that effective step size quietly shrinks until the network stops learning new features. Weight decay, Muon, spectral spheres and μP are all ways of controlling it.


2. Weight Decay 101: Penalty vs. Constraint

2.1 SGD: weight decay is an L2 penalty

For plain SGD weight decay is a penalty. Take the regularized loss:

$\hat f(x) = f(x) + \frac{\lambda}{2}\lVert x\rVert_2^2$ .

One SGD step on $\hat f$ gives

$x_{t+1} = x_t - \eta\big(\nabla f(x_t) + \lambda x_t\big) = (1-\eta\lambda),x_t - \eta\nabla f(x_t),$

which is exactly "shrink the weights by $(1-\eta\lambda)$, then take a gradient step”. So L2 regularization and weight decay are the same thing here. The optimizer accepts a slightly higher $f$ in exchange for a smaller $x$. This is a penalty.

2.2 Adam and other adaptive optimizers breaks the equivalence