All About Optimization: From Gradient Descent to Muon

How models actually get trained: the optimizer family tree, why Adam won and where it loses, the schedule and warmup decisions that matter more than the optimizer choice, and the practical machinery (clipping, accumulation, checkpointing, sharded states) that shows up in real training runs.

September 15, 2026 · 21 min · Abdullah Al Mamun