
Gemma-3: Architecture and Mathematical Foundations
A layer-by-layer mathematical description of the Gemma-3 decoder, written alongside a from-scratch reference implementation in MintEngine and used to check numerical parity of every layer against production inference engines. Covers embeddings and scaling, the pre-normalized transformer layer with explicit residual handling, multi-query attention, RoPE, the gated MLP, and Gemma's RMSNorm variant.
