Appendix F
Bibliography
Every cited source, with links to the primary papers.
[agarwal2023onpolicy] R. Agarwal et al. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes. 2023. arXiv:2306.13649
Cited in Chapter 38
[ainslie2023gqa] J. Ainslie et al. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. 2023. arXiv:2305.13245
Cited in Chapter 20, Chapter 24
[alayrac2022flamingo] J. Alayrac et al. Flamingo: a Visual Language Model for Few-Shot Learning. 2022. arXiv:2204.14198
Cited in Chapter 32
[anthropic2024agents] Anthropic. Building effective agents. Engineering blog, December 2024. https://www.anthropic.com/engineering/building-effective-agents
Cited in Chapter 42
[ba2016layer] J. L. Ba, J. R. Kiros, and G. E. Hinton. Layer Normalization. 2016. arXiv:1607.06450
Cited in Chapter 16, Appendix C
[bahdanau2014neural] D. Bahdanau, K. Cho, and Y. Bengio. Neural Machine Translation by Jointly Learning to Align and Translate. 2014. arXiv:1409.0473
Cited in Chapter 19
[baydin2015automatic] A. G. Baydin et al. Automatic Differentiation in Machine Learning: a Survey. 2015. arXiv:1502.05767
Cited in Chapter 10, Appendix C
[bengio2003] Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin. A neural probabilistic language model. Journal of Machine Learning Research 3, 1137–1155, 2003.
Cited in Chapter 18, Chapter 23
[bishop2006] C. M. Bishop. Pattern Recognition and Machine Learning. Springer, 2006.
Cited in Chapter 6, Chapter 7, Chapter 9, Chapter 13
[blitzstein2019] J. K. Blitzstein and J. Hwang. Introduction to Probability, 2nd edition. CRC Press, 2019. https://projects.iq.harvard.edu/stat110
Cited in Chapter 6
[bloc972023] bloc97. NTK-aware scaled RoPE allows LLaMA models to have extended context size without fine-tuning. Reddit r/LocalLLaMA post, June 2023.
Cited in Chapter 21
[bradley1952] R. A. Bradley and M. E. Terry. Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika 39, 324–345, 1952.
Cited in Chapter 35, Chapter 36
[bridle1990] J. S. Bridle. Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition. In Neurocomputing, Springer, 1990.
Cited in Chapter 12
[cai2024medusa] T. Cai et al. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads. 2024. arXiv:2401.10774
Cited in Chapter 39
[chen2016training] T. Chen et al. Training Deep Nets with Sublinear Memory Cost. 2016. arXiv:1604.06174
Cited in Chapter 10
[chen2021evaluating] M. Chen et al. Evaluating Large Language Models Trained on Code. 2021. arXiv:2107.03374
Cited in Chapter 43
[chen2023accelerating] C. Chen et al. Accelerating Large Language Model Decoding with Speculative Sampling. 2023. arXiv:2302.01318
Cited in Chapter 39
[chen2023extending] S. Chen et al. Extending Context Window of Large Language Models via Positional Interpolation. 2023. arXiv:2306.15595
Cited in Chapter 21
[chowdhery2022palm] A. Chowdhery et al. PaLM: Scaling Language Modeling with Pathways. 2022. arXiv:2204.02311
Cited in Chapter 11, Chapter 12
[christiano2017deep] P. Christiano et al. Deep reinforcement learning from human preferences. 2017. arXiv:1706.03741
Cited in Chapter 35
[cover2006] T. M. Cover and J. A. Thomas. Elements of Information Theory, 2nd edition. Wiley, 2006.
Cited in Chapter 7
[cs336] Stanford CS336: Language Modeling from Scratch. Course, 2026. https://cs336.stanford.edu/
Cited in Chapter 44
[dai2024deepseekmoe] D. Dai et al. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models. 2024. arXiv:2401.06066
Cited in Chapter 27
[dao2022flashattention] T. Dao et al. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. 2022. arXiv:2205.14135
Cited in Chapter 19, Chapter 26
[dao2023flashattention2] T. Dao. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning. 2023. arXiv:2307.08691
Cited in Chapter 19, Chapter 26
[deepseek2025v32exp] DeepSeek-AI. DeepSeek-V3.2-Exp. Source code and report, 2025. https://github.com/deepseek-ai/DeepSeek-V3.2-Exp
Cited in Chapter 25
[deepseekai2024deepseekv3] DeepSeek-AI et al. DeepSeek-V3 Technical Report. 2024. arXiv:2412.19437
Cited in Chapter 29, Chapter 41
[deepseekai2025deepseekr1] DeepSeek-AI et al. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. 2025. arXiv:2501.12948
Cited in Chapter 37, Chapter 38, Chapter 43
[dehghani2023scaling] M. Dehghani et al. Scaling Vision Transformers to 22 Billion Parameters. 2023. arXiv:2302.05442
Cited in Chapter 31
[dettmers2022llmint8] T. Dettmers et al. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale. 2022. arXiv:2208.07339
Cited in Chapter 40
[dettmers2023qlora] T. Dettmers et al. QLoRA: Efficient Finetuning of Quantized LLMs. 2023. arXiv:2305.14314
Cited in Chapter 33
[dosovitskiy2020image] A. Dosovitskiy et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. 2020. arXiv:2010.11929
Cited in Chapter 31
[efron1979] B. Efron. Bootstrap methods: another look at the jackknife. Annals of Statistics 7(1), 1–26, 1979.
Cited in Chapter 8
[fedus2021switch] W. Fedus, B. Zoph, and N. Shazeer. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. 2021. arXiv:2101.03961
Cited in Chapter 27
[frantar2022gptq] E. Frantar et al. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. 2022. arXiv:2210.17323
Cited in Chapter 40
[gloeckle2024better] F. Gloeckle et al. Better & Faster Large Language Models via Multi-token Prediction. 2024. arXiv:2404.19737
Cited in Chapter 39
[glorot2010] X. Glorot and Y. Bengio. Understanding the difficulty of training deep feedforward neural networks. AISTATS 2010.
Cited in Chapter 11, Chapter 14
[goldberg1991] D. Goldberg. What every computer scientist should know about floating-point arithmetic. ACM Computing Surveys 23(1), 5–48, 1991.
Cited in Appendix B
[goodfellow2016] I. Goodfellow, Y. Bengio, and A. Courville. Deep Learning. MIT Press, 2016. https://www.deeplearningbook.org
Cited in Chapter 9, Chapter 12, Chapter 13, Chapter 14, Appendix A
[grattafiori2024] A. Grattafiori et al. The Llama 3 herd of models. 2024. arXiv:2407.21783
Cited in Chapter 6, Chapter 9, Chapter 14, Chapter 44
[griewank2008] A. Griewank and A. Walther. Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation, 2nd edition. SIAM, 2008.
Cited in Chapter 10, Appendix C
[gu2023mamba] A. Gu and T. Dao. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. 2023. arXiv:2312.00752
Cited in Chapter 28
[harris2020] C. R. Harris et al. Array programming with NumPy. Nature 585, 357–362, 2020. arXiv:2006.10256
Cited in Appendix B
[he2015deep] K. He et al. Deep Residual Learning for Image Recognition. 2015. arXiv:1512.03385
Cited in Chapter 16
[he2015delving] K. He et al. Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. 2015. arXiv:1502.01852
Cited in Chapter 14
[hendrycks2016gaussian] D. Hendrycks and K. Gimpel. Gaussian Error Linear Units (GELUs). 2016. arXiv:1606.08415
Cited in Chapter 11
[higham2002] N. J. Higham. Accuracy and Stability of Numerical Algorithms, 2nd edition. SIAM, 2002.
Cited in Appendix B
[hinton2015distilling] G. Hinton, O. Vinyals, and J. Dean. Distilling the Knowledge in a Neural Network. 2015. arXiv:1503.02531
Cited in Chapter 13, Chapter 38
[hoffmann2022training] J. Hoffmann et al. Training Compute-Optimal Large Language Models. 2022. arXiv:2203.15556
Cited in Chapter 18, Chapter 22, Chapter 29, Chapter 44
[holtzman2019curious] A. Holtzman et al. The Curious Case of Neural Text Degeneration. 2019. arXiv:1904.09751
Cited in Chapter 23
[hu2021lora] E. J. Hu et al. LoRA: Low-Rank Adaptation of Large Language Models. 2021. arXiv:2106.09685
Cited in Chapter 33
[huang2018gpipe] Y. Huang et al. GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism. 2018. arXiv:1811.06965
Cited in Chapter 41
[huber1964] P. J. Huber. Robust estimation of a location parameter. Annals of Mathematical Statistics 35(1), 73-101, 1964.
Cited in Chapter 13
[ieee754] IEEE Standard for Floating-Point Arithmetic, IEEE 754-2019.
Cited in Appendix B
[ioffe2015batch] S. Ioffe and C. Szegedy. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. 2015. arXiv:1502.03167
Cited in Chapter 16
[jang2016] E. Jang, S. Gu, and B. Poole. Categorical reparameterization with Gumbel-softmax. ICLR 2017. arXiv:1611.01144
Cited in Chapter 6
[jimenez2023swebench] C. E. Jimenez et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? 2023. arXiv:2310.06770
Cited in Chapter 8, Chapter 43
[jordan2024muon] K. Jordan et al. Muon: An optimizer for hidden layers in neural networks. Blog post, 2024. https://kellerjordan.github.io/posts/muon/
Cited in Chapter 15
[kalamkar2019] D. Kalamkar et al. A study of BFLOAT16 for deep learning training. 2019. arXiv:1905.12322
Cited in Appendix B
[kaplan2020scaling] J. Kaplan et al. Scaling Laws for Neural Language Models. 2020. arXiv:2001.08361
Cited in Chapter 18, Chapter 29
[karpathy2023nanogpt] A. Karpathy. nanoGPT. Source code, 2023. https://github.com/karpathy/nanoGPT
Cited in Chapter 23, Chapter 44
[katharopoulos2020transformers] A. Katharopoulos et al. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. 2020. arXiv:2006.16236
Cited in Chapter 28
[kingma2013] D. P. Kingma and M. Welling. Auto-encoding variational Bayes. ICLR 2014. arXiv:1312.6114
Cited in Chapter 6
[kingma2014adam] D. P. Kingma and J. Ba. Adam: A Method for Stochastic Optimization. 2014. arXiv:1412.6980
Cited in Chapter 15, Chapter 23
[kudo2018sentencepiece] T. Kudo and J. Richardson. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing. 2018. arXiv:1808.06226
Cited in Chapter 17
[kudo2018subword] T. Kudo. Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates. 2018. arXiv:1804.10959
Cited in Chapter 17
[kullback1951] S. Kullback and R. A. Leibler. On information and sufficiency. Annals of Mathematical Statistics 22(1), 79–86, 1951.
Cited in Chapter 7, Chapter 36
[kwon2023efficient] W. Kwon et al. Efficient Memory Management for Large Language Model Serving with PagedAttention. 2023. arXiv:2309.06180
Cited in Chapter 40
[lambert2025reinforcement] N. Lambert. Reinforcement Learning from Human Feedback. 2025. arXiv:2504.12501
Cited in Chapter 34, Chapter 35
[lepikhin2020gshard] D. Lepikhin et al. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. 2020. arXiv:2006.16668
Cited in Chapter 27, Chapter 41
[leviathan2022fast] Y. Leviathan, M. Kalman, and Y. Matias. Fast Inference from Transformers via Speculative Decoding. 2022. arXiv:2211.17192
Cited in Chapter 39
[lewis2020retrievalaugmented] P. Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. 2020. arXiv:2005.11401
Cited in Chapter 43
[li2023blip2] J. Li et al. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. 2023. arXiv:2301.12597
Cited in Chapter 32
[li2024datacomplm] J. Li et al. DataComp-LM: In search of the next generation of training sets for language models. 2024. arXiv:2406.11794
Cited in Chapter 29
[lieber2024jamba] O. Lieber et al. Jamba: A Hybrid Transformer-Mamba Language Model. 2024. arXiv:2403.19887
Cited in Chapter 28
[lightman2023let] H. Lightman et al. Let’s Verify Step by Step. 2023. arXiv:2305.20050
Cited in Chapter 38
[lin2017focal] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár. Focal Loss for Dense Object Detection. 2017. arXiv:1708.02002
Cited in Chapter 13
[lin2023awq] J. Lin et al. AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration. 2023. arXiv:2306.00978
Cited in Chapter 40
[liu2023visual] H. Liu et al. Visual Instruction Tuning. 2023. arXiv:2304.08485
Cited in Chapter 32
[liu2025muon] J. Liu et al. Muon is Scalable for LLM Training. 2025. arXiv:2502.16982
Cited in Chapter 15
[liu2025understanding] Z. Liu et al. Understanding R1-Zero-Like Training: A Critical Perspective. 2025. arXiv:2503.20783
Cited in Chapter 37
[loshchilov2016sgdr] I. Loshchilov and F. Hutter. SGDR: Stochastic Gradient Descent with Warm Restarts. 2016. arXiv:1608.03983
Cited in Chapter 15
[loshchilov2017decoupled] I. Loshchilov and F. Hutter. Decoupled Weight Decay Regularization. 2017. arXiv:1711.05101
Cited in Chapter 15, Chapter 23
[mackay2003] D. J. C. MacKay. Information Theory, Inference, and Learning Algorithms. Cambridge University Press, 2003. https://www.inference.org.uk/mackay/itila/
Cited in Chapter 7
[maddison2014] C. J. Maddison, D. Tarlow, and T. Minka. A* sampling. NeurIPS 2014. arXiv:1411.0030
Cited in Chapter 6
[mcnemar1947] Q. McNemar. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12(2), 153–157, 1947.
Cited in Chapter 8
[mcp2025] Model Context Protocol. Specification, version 2025-11-25. https://modelcontextprotocol.io/specification/2025-11-25
Cited in Chapter 42
[micikevicius2017] P. Micikevicius et al. Mixed precision training. ICLR 2018. arXiv:1710.03740
Cited in Chapter 16, Appendix B
[micikevicius2022fp8] P. Micikevicius et al. FP8 Formats for Deep Learning. 2022. arXiv:2209.05433
Cited in Chapter 40, Chapter 41
[minimax2025minimax01] MiniMax et al. MiniMax-01: Scaling Foundation Models with Lightning Attention. 2025. arXiv:2501.08313
Cited in Chapter 28
[nair2010] V. Nair and G. E. Hinton. Rectified linear units improve restricted Boltzmann machines. ICML 2010.
Cited in Chapter 11
[olmo2] Team OLMo et al. 2 OLMo 2 Furious. 2025. arXiv:2501.00656
Cited in Chapter 44
[oord2018] A. van den Oord, Y. Li, and O. Vinyals. Representation learning with contrastive predictive coding. 2018. arXiv:1807.03748
Cited in Chapter 7, Chapter 30
[ouyang2022training] L. Ouyang et al. Training language models to follow instructions with human feedback. 2022. arXiv:2203.02155
Cited in Chapter 9, Chapter 33, Chapter 35
[parr2018] T. Parr and J. Howard. The matrix calculus you need for deep learning. 2018. arXiv:1802.01528
Cited in Appendix A, Appendix C
[penedo2024fineweb] G. Penedo et al. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. 2024. arXiv:2406.17557
Cited in Chapter 29
[peng2023yarn] B. Peng et al. YaRN: Efficient Context Window Extension of Large Language Models. 2023. arXiv:2309.00071
Cited in Chapter 21
[polyak1964] B. T. Polyak. Some methods of speeding up the convergence of iteration methods. USSR Computational Mathematics and Mathematical Physics 4(5), 1–17, 1964.
Cited in Chapter 15
[press2016using] O. Press and L. Wolf. Using the Output Embedding to Improve Language Models. 2016. arXiv:1608.05859
Cited in Chapter 17, Chapter 22, Chapter 23
[press2021train] O. Press, N. A. Smith, and M. Lewis. Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation. 2021. arXiv:2108.12409
Cited in Chapter 21
[pytorch-linear] PyTorch documentation.
torch.nn.Linear. https://docs.pytorch.org/docs/stable/generated/torch.nn.Linear.htmlCited in Appendix A
[qiu2025gated] Z. Qiu et al. Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free. 2025. arXiv:2505.06708
Cited in Chapter 24
[radford2019] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. OpenAI technical report, 2019.
Cited in Chapter 18, Chapter 23
[radford2021learning] A. Radford et al. Learning Transferable Visual Models From Natural Language Supervision. 2021. arXiv:2103.00020
Cited in Chapter 30, Chapter 32
[rafailov2023direct] R. Rafailov et al. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. 2023. arXiv:2305.18290
Cited in Chapter 36
[rajbhandari2019zero] S. Rajbhandari et al. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. 2019. arXiv:1910.02054
Cited in Chapter 41
[ramachandran2017searching] P. Ramachandran, B. Zoph, and Q. V. Le. Searching for Activation Functions. 2017. arXiv:1710.05941
Cited in Chapter 11
[robbins1951] H. Robbins and S. Monro. A stochastic approximation method. Annals of Mathematical Statistics 22(3), 400–407, 1951.
Cited in Chapter 9
[rumelhart1986] D. E. Rumelhart, G. E. Hinton, and R. J. Williams. Learning representations by back-propagating errors. Nature 323, 533–536, 1986.
Cited in Chapter 10, Chapter 14
[scalingbook2025] Google DeepMind. How to scale your model. Online book, 2025. https://jax-ml.github.io/scaling-book/
Cited in Chapter 29, Chapter 44
[schroff2015facenet] F. Schroff, D. Kalenichenko, and J. Philbin. FaceNet: A Unified Embedding for Face Recognition and Clustering. 2015. arXiv:1503.03832
Cited in Chapter 30
[schulman2015highdimensional] J. Schulman et al. High-Dimensional Continuous Control Using Generalized Advantage Estimation. 2015. arXiv:1506.02438
Cited in Chapter 34
[schulman2017proximal] J. Schulman et al. Proximal Policy Optimization Algorithms. 2017. arXiv:1707.06347
Cited in Chapter 35, Chapter 37
[schulman2020] J. Schulman. Approximating KL divergence. Blog post, 2020. http://joschu.net/blog/kl-approx.html
Cited in Chapter 7
[sennrich2015neural] R. Sennrich, B. Haddow, and A. Birch. Neural Machine Translation of Rare Words with Subword Units. 2015. arXiv:1508.07909
Cited in Chapter 17
[shah2024flashattention3] J. Shah et al. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision. 2024. arXiv:2407.08608
Cited in Chapter 26
[shannon1948] C. E. Shannon. A mathematical theory of communication. Bell System Technical Journal 27, 379–423 and 623–656, 1948.
Cited in Chapter 7
[shao2024] Z. Shao et al. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. 2024. arXiv:2402.03300
Cited in Chapter 6, Chapter 7, Chapter 37
[shao2024deepseekv2] Z. Shao et al. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. 2024. arXiv:2405.04434
Cited in Chapter 25, Chapter 27
[shazeer2017outrageously] N. Shazeer et al. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. 2017. arXiv:1701.06538
Cited in Chapter 27
[shazeer2019fast] N. Shazeer. Fast Transformer Decoding: One Write-Head is All You Need. 2019. arXiv:1911.02150
Cited in Chapter 20, Chapter 24
[shazeer2020glu] N. Shazeer. GLU Variants Improve Transformer. 2020. arXiv:2002.05202
Cited in Chapter 11, Chapter 22
[shinn2023reflexion] N. Shinn et al. Reflexion: language agents with verbal reinforcement learning. 2023. arXiv:2303.11366
Cited in Chapter 43
[shoeybi2019megatronlm] M. Shoeybi et al. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. 2019. arXiv:1909.08053
Cited in Chapter 41
[snell2024scaling] C. Snell et al. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. 2024. arXiv:2408.03314
Cited in Chapter 38
[stiennon2020learning] N. Stiennon et al. Learning to summarize from human feedback. 2020. arXiv:2009.01325
Cited in Chapter 35
[su2021roformer] J. Su et al. RoFormer: Enhanced Transformer with Rotary Position Embedding. 2021. arXiv:2104.09864
Cited in Chapter 21, Chapter 25
[sutton2018] R. S. Sutton and A. G. Barto. Reinforcement Learning: An Introduction, 2nd edition. MIT Press, 2018. http://incompleteideas.net/book/the-book-2nd.html
Cited in Chapter 34
[szegedy2015rethinking] C. Szegedy et al. Rethinking the Inception Architecture for Computer Vision. 2015. arXiv:1512.00567
Cited in Chapter 12
[touvron2023llama] H. Touvron et al. LLaMA: Open and Efficient Foundation Language Models. 2023. arXiv:2302.13971
Cited in Chapter 22
[vaswani2017attention] A. Vaswani et al. Attention Is All You Need. 2017. arXiv:1706.03762
Cited in Chapter 12, Chapter 14, Chapter 19, Chapter 20, Chapter 21, Chapter 22, Chapter 24, Chapter 26, Chapter 31, Appendix C
[wang2019neural] C. Wang, K. Cho, and J. Gu. Neural Machine Translation with Byte-Level Subwords. 2019. arXiv:1909.03341
Cited in Chapter 17
[wang2022selfconsistency] X. Wang et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. 2022. arXiv:2203.11171
Cited in Chapter 38
[wang2024auxiliarylossfree] L. Wang et al. Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts. 2024. arXiv:2408.15664
Cited in Chapter 27
[wang2024qwen2vl] P. Wang et al. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. 2024. arXiv:2409.12191
Cited in Chapter 32
[wasserman2004] L. Wasserman. All of Statistics: A Concise Course in Statistical Inference. Springer, 2004.
Cited in Chapter 8
[williams1992] R. J. Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning 8, 229–256, 1992.
Cited in Chapter 6, Chapter 34
[xiao2022smoothquant] G. Xiao et al. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. 2022. arXiv:2211.10438
Cited in Chapter 40
[xiao2023efficient] G. Xiao et al. Efficient Streaming Language Models with Attention Sinks. 2023. arXiv:2309.17453
Cited in Chapter 24
[xiong2020layer] R. Xiong et al. On Layer Normalization in the Transformer Architecture. 2020. arXiv:2002.04745
Cited in Chapter 16, Chapter 22
[yang2022tensor] G. Yang et al. Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer. 2022. arXiv:2203.03466
Cited in Chapter 29
[yang2024gated] S. Yang, J. Kautz, and A. Hatamizadeh. Gated Delta Networks: Improving Mamba2 with Delta Rule. 2024. arXiv:2412.06464
Cited in Chapter 28
[yang2024sweagent] J. Yang et al. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. 2024. arXiv:2405.15793
Cited in Chapter 42
[yang2025qwen3] A. Yang et al. Qwen3 Technical Report. 2025. arXiv:2505.09388
Cited in Chapter 24
[yao2022react] S. Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models. 2022. arXiv:2210.03629
Cited in Chapter 42
[yao2024bench] S. Yao et al. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. 2024. arXiv:2406.12045
Cited in Chapter 43
[yu2025dapo] Q. Yu et al. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. 2025. arXiv:2503.14476
Cited in Chapter 37
[yuan2025native] J. Yuan et al. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention. 2025. arXiv:2502.11089
Cited in Chapter 25
[zhai2023sigmoid] X. Zhai et al. Sigmoid Loss for Language Image Pre-Training. 2023. arXiv:2303.15343
Cited in Chapter 30
[zhang2019root] B. Zhang and R. Sennrich. Root Mean Square Layer Normalization. 2019. arXiv:1910.07467
Cited in Chapter 16, Chapter 22, Appendix C
[zhang2025kimi] Y. Zhang et al. Kimi Linear: An Expressive, Efficient Attention Architecture. 2025. arXiv:2510.26692
Cited in Chapter 28
[zhao2023pytorch] Y. Zhao et al. PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel. 2023. arXiv:2304.11277
Cited in Chapter 10
[zheng2023judging] L. Zheng et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. 2023. arXiv:2306.05685
Cited in Chapter 8, Chapter 43
[zheng2025group] C. Zheng et al. Group Sequence Policy Optimization. 2025. arXiv:2507.18071
Cited in Chapter 37
[zoph2022stmoe] B. Zoph et al. ST-MoE: Designing Stable and Transferable Sparse Expert Models. 2022. arXiv:2202.08906
Cited in Chapter 27