
LLM Inference Optimization: State-of-the-Art Research
Author(s): David Spuler (Author)
- Publisher: Aussie AI Labs
- Publication Date: 30 May 2026
- Language: English
- Print length: 2420 pages
- ISBN-10: B0H3FKR39T
Book Description
For research engineers, scientists and academics with cutting-edge new ideas in state-of-the-art LLM inference optimization. An extensive survey of LLM Inference optimization with research papers and clarity about what’s new, what’s widely used in industry, and what’s coming soon from research labs.
Highlights:
– 500+ LLM inference optimization techniques
– 100+ detailed chapters on the best ones.
– 50+ breakthroughs: new, emerging, and future research. Table of Contents:
Part I: Introduction to LLM Inference Optimization
1. Speed is the Need
2. Introduction to LLM Inference Optimization
3. SOTA Inference Stacks
4. Accuracy Versus Speed Part II: Top-Level Architecture Optimizations
5. Harness Engineering
6. Deployment Architecture
7. RAG Architectures
8. Agentic Architectures
9. Adaptive Inference Part III: Quantization
10. Quantization
11. Low-Bit Quantization
12. Block-Scaled Quantization
13. Fixed-Point & Block-Floating Point
14. Bitshift Quantization
15. Weight Clustering
16. Rare Quantization Part IV: Compute Optimizations
17. Prefill Optimizations
18. Serving Disaggregation
19. Mixture-of-Experts (MoE)
20. MoE Attention
21. FFN Optimizations
22. FFN Fusion
23. FFN Folding
24. Flash FFN Part V: Transformer Components
25. Tokenizer & Vocabulary
26. Unembedding
27. Model Loader
28. Encoders & Decoders
29. Activation Functions
30. Softmax
31. Normalization
32. Positional Encoding
33. Embedding & Unembedding Part VI: Attention Kernels
34. Attention Overview
35. Flash Attention
36. Paged Attention
37. Radix Attention
38. MLA Attention
39. QKV Compute
40. DeepSeek Sparse Attention
41. Other Attention Kernels
42. Hybrid Attention
43. Long & Infinite Context Part VII: Decoding Algorithms
44. Decoding Algorithms
45. Speculative Decoding
46. Optimizing Speculative Decoding
47. Model-Based Spec Dec
48. Parallel Drafting
49. Eagle, Medusa, & FR-Spec
50. Lookup Decoding
51. Generalized Spec Dec
52. Structured Decode
53. Aggressive Decoding
54. Edit Decoding Part VIII: Kernel Optimization
55. Kernel Optimizations
56. Bitwise Operations
57. Floating-Point Bit Tricks
58. Kernel Fusion
59. Fusion Examples
60. CUDA Optimizations
61. Grace CPU
62. Hopper & Blackwell
63. Rubin and Feynman
64. Vectorization
65. Branchless Code Part IX: Matrix Multiplication Kernels
66. MatMul/GEMM
67. MatMul Research
68. Strassen & Winograd Part X: Caching
69. Caching
70. KV Caching
71. RAG Caching
72. KV Cache Compression
73. KV Pruning and Fusion
74. Mooncake
75. KV Correction
76. Prefix KV
77. Non-Prefix KV Part XI: Pruning
78. Pruning
79. Unstructured Pruning
80. Structured Pruning
81. Depth Pruning
82. Layer Pruning
83. Early Exit
84. Layer Skipping
85. Shallow Decoder/Prefill
86. Layer Reordering
87. Width Pruning
88. Length Pruning
89. Prompt/Context Compress
90. Token Pruning
91. Zero Skipping
92. Weight Sharing Part XII: Data Structures & Algorithms
93. Radix Trees
94. LCS
95. Dot Product
96. Top-K
97. Vector Norms
98. Vector DB
99. Tensors
100. LUTs & Precompute
101. Conditional Compute
102. Weight Precomputation Part XIII: Advanced Research
103. Number Systems
104. Distillation
105. Zero-Multiplication
106. Additive Models
107. Arithmetic Optimization
108. Ensemble Architectures
109. Logarithmic Models
110. NAS Part XIV: Reasoning Model Optimizations
111. RIO
112. CoT Token Reduction
113. Small Reasoning Models
114. Concept Tokens
Appendix 1. 500 Techniques
– 500+ LLM inference optimization techniques
– 100+ detailed chapters on the best ones.
– 50+ breakthroughs: new, emerging, and future research. Table of Contents:
Part I: Introduction to LLM Inference Optimization
1. Speed is the Need
2. Introduction to LLM Inference Optimization
3. SOTA Inference Stacks
4. Accuracy Versus Speed Part II: Top-Level Architecture Optimizations
5. Harness Engineering
6. Deployment Architecture
7. RAG Architectures
8. Agentic Architectures
9. Adaptive Inference Part III: Quantization
10. Quantization
11. Low-Bit Quantization
12. Block-Scaled Quantization
13. Fixed-Point & Block-Floating Point
14. Bitshift Quantization
15. Weight Clustering
16. Rare Quantization Part IV: Compute Optimizations
17. Prefill Optimizations
18. Serving Disaggregation
19. Mixture-of-Experts (MoE)
20. MoE Attention
21. FFN Optimizations
22. FFN Fusion
23. FFN Folding
24. Flash FFN Part V: Transformer Components
25. Tokenizer & Vocabulary
26. Unembedding
27. Model Loader
28. Encoders & Decoders
29. Activation Functions
30. Softmax
31. Normalization
32. Positional Encoding
33. Embedding & Unembedding Part VI: Attention Kernels
34. Attention Overview
35. Flash Attention
36. Paged Attention
37. Radix Attention
38. MLA Attention
39. QKV Compute
40. DeepSeek Sparse Attention
41. Other Attention Kernels
42. Hybrid Attention
43. Long & Infinite Context Part VII: Decoding Algorithms
44. Decoding Algorithms
45. Speculative Decoding
46. Optimizing Speculative Decoding
47. Model-Based Spec Dec
48. Parallel Drafting
49. Eagle, Medusa, & FR-Spec
50. Lookup Decoding
51. Generalized Spec Dec
52. Structured Decode
53. Aggressive Decoding
54. Edit Decoding Part VIII: Kernel Optimization
55. Kernel Optimizations
56. Bitwise Operations
57. Floating-Point Bit Tricks
58. Kernel Fusion
59. Fusion Examples
60. CUDA Optimizations
61. Grace CPU
62. Hopper & Blackwell
63. Rubin and Feynman
64. Vectorization
65. Branchless Code Part IX: Matrix Multiplication Kernels
66. MatMul/GEMM
67. MatMul Research
68. Strassen & Winograd Part X: Caching
69. Caching
70. KV Caching
71. RAG Caching
72. KV Cache Compression
73. KV Pruning and Fusion
74. Mooncake
75. KV Correction
76. Prefix KV
77. Non-Prefix KV Part XI: Pruning
78. Pruning
79. Unstructured Pruning
80. Structured Pruning
81. Depth Pruning
82. Layer Pruning
83. Early Exit
84. Layer Skipping
85. Shallow Decoder/Prefill
86. Layer Reordering
87. Width Pruning
88. Length Pruning
89. Prompt/Context Compress
90. Token Pruning
91. Zero Skipping
92. Weight Sharing Part XII: Data Structures & Algorithms
93. Radix Trees
94. LCS
95. Dot Product
96. Top-K
97. Vector Norms
98. Vector DB
99. Tensors
100. LUTs & Precompute
101. Conditional Compute
102. Weight Precomputation Part XIII: Advanced Research
103. Number Systems
104. Distillation
105. Zero-Multiplication
106. Additive Models
107. Arithmetic Optimization
108. Ensemble Architectures
109. Logarithmic Models
110. NAS Part XIV: Reasoning Model Optimizations
111. RIO
112. CoT Token Reduction
113. Small Reasoning Models
114. Concept Tokens
Appendix 1. 500 Techniques
Wow! eBook


