Recorded session
Introduction to LLM Inference Engines
· 1 h 30 min
With Mohamed Rashad, Engineering Lead & Co-founder, Hyperion & DevisionX
- LLMOps
- Model serving
- Language
- Arabic & English
Recording
The recording will be posted here soon.
About this session
How inference engines turn a trained LLM into a fast, cheap, scalable and reliable service: token generation, memory, batching, scheduling, caching, quantization and hardware execution.
What you'll learn
- Where an inference engine sits between model weights and the application.
- Architectures of vLLM, TensorRT-LLM, SGLang, llama.cpp and others compared.
- KV cache management, continuous batching, speculative decoding, tensor parallelism and quantization.
- Choosing and configuring the right engine, from a single local GPU to high-throughput production clusters.

