Skip to main content
MLOps MENA Academy

Recorded session

Introduction to LLM Inference Engines

· 1 h 30 min

With Mohamed Rashad, Engineering Lead & Co-founder, Hyperion & DevisionX

  • LLMOps
  • Model serving
Language
Arabic & English

Recording

The recording will be posted here soon.

About this session

How inference engines turn a trained LLM into a fast, cheap, scalable and reliable service: token generation, memory, batching, scheduling, caching, quantization and hardware execution.

What you'll learn

  • Where an inference engine sits between model weights and the application.
  • Architectures of vLLM, TensorRT-LLM, SGLang, llama.cpp and others compared.
  • KV cache management, continuous batching, speculative decoding, tensor parallelism and quantization.
  • Choosing and configuring the right engine, from a single local GPU to high-throughput production clusters.

Speakers