Recorded session
The On-Prem MLOps Playbook
· 1 h 25 min
With Mohamed Rashad, Engineering Lead & Co-founder, Hyperion & DevisionX
- On-prem
- MLOps
- Language
- Arabic & English
Recording
About this session
Bridging traditional DevOps and private AI workloads. Banks, manufacturers and government environments run every inference on hardware they physically control: no elastic GPU clusters, no managed scheduling, one fixed pool of GPUs shared by training, inference and experimentation. Most MLOps content quietly assumes the cloud; this session covers what changes when that assumption disappears, from GPU partitioning and scheduling to CUDA and driver versioning.
Why attend
In MENA, on-prem is the default for banks, telecoms, government and healthcare due to data-residency rules, while most MLOps content assumes the cloud. A production playbook from regulated deployments, plus Q&A with a CTO. For MLOps/DevOps engineers moving into AI workloads, ML engineers, and tech leads weighing cloud vs on-prem. Hosted by MLOps MENA Community in partnership with DevisionX.
What you'll learn
- Partitioning and scheduling a fixed GPU pool across training, inference and experimentation.
- Managing CUDA/driver versions without breaking running jobs.
- Monitoring utilization to catch idle or contended hardware.
- Designing around power, cooling and bandwidth.
- When on-prem is mandatory: data residency, air-gapped environments, regulation.

