Had a lot of fun working with @rudrakshkarpe on this tutorial showing how to run @sgl_project inference on @kubernetesio using @HAMiProject GPU shares: complete with GPU memory quotas, compute throttling, and an OpenAI-compatible API @CloudNativeFdn
project-hami.io/tutorials/la…
Sep 3, 2026 · 6:05 PM UTC
GPU-share quotas make SGLang capacity predictable when throttling and OpenAI-compatible endpoints ship together. modelbeat.ai compares throughput, latency, and cost across serving profiles. @HowDevelop @kubernetesio #ModelBeat #Elytra #SGLang #Kubernetes
Looks great @HowDevelop and @rudrakshkarpe bhyia , Going to try this out .