teapawtteapawt
4chansoyjakawoo
Shivay Lamba
@HowDevelop
Sep 3
Had a lot of fun working with @rudrakshkarpe on this tutorial showing how to run @sgl_project inference on @kubernetesio using @HAMiProject GPU shares: complete with GPU memory quotas, compute throttling, and an OpenAI-compatible API @CloudNativeFdn project-hami.io/tutorials/la…

Lab 15: Run SGLang on HAMi GPU Shares | HAMi

Install HAMi on a GPU cluster and schedule SGLang inference services with GPU partitioning.

project-hami.io

Sep 3, 2026 · 6:05 PM UTC

3
2
1
12
1,366
RelevantRecentLikes
Shivay Lamba
@HowDevelop
Sep 3
Thanks to @rudrakshkarpe for leading this
1
2
145
Shivay Lamba
@HowDevelop
Sep 3
And @satyampsoni
1
1
108
Back to the original post: Shivay LambaHad a lot of fun working with @rudrakshkarpe on this tutorial showing how to run @sgl_project inference on @kubernetesio using @HAMiProject GPU shares: complete with GPU memory quotas, compute throttling, and an OpenAI-compatible API @CloudNativeFdn https://project-hami.io/tutorials/labs/hami-sglang
Aniket Tapre
@Aniketx
Sep 3
GPU-share quotas make SGLang capacity predictable when throttling and OpenAI-compatible endpoints ship together. modelbeat.ai compares throughput, latency, and cost across serving profiles. @HowDevelop @kubernetesio #ModelBeat #Elytra #SGLang #Kubernetes
7
Shivam Kumar@maishivamhoo
Sep 4
Looks great @HowDevelop and @rudrakshkarpe bhyia , Going to try this out .
2
64