@gpuemi

cofounder/ceo @wafer_ai. math @uchicago

san francisco
Joined December 2015
emilio andere retweeted
you'll know more about how KV caching, batching, and GPU communication affect transformer inference latency than 95% of people if you fully understand this article this is only the sixth resource in the ai performance engineering repo btw. follow and save to keep up with the series
we launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the series. links in thread 🧵 part 6: Transformer Inference Arithmetic Carol Chen (kipply) builds an approximate inference cost model around the work each token requires, the bytes the GPU moves, and communication between GPUs. it's a useful starting point for investigating latency and throughput on H200s, B200s, and B300s. Chen's examples use A100 GPUs. the applications below combine that model with NVIDIA's Hopper and Blackwell documentation: - prefill and decode expose different amounts of token parallelism. prefill processes prompt tokens together, giving projection and feed-forward matrix multiplies more weight reuse. small-batch decode has less reuse and can spend more time loading weights than computing with them. profile each phase. - batching can raise arithmetic intensity in projection and feed-forward matrix multiplies: more computation for each byte of weights loaded. that can improve throughput, but larger batches can increase key-value (KV) cache requirements and the time between output tokens. tune batch size against the latency target, using representative prompt and output lengths. - memory capacity determines which serving configurations fit. budget for weights, KV cache, and runtime workspace. H200, B200, and B300 have different memory capacities per GPU. a larger memory budget can reduce the number of model shards needed for capacity or leave more room for cached tokens. test how that changes concurrency and latency. - KV-cache capacity and KV-cache bandwidth need separate accounting. with full attention, decoding reads the retained keys and values at each step. a long context can fit in memory while its reads dominate latency. grouped-query attention shares keys and values across query heads, so use the KV-head count in the memory estimate. - precision can change memory traffic and compute throughput. Hopper supports FP8; B200 and B300 support native NVFP4. include scale metadata in memory estimates, and check the engine's supported combinations of weight, activation, and KV-cache precision. confirm which kernels execute, and evaluate the quantized model on the tasks it must perform. - tensor parallelism splits weights and matrix operations across GPUs and requires communication to combine partial results. NVLink bandwidth, message size, and synchronization affect the tradeoff. increasing the shard count can reduce computation per GPU while making communication a larger share of latency. - kernel fusion can reduce redundant memory accesses. FlashAttention uses tiling and fusion to reduce attention's intermediate reads and writes. Hopper's Tensor Memory Accelerator lets a thread block continue other work while transfers run between global and shared memory. register and shared-memory use can limit the number of active thread blocks on a streaming multiprocessor. - check estimates against measured kernel durations and achieved memory bandwidth, then investigate communication and kernel-launch overhead. compare configurations on the same model and representative workload. record time to first token, inter-token latency, and output-token throughput. reject throughput gains that violate the latency target. Chen checks her estimates against FasterTransformer measurements. for a modern deployment, update the hardware and model assumptions, predict the limiting resource, change the relevant serving or kernel configuration, and check the result against the workload's latency target.
9
14
218
15,035
emilio andere retweeted
we launched the most comprehensive ai performance engineering repo in the world follow and save to keep up with the series. links in thread 🧵 part 6: Transformer Inference Arithmetic Carol Chen (kipply) builds an approximate inference cost model around the work each token requires, the bytes the GPU moves, and communication between GPUs. it's a useful starting point for investigating latency and throughput on H200s, B200s, and B300s. Chen's examples use A100 GPUs. the applications below combine that model with NVIDIA's Hopper and Blackwell documentation: - prefill and decode expose different amounts of token parallelism. prefill processes prompt tokens together, giving projection and feed-forward matrix multiplies more weight reuse. small-batch decode has less reuse and can spend more time loading weights than computing with them. profile each phase. - batching can raise arithmetic intensity in projection and feed-forward matrix multiplies: more computation for each byte of weights loaded. that can improve throughput, but larger batches can increase key-value (KV) cache requirements and the time between output tokens. tune batch size against the latency target, using representative prompt and output lengths. - memory capacity determines which serving configurations fit. budget for weights, KV cache, and runtime workspace. H200, B200, and B300 have different memory capacities per GPU. a larger memory budget can reduce the number of model shards needed for capacity or leave more room for cached tokens. test how that changes concurrency and latency. - KV-cache capacity and KV-cache bandwidth need separate accounting. with full attention, decoding reads the retained keys and values at each step. a long context can fit in memory while its reads dominate latency. grouped-query attention shares keys and values across query heads, so use the KV-head count in the memory estimate. - precision can change memory traffic and compute throughput. Hopper supports FP8; B200 and B300 support native NVFP4. include scale metadata in memory estimates, and check the engine's supported combinations of weight, activation, and KV-cache precision. confirm which kernels execute, and evaluate the quantized model on the tasks it must perform. - tensor parallelism splits weights and matrix operations across GPUs and requires communication to combine partial results. NVLink bandwidth, message size, and synchronization affect the tradeoff. increasing the shard count can reduce computation per GPU while making communication a larger share of latency. - kernel fusion can reduce redundant memory accesses. FlashAttention uses tiling and fusion to reduce attention's intermediate reads and writes. Hopper's Tensor Memory Accelerator lets a thread block continue other work while transfers run between global and shared memory. register and shared-memory use can limit the number of active thread blocks on a streaming multiprocessor. - check estimates against measured kernel durations and achieved memory bandwidth, then investigate communication and kernel-launch overhead. compare configurations on the same model and representative workload. record time to first token, inter-token latency, and output-token throughput. reject throughput gains that violate the latency target. Chen checks her estimates against FasterTransformer measurements. for a modern deployment, update the hardware and model assumptions, predict the limiting resource, change the relevant serving or kernel configuration, and check the result against the workload's latency target.
16
11
3
108
18,329
wafer mention
I pulled the fastest growing startups on X by follower growth over last 90 days:
1
21
2,516
emilio andere retweeted
ts compute trading
selling fixed-price inference against floating GPU rental costs gives an operator exposure to compute prices. a five-year reservation fixes the rate, but also commits the operator to paying for capacity before demand is certain. @gpugene explores how compute derivatives could separate price protection from that capacity commitment. a call on the quarter's average rental index can cap effective rent before the option premium and financing, provided the average rental price matches the index and the hedge covers the same hours and period. the operator retains the choice of whether to rent, but still needs to secure physical capacity. Eugene works through the pricing assumptions, then uses a hypothetical simulation to compare annual rental renewals with and without calls. full piece in thread 🧵
4
7
39
3,954
emilio andere retweeted
selling fixed-price inference against floating GPU rental costs gives an operator exposure to compute prices. a five-year reservation fixes the rate, but also commits the operator to paying for capacity before demand is certain. @gpugene explores how compute derivatives could separate price protection from that capacity commitment. a call on the quarter's average rental index can cap effective rent before the option premium and financing, provided the average rental price matches the index and the hedge covers the same hours and period. the operator retains the choice of whether to rent, but still needs to secure physical capacity. Eugene works through the pricing assumptions, then uses a hypothetical simulation to compare annual rental renewals with and without calls. full piece in thread 🧵
11
6
2
51
8,089
please codex desktop app let me have multiple tabs open :(
7
1
24
1,721
emilio andere retweeted
It's hard to overstate how valuable @wafer_ai's stability has been. Previous providers caused latency spikes and eval aberrations. Wafer has eliminated a class of issues we used to lose sleep over. If something goes awry, they're the most responsive partner we've worked with.
This is an awesome customer story (I’m a proud @brilliantorg investor and a huge supporter of their mission). We heard the same themes over and over again when speaking to @wafer_ai customers.
4
4
21
3,042
emilio andere retweeted
realtime inference? wafer.
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC’s request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 🧵 link in thread
4
8
564
emilio andere retweeted
ultra fast inference? wafer.
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC’s request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 🧵 link in thread
3
5
15
1,230
emilio andere retweeted
ultra fast inference? wafer
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC’s request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 🧵 link in thread
7
6
60
4,893
realtime inference? wafer
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC’s request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 🧵 link in thread
8
5
1
40
9,987
emilio andere retweeted
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras. that's 44% lower latency with a much larger model. YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint. Wafer agents tuned the serving setup for YC’s request rate, cache usage, and prompt and response lengths. users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers. read how YC built the experience and landed on Wafer 🧵 link in thread
31
12
10
106
30,096
emilio andere retweeted
Ever since we moved our tutor to Wafer, the responses feel near-instant. We didn’t expect things to get this fast so soon. Thanks to @gpuemi and the @wafer_ai team! Love using a product built by a founder who grew up on Brilliant :)
i grew up in Mexico, where it wasn’t always easy to find really high-quality educational resources that pushed me intellectually. i knew pretty early that i wanted to eventually go to one of the best universities i could in the US, and i spent a lot of time trying to get sharper on my own. @brilliantorg was a very meaningful part of that. i still remember being in early high school and absolutely spamming their logic courses. i loved them. they genuinely made me better at thinking, especially around math and logic, and helped build a lot of the skills that eventually got me where i wanted to go. so there is something very surreal about @wafer_ai now serving Brilliant :) their AI tutor, Koji, needs to respond incredibly quickly while students are working through math and coding problems. before Wafer, Brilliant was speculatively prefetching AI-generated responses so students wouldn’t have to wait. on our dedicated GLM-5.2 endpoint, they’re now getting ~250ms time to first token and 300+ output tok/s, 3x the throughput of their previous provider! that let them remove the prefetching layer entirely and cut inference costs by 50%. it’s hard to describe how gratifying it is to help power a product that helped me become a better thinker when i was a kid. very full circle moment. 🧵 Brilliant’s full story in thread.
6
3
16
3,405
emilio andere retweeted
Replying to @gokulr @wafer_ai
👀
5
11
1,104
emilio andere retweeted
Replying to @gpuemi
It's been such a game-changer. We used to spend a lot of effort making the ~5s response times feel fast -- multiple layers of caching, loading animations, restructuring our prompts, etc. Switching to Wafer meant we could spend that time on the core parts of the learning experience instead!
5
11
453
emilio andere retweeted
A great learning experience has a sense of flow. Having a tutor take 5-6s to respond kills that. Switching to @wafer_ai for inference made Brilliant’s tutoring sessions feel fluid and fun again, and every session metric improved along with that.
i grew up in Mexico, where it wasn’t always easy to find really high-quality educational resources that pushed me intellectually. i knew pretty early that i wanted to eventually go to one of the best universities i could in the US, and i spent a lot of time trying to get sharper on my own. @brilliantorg was a very meaningful part of that. i still remember being in early high school and absolutely spamming their logic courses. i loved them. they genuinely made me better at thinking, especially around math and logic, and helped build a lot of the skills that eventually got me where i wanted to go. so there is something very surreal about @wafer_ai now serving Brilliant :) their AI tutor, Koji, needs to respond incredibly quickly while students are working through math and coding problems. before Wafer, Brilliant was speculatively prefetching AI-generated responses so students wouldn’t have to wait. on our dedicated GLM-5.2 endpoint, they’re now getting ~250ms time to first token and 300+ output tok/s, 3x the throughput of their previous provider! that let them remove the prefetching layer entirely and cut inference costs by 50%. it’s hard to describe how gratifying it is to help power a product that helped me become a better thinker when i was a kid. very full circle moment. 🧵 Brilliant’s full story in thread.
8
5
1
30
2,705
emilio andere retweeted
This is an awesome customer story (I’m a proud @brilliantorg investor and a huge supporter of their mission). We heard the same themes over and over again when speaking to @wafer_ai customers.
i grew up in Mexico, where it wasn’t always easy to find really high-quality educational resources that pushed me intellectually. i knew pretty early that i wanted to eventually go to one of the best universities i could in the US, and i spent a lot of time trying to get sharper on my own. @brilliantorg was a very meaningful part of that. i still remember being in early high school and absolutely spamming their logic courses. i loved them. they genuinely made me better at thinking, especially around math and logic, and helped build a lot of the skills that eventually got me where i wanted to go. so there is something very surreal about @wafer_ai now serving Brilliant :) their AI tutor, Koji, needs to respond incredibly quickly while students are working through math and coding problems. before Wafer, Brilliant was speculatively prefetching AI-generated responses so students wouldn’t have to wait. on our dedicated GLM-5.2 endpoint, they’re now getting ~250ms time to first token and 300+ output tok/s, 3x the throughput of their previous provider! that let them remove the prefetching layer entirely and cut inference costs by 50%. it’s hard to describe how gratifying it is to help power a product that helped me become a better thinker when i was a kid. very full circle moment. 🧵 Brilliant’s full story in thread.
2
7
2
25
12,120
i grew up in Mexico, where it wasn’t always easy to find really high-quality educational resources that pushed me intellectually. i knew pretty early that i wanted to eventually go to one of the best universities i could in the US, and i spent a lot of time trying to get sharper on my own. @brilliantorg was a very meaningful part of that. i still remember being in early high school and absolutely spamming their logic courses. i loved them. they genuinely made me better at thinking, especially around math and logic, and helped build a lot of the skills that eventually got me where i wanted to go. so there is something very surreal about @wafer_ai now serving Brilliant :) their AI tutor, Koji, needs to respond incredibly quickly while students are working through math and coding problems. before Wafer, Brilliant was speculatively prefetching AI-generated responses so students wouldn’t have to wait. on our dedicated GLM-5.2 endpoint, they’re now getting ~250ms time to first token and 300+ output tok/s, 3x the throughput of their previous provider! that let them remove the prefetching layer entirely and cut inference costs by 50%. it’s hard to describe how gratifying it is to help power a product that helped me become a better thinker when i was a kid. very full circle moment. 🧵 Brilliant’s full story in thread.
13
14
6
92
19,847
RT @gpusteve: @brilliantorg cut inference costs by 50% by making its AI tutor fast enough to answer in real time instead of pre-generating…
1
13