@keennayi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
M.S. CS ML @GeorgiaTech | @SegaKatanaCom
InfiniBand
Joined April 2017
- Tweets4.6K
- Following1.2K
- Followers11.4K
- Likes65.7K
Pinned Tweet
Introducing Yoniq Compute
This repo provides a quick and simple way to stage a Linux node for ML workloads, including scripts to build & serve inference cookbook recipes across SGLang / vLLM
Recipes were validated on 1x, 2x, 4x, and 8x NVIDIA H200s, with future support for additional GPUs & models
Recipes include model orgs from:
Ai2, Arcee AI, Cohere, Datalab, DeepSeek, Dots Studio, Google, IBM, InclusionAI, Inco AI, Inferact, Intel, Liquid AI, Meta, Microsoft, MiniMax, Mistral AI, Moonshot AI, Nanbeige, Nex-AGI, NVIDIA, OpenAI, Paradigma, Prime Intellect, Poolside, Qwen, RadixArk, Red Hat AI, StepFun, Tencent, Thinking Machines Lab, Xiaomi, Z Lab, Zai, and Zyphra
Congrats on Reflection's Beam launch team!
8 months ago I joined Reflection to build a frontier western open weight model. Today we're releasing Beam (501B total, 23B active with an Apache 2.0 license).
I worked in many roles over this time, but spent a majority of the time here leading our data platform team.
We started from zero and built out all of the: pipelines, storage, inference, infrastructure, and tooling needed to get this model built.
Processing + storage. Beam trained on 23.8T tokens. I'm most proud of the OCR pipeline we put together generating trillions of tokens from hundreds of millions of PDFs. We scaled the data team from working with single terabytes to multiple petabytes.
None of this happens without the team. The talent across research and engineering here is the best I've worked with, and it's why we could ship this on a first release. Looking forward to getting this model and future releases into everyones hands soon!
Yannick Monye retweeted
The answer here is to use the industry standard: AI Perf.
Then there are two choices, raw AI Perf or AI Perf with Claude traces.
For work, I’m doing the first for now with eventually adding in the second.
And it’s what you’ll see me do for local benchmarks as well
“we as community really need to come up with a standard measurement for decode/prefill, and start questioning any numbers people put out including mine”
I’m a huge fan of Local Inference Lab’s LLM Inference Bench tool for measuring pre-fill / decode throughput data across a matrix of concurrency levels & context lengths
Would love to see a much wider adoption and unified process across the entire local AI space
It’s the best tool l’ve seen easily replicable across multiple hardware setups while supporting both the SGLang and vLLM inference engines
We really need this in LocalMaxxing
github.com/local-inference-l…
Yannick Monye retweeted
Will propose something. Give me Tuesday.
That lets us both speak the same language as those in the industry, and give us the full range of values that lets us understand local speeds
“we as community really need to come up with a standard measurement for decode/prefill, and start questioning any numbers people put out including mine”
I’m a huge fan of Local Inference Lab’s LLM Inference Bench tool for measuring pre-fill / decode throughput data across a matrix of concurrency levels & context lengths
Would love to see a much wider adoption and unified process across the entire local AI space
It’s the best tool l’ve seen easily replicable across multiple hardware setups while supporting both the SGLang and vLLM inference engines
We really need this in LocalMaxxing
github.com/local-inference-l…
Yannick Monye retweeted
Agreed. Really worth following the work of @YourLocalAILab at local-inference-lab.ai.
Great resources for Nvidia Blackwell hardware.
> Prebuilt and validated dockers.
> Quantized and tested models.
Quality that sets the standard.
“we as community really need to come up with a standard measurement for decode/prefill, and start questioning any numbers people put out including mine”
I’m a huge fan of Local Inference Lab’s LLM Inference Bench tool for measuring pre-fill / decode throughput data across a matrix of concurrency levels & context lengths
Would love to see a much wider adoption and unified process across the entire local AI space
It’s the best tool l’ve seen easily replicable across multiple hardware setups while supporting both the SGLang and vLLM inference engines
We really need this in LocalMaxxing
github.com/local-inference-l…
Yannick Monye retweeted
This would be amazing tbh, we all learn more/improve more with more legible measures.
Having something people are familiar with, and new people can be pointed to would be ideal
thank you @YourLocalAILab
“we as community really need to come up with a standard measurement for decode/prefill, and start questioning any numbers people put out including mine”
I’m a huge fan of Local Inference Lab’s LLM Inference Bench tool for measuring pre-fill / decode throughput data across a matrix of concurrency levels & context lengths
Would love to see a much wider adoption and unified process across the entire local AI space
It’s the best tool l’ve seen easily replicable across multiple hardware setups while supporting both the SGLang and vLLM inference engines
We really need this in LocalMaxxing
github.com/local-inference-l…
“we as community really need to come up with a standard measurement for decode/prefill, and start questioning any numbers people put out including mine”
I’m a huge fan of Local Inference Lab’s LLM Inference Bench tool for measuring pre-fill / decode throughput data across a matrix of concurrency levels & context lengths
Would love to see a much wider adoption and unified process across the entire local AI space
It’s the best tool l’ve seen easily replicable across multiple hardware setups while supporting both the SGLang and vLLM inference engines
We really need this in LocalMaxxing
github.com/local-inference-l…
V41 Flash 6000 TP4
539 tok/s code recorded, 518 tok/s median
+350core/+6000mem stable OC
Now that I got your attention, we as community really need to come up with a standard measurement for decode/prefill, and start questioning any numbers people put out including mine.
Too often I see a headline, actually check their work to find it’s inflated 20-30% on top of using extremely predictable prompt can add +100 tok/s with drafters. Most people do not have 4 6000s to confirm these numbers, so it simply goes unnoticed and it bothers me greatly that they are preying on communities trust to simple believe what is being shown
I’d like to think local AI is all about trust and we should always do our due diligence when we share/receive data in this community. I personally will always spend the extra effort to measure every single data I put on here, be as transparent as humanly possible, and freely admit if I’m ever wrong about anything without hesitation to gain your trust.
Even then you should always question my work and numbers, and demand nothing but absolute honesty out of everything you see in this community so people can’t get away with shady business.
Anyways, good morning and happy Sunday
@alexocheema, @ivanfioravanti, @Prince_Canuma, what are you all using for benchmark throughput data on the MLX side?
Let's not limit @YourLocalAILab's efforts to only NVIDIA / AMD, further fragmenting the local AI space
This Codex proxy setup for locally run models was quick & seamless, check out Zach's repo!
I did promise you local model inside of ChatGPT!
Point your LLM at this to have it configure itself
One half of the repo is a proxy needed (could just be one host somewhere), the other half is the per-machine changes needed to your ChatGPT config github.com/muellerzr/local-c…
This is how I discovered yearly power outages in Japan are virtually non-existent, never felt more fourth world in my life
Nice work James! We need more DGX Station GB300 recipes out in the open
Yannick Monye retweeted
Today we're unveiling Trillium Labs @trillium_labs, a new non-profit to foster the open science of frontier AI. We're building open post-training recipes and will expand into open infra to study RSI, reward-hacking, multi-agent systems, and whatever comes next.
We're built around the theory of change that you need more eyes to solve hard technical problems. We have faith in the scientific methods and communities that humanity has built, and worry that AI is becoming too closed to utilize them.
Trilliums are wildflowers that bloom briefly in the spring, before the forest canopies fill out. Though they are small, they lay the foundation for the cycles of growth and nourishment through the rest of the year. At Trillium Labs, the recipes will be the slow nutrients for the seasons and the model releases will be the blooms. Building an institution dedicated to this is needed because, much as nature’s trilliums are slow to expand and grow, the open-ecosystem needs time and dedicated resources to catch up.
I co-founded with with a long-time friend and collaborator Tom Zick (@thesezickbeats). We're hiring (full time + student collabs/interns), we're fundraising, and we're looking for compute. Please get in touch if you're interested in helping out. Offices based in the Bay Area and Cambridge MA, remote okay.
I’m in the Bay Area until for The Curve and COLM to connect with people who are interested. We’re thankful to have initial support from Halcyon Futures and Schmidt Sciences with more funding en route to enable our ambitions of scaling. Our advisors @Thom_Wolf, @HannaHajishirzi, @gneubig and @ctnzr have been instrumental to building the ecosystem that exists today, and I’m stoked to get to keep working with them.
My only request for Moonshot AI’s next Kimi model(s) is to release one <= the size of the previous K2.6 / K2.7 Coder models
At least those were able to squeeze onto 8x RTX Pro 6000s
Very tight yet still got the job done
Pretty cool Super Mario Bros 3 3D demo
Looking forward to the release & VR mode!
Super Mario Bros. 3 in 3D now has a create mode!
Pick any piece from the level itself (hills, bushes, pipes, blocks, clouds) & place it anywhere in the open world. Everything you build shows up in the classic view too, right where you put it, just like it was always part of the original game.
It's still the real game underneath. Mario can walk into anything you build, stand on it, or wander off the path in any direction.
Built live from the original game file you provide. No custom assets used here (every asset has been live-constructed from the game's actual sprites in memory at runtime). Create mode is being improved every day & also coming to Mario Galaxy.
#Flat2VR #VRGaming #MVRH
This media is unavailable