@JacksonHPCi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
GPU Middle Class | Investor | Playlist King
Louisville, KY & Atlanta, GA
Joined October 2022
- Tweets3K
- Following1.7K
- Followers227
- Likes13K
Mark retweeted
We couldn't do what we do @typesafeai without our partners at @modal 🙏 thanks Modal team!
Mark retweeted
all neoclouds rent out gpus, and gpus are the most expensive part. it feels like the other parts like the land, power, and building itself would cost more as they feel so large but for a 1 GW data center, it’s said that the gpu servers cost almost 2x as much as everything else combined
neoclouds have different ways to go about this. a pretty stark difference is the nebius model and the coreweave model
- coreweave doesn't own many data centers. they rent space. their customers have paid ~15–25% of what they owe for the contract upfront
- nebius owns some data centers and rents space too. ~70% of the deals it signed recently included upfront payments. for large deals, these payments covered a much larger amount (50–60%) of what it needed to build clusters
this comes down to needing to borrow $. nebius doesn't have to borrow much money because their customers can fund most of it. and, since they own land and power, customers are probably more comfortable paying (even if they know the cluster hasn't been built yet).
neoclouds have different financing models, yet clearly both nebius and coreweave are doing great!!
(though if demand stays this high, neoclouds may have to rethink how they do financing)
Mark retweeted
Stoked to launch 𝚜𝚖𝚒𝚝𝚑𝚝𝚞𝚗𝚎 in public beta!
Post-training your agent's policy has never been easier, with training APIs from @baseten and @FireworksAI_HQ becoming increasingly accessible. But sourcing high-quality data to use them effectively remains a challenge.
LangSmith processes hundreds of millions of traces a day, making it a rich system of record for your agent’s behavior.
𝚜𝚖𝚒𝚝𝚑𝚝𝚞𝚗𝚎 turns those trajectories into the data you need for supervised fine-tuning, with an end-to-end pipeline for curation, preparation, training and deployment for models you want to SFT for specific tasks.
We care deeply about giving teams ways to own their own intelligence, and as the open model frontier continues to advance, it’s becoming increasingly clear that post-training on your own agent data will become a key part of that.
Excited to get this into people's hands!
Introducing LangSmith Fine-Tuning and the smithtune CLI.
LangSmith now handles the entire fine-tuning process. Use your traces to train specialized models that cut cost and latency.
Now in Public Beta. langchain.com/blog/langsmith…
Mark retweeted
Introducing typesafe/jev-router: a cache-aware model router powered by Jev and @typesafeai
The Jev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost.
Here's how it works 👇🏻
Open collaboration is moving GPU kernel development forward.
AMD and contributors from @Caltech, @Stanford, @togethercompute and the broader developer community teamed up to bring HipKittens support to AMD Instinct MI455X GPUs. The collaboration produced peak-performance GEMM kernels for the Helios GPU architecture using HipKittens.
The walkthrough starts with a baseline kernel and progresses through optimizations including asynchronous memory transfers, Tensor Data Movers (TDM) and workgroup-cluster multicast. Developers can follow the code and see how each technique shapes kernel performance.
Explore the technical breakdown and code: rocm.blogs.amd.com/software-…
512 Blackwell Ultra GPUs.
8 frontier models, 56B to 1T params.
BF16, FP8, NVFP4.
Every config met @NVIDIA Exemplar criteria on Crusoe Cloud: crusoe.ai/resources/blog/cru…
Mark retweeted
Want to RL-train a model on your own task without building the infra?
Use our new cookbook with @hud_evals:
→ Define your task + grader in HUD
→ Sample + train the model on Fireworks
Define the task once. Train and evaluate against the same environment.
docs.fireworks.ai/fine-tunin…
ultra fast inference?
wafer.
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours.
GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras.
that's 44% lower latency with a much larger model.
YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint.
Wafer agents tuned the serving setup for YC’s request rate, cache usage, and prompt and response lengths.
users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers.
read how YC built the experience and landed on Wafer
🧵 link in thread
Mark retweeted
Jev, now open source: Lev
A 4B open source System One model based on Qwen backbone
The best performance for it's small size
huggingface.co/interfaze-ai/…
Mark retweeted
We're announcing Elastic InfiniBand Partitions in SFC.
You can now get RDMA connected nodes reserved, change plans, and sell your whole contract on the fly. That’s not “sell to spot”, that’s “sell the whole contract”.
Launch the 30-day hero run on Monday, hit the snag on Tuesday, get out of the contract on Wednesday.
Running a PaaS and your customers need their own partitions? Create thousands of partitions, one for each customer. Lots of providers give IB, but all on the same partition, which means it’s not secure enough for production.
sfcompute.com/news/elastic-i…
Wafer beat Cerebras on latency for @ycombinator's AI Office Hours.
GLM-5.2 on Wafer averaged 379 ms versus 674 ms for Gemma 4 31B on Cerebras.
that's 44% lower latency with a much larger model.
YC wanted people to get startup advice from AI versions of its partners at conversational speed. after testing lightweight Gemma and OpenAI models, they moved to a dedicated Wafer endpoint.
Wafer agents tuned the serving setup for YC’s request rate, cache usage, and prompt and response lengths.
users spent 2.5 minutes longer talking to its AI partners on Wafer compared to other providers.
read how YC built the experience and landed on Wafer
🧵 link in thread
Mark retweeted
AI compute is not becoming a commodity.
The industry is building vertically integrated infrastructure, reaching from software and networking to the physical sites and power that make large-scale compute possible.
The stakes are physical.
Lambda Co-Founder & CTO Stephen Balaban makes that case on the Main Stage at Yotta 2026. He will also discuss our target of 3 GW of AI compute capacity by 2030.
A thousand hours of audio costs $7 to transcribe on the right GPU. $507 on the wrong one.
We benchmarked Whisper large-v3-turbo across 23 GPUs and the results aren't what most people expect.
Read the entire breakdown here:
runpod.io/articles/guides/be…
Mark retweeted
Last scoop of the night: AI chip startup DensityAI is raising hundreds of millions of dollars at a $10b valuation, on the back of a new AWS deal.
w/ @_pheebini @validapau
theinformation.com/articles/…
Mark retweeted
Stephen Balaban (@stephenbalaban) has been through 5 pivots in 14 years. He started with facial recognition and AI filters before building the workstation and GPU compute business.
The interview is a look at the long, non-linear path behind Lambda, and the decisions that shaped it along the way.
Listen to the full conversation with @LambdaAPI’s CTO below on the @FoundersInArms podcast with @immad and @rajatsuri.
foundersinarms.substack.com/…