@sudoingX

GPU/local LLM. more RAM and OSS... everywhere

Bangkok, Thailand
Joined August 2022
Pinned Tweet
this is what 12gb of vram builds in 2026, absolute magic > rtx 3060 12gb, #1 gpu on steam > bonsai 2 27b + mtp, 5.95 gb of weights > hermes agent, 5 hours, 328k tokens written > 8 js files, 2,368 lines, zero hand written code > 50 tok/s fresh, 22 tok/s average, 125k context watch the full video, 5 hours in 12 minutes of pure dance of a local ai model on rtx 3060 12gb vram, and stay till the end for the full gameplay. this entire game was built by bonsai2, a qwen 3.8 27b dense compressed to ternary, and this small model is punching way above its weight. it built multi file engineering work using hermes agent, sure it's not fast but perfect for overnights and routine work and the quality is insane, and context holding is another best one, it does not lose the thread. i ran PrismML bonsai 1 made from qwen 3.6 27b dense and i built things with it, but this time with the latest base model qwen 3.8 these results are insane, and because i loved building with it a lot i thought many more of you would run it because this gpu exists in almost every home. so i packaged mtp, doubled the speed from 26 tok/s to 50 tok/s, packed the prefill fix in and released it on huggingface, almost 4,000 downloads in 3 days. i'll leave a link below.
holy shit you won't believe what bonsai2 is capable of at just 12gb vram, results incoming
38
46
6
553
44,489
here's the order i'd go into local ai, anon: > 12gb: learn the stack on a compressed 27b, how a model gets served, how an agent runs overnight with features like hermes agent's /goal. you will see good results and see some light at 50 tok/s fresh. > 24gb: this is where you meet the king, none other than the one and only 27b dense q4. get 2x 3060 if you want hands-on tensor parallel, or one 24gb card if you want to scale correctly and efficiently. at 24gb i prototype about 30% of my daily work, that's where owning your thinking starts. > skip macs and the 5090 and the rtx 6000 pro for now, learn the stack small first and understand the difference. it's not a thing that happens overnight, it's a process and a lifestyle, you don't just start doing everything, get disappointed and then never realize how profound this is, that you can think and it stays on your metal. > dgx spark: go gentle, and if you don't want to learn pcie lanes, power distribution and bifurcation, the dgx spark path is a clean cuda box, plug in and go. you can stack more and with the connectx cable it spits well. i have 2 dgx sparks so i can say for sure scaling with dgx spark is the cleanest. sure you don't get super fast bandwidth for dense models, but moes like qwen 3.8 flash next, ling flash and stepfun 3.7 dance on a single dgx spark and go zoom. > but if you want to learn real infrastructure, from cable management to pcie lanes at x8 and x16, bifurcation, slimsas cables, backplanes, airflow and power distribution, that's what the dgx spark skips for you, and the price is bandwidth. the ultimate path is the rtx 6000 pro path if you dare touch fire, and the fire doesn't care about dense or moe, it just burns everything, which i only intend to. rejoice my fren
when you realize your 12gb gaming gpu runs a 27b ai model at 50 tok/s and does overnight agentic tasks
351
you have no idea how good qwen 3.8 flash next 180b is, and the part that matters most to me is that i own it, it runs on my desk and nobody can switch it off. i've been running the official fp8 on my 2x dgx spark for about 2 weeks, and i know a speed update has landed, nvfp4 and all, but my home wifi takes another week so i'm still on fp8 and honestly it doesn't matter. even if everything disappeared tomorrow, the cloud or every quant on huggingface, i don't think it would touch me, because i already have frontier level intelligence running on 256gb of unified memory on my desk right now. nobody changes the terms on me overnight or retires the model my workflow runs on, the weights sit on my own disk. that's what owning your cognition looks like, the thing you think with stays yours. once you know what i know, you'll see what i'm talking about.
qwen 3.8 flash next built this 3d octopus invaders game in about 2 hours on my 2x dgx sparks and i didn't even try hard. the model is soo tasteful you can feel it in responses. here are the data and specs: > qwen 3.8 flash next, 180b moe, official fp8 > 45 tok/s fresh, 35 tok/s mid-build with mtp > 2x dgx spark, 256 gb unified, tensor parallel > hermes agent, three.js and vite, a 3d voxel shooter it wrote every file and fixed what i called out. with these hardware tiers you just need to know what good looks like and say it plainly. this is the state of local ai right now, if you have the right hardware you won't even miss the frontier anymore. since i've been using this model i'm going to the cloud less and less. watch, this is the game played live, no speed up.
6
1
20
1,217
qwen 3.8 flash next built this 3d octopus invaders game in about 2 hours on my 2x dgx sparks and i didn't even try hard. the model is soo tasteful you can feel it in responses. here are the data and specs: > qwen 3.8 flash next, 180b moe, official fp8 > 45 tok/s fresh, 35 tok/s mid-build with mtp > 2x dgx spark, 256 gb unified, tensor parallel > hermes agent, three.js and vite, a 3d voxel shooter it wrote every file and fixed what i called out. with these hardware tiers you just need to know what good looks like and say it plainly. this is the state of local ai right now, if you have the right hardware you won't even miss the frontier anymore. since i've been using this model i'm going to the cloud less and less. watch, this is the game played live, no speed up.
if 12gb of vram is the entry point of the local ai quest, then 2x dgx spark with 256gb of unified ram or vram is the point where you don't miss the frontier anymore. now i am running qwen 3.8 flash next fp8 on my 2x dgx spark with vision and full context, it builds insane stuff you have no idea of and it's better at frontend design than any of the humans i've met. i know because i run this every day on my desk, run experiments and put it through things you could not imagine yet, and most of them i cannot post, so what i post is games and builds and stacks. so here is hermes agent running on telegram, working on the build i am working on to demonstrate.
1
1
9
2,468
this is not a joke, this is literally what happened this week, bonsai 2 on a 12gb 3060 built a whole game through hermes agent overnight, the full 5 hour build is below 🧵
when you realize your 12gb gaming gpu runs a 27b ai model at 50 tok/s and does overnight agentic tasks
9
5
1
166
13,938
for anyone with an rtx 30 or 40 series card, 3060, 3070, 3080, 3090, 4060, 4070, 4080, 4090, everything is open here, you can run 27b model today: github.com/sudoingX/bonsai2-… there's a serve script for every vram tier, 8gb at 64k context, 12gb at the full 262k window, 12gb with the mtp head that runs it at 50 tok/s, and 16gb with the vision tower on. the kernel patches and every benchmark sweep behind my numbers are in there too, so you can check my work instead of trusting it. one thing before you start, stock llama.cpp won't load this file, the repo shows you the fork to build.
2
7
1,156
and here is the model card, pull it now: huggingface.co/sudoingx/Tern… it's qwen 3.8 27b compressed to ternary by prismml, i grafted the mtp head back on and shipped a faster kernel, and together they doubled the speed from 26 tok/s to 50 tok/s. there's the full file with the head at 7 gb, a lean one at 6.3 gb, and a prebuilt bundle for linux on rtx 30 and 40 series, extract it, run the script for your card and you're up, no compiler needed. almost 4,000 downloads in its first 3 days, go pull it and tell me what your card does.
1
6
814
if 12gb of vram is the entry point of the local ai quest, then 2x dgx spark with 256gb of unified ram or vram is the point where you don't miss the frontier anymore. now i am running qwen 3.8 flash next fp8 on my 2x dgx spark with vision and full context, it builds insane stuff you have no idea of and it's better at frontend design than any of the humans i've met. i know because i run this every day on my desk, run experiments and put it through things you could not imagine yet, and most of them i cannot post, so what i post is games and builds and stacks. so here is hermes agent running on telegram, working on the build i am working on to demonstrate.
here is the video of qwen 3.8 flash next autonomously building a visually striking website, start to finish, 29 minutes at 9.7x speed. the stack: official fp8 on 2x dgx spark, tensor parallel over one cable, 256k context loaded, 45 tok/s sustained with mtp on, hermes agent driving it from my laptop. the input was one spec file. it came back with this and started a server on my tailnet so i could open it from the couch.
7
3
1
24
3,826
i don't see enough videos about hermes agent tho, if you are creating quality content with hermes put it in this community and i will help it reach people. i opened the hermes agent community on x back in march for setup help and configs, it's sitting at 9.9k members now. i've been using hermes agent since it had under 1k github stars and since then i run it on every local build, even the one you saw this week ran on hermes agent on a 12gb card, and that community is where people figure out exactly that kind of setup, local models, configs, tool calls, the bugs nobody has documented yet. if you want to run agents on your own hardware, come in and ask, someone in there has already hit your error. nitter.cf/i/communities/20362891…

2
2
24
1,769
what’s the biggest model you can run completely locally?
61
1
41
15,180
people asked how an agent drives tmux, it's two commands: - the orchestrator types into a session with tmux send-keys -t agent1 "fix the failing test" Enter - reads it back with tmux capture-pane -p -t agent1 -S -200 - and decides what to send next that's the whole loop, every agent in its own named session, the orchestrator reading panes and typing like you would, and the ssh can drop without killing anything.
if i could recommend the 3 most useful tools for any aspiring agentic dev, the ones that changed the way i work, ate my frustration and kept my privacy, the tools i use every day are: > 1. tmux. since the day i got in touch with tmux a few years ago it changed my world view, and now that we are in the agentic era it is more useful than ever. every agent runs in its own pane and a master orchestrator seeds and steers each tmux session itself, so everything stays detached yet connected. > 2. forgejo private git server. a private git server plays as the memory layer for each ai agent, and each agent pulls and pushes and merges its own memory. once you get this you will never worry about your memory and you own every byte of it. > 3. gpu. and i am not talking about an rtx 6000 pro or a dgx spark. once you have any 12 to 24gb of vram and start tasting local ai, your entire perspective will change on how you work and what you share with closed cloud ai. own these 3 as infrastructure and you will make it. miss it anon and ngmi.
5
1
42
3,519
when you realize your 12gb gaming gpu runs a 27b ai model at 50 tok/s and does overnight agentic tasks
65
41
11
1,649
86,360
a judge ordered openai to hand 20 million chatgpt conversations to the new york times lawyers, people type their health problems and company plans thinking it's private. anything you send to someone else's server can end up in someone else's hands, on your own gpu the only copy is yours. learn to own your cognition tools and have a safe place for your core ideas and alpha. this should be practiced as the first protection for your thinking. just like you're not okay having a surveillance camera in your bedroom, how can you be okay giving the door key to your mind? i have seen these ai companies claim other people's research and work as their own, and they see what's in your mind wide open, everything you type in that chat box is their training data you like it or not. wake up you people, this technology is revolutionary, don't get distracted by it, instead know the core and the stack and own every byte of it. start small and then grow gradually.
i think local ai is inevitable. it starts with businesses realizing they simply cannot send their clients’ data, internal documents, code, research or trade secrets to someone else’s servers. then institutions start realizing the same thing. they’ll want the intelligence, but they’ll also want control over where the data and compute live. once enough of those organizations start buying their own compute, building private inference infrastructure and running models locally, the hardware gets cheaper, the software gets better and the ecosystem matures. and then normal people hit the wave. today local ai feels like a nerd thing. eventually it’s just called ai.
4
10
2
62
4,450
i think local ai is inevitable. it starts with businesses realizing they simply cannot send their clients’ data, internal documents, code, research or trade secrets to someone else’s servers. then institutions start realizing the same thing. they’ll want the intelligence, but they’ll also want control over where the data and compute live. once enough of those organizations start buying their own compute, building private inference infrastructure and running models locally, the hardware gets cheaper, the software gets better and the ecosystem matures. and then normal people hit the wave. today local ai feels like a nerd thing. eventually it’s just called ai.
9
9
3
110
7,490
if your goal is to understand modern AI infrastructure, buying the machine that hides the hardware stack from you is certainly one strategy.
a mac studio is the most expensive way to avoid learning how ai actually works. if you're serious about learning the ai stack, buy the hardware the industry actually runs on. a mac runs models. a gpu teaches you how they run.
7
2
29
5,181
what’s something that feels impossible today but you think will be normal eventually? i’ll go first: local AI.
8
24
1,993
if i could recommend the 3 most useful tools for any aspiring agentic dev, the ones that changed the way i work, ate my frustration and kept my privacy, the tools i use every day are: > 1. tmux. since the day i got in touch with tmux a few years ago it changed my world view, and now that we are in the agentic era it is more useful than ever. every agent runs in its own pane and a master orchestrator seeds and steers each tmux session itself, so everything stays detached yet connected. > 2. forgejo private git server. a private git server plays as the memory layer for each ai agent, and each agent pulls and pushes and merges its own memory. once you get this you will never worry about your memory and you own every byte of it. > 3. gpu. and i am not talking about an rtx 6000 pro or a dgx spark. once you have any 12 to 24gb of vram and start tasting local ai, your entire perspective will change on how you work and what you share with closed cloud ai. own these 3 as infrastructure and you will make it. miss it anon and ngmi.
13
14
2
257
14,171
apple sold you the illusion of compute. NVIDIA sells you the compute.
a mac studio is the most expensive way to avoid learning how ai actually works. if you're serious about learning the ai stack, buy the hardware the industry actually runs on. a mac runs models. a gpu teaches you how they run.
20
2
3
85
14,327
what are you most optimistic about for the next 10 years?
11
1
5
2,362
dear rtx 3060, 3070, 3080 and every 8gb to 12gb rtx owner, your gpu can run 27b dense model all locally and cowork with you, a 12gb 3060 built this whole game and i want to see what your card does with the same prompt. give it a shot, it talks like an engineer and i am sure you will find it useful.
this is what 12gb of vram builds in 2026, absolute magic > rtx 3060 12gb, #1 gpu on steam > bonsai 2 27b + mtp, 5.95 gb of weights > hermes agent, 5 hours, 328k tokens written > 8 js files, 2,368 lines, zero hand written code > 50 tok/s fresh, 22 tok/s average, 125k context watch the full video, 5 hours in 12 minutes of pure dance of a local ai model on rtx 3060 12gb vram, and stay till the end for the full gameplay. this entire game was built by bonsai2, a qwen 3.8 27b dense compressed to ternary, and this small model is punching way above its weight. it built multi file engineering work using hermes agent, sure it's not fast but perfect for overnights and routine work and the quality is insane, and context holding is another best one, it does not lose the thread. i ran PrismML bonsai 1 made from qwen 3.6 27b dense and i built things with it, but this time with the latest base model qwen 3.8 these results are insane, and because i loved building with it a lot i thought many more of you would run it because this gpu exists in almost every home. so i packaged mtp, doubled the speed from 26 tok/s to 50 tok/s, packed the prefill fix in and released it on huggingface, almost 4,000 downloads in 3 days. i'll leave a link below.
5
3
50
4,047