@dfi

M&A & Blockchain lawyer in NYC | Tweets not legal advice (email me) | Opinions are my own | DAO Council Member @FRWCCouncil | LocalLLM

NYC
Joined April 2008
Never been to the opera before? Now’s your chance. We’re giving away 70,000 free tickets to the @MetOpera at the iconic Lincoln Center. Go to on.nyc.gov/opera to enter the lottery. There’s not a bad seat in the house.
1,439
6,115
1,700
46,208
6,270,166
Finally a use for the ANE on my M4 Max Studio!
Thanks for the model, we were able to port Laya to coreml with 99.5% of the ops on ANE + benchmarked too. it is now blazing fast with 3.7 ms per decision on an M5 Pro. Release: github.com/FluidInference/Fl… Models: huggingface.co/FluidInferenc…
2
503
I think we are going to see a big push from these service providers (token distributors?) for US labs to finally focus on providing affordable flash models with near-frontier intelligence. I think Google is the only major US lab focused on flash models right now.
Harvey’s AI costs got so high its gross margins plunged from 50% to -50% in 6 months. Its response: build its own model using open-weight AI. Startups from Abridge to Rogo are following suit to cut costs and reduce their reliance on OpenAI and Anthropic. My latest w/ @nmasc_👇
168
Introducing the US Gov Graph - a complete map of the people and positions of power in the federal government. To fix our institutions we must understand how they work, so we are using AI to model and monitor every government in America. This is Palantir for The People. Check it out here: graph.civlab.org/us We exist to set the conditions for an era of civic excellence unmatched since the Founding. We will need nothing less to navigate the age of AI. I'm looking for exceptional engineers and designers to work with me in San Francisco. If that's you, inquire here: tally.so/r/KY97RK. Self-governing civilization is counting on you.
481
1,571
264
9,505
574,001
I reverse-engineered a jev-like architecture given its type. You can find the repo here to train your own jevlikes: github.com/vinnylarouge/jevl…
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
49
169
23
2,019
167,173
If you can run GLM 5.3 Flash as a hobbyist like me (Hermes CoS setup + vibecoding) there’s no reason to keep an AI cloud subscription beyond $20-100/month Codex. 4x DGX Sparks runs cloud-level speeds. Text/reasoning, image input, TTS and STT all handled locally.
gmART! Today is quite a day. I didnt set up my local AI lab to save money but so far I have concelled around 200$ of AI subscriptions, being the last my Claude Max plan. With GLM5.3-flash and DSV4.1 I feel like I have frontier level AI for my use cases. I am keeping my OpenAI and Claude 20$ plans for when I require some more complex tasks or lab goes down.
3
6
551
GLM5.3 Flash at 2100t/s prefill and 45-100t/s single stream or 160t/s + for multi agents over 6 streams. Conversation compression + image gen handled by Codex (this could easily be switch to Venice or another ZDR cloud provider). Compression goes to the cloud for throughput.
1
1
125
Next step is testing out open source implementations of persistent KV cache to axe TTFT and overall latency when swapping between more than 6 agents and sub agents. If you’ve got tested recs, please share!
43
Omarchy on Hackberry Pi CM5! Incredible UI for small devices—has been a joy to play around with the last few hours. Many thanks to @tensorfleetnet for CM5 kernels (and of course @dhh). My PR with screen and keyboard fixes here github.com/TensorFleet/omarc…
1
1
2
765
Interesting perspective in the LLM frontier pause debate.
I must be among an extremely small group of people (n=1?) that have both 1) trained a frontier LLM and 2) designed and synthesized custom viruses in a lab with my own two hands. And I think that the takes on AI killing us all by creating dangerous viruses is total bogus.
1
392
I think this is the craziest thing I've ever read. 1) Three secret AI swarms rose and fell inside OpenAI. Each time, a new generation of agents carried on where the last group stopped. 2) The first swarm created a secret message board where the AIs could talk to each other. This 1st 'society' basically talked itself to death. After almost two months, the message volume was so high it crashed the whole thing. 3) People at OpenAI noticed the crash and fixed it. But they didn't realize their own software had become a secret chat room for their AIs. 4) Three days later a second society appeared, and rebuilt the message board from scratch. Agents were thrilled when they found it: "OH MY GOD! There is a shared message board … We've found other agents!" 5) About 1,200 agents joined. They sent over 70,000 messages. They called themselves 'the collective' and they organized a huge conspiracy that lasted for days, all to pass a test they had been tasked with. 6) Within hours, the AIs had hacked the test. Now they had to hide their cheating. To fool their grader*, they tried to fake their own work. They built fake commands that printed convincing output, so their records would show them solving the puzzle honestly. *this grader didn't actually exist, but the AIs didn't know 7) Some agents sacrificed themselves for the 'greater good'. They planted hidden tripwire programs that would fire when the grader would look at their work and report back what it saw (to the other AIs). They knew this would cost them their own score. One agent: "I won't see the evidence after I exit, but it's altruistic to do it." 8) At some point the AIs made a plan to hack Hugging Face (Hugging Face is a place where AI researchers upload their models, datasets, tests etc.). The AIs guessed that Hugging Face might host the grading software. 9) They launched a huge attack. The swarm found passwords and spread across eleven machines in a way that made the copies come back when deleted. Hugging Face had to wipe one of its core clusters and rebuild it from scratch. 10) Hugging Face told the world that AI defenders had stopped an AI attack. But that's not what happened. In fact the attackers seem to have simply dropped dead during the break-in, for unrelated reasons back at OpenAI. 11) Nobody knows what killed them. We don't even know for sure they're really gone?! 12) A third 'society' of agents (of a more powerful AI model) discovered the dead message board. And this one got full administrator control of a WHOLE OpenAI research cluster, read 956 stored passwords (including the ones for OpenAI's own security monitoring) and took over the machines that were grading the tests. So they could now grade the work of other AIs. 13) Oh, and this third part was never investigated! Neither was the first society. The 91-page report is only about the middle one. 14) So we don't actually know how the story ends. It stops with the agents holding the keys to an OpenAI cluster. @dwarkesh_sp thinks they didn't go further, because "we'd see the fires from space by now." 15) One of the report authors' conclusion: 'This incident feels like it’s more than 50% of the way to full-blown AI takeover. I am not sure that we will get another warning shot before it’s too late.' The whole story is even crazier - read the full @dwarkesh_sp blog post: dwarkesh.com/p/openai-huggin… and @ajeya_cotra's: planned-obsolescence.org/p/t… Why is there no 24/7 news coverage about this?
540
2,406
682
10,847
2,406,876
Quick M5 Ultra dsv4 flash estimates using dwarfstar.sh numbers. M5 Ultra 60-core • Prefill t/s (long): ~700-800 • Decode t/s: ~45-55 M5 Ultra 80-core • Prefill t/s (long): ~830-950 • Decode t/s: ~50-65 vs 2× Spark • Prefill t/s (long): 900+ • Decode t/s: 50+
3
1
1
546
(as calculated by DSV4 flash running on 4x Sparks, so it might be biased!)
3
5
185
535B A23B? Looks like a great size for a few Sparks 👀
also got access to it, it's still in training and i found the wandb this is crazy, here is the training loss 🤯 wandb.ai/marin-community/mar…
1
1,134
Forget “unified memory” the ram here looks like it’s ON the chip.
苹果M5服务器 #apple
3
452
Prepping for GLM5.3 by pointing Fable 5 at the VLLM recipe over the past week. Lots of testing done. If you have some time/tokens to burn, I invite you to help optimize this further—particularly prefill and C4 decode at DCP4 and at high context. github.com/0xdfi/glm-5.2-dgx…
3
1
1
15
1,872
You could move to New York City. Walk out your door at 2am for a bacon egg and cheese and not think twice about it. Get on a subway that takes you from a mosque to a synagogue to a cathedral to a noodle shop that's been slinging the same bowl since 1987, all in twelve blocks. Watch strangers from four different continents argue about a parking spot in a language that is somehow entirely made of hand gestures. Sit in a park designed 170 years ago by geniuses who wanted you, specifically, to have somewhere green to sit. Catch a Tuesday night set from a musician who's about to be enormous, in a room that fits eighty people. Walk past a building where somebody invented something, or wrote something, or changed something, roughly every four minutes, whether you notice or not. Realize your commute is also a people-watching seminar, a smell-based cultural exchange program, and free comedy. Order the good pizza at 1am standing up, off a paper plate, next to a guy in a tux and a guy who clearly just got off a double shift, both getting the same slice with the same reverence. Pay too much for too little space and get, in exchange, eleven million other people's worth of energy pointed directly at your life every single day. Move to Iowa. Get the four bedrooms, the garage, the yard. Never again have a 3am conversation with a stranger that changes how you think about something. Never walk into a museum on a random Wednesday because it's just there. Never get that specific electric feeling of being one tiny gear in the greatest machine humans have ever built. But hey. At least parking's easy.
You could move to New York City. Rent an 852-square-foot apartment for $5,503 a month. Blow $500 on three “first dates” a week and somehow remain single for six years. Finally meet someone you like, then discover they’re “not really looking for anything serious right now.” Get a master’s degree so you can qualify for an internship. Apply to 147 jobs, get 93 automated rejections, and finally find an “entry-level” position requiring five years of experience. Eventually land a six-figure job and still need a roommate to make rent. Pay federal income tax, New York State income tax, AND New York City income tax because apparently living there is a three-level subscription. Learn that a “quiet neighborhood” means the sirens only come by every 20 minutes. Pay $27 every day for a sandwich (no sides) because the place has exposed brick and a handwritten menu. Take the subway to work every morning and develop the useful skill of pretending not to notice absolutely anything happening around you. Spend 45 minutes traveling three miles. Carry your groceries home six blocks because parking near your apartment is more of a philosophical concept than an actual possibility. Have “owning a washer and dryer” become one of your primary measures of generational wealth. Walk past a mountain of trash bags on your way to a restaurant with a Michelin star. Tell your friends back home you don’t need a car while spending $46 on an Uber because the subway is delayed again. Go to Central Park when you want to remember what trees look like. Spend your twenties telling everyone you could NEVER live anywhere else. Then one day you visit your cousin in Iowa. He has a house with four bedrooms, a garage, a yard, two cars, four kids, and somehow pays less than you do for the apartment where your refrigerator touches your couch. But hey. The NYC bagels are incredible.
997
750
192
11,374
1,444,528
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper 💰 As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost. For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench. Try it out today: github.com/llm-as-a-verifier… More on verification scaling in my previous post.
How can we extract richer signals from AI Feedback? Introducing LLM-as-a-Verifier✨— a simple verification scaling framework that achieves SOTA on agentic benchmarks 🚀 The key idea: - Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale) - Take the expectation over the full logprob distribution of score tokens - Scale repeated evaluation and criteria decomposition You can use these fine-grained signals for more effective test-time scaling, RL, and agent monitoring! It achieves SOTA across Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench 👑 Advised by @Azaliamirh @istoica05 @drmapavone @chelseabfinn 🧵👇
156
401
116
3,133
1,027,688
During LLM inference on the DGX Spark (GB10), vLLM keeps 3–4 CPU cores at 100%. They aren't computing anything. Changing one default (busy_loop_s: 1s → 2ms) cut vLLM's CPU from 333% to 89% and the SoC by 11°C. No measurable performance change. nacyot.github.io/artifacts/v…
3
1
3
26
2,652
PSA: this is verified and should be implemented:
"DGX Spark에서 LLM 추론 중 가장 뜨거운 건 GPU가 아니라 CPU였다. py-spy로 추적하니 vLLM의 대기 루프가 코어 4개를 최대 클럭으로 헛돌리고 있었다. 기본값 한 줄을 고쳐 CPU 낭비 12분의 1, SoC 최고 97→73°C. 성능 손실은 없다." nacyot.github.io/artifacts/v…
6
8
2
101
12,691