@init_malachii
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
continually learning in a state of delight | ex sr member of technical staff | interested in ai epistemology | robotics infra & learning, ml sys | father of 2
synthesis
Joined May 2022
- Tweets25.1K
- Following4.4K
- Followers1.1K
- Likes314K
Pinned Tweet
i like this because i can deeply apply research in alignment with edifying incentives even as an entrepreneur
M retweeted
Wrote a short blog about the "shape" of language models, and the tradeoffs they may present in the future. I genuinely think it's a valuable research direction to start thinking about now, especially w.r.t. harness design.
alexzhang13.github.io/blog/2…
M retweeted
You're not "searching for meaning" lmao, you're searching for secure attachment/deep okayness
how do people serve dashboards? like is there a role-gated dashboard browser, book, filesystem? or do they just mysteriously float around
what would be the ideal way for this to happen
The economics of a Neolab.
A neolab is loosely defined as a startup of AI researchers who raises a lot of money pre-production to be able to finance GPU compute to take on a large AI problem.
To buy 1000 GB300s or ~14 NVL72 racks will set you back $125-150M for 3yrs with 15-30% upfront. That’s about ~2-2.5MW.
Thats about enough to do 10^25 flops a quarter and get to a GPT-4 level model which is 1-2 OOMs off frontier for pretraining.
If you post-train on a great open source model, you have a better chance of getting to frontier. The risks are a) you need to spend millions on RL environments too and b) being lapped by another model release while being tied to a base model.
For this to payback, you need to give your customers a better and ideally cheaper inference service than a base model and serve them for long enough to recoup your large investment. Even at 50% margin on inference, to recoup $10M in training means serving ~10T tokens (!) if you price like Fable / Astra given a standard cache read / input / output split ($2/M blended). And you have to justify being better than a release like Opus 5.5 which is even cheaper. Often, you end up charging your customers a huge premium in terms of platform fees and compute fees on top of pure inference.
Meanwhile, every hour you’re not utilizing your GPUs you are burning money so you typically resell this compute back to a broker or run inference for open models / resell spot instances. At below a ~60% utilization on spot, you will still lose money.
Add to that insane cost of talent.
So what can you do with the compute?
- Not play the model game at all.
- Play an entirely different model game (Jev, World Labs) that if big labs played, would either a) cannibalize their business or b) be incrementally not significant revenue c) would cause too much distraction from the main main thing
- Acquire a proprietary data set (Peridodic Labs) in enough volume in a domain of usefulness to eclipse frontier quality. Often happens in robotics, biology, chemistry.
If you do overcome the challenge of building a model that is useful and well priced beyond big labs models, given the huge price of compute, you still need to play in an area where the revenue / compute ratio is signficant and market demand is large enough to payback your compute spend.
It is a difficult game.
M retweeted
We didn't start the scaling.
Nine years of AI, from "Attention" to Opus 5.5, as an anime opening. It's one of six genres on the album, each with its own music video, and a mini-site to explain every reference.
Links below 👇
M retweeted
The seat / seatbelt in Tesla Semi is super cool. Produced in the same factory as the truck and with cool safety features!
I have a full factory tour video coming soon on Tesla Semi so stay tuned for a lot more info like this
M retweeted
Replying to @orrdavid
First off remember that there are two options if it's going out from IB - the payment for order flow (which IB may match internally or route to dealer/MM) or exchange. So as MM to get IB order you're already adverse selected on one hurdle (IB didn't match internal), rest of street could've seen it (and may have passed on it), and I'm now holding it. And, IB lot that id see as a dealer is probably someone like you - not HFT ping pong.
I need to look at 505/506 data - I can send you a chart pack when I dig it up - but my modeling of markups/PFoF used IB as the proxy/control versus dumb money retail, it's basically the opposite of robinhood flow all things considered, and partially because of their model.
As a MM, I'm going to have some buckets of predictable flow:
- Money manager type accounts who are one-way in a name; VWAP/TWAP orders, perhaps, or I'm talking to them and I know
- Option-driven, index-driven (which is the offset, here AP create-redeem more important)
- Much of what is thematic / the herd is one directional
- Pods, which all herd and I'm basically playing a momentum game against them and waiting for their blow ups to monetize, but which I can predict - if pod A is buyer at 11am, pod B will be at 11:15 (esp if they have any momentum-type pricing - usually it's clusters of them coming in off each other, it is very amusing in practice)
An IB order is the most likely to be informed, yes - because it has broad retail access AND funds can route through it.
If I'm at a large dealer, they're a broker. They're not a client. They have a sales rep, probably, but I'm going to treat them as an adversary, the brokers still pick you off. The client did not route to me - I do not get (necessarily) a promise of a future trade or relationship benefits - nor do I have a sales guy who has an idea of why you're executing.
Bursty random flow - even if "unintelligent" - can still be quite toxic, as well.
If one of my clients is doing a large trade - and large in market impact - and he says he's done, then he goes sells more on another venue, that's toxic/unpredictable, hurts my ability to hedge/distribute.
And a routing from a broker is max adverse selection because IB will just route across street as treated the same. So also scenario where if I'm not axed and getting hit by a broker, I'm going to widen out, usually you look into that & assume something went wrong, you were quoting poorly, there was some arb (again, broadly but maybe less the case in cash equities - brokers are actually the worse for dealers - because, here I'm talking less liquid markets/derivs/swaps experience, many of them pick you off when your quoting is off/look for arbs, because there's no trust relationship).
So if you were sending in your order, I may see this 50k lot from IB, but if it was my client, I may be able to have them TWAP/VWAP it and that is what makes it predictable. Basically intraday 15 minute reset TWAP/VWAP - distance from here is "is MM pissed off", to simplify.
I worked for one of the large dealers and sat by the algo sales traders on the equity side + covered some large etrading accounts on the treasury side.
M retweeted
One day it will be recognised that our mood is one way that we pollute the commons; that feelings are not private; that we have a responsibility to the commons & each other to cultivate joy & not be in a gigantic sulk at life all the time. its like defecating in the town water
M retweeted
It's becoming obvious that we are all spending a lot of time chatting with things that don't push back. Not good.
M retweeted
lots of categories emerging around the AGI stack:
- open model RLaaS
- data & evals
- inference
- coding harnesses
- GPUs
- agent tracing
- sandboxes
at @primeintellect, we agree. we do all of these things. people used to ask why we do so many things. they ask that less now.
M retweeted
"have a lot of fun and be high energy and control your cravings" are ideal end states but terrible advice
people who don't do these things don't *know* how to do them... if they did they'd be happy...
what i've found to be much more actionable and useful:
- learn to be ok with being misunderstood
- learn to enjoy boredom and the thrill of Facing the Void
- learn to embrace the freedom and grief of being alone
- develop emotional capacity
- take stimulants, in moderation
- give compliments generously
- celebrate tiny wins every day
- get good at asking for help
- use energy to get energy
most of it boils down to
1/ some form of shadow work,
2/ bottom-up nervous system regulation, and
3/ psyopping yourself into appreciating beautiful things, people, and experiences (authentically & in good faith)
"the amount of beautiful things in my life is directly proportional to my ability to see them"
M retweeted
if i had to steelman the general usage of OPSD, my recommendation would be:
- use judges to identify behaviors in rollouts which expressly violate guidance which is *already in context* (e.g. system prompt rules, tool schemas)
- insert a verbatim reminder of that guidance right before the failure occurs
- train only on the tokens immediately following the reminder
this avoids hint leakage, and combats the legitimate problem of rule adherence in long-context settings. however, you can also just do this with RL, and it's not clear why you should expect OPSD to be any better than folding those same failures into reward penalties. it's also training the model to take actions which might directly conflict with preceding reasoning, which could be an issue for CoT faithfulness if that's something you're concerned about.
tool schema failures are the clear win, but if your model is regularly messing up tool calls, there are probably more bigger problems in your post-training to solve. i think it's pretty unlikely that any serious frontier lab uses it in a real way other than as a last-minute band-aid on weird formatting bugs which show up in final testing states. it's not bitter lesson pilled.
some people like to think that *specialization* isn't bitter lesson pilled, and i very much disagree here. the most intelligent and successful humans are usually exceptionally specialized and spiky. somehow we convinced ourselves that being superhuman at *everything* ought to be a free lunch, but there's not really any historical precedent or theoretical argument for this being true in the limit.
the success of frontier foundation models isn't really evidence here, they're *way* more expensive to train than specialized counterparts, but final-run training costs are dwarfed by research and inference, and it's much easier to ship a one-size-fits-all product. jev is a hit because nobody ever made BERT finetuning easy enough for non-expert developers. real continual learning has never been tried. we don't even have the cognitive core nailed yet.
M retweeted
Apple’s GPU register file (marketing term = Dynamic Caching) is super nice for ray-tracing. Full cache hierarchy + RAM supporting it. Can keep occupancy much higher as waiting waves don’t block other work. Most people don’t understand how advanced GPU tech Apple has.
M retweeted
using philosophy to train LLMs: distinguish universal facts (true in all possible worlds) from contingent facts (true in this world), note that a model can self-generate universal facts, use this for pre-training… profit!
Can an LM, starting from random init (!!), learn to generate all of its pretraining data?
Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities.
A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
M retweeted
One easy thing to do is to look for frequency of multiordinal words, words that have varying definitions at different levels of abstraction. The more someone is using such language, the more they are exploiting the lack of type safety in your abstraction compiler.