@onthexitter69

AI & Pizza Enthusiast, Software Engineer

Joined July 2026
what exactly are they banning? we don’t have a frontier lab…
BREAKING: 40 British MPs have signed a letter to Prime Minister Andy Burnham calling for a ban on the development of superintelligent AI.
218
108
23
3,601
184,122
Changed direction. Instead of trying to predict multiple tokens into the future, predicting 8 layers into the future instead. Less room for improvement, but much more readily attainable on limited training capacity and data. Now at +16% TG throughput JIT Loads 4016 -> 1621 (-59.7%) Physical loads 4016 -> 2227 The new model is also applicable to PP, not yet wired up or benchmarked. We should expect greater relative throughput as we have the full prompt ahead of processing. We now have an 83% apparent waste overhead that training should be able to reduce to near-zero. Depth specialised heads and budgets could narrow budget where confidence is higher. Latent state prediction could bring load-level confidence a full 8 layers ahead of time, currently most d=8 predictions are hold confidence.
First notable net-positive results from speculative expert streaming experiment (0.8 hold confidence / 0.9 load confidence): JIT Loads: 4546 Baseline -> 3111 w/ speculative streaming (31.6% reduction) AOT loads: 239, 140-141 used (~59% acceptance) Combined load/hold route coverage: 61.23% TG Throughput +4.48% versus baseline Not much, but it's a positive start. I think tweaking the load confidence might yield slightly better results, but further training of the speculative router should yield several orders of magnitude higher TG throughput. I haven't even started training on the 4090 yet, and the M3 Ultra's training rounds are still lasting <1 minute each. Don't ask why I'm training on an M3 Ultra when I have an Ada card. I don't know why.
1
1
17
1,573
First notable net-positive results from speculative expert streaming experiment (0.8 hold confidence / 0.9 load confidence): JIT Loads: 4546 Baseline -> 3111 w/ speculative streaming (31.6% reduction) AOT loads: 239, 140-141 used (~59% acceptance) Combined load/hold route coverage: 61.23% TG Throughput +4.48% versus baseline Not much, but it's a positive start. I think tweaking the load confidence might yield slightly better results, but further training of the speculative router should yield several orders of magnitude higher TG throughput. I haven't even started training on the 4090 yet, and the M3 Ultra's training rounds are still lasting <1 minute each. Don't ask why I'm training on an M3 Ultra when I have an Ada card. I don't know why.
1
1
1
12
2,430
Fabian retweeted
running linux on a mac is the ultimate i can fix him relationship
1
106
Could someone with 128+GB memory on an M1-M4 preferably Ultra chip please test this out, tune and benchmark? It brings ANE offload support to Qwen4 architecture. I was able to achieve 33.8% speedup in GDN projection versus GPU alone. github.com/jundot/omlx/pull/… Loading Qwen3.8-flash on 96GB is just too unstable to benchmark.
10
6
1
82
6,051
I don't know who I was calling a masochist when I was benchmarking on a 0.7GiB/s USB 3 SSD. Expect significantly better, but still bad, performance.
For the low-memory masochists who keep asking for SSD expert streaming in oMLX: github.com/jundot/omlx/pull/… Supports pinning override (SoftREAP), although whether that's a useful feature is yet to be determined. Huge thanks to DwarfStar - a lot of the optimisations here are based on that project. Now, to turn this cursed feature off and never use it again...
3
1
1
7
2,121
For the low-memory masochists who keep asking for SSD expert streaming in oMLX: github.com/jundot/omlx/pull/… Supports pinning override (SoftREAP), although whether that's a useful feature is yet to be determined. Huge thanks to DwarfStar - a lot of the optimisations here are based on that project. Now, to turn this cursed feature off and never use it again...
7
4
1
38
4,455
(Not claiming credit for this, I’ve just added a tuner and guarding onto Jundot’s efforts to port the changes to Deepseek. I’m still experimenting with further offloads)
6
329
Probably going to be the last push till the weekend and it's quite practical... Tail padding. github.com/jundot/omlx/pull/… Extends some of the performance enhancements of all prior optimisations to the "tail block" - however many tokens over a multiple of 2048 the scheduler is processing. Configurable (tuned) threshold based on a profitability calculator. In practice, this means that agentic workloads will benefit substantially more from the PP changes on small incremental (uncached) prompts of ~1100 tokens or more, instead of beginning optimisation at 2048.
1
18
934
Also removes the need entirely for the "ANE-aligned prompt" checkmark in the benchmark tool - it will just process, then truncate, 1 extra token.
2
313
Very marginal improvement found by offloading attention to the ANE. Offloading GDN output had negative results but included in the PR as experimental anyway, I think in some lower-end chips it may yield a positive result. github.com/jundot/omlx/pull/…
7
16
1,083
M5 support is ready (in theory). Could someone with an M5 Pro or M5 Max test? Fused approach only supports Q4 at the moment.
5.1% relative speedup from yesterday's 45.8% isn't the bump I was looking for, but I'll take it. Non-M5 users, give it a try github.com/jundot/omlx/pull/… M5 support WIP
9
2
38
4,410
5.1% relative speedup from yesterday's 45.8% isn't the bump I was looking for, but I'll take it. Non-M5 users, give it a try github.com/jundot/omlx/pull/… M5 support WIP
Initially I'd discontinued efforts to offload down-projection to the ANE despite the significant performance gains to be found there (potentially 16.6%) due to accuracy issues. It seems those accuracy issues may now be caused by procedure index. Iterating now, another huge bump may be coming.
5
2
1
25
6,321
Initially I'd discontinued efforts to offload down-projection to the ANE despite the significant performance gains to be found there (potentially 16.6%) due to accuracy issues. It seems those accuracy issues may now be caused by procedure index. Iterating now, another huge bump may be coming.
4
1
1
42
3,252
Wrapping up on this, we're at a net 45.8% performance increase against GPU alone on M3 Ultra at a staggering 517.9 PP tok/s using just 3 CPU cores + ANE. Now featuring a new and improved tuner which takes just a couple of minutes instead of 10+ - which now balances 5 dimensions in a fraction of the time the old one balanced 2, at a much higher arbitrary precision level. Remaining work is to add options to the tuner, so users can opt-out of specific offloads if they don't want them, then will mark PR as ready for review. Feel free to test it out.
For anyone who's interested.... It turns out Apple's CPU has a few tricks up its sleeve too. +7.51% PP performance on top of the existing GPU/ANE split gains. Now at nearly 500PP on an M3 Ultra running 27b dense. Huge memory overhead though. github.com/jundot/omlx/pull/…
19
9
11
121
37,996
For anyone who's interested.... It turns out Apple's CPU has a few tricks up its sleeve too. +7.51% PP performance on top of the existing GPU/ANE split gains. Now at nearly 500PP on an M3 Ultra running 27b dense. Huge memory overhead though. github.com/jundot/omlx/pull/…
6
10
4
117
27,326
Now at 512PP and I should be able to get this up another 1-2%
7
1,388