@HumanIntellig

AI is based idc.

Joined July 2014
Anthropic will have to keep this up for a while. OpenAi is valued 1.2 - 1.5 Trillion, Anthropic wants a 2T IPO in the next 1-2 months. They have to show they are ~30% more valuable than OpenAi and the beat way is to overshadow them at every step OpenAi will do. Let’s even look at what Anthropis learned from Fable releases, according to Openrouter, Fable did not increase Anthropic wallet share much and Opus was still the main model, what happened? Every release of Opus came after Fable and was better (generally) than Fable. Luna took a big token share and Astra took a big wallet share, the response? Haiku is back to take a punch at Luna, Opus 5.5 they fixed what 5 lacked and I am sure we will se that at the wallet share soon. Opus 5.5 being compared to Astra is the best thing for Anthropic since now it feels like Astra isn’t “Fable level” bust just Opus level and now they’ll need Bel or next Astra to he amazing. Astra - Opus Sol - Sonnet Luna - Haiku
anthropic is trying to steal openai's spotlight. look at the timing. opus 5.5 dropped the same day as gpt-6 sol and luna. and now partners and some users are testing a new sonnet 5.5 checkpoint, with release reportedly set for tuesday next week. the same day as openai's devday. devday has been on the calendar since april, anthropic knows exactly when it is. it worked with opus 5.5 as it took almost all of the attention from gpt-6 sol/luna launch.
1
3
255
Real question especially to those who don’t like the idea of OpenAi new subscription (500$ - 600$). If the next Sol will be at the same price of 6-Sol with similar Opus 5.5 capabilities and the new sub will give you 750t/s and more usage to where the ultra fast mode will affect your usage the same as normal on the 100$/200$. Would you be ok with this? You’ll effectively get the same usage as lower tiers but with ultra fast mode, this sub wouldn’t be for everyone but mainly for those who have 2+ 200$ subs or where speed can actually make a difference for them.
19
If true, that means that Sonnet 5.5 release would probably come on DevDay.
Partners got access to another Claude Sonnet 5.5 checkpoint, which is even better than the first one. Release is upcoming week :*
1
176
This is huge if works properly, if this can be scaled for larger models it seems like we will basically have a “swarm” in a sense in a single agent. Hopefully we will see this work and labs take this approach and we will see another jump in AI capabilities for research.
We achieved a breakthrough at BTL. AI models are very linear. They think in one direction at a time, and when they're wrong they scramble and start again. For something meant to replace humans, that's way too human. We changed that. We made LLM reasoning and execution work more like a quantum computer: many paths at once, the ones that meet merge into one, the dead ends cancel out, and everything left moves forward together. On the same hard problems, a 1.7B model thinking the normal way solved 3 out of 30. Interference Search(our new architecture)solved 23. We watched the normal model find the right answer at token 1,313, check it 9 more times, wander off and run out of budget without ever answering. Ours got there in 3 steps. Parallel thinking and execution. Subagents were a terrible way to tackle this. More technical details soon.
21
US Labs: we need to pace AI. Sakana: we are actively trying to reach RSI fuck off
Replying to @SakanaAILabs
We announced our RSI Lab earlier this year: sakana.ai/rsi-lab/ Over the last two years, we have systematically shipped the foundations for autonomous R&D: ▪ LLM²: AI automating research to invent new optimization algorithms. ▪ Darwin Gödel Machine: Agents rewriting their own codebase to double performance. ▪ ShinkaEvolve: Hyper-sample-efficient program evolution. ▪ ALE-Agent: Self-learning agents beating hundreds of human experts. ▪ Digital Red Queen: Open-ended adversarial coevolution. ▪ The AI Scientist: End-to-end automated research, published in Nature. Now we are unifying them into a single mission: open-ended, adaptive architectures that collectively self-improve. Human intelligence did not emerge from unlimited resources. It was forged through open-ended evolution under strict constraints. We believe the same principle applies to AI. Recursive self-improvement should not be confined to a hyperscale cluster, but should enable vastly more efficient AI systems. Under Jürgen's guidance, we are taking our foundation of shipped research, from the Darwin Gödel Machine to The AI Scientist, to the next level. We are building world models an agent can plan inside, and systems that design and run their own experiments. We are seeking a select group of highly driven Frontier Research Scientists and Advanced Core Engineers. If you have a proven track record at top labs but want to break away from standard benchmarking to discover fundamental new laws of machine intelligence, apply here: sakana.ai/careers/member-of-… Join us in Tokyo.
48
LEAKS ! GPT-69 Terra: > Better than GPT-5 Astra > Nice name > 0.1$/M in > 0.01$/B out > 1 tps > Closed source > Built by Anthropic > Mass produced in China > Probably won’t come out > Science fiction > Hi!
13
Wtf just happened?
Its so overr our team just coooookeddd so baddd
16
wtf happened lol, they added and removed these models twice in a span of ~3 minutes.
1
Since Haiku is going to be back, here is something I'm excited about it. Haiku is fast, if Anthropic are able to cut the price by half to 0.5$/2.5$ while keeping the speed and having a quality higher than Luna (especially for tasks we delegate to sub-agents) Haiku will be a killer. Compared to Luna, Haiku is able to perform ~40tps more. If we are able to utilize this 100-agent team strategies better with Opus 5.5, I can already see Haiku and Sonnet being a big part here.
13
Might be too soon to say but 6-Sol feels FAST!
If Haiku and Sonnet are back on the table with Haiku 5.5 at a close to Luna price, I am going back to Anthropic. Honestly, this last drop made OpenAi act similar to Google, building fast, cheap and mid models. A 1 - 2 jump in score for a whole new model while your competitor was able to make 10+ points jumps is CRAZY.
6
HumanIntellig retweeted
Hey OpenAI, I think you missed the memo Anthropic is at Opus 5.5 and Fable 5.1
27
9
3
774
39,330
Now waiting for OpenAi releases
Claude Opus 5.5 JUST DROPPED in Claude Code!! We are so back.
12
HumanIntellig retweeted
We might be getting a DOUBLE DROP today. Opus 5.5 is dropping. Now GPT 6 Sol, GPT 6 Luna, and GPT 6 Astra Minor just showed up in Microsoft's Azure config overnight. OpenAI demoted GPT 5.6 Sol to "workhorse" in the same commit. You do not do that unless the replacement is ready. We are finally about to get intelligent models we can actually use on our subscriptions. The biggest day in AI this year might be today.
152
120
99
3,284
599,432
This'll probably be the next goal for pricing for other companies as well.
yo MiMo said 200 a month? that's nothing, how about 400$ a month WOW
Grok 4.7: 500k context window 2$/M in 6$/M out docs.x.ai/developers/grok-4-…
29
Opus 5.5 is expected to cost 20% less than Opus 5. That means that the 17% usage decrease wouldn’t even be felt. We went from 150% usage to 125% usage, with 20% price reduction, the usage would be 120% more. 125% x 120% = 150%. Opus 5.5 usage should be the same as Opus 5 before the usage reduction.
Opus 5.5 Pricing: Input per Mtok: 4$ Output per Mtok: 20$ Cache Read: 0.2$ Cache Write: 5$
1
89
HumanIntellig retweeted
i propose we get a banked reset every day starting tuesday if we don't get anything from openai.
9
5
1
96
4,364
"Terra Max matches Astra at half the price. I'll just leave that here" AND WE WERE SHITTING ON TERRA!
Okay, woke up this morning and my Sol benchmark was finished, so I have now completed my full run, on the five core OpenAI models available today, across all effort levels. Through this process as I shared a bit last week, I also decided to create a third eval suite, focused on routine engineering tasks. This current suite, which I'm now calling VulcanBench Frontier, is really, most likely, harder tasks that regular engineers on engineering teams are giving models on a normal day. What I'm testing with this eval suite is how these models do with hard stuff, things you might give a model, but not likely on a daily basis. Which means I would look at these results as providing signal for what model/effort to use for your hardest tasks. Also as a quick reminder, or heads up for those just joining me here. With this v4 eval suite I also included a code quality scoring system that uses both Muse and Grok to score the code quality and include this as 33% of the score. I think code quality output for models is going to matter more and more over time, not because models need to write good code for humans, but because models need to write code other agents can understand and work on too. As for key insights I got from this, here's three: 1. Increasing effort level does not seem to impact code quality. So anyone thinking Low effort writes crappy code and Max writes beautiful code, that doesn't seem to be the case. Low and Max in pretty much every OpenAI model writes the same quality code. 2. Terra Max matches Astra at half the price. I'll just leave that here 👀 3. You never need to use GPT 5.5, and if you are, in any workflows, stop, you're wasting money. Comparison model cards below, and if you want to see detailed reports on each specific model run, you can find those here: vulcanbench.com/benchmarks.h…
14