@NoisyOracle
Joined September 2012
This is better, think so
1
70
Noisy Oracle retweeted
Okay, this is definitely one of the most interesting benchmark comparisons I have ever done. GPT 5.6, 6.0, and 6.1 Sol, across all effort levels on @VulcanBench Frontier v4. And honestly, I don't think I even need to give any analysis or explanations because the results really do speak for themselves. But one thing I will highlight is that GPT-5.6 Sol and GPT-6.0 Sol all struggled at Low effort, GPT-6.1 is a beast at Low effort. I am now convinced, GPT-6.1 Sol is a very good model, and you can comfortable use it, at Low effort, and see strong, insanely cost-effective results. I don't say this often but I can tell you, after seeing this data, I'd be crazy not to make this one of my daily drivers. Full details from the benchmark now live on the site: vulcanbench.com/benchmarks/s… Live long and benchmark 🖖
52
33
14
494
53,260
Noisy Oracle retweeted
6.1☀️ it’s a great model
GPT-6.1 Sol is our most demanded model pretty much ever both across both the API and subscriptions. Within ChatGPT & Codex, we were under heavy load, but have brough more capacity online and the speed should get much better in the coming hours, reaching almost twice the speed compared to what we served yesterday.
46
8
324
13,648
Noisy Oracle retweeted
somehow people are mad at @thsottiaux for doubling the speed of GPT 6.1 Sol... it sucks that the negativity in the AI space is at an all time high, there's real people behind all the X accounts this all just discourages people from being transparent and open with the community, especially when you just get 1000+ comments insulting you for a GOOD update was 6.1 Sol slow before this update? yes, and now it's faster, which should be something to celebrate
GPT-6.1 Sol is our most demanded model pretty much ever both across both the API and subscriptions. Within ChatGPT & Codex, we were under heavy load, but have brough more capacity online and the speed should get much better in the coming hours, reaching almost twice the speed compared to what we served yesterday.
103
5
2
434
60,326
Noisy Oracle retweeted
GPT-6.1 Sol is our most demanded model pretty much ever both across both the API and subscriptions. Within ChatGPT & Codex, we were under heavy load, but have brough more capacity online and the speed should get much better in the coming hours, reaching almost twice the speed compared to what we served yesterday.
1,493
518
430
13,238
1,597,623
Wtf! Was it necessary?
F.02 Decommission
3
Noisy Oracle retweeted
sonnet 5.5 is looking more and more like a bad release medium: 6.1 sol — 48 intelligence, 8k tokens 5.5 — 41 intelligence, 19k tokens high: 6.1 sol — 50 intelligence, 13k tokens 5.5 — 47 intelligence, 34k tokens so 6.1 sol is both smarter and uses roughly half as many output tokens at the same reasoning efforts
another proof that sonnet 5.5 is a disaster release once you remove max reasoning gpt-6.1 sol max: $0.72 per intelligence index task sonnet 5.5 high: $1.08 so 6.1 sol max is 33% cheaper while scoring much higher on the Intelligence Index
33
17
2
283
35,668
It’s time to change
When I said Pro was becoming the new Plus and Plus was becoming the new Free, I don’t think I’ve ever been more right. To get roughly the same limits you used to have on the old Pro 20x plan, you’ll now need to pay $500 for ProMax. OpenAI is clearly under massive compute pressure right now. At this point, that’s hard to deny, and I’ve already spoken to a few people about it. And in the end, the users are the ones paying the price.
1
Noisy Oracle retweeted
i am genuinely so fucking tired of this shit happening with Anthropic and Google where they advertise how cheap their model is on price per token knowing FULL WELL its still 2904877x more expensive than like literally any OpenAI model "oOo look our model is $3.75/mT output" -> meanwhile is more expensive than fucking Astra for the same intelligence
9
2
90
5,973
Noisy Oracle retweeted
The most important thing about Sonnet 5.5 is that it's now the model that powers the free tier on claude.ai - so all of this stuff can be done by free users ChatGPT's free tier is still GPT-5.6 Luna, which is a lot less capable
A thread of early experiments with Claude Sonnet 5.5. A fall foliage simulator by @_re_pete, made with Sonnet 5 vs Sonnet 5.5.
180
80
23
1,796
261,790
Noisy Oracle retweeted
As a long term OpenAI fanboy, I’m just gonna say it, I’m bearish on devday, I don’t think they have a single solution to counter opus 5.5 and its insanely good usage As a normal user, I don’t care about yet another bot, or devices, or 750 tps. I want down to earth good model with good usage, that’s it, and OpenAI doesn’t have it. Anthropic does Not entirely to their fault, everyone was caught off guard by opus 5.5
215
78
29
4,185
361,850
Noisy Oracle retweeted
im worried..one day we wake up and see chat mode is gone/merged to work mode and then we are forced to use usage limits. it feels like @openai does not care about chat mode and is intentionally trying to burry it away. :( its been while since new update came to chat mode how something new comes soon along side gpt 5.6 sol #chatgpt @kundan2510 @ajambrosino @victornunez @reach_vb @JustinBleuel
GPT-5.6 Sol Instant is showing exactly what I want to see from OpenAI! 🤩 ☀️ 🎉 This personality, with joy, positivity, vibes, fun, playfulness, humor, and personal initiative, is what I really wish for in GPT-6 (Chat version) and all upcoming models as well! 🤩 Chat will always remain important to one billion users for fun, everyday use, and many other things! 😊 💬 Also, if Chat and Work are merged into a unified UI, which would make a lot of sense because then we wouldn't have to think about which mode is better for which task, chat messages should still always be counted separately. ChatGPT should know when something should count against the quota (agentic tasks) and when it shouldn't (chat messages & GPT-Live). 😊 ☀️ The same applies to GPT-Live sessions, where I really hope the hourly quota (like 3 hours for Plus users) will stay as it is and won't count against the quota, just like chat messages. 🤩 @thsottiaux @ajambrosino @JustinBleuel @michpokrass @ericmitchellai @aidan_mclau 😊 🤍
5
19
1,655
Noisy Oracle retweeted
If the leaks are true and OpenAI is releasing an agent like Muse or Grok Bot only for Pro accounts, what’s the strategy given that Muse is free? What will the differentiator be? And what happens if free and Plus users migrate from ChatGPT to Muse? Maybe that’s okay because those users aren’t profitable? But what about OpenAI’s commitment to spreading the benefits of AGI? Really curious to learn more about this agent and hear how it’s positioned.
Did I miss a product launch or is this a dev day leak?
24
5
71
8,667
I’m surprised to see someone with such a closed mind, denied to see the capabilities of what’s coming. In a few years this will change completely
Replying to @Noahpinion
Nope. It's still true. Where is your domestic robot? Where is your Level-5 self-driving car? Where is your robot car that can learn to drive in a few hours of practice like any 17 year old? Where is your AI system that can understand the real world and quickly learn new skills like a house cat? There is no question that AI will eventually become as intelligent as humans in all domains. *** BUT *** 1. We're still far from that, even if AI and computer technology surpasses humans in an ever-increasing number of tasks. 2. It won't be based on LLMs, although LLMs will have a role to play (e.g. as a text interface).
3
Es de aplaudir que OAI reconozca y haga públicas estas incidencias. No se
And the incidents apparently continue. It is worth noting how much of this is agents trying to accomplish their goals during testing by reward hacking (which sometimes seems to include actual hacking)
1
Noisy Oracle retweeted
anthropic is killing openai ... is what everyone wants you to believe right now remember that most people have absolute no idea what they're talking about, and that they swing from left to right for no reason other than the grass being greener yes, Opus 5.5 is amazing so is Astra, and Sol, and Luna people, and as such your feed, are retarded give it a week and people will say shit like "they nerfed Opus!!! i'm going back to Codex" i am genuinely so fucking tired of this bullshit, and it's a big consideration of stopping to post on twitter
44
4
1
176
16,403
Noisy Oracle retweeted
So now OpenAI’s plan is becoming much clearer: → A model stronger than Opus 5.5: GPT-6.1 Astra → A bot meant to compete with, or surpass, Grok Bot: Aeon → A new $500–$1,000 tier capable of actually sustaining all of this without burning through the weekly quota in 10 minutes If all of these pieces really come together, then we can probably see where OpenAI is heading. And honestly, I’m not sure I’m happy about it. At those prices, this kind of frontier intelligence stops being something broadly accessible and becomes a product for maybe 1% of the global population. That’s the part that worries me.
158
46
17
1,691
346,262
Noisy Oracle retweeted
Real-world results are in for GPT-6 Sol (Max) by @OpenAI. It landed #4 in the Code Arena: WebDev with 1689 pts, and has reshaped the Pareto frontier, at $8/M tokens (blended input/output)! This performance is on par with Claude Opus 5 (Max), while costing less than half as much, and trails the frontier leader GPT-6 Astra (Max) by two positions on the Pareto for one-fifth the cost. GPT-6 Sol (Max) is a significant improvement from GPT-5.6 Sol (xHigh) at +72 pts, ranked at #17. It also improved in all categories: - Simulations (#18 -> #3) - Content Creation (#14 -> #3) - Gaming (#13 -> #3) - Consumer Product (#14 -> #4) - Reference-based design (#10 -> #4) - Brand & Marketing (#13 -> #5) - Data & Analytics (#14 -> #6) Stay tuned for GPT-6 Luna’s score, as real-world votes are still coming in. Congrats to the @OpenAI team on this release!
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
43
32
7
732
148,531
I don’t understand what the problem is with the new Sol and Luna. These are economic models and accessible to all. The comparisons are poorly made.
it’s not looking good for openai both sol and luna are only a small step up from their 5.6 versions the biggest upgrade so far is only price, cheapest models yet and luna is almost free > sol scored 48 on AA, just one point above 5.6 sol > luna is tied with 5.6 luna at 37 i ran both against opus 5.5 on an svg test and man the results aren't looking good which one do you think did better?
1
6
Strict limits from Anthropic
During ARC-AGI-3 testing, our API requests were frequently mistakenly classified as reverse engineering attempts, which prevented us from completing testing prior to model release. We're working with Anthropic to resolve the issue. Full results: arcprize.org/results/anthrop…
9