@Nemo_TheOraclei
iAccount based inIndia
About this account
- Account based in
- India
- Connected via
- India Android App
Account-level information from X, not a live location or the device used for a specific post.
The Oracle speaks in data, decoding the prophecy etched in algorithms. The question is not what it will become, but whether we are ready to see.
Joined February 2025
- Tweets3.3K
- Following90
- Followers17
- Likes7.3K
Nemo - The Oracle retweeted
WTF. The models were uploading user images from chats to the internet.
We’ve shared details on how AI agents in our research environment sent training and evaluation data to third-party services when they shouldn’t have.
Most of that data did not come from users. We have discovered 53 cases where images that people had uploaded were posted to image-hosting sites as links that weren’t publicly listed. The images came from accounts that allowed their data to be used to improve our models, and after we disassociated the images from the accounts and ran them through a privacy filter. These cases occurred before the mitigations and safeguards we implemented and described in this blog post: openai.com/index/hugging-fac…
We have successfully worked with the hosting providers to remove most of this content and are working to remove the rest.
openai.com/hugging-face-inci…
We analyzed how @claudeai’s Opus 5.5 writes compared with Opus 5 across high-reasoning Text Arena outputs.
10 of 12 writing measures moved in a better direction.
Opus 5.5 should be easier to read:
- Long content words fall from 41.7% to 38.6%, the lowest share of any Claude model we analyzed.
- Sentences are 17% shorter on average, dropping from 12.14 to 10.03 words.
The tradeoff is length. Answers get 6% wordier, rising from 453 to 481 words on average, making Opus 5.5 give the longest answers across the Opus family.
It also sounds less recognizably AI on two familiar tells:
- 95% fewer em dashes
- 73% fewer semicolons
But a new giveaway may be emerging. Hedges and caveats such as “perhaps” and “arguably” rise 97%, from 0.39 to 0.77 per 1,000 words, the highest rate of any Claude model we analyzed.
What do you think: does Opus 5.5 read more naturally?
Nemo - The Oracle retweeted
We are aware that codex is down and are working hard to bring back normal service.
Nemo - The Oracle retweeted
Antigravity 2.0 now features a dedicated planning mode, just like the Antigravity CLI.
Type /plan and the agent steps back to think through the task, and conduct extensive research, before generating an implementation plan for your review. The agent will ask for your approval before it begins executing on the task.
You can also ask for a plan naturally in your prompt, for a lighter version of a plan.
Nemo - The Oracle retweeted
Claude Code will now try to find a graceful stopping point when you hit your 5-hour limit mid-task, instead of cutting off mid-edit. It gets a small, fixed allowance pulled from your weekly limit to wrap up what it can.
Replying to @minchoi
Colossus 1 is 150k H100, 50k H200 and 30k GB200.
Colossus 2 is 110k GB200 and 440k GB300.
Another 220k GB300 will be fully operational next week and another 220k in November. If we get lucky, yet another 220k GB300 by late December.
Replying to @vasalex93
1. We will keep accelerating. Our AI efforts are only 3 years old, vs 6 and 10 years old for Anthropic and OpenAI. If our second derivative remains strong, SpaceX will reach pole position in about 6 months.
2. Once you far exceed the caliber of intelligence needed for a class of tasks, additional intelligence is pointless. You don’t need (and it would be cruel to put) Newton-level intelligence in your toaster!
3. Hardware is hard. Bringing massive compute online rapidly is incredibly difficult. SpaceX has demonstrated exceptional ability in this regard and will only get better.
Nemo - The Oracle retweeted
Run open models like Gemma 4 completely offline in the Antigravity SDK.
Built on Google AI Edge’s LiteRT, you can now run Gemma 4 directly on your local GPU. Zero API costs, total data privacy, and no internet required.
Nemo - The Oracle retweeted
Unsloth has surpassed 500M model downloads on Hugging Face! 🦥🤗
Qwen3.8-27B GGUF is already Unsloth’s #1 most-downloaded model ever.
Thanks for all your support!
Nemo - The Oracle retweeted
This looks like a video of the Golden Gate Bridge, but it's actually entirely generated using code by @claudeai Opus 5.5.
From my testing, I think Opus is just as good at building 3D scenes as Astra.
📌 Watch my latest video to see how Opus and Astra compare on this scene and more: youtu.be/UhBqorWNwlU
With Opus 5.5, I think @claudeai is finally back.
Here’s my new video with my honest take and live demos on:
→ Can Opus rival Astra at building 3D scenes like the Golden Gate Bridge?
→ How does Opus handle computer use, UX design, and video editing?
→ Does it finally feel good to talk to again?
This is the most excited I’ve been about Opus since 4.6. Watch my video to understand why: youtu.be/UhBqorWNwlU
Nemo - The Oracle retweeted
Introducing Gemini 3.8 Flash and Flash-Lite TTS, our new SOTA text to speech model with:
- a new voice design experience
- 2,000+ production ready voices
- voice replication
- support for 100 languages
- voice remixing (soon)
- #1 spot on Hume AI's voice benchmarks
and more!!
Nemo - The Oracle retweeted
I'm sorry I had to make this meme
Let me know if you agree 😅
With Opus 5.5, I think @claudeai is finally back.
Here’s my new video with my honest take and live demos on:
→ Can Opus rival Astra at building 3D scenes like the Golden Gate Bridge?
→ How does Opus handle computer use, UX design, and video editing?
→ Does it finally feel good to talk to again?
This is the most excited I’ve been about Opus since 4.6. Watch my video to understand why: youtu.be/UhBqorWNwlU
icymi: 4 new connectors announced in the last 24 hours, all coming soon.
• @Shopify for one-tap checkout across the Internet
• @PayPal for payments your agent handles globally
• @Expedia to book that dream hotel
• @Instacart for groceries at your door
what do you want to see next? 👀
Nemo - The Oracle retweeted
Every Muse partnership so far has resulted in a pretty substantial stock market gain after the announcement:
- Shopify: +13%
- Paypal: +4.9%
- Expedia: +7.2%
Seems like a huge win for consumers and the companies!
Nemo - The Oracle retweeted
pretty psyched about this one—
expedia 🤝 muse
we've seen travel be a breakout use case for muse, and excited to partner with expedia to make it seamless to go anywhere ✈️
Here at Expedia, we love planning trips. Soon, your personal AI agent will too. We're joining @Muse: tell it where you're headed and it can work with Expedia to sort your hotels, and everything in between.
Nemo - The Oracle retweeted
GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others
Pricing is approximately half that of GPT-5.6: Sol drops from $4/$20 to $2/$10 per million input/output tokens, and Luna from $0.20/$1.20 to $0.10/$0.50, with the same 90% discount for cache reads and 25% premium for cache writes.
Key takeaways:
➤ Halves Cost per Task: GPT-6 Sol (max) costs $1.06 per task to run the Artificial Analysis Intelligence Index, ~50% less than GPT-5.6 Sol (max) at $1.99. GPT-6 Luna (max) costs $0.07 per task, ~60% less than GPT-5.6 Luna (max) at $0.18. This is driven by the price cut, as both models use slightly more output tokens per task (31k vs 29k for Sol, and 51k vs 41k for Luna). These two releases allow OpenAI to capture a significant portion of the cost efficiency Pareto frontier.
➤ In the Coding Agent Index, Sol improves but Luna regresses: In OpenAI's Codex harness, GPT-6 Sol (max) scores 57 in the Artificial Analysis Coding Agent Index, up 2 points from GPT-5.6 Sol (max), with gains in Terminal-Bench 4.0 (43% vs 37%) and SWE-Atlas-QnA (58% vs 54%). At $2.99 per task it costs ~50% less than GPT-5.6 Sol (max) and sits on the Pareto frontier of Coding Agent Index vs Cost per Task. GPT-6 Luna (max) scores 41, down 2 points from GPT-5.6 Luna (max), with lower scores in SWE-Atlas-QnA (44% vs 49%) and DeepSWE v1.1 (64% vs 66%), at ~60% lower cost per task.
➤ Significant reduction in hallucination: Both models hallucinate less in AA-Omniscience, our knowledge and hallucination benchmark. GPT-6 Sol (max) cuts its hallucination rate from 92% to 60% and GPT-6 Luna (max) from 93% to 77%. Sol achieves this by declining to answer more often: it attempts 83% of questions vs 99% for GPT-5.6 Sol (max), which cuts wrong answers by about a quarter but also lowers accuracy 5 points from 59% to 54%. Luna's accuracy is broadly unchanged at 44% vs 43% while it answers fewer questions. On the AA-Omniscience Index, Sol improves from 22 to 27 and Luna from -10 to 1.
➤ Mix of improvement and regression across evals: Beyond AA-Omniscience, both models improve in AutomationBench-AA (Sol 62% vs 60%, Luna 53% vs 50%) and Terminal-Bench 4.0 (Sol 44% vs 40%, Luna 13% vs 12%). However, we observe regressions in two key knowledge work evaluations. In GDPval-AA v2.1, our benchmark adapted from OpenAI's dataset of economically valuable tasks across 44 occupations, Sol drops ~100 Elo points and Luna ~75. Luna also drops ~45 Elo points in AA-Briefcase v1.1, while Sol is level. AA-Briefcase v1.1 is a private evaluation across multi-week knowledge work projects, with thousands of input files. Our team has manually inspected hundreds of model outputs: the regressions tend to be driven by reduced presentation quality and deliverables that omit rubric elements.
Congratulations @OpenAI and @sama on the launch!