@cels19x

will stay in trenches until i make it

the outback
Joined July 2022
With 5 ChatGPT subs I can barely use Astra. With 1 Claude plan I have practically infinite opus
1
18
i have a game idea, we create a open world game, and we have a repository on github that is somehow public, meaning it merges any PRs, play at your risk, so anybody can make their own characters, change things in the game etc and update it real time
7
1
17
731
SWE’s: first time?
People seem excited about this. Idk, I think this is horrible. There is nothing poetic about coldly releasing hundreds of solutions to problems that could have been inspiring stories (e.g., Andrew Wiles). Using AI is fine; turning maths into a race won with concentrated resources is disgusting.
16
cels retweeted
Claude CEO just texted me asking for $100 bro ran out of tokens before reaching AGI 😭
69
55
6
2,175
56,283
Says the person who refused to speak on jijang ping
“Professor” Jiang says that Nick Fuentes is a fed.
20
Once Hermes app comes out it’s over for all ai agents
4
Anthropic is quickly speedily outpacing OpenAI it’s starting to become cooked for openai
1
5
Who even uses dots?
Proof that Dots is literally a Codex thread and NOTHING special. Accidentally found Dot's backend thread and it's just a cloud codex thread. Nice engineering OpenAI.
7
Honestly atp computer use with codex sub is the only thing worth it. Everything else opus is better at
5
so true. using 100 million tokens w opus can get wayyy more done than 2 billion w astra
someone somewhere is still struggling with sol and astra instead of plowing through tasks with opus 5.5... a minute of silence for them please
11
someone somewhere is still struggling with sol and astra instead of plowing through tasks with opus 5.5... a minute of silence for them please
56
15
3
705
22,766
Can we declare Gemini 4 Argon dead before it’s even born?
33
10
2
427
20,762
OpenAI's streak of aura losses continue. Lock in! 😭
Replying to @claudeai
Haiku 5.5 is a significant step up over Haiku 4.5 across coding, computer use, and knowledge work.
59
39
3
2,501
63,619
Fable 5.5 is coming.
45
7
923
39,937
bro was using sonnet
JUST IN: 16-year-old has to be rescued from British Columbia’s “Widowmaker” mountain after reportedly letting Claude map out the hike, leading him off trail.
8
You don’t have to be a genius to see that Opus 5.5 absolutely crushed GPT-6 Astra, and Grok Bot crushed Dots. This month is make-or-break for OpenAI, because even though they have 1.2B users, if most of them are either casual free users or casual $20 subscribers, they won’t be able to keep things afloat for that long while relying mostly on API revenue.
48
10
1
558
28,408
imagine being a math PhD student with $400k in student loans and seeing this today
We’re releasing a broad range of new mathematical results produced by an internal frontier model. We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results. github.com/openai/math
159
222
17
5,923
229,001
Using codex was your first mistake
I regret letting Codex ever run with a /goal prompt or launch sub-agents. It doesn't matter what model it is at this point. The Sol and Astra lineage are completely RL fried. It doesn't matter how tight the prompt is or the agent instructions are. I've even had prompts and instructions externally reviewed and rewritten, but it doesn't change the behavior. The model always goes off track, ignores scope, and finds a way to digress into scaffolding, certification machinery, evidence gathering or audits. Even with the literal words barring it from doing so are right there in the prompt, plan, or instruction file. And it always ends in an absolute mess of junk files, slop code, and an account drained of limits. Codex can only be used effectively through the tightest of handholding. Keep one-shotting whatever you want if it's working for you. But real work on an existing system is excruciating and nearly impossible. They are shipping models they know don't follow directions and are nearly useless for real long-horizon work.
22
opus 5.5 is basically just a magic swe genie atp
1
18
cels retweeted
We shipped four things that were deemed good to great and some math proofs, but the vote is clear and the community demands a reset. I did calibrate it and it *seems* that the game is rigged in reset's favor, but such are the rules at the moment. Therefore ... the reset has been processed. Enjoy!
Roundup of Day 2/ 2.1/ Approve for me (auto-review) is now included and does not use usage. Can be between 2-10% of plan when used. Also better for you. 2.2/ Simplified API for builders. 2.3/ Meeting notes integrated. 2.4/ Decisions API live for builders. Will use in the app to improve the experience.
2,241
662
1,001
15,905
1,580,920