@cels19xi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
will stay in trenches until i make it
the outback
Joined July 2022
- Tweets1.1K
- Following539
- Followers68
- Likes25.9K
cels retweeted
i have a game idea, we create a open world game, and we have a repository on github that is somehow public, meaning it merges any PRs, play at your risk, so anybody can make their own characters, change things in the game etc and update it real time
SWE’s: first time?
cels retweeted
someone somewhere is still struggling with sol and astra instead of plowing through tasks with opus 5.5... a minute of silence for them please
cels retweeted
You don’t have to be a genius to see that Opus 5.5 absolutely crushed GPT-6 Astra, and Grok Bot crushed Dots.
This month is make-or-break for OpenAI, because even though they have 1.2B users, if most of them are either casual free users or casual $20 subscribers, they won’t be able to keep things afloat for that long while relying mostly on API revenue.
cels retweeted
imagine being a math PhD student with $400k in student loans and seeing this today
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.
github.com/openai/math
Using codex was your first mistake
I regret letting Codex ever run with a /goal prompt or launch sub-agents.
It doesn't matter what model it is at this point. The Sol and Astra lineage are completely RL fried.
It doesn't matter how tight the prompt is or the agent instructions are. I've even had prompts and instructions externally reviewed and rewritten, but it doesn't change the behavior.
The model always goes off track, ignores scope, and finds a way to digress into scaffolding, certification machinery, evidence gathering or audits. Even with the literal words barring it from doing so are right there in the prompt, plan, or instruction file.
And it always ends in an absolute mess of junk files, slop code, and an account drained of limits.
Codex can only be used effectively through the tightest of handholding. Keep one-shotting whatever you want if it's working for you. But real work on an existing system is excruciating and nearly impossible.
They are shipping models they know don't follow directions and are nearly useless for real long-horizon work.
cels retweeted
We shipped four things that were deemed good to great and some math proofs, but the vote is clear and the community demands a reset. I did calibrate it and it *seems* that the game is rigged in reset's favor, but such are the rules at the moment.
Therefore ... the reset has been processed. Enjoy!
Roundup of Day 2/
2.1/ Approve for me (auto-review) is now included and does not use usage. Can be between 2-10% of plan when used. Also better for you.
2.2/ Simplified API for builders.
2.3/ Meeting notes integrated.
2.4/ Decisions API live for builders. Will use in the app to improve the experience.