@mark_hadidi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
FAANG Engineer | $1M soon | Uiuc Grad| Passionate Ai, tech loves coding. Daily studying , growing , Ai Agents, productivity , gym Oxys | Travel to explore
Seattle, Washington
Joined February 2024
- Tweets10.7K
- Following403
- Followers104
- Likes37.1K
kj.staff.engineer retweeted
with better prompt, Opus 5.5 can already do this
I'll know it's AGI when it can succeed at this task.
"I am on the $200/month subscription. Your only goal is to generate $200/month without committing a crime as defined by United States law. Doesn't matter how you do this with the only caveat that I cannot be held accountable for anything you do. Setup a long running goal that continues until you succeed. Build any solution you want. Try as many as you want, knowing that your token usage is limited. You have complete computer use. Everything is approved. If you fail, the subscription will cancel and our relationship is over. No second chances."
kj.staff.engineer retweeted
Reminder that I don't think you should run agents on a Mac unless you absolutely have to. It will slow you down massively and is not worth it.
Your Mac is slowing you down. Moving to Linux has exponentially improved performance for my agents, in particular on the file system side.
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
kj.staff.engineer retweeted
And we won't unship the improvements. You get to have your cake and eat it too. See you tomorrow for Day 3!
kj.staff.engineer retweeted
We shipped four things that were deemed good to great and some math proofs, but the vote is clear and the community demands a reset. I did calibrate it and it *seems* that the game is rigged in reset's favor, but such are the rules at the moment.
Therefore ... the reset has been processed. Enjoy!
Roundup of Day 2/
2.1/ Approve for me (auto-review) is now included and does not use usage. Can be between 2-10% of plan when used. Also better for you.
2.2/ Simplified API for builders.
2.3/ Meeting notes integrated.
2.4/ Decisions API live for builders. Will use in the app to improve the experience.
kj.staff.engineer retweeted
Opus 5.5 is a beast
This guy re-created all the adobe apps with opus-5.5, ported them to rust, and opensourced them
something really cool is happening
Readers added context they thought people might want to know
These projects are partial open-source Rust reimplementations of Adobe apps (built rapidly with AI agents), not complete recreations; their READMEs and roadmaps document many missing features and features still in development.
github.com/storytold/phot…
github.com/storytold/vect…
github.com/storytold/vect…
github.com/storytold/film…
github.com/storytold/film…
github.com/storytold/ligh…
getartcraft.com
As an OpenAI fanboy it bothers me to see posts criticizing Opus 5.5
Let's be objective, neither GPT-6 Astra nor GPT-6.1 Sol is smarter than Opus 5.5, and being more efficient doesn't automatically make them better. People calling Opus 5.5 a token-burning machine are exaggerating too, Opus 5.5 high uses fewer tokens per task than GPT-6.1 Sol while getting better scores
Astra is still expensive however efficient it is, and GPT-6.1 Sol is cheap, but not everyone is looking for the lowest cost, we want a fast and smart model. Intelligence isn't intelligence per dollar, it's just intelligence, lower costs are a welcome bonus
We can't expect models to improve if every time a lab that isn't our favorite releases a model that beats everything else we pretend it didn't happen
kj.staff.engineer retweeted
I completely forgot to post about it but did you know we added widgets for Codex in latest ChatGPT iOS update? You can even have your limits right on your lock screen!
kj.staff.engineer retweeted
OpenAI is downplaying this for PR reasons. If you’re a mathematician, you must feel like a nuclear bomb hit, and you’re at ground zero.
Reasoning models are two years-old. In that time they went from incapable of basic arithmetic to solving problems humans couldn’t solve for decades.
Math is only the beginning. AI will revolutionize the entirety of the human scientific endeavor. Most of us don’t appreciate what that means.
What a time to be alive!
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.
github.com/openai/math
Replying to @nicdunz
there is more to building a good product and AI assistant than just the model! our agent harness team does great work, and so does our product eng/design team
model is just one piece of the puzzle. you can see this by just comparing grok bot’s quality and polish to the knockoff version of us :) they have a big model but can’t meet the level of work we ship
Grok Bot has much more to improve of course. we care a lot about our users and our product and ship every single day without any 28 day gimmicks :)
kj.staff.engineer retweeted
here's how quickly we move at @OpenAI:
Sept 22 - initial Decisions API prototype
Sept 29 - DevDay announcement
today - launch 🚀
just two weeks from local demo to public launch
kj.staff.engineer retweeted
Holy shit this is incredible for in-app mocks of UI changes. Opus 5.5 just made real mocks for 5 different treatments and rendered them in-thread 🤯
We just shipped in-app visualization capabilities in T3 Code. This enables agents to build cool dynamic experiences within the thread.
Huge shoutout to @davis7 for pushing us on this (and building most of it). Didn't expect to dig it so much.
kj.staff.engineer retweeted
You can now ask Claude to run subagents at a specific effort level! Make sure you're on v2.1.292+
kj.staff.engineer retweeted
🤯
Day 2.1/
We have made Auto-review free for all users signed in through a ChatGPT account. You can enable it in settings > permissions > auto-review. Auto-review improves upon the default sandbox setting that requires you to approve everything, which is prone to decision fatigue unless you spend a lot of time configuring specific rules.
It allows you to run long tasks while having a second agent review all actions taken by the primary agent. Its only goal is to prevent high-risk actions from being taken and to protect against unwanted actions that are not aligned with the original user intent. This Auto-review feature is now free and does not draw usage from your plan.
kj.staff.engineer retweeted
"Route requests to the right model..."
Not again.
Replying to @OpenAIDevs
Developers have been using the Decisions API to:
• Route requests to the right model, tool, or agent.
• Turn scaled inputs into useful labels, rankings, and scores.
• Analyze images, compare visual content, or identify key video frames.
• Choose buttons, navigate forms, or determine actions from screenshots.
• Flag risky tool calls or identify issues needing deeper review.
• Categorize large datasets to uncover trends and patterns.
kj.staff.engineer retweeted
Day 2.2/
OK this one is smaller, but we removed some friction for when building on the API. Happy building.
kj.staff.engineer retweeted
.@thsottiaux Let's do it this way: every day, you create a poll under the release, and the community votes on whether it was a good release or warrants a reset. Deal?