@27upon2i
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
post-training research @chakra_ai prev harvard
NYC
Joined July 2016
- Tweets2.5K
- Following4.5K
- Followers2.3K
- Likes12.6K
Pinned Tweet
Introducing Gemini Cursor ✨ – a second multimodal AI cursor for your desktop that's open-source and free! Link below 👇
This experiment 🧪 reimagines how we interact with our computers because visual cues 👀 help us make sense of what we see on a screen.
In this demo, I had my friend test it out by trying to add a payment method 💳 to Amazon. The cursor walks through the entire process 💬 while talking and pointing 🖱️ to the right parts of the website.
Powered by Gemini 2.0 Flash (Experimental)⚡ from @Google and their live multimodal API.
Shoutout to @alexanderchen for sharing the starter code that powers most of this app 🙌🔥
I don’t see OpenAI and Google/Deep Mind in here hmmm
Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry.
Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come.
But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility.
This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems.
Together, we are building the foundation of the AI economy.
Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. nvda.ws/4hcoq7m
Real life LLM as a Judge 😭😭😭
BANGALORE DISTRICT COURT POSTED THE CHATGPT PROMPT IN THEIR JUDGMENT 😭😭 SON😭😭😭 indiankanoon.org/doc/1955248…
Dumb question: why doesn’t OpenAI pause the rollout sandbox when P0 alerts are raised? Surely they have the infrastructure to resume sandbox states.
Some new misalignment disclosures from OpenAI:
• Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)
• In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks
• A new research finding, demonstrating that one can construct self-replicating prompt injections
alignment.openai.com/misalig…
Claude Opus 5.5 made this 3D clay stop motion animation from scratch talking about itself.
Even though the sandbox had network policy constraints it found an open source TTS model Kokoro on github and used it to talk.
It's also really cool how it can make simple music with code and iterate on issues with talking in different languages with audio spectrogram analysis but still messed up Japanese.
Formal verification for distributed systems would be so useful and just cool
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached.
TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt.
I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted.
Is formal verification the future of coding (or at least, bug finding)?
Hmm World models as RL envs might be an interesting research direction. Not just video but text based world models too for tool calls