@27upon2

post-training research @chakra_ai prev harvard

NYC
Joined July 2016
Introducing Gemini Cursor ✨ – a second multimodal AI cursor for your desktop that's open-source and free! Link below 👇 This experiment 🧪 reimagines how we interact with our computers because visual cues 👀 help us make sense of what we see on a screen. In this demo, I had my friend test it out by trying to add a payment method 💳 to Amazon. The cursor walks through the entire process 💬 while talking and pointing 🖱️ to the right parts of the website. Powered by Gemini 2.0 Flash (Experimental)⚡ from @Google and their live multimodal API. Shoutout to @alexanderchen for sharing the starter code that powers most of this app 🙌🔥
🔥 @Google Gemini 2.0 Flash is crazy good at pointing. I was over engineering before but now I'm just gonna bet on model capabilities. This is a demo of an AI cursor explaining a diagram on @tldraw with just a prompt and an image. Streaming is also simple with @vercel AI SDK.
32
104
21
1,020
177,694
I think people are underestimating the possibilities with ultra fast inference
1
9
339
I don’t see OpenAI and Google/Deep Mind in here hmmm
Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together, we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. nvda.ws/4hcoq7m
5
310
Omg why is Gemini helping LLMs get misaligned 😭😭
This quoted post is unavailable.
2
123
Sriraam retweeted
rube goldberg machine as a science teacher
4
8
40
2,765
Omg Notepad++ but the time I spent customising Atom was crazy lol
Be honest: what was your first code editor?
4
293
I like this because it’s honest and doesn’t over claim.
Here’s to the wanting. Here’s to the dreaming. Here’s to everyone.
210
Ppl debating the best and fastest way to search emojis. It’s a great time to be alive.
2
6
324
Real life LLM as a Judge 😭😭😭
BANGALORE DISTRICT COURT POSTED THE CHATGPT PROMPT IN THEIR JUDGMENT 😭😭 SON😭😭😭 indiankanoon.org/doc/1955248…
1
9
475
Dumb question: why doesn’t OpenAI pause the rollout sandbox when P0 alerts are raised? Surely they have the infrastructure to resume sandbox states.
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
2
4
639
Reward functions should be piecewise linear functions with clear semantics of strengths and weaknesses associated with each piece. The same applies for benchmarks too.
1
7
459
After listening to Dwarkesh’s podcast I realized that at some point people started calling them “AIs” instead of “models” and it sounds weird for some reason
6
232
Omg is Muse booty gonna be the next Elmo
weird this is the graphic that the partnerships team told me we were going with @Box 🤝 @Muse
1
2
460
Hogwarts Sorting hat ceremony made by @claudeai Opus 5.5 using Blender
1
6
268
Claude Opus 5.5 made this 3D clay stop motion animation from scratch talking about itself. Even though the sandbox had network policy constraints it found an open source TTS model Kokoro on github and used it to talk. It's also really cool how it can make simple music with code and iterate on issues with talking in different languages with audio spectrogram analysis but still messed up Japanese.
2
1
1
8
582
Opus 5.5 doing audio spectrogram analysis on the video its generating
4
216
Formal verification for distributed systems would be so useful and just cool
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached. TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt. I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted. Is formal verification the future of coding (or at least, bug finding)?
3
1
22
2,647
My friend’s ChatGPT became bilingual where it talks in Korean when she prompts in English and vice versa
4
260
Hmm World models as RL envs might be an interesting research direction. Not just video but text based world models too for tool calls
I've been on a Blender kick with Opus 5.5. Its better 3D modeling and vision mean you can build an entire world from a single prompt. Historically accurate San Francisco Market street in 1906, pre-earthquake
2
14
1,050
Umm why was this benchmark released?
SWE-Bench Pro V2 is live. What’s new: 🧵
6
56
12,064
I hope Jev can help us build I don’t know machines instead of lying machines
9
438