@loganzevi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
building @thoughtfullab | hill climbing | prev @cornell
NYC
Joined August 2018
- Tweets582
- Following1.6K
- Followers1.1K
- Likes1.8K
- no mention of ads (spoiler: they are coming)
muse announcements so far
- free for users, but may eventually take a cut of transactions
- adding computer use
- muse email addresses coming soon
- walmart, best buy, gap, sephora, instacart, and more integrating with muse
- integrations with box, github, granola, notion
- 1500+ applications for connector platform (lovable, eleven labs, and more)
wonder if the real AI risk is that we slowly run out of genuine human knowledge and taste that hasn’t already been influenced by the models. what is the difference between human-authored and synthetic data when the human is using AI to do the thinking..
if our own sense of what’s good is increasingly shaped by AI then we’re judging the models against preferences they helped create
human taste / knowledge (unaffected by AI models) is more valuable than ever
Friday: Anthropic report explains that China is actually not that close with respect to the AI race without the help of distillation + model routing to Claude
Saturday: Dario publishes his "We Must Pace the Frontier"
Sunday: Trump says we must keep pushing the frontier to compete with China
Questions:
1. Is China actually competing with respect to the frontier?
2. Why are we not just forcing KYC when using the frontier models? (if they were serious about pacing the frontier this would happen)
3. In the history of the world, have we ever had companies asking for more regulation + intervention with an active administration essentially saying "no keep going"
Sooo
We're publishing our most detailed threat intelligence report to date.
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies.
These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve.
We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop.
Read the report: anthropic.com/threat-intelli…
Codex needs this plus an easier way to interact remotely cc @thsottiaux please sir
highly recommend reading the system card for these new model releases: just gives a sense for how much work is done for each release and by so many talented individuals. they are releasing new iPhone level updates at a remarkable cadence
System Card: Fable 5.1 and Mythos 5.1. They are the exact same model, with different safeguards.
anthropic.com/claude-fable-5…
front ends as diffusion models are coming
whole new category of evals to be built here
Interfaces that are generated, not coded.
Excited to finally share Solaris. Video is the universal interface. Eventually, all UI will be chat and gestures in, chat and video out.
Solaris opens new ways to build websites, apps and other interfaces. And most importantly, it gives new ways to train agents in much more dynamic environments.
age of post training is arriving and we will see a whole new set of products created to support this
If you lead AI for an enterprise, the highest leverage thing you can do is to make your AI stack model agnostic.
There are two things I would invest in.
1/ Today: build an eval suite that fully captures your use cases and business outcomes. Most companies I’ve seen do not do a good job of this.
2/ Within the next year: build the capability to post-train open models. The talent gap here is even more acute, so start planning today.
This gives you the flexibility to switch models, customize them, compare them on your own workloads, and optimize for the best combination of quality, cost and latency.
own the evals, own the models.
The first, the one and the only. Building the best product in the batch
@KupfermanAsaph
Bruh @garrytan @ycombinator can you please confirm or deny that @KupfermanAsaph was the first to do this during the batch
Instinct launch strategy of generating fomo is great (seems to be working at least from a timeline perspective) but for the uninitiated out there, they really should be collecting emails or creating a waitlist
Attention is short. If people are trying to get invited my take is you should give them a way to express interest