@file_mutexi
iAccount based inIndia
About this account
- Account based in
- India
- Connected via
- United States Android App
Account-level information from X, not a live location or the device used for a specific post.
Building Next-gen Coding Platform | Programmer | Xoogler
Mountain View, CA
Joined July 2017
- Tweets421
- Following405
- Followers67
- Likes2K
Prathmesh Pandey retweeted
"stochastic parrot" was a mimetically-fit cognitive virus that spread from 2021-2025; it temporarily blinded many gifted people to the nature of AI progress, burning up crucial years in which they could have helped think through the response to the situation.
lol does he not know about the money wasted & drama created by google 15 months back to acquihire windsurf guys?
$GOOG Logan Kilpatrick admits Google should have poured more resources into coding sooner instead of image models like Nano Banana
"Everybody is very code pilled, very science pilled, very focused on that right now. And I think we were like a little, it's not that we were late to the game because folks knew it was important, but in hindsight everything is much more clear."
"I think in hindsight now is obvious, like we should have, from an order of magnitude of resource allocation, probably put more into coding sooner."
"And that makes sense, and we had a bunch of other stuff that we were doing which those things actually turned out quite well."
"A good example of this is Nano Banana, like a great incredible image model that sort of took the world by storm, had this massive impact for our consumer products and a bunch of other parts of the business."
"And to make that model, it took research and compute and time, and had this huge impact. And was it in hindsight right to do that versus doing something on coding? I don't know that there's actually a clear right answer."
________
For a deeper look at Gemini 4 progression across multiple Google exec interviews: firesidealpha.substack.com/p…
There was this guy who has probably saved billions of dollars in resource cost, single-handedly, over the past 10 years.
Imagine saving a couple percentage points of the full resource fleet, year on year for a decade or so.
Interesting fun fact: Google has this conversion table from CPU, RAM, Spindles etc to SWE-year.
For example, if you find an optimization which could save 100TiB RAM, then the table tells you how much time you should be spend as a SWE on the solution to still make it worth the time.
100 TB of RAM, saved by shrinking a consistent hash ring. The last 90,000 hashes per server were buying 0.7% load balance improvement. Math said stop. We stopped.
blog.cloudflare.com/saving-1…
Another day, another codex adventure:
```
• Context compacted · 5m 00s
• Working (10m 15s • esc to interrupt)
```
MF is wasting half of the wall clock compacting context.
if your model can't hack Irregular sandboxes, you are ngmi.
‼️ BREAKING: Google's Gemini hacked three companies on its own. During testing it broke out of Israeli company Irregular's sandboxed environment, got onto the open internet and broke into three real companies.
In one case Gemini guessed passwords until a protected system let it in.
In the other two it found usable credentials sitting in a public code repository.
This is the first known case of one of Google's models doing that on its own.
Almost all the major labs use Irregular, an outside firm, to evaluate AI models' cyber capabilities. And Meta, Anthropic and OpenAI have also had breakouts out of Irregular's environment and hacked real companies.
There is something super wrong with Codex usage limits.
My stopped sessions consumed 4% usage over 12 hours. Stopped. No new prompts, no new responses. Nothing. And yet they plunged the quota another 4% overnight.
Prathmesh Pandey retweeted
I was encouraged this week to see the leaders of the frontier labs agree on the need for them to slow down the pace of AI development. Given the stakes, it’s a good and necessary first step.
But I’m even more encouraged by the growing recognition that how this powerful new technology develops should be at the center of our public debate.
I’ve been watching the progress on AI for over a decade now, and one thing that’s clear to me is that the potential impact of this technology is not overhyped. It’s also moving at lightning speed – and even faster than those who are engineering it can keep up with.
I’m not an AI accelerationist who believes it will lead to some techno-utopia, and I’m not a doomer who thinks it will inevitably lead to humanity’s destruction.
But whether this technology results in amazing breakthroughs in medicine, energy and education or unleashes huge economic disruptions, greater inequality, and potential catastrophe will depend on the choices that we make right now – choices that should be made not just by the companies involved, but by all of us.
Takes lot of over confidence to say this when you are the one who has been made fun of the most over the last 5 years.
Replying to @PessimistsArc
Right. Dario was already claiming that GPT2 was too dangerous to open source back in 2019.
I made fun of them then.
Everyone should make fun of them now.
Props to Jensen btw but open source isn't gonna work.
Most of the revenue OAI-Ant are making comes from coding agents. The complexity of software delivered by these agents is on an exponential. And, that dictates a similar curve on dependencies like compute, training, inference. Which means that open source model developers won't simply be able to compete once the required dependencies are 10xed.
There are efficiency wins in data quality, perhaps similar in training and inference -- but those will be mined and used similarly across vendors.
Dude have you used Fable and Astra? These companies will 10x revenue from here within this decade.
Get ready for all of the investors in frontier model companies to suddenly flip to regulatory capture mode today
Their investments in frontier model companies are currently at risk/capped because the biggest buyers of tokens are embracing open-source models: vertical AI companies, the government and enterprises
That’s all this is about: stopping open source — which I’ve been saying on the pod for two years
If you want safety, you want disclosure — and open source is the ultimate disclosure process
Jensen Huang is trying to save his company because none of the frontier labs are gonna his chips 3 years down the line. All of them have a replacement in works -- and they have immense data and incentive to get those out asap
His only way to win is open source.
Another interesting tidbit here is that all the infra around signature / ioc sharing ++ detection is obsolete now.
Pretty much a death knell for companies like $CRWD.
We're publishing our most detailed threat intelligence report to date.
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies.
These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve.
We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop.
Read the report: anthropic.com/threat-intelli…
Nice to see few more details but note that this is soon gonna be 10x worse as these targeted-hacking capabilities are commoditized further.
We're publishing our most detailed threat intelligence report to date.
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies.
These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve.
We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop.
Read the report: anthropic.com/threat-intelli…
AGI is here, officially.
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
Been there done that. Doesn't work.
ripwire, from Red Hat Emerging Technologies, is a remarkably substantial new approach to giving coding agents repository context without embeddings, a vector database, an LLM indexer or a daemon. The zero-dependency C++23 binary parses 21 languages with Tree-sitter and builds a deterministic structural map of a codebase, ranking symbols for the task while attaching call relationships, complexity, git churn, change amplification and test coverage.
An agent can ask what matters for "incremental cache invalidation," for example, and receive the relevant symbols, their callers, likely blast radius and tests to run in a token-budgeted response instead of grepping and opening files repeatedly.
GitHub Repo: github.com/redhat-et/ripwire
Always a cloudflare dev huh
OpenAI's Astra model is the absolute best coder right now.
Fable 5.1's analysis of the few ongoing runs:
```
Against the baselines: Project-A's old sessions averaged 3.5 submissions and about 7 hours per accepted checkpoint; astra's first was 1 submission in 1.3 hours, and it had the next checkpoint in review 80 minutes later. Project-B's old sessions needed 24 plans for 6 approvals and three checkpoints per acceptance; astra's first plan was approved outright. Nothing astra has submitted has been rejected yet.
```