@recursiveagi
San Francisco, CA
Joined May 2012
Lately, I have been trying to beat Google. Ok not really xD I have been building a search engine over niche web data and it's been one of the most fun projects I've ever worked on. The idea of querying the internet exactly the way I want, at scale, has always fascinated me. Especially what I could do with it: look at 1,000 companies' job postings and find hiring trends using natural language, map competitive landscapes including the lesser known players, find alternative data for investing. Or just a personalized corner of the web, curated for me. It's been an interesting journey. Started out with 200 workers on localhost, rewrote the crawler from Python to Rust to get up to 1,600 workers in parallel. Moved it to the cloud. Hit challenges that made the whole Rust rewrite useless xD Then I woke up to an abuse report from a French website with the words "European Commission", "appeal", "court" in it. For a second I genuinely thought I was about to have legal proceedings against me. Phew. Nevertheless it's been the most fun project I've worked on. It's surprising how you can cover a lot of information starting from a modest number of seed URLs. I'll be sharing everything I learned in the open. Code will be open sourced soon and the search engine will be live too. If you're curious how to build a distributed web crawler that covers 10M+ pages under $200/month, do read my blogpost. Blogpost Link: lnkd.in/g7MsaXYA
9
4
2
133
38,552
Anthropic usage reset scheduled for today or tomorrow? Should have a prediction market for this
1
85
hrishikesh kamath retweeted
OpenAI serving open-source models through Baseten is a big deal. They are embracing open source. Until now, OpenAI's pitch was "use our models." Now it is "commit your annual AI budget to us". We will get you the best model for each job, whether it's ours, someone else's, or open source. So if you're a large bank looking to commit $100M in AI spend, you centralize with OpenAI. One throat to choke. GPT handles the credit memos. Open-source handles the KYC volume at a fraction of the cost. A post-trained / specialized model handles fraud. This changes OpenAI's position. It stops competing only on whose model is best this season and starts competing to own the enterprise AI budget. A layer above being just an AI lab. Also proves AI is positive sum, not zero sum between closed and open source!
12
19
10
269
32,337
hrishikesh kamath retweeted
Shower thought: Maybe PDFs should come with instructions for how they should be parsed — obviously, having the raw data is the best kind of instruction, but can be a pain to include correctly + there are ambiguities that instructions could easily resolve.
3
1
8
503
Told yea
“Pause AI development.” Translation: “We’re about to IPO, public markets hate giant negative PAT margins, and we’d like to cut R&D without giving up our competitive advantage.”
1
348
hrishikesh kamath retweeted
For Reference, Zhipu AI has -439% PAT margins and Minimax -2368% for FY2025. The two publicly listed pure play foundational model companies.
“Pause AI development.” Translation: “We’re about to IPO, public markets hate giant negative PAT margins, and we’d like to cut R&D without giving up our competitive advantage.”
1
1
393
hrishikesh kamath retweeted
Everyday intelligence as @noahrshinn put it is an interesting way to view the contrast of what you would use for a personal assistant vs. frontier models you would use at work
3
1
15
1,233
hrishikesh kamath retweeted
I am most afraid of us eating the productivity gains of agents by just becoming lazier.
594
151
136
5,781
442,479
General purpose auto routing is the biggest scam. It’s impossible to have a low latency model to know nuances of every single kind task to route it right. Big believer in domain specific routing, most applied ai companies are doing it right. Cursor being the earliest.
2
2
187
Damn Claude code sub is a luxury in India
gifted myself a claude sub for birthday :P
189
hrishikesh kamath retweeted
we’re probably the last humans who will ever be the smartest things alive everything after this is just us managing something smarter
11
2
2
61
2,198
$META makes like $20 per month per user for US users only. Not sure if this can sustain browser use actions regularly. On the brighter side I am sure a lot of sites will become more agent friendly!
134
Best example for an asset that doesn’t go up much even after great news: robinhood:0xd0601ce157db5bdc3162bbac2a2c8af5320d9eec, one that goes up even with slightest good news and not down too much: $META
This is a heuristic I learned early in my career that has saved me a good deal of money. Generally speaking: If you are short an asset and it experiences a stream of negative news that fails to push the asset lower, get out. If you are long an asset and it experiences a stream of positive news that fails to push the asset higher, get out.
199
hrishikesh kamath retweeted
i think everyone should build something at least once - an app, a newsletter, a tiny business, a youtube channel, anything creating something nobody is obligated to care about teaches you more about people than a surprising amount of formal education does
9
11
2
136
4,082
hrishikesh kamath retweeted
Losses are a part of life Missing out on the next big one is even worse than taking a loss. Best of the investors have a 60-65% hit rate. Many retail folks need to internalise this
Druckenmiller on Soros: “He’s the best loss taker I’ve ever seen. He doesn’t care whether he wins or loses a trade. If a trade doesn’t work, he’s confident enough about his ability to win on other trades that he can easily walk away from a loss.” This is where you want to be.
9
33
1
465
42,847
Yep it’s simple as that
This is a heuristic I learned early in my career that has saved me a good deal of money. Generally speaking: If you are short an asset and it experiences a stream of negative news that fails to push the asset lower, get out. If you are long an asset and it experiences a stream of positive news that fails to push the asset higher, get out.
348
hrishikesh kamath retweeted
Just spoke to one of the big data labeling businesses. Few interesting insights: - They predict the majority of their revenue will come from Fortune 1000 enterprises, not labs in a few years - They believe every company will want to own their intelligence, but owning intelligence does not necessarily mean using open source models - A company’s evals will become their main proprietary IP given the improvement in agent performance after properly setting up & running internal eval environments - Most enterprises haven’t graduated from coding agents and it’s largely due to not having the proper eval infrastructure to make non-Eng agents performant
84
43
21
895
225,558
hrishikesh kamath retweeted
Rumor mill saying there’s some pretty big tokenization news with institutional asset managers coming very very soon. The biggest of the big players are supposedly deepening their ties with issuers is what our sources are telling us.
98
80
21
1,011
115,360
SLMs are everywhere! Took me a Jev to realize.
This instance was caused by a hallucination (the model fabricated a proper noun) that was further amplified by its causal thinking trace. This was not a data breach, no user data was shared, and no user isolation boundary was violated. Our users place a lot of trust in Instinct to keep their data private and secure. We've done a lot of work to make sure that users’ data are truly partitioned, from isolated sandboxes to short-lived local credentials to identity-signed tool execution. Regarding hallucination prevention: We worked over the last 48 hours to build an active hallucination detection system that now scans and verifies every token that flows through the platform. This layer is powered by small models that are trained to predict and detect hallucinations caused by ungrounded claims, creative brainstorming in thinking traces, or rare random sampling errors. Our systems are adversarially trained and battle-tested by world-class agent exploitation security teams to ensure that they can detect very subtle and nuanced hallucination cases. This layer has the ability to steer or intercept Instinct from proceeding with the next thinking trace or tool call execution before it is generated or executed. In the coming weeks, we’ll share a deeper dive into how we’re building a proactive engine to protect Instinct from hallucinations. We take every opportunity to learn and create, and we’re excited by the performance and potential for this line of work. These models are going to continue to be hardened every day through continuous testing, sampling, and creative adversarial training.
293
hrishikesh kamath retweeted
Druckenmiller on Soros: “He’s the best loss taker I’ve ever seen. He doesn’t care whether he wins or loses a trade. If a trade doesn’t work, he’s confident enough about his ability to win on other trades that he can easily walk away from a loss.” This is where you want to be.
45
279
23
3,415
210,462