@saen_devi
iAccount based inSouth Asia
About this account
- Account based in
- South Asia
- Connected via
- West Asia App Store
Account-level information from X, not a live location or the device used for a specific post.
Automating the boring stuff and sharing the insights building : LumaSleep , SpaceFlip Ai , Optify
UAE , Dubai
Joined October 2023
- Tweets43.7K
- Following276
- Followers2K
- Likes1.8K
Pinned Tweet
what do you think if you got 2k users in a week on your app without doing any sort of marketing even not x post
Grok 4.7 at the same price with longer RL runs and better self-verification. xAI is betting on training longer rather than training differently.
LAB: Grok 4.7 is xAI's new coding model
DeepSWE v1.1 scores 71.0%, up from 65.2% on Grok 4.6. CursorBench 4.0 lands at 46.3%, behind Fable 5.1 Max at 51.8% but ahead of GPT 5.6 Sol Max at 41.7%. Priced at $2/$6 per 1M tokens, same as 4.6. Rolling out in GitHub Copilot now.
What demos skip: xAI claims 38% on Terminal Bench, independent tests say 26%. Trust the benchmark you can run.
Developer tool winners are picked by the community, not the product team. Agent tools will be no different, and Firecrawl already understood that early.
Top developer tools companies were built democratically. I think top agent tools companies will be too.
AI agents are only as good as the context they can access. Firecrawl is becoming their default tool for this.
Alexandria gives them one interface to the live web, official data providers, and specialized indexes.
Proud to double down in this round and continue backing @firecrawl. Congrats to @CalebPeffer, @ericciarla, @nickscamara_, and the team on the $75M Series B!
@nexusvp
9.5 hours to port HAProxy from C to Rust. The conversation about what stays in the engineering job description just got more specific.
AI agents failing at recall rather than reasoning means context access is the actual bottleneck, not the model. The intelligence problem is mostly solved.
What if the biggest problem with AI agents isn’t intelligence?
What if it’s simply getting the right information?
Imagine asking an AI:
Find me a $150K+ job in this city and an apartment under $4K.
It can search the web.
But the information you need is scattered everywhere.
Jobs are on one platform.
Apartments are on another.
Financial data lives somewhere else.
Government records are somewhere else entirely.
The agent has to piece it all together.
That’s what Firecrawl is trying to change with Alexandria.
The idea is simple: give AI agents one place to access official data providers, specialized indexes, structured datasets and the live web.
Firecrawl says Alexandria already connects:
→ 88 official providers, registries & publishers
→ 504 capabilities across 28 categories
→ 113M+ indexed sources
In Firecrawl’s own internal evaluation, agents using Alexandria scored 21% higher on answer quality than agents using built in web tools.
But the bigger idea is what really caught my attention.
What happens when information isn’t just published for humans, but becomes something AI agents can discover, use and eventually pay for?
Firecrawl says it wants researchers, publishers, experts, creators and data providers to share in the value their information creates.
If that works, it could change how the web’s information economy works.
The internet was built for people to read.
Search engines helped us find it.
Now we’re building infrastructure that lets AI agents actually use it.
Maybe the next great library won’t be designed for humans.
Maybe it’ll be designed for machines.
Welcome to Alexandria.
Better models still need the right source material. Smarter reasoning on bad context just fails with more confidence.
Congrats to @firecrawl on the $75M Series B and the launch of Alexandria! AI agents using Alexandria scored 21% higher on answer quality than built-in web tools - better models still need the right source material. Glad to be on the cap table with @CalebPeffer at the helm of this great team and movement!
Writing quality has no benchmark row and silently regresses across model generations. Someone actually tracking it across evals is doing real work.
Dropping a model without a press tour and letting the benchmark speak is how you signal you are not competing for headlines anymore.
Anthropic just shipped Claude Opus 5.5 and didn't bother with a press tour - just dropped a benchmark table against everything else on the market
The shift here isn't a new modality or a new trick. It's the same model, still faster, still cheaper, now just further ahead on the numbers that actually predict whether an agent finishes the job
here's the breakdown:
1 - agentic coding, three ways → 66.4% on Terminal-Bench 4.0 (Astra 57.9%, Fable 5.1 55.8%, Sol 37.3%), 54.4% on FrontierCode v1.1, 57.8% on CursorBench 4.0. Not a one-benchmark fluke - it leads on every coding harness that got tested
2 - knowledge work → 1846 on GDPval-AA v2.1, ahead of Fable 5.1 (1735), Opus 5 (1708), and well clear of GPT-6 Astra (1542) and Sol (1588)
3 - reasoning → 67.7% on Humanity's Last Exam with tools, the top score on the table, edging out Fable 5.1 (65.6%) and beating Astra by 10.5 points (57.2%)
4 - computer use → 81.8% on OSWorld 2.0 (partial), ahead of Fable 5.1 at 80.7% - Astra didn't even report a number here
5 - the honest catch → it's not a clean sweep either. Astra actually wins two categories outright: AutomationBench business workflows (41.4% vs Opus 5.5's 40.0%) and Terminal-Bench-Science (64.6% vs 58.7%). Worth knowing before you pick a model by vibes
6 - the stack worth building → pair Opus 5.5 with Jev (TypeSafe AI's new "System One" model, out Sept 15). Jev doesn't write text - it takes unstructured input and returns a typed decision with a calibrated probability in 70-500ms, at a fraction of a cent per call. So Opus 5.5 does the actual thinking - the agentic coding, the long research runs, the judgment calls - and Jev sits inside the loop doing the thousand small classifications and routing decisions the agent doesn't need a full model turn for. System 2 for the hard parts, System 1 for everything that repeats
7 - why this matters → most teams are running one model for both jobs right now - reasoning through a plan AND deciding "is this ticket urgent, yes or no" with the same expensive call. That's the gap Jev is built for, and it's the gap Opus 5.5's own numbers make obvious: the model is good enough that burning it on trivial decisions is waste
Full benchmark table and the Jev docs are worth five minutes before you touch your stack
60% cheaper Luna output on top of a previous 80% cut in two months. The inference cost floor is not visible yet.
80% lower cost while nearly matching Fable 5 on long-horizon engineering tasks. That is the benchmark result that makes enterprise CTO conversations uncomfortable.
Replying to @OpenAIDevs
On DeepSWE v1.1, which tests coding agents on long-horizon engineering tasks, GPT-6 Sol (max) nearly matches Claude Fable 5 (xhigh) at ~80% lower cost per task.
GPT-6 Luna (max) is comparable to Fable 5 (medium) at 96% lower cost per task.
An evals skill that saves hours and prevents mistakes should be mandatory, not optional. No eval harness means production is your debugging environment.
Pro tip: Install this new evals skill from @HamelHusain and @sh_reya, it'll save you many hours and a lot of mistakes
github.com/ai-evals-course/e…
Per-task pricing is the metric token-per-dollar tables hide. Cost per finished job is what actually determines which model ships in production.
Sol getting close to Astra at 80% lower cost is the comparison OpenAI's pricing team wanted developers to make. And every developer will.
This quoted post is unavailable.
Funding your dev team before your first customer is terrible financial advice and exactly how you build something nobody else will bother to build.
This dude Tom who makes $50K/month from a website widget says: he funded his entire dev team before he had a single customer...
"Yet we had zero revenue. So on paper, terrible idea. Don't recommend it at all. But with that gave us the confidence to sell into these bigger companies that we kind of needed to get on board to pay us."
"I did run a marketing agency for 10 years, which still had like 70 plus clients paying like website hosting. So I used my like very precious income that I did have to live to basically fund the start of this."
Clinical expertise as a validation bottleneck is not a tooling problem. It is a trust-building problem that no agent framework ships a solution to.
Healthcare organizations are using agents to transform patient care across workflows. But one persistent bottleneck is that clinical expertise is a scarce resource, and these experts are typically ones that must vet the accuracy of these agents.
LangSmith is helping transform expert review into durable + reusable evaluations, compounding the value of human judgment.
Read on to see real-world examples of this loop in action.