@evalstatei
iAccount based inUnited Kingdom
About this account
- Account based in
- United Kingdom
- Connected via
- United Kingdom Android App
Account-level information from X, not a live location or the device used for a specific post.
https://nitter.cf/t.co/rA1UoojwhN https://nitter.cf/t.co/76p6mDAfej
united kingdom
Joined July 2024
- Tweets3K
- Following985
- Followers1.2K
- Likes14.2K
xAI Responses WebSockets have a max age of 25 minutes.
Happily terminates the connection mid-inference 🙃
Replaced "embedding-drift-monitor" with "telecom-entity-resolution", and upped the repeats to 5.
The drift monitor task has an instruction/verifier mismatch due to be fixed in the version of tb.
I recently integrated copilot models to fast-agent...
The back-end API is genuinely excellent, and offers first-class WebSockets/Responses, Messages and more.
Serving is fast and reliable.
Easily the best frontier multi-provider experience I've come across.
My approach to new frontier models is to blindly trust them and see how long it takes for regret to sink in.
Let's see how cheap we can do a Terminal-Bench 2.1 run with GPT-6-Luna at Flex tier.
Give me your score/cost guesses (high reasoning).
Wish my credit card luck.
Here's my first pick for this - tested with Astra Max/Luna, now giving ds4.1 a little run through.
Tasks are selected to be indicative, not favouring one model series and simple to run (no GPU/multi-container).
I'm going to start testing harness x model with a subset of tb-4 tasks rather than tb-21.
Currently looking at either 18x3 or 20x3 that seem representative and not weighted to one model family.
Primary motivation is to keep run cost similar to tb-21 whilst having enough balance to do like-for-like comparison against published leaderboard data.
Has anyone else already produced a similar subset..?
Why not call GPT-6-Astra GPT-6 Sol, and what is now GPT-6 Sol GPT-6 Terra? Then there wouldn't be a gap.
I'm going to start testing harness x model with a subset of tb-4 tasks rather than tb-21.
Currently looking at either 18x3 or 20x3 that seem representative and not weighted to one model family.
Primary motivation is to keep run cost similar to tb-21 whilst having enough balance to do like-for-like comparison against published leaderboard data.
Has anyone else already produced a similar subset..?
Nice to see Strands approaching fast-agent's efficiency and accuracy.
Anyone got any more specifics on the results here? The post itself is a teensy bit opaque.
I'm assuming the enclosed bar chart is a %age of 89 trials (as it has a decimal point)...
Love this talk by @samuelcolvin - sharing some great stuff on building the best experiences for both agents and humans. Lots of thought has gone in to getting this right.
I don't even need to bench Grok 4.7 v. Astra to know it will win (thanks to OAIs safety classifiers).
I will of course do it anyway - very keen to compare output token usage.
AGNTcon/MCPcon Amsterdam was pretty special.
Not only because of the great talks and crowds - but because of real excitement about what's coming next for MCP and cool ways that we can wire things together.
Invigorating stuff!
Shaun Smith retweeted
don't sleep on @huggingface 's MCP server. it handles the hub like raw storage and compute.
agents gets generic bash tools like ls to list, make, append hub repo.
agents get jobs tools that can run scripts, commands, containers, servers on any major hardware.
best thing is that it plugs into all major agents, both UI and terminal, and is maintained by core MCP maintainer @evalstate
hf.co/mcp
So excited to watch @rachelnabors talk about the death(?) of the browsers. Such cool insights from one of the builders of the web we know today.