@evalstate

https://nitter.cf/t.co/rA1UoojwhN https://nitter.cf/t.co/76p6mDAfej

united kingdom
Joined July 2024
xAI Responses WebSockets have a max age of 25 minutes. Happily terminates the connection mid-inference 🙃
5
265
Replaced "embedding-drift-monitor" with "telecom-entity-resolution", and upped the repeats to 5. The drift monitor task has an instruction/verifier mismatch due to be fixed in the version of tb.
Here's my first pick for this - tested with Astra Max/Luna, now giving ds4.1 a little run through. Tasks are selected to be indicative, not favouring one model series and simple to run (no GPU/multi-container).
290
Wow! Holy turnaround Anthropic, Opus 5.5 slaps.
Right, and now I'm basically bankrupt. Hits hard when you pay for this by the token from your own wallet. Not a had a session without this pathology yet, not enjoying it any more. Waiting for Opus 5.1 I guess.
1
11
618
Why is the ChatGPT app telling me to download the ChatGPT app?
10
13
598
🤫🤫 openai don't seem to force responses lite through the codex auth route any more🤫🤫
1
2
384
I recently integrated copilot models to fast-agent... The back-end API is genuinely excellent, and offers first-class WebSockets/Responses, Messages and more. Serving is fast and reliable. Easily the best frontier multi-provider experience I've come across.
3
6
287
My approach to new frontier models is to blindly trust them and see how long it takes for regret to sink in.
2
1
1
22
828
Let's see how cheap we can do a Terminal-Bench 2.1 run with GPT-6-Luna at Flex tier. Give me your score/cost guesses (high reasoning). Wish my credit card luck.
3
249
Here's my first pick for this - tested with Astra Max/Luna, now giving ds4.1 a little run through. Tasks are selected to be indicative, not favouring one model series and simple to run (no GPU/multi-container).
I'm going to start testing harness x model with a subset of tb-4 tasks rather than tb-21. Currently looking at either 18x3 or 20x3 that seem representative and not weighted to one model family. Primary motivation is to keep run cost similar to tb-21 whilst having enough balance to do like-for-like comparison against published leaderboard data. Has anyone else already produced a similar subset..?
5
1
10
908
Why not call GPT-6-Astra GPT-6 Sol, and what is now GPT-6 Sol GPT-6 Terra? Then there wouldn't be a gap.
3
6
451
I'm going to start testing harness x model with a subset of tb-4 tasks rather than tb-21. Currently looking at either 18x3 or 20x3 that seem representative and not weighted to one model family. Primary motivation is to keep run cost similar to tb-21 whilst having enough balance to do like-for-like comparison against published leaderboard data. Has anyone else already produced a similar subset..?
1
1
4
1,029
Nice to see Strands approaching fast-agent's efficiency and accuracy. Anyone got any more specifics on the results here? The post itself is a teensy bit opaque. I'm assuming the enclosed bar chart is a %age of 89 trials (as it has a decimal point)...
Today we're excited to share Strands Harness. Strands Harness makes it easier to build reliable, cost-effective, agents with any model. It's open source, too! And equal or better performance to proprietary harnesses with 28% fewer tokens!
6
362
Installed Linux (Omarchy) and now my webcam doesn't work. AMA.
6
10
426
Love this talk by @samuelcolvin - sharing some great stuff on building the best experiences for both agents and humans. Lots of thought has gone in to getting this right.
3
1
11
919
I don't even need to bench Grok 4.7 v. Astra to know it will win (thanks to OAIs safety classifiers). I will of course do it anyway - very keen to compare output token usage.
5
258
AGNTcon/MCPcon Amsterdam was pretty special. Not only because of the great talks and crowds - but because of real excitement about what's coming next for MCP and cool ways that we can wire things together. Invigorating stuff!
6
6
42
1,575
Night out.
2
7
358
don't sleep on @huggingface 's MCP server. it handles the hub like raw storage and compute. agents gets generic bash tools like ls to list, make, append hub repo. agents get jobs tools that can run scripts, commands, containers, servers on any major hardware. best thing is that it plugs into all major agents, both UI and terminal, and is maintained by core MCP maintainer @evalstate hf.co/mcp
11
5
36
2,013
So excited to watch @rachelnabors talk about the death(?) of the browsers. Such cool insights from one of the builders of the web we know today.
1
1
18
537