@pandasmamasi
iAccount based inSingapore!
About this account
- Account based in
- Singapore
- Connected via
- Croatia App Store
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
hi. i like computers n create things i like to use. https://nitter.cf/t.co/EW6iYly9lM & https://nitter.cf/t.co/7adg8hl2W5
Joined May 2023
- Tweets15
- Following259
- Followers26
- Likes1.3K
introducing an open source alternative for this kind of agent workflow, with no required monthly subscription for the tools and session data you can export, inspect, and use however you want,
I have nothing against Cursor. I’ve used it for a long time.
but I am serious about avoiding lock-in.
I expect capable models to keep getting ridiculously cheap. I want to be able to move between providers, try new models, and choose what makes sense for the task.
what I don’t want is to spend months improving my workflow, then feel reluctant to try something else because so much of how I work is tied to one product.
that’s why I use herdr + pi.
herdr manages the terminals, workspaces, and agents. pi is the agent harness. both are open source.
on top of that, I use two skills:
1. a herdr skill that teaches the agent how to operate the CLI
2. an orchestrator skill that teaches it how to coordinate other agents
I talk to the orchestrator. it investigates the task, prepares briefs, assigns work, and launches workers in separate git worktrees through a helper called hspawn.
the workers implement. the orchestrator checks their actual changes, handles corrections, and keeps the main conversation available while they work.
I can choose the workers models independently from the model running the orchestrator. inside a pi session, I can also switch to another configured provider and model and keep going.
the skills and coordination process stay in place.
if I want workers briefed differently, I edit the skill. if I want stricter verification, I change the instructions. if I need different behavior from the harness, the source is available.
herdr and pi are beast open source projects that everyone should use imo.
you talk in a single chat about your project to orchestrator agent and only through that
Introducing Projects, a new way of working in Cursor.
Rather than creating a chat for every task, you work with a coordinator agent in a single, persistent thread.
Like @bot, your agent is always on, proactively manages work with subagents, and improves over time.
this is massively overcomplicated for the coding workflow most people actually need.
use orchestrator agent that will manage your herd of agents,
that is all you need
1 orchestrator agent that will run all your agents, i've been running this setup for months,
it only recently came out to cursor and claude,
avoid being locked into subscriptions from claude, grok, openai etc,
my setup is herdr + pi, two skills, and a helper that launches workers in separate git worktrees,
orchestrator agent manages the whole setup, from talking to agents, exploring and cleaning up,
once you use claude, grok or any other provider software daily it becomes so difficult to switch, especially if it's your business and how you make money,
real life picture of my orchestrator agent managing herd of agents
and this is how I work with agents now.
- I speak with Grok bots
- they coordinate Pi sessions
- Pi sessions interact with each other
- the Pi extension manipulates each worker's context
- the Pi extension guides workers to keep the journal up to date
- the Grok bot routine scans thread histories to improve the system
read more at:
mega.dev/autonomous-product-…
since this is becoming mainstream, I’d like to share my setup for doing this in a much better way without getting locked into single provider.
I use herdr + pi with two skills: one for operating the herdr CLI, and one for orchestration.
the basic idea is simple: you talk to an orchestrator agent, and it delegates work to other agents.
herdr manages their terminals and workspaces. pi runs the agents. the skills explain how to operate the setup and how to coordinate the work.
you give the orchestrator a task, it breaks down the work, prepares briefs, and launches workers in separate git worktrees. each worker gets a clear assignment, and the orchestrator checks the actual changes when they’re finished.
the main conversation stays available while they work, so I can add context, change priorities, or discuss what comes next.
but the part I care about most is being able to change providers and models while keeping this workflow.
I can choose different models for the orchestrator and the workers. I can also switch models within a pi session and keep going with the same conversation and skills.
once you’ve spent time building a workflow around an AI product, moving away becomes harder. you’ve tuned your instructions, learned its behavior, and organized your work around how it does things.
both herdr and pi are open source,
trust me you want this freedom, i have been running this setup for months and it only came out yesterday inside claude, possibilities are endless,
everything i do goes through single orchestrator agent or i group work per project and delegate multiple orchestrator agents
OPEN SOURCE AI needs to win on the boring stuff too.
Docs. Deployment. Upgrades. A model people can actually run is a stronger alternative to closed labs than a leaderboard screenshot.
>build an AI meeting summarizer
>everyone uses it
>nobody reads the summaries
>add an AI summary of the summaries
enterprise software has achieved recursion
I think an export button is underrated marketing for a SaaS
If people can leave with all their data they have one less reason to be scared of trying it
Make the exit easy and then build something they wanna stay for
The AI result I would pay attention to: the same improvement across several unfamiliar tasks, with failed runs included.
Less dramatic than a perfect demo. Much more useful for deciding what to trust.
an ai agent calling something a quick cleanup is how you end up reviewing a new religion in package.json
everyone's celebrating 75% cheaper cache reads while cost-per-task went up. the unit that matters flipped and half the timeline hasn't noticed.
Claude Fable 5.1 tops the Artificial Analysis Intelligence Index but costs 20% more per task than Fable 5 despite a 75% cache read price cut
We supported @AnthropicAI with pre-release evaluation of Claude Fable 5.1. At max effort it scores 66 on the Artificial Analysis Intelligence Index, the highest score we have measured, ahead of Claude Opus 5 (max, 63), Claude Fable 5 (max, 62), GPT-5.6 Sol (max, 61) and Grok 4.6 (high, 61). We evaluated the model with Anthropic's ‘default’ server-side fallback, which routes safety-flagged requests to Claude Opus 4.8 or Claude Opus 5; fallback served ~4% of output tokens across the Intelligence Index.
Key takeaways
➤ Frontier Intelligence with improvements across benchmarks: Fable 5.1 gains +4 points on the Intelligence Index over Fable 5. On HLE, Fable 5.1 scores 59.1%, ahead of the previous best of 55.5% from Claude Fable 5. It posts the narrowly highest scores we’ve seen on Terminal-Bench v2.1 (91.4%) and SciCode (62.0%), and on τ³-Banking it gains 9 points over Fable 5
➤ 75% cache read price cut, but Fable 5.1 still costs more per task: Anthropic has cut the cache read price from $1 to $0.25 per 1M cached input tokens, with standard pricing unchanged at $10/$50 per 1M input/output tokens. Fable 5.1 (max) costs $3.76 per Intelligence Index task, 20% more than Fable 5 (max), because it uses ~1.7x the output tokens. The cache cut saves ~$1.40 per task, concentrated in the agentic evaluations where the majority of input tokens are cache reads. At xhigh effort Fable 5.1 scores 65 at $2.72 per task, $1.04 less than max, but still above Claude Opus 5 (max, 63) at $2.34
➤ Claude Fable 5.1 holds the upper end of the Intelligence vs Output Tokens per Task Pareto frontier: every model variant scoring higher than GPT-5.6 Sol (medium) on the Intelligence Index is matched or beaten by a Fable 5.1 effort level on both intelligence and token usage
➤ Highest scores on agentic work tasks, but effectively tied with Opus 5: Fable 5.1 sets the highest scores we have measured on GDPval-AA v2 (1,853 Elo, +130 over Fable 5) and AA-Briefcase (1,694 Elo, +122 over Fable 5), our agentic knowledge work evaluations. Against Claude Opus 5 the GDPval-AA v2 lead is within the confidence interval and AA-Briefcase (1,685) is effectively tied, with Fable 5.1 ahead on analytical quality and rubric correctness, but behind on presentation
Other model details:
➤ Context window: 1 million tokens, supporting image and text inputs as with Anthropic’s other recent launches
➤ Pricing: Fable 5.1 retains the $10/$50/$12.5 input, output, and cache write prices per million tokens from Fable 5, but cache hits have been reduced to $0.25 per million tokens, a 75% relative reduction from before that will materially reduce agentic workload costs