@CamelAIOrgi
iAccount based inUnited Kingdom
About this account
- Account based in
- United Kingdom
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
https://nitter.cf/t.co/FmX1B3nzjA is working on finding the scaling laws of agents. The first and the best multi-agent framework. Discord: https://nitter.cf/t.co/DRweXf0nOl. Product @Eigent_AI
Joined June 2023
- Tweets1.5K
- Following96
- Followers8.6K
- Likes1.6K
Pinned Tweet
Open AI and Anthropic just release GPT-5.3-Codex and Opus 4.6 model, terminal capability is now on top of their list evaluating modal capability. But terminal training hits a wall fast: there aren’t enough high-quality environments.
In SETA, we just shipped 1,376 validated terminal environments across: SE • sysadmin • security • debugging • networking • DevOps
Compatible with Terminal Bench & Harbor. @Mike_A_Merrill @alexgshaw
And we’re scaling fast 👀
Find it in: github.com/camel-ai/seta-env or search for seta-env in on harbor registry
Thanks for highlighting our work. Great work @AtriaASI!
We’re excited to highlight SETA, an open-source RL environment for training and evaluating AI agents from the @CamelAIOrg .
A valuable resource for researchers and developers exploring agentic reinforcement learning and complex task environments.
Check it out and support the CAMEL-AI community! 🚀
🔗 github.com/camel-ai/seta
CAMEL-AI.org retweeted
Replying to @OpenAI
@OpenAI GPT 6 Astra is now on Eigent! ⚡️⚡️⚡️
Our design engineer @douglas_ym used Eigent with GPT 6 Astra to rebuild his old master’s project, TaSH, from just a few photos of the original prototype.
Eigent recreated the full product in Blender, filled in details missing from the original prototype, and automatically rendered a 30-second @Apple style product showcase video.
Download Eigent today. Give Eigent a few images from one of your forgotten projects and see what GPT 6 Astra can rebuild.
Fully open source. Self-hostable on your desktop. Even for confidential projects, your data can stay completely under your control.
We are starting in 5 mins! Come join us 🐫
Zoom Meeting ID: 872 7159 9220
Passcode: 886353
This Friday on 🐫 CAMEL-AI Live Talk, Ziyi Wang @Ziyi0_0 is sharing Trajectory2Task: a trajectory-first way to train stronger tool-calling agents with synthesized, verifiable data.
Instead of writing tasks first and hoping agents can solve them, Trajectory2Task starts with executable tool-use trajectories, then turns their outcomes into realistic user intents. The result: training data that is diverse, grounded, and actually checkable.
If you care about LLM agents, tool use, or synthetic data for post-training, don’t miss this one!
🗓️ Sep 4 · 4 PM UK Time / 8 AM PT · Zoom
📝 Register: forms.gle/aqTzxMy9KqauHQj28
This Friday on 🐫 CAMEL-AI Live Talk, Ziyi Wang @Ziyi0_0 is sharing Trajectory2Task: a trajectory-first way to train stronger tool-calling agents with synthesized, verifiable data.
Instead of writing tasks first and hoping agents can solve them, Trajectory2Task starts with executable tool-use trajectories, then turns their outcomes into realistic user intents. The result: training data that is diverse, grounded, and actually checkable.
If you care about LLM agents, tool use, or synthetic data for post-training, don’t miss this one!
🗓️ Sep 4 · 4 PM UK Time / 8 AM PT · Zoom
📝 Register: forms.gle/aqTzxMy9KqauHQj28
CAMEL-AI.org retweeted
Gemini 3.8 now on Eigent! ⚡️⚡️⚡️
We handed it a real finance analyst's job: value Equinix vs Digital Realty, Damodaran-style, straight from SEC filings.
A team of agents researched, checked FX, and built the DCF, then handed back an Excel file you can actually poke at.
Click a number → see the filing it came from. Change an assumption → watch the valuation shift live.
Gemini 3.8 is now supported on Eigent via BYOK and cloud API. Try it today!
Huge congrats to the incredible @radixark team on the launch of Miles v0.1! We've really enjoyed using it in our research projects such as SETA: github.com/camel-ai/seta/tre…
Today we're launching Miles v0.1, an open-source RL framework for LLMs and multimodal models.
RL training is easy to start and hard to debug. Miles helps you ensure your run is correct, use hardware efficiently, and keep RL running at scale.
Over the past 9 months, 72 contributors have landed 1,326 commits, 85 GPU E2E CI tests, battle-testing Miles on frontier open models like Kimi K3, DeepSeek V4, Qwen 3.8, GLM 5.2, Inkling, MiniMax H3, etc.
Miles powers frontier-model development and production RL workloads at @humansand, @periodiclabs, @modal, @DecagonAI, @Eigent_AI, @nebiusai, @IBM and more, on both @NVIDIAAI and @AIatAMD hardware.
Here is what we built, and why teams picked Miles🧵
CAMEL-AI.org retweeted
Gemini 3.7 Flash now live on Eigent on both byok and cloud api ⚡️
Lightning speed meets unbelievable precision:
> Video-to-3D understanding
> Ultra-detailed, articulated CAD/mesh models (.glb)
> Full PBR materials & emissive lighting
From video reference to production-ready 3D in 3 mins! We also did a comparison with 3.6 Flash it's 2x faster and delivers far richer, more accurate details! 🚀
Download Eigent with free credits to build your own 3D models today!
We’re starting in 30 minutes! If you're working on model evaluation, LLM testing, or agent diagnostics, come join CAMEL-AI Community Live Talk by @NishantBalepur! 🐫
Meeting Info👇 us05web.zoom.us/j/8573785877…
Meeting ID: 85737858776 Passcode: Hz3cR1
This Friday on 🐫 CAMEL-AI Live Talk, Nishant is unveiling BenchMarker, and it might change how much you trust your favorite benchmarks. He is an NLP PhD candidate at the University of Maryland who cares about making LLM evaluation actually mean something.
The idea is simple but sharp: grade benchmarks the way teachers grade exam questions. BenchMarker uses LLM judges to catch three sneaky flaws in multiple choice tests, internet contamination, shortcuts models quietly exploit, and writing errors scored against a 19 rule rubric. Validated against human annotations and run across 12 popular benchmarks, the flaws showed up everywhere.
If your work touches evaluation, LLM testing, or agent diagnosis, this one is for you.
🗓 Jul 31 · 1 PM UK (BST) / 8 AM US East (EDT) · Zoom
📝 Register: docs.google.com/forms/d/e/1F…
This Friday on 🐫 CAMEL-AI Live Talk, Nishant is unveiling BenchMarker, and it might change how much you trust your favorite benchmarks. He is an NLP PhD candidate at the University of Maryland who cares about making LLM evaluation actually mean something.
The idea is simple but sharp: grade benchmarks the way teachers grade exam questions. BenchMarker uses LLM judges to catch three sneaky flaws in multiple choice tests, internet contamination, shortcuts models quietly exploit, and writing errors scored against a 19 rule rubric. Validated against human annotations and run across 12 popular benchmarks, the flaws showed up everywhere.
If your work touches evaluation, LLM testing, or agent diagnosis, this one is for you.
🗓 Jul 31 · 1 PM UK (BST) / 8 AM US East (EDT) · Zoom
📝 Register: docs.google.com/forms/d/e/1F…
See you at #AGIPlayground 2026 in Singapore, where @GeekParkHQ is gathering ~500 AI builders 🇸🇬
A playground for AI founders & builders. Aug 3–4. Gardens by the Bay, Singapore.
🔗 luma.com/n9r72dc9
CAMEL-AI.org retweeted
"behavioral state decay": the failure mode that long-horizon agents forget what matters. meta ai's fix: a memory agent that updates the memory bank and then decides whether to emit a proactive intervention. they train qwen3.5-27b on seta using sft and grpo and show gains on terminal-bench 2.0. great to see seta terminal agent rl envs used this way
- remember when it matters by meta ai @yifannnwu @zhuokaiz: arxiv.org/abs/2607.08716
- our seta project: github.com/camel-ai/seta
CAMEL-AI.org retweeted
we added grok 4.5 to eigent cloud and it is free for new registered users! also check out a quite smooth demo that builds an interactive rocket explorer with eigent x grok 4.5. ofc i know nothing about rocket. probably need someone from @SpaceXAI @SpaceX @elonmusk to verify if the result is good 😆
Grok 4.5 is now live on Eigent Cloud! @SpaceXAI
Download Eigent and get started with 500 free registration credits!
CAMEL-AI.org retweeted
Sharing our ICML’2026 paper Gecko: a stateful simulated environment for agentic tool calling, lead by our intern Zeyu Zhang @zeyu_au at @Eigent_AI x @CamelAIOrg from ANU. It was a great partnership with Prof. @LiangZheng_06 from ANU, @boltzmanns0ul from @AipoLabs. Come to our poster in Seoul to chat more!
Simulated stateful tool calling will become a common technique for training tool calling agents at large scale since the instability of real-world tools. This work shares the same spirit as Toolathlon-GYM we released earlier this year: github.com/eigent-ai/toolath…
If you are interested in scaling tool-use environments or want to partner on training tool-use agents, please dm me!
Will share more progress in training tool-use agents and scaling tool-use environments. Stay tuned!
Execution feedback from real tools can help agents refine tool-use plans. However, repeated real-tool calls may incur costs, hit rate limits, or cause side effects.
Our ICML 2026 paper introduces Gecko, a pre-execution simulation sandbox for tool-use agents.
Given tool schemas or descriptions, Gecko creates simulated tool interfaces, validates tool calls, simulates responses, tracks task state, and returns task-level feedback.
Project: camel-ai.github.io/gecko/
Code: github.com/camel-ai/gecko
Paper: arxiv.org/abs/2602.19218
CAMEL-AI.org retweeted
task state matters. ReAct bases its action on the observation. but the observation is global and does not precisely reflect the consequences of the action. Gecko gives a good task state estimation which better informs actions - all happening in simulation. no pains in real API costs. LLM agnostic.
real gains over state-of-the-art LLMs like GPT5.5 and Gemini-3.0-Pro.
camel-ai.github.io/gecko/
can we make it more efficient? can we generate RL trajectories with it? open questions are there
Execution feedback from real tools can help agents refine tool-use plans. However, repeated real-tool calls may incur costs, hit rate limits, or cause side effects.
Our ICML 2026 paper introduces Gecko, a pre-execution simulation sandbox for tool-use agents.
Given tool schemas or descriptions, Gecko creates simulated tool interfaces, validates tool calls, simulates responses, tracks task state, and returns task-level feedback.
Project: camel-ai.github.io/gecko/
Code: github.com/camel-ai/gecko
Paper: arxiv.org/abs/2602.19218
Execution feedback from real tools can help agents refine tool-use plans. However, repeated real-tool calls may incur costs, hit rate limits, or cause side effects.
Our ICML 2026 paper introduces Gecko, a pre-execution simulation sandbox for tool-use agents.
Given tool schemas or descriptions, Gecko creates simulated tool interfaces, validates tool calls, simulates responses, tracks task state, and returns task-level feedback.
Project: camel-ai.github.io/gecko/
Code: github.com/camel-ai/gecko
Paper: arxiv.org/abs/2602.19218
We see Gecko as useful beyond test-time refinement:
• improving tool-use training data
• building interactive RL environments
• reducing real-tool trial-and-error before deployment
Zeyu Zhang @zeyu_au, the first author of the Gecko paper will be at ICML. If you have any questions about Gecko or interested in the topic, feel free to find him there and have a chat!