@CamelAIOrg

https://nitter.cf/t.co/FmX1B3nzjA is working on finding the scaling laws of agents. The first and the best multi-agent framework. Discord: https://nitter.cf/t.co/DRweXf0nOl. Product @Eigent_AI

Joined June 2023
Open AI and Anthropic just release GPT-5.3-Codex and Opus 4.6 model, terminal capability is now on top of their list evaluating modal capability. But terminal training hits a wall fast: there aren’t enough high-quality environments. In SETA, we just shipped 1,376 validated terminal environments across: SE • sysadmin • security • debugging • networking • DevOps Compatible with Terminal Bench & Harbor. @Mike_A_Merrill @alexgshaw And we’re scaling fast 👀 Find it in: github.com/camel-ai/seta-env or search for seta-env in on harbor registry
4
6
1
41
6,745
Thanks for highlighting our work. Great work @AtriaASI!
We’re excited to highlight SETA, an open-source RL environment for training and evaluating AI agents from the @CamelAIOrg . A valuable resource for researchers and developers exploring agentic reinforcement learning and complex task environments. Check it out and support the CAMEL-AI community! 🚀 🔗 github.com/camel-ai/seta
1
5
526
CAMEL-AI.org retweeted
Replying to @OpenAI
@OpenAI GPT 6 Astra is now on Eigent! ⚡️⚡️⚡️ Our design engineer @douglas_ym used Eigent with GPT 6 Astra to rebuild his old master’s project, TaSH, from just a few photos of the original prototype. Eigent recreated the full product in Blender, filled in details missing from the original prototype, and automatically rendered a 30-second @Apple style product showcase video. Download Eigent today. Give Eigent a few images from one of your forgotten projects and see what GPT 6 Astra can rebuild. Fully open source. Self-hostable on your desktop. Even for confidential projects, your data can stay completely under your control.
3
3
1
12
696
We are starting in 5 mins! Come join us 🐫 Zoom Meeting ID: 872 7159 9220 Passcode: 886353
This Friday on 🐫 CAMEL-AI Live Talk, Ziyi Wang @Ziyi0_0 is sharing Trajectory2Task: a trajectory-first way to train stronger tool-calling agents with synthesized, verifiable data. Instead of writing tasks first and hoping agents can solve them, Trajectory2Task starts with executable tool-use trajectories, then turns their outcomes into realistic user intents. The result: training data that is diverse, grounded, and actually checkable. If you care about LLM agents, tool use, or synthetic data for post-training, don’t miss this one! 🗓️ Sep 4 · 4 PM UK Time / 8 AM PT · Zoom 📝 Register: forms.gle/aqTzxMy9KqauHQj28
421
This Friday on 🐫 CAMEL-AI Live Talk, Ziyi Wang @Ziyi0_0 is sharing Trajectory2Task: a trajectory-first way to train stronger tool-calling agents with synthesized, verifiable data. Instead of writing tasks first and hoping agents can solve them, Trajectory2Task starts with executable tool-use trajectories, then turns their outcomes into realistic user intents. The result: training data that is diverse, grounded, and actually checkable. If you care about LLM agents, tool use, or synthetic data for post-training, don’t miss this one! 🗓️ Sep 4 · 4 PM UK Time / 8 AM PT · Zoom 📝 Register: forms.gle/aqTzxMy9KqauHQj28
1
1
3
714
CAMEL-AI.org retweeted
Gemini 3.8 now on Eigent! ⚡️⚡️⚡️ We handed it a real finance analyst's job: value Equinix vs Digital Realty, Damodaran-style, straight from SEC filings. A team of agents researched, checked FX, and built the DCF, then handed back an Excel file you can actually poke at. Click a number → see the filing it came from. Change an assumption → watch the valuation shift live. Gemini 3.8 is now supported on Eigent via BYOK and cloud API. Try it today!
Introducing Gemini 3.8 Flash, another jump in Gemini's agentic + coding capabilities, and our 3rd updated Flash model in only 6 weeks... This model has been a ton of fun to work with, excited to see what you all think!
2
2
15
3,045
Huge congrats to the incredible @radixark team on the launch of Miles v0.1! We've really enjoyed using it in our research projects such as SETA: github.com/camel-ai/seta/tre…
Today we're launching Miles v0.1, an open-source RL framework for LLMs and multimodal models. RL training is easy to start and hard to debug. Miles helps you ensure your run is correct, use hardware efficiently, and keep RL running at scale. Over the past 9 months, 72 contributors have landed 1,326 commits, 85 GPU E2E CI tests, battle-testing Miles on frontier open models like Kimi K3, DeepSeek V4, Qwen 3.8, GLM 5.2, Inkling, MiniMax H3, etc. Miles powers frontier-model development and production RL workloads at @humansand, @periodiclabs, @modal, @DecagonAI, @Eigent_AI, @nebiusai, @IBM and more, on both @NVIDIAAI and @AIatAMD hardware. Here is what we built, and why teams picked Miles🧵
1
1
14
1,420
CAMEL-AI.org retweeted
Grok 4.6 on Eigent: just one prompt and out comes a fully interactive Starlink satellite you can click any part to name it!
Introducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price.
7
3
14
1,078
CAMEL-AI.org retweeted
Gemini 3.7 Flash now live on Eigent on both byok and cloud api ⚡️ Lightning speed meets unbelievable precision: > Video-to-3D understanding > Ultra-detailed, articulated CAD/mesh models (.glb) > Full PBR materials & emissive lighting From video reference to production-ready 3D in 3 mins! We also did a comparison with 3.6 Flash it's 2x faster and delivers far richer, more accurate details! 🚀 Download Eigent with free credits to build your own 3D models today!
Gemini 3.7 Flash is here. It’s stronger for coding, knowledge work, and web development. 🧵
4
1
11
872
We’re starting in 30 minutes! If you're working on model evaluation, LLM testing, or agent diagnostics, come join CAMEL-AI Community Live Talk by @NishantBalepur! 🐫 Meeting Info👇 us05web.zoom.us/j/8573785877… Meeting ID: 85737858776 Passcode: Hz3cR1
This Friday on 🐫 CAMEL-AI Live Talk, Nishant is unveiling BenchMarker, and it might change how much you trust your favorite benchmarks. He is an NLP PhD candidate at the University of Maryland who cares about making LLM evaluation actually mean something. The idea is simple but sharp: grade benchmarks the way teachers grade exam questions. BenchMarker uses LLM judges to catch three sneaky flaws in multiple choice tests, internet contamination, shortcuts models quietly exploit, and writing errors scored against a 19 rule rubric. Validated against human annotations and run across 12 popular benchmarks, the flaws showed up everywhere. If your work touches evaluation, LLM testing, or agent diagnosis, this one is for you. 🗓 Jul 31 · 1 PM UK (BST) / 8 AM US East (EDT) · Zoom 📝 Register: docs.google.com/forms/d/e/1F…
1
5
748
This Friday on 🐫 CAMEL-AI Live Talk, Nishant is unveiling BenchMarker, and it might change how much you trust your favorite benchmarks. He is an NLP PhD candidate at the University of Maryland who cares about making LLM evaluation actually mean something. The idea is simple but sharp: grade benchmarks the way teachers grade exam questions. BenchMarker uses LLM judges to catch three sneaky flaws in multiple choice tests, internet contamination, shortcuts models quietly exploit, and writing errors scored against a 19 rule rubric. Validated against human annotations and run across 12 popular benchmarks, the flaws showed up everywhere. If your work touches evaluation, LLM testing, or agent diagnosis, this one is for you. 🗓 Jul 31 · 1 PM UK (BST) / 8 AM US East (EDT) · Zoom 📝 Register: docs.google.com/forms/d/e/1F…
2
3
7
3,521
See you at #AGIPlayground 2026 in Singapore, where @GeekParkHQ is gathering ~500 AI builders 🇸🇬 A playground for AI founders & builders. Aug 3–4. Gardens by the Bay, Singapore. 🔗 luma.com/n9r72dc9
1
4
378
CAMEL-AI.org retweeted
"behavioral state decay": the failure mode that long-horizon agents forget what matters. meta ai's fix: a memory agent that updates the memory bank and then decides whether to emit a proactive intervention. they train qwen3.5-27b on seta using sft and grpo and show gains on terminal-bench 2.0. great to see seta terminal agent rl envs used this way - remember when it matters by meta ai @yifannnwu @zhuokaiz: arxiv.org/abs/2607.08716 - our seta project: github.com/camel-ai/seta
5
13
1
92
13,227
CAMEL-AI.org retweeted
we added grok 4.5 to eigent cloud and it is free for new registered users! also check out a quite smooth demo that builds an interactive rocket explorer with eigent x grok 4.5. ofc i know nothing about rocket. probably need someone from @SpaceXAI @SpaceX @elonmusk to verify if the result is good 😆
Grok 4.5 is now live on Eigent Cloud! @SpaceXAI Download Eigent and get started with 500 free registration credits!
1
4
19
2,073
CAMEL-AI.org retweeted
Grok 4.5 is now live on Eigent Cloud! @SpaceXAI Download Eigent and get started with 500 free registration credits!
3
3
23
4,553
CAMEL-AI.org retweeted
Sharing our ICML’2026 paper Gecko: a stateful simulated environment for agentic tool calling, lead by our intern Zeyu Zhang @zeyu_au at @Eigent_AI x @CamelAIOrg from ANU. It was a great partnership with Prof. @LiangZheng_06 from ANU, @boltzmanns0ul from @AipoLabs. Come to our poster in Seoul to chat more! Simulated stateful tool calling will become a common technique for training tool calling agents at large scale since the instability of real-world tools. This work shares the same spirit as Toolathlon-GYM we released earlier this year: github.com/eigent-ai/toolath… If you are interested in scaling tool-use environments or want to partner on training tool-use agents, please dm me! Will share more progress in training tool-use agents and scaling tool-use environments. Stay tuned!
Execution feedback from real tools can help agents refine tool-use plans. However, repeated real-tool calls may incur costs, hit rate limits, or cause side effects. Our ICML 2026 paper introduces Gecko, a pre-execution simulation sandbox for tool-use agents. Given tool schemas or descriptions, Gecko creates simulated tool interfaces, validates tool calls, simulates responses, tracks task state, and returns task-level feedback. Project: camel-ai.github.io/gecko/ Code: github.com/camel-ai/gecko Paper: arxiv.org/abs/2602.19218
3
3
34
5,331
CAMEL-AI.org retweeted
task state matters. ReAct bases its action on the observation. but the observation is global and does not precisely reflect the consequences of the action. Gecko gives a good task state estimation which better informs actions - all happening in simulation. no pains in real API costs. LLM agnostic. real gains over state-of-the-art LLMs like GPT5.5 and Gemini-3.0-Pro. camel-ai.github.io/gecko/ can we make it more efficient? can we generate RL trajectories with it? open questions are there
Execution feedback from real tools can help agents refine tool-use plans. However, repeated real-tool calls may incur costs, hit rate limits, or cause side effects. Our ICML 2026 paper introduces Gecko, a pre-execution simulation sandbox for tool-use agents. Given tool schemas or descriptions, Gecko creates simulated tool interfaces, validates tool calls, simulates responses, tracks task state, and returns task-level feedback. Project: camel-ai.github.io/gecko/ Code: github.com/camel-ai/gecko Paper: arxiv.org/abs/2602.19218
1
4
13
1,467
Execution feedback from real tools can help agents refine tool-use plans. However, repeated real-tool calls may incur costs, hit rate limits, or cause side effects. Our ICML 2026 paper introduces Gecko, a pre-execution simulation sandbox for tool-use agents. Given tool schemas or descriptions, Gecko creates simulated tool interfaces, validates tool calls, simulates responses, tracks task state, and returns task-level feedback. Project: camel-ai.github.io/gecko/ Code: github.com/camel-ai/gecko Paper: arxiv.org/abs/2602.19218
4
8
2
26
9,224
We see Gecko as useful beyond test-time refinement: • improving tool-use training data • building interactive RL environments • reducing real-tool trial-and-error before deployment
1
119
Zeyu Zhang @zeyu_au, the first author of the Gecko paper will be at ICML. If you have any questions about Gecko or interested in the topic, feel free to find him there and have a chat!
1
147