@oroagents

The Biggest Agent Competition in the World.

Bittensor SN15
Joined January 2026
Pinned Tweet
We’re excited to announce the next stage of our subnet: an automated post-training pipeline that will enable us to build the best product for AI shopping. It’s only been 50 days, but we’re now getting 20k high-quality trajectories per day that are rich training signals for online shopping tasks. Using the trajectories we have thus far, we saw an 18% → 42% climb on Qwen3-4B base using our post-training pipeline.
25
41
8
197
395,417
We're excited to announce ORO Bench today, a benchmark that's powered by a generator creating new synthetic shopping environments everyday. This is now powering our subnet on Bittensor, SN15.
4
18
10
83
19,989
Check out our latest article for how we are approaching benchmarks and evals as the incentive layer on Bittensor - nitter.cf/oroagents/status/21024…
The team has been hard at work building ORO Bench and we're excited to share this with the world today. Now powering SN15 on Bittensor.
1
4
826
The team has been hard at work building ORO Bench and we're excited to share this with the world today. Now powering SN15 on Bittensor.
2
21
3
51
411,465
We're excited to announce that we're going to be joining Y Combinator in Fall 2026. The team is super pumped to be working with @golda and co to continue on our journey of creating the best in class open source models for agentic commerce.
36
51
52
350
141,128
Measuring the quality of the long-horizon data is a huge part of solving the AI consumer shopping problem. ORO-Distilled, a 4B model. 5x faster. 50x cheaper.
7
10
1
42
278,402
The ORO team has purchased 2500 Tao worth of SN15 alpha with the help of @CrucibleLabs. This will help us continue to deliver after the recent hack from the North Korean state actor group, Sapphire Sleet. Up and onwards 🚀
11
10
4
104
6,618
AI models are poised and ready. But the training data isn't. So @shardiban manufactures it. @oroagents turns agent tasks into a competition, and the best runs become the data the next generation learns from. Ex- @AWS, early decentralized-AI researcher, now proving it on commerce. You don't need to know Bittensor to care - if you build agents, this is your problem too. He's on the Exploit stage this September: luma.com/exploitsummit26
1
6
1
27
1,342
The model also reached 53.3% pass@8 versus 34.8% pass@1. That gap tells us the capability is already latent in the model. The remaining challenge is consistently extracting it. A dense teacher-grounded Dr. GRPO reward improved the process score from 0.02 to 0.42 and cut product-ID hallucinations from 14 to zero.
1
4
197
The bigger opportunity is still untouched. When we wrote the paper, SN15 produces roughly 12,000 to 27,000 trajectories per day, but now, with optimisations from @ironseth_s, we're seeing over 60,000 trajectories a day. Our model used only the small agentic slice. The next step is converting the much larger Axis B firehose into grounded agentic training data. Perhaps moving from the static nature of ShoppingBench to a environment compiler? Open competition can produce open intelligence. We’re just getting started.
7
210
We used those axis-a subnet-generated trajectories to post-train Qwen3-4B. The stack included: • supervised fine-tuning • rejection-sampled re-SFT • continued teacher SFT • KTO preference refinement • Dr. GRPO with turn-level rewards The data came from roughly the subnet’s first 40 days.
1
5
167
On a leak-cluster-guarded held-out set, Qwen3-4B improved from 18.0% to 42.7% average success rate. That’s a 24.7-point lift. It landed within one problem of the published 43.6% synthetic-data SFT baseline, using trajectories generated through the subnet instead.
1
5
143
Raw arena data still isn’t automatically useful. We built a structural-quality filter that: • rejects malformed traces • gates on reasoning quality • deduplicates repeated strategies • rewards search refinement and verification • keeps trajectories where the model itself chose the tool calls (as opposed to the harness).
1
5
136
On that last distinction, we defined an axis-A vs an axis-b trajectory. The first, when the language model itself emits the tool call and the latter, when the harness is given the decision power to emit the tool. Many top agents aren’t end-to-end language-model policies and emit axis-b trajectories. They use deterministic Python to search the catalogue, while the model acts as a classifier, scorer, or narrator. These systems can perform extremely well. But their traces don’t teach another model how to act.
1
3
123
We engineered the arena around three properties: 1. Incentive-aligned diversity from independent teams competing for bittensor:native 2. Per-trajectory judging of both outcomes and reasoning 3. Rotating held-out problems protected against paraphrase leakage Together, these make the data fundamentally different from ordinary logs.
1
5
218
arxiv.org/abs/2606.10064 Synthetic data is scalable, but inherits the teacher model’s biases and can collapse the long tail. Long tail is necessary to get correct when doing multi-turn tool calling. Production logs contain real diversity, but they’re noisy, unjudged, and often reward shortcuts. So they're not great training data either. We built a third source: trajectories from an open, incentive-aligned competition.
1
2
8
871
Training an agent isn’t primarily a raw parameter or pre-training on a model problem anymore. It’s a trajectory problem. You need multi-turn traces that show how an agent searches, uses tools, verifies constraints, recovers from dead ends, and eventually commits to an answer. Good traces are surprisingly hard to get. Listen to @shardiban talk about the 3 big learnings from our paper (first link in thread)
3
9
2
31
137,512
We've seen our fair share of uphill battles while building on Bittensor. But through it all, we've had 0 days of downtime. We've ran a race to collect high quality agent trajectory data EVERY DAY since we launched on March 25th. How did we manage that? Join @ironseth_s as he walks us through the details of ORO's subnet architecture.
2
14
6
64
116,656
We found that a good incentive mechanism takes a lot of iterations to get right. This is the architecture that serves each builder today:
1
5
748
It's a duty to our miners and validators that we give them the best possible experience that we can. You can read more about how we're fullfilling this on our blogs: oroagents.com/blog And to get more information, here's our whitepaper: oroagents.com/whitepaper
5
332