@seldon_techi
iAccount based inSwitzerland
About this account
- Account based in
- Switzerland
- Connected via
- Switzerland App Store
Account-level information from X, not a live location or the device used for a specific post.
Advancing multimodal AI & computer-use
San Francisco, CA
Joined April 2026
- Tweets37
- Following14
- Followers245
- Likes276
Our Co-Founders @PascalBtd and @atoof_sh sat down on the Europe on Edge podcast and talked about the current state of multimodal AI and the company's mission. If this mission resonates with you, have a look at our open job listings (links are below).
Podcast episode: open.spotify.com/episode/60H…
Career page: seldon.global/careers
New SOTA on CADBench: Gemini 3.8 Flash.
30% pass rate on 105 expert-authored CAD tasks in Autodesk Fusion, up from 24.6% for Gemini 3.7 Flash.
Muse Spark 1.3 also improves its mean verifier score by 54% over 1.2, though full task completion remains a challenge.
Muse Spark 1.3 improves despite pass rate staying at 0%:
• Mean verifier score: 13.44 → 20.66 (+54%)
• Average runtime: 60.8 → 23.8 minutes (61% shorter)
• Valid tool calls: 76.8% → 93.0%
Better intermediate results and faster runs, but full task completion remains unsolved.
Read more: seldon.global/blog/cadbench
We’re expanding into more real-world engineering workflows across multiple domains, with new tasks and benchmarks coming soon.
Seldon retweeted
Can computer-use agents use complicated graphical software? You’re watching GPT-5.6 Sol operate Autodesk Fusion for 250 turns as it attempts a real mechanical-design task. The run is part of a benchmark aimed at the question "how good are agents at CAD?"
@seldon_tech's CADBench contains 105 mechanical-design tasks tested across 10 frontier models, and the results show how early reliable, autonomous CAD work remains:
More than two-thirds of the tasks are still unsolved.
That gap matters as agents make their way into the professional tools used to design the physical world. Getting around Autodesk Fusion is one thing, but producing a correct, editable model an engineer can pick up and use is much harder.
CADBench gives us a way to see and measure that frontier.
Great work from the Seldon team! Explore CADBench: seldon.global/blog/cadbench
Gemini 3.7 Flash delivers performance comparable to Fable 5 at a fraction of the cost
We see that many models finish the tasks prematurely, resulting in lower scores and costs
Read the full report here:
seldon.global/blog/cadbench
Seldon retweeted
drop by our meetup today if you're in SF!
luma.com/g7ludajl
Infinibench-v0: part of the benchmark was created by a continuously exploring agent, rewarded for identifying new failure modes in VLMs.
It produces non-trivial, out-of-distribution problems — trivially easy for humans, yet SOTA VLMs fail — and diagnoses failure modes that were previously unknown.
To celebrate the release of this benchmark we'll host a meetup in SF. If you're in town, sign up here: luma.com/g7ludajl
Full blog-post about this benchmark: seldon.global/blog/vgi-bench
Introducing VGI-Bench: a multimodal, holistic benchmark probing 12 distinct visual and audio-visual skills.
550 human-curated questions, designed to mitigate the common mistakes in today's video benchmarks and expose pragmatic failures of state-of-the-art models.
Best model: 64.73%. Humans: 84.5%.
Every question in our final dataset must pass two gates, in order:
-> Not solved in pass^3 by a blind, text-only model
-> Not solved in pass^3 by an older baseline VLM (Gemini 2.5 Flash-Lite)
This ensure that every questions tests actual visual understanding
VLMs are today's de-facto standard for visual reasoning. They are widely applied in embodied systems like humanoid robots and self-driving cars, and serve as the visual backbones for computer-use and design agents.
This demands diagnostic benchmarks that show exactly where a model can be trusted and what remains difficult.