@PascalBtd

physics @eth co-founder @ETH_agent_lab https://nitter.cf/t.co/xnO7yUUlr4

Joined February 2020
🇨🇭🇨🇭
Some data we recently assembled on entrepreneurship/compute in Europe: eudata.vercel.app. We hope that one of the useful roles that Stripe can play is in collecting and publishing empirical data pertaining to entrepreneurship and industry in Europe. There's growing appetite to get Europe on a better footing, and cross-sectional comparisons can often shine light on where opportunities lie. If you're interested in this kind of thing, we publish more at stripeeconomics.substack.com.
10
480
Pascal Bertrand retweeted
Our Co-Founders @PascalBtd and @atoof_sh sat down on the Europe on Edge podcast and talked about the current state of multimodal AI and the company's mission. If this mission resonates with you, have a look at our open job listings (links are below).
1
2
11
220
Thanks for having us!
Still a student? Want to launch a startup and raise your first angel round? Here’s the blueprint. Pascal Bertrand (@PascalBtd) & Atoof Shakir (@atoof_sh) join Europe on Edge to talk about building an AI startup at ETH, raising without a pitch deck, and teaching AI to see. 🎥 Trailer below 🎧 Full episode in the first reply
1
14
759
If someone asks me to design a politician in CAD
The head of Russias Rossotrudnichestvo in a press conference today talking about Russias deteriorating relationship with Armenia…
3
222
New results on CADBench: GPT-6 Astra numbers coming soon ;)
New SOTA on CADBench: Gemini 3.8 Flash. 30% pass rate on 105 expert-authored CAD tasks in Autodesk Fusion, up from 24.6% for Gemini 3.7 Flash. Muse Spark 1.3 also improves its mean verifier score by 54% over 1.2, though full task completion remains a challenge.
7
294
the vatican’s ai advisor???
JUST IN: King Charles to host exclusive AI gathering this month with Jensen Huang, Demis Hassabis, and the Vatican’s AI adviser, per Politico.
2
3
285
Interesting thoughts! This thread explains the reasoning behind us leaning into multimodality with robust environments very well. To be useful in the real world, models will need to rely on audiovisual inputs which isn’t yet feasible, as we have shown in VGI-Bench
This quoted post is unavailable.
5
194
Pascal Bertrand retweeted
a bit of a longer post with my experience training recursive language models with rl. it was quite hard and i had to fail a bunch of times to achieve a somewhat interesting final result. hope you'll enjoy it!
14
3
3
32
3,734
Pascal Bertrand retweeted
Can computer-use agents use complicated graphical software? You’re watching GPT-5.6 Sol operate Autodesk Fusion for 250 turns as it attempts a real mechanical-design task. The run is part of a benchmark aimed at the question "how good are agents at CAD?" @seldon_tech's CADBench contains 105 mechanical-design tasks tested across 10 frontier models, and the results show how early reliable, autonomous CAD work remains: More than two-thirds of the tasks are still unsolved. That gap matters as agents make their way into the professional tools used to design the physical world. Getting around Autodesk Fusion is one thing, but producing a correct, editable model an engineer can pick up and use is much harder. CADBench gives us a way to see and measure that frontier. Great work from the Seldon team! Explore CADBench: seldon.global/blog/cadbench
1
2
7
506
This is a cool series! Worth a read
How to build great evals - part 6 Hill climbing on evals is just a fancy way of saying: pick a dimension that matters and optimize for it. This could be improving the quality of existing features based on your latest production data on high value user journeys, expanding to adjacent use cases, lowering cost or latency. The actual work boils down to better harnesses and model selection through methods like prompt eng, context eng, memory, post training, deterministic old school code etc. Your failure mode taxonomy (from part 3) is a good compass for where your product struggles and needs some love. E.g. maybe tool calling failures are your most common problem. You dig in and notice you stuff 20 tools in context, when each task really only needs 3-5. hill climbing here involves context eng to give it the right tools at the right stage and iterate until you get it to good. Or take the cost reduction goal.. I’ve written about how I advise launching your product with the best model first. Get the quality as high as you can. Once you know users love the experience, hill climb to get similar quality with a smaller, cheaper, faster model. Same methods - harness, models. The important thing is to have evals that tell you whether you are actually moving in the right direction. More tomorrow. Send this to your teammates! Drop your questions in the comments and I will answer in future posts.
5
1,475
This will be fun!
We're hosting another event in San Francisco coming monday, this time a research talk with @samsja19 as guest speaker Drop by if you're in town! Sign up link is in the comments
9
280
Pascal Bertrand retweeted
why haven’t other engineering practices seen the same model improvements as coding? we measured models’ ability to interact with a cad environment, and the results were super interesting, with some curious upsets in model and harness rankings. i’m super excited for what’s to come in computer use
How good are agents actually at CAD? Today we release CADBench, evaluating 10 Models on 105 realistic tasks in Fusion 360 Fable 5 and Gemini 3.7 Flash are #1 with only 24.6% pass rate
2
8
475
This is pretty surprising
Replying to @seldon_tech
Gemini 3.7 Flash delivers performance comparable to Fable 5 at a fraction of the cost We see that many models finish the tasks prematurely, resulting in lower scores and costs
9
271
Pascal Bertrand retweeted
How good are agents actually at CAD? Today we release CADBench, evaluating 10 Models on 105 realistic tasks in Fusion 360 Fable 5 and Gemini 3.7 Flash are #1 with only 24.6% pass rate
27
61
17
540
147,417
Pascal Bertrand retweeted
Replying to @AIatMeta @xiaolonw
Teach my guy some CAD
1
1
7
637
Pascal Bertrand retweeted
Replying to @PascalBtd
thx for this link! we want to start evaluating VLMs on videos soon. articles like this are super inspiring for me.
1
1
91
professional situation monitoring
1
17
932