@guyd33

Machine learning researcher @JaneStreetGroup. PhD @NYUDataScience in AI & CogSci. Married to Sarah, dad to Nina (👶) and Lila (🐕‍🦺), amateur hot sauce maker.

New York, USA
Joined April 2019
New preprint alert! We often prompt ICL tasks using either demonstrations or instructions. How much does the form of the prompt matter to the task representation formed by a language model? Stick around to find out 1/N
1
36
5
277
52,180
Excited for #ICML2026 and looking forward to seeing everyone!
Hi friends (and friends I haven’t met yet)! I’ll be at ICML this year with some other folks from ML @ Jane Street. We do (really) exciting work across various machine learning domains and we can’t wait to see what the community is working on… 1/N
1
34
2,567
Starting my trip to Seoul for #ICML2026 off right
11
752
Hi friends (and friends I haven’t met yet)! I’ll be at ICML this year with some other folks from ML @ Jane Street. We do (really) exciting work across various machine learning domains and we can’t wait to see what the community is working on… 1/N
9
3
1
280
32,079
See our open roles at janestreet.com/join-jane-str… (and if you can’t find the right role for you, don’t hesitate to ping me) 3/N
1
1
13
3,035
Oh, and we have a PhD fellowship (no strings attached funding for a year: janestreet.com/join-jane-str…), too! If any of that sounds relevant to you, come find me at our booth or let’s set up a coffee chat. 4/N=4.
1
10
2,513
If “getting started with agents” feels like setup hell — same. So we made a starter tutorial: First agent running in <14 minutes, no Docker/AWS. Laptop + API key only. 👇 youtube.com/watch?v=gzNW_LXE…
3
13
2,410
Like ~everyone, I'll also be at #NeurIPS this week! Please reach out to chat about past (goal representations, cognitive science, intrep) or current interests (LLM mental state inference, social environments for RL). Also if you have leads on great coffee, craft beer, or tacos.
3
3
52
4,480
We're also presenting some work! Our (@adinamwilliams @LakeBrenden @todd_gureckis ) interpretability work on task representations from different prompting forms will be poster #1016 on Friday's afternoon session (4:30-7:30, hall C/D/E) nitter.cf/guyd33/status/19259677…
New preprint alert! We often prompt ICL tasks using either demonstrations or instructions. How much does the form of the prompt matter to the task representation formed by a language model? Stick around to find out 1/N
1
1
12
1,183
@jcyhc_ai will present SAGE-Eval, our (w/ @LakeBrenden) systematic generalization safety benchmark at poster #1104 on Friday AM (11-2). John does fantastic work and he's open to RE/RS roles or PhD positions in AI Safety. If you're hiring, talk to him! nitter.cf/jcyhc_ai/status/192811…
Do LLMs show systematic generalization of safety facts to novel scenarios? Introducing our work SAGE-Eval, a benchmark consisting of 100+ safety facts and 10k+ scenarios to test this! - Claude-3.7-Sonnet passes only 57% of facts evaluated - o1 and o3-mini passed <45%! 🧵
2
5
1,681
Stop by the Meta booth tomorrow, Wednesday Dec 3rd at #NeurIPS in San Diego! 🤖📱 We demo our new research environment, OpenApps, for digital agents. Generate thousands of app versions to train and evaluate multimodal agents to use apps like humans do. Not attending? Stay tuned
1
2
10
1,060
Guy Davidson retweeted
In San Diego for #NeurIPS Happy to chat about open-endedness, self goal-generation, intrinsic motivations, self-improvement, human-machine collective intelligence Open to hear about research scientist opportunities too Don't hesitate to reach out!
3
3
29
2,500
My team at FAIR at Meta is recruiting interns for next summer! If you're a PhD student interested in questions around theory of mind in language models for social, multi-agent settings, and have relevant background and/or experience: metacareers.com/jobs/1821713…
5
21
186
12,788
Other great humans you might end up working with include @bvp22294, @AnsongNi, and @real_asli (in addition to several twitter-less folks). Feel free to reach out with any questions! (though it may take me a bit to reply)
3
954
Belated update #2: my year at FAIR @AIatMeta through the AIM program was so nice that I’m sticking around for the long haul. I’m excited to stay at FAIR and work with @real_asli and friends on fun LLM questions; I’ll be working from the New York office so we’re sticking around.
1
1
77
7,014
Belated update #1: I defended my PhD about a month ago! I appreciate the warm reception from everyone who made it in-person and virtually. Thanks to my committee, @LerrelPinto, @togelius, and Mark Ho for your feedback and fun questions.
11
1
71
7,296
I owe tremendous thanks to many other people, all (or, hopefully, at least most) of whom I mentioned in my acknowledgments. I’m also so grateful my dad could represent my family, and for my wife, Sarah, for, well, everything.
1
3
605
Tune in tomorrow for belated update #2, on post-PhD plans!
1
481
do you use Letterboxd? would you be willing to participate in a 30-min research study where you use movie recommenders based on your Letterboxd ratings? DM me! (you will receive $20 for participating)
3
13
2,217