@meetshahdev
Mountain View, CA
Joined February 2016
Meet Shah retweeted
Gemini Omni Flash is officially out! 🚀 It has been incredible working on the model pretraining, landing cool features in video editing, and pushing the 3D/camera capabilities in Omni. Seeing it live at #GoogleIO is surreal. 😍😇 Took countless days of no to little sleep. But sometimes it takes a little crazy to make something truly amazing. ✨✨✨ Go try it out! deepmind.google/models/gemin… #googleemployee @GeminiApp @GoogleDeepMind #gemini #omni
🤖 Made with AI
9
3
2
41
7,490
(1/10) Excited to share our new paper: "Towards Conversational Medical AI with Eyes, Ears and a Voice" arxiv.org/abs/2605.09272 Clinical medicine depends on more than words; it relies on real-time visual cues (rashes, distress, gait), auditory signals (prosody, breathing), and guided physical exams (e.g., guiding someone through correcting their inhaler technique, or walking them through shoulder maneuvers to work up a rotator cuff injury). But existing conversational medical AI systems are blind and deaf to this crucial information and are constrained to rigid, turn-by-turn user experience. As an important ingredient in AI co-clinician, our new healthcare initiative at @GoogleDeepMind, we developed a first-of-its-kind real-time multimodal system that conducts natural clinical consultations over video calls while continuously watching, listening, and reasoning. The preprint details the research progress we’ve made so far together with our collaborators at @StanfordMed and @harvardmed (@AdamRodmanMD @euanashley @DrJackOSullivan @jasongusdorf). More details below!
Since joining @GoogleDeepMind I’ve dreamt for a decade that AI can give clinicians superpowers: amplifying reach as an always-available, trustworthy member of the care team. Delighted to share our strategic research initiative at @GoogleDeepMind towards this vision: AI co-clinician deepmind.google/blog/ai-co-c… (1/n)
18
46
4
339
966,924
1/ Introducing AI co-clinician: latest research milestone I’ve been a part of at @GoogleDeepmind, researching the evolution of care delivery where AI can support both patients & doctors. tech report: gstatic.com/vesper/ai_coclin… blogpost: deepmind.google/blog/ai-co-c…
AI co-clinician is our new research initiative to help explore how multimodal agents could better support healthcare workers and patients. 🩺 Here’s a snapshot of our progress 🧵
1
3
18
2,088
8/ We utilize a dual-agent architecture to decouple fluid multimodal interaction from deep clinical reasoning. A low-latency "Talker" handles empathetic dialogue, while a supervisory "Clinical Planner" monitors the encounter state.
1
1
77
9/ We are currently in a phased research approach with global collaborators ( 🇺🇸 🇦🇺 🇳🇿 🇦🇪 🇸🇬 🇮🇳) to ensure this is developed responsibly. Huge congratulations to everyone on the team and our partners for making this possible.
1
44
AI co-clinician is our new research initiative to help explore how multimodal agents could better support healthcare workers and patients. 🩺 Here’s a snapshot of our progress 🧵
83
232
93
1,223
366,642
Meet Shah retweeted
I am happy to introduce AI co-clinician, @GoogleDeepMind's research initiative to explore how AI could better amplify doctor's expertise and help deliver higher quality care to patients. We’re excited about our early results, and are taking a phased approach to our research explorations with academic and research collaborators. Read more in our blog: deepmind.google/blog/ai-co-c…
21
98
13
607
40,187
VQA is consistently seeing tremendous progress year by year. Excited to see further progress as we now focus on specific sets of problems (reading text, consistency, external knowledge etc.). Planning research on V+L? Do have a look at Pythia github.com/facebookresearch/…
VQA performance on a standard benchmark (VQA v2 dataset) has gone up 20% (absolute) in the last ~4 years. You can really tell the difference when interacting with these models! Check the VQA demo out here: vqa.cloudcv.org.
1
2
14
Checkout and discuss our @cvpr2019 poster #196 "Towards VQA Models That Can Read" at 3:20PM. Learn how current VQA models fail to answer questions about image text and our steps toward solving it. More details + TextVQA dataset + Starter Code available at textvqa.org.
8
15
Meet Shah (@Jalebi_Fafda) presenting his @facebookai AI Residency work at #CVPR2019 on a new cycle consistecy based training framework to make VQA models robust to linguistic variations in the input question. Code + Dataset are publicly available facebookresearch.github.io/V…
3
35
If you're at #CVPR2019 kindly visit me present our work on "Cycle Consistency for robust VQA" in Grand Ballroom at 14.10. My @facebookai co-authors will be there to answer all your questions in the poster session #184 that follows. Dataset + Code + Paper: facebookresearch.github.io/V…
6
30
Meet Shah retweeted
VQA models need to be able to read and reason about text in an image to answer questions about them. But even today’s state-of-the-art models struggle to do this. To advance VQA and research, we're hosting a challenge: textvqa.org/challenge. Enter by 5/18!
Congrats to the 36 teams that participated in the VQA challenge! Now use your best VQA models in conjunction with LoRRA (github.com/facebookresearch/…) to participate in the TextVQA challenge (textvqa.org/)! Deadline May 18th (ping us if you want an extension).
1
8
2
66