@meetshahdevi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
Research Engineer @GoogleDeepmind prev: @Waymo, @Uber_ATG, @metaai, @iitbombay
Mountain View, CA
Joined February 2016
- Tweets31
- Following531
- Followers186
- Likes35.8K
Meet Shah retweeted
Gemini Omni Flash is officially out! 🚀
It has been incredible working on the model pretraining, landing cool features in video editing, and pushing the 3D/camera capabilities in Omni. Seeing it live at #GoogleIO is surreal. 😍😇
Took countless days of no to little sleep. But sometimes it takes a little crazy to make something truly amazing. ✨✨✨
Go try it out! deepmind.google/models/gemin…
#googleemployee @GeminiApp @GoogleDeepMind #gemini #omni
🤖 Made with AI
Meet Shah retweeted
(1/10) Excited to share our new paper: "Towards Conversational Medical AI with Eyes, Ears and a Voice"
arxiv.org/abs/2605.09272
Clinical medicine depends on more than words; it relies on real-time visual cues (rashes, distress, gait), auditory signals (prosody, breathing), and guided physical exams (e.g., guiding someone through correcting their inhaler technique, or walking them through shoulder maneuvers to work up a rotator cuff injury).
But existing conversational medical AI systems are blind and deaf to this crucial information and are constrained to rigid, turn-by-turn user experience.
As an important ingredient in AI co-clinician, our new healthcare initiative at @GoogleDeepMind, we developed a first-of-its-kind real-time multimodal system that conducts natural clinical consultations over video calls while continuously watching, listening, and reasoning.
The preprint details the research progress we’ve made so far together with our collaborators at @StanfordMed and @harvardmed (@AdamRodmanMD @euanashley @DrJackOSullivan @jasongusdorf).
More details below!
Since joining @GoogleDeepMind I’ve dreamt for a decade that AI can give clinicians superpowers: amplifying reach as an always-available, trustworthy member of the care team. Delighted to share our strategic research initiative at @GoogleDeepMind towards this vision: AI co-clinician deepmind.google/blog/ai-co-c… (1/n)
1/ Introducing AI co-clinician: latest research milestone I’ve been a part of at @GoogleDeepmind, researching the evolution of care delivery where AI can support both patients & doctors.
tech report: gstatic.com/vesper/ai_coclin…
blogpost: deepmind.google/blog/ai-co-c…
8/ We utilize a dual-agent architecture to decouple fluid multimodal interaction from deep clinical reasoning. A low-latency "Talker" handles empathetic dialogue, while a supervisory "Clinical Planner" monitors the encounter state.
Meet Shah retweeted
AI co-clinician is our new research initiative to help explore how multimodal agents could better support healthcare workers and patients. 🩺
Here’s a snapshot of our progress 🧵
Meet Shah retweeted
I am happy to introduce AI co-clinician,
@GoogleDeepMind's research initiative to explore how AI could better amplify doctor's expertise and help deliver higher quality care to patients.
We’re excited about our early results, and are taking a phased approach to our research explorations with academic and research collaborators.
Read more in our blog: deepmind.google/blog/ai-co-c…
Meet Shah retweeted
VQA is consistently seeing tremendous progress year by year. Excited to see further progress as we now focus on specific sets of problems (reading text, consistency, external knowledge etc.).
Planning research on V+L? Do have a look at Pythia github.com/facebookresearch/…
VQA performance on a standard benchmark (VQA v2 dataset) has gone up 20% (absolute) in the last ~4 years. You can really tell the difference when interacting with these models! Check the VQA demo out here: vqa.cloudcv.org.
Meet Shah retweeted
Checkout and discuss our @cvpr2019 poster #196 "Towards VQA Models That Can Read" at 3:20PM. Learn how current VQA models fail to answer questions about image text and our steps toward solving it. More details + TextVQA dataset + Starter Code available at textvqa.org.
Meet Shah retweeted
Meet Shah (@Jalebi_Fafda) presenting his @facebookai AI Residency work at #CVPR2019 on a new cycle consistecy based training framework to make VQA models robust to linguistic variations in the input question.
Code + Dataset are publicly available
facebookresearch.github.io/V…
If you're at #CVPR2019 kindly visit me present our work on "Cycle Consistency for robust VQA" in Grand Ballroom at 14.10. My @facebookai co-authors will be there to answer all your questions in the poster session #184 that follows.
Dataset + Code + Paper: facebookresearch.github.io/V…
Meet Shah retweeted
VQA models need to be able to read and reason about text in an image to answer questions about them. But even today’s state-of-the-art models struggle to do this. To advance VQA and research, we're hosting a challenge: textvqa.org/challenge. Enter by 5/18!
Meet Shah retweeted
Congrats to the 36 teams that participated in the VQA challenge! Now use your best VQA models in conjunction with LoRRA (github.com/facebookresearch/…) to participate in the TextVQA challenge (textvqa.org/)! Deadline May 18th (ping us if you want an extension).