@_Guz_i
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Research Scientist @Spotify · Working with IR, RecSys, NLP · PhD from @tudelft · ex @AmazonScience · https://nitter.cf/t.co/SMu8BlyfIb
Holanda (Países Baixos)
Joined January 2009
- Tweets650
- Following561
- Followers810
- Likes9.8K
Pinned Tweet
We wrote a post summarizing our #RecSys2024 paper on bridging search and recommendation with generative retrieval 🧵 (1/N)
research.atspotify.com/2024/…
w. @AliVardasbi, @denadai2, @enricopalumbo91, Hugues Bouchard
Gustavo Penha retweeted
For anyone worried their LLM might be making stuff up, we made a budget‐friendly truth serum (semantic entropy + Bayesian). See for yourself: youtube.com/watch?v=x_8ORGLD…
Paper: arxiv.org/pdf/2504.03579
Gustavo Penha retweeted
I doubt to what extent improvements on these datasets would translate to improvements in today's real-world recommendation settings. Reference: arxiv.org/abs/2508.19399v1
Happy to share our #recsys25 paper: “Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge”.
🧠 90 days of listening → natural-language user profiles → LLM judges alignment
📊 Aligns with human eval.
With amazing Spotify co-authors.
📄 arxiv.org/abs/2508.08777
Excited to share our paper “Semantic IDs for Joint Generative Search & Recommendation” @ RecSys'25
🧠 Jointly fine-tuning embeddings for both tasks → shared Semantic IDs that work for search and recs ⚖️
📦 No more task-specific trade-offs!
Gustavo Penha retweeted
Semantic IDs for Joint Generative Search and Recommendation
@_Guz_ et al. at Spotify introduce a bi-encoder model fine-tuned on both search and recommendation tasks to obtain item embeddings, followed by construction of unified Semantic ID space.
📝arxiv.org/abs/2508.10478
Gustavo Penha retweeted
Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge
Spotify introduces a profile-aware LLM framework for evaluating personalized podcast recommendations using natural-language user profiles distilled from listening history.
📝arxiv.org/abs/2508.08777
Gustavo Penha retweeted
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
@denadai2 et al. at Spotify use multimodal LLMs to generate natural-language descriptions of video content for better recommendations
📝arxiv.org/abs/2508.09789
👨🏽💻huggingface.co/datasets/marc…
Gustavo Penha retweeted
What if we could use off-the-shelf Multimodal Large Language Model to enrich current video recommendation models?
This is what we asked ourselves in our recent #recsys2025 paper arxiv.org/pdf/2508.09789
🧵
🔎 LLM alignment techniques can enhance query expansion by eliminating the need for multiple generations followed by re-ranking/filtering steps.
Check out this work led by @adam_x_yang during his internship with us at @SpotifyResearch w. @enricopalumbo91 and Hugues Bouchard⬇️
Gustavo Penha retweeted
Adaptive Repetition for Mitigating Position Bias in LLM-Based Ranking
Spotify introduces a dynamic early-stopping method that adaptively determines repetitions needed for each ranking instance, reducing LLM calls by 81% while preserving accuracy.
📝arxiv.org/abs/2507.17788
Gustavo Penha retweeted
Aligned Query Expansion: Efficient Query Expansion for Information Retrieval through LLM Alignment
@adam_x_yang et al. leverage LLM alignment techniques to fine-tune models for generating query expansions that directly optimize retrieval effectiveness.
📝arxiv.org/abs/2507.11042
Gustavo Penha retweeted
Contextualizing Spotify's Audiobook List Recommendations with Descriptive Shelves
Spotify introduces a pipeline that generates personalized audiobook recommendations with descriptive shelves to help users explore content based on their interests.
📝arxiv.org/abs/2504.13572
We just published this blog post about our research on music track search with generative retrieval. 🧵
With @enricopalumbo91 @adamianou @peputo Timothy Christopher, Alice Wang, Hugues Bouchard, @mounialalmas
The best-performing ID strategy was to use collaborative-filtering embeddings as input to the discretization approach for semantic IDs
Blog post: research.atspotify.com/2025/…
Paper: arxiv.org/pdf/2503.24193
I am attending #ECIR25 at Lucca 🇮🇹 if you are interested and want to discuss this position!
We have an open research scientist position in our lab at Spotify, Personalization ! The areas of expertise are: Information Retrieval, Recommendation System, Language Technologies, Foundational Models, Generative AI Technologies, and Machine Learning.
lifeatspotify.com/jobs/resea…
We have an open research scientist position in our lab at Spotify, Personalization ! The areas of expertise are: Information Retrieval, Recommendation System, Language Technologies, Foundational Models, Generative AI Technologies, and Machine Learning.
lifeatspotify.com/jobs/resea…