@SFResearchi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
We advance state-of-the-art #AI techniques paving the path for innovative products at @Salesforce. Focus areas: #AIAgents, #EnterpriseAI, #EGI, and #TrustedAI.
Palo Alto, CA
Joined September 2014
- Tweets2K
- Following431
- Followers19.8K
- Likes1.7K
Pinned Tweet
Looking for the cutting-edge of AI research? Follow Salesforce AI Research to see how we're transforming enterprise technology through advanced innovations. From world models to agentic systems, discover the future of AI before it hits the market.
Salesforce AI Research retweeted
📢 Excited to share our latest work from @SFResearch: Learning Generalizable Behaviors for Terminal Agents
🤔 Q1: How does RL actually help terminal agents generalize?
→ Not by teaching new skills. It reshapes behavior.
• Skills (write the regex, call the tool) are already there from pre-training + SFT
• Behaviors (inspect first, verify, stop looping) are what RL shapes, and what predicts success
🔍 Q2: Can you trust your synthetic RL environments?
→ Often not. The reward can be wrong in both directions.
• The cleanest public collection we audited: 35.8% clean
• Reward 1 for copying a leaked answer. Reward 0 for a correct solution
🌊 So we built RIVER
→ Filter defective envs, penalize repetitive loops, then run RL.
• River-8B: best among evaluated open RL-trained 8B models on all 4 terminal benchmarks
• 2B→27B with <30% of the envs: RL gains +106% on Terminal-Bench-Lite, +30% on Terminal-Bench-v2.1
• Same 3.5K-env budget: filtered 19.4 vs. randomly sampled 17.7
✅ Takeaway: audit your verifiers before you scale your envs.
🙌 Led by @YYYao45 and @bo_pang0 , with @xuanphinguyen , @zhao__ding and @JotyShafiq
📄 Paper: arxiv.org/abs/2608.22631
🌐 Website: self-evolving-agents.salesfo…
🧵👇
Salesforce AI Research retweeted
We announced Koa at #DF26 this week. #EnterpriseAI runs on governed, repeatable work, and that takes domain-specific reasoning. Built on @Nvidia's model, @SFResearch trained Koa on 27 years of enterprise workflows. @justjayesh's and my piece for more: sforce.co/4rhI9WA
Announcing Koa: @Salesforce's first CRM reasoning model, built on @NVIDIA Nemotron 🧠
It builds on a portfolio of domain-specialized models developed by @SFResearch and Salesforce Engineering
By orchestrating these purpose-built models alongside frontier LLMs, we route targeted jobs — like intent classification, toxicity screening, and search reranking — to the exact model built for them, while Koa handles multi-step enterprise reasoning.
→ Domain-specialized models: sforce.co/4AfWkPQ
→ Moirai: sforce.co/4raX1G2
→ Press Release: sforce.co/4y6xXCT
#FutureOfAI #EnterpriseAI #AgenticAI
Introducing Koa, built on @NVIDIA Nemotron
Salesforce’s first CRM reasoning model for Agentforce brings 27 years of CRM intelligence into the model itself
→ Matches or exceeds leading model performance on CRM actions with 3x fewer errors
→ Built for complex, multi-step workflows
→ No customer data used to train it
→ Now in pilot
And our work with @NVIDIA goes further — bringing NVIDIA Nemotron-based models + accelerated computing to Missionforce for mission-specific AI.
Read more 👇 sforce.co/4xugjb9
🧵 (1/3) We're heading to #ECCV2026 in Malmö, Sweden. 🇸🇪 This week, our team will present two new papers spanning physical generative reasoning and video question answering. More below ⬇️
(2/3) DreamHouse asks whether vision-language models can build the real world, not just picture it. Our new benchmark grounds physical generative reasoning in timber-frame construction.
arxiv.org/abs/2603.24866
#ECCV2026
(3/3) Evidence-Backed Video QA asks video models to show their work: an answer paired with precise spatio-temporal evidence, tracked masks, not just text. We introduce ST-Evidence, plus a 160k-scale training set.
arxiv.org/abs/2607.11862
#FutureOfAI #EnterpriseAI #ECCV2026
Traditional monitoring tells you what happened last week; #OperationalIntelligence tells you what's happening now and why.
Our EVP & Chief Scientist @silviocinguetta breaks down the shift to continuous, proactive systems that surface what matters before it becomes a problem.
sforce.co/4cquhTH
#FutureOfAI #EnterpriseAI
(1/13) 🎉 We are pleased to announce our participation in #EMNLP2026, the Conference on Empirical Methods in Natural Language Processing, in Budapest, Hungary, October 24–29. 🇭🇺 Our researchers will present 12 accepted papers spanning agentic reasoning, evaluation, and multimodal AI. Full list below ⬇️ @EMNLPmeeting #NLProc
(12/13) UserBench: An Interactive Gym Environment for User-Centric Agents: a gym environment testing whether LLM agents can uncover and align with evolving user preferences, not just complete tasks.
arxiv.org/abs/2507.22034
Authors: @qiancheng1231, @LiuZuxin, @aksh_555, Zhiwei Liu, Jianguo Zhang, Haolin Chen, @hengjinlp, @iscreamnearby, @shelbyh_ai, @silviocinguetta, @CaimingXiong, @huan__wang
#EMNLP2026
(13/13) UserRL: Training Interactive User-Centric Agent via Reinforcement Learning: a unified RL framework with eight gym environments for training agents to better assist users in multi-turn interactions.
arxiv.org/abs/2509.19736
Authors: @qiancheng1231, @LiuZuxin, @aksh_555, @_Jason_Q, Zhiwei Liu, Haolin Chen, @KokaneShirley, @hengjinlp, @iscreamnearby, @shelbyh_ai, @silviocinguetta, @CaimingXiong, @huan__wang
#EMNLP2026