All things open source at Ant Group. We aim to bring high caliber infrastructure FinTech OSS to the community.

Joined March 2024
See you in Bay Area! Ant Open Source and inclusionAI are hosting a happy hour for AI builders! Expect sharp technical conversations, food, drinks, and plenty of time to meet the community. Join us during #OpenSourceAIWeek 2026: luma.com/xs9sn0lp
Can we trust what the AI benchmarks says? Can an agent system be smarter than its best model? Join inclusionAI at #OpenSourceAIWeek 2026 for tech talks, happy-hour drinks, food, and Token Shots 🍸 Palo Alto · RSVP: luma.com/xs9sn0lp
There's a new version of this post
1
114
📰Latest updates: welcome the new members of the @AntLingAGI family, meet Ming-Image-0.1-Design! 1️⃣ Ming-Image-0.1-Design: a 6B text-to-image model for UI, infographics, posters, and other text-rich visual designs huggingface.co/inclusionAI/M… 2️⃣ Ming-Image-0.1-Design-Layer: decomposes a flattened design image into a requested number of RGBA layers using an image and a layer plan huggingface.co/inclusionAI/M… ➕ and 2 Agent Skills for visual coding and presentations: the Ling UI Design Skill and the Image-to-Editable-PPT Skill. 🤗💗Explore and discuss with us! Join the Ant Ling community on Discord: discord.com/invite/GNaQc8WC5… #inclusionAI #opensource #AgentSkills #multimodal
We’re open-sourcing the Ming-Image-0.1-Design family: • Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard. 🧵
6
375
Three quantized versions of @AntLingAGI Ling-3.0-flash-Fin are now available: FP8, FP4 and INT4. If you’re building real-world financial workflows, you can choose the version that best fits your infrastructure, memory and efficiency needs. 🧵 Now available on HuggingFace and ModelScope: Hugging Face: 🔗huggingface.co/inclusionAI/L… 🔗huggingface.co/inclusionAI/L… 🔗huggingface.co/inclusionAI/L… ModelScope: 🔗modelscope.cn/models/inclusi… 🔗modelscope.cn/models/inclusi… 🔗modelscope.cn/models/inclusi…
1
233
Ant Open Source retweeted
Today, we are releasing Ling-3.0-flash-VL, built on Ling-3.0-flash with visual understanding and visual agent capabilities. It performs well across visual perception, STEM reasoning, document intelligence, multimodal agent tasks, frontend coding, and medical report interpretation.
84
112
47
607
737,866
👋 LLaDA-Image combines the LLaDA2.0-mini dLLM understanding backbone with a 6B DiT for high-quality image generation and instruction-guided editing. Both backone and Image-Gen are diffusion models, trained in a unified framework. 🔑 Image-first training at scale: Of ~220M cumulative generation-training samples, >90% use image-only supervision. Image-text pairs are introduced later for language alignment. The result is a model family that covers high-quality generation, instruction-guided editing, and fast 2-4-step inference with LLaDA-Image-Turbo. #OpenSource #inclusionAI #dLLM
9
21
526
🔑One model for generation and editing, plus a faster Turbo version: LLaDA-Image combines a dLLM-based VLM, built on LLaDA2.0-mini, with a single-stream 6B DiT for image generation. In editing, reference-image features bypass the VLM and go directly to the DiT. Generation and editing share one backbone. TwinFlow further distills the 50-step Base model to 2-4 steps, providing quality-first and speed-first deployment profiles within the same model family.
5
7
291
🔑 Benchmark results and open-source resources: On Qwen-Image-Bench, LLaDA-Image scores 53.53 on the English track and 53.38 on the Chinese track, advance rankings on both among the open-source models listed in the technical report. 🔧Try out and explore LLaDA-Image with models, code, and technical report: 🤗 Hugging Face: huggingface.co/collections/i… 💻 Code: github.com/inclusionAI/LLaDA… 📄 Technical Report: arxiv.org/pdf/2609.03796
163
Given human demo videos, it infers objects, interactions & task flows to generate robot‑executable motions. Check out the latest ICL work from Robbyant Research.
Can robots learn new tasks directly from human‑operation videos?✅ Meet Zero‑WAM: video‑based in‑context learning for robotics. No heavy retraining, no verbose text instructions. Given human demo videos, it infers objects, interactions & task flows to generate robot‑executable motions. - HumanGen dataset: 8.6K tasks / 74.2K human‑robot samples - IFP training for long‑range video task understanding - 46.95% avg success rate on 7 unseen RoboTwin 2.0 tasks Paper out now, code releasing soon. 📄 arxiv.org/abs/2608.26103 🌐 robbyant-research.github.io/…
3
6
447
Sharing a new collaboration with the #SGLang community. lmsys.org/blog/2026-08-28-in… Infer-forge uses Harness + Loop + Graph to keep 28-hour median task lifetimes fully traceable. Peak tasks in flight: 2→9. A full DeepSeek-V4-Pro serving project split into 38 pieces, each verified like a human engineer. And the harness catches kernel silent corruption on its own.
🚀 Infer-forge: Harness, Loop, and Graph Engineering Around SGLang @ant_oss built a three-layer system for running long inference optimization work through agents without losing provenance: Harness for execution, Task Loop for one Task's contract, Task Graph for verified Handoffs. Highlights: - Peak Tasks in flight rose from 2 to 9, median Task lifetime from 10h to 28h - The agent ran a full serving project on DeepSeek-V4-Pro by itself. Splitting the work into 38 pieces and verifying each one like a human engineer. - The harness catches kernel silent corruption on its own. Read the full blog below 👇
5
499
Meet Ling-3.0-flash-dspark➡️🤗We adapted the public DSpark recipe to Ling with distribution aligned data, architecture ablations, and an acceptance aware loss. The checkpoint and serving recipe are now available. 🔗Full engineering details: lmsys.org/blog/2026-08-21-li… 🔗Hugging Face: huggingface.co/inclusionAI/L… 🔗ModelScope: modelscope.cn/models/inclusi…
Today we are open sourcing Ling-3.0-flash-dspark, a DSpark draft model built specifically for Ling-3.0-flash. On 4 NVIDIA Blackwell GPUs at batch 1, it delivered 1,120 tok/s, 0.78 ms mean TPOT, and an accept length of 9.95 across 1,000 requests. 🧵
7
13
602
Ant Open Source retweeted
🧵 We’ve open-sourced 6 Base Model checkpoints for Ling-3.0-tiny & Ling-3.0-flash, covering pre-trained, mid-trained, and WSM-merged stages. None has undergone post-training, giving researchers flexible starting points for continued pre-training, fine-tuning, and further research. Two key highlights: - We use WSM to replace LR decay with weighted checkpoint merging, making the training process better suited for continual pre-training while enabling offline exploration of different LR decay strategies. - With one shared training recipe, the community can validate strategies on tiny-base, then scale them to flash-base.
22
43
17
351
521,789
Had a 💥blast building this with the SGLang team, the collaboration shows what scenario-specific optimization open-source can do. The Humming + DSpark optimizations are already upstreamed. 📖 Full breakdown: lmsys.org/blog/2026-08-19-de…
🚀 New blog: Pushing the Limits of Serving DeepSeek-V4-Pro DeepSeek-V4-Pro (1.6T MoE) on H20 reaches 271 output tokens/s at batch size 1, just 1.42× off B300 on hardware with no native FP4 Tensor Cores. Together with @ant_oss, we built a scenario-specific serving stack on SGLang: - 74.8%–78.0% peak TPOT reduction at batch size 1 from optimized DSpark - 1M-token prefill in 43.7s, 36.5% geomean prefill throughput gain - 10.14× full-token KV capacity from Humming MXFP4AFP8 + Online C128 - 2.20× per-GPU decode throughput at 4K (319.9 to 703.2 tok/s/GPU)
2
9
586
Happy 1st Open-Source Anniversary to the JVM JIT compiler Jeandle (github.com/jeandle)! 🎉 In the past year, we didn't just give Jeandle full-fledged Java compilation capabilities; we also rolled out a bunch of key optimizations—like Inlining, OSR, Partial Escape Analysis, and Type Analysis—giving Jeandle a massive performance boost! 🚀 🤗A huge shoutout and thank you to all of our contributors for the amazing work! In the future, we’ll be bringing more optimizations, including Profile-Guided Optimization (PGO), intrinsics, powerful inlining algorithms, and enhanced loop optimization capabilities. We're committed to taking Jeandle's performance to the next level! ⚡️Stay tuned❤️️
17
17
361
🤔Can a lightweight model not only run locally, but also train locally? With AReno @ARenoTeam, we post-trained @AntLingAGI Ling-3.0-tiny on DGX Spark using an Agentic RL tic-tac-toe task. The model learned from tool calls, environment feedback, and rewards: ✅rewards_mean: ~-0.5 → ~0.4 ✅response_len dropped to ~850 tokens ✅behavior became more stable after training #inclusionAI #OpenSource #ReinforcementLearning From running, to training, to adapting to your own task. Check out Areno github.com/inclusionAI/AReno!
20
23
387
Introducing the new member of the @AntLingAGI family, 👏Ling-3.0-tiny, delivering balanced coverage across reasoning, coding agents, instruction following and long-context tasks. 📣Now open weights on Hugging Face & ModelScope 🤗Hugging Face weights: 🔗BF16: huggingface.co/inclusionAI/L… 🔗FP8: huggingface.co/inclusionAI/L… 🔗INT4: huggingface.co/inclusionAI/L… ⚙️ModelScope weights: 🔗BF16: modelscope.cn/models/inclusi… 🔗FP8: modelscope.cn/models/inclusi… 🔗INT4: modelscope.cn/models/inclusi…
Ling-3.0-tiny is now available as an open-weight model in BF16, FP8 and INT4. On Artificial Analysis, it scores 25 on the Intelligence Index and 16 on the Agentic Index, with 772 Elo on GDPval-AA v2 and 20.80 on τ³-Banking—built for real task execution. 🧵
12
1
29
906
After more than a year‑long research journey spanning LLaDA‑MoE to newly‑launched LLaDA2.2, LLaDA series lead Jake Zhao shares his technical insights and forward‑looking outlooks for dLLM. #llada #dllm #opensource
Some Theoretical and Practical Thoughts on Diffusion Language Models jzhao2024.github.io/notes/20…
10
2
15
834
Ant Open Source retweeted
Today, we’re releasing the open weights for Ling-3.0-flash. 🎉 Official BF16 and FP8-quantized versions are now available, so you can choose the option that best fits your hardware, performance requirements, and deployment needs.
77
138
77
1,109
823,468