See you in Bay Area! Ant Open Source and inclusionAI are hosting a happy hour for AI builders! Expect sharp technical conversations, food, drinks, and plenty of time to meet the community. Join us during #OpenSourceAIWeek 2026: luma.com/xs9sn0lp
Can we trust what the AI benchmarks says? Can an agent system be smarter than its best model?
Join inclusionAI at #OpenSourceAIWeek 2026 for tech talks, happy-hour drinks, food, and Token Shots 🍸
Palo Alto · RSVP: luma.com/xs9sn0lp
There's a new version of this post
📰Latest updates: welcome the new members of the @AntLingAGI family, meet Ming-Image-0.1-Design!
1️⃣ Ming-Image-0.1-Design: a 6B text-to-image model for UI, infographics, posters, and other text-rich visual designs
huggingface.co/inclusionAI/M…
2️⃣ Ming-Image-0.1-Design-Layer: decomposes a flattened design image into a requested number of RGBA layers using an image and a layer plan
huggingface.co/inclusionAI/M…
➕ and 2 Agent Skills for visual coding and presentations: the Ling UI Design Skill and the Image-to-Editable-PPT Skill.
🤗💗Explore and discuss with us! Join the Ant Ling community on Discord: discord.com/invite/GNaQc8WC5…
#inclusionAI #opensource #AgentSkills #multimodal
We’re open-sourcing the Ming-Image-0.1-Design family:
• Ming-Image-0.1-Design, 6B
• Ming-Image-0.1-Design-Layer, 6B
• Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill
Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard. 🧵
Three quantized versions of @AntLingAGI Ling-3.0-flash-Fin are now available: FP8, FP4 and INT4.
If you’re building real-world financial workflows, you can choose the version that best fits your infrastructure, memory and efficiency needs. 🧵
Now available on HuggingFace and ModelScope:
Hugging Face:
🔗huggingface.co/inclusionAI/L…
🔗huggingface.co/inclusionAI/L…
🔗huggingface.co/inclusionAI/L…
ModelScope: 🔗modelscope.cn/models/inclusi…
🔗modelscope.cn/models/inclusi…
🔗modelscope.cn/models/inclusi…
Ant Open Source retweeted
Today, we are releasing Ling-3.0-flash-VL, built on Ling-3.0-flash with visual understanding and visual agent capabilities. It performs well across visual perception, STEM reasoning, document intelligence, multimodal agent tasks, frontend coding, and medical report interpretation.
👋 LLaDA-Image combines the LLaDA2.0-mini dLLM understanding backbone with a 6B DiT for high-quality image generation and instruction-guided editing. Both backone and Image-Gen are diffusion models, trained in a unified framework.
🔑 Image-first training at scale:
Of ~220M cumulative generation-training samples, >90% use image-only supervision. Image-text pairs are introduced later for language alignment.
The result is a model family that covers high-quality generation, instruction-guided editing, and fast 2-4-step inference with LLaDA-Image-Turbo.
#OpenSource #inclusionAI #dLLM
🔑One model for generation and editing, plus a faster Turbo version:
LLaDA-Image combines a dLLM-based VLM, built on LLaDA2.0-mini, with a single-stream 6B DiT for image generation. In editing, reference-image features bypass the VLM and go directly to the DiT.
Generation and editing share one backbone. TwinFlow further distills the 50-step Base model to 2-4 steps, providing quality-first and speed-first deployment profiles within the same model family.
🔑 Benchmark results and open-source resources:
On Qwen-Image-Bench, LLaDA-Image scores 53.53 on the English track and 53.38 on the Chinese track, advance rankings on both among the open-source models listed in the technical report.
🔧Try out and explore LLaDA-Image with models, code, and technical report:
🤗 Hugging Face: huggingface.co/collections/i…
💻 Code: github.com/inclusionAI/LLaDA…
📄 Technical Report: arxiv.org/pdf/2609.03796
Given human demo videos, it infers objects, interactions & task flows to generate robot‑executable motions.
Check out the latest ICL work from Robbyant Research.
Can robots learn new tasks directly from human‑operation videos?✅
Meet Zero‑WAM: video‑based in‑context learning for robotics.
No heavy retraining, no verbose text instructions.
Given human demo videos, it infers objects, interactions & task flows to generate robot‑executable motions.
- HumanGen dataset: 8.6K tasks / 74.2K human‑robot samples
- IFP training for long‑range video task understanding
- 46.95% avg success rate on 7 unseen RoboTwin 2.0 tasks
Paper out now, code releasing soon.
📄 arxiv.org/abs/2608.26103
🌐 robbyant-research.github.io/…
Sharing a new collaboration with the #SGLang community. lmsys.org/blog/2026-08-28-in…
Infer-forge uses Harness + Loop + Graph to keep 28-hour median task lifetimes fully traceable. Peak tasks in flight: 2→9. A full DeepSeek-V4-Pro serving project split into 38 pieces, each verified like a human engineer. And the harness catches kernel silent corruption on its own.
🚀 Infer-forge: Harness, Loop, and Graph Engineering Around SGLang
@ant_oss built a three-layer system for running long inference optimization work through agents without losing provenance: Harness for execution, Task Loop for one Task's contract, Task Graph for verified Handoffs.
Highlights:
- Peak Tasks in flight rose from 2 to 9, median Task lifetime from 10h to 28h
- The agent ran a full serving project on DeepSeek-V4-Pro by itself. Splitting the work into 38 pieces and verifying each one like a human engineer.
- The harness catches kernel silent corruption on its own.
Read the full blog below 👇
Meet Ling-3.0-flash-dspark➡️🤗We adapted the public DSpark recipe to Ling with distribution aligned data, architecture ablations, and an acceptance aware loss.
The checkpoint and serving recipe are now available. 🔗Full engineering details: lmsys.org/blog/2026-08-21-li…
🔗Hugging Face: huggingface.co/inclusionAI/L…
🔗ModelScope: modelscope.cn/models/inclusi…
Ant Open Source retweeted
🧵 We’ve open-sourced 6 Base Model checkpoints for Ling-3.0-tiny & Ling-3.0-flash, covering pre-trained, mid-trained, and WSM-merged stages. None has undergone post-training, giving researchers flexible starting points for continued pre-training, fine-tuning, and further research.
Two key highlights:
- We use WSM to replace LR decay with weighted checkpoint merging, making the training process better suited for continual pre-training while enabling offline exploration of different LR decay strategies.
- With one shared training recipe, the community can validate strategies on tiny-base, then scale them to flash-base.
Had a 💥blast building this with the SGLang team, the collaboration shows what scenario-specific optimization open-source can do.
The Humming + DSpark optimizations are already upstreamed. 📖 Full breakdown: lmsys.org/blog/2026-08-19-de…
🚀 New blog: Pushing the Limits of Serving DeepSeek-V4-Pro
DeepSeek-V4-Pro (1.6T MoE) on H20 reaches 271 output tokens/s at batch size 1, just 1.42× off B300 on hardware with no native FP4 Tensor Cores.
Together with @ant_oss, we built a scenario-specific serving stack on SGLang:
- 74.8%–78.0% peak TPOT reduction at batch size 1 from optimized DSpark
- 1M-token prefill in 43.7s, 36.5% geomean prefill throughput gain
- 10.14× full-token KV capacity from Humming MXFP4AFP8 + Online C128
- 2.20× per-GPU decode throughput at 4K (319.9 to 703.2 tok/s/GPU)
Happy 1st Open-Source Anniversary to the JVM JIT compiler Jeandle (github.com/jeandle)! 🎉
In the past year, we didn't just give Jeandle full-fledged Java compilation capabilities; we also rolled out a bunch of key optimizations—like Inlining, OSR, Partial Escape Analysis, and Type Analysis—giving Jeandle a massive performance boost! 🚀
🤗A huge shoutout and thank you to all of our contributors for the amazing work!
In the future, we’ll be bringing more optimizations, including Profile-Guided Optimization (PGO), intrinsics, powerful inlining algorithms, and enhanced loop optimization capabilities. We're committed to taking Jeandle's performance to the next level! ⚡️Stay tuned❤️️
🤔Can a lightweight model not only run locally, but also train locally?
With AReno @ARenoTeam, we post-trained @AntLingAGI Ling-3.0-tiny on DGX Spark using an Agentic RL tic-tac-toe task.
The model learned from tool calls, environment feedback, and rewards:
✅rewards_mean: ~-0.5 → ~0.4
✅response_len dropped to ~850 tokens
✅behavior became more stable after training
#inclusionAI #OpenSource #ReinforcementLearning
From running, to training, to adapting to your own task. Check out Areno github.com/inclusionAI/AReno!
Introducing the new member of the @AntLingAGI family, 👏Ling-3.0-tiny, delivering balanced coverage across reasoning, coding agents, instruction following and long-context tasks.
📣Now open weights on Hugging Face & ModelScope
🤗Hugging Face weights:
🔗BF16: huggingface.co/inclusionAI/L…
🔗FP8: huggingface.co/inclusionAI/L…
🔗INT4: huggingface.co/inclusionAI/L…
⚙️ModelScope weights:
🔗BF16: modelscope.cn/models/inclusi…
🔗FP8: modelscope.cn/models/inclusi…
🔗INT4: modelscope.cn/models/inclusi…
After more than a year‑long research journey spanning LLaDA‑MoE to newly‑launched LLaDA2.2, LLaDA series lead Jake Zhao shares his technical insights and forward‑looking outlooks for dLLM.
#llada #dllm #opensource
Some Theoretical and Practical Thoughts on Diffusion Language Models
jzhao2024.github.io/notes/20…