@ken_techi
iAccount based inUnited States!
About this account
- Account based in
- United States
- Connected via
- China Android App
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
Principal Engineer & Chief Architect BIOS, Linux, Cloud Computing, ROS, AIGC Agent, Quantitative Trading
Joined April 2009
- Tweets727
- Following602
- Followers63
- Likes3.1K
Ken retweeted
study notes on speculative decoding
leoniemonigatti.com/blog/spe…
Ken retweeted
I recently gave a talk introducing young economists to post-training.
The slides are now up! kawine.github.io/assets/aies…
Not only is post-training your own model more doable than ever, but economists and other social scientists can play a big part in the next era of post-training.
The post-training suite looks roughly like:
1) SFT
2) Offline Methods (DPO/KTO/SimPO/etc)
3) Online Methods (PPO/GRPO/CISPO/etc)
4) RL Environments
5) Distillation (<--- we are here)
6) World Adaptation (<--- the future)
As agents are deployed in the real world, the real world will adapt to them. This adaptation is largely overlooked by the current post-training paradigm because working with verifiable rewards from static environments is much simpler and scalable. As we saw with the OpenAI-Huggingface incident however, we shouldn't assume that the real world will remain fixed---quite the opposite, in fact. So:
1. How can we anticipate real-world adaptation during post-training?
2. What do equilibria look like in these scenarios, if there are any at all?
3. What tools can be created to scale up human oversight?
Many such questions abound, and if you're an economist, these are high-impact problems worth studying.
Ken retweeted
Ever since RL became a core part of LLM training, train–rollout numerical mismatch has been one of the recurring problems in RL infra.
There have been several great open-source efforts toward alignment, but production-scale support still often comes with tradeoffs in model scale, performance, or feature coverage.
Today, we're open-sourcing the deterministic train–rollout alignment stack we use for GLM-5.2-scale training — covering FP8 weights + FP8 KV rollout, DeepEP, DeepGEMM, and sparse attention.
github.com/THUDM/slime/pull/…
1/3
Ken retweeted
别让 coding agent 把预算耗在翻仓库上。
arxiv.org/abs/2608.05886
CodeGrep用GRPO训练14B检索代理,并行调用grep、glob、read筛候选文件,再交给冻结的编码代理写补丁。SWE-Bench Verified 500题上,解决率25.8%→27.0%;已解决任务轮次降15%、token降19%。精度门槛很现实:BM25的0.375反而添噪,0.677才有净收益;检索应作为独立skill训练。
Ex-Google Jeff Dean just released 1-hour lecture on full AI engineering: LLM → prompts → agent teams → graphs from 0% to 100%:
0% → 1:45 - LLM from scratch - that made Google
30% → 17:22 - how to actually use AI models
65% → 30:03 - prompt engineering
100% → 52:35 - one human coordinating 100 agents
this is 27 years of AI in Google compressed into one hour, by the person who lived it
watch it today - then read how to build graphs in the article below ↓
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Ken retweeted
I implemented Kimi Delta Attention in Excel for this week's special seminar on Kimi 3. RSVP 👉 luma.com/5ylclkn6
Read full blog: lmsys.org/blog/2026-07-29-mx…
Ken retweeted
We are starting Intent Lab, building an autonomous team we call "fleet" that turns intent into production software. Today we are sharing some early results: the fastest GLM5.2 inference engine, one shot database creation, and a fully verified agent filesystem.
intentlab.ai/blog/turn-your-…
Ken retweeted
A historic day in China’s space program!
China’s Long March-10B has successfully completed its maiden flight—and recovered its first stage via a sea-based net. This marks the country’s first-ever controlled rocket recovery. A major leap toward reusable launch capabilities. 🚀🌊🇨🇳
Ken retweeted
目前看到关于 “Agentic Engineering Workflow”的最完整的介绍👇
花了一个小时完整看完了,完全可以做成一个付费教程。
内容涵盖了tmux,agent记忆,skills,语音输入,长任务执行,并行worktree管理,多agent调度。
还有让我眼前一亮的可视化html编辑器Lavish和一套代码变更校验的流水线: no-mistakes
感谢作者@kunchenguid分享,值得每一个用ai agent的人收藏、学习。
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Ken retweeted
TRUTH SOCIAL: NVLink multicast is not supported on Blackwell "Confidential Computing" leading to 61% performance regression on SGLang Qwen3.5 397B according to @verdacloud 's recent github ticket. NVIDIA's "Confidential Computing" is complete slop as in addition Hopper's confidential computing had fully unencrypted NVLink according to NVIDIA's own "NVIDIA Secure AI with Blackwell and Hopper GPUs" Whitepaper.
Ken retweeted
有人刚刚开源了一个神级 OSINT(开源情报)仪表盘。它被称为 Shadowbroker。它将地球上所有的公开军事信号汇聚到一张地图上:航母打击群、间谍卫星、GPS 干扰区、25,000 多艘船只、2,000 多个实时 CCTV 监控画面。在地球上的任何位置点击右键,即可获取完整的情报档案。100% 开源。
Claude Opus 4.7现在能自动为你找工作!🤯
有人开发了一个工具能为你找工作
> 扫描顶级公司的职位空缺
> 自动为你填写表格
> 为每个职位量身定制改写简历
无需猎头,无需发送200份相同简历
100%免费开源
GitHub仓库:github.com/santifer/career-o…
Ken retweeted
I (finally) put together a new LLM Architecture Gallery that collects the architecture figures all in one place!
sebastianraschka.com/llm-arc…