@XChe16

Scientist | Pave the way to scientific AGI

Niskayuna, NY
Joined January 2020
XChe retweeted
You can now generate an entire 3blue1brown style video from any research paper with Opus 5.5. Here’s a 8min video summary of “Regularized Recursive Self Improvement of Agent Harnesses”. The 90%ile educational YouTuber is fully automated.
175
394
68
5,554
327,785
A small molecule turns an onco-protein into a polymer and thereby triggers its destruction. This striking route to targeted protein degradation is exemplified by BI-3802. It binds transcription factor B cell lymphoma 6 (BCL6) and induces its reversible assembly into filaments. The resulting supramolecular structure promotes ubiquitination by the SIAH1 E3 ligase and proteasomal degradation of BCL6. Cryo-EM revealed how the drug becomes part of the protein–protein interface that drives polymerization. 1/
3
38
2
176
10,408
Google Brain founder, Andrew Ng: "100% of my tasks are done by ai agents, self-improving loops are next. Give it 3-6 months and prompting is gone." 31 minutes of clear explanation on building self-improving agents from scratch. Worth more than any $500 agentic course. Watch it, then read the full guide on loops below.
25
76
1
421
73,145
XChe retweeted
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5… API: docs.z.ai/guides/llm/glm-5.3… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai
1,003
2,618
2,011
23,862
6,937,325
XChe retweeted
🔥 高德团队刚开源了一个狠东西,能让 Claude Code、Codex 这些 AI 接了活之后,连着干几十个小时不跑偏。 论文挂在 arXiv 上,还上过 Hugging Face 每周论文榜的第一名。 用 AI 干过长活的都懂那种崩溃:你让它把一个东西从头做完,前两个小时挺像样,越往后越离谱,最后甩你一句「已完成」,点开一看一半是空的。它不是不努力,是干着干着把自己一开始要干嘛给忘了。 LongHorizon-Harness 的解法很像一个靠谱的小团队:一个人管进度,只负责想下一步做什么;一个人埋头干这一步,每次都是清清爽爽的脑子上工;还有一个人专门当验收员,不听干活的自己汇报,直接去电脑上翻文件、点界面、看日志、跑测试,对得上才算数,对不上就把证据记下来重做。 同一个模型、同一个工具,只是套上这层壳,官方测下来那种要跨软件折腾一两个小时的任务,能做完的从一半左右提到了八成;命令行和写代码那类活儿也稳了一截,而且 token 还比原来少烧 24%。干得更多,花得更少,这个是真少见。 它管的也不只是敲代码。浏览器、表格、文档、设计软件、3D 软件都能上手,一个任务可以从浏览器开始,中间转到命令行处理数据,再回到桌面软件里出成品,最后回终端跑一遍验证。 而且它不挑模型也不挑工具,Claude、GPT、Qwen 都能接,Claude Code、Codex CLI、OpenCode 这些现成的直接就能用,管进度的、干活的、验收的还能各用各的模型。 GitHub:github.com/AMAP-ML/LongHoriz… 项目团队在这儿 @loopcua
32
134
3
610
50,057
Can molecular glues be designed rather than discovered by chance? We built EvoBind-multimer to design small (6-10 AAs) cyclic peptide glues directly from protein sequences. We validate de novo VHL-KRAS and VHL-BRD4 glues in cells and patient-derived tumoroids.
3
44
1
224
13,141
XChe retweeted
The lead engineer at Anthropic who replaced 300 people with one agent and makes $2.2M - came to Stanford to share how: 03:20 - how she fired 300 people and replaced them with one agent 11:34 - how an agent replaces an entire department and works 10x faster 29:47 - $2.2M a year from one agent - the full system from scratch after watching I replaced my team with one agent - and saved $20k in the first month. Save & watch - the article below shows how one agent replaces the Anthropic team and generates $2.2M a year.
19
69
1
443
84,814
Tomorrow will be my last day at Google after 27 years, and watching it grow from 25 people to 190,000+ has been an amazing journey. Below is a note I shared with many people internally at Google today. An excerpt is: It has been an absolute pleasure to work with you and to help build some of the most widely used and impactful products of all time. As a kid, I dreamed of helping build software that would be used by many people, and Google now has thirteen products used by more than a billion people (amazing!). Our work has had a tremendous impact in the world, and I have been lucky enough to collaborate and form friendships with many colleagues that I deeply admire, respect, and enjoy. It still brings me joy every time I see people out in the world using our products to find information, handle email, translate documents, watch videos, learn new things, navigate and understand the physical world, browse the web, use their phone, run large-scale computations on our infrastructure, ride in an autonomous vehicle, or perform complex tasks with the help of our AI systems. I hope you all share this sense of joy, because it is a shared accomplishment! Thank you to all of my colleagues at Google over many years! Now I'm excited to go start @DiscoLoopAI with my longtime friends and colleagues @Sanjay_Ghemawat, @OriolVinyalsML, and @quocleix. (Updated post: slightly redacted to not have some personal info)
848
2,767
757
33,753
11,891,813
I had so much fun creating the Self-Improving AI Agents course with @achowdhery, and teaching it twice in one year at Stanford! We also collaborated with Stanford Online to make the course available online: YouTube: youtu.be/6YnLB0XbTnI?si=MVwR… The field is moving incredibly fast, but we tried to focus on the core concepts that help us build better AI systems. I hope you enjoy it!
41
119
6
1,142
107,019
This paper is f*cking insane. Prompt Engineering just got replaced by Graph Engineering. A new 20-page paper formalized what the best AI builders are already doing: Stop writing one giant prompt. Build a graph of agents instead. Planner → specialists → verifier → feedback loop. The authors test this definition against LangGraph, DSPy, AutoGen, CrewAI, Prompt Flow, and Claude Code subagents. The prompt is no longer the system. The graph around it is. Bookmark this, then read the full Graph Engineering guide below.
43
171
9
1,020
196,958
🚀 OpenRSI is a new open research series from @FrontisAI for concrete, testable progress toward recursive self-improvement (RSI). As its first project—and also my first work as first author—I’m proud to present OpenMLE: an open full-stack AI4AI system for autoresearch, where evolutionary agents improve ML solutions through executable feedback. OpenMLE has three components: - OpenMLE-Gym: 5,758 executable tasks + evaluators - OpenMLE-ERL: execution-grounded SFT + RL - OpenMLE-Evo: experience-guided long-horizon search 🏆 The full system—our trained Frontis-MA1-35B model paired with OpenMLE-Evo-Max—reaches 71.21% Medal Average on MLE-Bench Lite: surpassing GPT-5.5 + Codex (68.18%) and just 1.52% from GPT-5.6 Sol + Codex and the 2.8T Kimi K3 + Claude Code (72.73%). Budget: 12 hours/task on one RTX 4090 capped at 12 GB VRAM. 🌍 On 10 held-out NatureBench Lite tasks, both components transfer: • same framework, model swap: Match-SOTA 50% → 70% • same base model, framework swap: Match-SOTA 20% → 50% 🔓 Paper, code, models, data, and analysis below. 🧵
11
72
8
335
54,256
XChe retweeted
Google just released free 2-hour course on full Graph engineering: 1 prompt → 100 agents → loops → graphs from 0% to 100%: 10% → 17:44 - build your first agent 30% → 39:30 - Loop engineering: iterate, check, break 60% → 1:12:38 - Graph engineering 75% → 1:34:26 - agents that throttle themselves 100% → 1:55:05 - full graph for multi-agentic systems everyone builds one agent and calls it done - this is the full system where agents wire themselves into a graph watch the course, build the graph - then read the full architecture below ↓
55
678
19
3,525
467,679
Anthropic engineer: “90% of our engineers were building self-improving loops. now we’ve moved to agentic graphs.” “we’re not prompting anymore.” In just 10 minutes, she walks through her entire Claude Code workflow live starting from a blank terminal. this is the kind of knowledge people charge $500+ for. watch the video first, then save the article below to learn how to become a graph engineer. 👇
41
111
7
879
224,597
An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science. We believe it will be a major step for scientific reasoning. openai.com/index/ten-advance…
10 proofs from our next major model Astra on long-standing open problems in mathematics and theoretical computer science (also including new circuit lower bounds for computing the permanent!) GPT-5.6 has already enabled so much exciting work in math and science. Can’t wait to see what comes next!
840
2,531
1,516
17,528
13,347,879
🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇 🔷 The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex! Check out the configuration details in our official API docs: api-docs.deepseek.com/quick_…
1,700
3,512
2,168
29,962
9,582,156
Raygun 1, our model for redesigning proteins the way evolution does, is out today in @Nature. For an arbitrary protein, it can: - miniaturize (eGFP, shorter than any natural FP) - magnify (EGF, into a tighter EGFR binder) - modify (dozens of edits at once-- indels and subs)
16
102
6
453
34,776
I spent 48 hours with the Kimi K3 modeling code. It took: - 650 mg of caffeine (mandatory) - 40 cans of LaCroix (optional... world record (?)) - 8 papers - 6 months off my lifespan Finally grokked the entire lineage of Kimi K3 and how we got here... every single step, since 2019 GPT-2
247
717
83
9,692
2,074,458
Is physics necessary for building foundation models for chemistry or data is all you need? Our Orbitall foundation model uses 35× less molecular training data and is a 50× smaller model than frontier UMA model but outperforms and is 100x faster and accurate in reactions with solvents. arxiv.org/abs/2507.03853 Unlike language and image models, generating training data of larger chemical systems through density functional theory (DFT) is incredibly expensive. Instead Orbitall trains on smaller systems and incorporates physics as: (1) input orbital features using cheaper semi-empirical calculations (2) symmetries. Orbital features natively incorporate spin and charge allowing us to model open shell systems cleanly and use cheaper implicit solvation methods that UMA is unable to. This allows it to beat Meta's UMA model in chemically important solvent-based reactions and be 100x faster. OrbitAll accurately reproduces solvent-dependent reaction energetics and transition-state structures. Orbitall is a molecular foundation model with physics-grounded representations that handle charges and spin states and environmental effects such as solvation and external electric fields in an elegant manner. Foundation models for science are not built blindly on data and certainly not just with LLMs. Physics is key!
7
44
2
242
14,064
The Kimi K3 architecture figure for yesterday's big open-weight model release, along with some observations and thoughts. 1. Yes, it looks relatively complicated, but it's essentially a scaled-up production version of their Kimi Linear model they released last year (scaled up from 48B -> 2.8T; K3 is by far the biggest open-weight model right now) 2. The one new component compared to Kimi Linear is the LatentMoE. I omitted it in the figure below since it's already very crowded, but that's essentially the same LatentMoE as in Nemotron 3 Ultra (you can find it in my LLM Architecture Gallery if you are curious). The idea here is to compress (down-project) large linear layers similar to multi-head latent attention. 3. Kimi K3's overall trend (similar to Nemotron 3, DeepSeek V4, and others) is also towards better inference efficiency. That is, there are many components that replace existing components with efficiency-tweaked versions. I.e., MoE -> LatentMoE, regular attention -> multi-head latent attention and Kimi Delta Attention. (I also have short tutorials and write-ups in my gallery if you are curious about additional details). 4. The one component change that is not an efficiency tweak is attention residuals. Like DeepSeek V4 improved the residual path with mHC (manifold-constrained Hyper-Connections), attention residuals are a way to improve the residual path, but it works a bit differently. I.e., mHC made the residual path wider. Attention residuals (also already part of Kimi Linear) connect the residuals across layers; the connection itself uses an attention score for an important/contribution weight. According to the report, it improves the validation loss and downstream performance (a bit) consistently and adds about 4% in training cost and 2% in inference cost. 5. Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead. (Again, this is inherited from Kimi Linear). In other architectures, the recent trend was towards RoPE in local attention layers (like sliding window attention) and NoPE in the global layers. There were a few architectures that only used NoPE everywhere, but this is the first frontier-level one as far as I know. 6. Kimi K3 now also has native multimodal support, which is great! There are several other interesting training tidbits in the technical report, but that's it from the architecture front so far. A really great release overall.
106
618
30
4,089
238,649