@vitaestfes3873i
iAccount based inKorea
About this account
- Account based in
- Korea
- Connected via
- Korea App Store
Account-level information from X, not a live location or the device used for a specific post.
Joined July 2024
- Tweets813
- Following269
- Followers9
- Likes55
苹果在 AI 领域憋了个真·大招。很多人以为苹果只在端侧修修补补,其实他们刚发了一篇直击目前 AI Agent 致命软肋的重磅论文(arXiv:2609.32391)。
这篇被严重低估的论文到底讲了什么?帮你划 4 个炸裂重点:
1. 直戳痛点:现在的 AI Agent 都有“严重健忘症”
现在的 AI 助手基本上是单次会话的“短命鬼”。现实中一个真正好用的个人助理,需要跨几天、几星期记住你的习惯、定时整理记忆(Memory Consolidation)、跑后台 Cron 任务。目前业界的 Benchmark 几乎测不出这种能力。
2. 黑科技“时钟”:把 1 个月的真实任务压缩到几个小时跑完
苹果搞出了一个叫 SCLATE 的底座,最绝的是它的「混合模拟时钟」。AI 在操作思考时走真实物理时间;一旦进入等待、跨天或休眠,时间直接“快进跳过”。原本需要真实跑一个月的复杂长周期测试,几个小时就能无损跑完。
3. 打脸行业共识:外挂记忆库不一定比自带记忆强
苹果横评了 Claude Code、Hermes、Codex 等主流架构以及 Mem0 等主流外挂记忆组件,得出一个反常识结论:强行加装外部记忆库,经常还不如系统原生记忆。模型和记忆系统必须深度绑定、联合微调,否则纯粹是负优化。
4. 后训练逆天提效:4B 小模型暴打大任务
通过这套系统对 Qwen3.5-4B 做了长程后训练后,模型不仅学会了怎么写高质量记忆,SWE-bench Verified 通过率直接暴涨 16.7 个百分点,读代码行数甚至直接骤降 6.8 倍。
真正的下一代 Siri / Apple Intelligence 显然不是简单的聊天框,而是在设备后台默默长线运行、懂你习惯的持续学习代理。
苹果的底层基建已经搭完了。你觉得未来个人 AI 是拼大模型参数,还是拼这种“跨天跨周的长程记忆力”?
vitaestfestum retweeted
Sol 6.1 is a great model and an incredible value. Is it as good as Opus 5.5?
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
vitaestfestum retweeted
designeer.xyz wtf it's free??? how? what a tool
vitaestfestum retweeted
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions.
CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks.
With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks.
We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed.
Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size.
📄 Blog: contrastive-lm.notion.site
💻 Code: github.com/Contrastive-LM/CL…
🗣️ Discord: discord.gg/5dAQEDJBs
🤗 Data & Models: huggingface.co/Contrastive-L…
More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
vitaestfestum retweeted
Rumors I’ve been hearing, not here on X.
First, let’s start with OpenAI and I’ll go towards Anthropic.
GPT-6 Sol is coming Tuesday. It’s both cheaper and more intelligent than 6 Astra, think of it like a 6.2 jump. The internal model “significantly more capable than Astra,” named Bel internally, helped with this release. Bel is considered “AGI” within OpenAI. They are very impressed with this model.
OpenAI is growing very confident that their internal lead is so big that no other lab can catch up. Unbelievably confident.
Anthropic is currently not in, let’s say, a “code red,” but is aware of OpenAI’s lead and doing everything in their power to catch up. Their new model Opus 5.5 is coming probably Monday rather than Tuesday due to OpenAI releasing on Tuesday.
vitaestfestum retweeted
stock market 2027
Because of huge digital workforces, even better AI video models with much longer outputs, AI able to create AAA games (billions of tokens needed for one output), 3D AI, Virtual Reality/world building AI, AI creating complex software, super math AI, physics AI, biology AI, and generally science AI, token usage will explode next year.
Demand for compute, data centers, memory, and data center infrastructure will be ginormous.
The stock market (AI related stocks) will see a bull run like never before in history.
2027 - economy will get stranger
As soon as next year, we will almost certainly see a boom in extremely lean startups where small teams of just 10-50 people orchestrate 10 000-50 000 AI agents (digital workforce) and generate the kind of output that previously required medium size corporations.
This will emerge first in knowledge intensive industries: software, research, finance, design, media, biotech, legal work, consulting, and other fields where most of the value comes from work done mainly on computers rather than through physical activity.
Company size, headcount, and productive capacity begin to decouple.
And once that model works, it will spread incredibly fast.
vitaestfestum retweeted
2027 - economy will get stranger
As soon as next year, we will almost certainly see a boom in extremely lean startups where small teams of just 10-50 people orchestrate 10 000-50 000 AI agents (digital workforce) and generate the kind of output that previously required medium size corporations.
This will emerge first in knowledge intensive industries: software, research, finance, design, media, biotech, legal work, consulting, and other fields where most of the value comes from work done mainly on computers rather than through physical activity.
Company size, headcount, and productive capacity begin to decouple.
And once that model works, it will spread incredibly fast.
vitaestfestum retweeted
The United States has been leaping from one moral panic to the next: Black Lives Matter, MeToo, George Floyd, and now data centers plus the specter of AI-driven human extinction. In every case a genuine problem exists, but the threat is inflated out of all proportion and the costs of overreaction are overlooked. The resulting hysteria blocks clear thinking and closes off rational discussion. Consequently, individuals never notice that self-appointed experts are distorting the facts for their own gain.
Companies with billions in AI stakes are perfectly capable of staging “rogue LLM” incidents to stampede politicians into regulation that keeps out competitors or simply reduces their own liability.
Previous panics have already damaged trust in institutions and among ourselves. They have also intensified the political polarization that now endangers American democracy. We should resist the current panic rather than feed it. The costs of getting this one wrong dwarf those of the last several.
vitaestfestum retweeted
I'm not worried about evil AI.
I'm very worried about evil humans using AI.
The only way to defend against it, is for good humans to use the most advanced AI, and improve and accelerate it as fast as possible.
vitaestfestum retweeted
We have to defeat doomerism. Rage, rage against the dying of the light. I stand on the side of a humanity armed with strong AI.
vitaestfestum retweeted
I disagree with the doomer position: I believe that we need to build powerful AI to solve the problems our civilization is facing, and that we will be able to control it. Also, most of the doomers are sincere, intelligent and have the best intentions, I regard many as friends.
vitaestfestum retweeted
Replying to @haider1
The answer isn't slowing down AI, it's collapsing the fake dreams of a few losers running US AI labs who are willing to salt the mathematical earth if it gives their IPOs a 2% higher chance of succeeding. Pop the bubble and this fixes itself.
vitaestfestum retweeted
🚨 OBAMA ON AI: "If we are thinking about AI just in terms of how do we cure cancer or get better energy, you can do that without having agentic AI and having it just roaming free in the internet. The reason you are doing that is because you have to market a product that people will pay money for.
That’s a misalignment between what our society needs and the commercial imperatives that these companies are facing, not because necessarily they’re trying to do bad things, but because they’ve got to justify these valuations.
So, that’s one more reason why it is really important for us to have a competent government and a serious bipartisan conversation around this issue, and we have to do it fast. And I would encourage voters to pay attention to this. If somebody does not have a serious plan for how to deal with this, then they’re not meeting the moment, and you should probably look for somebody else."
vitaestfestum retweeted
I built a Photoshop-like image editor. It’s called Compositor.
robbietilton.com/compositor
I originally built it for myself to get off my Adobe subscription, but decided to release it for free and make it open source. It has all the essential tools I need for compositing, with none of the BS.
I know this workflow is kind of archaic in 2026, but I’m still using it until AI can actually get pixel-perfect on some of the details.
The entire app is 12MB. Photoshop is 6,455MB on my machine.
vitaestfestum retweeted
After our viral demo comparing Jev to GLiNER for browser use, there was a lot of demand for the local open source browser use so here is the GitHub repo github.com/sahibzada-allahya…
vitaestfestum retweeted
I was skeptical about a single model being well calibrated across many domains without per domain calibration so I tested it with my own open-jev. And it does seem to bear up.
I really think a good small model, tuned and calibrated per domain is ideal here.
vitaestfestum retweeted
All the grifters are completely wrong about Jev’s architecture so I decided I’d release an open-weight version. BUT training takes time, so while we all wait I decided I’d drop the sauce.
archerhume.com/posts/jevs-ar…
vitaestfestum retweeted
Jev 发布没几天,开源社区已经开始疯狂复刻了🔥
最值得推荐的五个模型:
1、Laya 421M:原生决策模型,支持 Mac
2、Decider-2B:最像 Jev,基于 Qwen3.5
3、NanoJev 0.6B:专门的 Decision Head
4、Reflex:Qwen3.5 + Direct Logits
5、System-One 4B:专门做概率校准
如果和我一样是苹果的芯片,我推荐: Laya 和 Reflex
下一步我准备选两个在本地运行 然后测试一下和Jev的差距
Jev 刚发布没几天,开源社区就出现了同款🔥
Decider-2B模型,是基于 Qwen3.5-2B 做了特殊调整
它和 Jev 模型是一样的 只做选择 评分和判断 不是文本类的 LLM 模型
但两者还是有几个明显区别:
1、模型
Jev:闭源 System One Model
Decider:Qwen3.5-2B,约 1.9B 参数,Apache 2.0 开源
2、价格
Jev:$0.042 / 100万输入 Token,输出免费
Decider:本地部署,没有 API Token 费用,只算机器成本
3、速度
Jev:官方约 70–500ms,吞吐 25万 Token/s
Decider:GH200 测试单次约 3–8ms,但硬件环境不同,不能直接横比
4、能力
Jev:重点是 RLCD + 概率校准,产品化更成熟
Decider:同样支持 Choice / Score / Noul,也专门做了 Calibration 训练
5、使用方式
Jev:直接调 API
Decider:自己部署,更适合直接塞进 Pi、Hermes 这类 Agent Harness
如果感兴趣 可以去看一下
模型地址:huggingface.co/Mapika/decide…