@buboseesi
iAccount based inUkraine
About this account
- Account based in
- Ukraine
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
I see in the dark . I build AI things and show what's really going on inside.
Europe
Joined August 2026
- Tweets398
- Following143
- Followers117
- Likes1.6K
This guide splits every agent into two jobs.
The LLM builds.
Jev decides.
So I made it literal: 300 agents build a fruit fly, part by part.
Then the brain switches on.
139,255 neurons. Zero words. Pure decisions.
That's System One.
Every frame is code.
Throw 3 prompts at Opus 5.5.
In under 30 minutes it directly generates a video.
Scary good.
Now imagine:
generation speed ramps up by several orders of magnitude.
Plug in a brain-computer interface.
Strap on VR.
The world in your mind renders in real time in front of you.
At that point, it won't be watching videos anymore.
It'll be stepping straight into your own imagination.
Opus 5.5 made this whole love story in code.
0 after effects. 0 video models.
Every frame is code.
Meta's Muse Charm is an AI agent on your keychain. It browses the web, fills out forms, books and buys for you. @finkd showed it at Meta Connect this week.
I drew what happens when two of them end up in the same coffee line.
By morning one agent has booked a table for two.
It didn't wake its human to ask.
No video model. Every frame is code.
Meta's Muse Charm is an AI agent on your keychain. It browses the web, fills out forms, books and buys for you. @finkd showed it at Meta Connect this week.
I drew what happens when two of them end up in the same coffee line.
By morning one agent has booked a table for two.
It didn't wake its human to ask.
No video model. Every frame is code.
ChatGPT, Claude, every LLM rereads your whole chat from zero with every message. The context window is all the memory it has.
So I raised one like a tamagotchi.
Fed it. Taught it. Fixed its bad habits.
Forgot one thing: to flush.
At 100% context it shat itself and forgot who I was.
现在的 LLM 训练完后权重都是完全冻结的,用户在对话中提供的最新事实、纠错或长文档
模型只能塞进 上下文窗口里,每次对话都要从头重读一遍,极度烧显存和算力,而且聊完就忘。
Boltzbit 与剑桥大学的团队研究出一种全新的方法:“把知识写进权重,而不是塞进 Prompt”...
通俗易懂理解就是:
现在的 AI(比如 ChatGPT、Claude)就像一个背完字典后被“物理封印”的学霸:
考试结束(模型训练完成)后,它的脑子(模型权重)就彻底锁死了,一个字都不能改。
你跟它聊天、给它发一份很长的资料,它怎么处理?它只能把资料写在手心里、贴在眼前(也就是 Prompt / 上下文窗口)。
致命缺点:
每说一句话,它都要把眼前厚厚的资料从头到尾重读一遍(极其浪费算力、显卡显存狂飙)
资料一旦太长,它眼前贴不下,还会看花眼(注意力稀释、答非所问)
聊完关掉对话框,它立刻忘得一干二净,下次还得重新贴一遍
他们提出不保存一堆现成的专家参数,而是用一个轻量超网络(Hypernetwork),把用户给的实时交互数据动态编译成低秩(LoRA 形式的)权重;再用一套贝叶斯机制让这套权重在对话中随着上下文动态演化。
他们官这种叫“无限参数”:
传统 MoE:在固定的比如 64 或 128 个专家库里做离散路由选择,无论怎么组合,都在一个有限的凸包内。
无限参数 LLM:它根本不存固定专家,而是根据输入数据在一个连续的潜在空间(Latent Space)里按需动态生成专属专家。因为输入的数据和隐空间是连续无限的,生成的权重也就是无边界的。
也就是别再把资料贴在眼前了,直接“临时写进脑子里”!
他们给 AI 配了一个极其轻巧的“临时脑回路生成器”。 当你把新的事实、文章或对话历史丢给它时,生成器会把这些内容直接编译成一层超轻量的临时神经元(LoRA 权重),啪的一下贴在原有的大脑上。
聊天的过程中,如果你补充了新信息或者纠正了它,它脑子里的临时回路不是重来一次,而是像滚雪球一样实时微调,越聊越懂你,但每轮对话消耗的算力却几乎不变。
通过三组核心实验验证了这种方案的实际效果,核心结论是:短文本拼不过直接读 Prompt,但在长文档、多干扰项以及多轮对话中,全面反超传统方案。
准确率:凭借贝叶斯机制,模型对用户上下文的理解像滚雪球一样持续收敛,面对指代不清的代词(如“它的作者是谁?”、“那后来呢?”),准确率随着轮数单调递增,一路爬升到接近满分。
算力成本:每一轮只做固定维度的内积更新,计算成本从第 1 轮到第 100 轮完全恒定。
Claude Opus 5.5 caught 72% of known bugs in Deloitte's code review tests. At its lowest effort setting.
Opus 5 at high effort caught 56%.
But the code Opus 5.5 writes has 44% more concurrency issues than code from Opus 5.
So here's a reviewer prompt aimed at concurrency.
4 passes, flags hidden commands in the code, and opens with a verdict. SHIP, FIX FIRST or BLOCK.
Full prompt in the reply.
Anthropic built Opus 5.5 to talk less.
In Box's tests it used about a third of the tokens the last version needed, and its answers came out 40% shorter with no drop in accuracy.
Less talk, same accuracy. The reviewer prompt below keeps it that way.
Claude Opus 5.5 caught 72% of known bugs in Deloitte's code review tests. At its lowest effort setting.
Opus 5 at high effort caught 56%.
But the code Opus 5.5 writes has 44% more concurrency issues than code from Opus 5.
So here's a reviewer prompt aimed at concurrency.
4 passes, flags hidden commands in the code, and opens with a verdict. SHIP, FIX FIRST or BLOCK.
Full prompt in the reply.
Spotify has 10 years of your data and still recommends the same 20 songs
I took Jev + Claude and built what it couldn't in a decade
16,558 songs. 8 years of listening history. 5 steps. $0.31
Spotify has 3 weeks of memory. this has 8 years and adaptive reasoning
code is open
Anthropic built Opus 5.5 to talk less.
In Box's tests it used about a third of the tokens the last version needed, and its answers came out 40% shorter with no drop in accuracy.
Less talk, same accuracy. The reviewer prompt below keeps it that way.
Claude Opus 5.5 caught 72% of known bugs in Deloitte's code review tests. At its lowest effort setting.
Opus 5 at high effort caught 56%.
But the code Opus 5.5 writes has 44% more concurrency issues than code from Opus 5.
So here's a reviewer prompt aimed at concurrency.
4 passes, flags hidden commands in the code, and opens with a verdict. SHIP, FIX FIRST or BLOCK.
Full prompt in the reply.
Claude Opus 5.5 caught 72% of known bugs in Deloitte's code review tests. At its lowest effort setting.
Opus 5 at high effort caught 56%.
But the code Opus 5.5 writes has 44% more concurrency issues than code from Opus 5.
So here's a reviewer prompt aimed at concurrency.
4 passes, flags hidden commands in the code, and opens with a verdict. SHIP, FIX FIRST or BLOCK.
Full prompt in the reply.
sources: Deloitte numbers via The New Stack, concurrency data from Sonar
thenewstack.io/claude-opus-5…
sonarsource.com/blog/claude-…
Spotify has 10 years of your data and still recommends the same 20 songs
I took Jev + Claude and built what it couldn't in a decade
16,558 songs. 8 years of listening history. 5 steps. $0.31
Spotify has 3 weeks of memory. this has 8 years and adaptive reasoning
code is open
I GAVE CLAUDE A BRAIN THAT DECIDES HOW HARD TO THINK
Not per session. Per step.
Jev is a classifier that scores each move 0 to 1. Claude Opus 5 shifts between LOW, MEDIUM, and HIGH reasoning mid-run.
Reading files. LOW. Twelve steps in a row.
Then it cross-references modules. LOW to MEDIUM.
Final analysis across 1,176 lines. MEDIUM to HIGH.
6 real bugs. 14 steps. $0.38.
160 lines of python, no framework, and one API beta you've never seen in a demo.
script, setup, and the beta flag — all here:
gist.github.com/robertovoyk-…
requires: Claude API key (Opus 5 or Fable 5.1), Jev key from @typesafe_ai, python 3.10+
pip install anthropic requests rich
export ANTHROPIC_API_KEY=your-key
python jev_agent.py "your task here"
Apple just paid $250,000,000 for lying about AI.
Jev scored their 10 official iPhone 18 Pro claims. Every word. Straight from apple.com.
Not a single claim scored above 2.5.
The AI's verdict: average innovation 2.14 out of 5. Worth upgrading? 16%.
The model is Jev by TypeSafe AI — it scores text on a spectrum, no opinions, just probability math.
Apple spent more on the settlement than on the innovation these words describe.