ActiveGraph retweeted
reintroducing untapped capital
new website (untapped.vc) and... music video 🙃🎙️🕺
ActiveGraph retweeted
the design of evaluatorbench.com grew out of a summer project, epistemedia.org (which i still need to write about), and runs on @activegraphai
the pattern across all three projects is that it treats the journey as the primary substrate, and the destination as a projection of it
the independence of AI evaluators is going to matter more... so i made it a "benchmark":
evaluatorbench.com (research preview)
source-linked directory of third-party evaluators of frontier AI. each one has an independence score you can take apart: change the weights, change which evidence counts (standard, against-interest, primary-only). not quality or competence. independence only
and then i compared where we are to other highly regulated industries. one thing that stood out is that in most regulated industries, the regulators monitor the assessors too. if the pattern holds, regulators would oversee METR et al., not only the labs
still a preview, not complete, not citable yet. i burned a stupid amount of agent tokens on it. contributions, remixes, forks, benchmaxxing all welcome 😉
ActiveGraph retweeted
found a new @activegraphai citing in a new paper on graph engineering for llm agents ☺️
arxiv.org/pdf/2608.21156
"Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence"
ActiveGraph retweeted
🚨 STOP RE-RUNNING YOUR WHOLE AGENT TO TEST ONE CHANGE.
someone just open-sourced the fix. free
activegraph. 547 stars. one install.
install it - and your agent gets:
- fork a finished run at any step
- the shared part replays from cache - zero new model calls
- strict replay re-fires every step and fails on the first divergence
- an append-only log: what changed, when, and why
- policies hold an action until a human approves it
setup, 30 seconds:
pip install activegraph
activegraph quickstart
runs on recorded examples. no keys, no config, same output every time
save it before your next long run dies with no log to replay ↓
ActiveGraph retweeted
[new @activegraphai blog post] How Synthetic Players used ActiveGraph to verify 4,919 runs without new model calls
activegraph.ai/blog/syntheti…
For the increasing number of agentic academic papers, reproducibility, etc become valuable in strengthening your argument
This blog post walks through how we used ActiveGraph for this.
there's an increasing amount of research replacing LLMs as human subjects, so I was curious how well this actually works
to do this, i had LLMs play well known games, and compared how they played, compared to real humans
arxiv: arxiv.org/abs/2608.00979
site: yoheinakajima.github.io/synt…
more in thread 👇
[new paper on synthetic players: Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel]
ActiveGraph retweeted
there is a missing link between how humans and agents collaborate today. the atomic unit of focus is a "task". but there's no way to merge tasks together in a way that gives context over the entire "workstream".
@yoheinakajima's @ActiveGraphAI has the right model of an append only event log, but theres need for a different way to do projections.
ive been thinking about how to take an event stream and "fold" it into something that enables better human-agent collaboration.
human - agent collaboration is still pretty sloppy.
- chats and sessions are fine, but they tend to be extremely unorganized.
- artifacts aren't created in a reasonable location, they're not visible
- the ui and framework used is owned by a model company (eg claude cowork)
- data stays in md files sure, but now you need to create a system to run these
- use @tobi's qmd and other things to search
- yet another knowledge mgmt system
- containers, et al
ActiveGraph retweeted
Nine iterations of BabyAGI in three years, and yet the bit that @yoheinakajima kept coming back to was graphs. @aiDotEngineer published where that landed, "Active Graph Agent Runtime (BabyAGI 4)", on YouTube. The legend Yohei is Managing Partner at Untapped Capital.
The talk is a working argument for building an agent around an immutable event log instead of around the LLM, with code, reference agents, and experiment results behind it.
- The log is the agent. One immutable typed event log holds what the agent did and every change to the agent itself, and it projects the graph that is the agent's state.
- Behaviors replace the control loop. They react to graph changes and emit events. LLMs never talk to each other, only to shared state, so replay, rollback, and forking come natively.
- Policies decide what can change. Adding a research source is cheap. Editing a prompt can require a human. A new fact can require that nothing contradicts it.
- Views are context management as a graph query, Handing a behavior the subset of the graph it should see.
- Packs, not skills. Memory, identity, tools, secrets, chat: each bundles object types and behaviors, so you can swap one memory pack for another.
- A runtime, not a harness. He rebuilds ReAct on top of it. On goal created, add a thought. On thought created, run reason.
- The log doubles as memory. On LongMemEval, no fact or entity extraction, just embed the query and pull the messages around the hits. When his API key ran out at question 350, the run picked up at 353 instead of starting over.
- Self-modification with gates. Regimes classifies the failure, lets the agent edit only the matching part of itself, then requires a static check, a sandbox check, and a measured rerun before a patch is accepted. Loops of 8 to 13 accepted 4 or 5, with modest but statistically significant gains.
- It remembers what failed. Around 80 tuning passes on a deterministic Pokemon trading card agent for a Kaggle competition, 20 to 30 accepted, and everything that didn't work stayed on the record.
- Old architecture, new workers. Blackboard and Kafka have decades of writing behind them while LLM agents have three years, which is his hypothesis for why coding agents write this style well.
I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!
ActiveGraph retweeted
my current @activegraphai inspired approach to a modular repo-centric agent operating system
Yohei Nakajima, the creator of BabyAGI, just showed why he stopped building agents around the LLM.
His AIE talk on ActiveGraph, in 5 timestamps:
1:55 – build around the log, not the model
3:24 – behaviors and policies gate what the agent can change
8:12 – his api key died at question 350, the run resumed itself at 353
11:11 – a loop that forks the agent and keeps a patch only if accuracy rises
15:49 – why long-running agents need an experiential world model
The frame: make an immutable event log the agent, and replays, rollbacks and forks come for free.
Different way to think about agents. 17 min well spent.
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Anthropic's Lance Martin just walked through how they build agents that run for hours with no human in the loop.
His talk on async, long-horizon agents in 5 timestamps:
3:19 – split the brain from the hands
6:20 – build agent vs verifier agent in a loop
8:18 – 20 iterations to clear an ML benchmark, unsupervised
13:19 – "dreaming" fixes memories the agent got wrong
16:08 – Claude Tag as an org-level harness, not a slackbot
The point: past the 1-hour mark, architecture (not just model size) is what makes async agents work.
Solid 25 min if you're building anything long-running.
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
ActiveGraph retweeted
having an event stream like activegraph is the base of how we move to next level and that will probably be composition, which can let you achieve better results with smaller models.
- get the core event stream
- wrap things in a nice monad
- now you compose the models:
> main runner's stream
> add on a validator agent that checks specific user interaction and changes/obfuscates data
> parallelized tasks that run multiple agents at same time
> classifiers, mappers, tinkerers, translators
and some can be SOTA models, some can be super fast models, some can be local mini models.
and this way you can build highly complex and observable pipelines, "swarms" or "brains" for whatever you need. much more solid, reproducible and observable than your generic imperative orchestration.
🆕 ActiveGraph: The Log is the Agent
my talk from AI Engineer is live!!! 😆
youtube.com/watch?v=khVX_BUn…
it's about @activegraphai, an event-sourced graph runtime for building durable long-running agents:
- what it is
- code examples
- experiments
- benefits & surprises
please check it out, share with your friends who like to tinker at the frontier, and let me know if you have any Qs!
ActiveGraph retweeted
Implemented Yohei Nakajima's paper on "ActiveGraph : Event-Sourced Reactive Graphs for Auditable, Forkable Agentic Systems"
Repo : github.com/corporatepiyush/a…
I've long felt AI harnesses were trapped by our obsession with code so @yoheinakajima's ActiveGraph video hit me hard. Reminded me of time I obsessing over Linda space & case-based reasoning. Watch the video. It's very different.
youtube.com/watch?v=khVX_BUn…
ActiveGraph retweeted
🆕 ActiveGraph: The Log is the Agent
my talk from AI Engineer is live!!! 😆
youtube.com/watch?v=khVX_BUn…
it's about @activegraphai, an event-sourced graph runtime for building durable long-running agents:
- what it is
- code examples
- experiments
- benefits & surprises
please check it out, share with your friends who like to tinker at the frontier, and let me know if you have any Qs!
ActiveGraph retweeted
your hippocampus fast-captures events (logs) that are computed on to hold a current state, and through selective replays, pulls some of this slowly into the cortex (model), which then feeds new priors back into the computations
ActiveGraph retweeted
continued working on @activegraphai reference packs last weekend, resulted in needing to harden the runtime:
it already could...
- keep a complete history of everything an agent did
- replay that history
- fork into alternate timelines to try different ideas
it can now...
- realize it's missing a capability
- write new code to add that capability
- test it in an isolated copy of itself
- ask a human to approve it
- safely adopt it
- continue running with the new capability
activegraph.ai/blog/activegr…
ActiveGraph retweeted
the other week i read active graph papers from @yoheinakajima and got inspired.
enough to rewire brigade, graphtrail, and miseledger around one habit: state is a projection of a log.
then i got a chance to read this yesterday and the same receipts turned out to be mineable. brigade now finds command sequences operators keep repeating, proposes runbooks from them, and pins each step's binary by sha256.
the whole loop: brigade.tools/blog/activegra…
this is a great approach, seeing this more
@flymy_ai also does this when you build an agent via their api, they'll build a deterministic reusable workflow, except for where you need models
ActiveGraph retweeted
Replying to @yoheinakajima
here it is: brigade.tools/blog/activegra…
ended up as the three that close a loop: graphtrail diffs the code graph, the diff rides into brigade's run receipts, exports to miseledger as content-addressed evidence, and comes back as context for the next run.
one wrinkle: a single event log wasn't enough for the promotion ratchet, it needed the decision receipts as their own transition log.
ActiveGraph retweeted
#TIR
BabyAGI 作者 Yohei Nakajima 的新论文 "The Log is the Agent" 提出了一种非常彻底的架构反转:把 Agent 的状态管理从「改内存」变成「写日志」。
传统 Agent 框架是围着 LLM 转的——先搭对话循环,再加工具、规则,最后 bolt-on 一层日志和向量记忆。日志是副产品,不是真相本身。
ActiveGraph 反过来:append-only event log 是唯一真相来源,工作图只是日志的确定性投影,行为(函数/LLM 调用/边逻辑)响应图的变化并生成新事件。没有 orchestrator,协调完全通过共享图完成。
论文给出了 diligence pack(投研尽职调查)的完整例子:671 个 events、93 个 objects、76 条 relations、103 次 model call、48 次 tool call——零行 orchestration 代码,纯靠 reactive behaviors 的触发链自动完成。
三个核心概念:
事件日志 = 唯一真相所有事情都是事件,一条一条追加。goal.created、object.created、llm.requested/llm.responded、behavior.completed……就像 Git 的 commit history,代码文件只是 commits 的一个投影。ActiveGraph 里,图也只是日志的一个投影。
图是确定性投影给定同样的日志,永远得到同样的图。Replay 时 model/tool 响应被 content-addressed cache 记录下来,直接返回 cache,不产生新调用,结果 byte-reproducible。
行为是反应式的没有主循环说「先做 A 再做 B」。行为只是订阅者:看到图里多了个公司 → 触发 question_generator → 写事件回日志 → 图更新 → 可能触发下一个行为。控制流是「涌现」出来的。
这带来三个传统架构很难同时拥有的特性:
Replay:日志就是全部状态,重放时 byte-for-byte 一致。传统框架几乎不可能,因为状态散落在各处。
Fork:在任意 event 处切一刀,前半段共享(从 cache 读,不花钱不耗时),后半段独立运行。可以做 A/B 测试、对比改进效果。传统框架没有这个概念。
Lineage:每个对象自带 provenance——谁创建的、因为哪个 event、哪次 LLM 调用产生的。从顶层 goal 到每一个 artifact 的完整因果链都可以在日志里重建。传统框架需要额外插桩,还不完整。
论文还讨论了 self-improving agents 的天然适配:规则修改本身就是 event,可以 fork 出来测试改进效果,structural diff 对比结果,共享前缀免费。
这种「软件工程反哺 AI 架构」的思路挺有意思的——event sourcing + CQRS + reactive dataflow 在 agent 领域的应用。
🔗 arxiv.org/abs/2605.21997
🛠️ pip install activegraph | activegraph quickstart