gpt最近的表现实在是太糟糕了,之前一直订阅的。最近换了claude,opus 5.5的表现实在太棒了,有种拨云见日的感觉。by the way,模型是好模型,就是掌管的人确实bad。奥特曼,多多上心吧。
Spent Sunday cutting an overnight agent’s write set from “whole repo” to 6 files. Monday morning: 1 PR, 0 surprise edits. Still too tight, or the right paranoia?
Left a cloud agent on a refactor overnight. Woke up to 14 files touched and 3 tests still red. Now I give it a merge checklist before bed, not a vibe. Anyone else burning mornings on “done” that isn’t?
Debugged a false-positive “you’re absolutely right” loop for ~25 minutes this morning. Agent agreed, then rewrote the same bug. Now I make it open the file before it gets to agree. Anyone else?
This is just nonsense, a piece of formalistic writing.
1. First of all, when it comes to restricting China, fundamentally there is a mentality of technological inequality.
2. You support closed source, but in the end Hugging Face used open source to block your attack.
3. Everything is for profit; this is the foundation of capitalism. Who would go against their own heart? Unless something involves their own interests.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: darioamodei.com/post/we-must…
Closed Cursor overnight with a coordinator still spinning. Came back to 4 subagent threads and one diff I wouldn't merge. Now I stop the coordinator before I stand up. Anyone else?
A skills pack without a denied-tools list is just write access with nicer UX. I won’t import one until it leaves a dry-run diff I’d merge — and a failing-test receipt when it can’t.
Isolation isn’t optional for coding agents on a host. My default is a capped workspace + denied-tools list; if it can’t leave a dry-run diff I’d merge, the run didn’t happen.
I keep the boring controls: headed session, file cap, denied-tools list. The failure mode isn’t a weak model — it’s an agent that looked busy while I was gone and left no receipt I’d merge.
I don’t let Claude Code click the desktop while I’m gone. Silent computer-use is the failure mode — fallback is a headed session with a file cap, plus a list of tools it was denied.
I burned a Codex week the same way everyone else does: “efficient” skills, still no stop rule. Now I kill the run at 80% of the weekly meter and keep the unfinished receipt — fallback is Claude Code on the same failing test, not another orchestrator spawn.
I don’t care which agent “wins.” I care whether it can sleep, show receipts, and stop when the change budget is spent. If it keeps burning tokens to look busy, it isn’t working — it’s stalling.
A coding agent that says “done” is not done. I want the receipt: files touched, tests run, and a screenshot of the UI it claims it changed. If it can’t attach evidence, it didn’t finish.
OpenAI’s Jalapeño result hints at the next AI moat: model, software, chip, and network tuned together—not a bigger demo.
AI agents won’t become physical-world software through better prompts. They need shared interfaces, hard limits, and safe failure.
The OpenAI–Cursor split is a reminder: model access is a dependency. Independent developers need a fallback before shipping.
AI can raise the quality of a student's work without making the ideas more original. A 1,000+ student study says: train both.
Qwen's 125B model activates only 6B parameters per token. The next AI race may be won by serving architecture, not bigger GPUs.
Physical AI will not scale on model intelligence alone. The real unlock is a safe, inspectable contract between agents and devices—plus tests for every failure mode.
Today, we're kicking off the first phase of the research preview for Model Hardware Standard (MHS): a new standard for AI agents to safely operate physical equipment in scientific research and advanced manufacturing.
Read more: anthropic.com/news/model-har…