@a_simpson

I do open source 📊 at @grafana. I tweet about the NBA before Christmas. I recognize cookies as currency. I love my wife, Christi. ❤s Magit.

Dayton, OH
Joined April 2007
Yesterday we launched a project in collaboration with the @okcthunder: ⚡thunder.run I couldn't be more proud of how it turned out. 🏃‍♂️💨🕹
1
6
Adam Simpson retweeted
If you remember this, your back probably hurts.
262
164
112
2,533
119,328
💯 To put it another way: have an agent write what you know, build what you don’t understand yourself.
recommended reading, with the caveat that i think the article goes too hard on bend 2. the kernel: you often only understand a problem by experiencing the journey, i.e. building a solution for it. if you hand that journey off to an agent, you will get something. but not necessarily something good, or even correct. blog.liampwll.com/posts/bend…
2
53
Adam Simpson retweeted
a data point, grain of salt.
Does your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge: 1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost 2️⃣A simple harness can be competitive 3️⃣The native harness isn’t always the best. Millions of people are using coding agents, but the impact of harness choice remains unclear. (1/n) More details in the thread. 🧵
2
5
126
19,441
It’s unreal how poorly the “AI labs” communicate and how high on their own supply they are. Never forget Anthropic employs a psychiatrist…for LLMs 🙈 This on the other is clear, sensible, and contains no hysterics about human extinction.
Last month I wrote about how we can build a positive and safe future for everyone: meta.com/thefutureisforevery… Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens. The reality is: - People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned. There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind. - Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well. Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would. We just did it as part of our day-to-day work because it was clearly the right thing for people and for us. I'm proud of the security foundations we've built. - Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas because it helps produce better work. Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators. - Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well. I believe the key to building a positive future for everyone is maintaining the right balance of power. This is within our power to do.
1
3
466
Adam Simpson retweeted
You may not like it, but this is what expert prompting looks like
34
26
6
593
42,777
Thanks @Tailscale for the opportunity to speak at TailscaleUp, it was a blast!
4
268
I’ve experienced @Starlink on @united. I can’t go back now.
3
186
I don't run @OmarchyLinux since @nixos_org is home but that doesn't stop me from stealing some of the great ideas for my own config 😏 Quickshell in particular seems really really good.
1
48
"The real value I think never really was the code, it was the learnings from the journey to get to the code" from news.ycombinator.com/item?id…
1
1
67
Paired nicely with: "[...]models are far more diligent than humans at rote correctness. Their errors are isolated to architecture, unexpected use cases, visual output their test environment is not feeding back to them, etc." -blog.exe.dev/devtools-must-b…
28
Thanks to @pidotdev I want real context and system prompt control everywhere. On mobile, ChatGPT is a not very good harness (has basic branching at least) and Claude is awful.
1
1
106
Performance is so much better when you can move any interaction from a long-running chat to: 1. Clear, efficent system prompt 2. Data and/or tools 3. Request/task
22
The OpenAI app confusion makes me wonder at what point do people vibe-code their own LLM apps and OpenAI/Anthropic/Gemini are relegated to "just" API providers?
1
122
Adam Simpson retweeted
Got em. I poison my AGENTS.md (and other things like code comments) all over the place with prompt injections like this to find people who don't review their code and sling it off to another human. Catches folks all the time and then its an instant ban. As I've said, I don't care if you don't review your own code. But if you're submitting code to an OSS project and crossing a human boundary, it is simple courtesy to do some human review.
183
434
101
7,024
780,863
I'm convinced this Fable fiasco could have been avoided if Anthropic were honest about what LLMs are and not constantly hyping each release as AGI.
2
46
Been working with Go's template library a lot recently, this slide deck (I know right?!) has been super useful docs.google.com/presentation…
1
2
87
Levie’s Law of AI Psychosis: The farther away you are from the actual work the more confident you are that humans are no longer needed I like it
CEOs are uniquely prone to AI psychosis because they’re sufficiently distant from the last mile of work that still has to happen to generate most value with AI. So when they play with AI, they see the happy path results, often not considering the next 10 or 20 things that have to happen to get sustainable results from agents. “Look I made this awesome product prototype”. Yes but you didn’t have to review the code before it went into production and fix a bunch of issues. “Look I generated a contract”. Yes but you didn’t verify all the terms before it goes out to the counterparty and didn’t have to wire up all the past contracts to work with. The best thing you can do as a CEO is to use AI a *ton* to figure out the real implications of agents in the enterprise, and come out the other side with an appreciation for both the upside and the real work that goes into them.
74
861
33
7,574
445,205