@MaxBessler

Researching how tech impacts our world, from big geopolitics to our small daily lives. @BayAreaEconomy and @Oxford's MPhil in International Relations

San Fransisco
Joined July 2012
Authenticity pays off in the long run. I won't let AI write my sentences here. I think I’d rather be wrong more often, post less with more typos, and lose out on notification dopamine... than outsource my thinking; the upside is I learn more in struggling through the process.
7
1,405
I just went shopping and the cashier couldn't divide 360/90 (it's 4 btw). AI is another milestone in technology abstracting our critical thinking one layer higher in the real world, and forgetting the ones below it. So school teaching fundamentals becomes more important.
Nothing will make you more against AI than hearing what kind of world the people behind it are trying to create
54
I’m not following this closely, but it sure does strike me as a similar time where there was a group of folks who rejected the the printing press.
Bernie Sanders and Greg Casar introduced the Ban Artificial Superintelligence Act today. It will: - ban the development or deployment of Artificial Superintelligence, aka any AI that 'exceeds human cognitive performance and capabilities across most domains, or has sufficient capabilities to destroy or disempower humanity, including by overthrowing the federal government' - pause 'Advanced AI development' until a new federal AI regulatory body is up and running and has established clear rules and model-review processes - establish a new cabinet-level federal Department of Artificial Intelligence (this should be the Department of Super Intelligence, there must be some mistake) -set penalties for any person or entity that attempts to violate or circumvent the pauses and prohibitions in the bill: 'Entities shall be subject to the corporate death penalty, and persons shall be subject to not more than 20 years in prison' - set the international policy of the United States to 'pursue international agreements, allied coordination, and policies such as export controls to prevent the development of artificial superintelligence anywhere in the world'
2
66
Data is the moat and always will be!
muse is not a great AI app, it’s a fantastic *personal* app. the ai model meta uses is not even frontier and yet you don’t care because the experience is so damn good. the secret sauce is meta built a much, much better open claw trained on your habits, preferences and decades of social data that’s why the agent’s intuition is so on-point. it knows you so well. i haven’t found myself coming back to an app so much since chatgpt launched. muse is currently building a tinder-style shopping app for me based on real clothing items available across top retailers. it’s doing this *while* giving me my daily brief, checking some stuff out on marketplace and much more. it’s like having 100+ personal assistants.
1
86
Max Bessler retweeted
muse is not a great AI app, it’s a fantastic *personal* app. the ai model meta uses is not even frontier and yet you don’t care because the experience is so damn good. the secret sauce is meta built a much, much better open claw trained on your habits, preferences and decades of social data that’s why the agent’s intuition is so on-point. it knows you so well. i haven’t found myself coming back to an app so much since chatgpt launched. muse is currently building a tinder-style shopping app for me based on real clothing items available across top retailers. it’s doing this *while* giving me my daily brief, checking some stuff out on marketplace and much more. it’s like having 100+ personal assistants.
24
19
2
365
37,202
A true masterpiece. Way ahead if it's time.
Name a TV show that was absolutely perfect from start to finish. No bad season
241
1,015
180
5,740
939,470
I'm increasingly convinced @Instinct is running on the APIs of Chinese models. Is it time to switch to @Muse? Who has done a security audit of this stuff? Shouldn't be hard to transfer context with Claude coding some easy scripts that download my iMessage convos.
1
69
Wild graphic out of the FT this morning on how top heavy midterm megadonations lean Republican
1
1
32
Lots of different reads, I think we're seeing an oligopoly forming in American frontier development > The top labs are pre-empting the wave of legislation by promising to self-police themselves in a way they—and their employees want. By doing so, it avoids more employees starting their own startups, blabbing about p(doom) on X, and keeps researchers happy about the mission. > This also preserves helps ease investor qualms ahead of their IPO. Introducing “independent” evaluators that are based in Berkeley, rather than DC, means they control the pipeline of how they’re regulated (Effective Altruist types and METR/Redwood to name a few). AI Safety is an industry that is BOOMING and will soon become a revolving door, akin to banking as cited in the article. If you’re any leading lab, you want that mindshare closer to what your team already believes, reducing uncertainty. By pacing research, it also allows the finance teams to better approximate compute spend and runway, even if they say their TAM is going to be $20T or whatever in their forthcoming S-1. > Lastly, it’s a private company pushing for more US techno-nationalism to further insure US AI survives the Chinese open threat. “Pacing” could be viewed as Anthropic confident that AI development could slow and China still wouldn’t surpass us on the frontier, given just how much of its progress rests on distillation and corporate espionage. Coupled with stricter export controls of hardware (GPUs) and people (talent), coordinating releases could even widen the gap by plugging holes before they leak. “Evaluators” embedded into labs allows the NSA, our IC, and private monitors to audit threats. I’m sure their threat intelligence team is completely overworked right now, just from reading their most recent report. > I fixated mostly on the independent evaluators point because I think that’s where the AI labs actually have the most sovereignty and influence. The higher you climb in politics, from within our democracy to global geopolitics, the influence of any one lab falls. So his points to #2 and #3 require more buy-in from OAI/GDM/X/MSL/etc and then the US/G7/G20 respectively. m
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
114
Max Bessler retweeted
History rhyming.... * When the US pushed Canada away with the Smoot-Hawley tariff in 1930, Canada retaliated with tariffs and also adopted Imperial Preferences with UK & Commonwealth... * Facing higher US tariffs today, Canada retaliated again and seeks closer ties with Europe since access to US market has become unreliable
22
33
2
216
30,634
Max Bessler retweeted
so, if dario thinks that us labs can still maintain a lead over chinese labs, while pacing the frontier, i think it probably means that he thinks most of their progress is distillation otherwise, china seems about 4 months behind the frontier and you can't afford much pacing
10
7
111
4,899
AI labs with the resources are Pacing the Frontier of Innovation— and Regulation. Embedding independent evaluators preempts the US Government stepping in. The writing was on the wall there.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
1
57
Robotaxis won’t just kill short haul flights. Once you can sleep in them, they’ll cannibalize long routes too. Imagine getting into your Tesla in Austin at 10 PM, make the lie flat seat/bed comfortable, and going to sleep. While you sleep, the car drives, charges itself, and keeps going. You wake up at 9 AM in Aspen, refreshed. The key is perceived travel time. A 10-hour robotaxi trip where you sleep for 8 hours may feel shorter than a 3-hour flight that requires an Uber to the airport, arriving early, security, boarding, the flight, baggage claim, and another Uber. Autonomous cars won’t just compete with driving. They turn sleep into transportation.
1,241
1,262
330
18,497
1,936,559
Am I going crazy here? Is this a fair comparison? Imagine if any other F500 public employee at another company that sells a GPT— like electricity or oil— were to say "yep! we believe our product could kill everyone!"
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
1
143
I took a break from the internet, slowed down, and revisited my own life, and remembered how much of the news and our desire for information is entertainment. I can continue to do things like daily briefings and constant commentary, but to what end?
1
1
92
Max Bessler retweeted
The METR/Redwood report is the top story on the New York Times home page the last few hours. Piece takes as a given that the report is very troubling and notes that the reality could be even worse as the investigators weren't given enough access to uncover everything.
9
38
6
625
20,151
Max Bessler retweeted
Across the US, workers are souring on AI. The tech industry has insisted that AI would make workers’ lives easier and more productive, that the more people understood and used AI, the more they would like it. Something closer to the opposite has happened.
18
272
51
1,082
153,070
The amount of techlash to Astra’s release on X—notorious for its technoptimism— means once these agents hit the general public, they’ll have an even worse narrative battle to fight.
1
43
In Amsterdam, I rode a BYD. We are behind.
1
76
My friends and I trade notes on what we're catching each morning. Here's mine for Saturday, August 29: Good morning, I missed yesterday with travels so today is a bonus Saturday makeup note. > China’s reaping the rewards of its large oil stock pile, purported to be 600 million barrels deep last year. This is an effective buttress against the US-West led Iran War that’s disrupting all sorts of supply chains. Analysts today suspect China has between 1 to 1.4 billion barrels; were everything to shut off tomorrow, that’s 120 days of imports. Still, China relies on 70% of imports for its crude supply. (WSJ) Energy security underpins manufacturing might, and with it, geopolitical. > OpenAI announced its plan to end access of its models on Cursor come November 12; this follows the SpaceX acquisition of the IDE basically because OpenAI cannot be confident that “SpaceX will use our technology within our terms of service, based on our experience with Elon Musk's companies violating contracts.” (OpenAI). > X revealed there’s a bot farm of ~200,000 something thousand users, of which ~200 were posting disinformation about energy and AI American policy, fomenting division. (Fox Business / X) Data centers need the narrative to change in its favor but technologists do not have the marketing skill for It to happen— that’s for sure. > Trump has struck a deal to gain 65 billion barrels of oil in Venezuela. That’s 55% of new private joint venture, the second largest of any private venture (behind Saudi Aramco). This is going to control more than a 5th of Venezuela’s total oil reserves. (FT)
1
1
93
Max Bessler retweeted
I want everyone who has ever even *thought* about Congressional oversight to read the entire linked post very slowly & absorb every word. These are technical experts explaining that the only way for them to understand & parse the complex/ massive amount of information about what happened in the Hugging Face incident is by using (imperfect) AI tools. THERE IS NO way for Congress to understand & oversee this industry & the others that are changing due to AI without investing in its own tools and technical talent. Just as the best Congressional staffers in the world could not have conducted adequate oversight of Internet companies without web access, AI literacy & AI-enabled tools are no longer optional for our legislative branch. They are essential to its operations & relevance. This is not a question of procurement. It needs to be a complete rethink of what it means to be a functional democratic institution in the age of AI. I truly believe that this upgrade is required to maintain our precious checks and balances and our right to self-governing. We will be releasing more about what it could look like in the coming weeks & months leading up to a new Congress in January 2027. This is not a drill!!
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'. I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident. Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them. We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation. Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why! The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing. While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future: - Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations. - While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies). - The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities). - We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation. In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
13
32
6
241
26,548