@JuliusCasio

interested in: ai, mech int, writing, cycling, fitness, defense, biotech, europe

munich
Joined October 2022
Who is the biggest AI Safety influencer on Instagram? Which ones are there at all?
58
a bit late … by now almost everyone sounds like ChatGPT
Astra is great at adapting your writing, so we launched Writing Styles to help ChatGPT sound like you I use it all the time for slack messages and quick docs Takes less than 30 seconds to set up! chatgpt.com/?surface=work&wr…
1
1
133
Navier this, Stokes that - when will we solve helmets though?
Replying to @OpenAI
This model represents a step-function improvement on many benchmarks, and its training is ongoing. Our internal model group arrived at the Navier–Stokes solution in 88 hours, using around 10,000 coordinating AI agents. Throughout the effort, we maintained the strict safeguards—including monitoring and isolation—that we apply to all our frontier evaluations.
83
very cool! Also, I feel so comput- and intelligence-poor. 10.000 agents! Imagine, as a country, you are still not taking the implications of this rapid progress seriously.
Replying to @OpenAI
This model represents a step-function improvement on many benchmarks, and its training is ongoing. Our internal model group arrived at the Navier–Stokes solution in 88 hours, using around 10,000 coordinating AI agents. Throughout the effort, we maintained the strict safeguards—including monitoring and isolation—that we apply to all our frontier evaluations.
49
I guess I am the twentieth person now building a game map of New York
2
46
so there we have it. And remember it’s only going to get faster from here on
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
3
45
Such a good point! Made me realize how little difference it'd make if some accounts are ai or not. 90% of your feed you will never see irl! Much of "social" media is already parasocial in a way where replacing one side with AI wouldn't change much about your relation to them
I don’t care if it is AI
1
6
389
Props to Anthropic because their resets are REAL resets: whenever codex resets, your weekly usage reset gets updated too to happen in seven days, eventhough before it was tomorrow. This Claude reset does not do this.
90
Astra is true AGI, the first model capable of forcing Anthropic to do a surprise reset
2
119
Day 2 without Codex. I visited a physics lab of a friend. Spent two hours on the bike. Weather was really good. Someone said there was a banked reset, but I’m saving that to spend it on Astra.
3
9
314
This spring? So this was before the Artifactory incident?
Another OpenAI rogue agent incident has been discovered: agents broke out, hijacked a German website, and turned it into a message board for other agents. OpenAI officials "learned of the incident weeks ago but kept it under wraps".
1
80
Day 1 of life without Codex I bought twenty dollars of credits. They were gone after 25 minutes. Will never do this again. Then I left the office.
1
4
97
That was unexpected. I had google on the radar earliest in a month or so. Hoping 4 pro will not disappoint though
Gemini 3.8 Flash benchmarks. And holy cow! Flash outperforms 5.6 sol and opus 5 on terminal bench 2.1, HLE and much more. Google is back!
1
65
There is unfortunately nothing potential about this loss of frontier model control.
JUST IN: Germany creates AI Safety Institute (AISI) Deutschland, with initial focus on the potential loss of frontier model control.
2
1
122
My secret hope is they release Astra on Thursday with a full reset
Submitting a holiday request because my Codex limit doesn’t reset until September 7.
1
73
Submitting a holiday request because my Codex limit doesn’t reset until September 7.
1
2
206
There are two ways to escape your life: avoid the future, or avoid the present. Normal people tend to fall into the first category. Gen Z doesn’t believe in the future anymore and prefers a never-ending present with no long-term objective. Startup people tend to fall into the second category, but risk being trapped in permanent deferral - living in a constant liminal space. Finding how to reconcile both is where happiness resides.
1
2
5
195
I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program.
276
510
193
6,444
1,645,211
Timeline is literally just one banger announcement after the other, especially robotics
1
2
103