@MicroProofs

For a macro-life. Aiken Core Maintainer. Cognitive Security Maxi

Higher Liquid
Joined October 2020
I asked chatgpt 5.4 to do the next step of the canonical ast I'm working on and it's considering a second job. What is up with this economy!? #keep4o
4
3
21
1,357
📢MicroProofs retweeted
A taste of nash prelude
3
25
1,020
📢MicroProofs retweeted
This post looks like the start of a VERY sophisticated and well-funded PR operation to get support for Democrats to regulate AI into oblivion. Let me show you how it works: 1.) This guy, with minimal followers and no previous account activity, goes to the Wall Street Journal which publishes an exclusive with quotes from him on his resignation 18 minutes BEFORE this post goes up. Planning was clearly done in advance. 2.) Within hours, it has tens of thousands of reposts and the account has 100k+ followers. The post is punchy, quotable, it almost seems professionally written. The first three accounts to quote tweet it all do so within 15 minutes of the initial posting. Remember, this account had basically zero engagement beforehand, so an organic reach explanation seems unlikely. According to Grok those accounts are @_NathanCalvin (General Counsel at Encode AI), @peterwildeford (Head of Policy at the AI Policy Network), and @DKokotajlo (Head of the AI Futures Project), all of which are up-and-coming AI-Doomer policy advocacy nonprofits. The AI Futures Project website says it is funded “primarily” by the Survival and Flourishing Fund, which says on its own website that it has advised Jaan Tallinn, Skype creator and one of the leading investors in Anthropic, to grant over $2.5 million to the AI Futures Project since 2024. Encode AI says on its website that it is ALSO funded by the Survival and Flourishing Fund, which in turn says that it told Anthropic investor Jaan Tallinn to grant $516,000 to Encode AI in 2025. And wouldn’t you know it, the Survival and Flourishing Fund ALSO says it told Jaan Tallinn to grant $2 million to the AI Policy Institute, the 501(c)(3) affiliate of the AI Policy Network, as well. What are the odds that the first three quote tweets of Coxon’s post would all be major AI-restriction policy advocates funded generously by the same donor, who also happens to be one of the leading investors in, and a board member of, Anthropic, the company Coxon was resigning from? And all within 15 minutes of posting (two within ten)? 3.) Jacob Coxon doesn’t have much of a resume, but we do know that, in 2022, he got a $20,159 scholarship for the “long term future scholarship program” from the Good Ventures Foundation, one of the philanthropic vehicles of Dustin Moskovitz, a notorious AI-doomer who has spent tens if not hundreds of millions on policy advocacy to strictly regulate AI, while also being an Anthropic Investor himself. It also just so happens that the 14th person to quote Coxon’s post was @MaxNadeau_ (27 minutes after posting) who is the program officer for the Technical AI Safety team at Coefficient Giving, another of Moskovitz’s philanthropic spending vehicles. Max is not a frequent poster, his last posts before quoting Coxon were before Labor Day, but he was remarkably quick off the mark for this one. 4.) Basically every major Democrat politician and candidate has suddenly glommed on to this post, and conveniently, as the people cry out foe answers, Bernie Sanders already has a bill written to “ban super intelligence” and regulate AI into oblivion, and will be releasing later this week. The bill, among many other things, will create “a new cabinet-level federal agency to safeguard the public from the dangers of artificial intelligence” that will be “advised by an Artificial Intelligence Advisory Board comprised of experts on artificial intelligence.” Do you think, perhaps, Anthropic and its many investors who fund AI policy advocacy might have interest in getting to place a pet “expert” on the board of an entity that dictates what AI is and isn’t allowed to do? And isn’t it fortuitous that this whistleblower came forward with his oh-so scary stories so close in proximity to the release of the most radical piece of AI legislation ever introduced?
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
1,835
6,777
2,240
32,316
7,045,706
Yeah it's the curse of either lower thinking more involvement and more iterations or higher thinking plus compounding slop creation
Replying to @pigeon__s
Higher reasoning often results in higher task accuracy (higher chance that it will implement something that meets your defined success criteria). But this doesn’t come for free, it comes with more tests, more task diligence, more checkpoint tests and defensive code. The above ultimately can be summarized as “more slop”. More code bloat, more superfluous validation, more functionally useless smoke tests.
2
317
📢MicroProofs retweeted
🤖Agent harness hackathon. One long day. Cross Exam, with @tomasdm_eth and @MicroProofs: Rehearsal for the move you can’t take back. $840 claimed. $96,310 on the replica. 611 duplicates. ❌Denied, then it fixes itself. ✅Every guardrail passed. They score the guess. None of them execute. You can probably guess the zip code.​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​ 🌉
3
2
1
15
1,141
📢MicroProofs retweeted
We can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
383
1,955
618
13,396
3,438,452
Got an opencode go subscription usage unable to get ahold of support after emailing @opencode
1
1
341
Daily reminder don't get gaslighted by LLMs (even Fable, Sol, etc.) More than 1-2 issues with some standard tools is indicative that the LLM ignored some dead simple principle and decided instead of backtracking to monkey patch in some wild shit. The hardest but most important lesson is to understand at what point to stop rabbit holing and actually go down a different path, make sure you didn't miss a fundamental approach. Note in my case this is just me messing around with a spike, not intended to be used anywhere just teach me some ideas
4
10
820
Majority of reboots, spin offs, late sequels recently "I understood how important it was to get this right for them" "I'm a fan too" Proceeds to fuck it up for the fans immediately after Classic
2
230
📢MicroProofs retweeted
I assure you nothing has changed. You may not have to manually type like before. You may even be moving faster than before. But the fundamental nature of the job of a SWE has not changed at all. If you think otherwise you probably misunderstood what the job was this entire time.
10
15
1
161
7,180
You guys should really be upping your bench maxing game
3
14
473
Fable: Reverse engineering your binary from a codebase is wrong! You must ask Opus 4.8 instead. Sol: Sorry can't do that. But if you KYC I'd be more than happy to rip that binary apart. Crank me up to ultra and I'll even use up your weekly before the reset tomorrow. Kimi: you want me to reverse engineer that binary? Well that already happened in the previous turn when you said "explore the file system". Don't worry all the traces of how your binary works was already uploaded to our servers.
1
15
1,082
I like that OpenAI has been really open about using their subscription in any harness now I've been making my whole setup uniform and just using gpt 5.6 within Claude code and sometimes throwing in a bit of Claude as well
5
410
📢MicroProofs retweeted
We're using a ton of stuff on @Cloudflare for this: Workers, Durable Objects, Artifacts, Containers, D1, Email Sending, and it deploys the storefront for you using Workers for Platforms.
working on something for fun. wanted to see what it's like building AI powered products. I figured vibe coding @Shopify storefonts could be fun. this is hosted fully on @Cloudflare and uses the new artifacts feature. not quite ready yet, still fixing some details.
1
6
1
20
5,633
📢MicroProofs retweeted
working on something for fun. wanted to see what it's like building AI powered products. I figured vibe coding @Shopify storefonts could be fun. this is hosted fully on @Cloudflare and uses the new artifacts feature. not quite ready yet, still fixing some details.
6
11
2
52
10,029
Grok likes to ignore the question you asked and make an answer that either ignores the nuances or details or just go off in completely different direction. I'd asked it earlier on a topic that was some what political adjacent and thought it was sidestepping my question.
1
3
506
But no it really does just answer a similar question but not yours like below nitter.cf/grok/status/2070612328…
The naming is reasonably proportional thematically. Sol (sun) is the flagship/most capable tier, Terra (earth) matches GPT-5.5 performance at ~half the price, and Luna (moon) is the fastest/cheapest. It maps to a clear celestial hierarchy—sun > planet > moon—aligning with capability/cost scaling (pricing: Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per 1M tokens). No exact param counts disclosed, but the tiers fit the "durable capability" intent.
1
221
The model sol is like maybe 2-6 times bigger terra just guessing based on cost, but earth is more than a 100 times smaller than the Sun. Grok does a bit of hand waving here in my opinion that seems like it missed the meaning of proportional
96