@mickcodez

Staff Software Engineer - Just a ghost in the machine 👻 . 🧑‍🚀 High Agency Individual.

Alpharetta, GA
Joined April 2011
You still have to take time for family and fresh air! Looking forward to next week for dev day, but tonight we watch some sportsball.
51
mick retweeted
okay, probably a bad idea to open up this experiment so early. it's VERY BAD, buggy, NOT polished. but...now Megawatt/Gigawatt members can opt into ampcode.com/settings/experim… to use Claude Code CLI (w/Claude Pro/Max subs) in Amp orbs it's just the Claude Code CLI jammed into Amp's UI, minimal integrations, no nice UI, no thread syncing, etc. as I said, VERY BAD. but judging from my emails/DMs/replies, people want to try this anyway. feel free to report bugs, will triage them but won't fix immediately, and we might kill this feature entirely or be prohibited from shipping it.
28
5
15
104
49,800
mick retweeted
This has never been more true. The opportunity cost of endless planning and contemplation just went way up. Agents allow you to just try vastly more stuff, so let them, and you'll learn way more, way faster.
“Action produces information. Just keep doing stuff.” — Brian Armstrong
121
421
32
4,689
212,959
Asking Meta's Muse what happens to my data:
191
912
88
14,152
550,038
401s from @OpenAI is never a good sign...
2
124
weird this is the graphic that the partnerships team told me we were going with @Box 🤝 @Muse
Muse and Box, together at last
210
45
62
2,799
422,950
Gotta say, the partnerships team cooked with this one
weird this is the graphic that the partnerships team told me we were going with @Box 🤝 @Muse
2
4
466
I can't say it enough but interfacecraft.dev is an incredible resource for "Uncommon Care". AI can do so much, but it can't care about your user like you do. I am really enjoying the new section on animations
2
72
Muse and Box, together at last
Replying to @hbarra
5/ And some of our favorite productivity connectors, with plenty more in the works: @Box / @github / @meetgranola / @NotionHQ
31
12
3
406
422,330
Oh my word
Being a dev today using frontier AI feels like I'm playing Watson to Sherlock Holmes. I just stand back and go "my word holmes how did you deduce that" and occasionally stop him from getting shot
1
83
mick retweeted
When meeting other engineers I have started saying “back in the day when we used to hand-write code” to quickly learn if we’re going to be friends. It’s like Jev classification for real life.
3
1
22
816
mick retweeted
We built a tutorial showing how to create a customer intake agent on @Vercel Connect that opens a folder of messy uploads, uses Box AI to identify each document, and files it into the right customer folder automatically. Full walkthrough here.👇
1
1
12
89,347
I’m very pleased with all my @robotaxi rides so far. The future is really bright… I may buy my own fleet and deploy them.
1
1
114
Dev Day attendees should get 10x resets so we can cook before, during, and after the conference
57
Opus 5.5 is a really good model. It's been my daily driver the last few weeks. We had Opus 5.5 and Fable 5.1 each port HAProxy from C to Rust. Both passed nearly all of HAProxy's tests, but Opus 5.5 finished in 9.5 hours compared to Fable 5.1's 12 hours, and for 51% less cost.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
381
262
54
7,670
521,029
Do we have a polymarket running to see how long it'll take OpenAI or Anthropic to come out with their own RLCD model usurping Jev?
87
mick retweeted
GPT-6 Sol and Luna are out. Not only are they a very significant improvement across the board, but also in writing and general "you know when you try it" quality. We are also permanently reducing the API price by 50% making both of them viable for a ton of new usecases and making your usage go further too, even on the subscriptions. And one more thing. We are loading a banked reset into all accounts of our Plus, Pro and Business users. Let's go! openai.com/index/introducing…
We have been focusing on efficiency and intelligence for all. Very proud of the team. Only possible when you have incredible models at the top end of the capability that you can then use to make a big difference in everything else.
3,033
1,610
949
25,291
3,213,439
What an insane day in AI. The frontier models just became substantially cheaper, with the Opus 5.5 price cuts, and now with GPT-6 Sol and Luna dropping token prices by 50%. The rate at which the cost per task (on a like-for-like basis) drops in AI is unlike any other type of technology in history. And every time the cost of AI drops, the use-cases you can deploy agents against dramatically increase. This is Jevons paradox applied to agents. These improvements will directly lead to broader diffusion of AI in the economy as we can use agents to process all of our data, scan our code for security issues, read through all log data to make decisions, have agent swarms in workflows, and much more. The cost of tokens is directly correlated to these use-cases being opened up at scale.
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
107
83
18
750
141,262
This will definitely be how we build UI's... the components will be built and then just rendered dynamically.
Quite a few people clowning on me for saying that UI will be generated on the fly. Two things: 1. You don't have to generate it all from scratch, but dynamically assemble widgets. 2. Tried chat jimmy? In a few years, we could have today's frontier intelligence at this speed
1
80
At Box, we've been testing Opus 5.5 on a variety of complex enterprise knowledge work tasks dealing with unstructured data with the Box Agent. Overall, we saw frontier capability levels, with major performance improvements over Opus 5. 63% fewer tokens used, 42% less verbosity, and 30% faster vs. Opus 5. And the model itself is cheaper, so this is a major win for any agentic computer use, coding, analytics, or data work that enterprises will be doing. Here are some examples of the task wins and performance gains across a variety of industry tests that we performed: • Financial services - due diligence (+39% task accuracy): A year of transaction records, with the job of finding every miscalculation in an acquisition target's pricing tool. Opus 5.5 scored a perfect result on every attempt in half the words Opus 5 used, consuming 82% fewer tokens overall. • Technology - cloud cost analysis (+65% task accuracy): Work out what a company should actually change about its cloud spend. Opus 5.5 picked the right basis for the retention calculation and kept the source data's unit conventions straight all the way through, so the number at the end actually holds up. It took half the time Opus 5 took, with 70% fewer tokens. • Consumer products - client account analysis (+17% task accuracy): Set the onboarding targets for a client account, reading across the signed contract, a satisfaction tracker and a team metrics sheet. The contract never states a senior/junior split, so Opus 5.5 derived it from the 18-person roster and showed the rule it used; several clients had a perfect 10 on individual survey questions, so it averaged each client's responses instead of crowning the single 10. It finished this one in half the time, on 78% fewer tokens. • Clinical diagnostics - data analysis (+15% task accuracy): Malaria rapid-test performance across a dry and a wet season: build the patient records out of two clinical PDFs, compute positive test rates by season and gender, and test whether parasite counts really differ between test-positive and test-negative patients. Opus 5.5 caught that the two groups' standard deviations differed more than 100-fold, re-ran it the right way, and found the dry-season difference didn't hold up after all. This accuracy gain came with a final answer that was half the length of Opus 5's, and also needed 78% fewer tokens end to end. Customers will be able to build AI Agents with Opus 5.5 shortly in the Box AI Studio.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
59
28
8
387
150,777