Chief Architect - AI Transformation @Microsoft https://nitter.cf/t.co/JAVYVd3M2R https://nitter.cf/t.co/jjfH1KydZp - views expressed are my own.
Northern Virginia
Joined January 2009
- Tweets2.1K
- Following298
- Followers203
- Likes2.7K
๐ ๐ ๐ณ๐ฎ๐ฐ๐๐ผ๐ฟ๐. ๐ ๐ ๐ฟ๐๐น๐ฒ๐. ๐ ๐ ๐๐.
Iโve built what I think of as a ๐๐ผ๐ณ๐๐๐ฎ๐ฟ๐ฒ ๐ณ๐ฎ๐ฐ๐๐ผ๐ฟ๐ around Codex: processes, principles, checks, conventions, agent instructions, testing approaches, and guardrails that let me build software very quickly.
Iโm not a software engineer. Iโm a business technologist with some software chops from many years ago.
What I bring today is mostly requirements, domain knowledge, judgment, taste, and an understanding of how things actually work in the real world.
And Iโve realized there are three very different ways I use AI.
๐ญ. ๐๐ฒ๐น๐ฒ๐ด๐ฎ๐๐ฒ.
Sometimes I just want the work done.
Iโve already put my principles, rules, processes, and standards into the factory. So I can describe the outcome I want and let Codex work.
I donโt need to supervise every decision because it is operating inside a system I already designed.
That is the whole point of the factory.
๐ฎ. ๐๐ผ๐น๐น๐ฎ๐ฏ๐ผ๐ฟ๐ฎ๐๐ฒ.
Here I want AI to think with me.
Stay within the architecture and principles weโve established, but challenge the implementation.
Maybe there is a better design. A simpler architecture. A cleaner workflow. Maybe my first idea was wrong.
We go back and forth. I contribute what I know. The model contributes what it knows. And together we usually produce something better.
๐ฏ. ๐๐ฒ๐ฐ๐ถ๐ฑ๐ฒ.
This is different.
Sometimes I know exactly what I want.
Iโve considered the alternatives. I understand the business process. Iโve made the judgment call.
At that point, I am no longer asking AI what it thinks.
I am telling it what to build.
Maybe it created a technically elegant UI that I know a real user will hate.
Maybe I donโt like the naming.
Maybe its approach is perfectly reasonable, but I understand something about the real-world workflow that it doesnโt.
Or maybe I simply want it done another way.
Then the instruction is simple:
๐๐ผ ๐ถ๐ ๐๐ต๐ถ๐ ๐๐ฎ๐.
I donโt care what I told you yesterday.
I donโt care what our normal convention is.
I donโt care what AGENTS.md says.
๐ ๐๐ฟ๐ผ๐๐ฒ ๐๐๐๐ก๐ง๐ฆ.๐บ๐ฑ.
Those rules exist because I created them to make the AI effective. They are not there to overrule me.
That is where human judgment comes in.
In Mode 1, I trust the system.
In Mode 2, I improve the answer with the system.
In Mode 3, ๐ ๐บ๐ฎ๐ธ๐ฒ ๐๐ต๐ฒ ๐ฑ๐ฒ๐ฐ๐ถ๐๐ถ๐ผ๐ป ๐ฎ๐ป๐ฑ ๐๐ต๐ฒ ๐๐๐๐๐ฒ๐บ ๐ฒ๐
๐ฒ๐ฐ๐๐๐ฒ๐ ๐ถ๐.
Sometimes I want an autonomous worker.
Sometimes I want a brilliant collaborator.
And sometimes I want an extraordinarily capable implementation engine that stops arguing and does exactly what I told it to do.
Within the hard limits of the underlying model, the hierarchy is pretty simple:
๐ ๐ ๐ณ๐ฎ๐ฐ๐๐ผ๐ฟ๐.
๐ ๐ ๐ฟ๐๐น๐ฒ๐.
๐ ๐ ๐๐.
๐ค Made with AI
Mere mortals were not meant to have this much power.
I can wake up at two in the morning, look at my phone, and steer three or four AI agents that are building software and infrastructure while I lie in bed.
The power of AI isnโt that it makes a programmer more efficient.
I havenโt gone from being a 1X programmer to a 10X programmer.
Iโve gone from being a business technologist to having an entire fleet of software engineers available whenever I need them.
I donโt remember which AI lab said that software engineering has been โsolved.โ And Iโm not sure it has been solved for the average person walking down the street.
But if youโve spent years adjacent to software engineeringโif you know the language, understand architecture and systems, and can reason conversationally about requirements and trade-offsโthen something fundamental has changed.
For that person, software engineering is very close to solved.
Not because you suddenly became a great programmer.
Because you can now direct great programmers.
I really like ๐๐ฃ๐ง-๐ฒ ๐๐๐๐ฟ๐ฎ.
It is smart, capable, persistent, and very good at engineering work. It also appears to have absolutely no self-control. ๐
Give it a well-bounded task and, if you are not careful, it may decide you also need:
โข a broader architecture
โข extra hardening
โข another abstraction
โข more edge cases
โข more tests
โข another verification pass
The work can be excellent. The token usage can be brutal.
I tried fixing this in AGENTS.md with increasingly explicit instructions:
โChoose the smallest sufficient solution.โ
โDo not add speculative hardening.โ
โAdd tests only for meaningful behavior.โ
โTrust operation results.โ
โStop when the work is done.โ
That helps. It does not solve the underlying behavior. So I stopped asking GPT-6 Astra to police itself.
๐ ๐ด๐ฎ๐๐ฒ ๐ถ๐ ๐ฎ ๐ฟ๐ฒ๐๐ถ๐ฒ๐๐ฒ๐ฟ.
Before consequential work begins, the primary agent must submit the proposed action to an independent reviewer.
The reviewer approves only if:
โข the work is necessary
โข the implementation is the smallest sufficient solution
โข verification is meaningful
โข new tests close a real gap
โข speculative edge cases are not driving scope
โข successful results are accepted and the work stops
I also explicitly banned these justifications:
โBest practice.โ
โExtra confidence.โ
โSomething might go wrong.โ
The reviewer returns only:
๐๐ฃ๐ฃ๐ฅ๐ข๐ฉ๐
๐ฅ๐๐ฉ๐๐ฆ๐
๐ฆ๐ง๐ข๐ฃ
And the primary agent cannot override it.
It is basically separation of powers for agentic software development. So far, this works much better than another paragraph telling Astra not to over-engineer.
The broader lesson:
As models get stronger, the problem is not always getting them to do more. It is getting them to know when they have done enough.
๐๐ป๐๐ฒ๐น๐น๐ถ๐ด๐ฒ๐ป๐ฐ๐ฒ ๐ถ๐ ๐ฒ๐
๐ฝ๐ฒ๐ป๐๐ถ๐๐ฒ ๐๐ต๐ฒ๐ป ๐ถ๐ ๐ต๐ฎ๐ ๐ป๐ผ ๐๐๐ผ๐ฝ๐ฝ๐ถ๐ป๐ด ๐ฐ๐ผ๐ป๐ฑ๐ถ๐๐ถ๐ผ๐ป.
GPT-6 Astra is a terrific model. It just needs to wear this sign for a while:
โI cant help myself. I over engineer and over test.โ ๐
---
@pvncher @thsottiaux - fixing this would probably reduce your load on the servers big time! We business technologists are running wild on token usage!
๐ค Made with AI
AI is rapidly making cognitive production abundant. That does not remove the ๐ก๐ฎ๐ฆ๐๐ง ๐๐๐ฏ๐๐ง๐ญ๐๐ ๐. It increases its value.
The opportunity now is to combine AI capability with human expertise, judgment, taste, context, relationships, trust, authority, and accountability - and turn that combination into governed, real-world operating capability.
The real measure of success is not output. It is ๐๐๐๐๐ฉ๐ญ๐๐ ๐จ๐ฎ๐ญ๐๐จ๐ฆ๐๐ฌ: customer value delivered and businesses that actually work in the real world.
That is the shift we are in.
That happened at while I was out. I didn't see it. I made the mistake of waking up at 2a and saw it then. Two hours later. I now have full control over a Linux workstation in the cloud...wait for it... Using my voice over my iPhone. It's a completely interactive experience and connected to my personal knowledge base (GBRAIN). What's next? Put in in droid form and I have a personal R2!
Thanks @thsottiaux and @garrytan
Claude Code has taught me to be much more expressive... on-hire engineers get a thesaurus+RSUs. Smooshing... Gesticulating... Discombobulating... Reticulating... Perusing... Enchanting... Wandering... Mulling... Spelunking... Deciphering... Forging... Actualizing... Moseying...
Kevin Tupper retweeted
It may be time to prepare for the next level of support at 455. Market is not bouncing out of this hole and the fact that we are glued to the weekly 50 sma is the worst possible news. #SPY #StockMarketCrash
Kevin Tupper retweeted
OpenAI on Tuesday announced its biggest product launch since its enterprise rollout. Itโs called ChatGPT Gov and was built specifically for U.S. government use.
More here: cnb.cx/4aIl872
Kevin Tupper retweeted
MSFT CEO up late tweeting a link to the Wikipedia article on Jevonโs Paradox. This is getting serious.
Jevons paradox strikes again! As AI gets more efficient and accessible, we will see its use skyrocket, turning it into a commodity we just can't get enough of. en.m.wikipedia.org/wiki/Jevoโฆ
For DeepSeek R1 fans . . . do you feel that showing the <thinking> leads to explainability from a user perspective?
Kevin Tupper retweeted
Notice how itโs always โvote according to Biblical values!โ until it comes to welcoming the immigrant, helping the poor, feeding the hungry, bringing healthcare to the sick, forgiving debts, caring for our planet, laying down our swords, or loving our neighbors.
Kevin Tupper retweeted
Steve Jobs email he sent himself 13 months before he died.
Whenever I re-read this, I regret waiting so long to have read it again.
Marking myself safe from today's carnage thanks to
@leadlagreport
and leadlagreport.substack.com.
Kevin Tupper retweeted
We're thrilled to partner with @GitHub to bring Azure AI's industry leading model selection to a community of more than 100 million devs. The latest Azure AI integration brings embedded safety features and simple APIs to unlock AI development. Learn more: msft.it/6019leVbz
The NVDA story reminds me of a quote I heard some time ago on a documentary about IBM vs APPLE in the early days and how MSFT made bank.
During the war for the PC between IBM and APPLE, MSFT sold bullets (software.)
During the AI war, NVDA sells bullets (chips.)
A lot of people jumped the gun, and some apologies are in order.
washingtonpost.com/technologโฆ
Kevin Tupper retweeted
Register for the upcoming series of NISS/FCSM AI in Federal Government by April 17.
Speakers: Reva Schwartz from @NIST, Kevin Tupper
@kevintupper from @Microsoft, and Travis Hoppe @metasemantic from @NCHStats
Moderator: Bob Sivinski from @OMBPress