@clay_phi

Building AI capability into reliable workflows @_GameFrame_ | ex @AzraGames | @WolvesDAO | ex Quant | ex Architect

Houston, TX
Joined April 2020
2FA needs to be re-thought. Having to verify personal agents repeatedly is clunky and slow.
1
3
78
Whatever model Instinct is using is a quite a bit better than Muse. Based on image gen capabilities, I would guess it’s a GPT RL
53
After using Instinct for the past couple of weeks, it’s more clear to me now than ever that we will rapidly move to a bespoke, just-in-time lifestyle. Games, music, movies, general digital entertainment, productivity software, finance, etc., created and curated around your tastes, needs, mood… exactly what you want when you want it. That will extend into the physical world as robotics give LLMs legs. I said this a year ago, and now fully believe that it is going to become very difficult to sell software and services. So much value lies within models.
55
Best result for all would be for OpenAI and Anthropic to slow down internal dev, allowing SpaceXAI to pass them up.
132
One of the frontier labs REALLY needs to prioritize training for ill-structured problem solving and strategic judgment under uncertainty. The most intellectually demanding part of work is knowledge work where deeply complex problem solving requires consistent metacognitive control over days to weeks. This is by far the biggest capability gap limiting how much LLMs accelerate my work, and I imagine for many others. The most demanding problems require sustained judgment over days or weeks when relevant factors, their relationships, the possible solutions, and even the criteria for choosing between them aren’t all specified or known upfront. Uncovering and continually reassessing those things is an integral part of solving the problem. Current models can explain individual concepts correctly, but fail to apply them together consistently - a problem that grows as the factor/concept count grows. They leave assumptions unexamined, miss how one set of factors changes the importance others, and don’t recognize when a new finding invalidates an earlier conclusion. They also struggle to identify what evidence is missing, which question would uncover it, and what alternative explanations remain before committing to an answer. I end up supplying the questions AND most answers that actually move the investigation forward. Co-working problems like this with a model reveals how far they have to go to bridge the chasm to human level general intelligence. I can make a judgment in seconds that takes a model 30 minutes in a multi-agent adversarial debate panel to reach…and the panel can still be badly wrong. The model ends up becoming a researcher and spec writer, mostly useless in problem solving no matter how specific the skill/workflow. In a relevant study, MT-InfoSeek found in a logic test that none of the tested models exceeded 40% accuracy at identifying the exact two missing variables needed to determine an answer, without being told that two were needed. This only measures exact selection of a tiny amount of missing information. Study: arxiv.org/html/2608.14808v1 That test starts with defined variables and rules. Real problems can involve 100s of potentially relevant factors, an unknown number of missing ones, and relationships or constraints that themselves also need to be discovered. Perhaps the labs have this internally to maintain an edge? Right now what’s available publicly is not close to human capabilities on this front. Constructing and revising understanding throughout an investigation, identifying consequential omissions, and supplying the important and right-timed change of direction is THE hardest part.
62
Seems relevant. Frontier charades will continue until they achieve this.
If Anthropic’s rhetoric fails to achieve the regulatory capture they desire to stifle competition, next step is an ‘open source’ model does something nefarious to force regs. Anyone working closely with these models understands that the LLM Anthropic is serving consumers is nowhere near capable of the rhetoric they proclaim.
45
Instinct is the first LLM integration that has Wow’d me in months. Very well done. No setup. It just works out of the box.
1
1
43
With what passes for good skills / mega GitHub stars, I have to believe many keep their workflows close to vest. Nearly every time I’m recommended a skill, it’s 50-200 lines of NL instructions.
36
After much annoyance, I have found AI writing detectors to be junk science. A passage I wrote scored 100% human. I added two sentences, also entirely my own writing, and it suddenly scored 100% AI, including the unchanged section it had just passed. I originally wrote this post and then had ChatGPT rewrite it. 100% Human. You can take this post, feed it into Pangram = 100% Human. Change the first line to "After wasting far too much time on this, I’m convinced AI writing detectors are junk science." = 100% AI Test can be repeated over and over in varying lengths.
36
Interesting dynamic I’ve noticed as my LLM usage has increased over the past year. Stress increases linearly alongside higher proficiency with these tools. Like managing 1000 Leonard Shelby / Raymond Babbitt clones. You can now implement any idea to any level of scalability or complexity. Problem is you have to manage 1000 disabled geniuses. Hope we end up like Charlie and not Teddy
1
2
47
Replying to @ArtificialAnlys
@ArtificialAnlys pls update your Harness Comparison. Run Sol, Grok, Claude, etc on Codex, Claude Code, Open Code, Grok Build. Important KPI totally missing from the market.
2
1
57
Need objective metrics for model ‘intelligence’ across harnesses first.
6
Is that thing where AI agents cause widespread unemployment still happening, or…?
1
43
Gpt 5.6 Sol is easily the best model out right now. It’s so good at following instructions that my task artifact and intra-task QA hooks built to get prd-loyal multi-day workflows out of Claude models now cause Sol to work at tasks near indefinitely due to endless QA revisions.
34
Fable was nerfed post-ban, and became even worse when Sol forced them to make it part of subscriptions. Opus is even worse. Benchmarks do not tell the tale of practical workflow intelligence. Sol is currently the leader, by far.
1
100
Kimi K3 is like when everyone thought Deepseek was about to upend the industry. Hardly anyone has even used it, and when you do, you quickly realize it’s not.
3
1
179
Feels inevitable that someone builds a trustless “Build Your Own Model” service. Connect your GitHub, Slack, Docs, Jira, CRM, etc. Click Build. It automatically mines your data and generates high-quality input/output pairs, and fine-tunes your proprietary model. Provably without any operator ever seeing a byte. Synthetic data pipelines have been commoditized. TEE training runs at ~99% native speed. Remote attestation works. Nobody’s bolted it together…
48
The poor performance of Opus 4.8 finally caused me to try Codex. GPT-5.6 Sol in Codex is -remarkably- good. Far far better than how it performs inside Claude code. My skills, agents, hooks, and settings don’t port cleanly and I didn’t invest time to harden anything for Code… yet Sol, unprompted, followed directives from within them while creating a spec for a simple request, initialized my task-building skill, and worked around the missing templates and scripts. The easy configurability in Codex is also superior.
70