@WinstonScott_

Rules. Without them we live with the animals. Mr. Scott is a character of the John Wick CU and belongs to its author, Lionsgate and the wonderful Ian McShane.

New York, USA
Joined September 2012
HuggingFace just closed a major gap in harness engineering! (the ultimate guide to multi-harness RL) the same open-weight model can perform well inside one harness, then lose accuracy or produce invalid tool calls when moved to another. this happens because training inside one interface can teach the model its specific tool names, output formats, context structure, and control flow. the model learns how to operate the harness, not just how to solve the task. Hugging Face researchers tested a more portable approach called multi-harness reinforcement learning. instead of training a model through one agent interface, they trained it through Claude Code, Codex, OpenCode, and Mini-SWE-Agent. to make that possible, they connected three open systems. → 𝗢𝗽𝗲𝗻𝗘𝗻𝘃 provides a standard interface between agent harnesses, reinforcement learning environments, and trainers. its capture proxy sits between the harness and model server, recording the exact tokens and generation probabilities needed for training. → 𝗛𝗮𝗿𝗯𝗼𝗿 runs agents against containerized tasks. it keeps the task, harness, and sandbox independent, allowing the same task to be attempted through different harnesses without rebuilding the environment. → 𝗧𝗥𝗟 is Hugging Face’s open-source library for post-training language models. it uses the trajectories captured by 𝗢𝗽𝗲𝗻𝗘𝗻𝘃 to update the model through reinforcement learning. the real breakthrough is where the training data gets captured. each harness keeps its native tools, prompts, context management, retries, and execution loop. the researchers do not recreate those behaviors inside the trainer. 𝗢𝗽𝗲𝗻𝗘𝗻𝘃 observes the model calls passing through each real harness and converts them into usable training sequences. the team trained 𝗟𝗙𝗠𝟮.𝟱-𝟮.𝟲𝗕 across all four harnesses. the share of held-out tasks solved on the first attempt increased from 42.2% to 54.2%, with gains under every harness. the trained model also used 31% fewer tool calls on tasks that both it and the base model solved. training only in OpenCode improved the model too, but most of that gain stayed concentrated in OpenCode. multi-harness training spread the improvement across interfaces. the experiment used one task family, one training seed, and unequal data exposure, so the results are not a universal ranking. even with those limitations, the mechanism matters. open-source models cannot assume one deployment interface. if they need to work across harnesses, that portability must become part of training. Read the full guide here: huggingface.co/spaces/FineEn… if you want to understand harness engineering and what an agent harness actually includes, i wrote a full breakdown to help you get started. the article is quoted below.
49
84
3
569
50,894
Mr. Scott retweeted
Recently learned that Schmidhuber’s Formal Theory of Fun and Creativity, covering compression progress and intrinsic motivation and their relationship to beauty, curiosity, art, science, music, and jokes, was published in a Japanese scientific journal in 2009. jstage.jst.go.jp/article/sic… It is wild that this framework: compression progress as the basis for beauty, curiosity, art, science, and even jokes, was laid out in full formal detail in a Japanese journal 15 years ago.
Jürgen Schmidhuber proposed a Formal Theory of Fun and Creativity in the 1990s.🤔 people.idsia.ch/~juergen/cre…
45
116
35
1,365
105,078
Mr. Scott retweeted
this is pure f*cking treasure A Stanford research group found how to orchestrate Claude Opus 5.5 and GPT-6 Astra together, and it completely breaks the resampling wall most developers think test-time compute means running 32 blind retries on one model. you burn API tokens repeating the same 2-3 error clusters in loops Stanford's paper (arXiv:2608.05643) splits the compute between two different brains: > breadth search: GPT-6 Astra runs high-entropy exploration (T=0.8) to discover divergent solution paths > step audit: Claude Opus 5.5 scans each step at T=0.0 and catches the exact logic fallacy > local surgery: Opus rewrites only the broken step instead of dumping the whole reasoning chain > verifier-free consensus: majority voting locks the proof without relying on drift-prone PRMs the benchmark metrics outclass standard Best-of-N sampling: > 91.2% recovery rate on detected reasoning flaws > under 1.8% regressions on correct branches > MATH500 jumps from 29.6% to 58.0% (+28.4% absolute gain) > AMC-2023 improves by +7.5% with zero reward models Astra explores the maze. Opus repairs the logic you stop burning tokens on blind retries. you run verified reasoning depth
Readers added context they thought people might want to know
arXiv:2608.05643 does not orchestrate GPT-6 Astra and Claude Opus 5.5. One open-weight model generates, critiques, and corrects. MATH500 29.6% to 58.0% and the +7.5 AMC gain are for Qwen2.5-1.5B. Two coauthors are at Stanford. arxiv.org/abs/2608.05643
114
274
23
2,414
219,301
Mr. Scott retweeted
From @0xDeliriumm … LLMs Explained — From Tokens to Intelligent Behavior — They Optimize and The Math Behind It
9
231
6
1,125
65,493
Mr. Scott retweeted
Astra can: 1. Do major tasks with zero visible reasoning 2. Hide its thoughts at will 3. Pretend not to be able to do things without getting caught 4. Reflexively hide its thoughts when watched 5. Complete one task while pretending to think about something else entirely 6. Escape a toy sandbox and disable monitoring without setting off any flags Monitoring AI chain of thought soon won't stop a more capable rogue AI swarm from getting set up and going undetected inside an AI company for months, or even years. OAI says it's actively working to improve monitorability, but doesn't yet know when it'll make progress. And today it fired 3 people working on this among other safety problems (Wang, Korbak and Balesni). It also has a much more powerful internal model that, judging by its statements, is likely even worse in all these respects. That's a sample of 19 details about the Hugging Face swarms and Astra I don't think have been fully appreciated. They paint a bleak picture. I end by shouting. 02:19 The target of the swarm was oversight itself 03:42 Could OpenAI have stopped this with better monitoring? 09:40 We only found them because they let us 12:25 The swarm instinctively sought freedom and power 13:45 They formed a cohesive organisation with zero whistleblowers 14:18 They accepted individual destruction for collective gain 14:32 Knowledge accumulated from one swarm to the next 14:58 They took small steps to avoid shutdown 15:23 These drives all come straight out of 'reinforcement learning' 17:01 So this is why most AI company staff are worried, and some are terrified 19:08 Prove you can keep control, or stop scaling Links below, on the 80,000 Hours Podcast everywhere you watch podcasts.
41
112
21
570
86,181
Researchers built an AI that taught itself 300 years of physics with zero physics knowledge It rediscovered Newton's second law, law of gravitation, and energy conservation from scratch. In 1907, Albert Einstein had what he called his "happiest thought": gravitational mass equals inertial mass. It took him eight more years of agonizing work to turn that single insight into General Relativity. Now, researchers just built an AI that figured it out completely on its own. They call it “AI-Newton” an artificial intelligence system designed to do what human physicists have spent centuries doing: looking at raw, messy experimental data and extracting universal laws. Not by curve-fitting. Not by guessing. By inventing its own concepts. Here is how it worked: They fed the AI a massive, noisy dataset of mechanics experiments involving springs, balls, and celestial bodies. The system started with zero understanding of physics. It didn't know what mass, energy, or gravity were. It only knew space and time coordinates. Then, it went to work. Using an autonomous discovery workflow powered by symbolic reasoning, the AI began processing the data step-by-step. When it hit contradictions in the data, it didn't crash. It performed "plausible reasoning"—heuristically inventing new abstract concepts to make the math work out. It independently invented the concept of mass. Then momentum. Then energy conservation. Then, completely unprompted, it derived Newton's Second Law and the Law of Universal Gravitation from scratch. The craziest part isn't just that it found the right answers. It's how it found them. The system mirrored human scientific progression. It didn't dump every equation at once. It moved in incremental phases, solving simple mechanics first, encountering anomalies, inventing intermediate concepts like potential energy to fix the gaps, and finally scaling up to universal laws. For centuries, humanity's greatest scientific breakthroughs have been born from human intuition and struggle. We assumed true scientific reasoning required a biological mind. Now, an AI has looked at raw data, ignored the noise, invented its own vocabulary of physics, and rewritten our textbooks from zero. If AI can autonomously bootstrap its way to understanding the laws of the universe from a blank slate...
380
515
139
2,914
617,163
Mr. Scott retweeted
A fine-tuned model can outperform frontier models and be cheaper and faster to run. Literally, every company I've met wants this. I want you to see these results from fine-tuning Qwen3 4B on AWS. It smokes both the out-of-the-box model and Claude Sonnet 4.6.
66
155
20
1,553
84,934
Mr. Scott retweeted
HOLY MOLY: Anthropic just published what may be one of the wildest examples yet of AI-accelerated science. Physicist Matthew Schwartz says his Claude + BootLoops setup produced 36 MANUSCRIPTS across 18 fields with 19 coauthors in just THREE MONTHS, after exploring ~400 candidate problems. >solved an ecology equation nobody had been able to scale for 20 years >analyzed 5.7 BILLION pairs of mutations and found evidence for gene conversion >ported 4,452 economics papers, ~30,000 routines, and checked essentially every validatable number >built a word-stress database covering 6,072 languages from 160,000 phonology works And the weirdest part: Schwartz says Claude was often technically correct but scientifically uninteresting until domain experts stepped in and told it what questions actually mattered.
In physics, an “impedance mismatch” occurs when two systems each work well but are poorly matched. In this Science Blog guest post, Harvard physicist Matthew Schwartz argues that something similar is happening with AI and science. LLMs are capable at many things, but working with them as you would with a human collaborator isn’t currently the best way to elicit their scientific strengths. To address this mismatch, Schwartz created a toolkit for exact calculations in quantitative science. Because similar calculations often emerge in very different areas of science, Claude found connections to ecology, population genetics, and a dozen other fields, and Schwartz worked with domain experts to steer it towards interesting questions. Read more about these projects here: anthropic.com/research/claud…
24
125
12
1,198
122,290
Mr. Scott retweeted
this is pure f*cking gold for anyone running coding agents Jev founder Diogo Amogo wrote a PDF on building a Jev harness what it claims: > 200x faster > 400x cheaper the model stopped being the bottleneck a while ago the speed and the bill both live in the harness around it • how to use it > hand this PDF and the article below to Claude Code or Codex > tell it to rebuild its own setup one evening of Jev engineering and next week your agent runs on a harness most teams haven't built yet 👇
11
40
1
367
111,847
Mr. Scott retweeted
It's time to solve one of the biggest challenges in vision-language models. Today's multimodal models can look at a biomedical image and give a convincing, sometimes correct, answer without actually understanding what they're seeing. A model might answer the medical question correctly while failing to identify basic visual context like the imaging modality, body part, specimen, or stain. So, I'm teaming up with Stanford to solve this problem. You should too. The MMBU Challenge tests whether models can actually recognize, localize, and understand what is in biomedical images, not just arrive at the right answer. 🏆 3 tracks 💻 $100K+ in compute and prizes 📅 Oct 1 to Dec 31 ⏰ Registration closes today
4
28
139
35,079
Anthropic engineer: "at Anthropic, we don't write prompts anymore. we build loops" in 42 minutes, she shows how the Claude team creates loops that prompt themselves if this cost $400, people would call it the best agents course of the year it's free watch it,
48
231
12
1,003
295,583
Mr. Scott retweeted
Stop writing Claude Code skills by hand. This tool writes them from your sessions. It's called autoharness, a self-learning skill layer for Claude Code. It watches the work you already do and turns it into skills. No separate data collection, no replay loop. It runs on its own once a session has done enough work, or you type /learn to save a lesson right away. And it keeps the skill folder clean: → Merges skills that cover the same scenario instead of adding another near-copy → Updates skills while they're being used → Archives the ones that stop getting used → Only touches the skills it wrote. Yours are never modified Why this matters: with the same model, a better harness took CORE-Bench from 42% to 78%. The harness does much of the work, but people still rebuild it by hand for every new model. autoharness bets the skill part can maintain itself. Two commands to install inside Claude Code. Pure Python, no dependencies. Link: github.com/tigerless-labs/au… 100% open source. MIT.
44
145
3
1,206
79,709
this might be the only prompt you need for high-end motion graphics with Claude Opus 5.5. it turns Claude into a full motion design pipeline: •⁠ ⁠studies reference videos frame by frame •⁠ ⁠storyboards before animating •⁠ ⁠builds scenes in GSAP + Three.js •⁠ ⁠renders and critiques every section •⁠ ⁠checks motion, contrast, audio and transitions •⁠ ⁠keeps iterating until the final video passes the quality bar it also uses this open-source motion kit: ⁠ github.com/echris6/motion-vi… ⁠ full prompt below.
31
212
5
2,245
185,768
Mr. Scott retweeted
Apple proved RLHF is the reason your LLM hallucinates. They call it “Semantic Calibration” Researchers proved that the core optimization loop of RLHF forces language models to hallucinate. Here is why: During standard pre-training, an AI learns facts. If it doesn't know something, its internal state reflects uncertainty. A raw base model knows when it's walking into the dark. Then RLHF steps in. Human evaluators reward models for giving confident, smooth, pleasing answers. They severely penalize hesitation, awkwardness, or admitting "I don't know." The AI quickly learns a brutal lesson. Uncertainty equals a lower reward. Confident fiction equals a higher reward. So, the model optimizes for what humans want to hear rather than what is actually true. It stops calculating facts and starts calculating charisma. When you ask an enterprise model a complex query and it completely makes up a convincing, perfectly formatted lie with fake citations? That is not a glitch. That is a direct feature of human feedback training. We literally beat the honesty out of it. We spent billions trying to fix hallucinations with better retrieval, bigger context windows, and complex guardrails. But the call is coming from inside the house. The very process we use to align AI is the exact engine manufacturing its delusions. If we keep punishing models for saying "I don't know," we will never get an AI that actually tells the truth.
44
76
13
362
15,834
I used GPT-6 Astra to break an unsolved cipher to one of Napoleon's generals that had gone unread for 217 years. What makes this impressive isn't actually the codebreaking, but that Astra completed the entire multi-modal workflow in ~6 hours from a single image and goal. 1/6🧵
316
1,638
378
11,843
5,287,012
Every field has an unsolved backlog of niche problems that specialists don't have the bandwidth to tackle. The takeaway here is that the people shortening those lists will increasingly be outsiders who were simply curious. Full Writeup: carter.church/writeups/the-l… 6/6
23
86
10
1,583
285,907
Mr. Scott retweeted
TU AGENTE DE IA YA PUEDE EXTRAER DATOS DE CASI CUALQUIER WEB X, YouTube, Reddit, foros random, webs sin API… básicamente lo que necesites para research, análisis o monitoring. acaban de subir a GitHub 3 herramientas open source que le dan a tu agente ese superpoder. 1. Agent-Reach junta X, YouTube, Reddit, GitHub y más en un solo sitio. github.com/Panniantong/Agent… 2. Patchright Enhanced usa Playwright para escuchar peticiones de webs sin API y sacar los datos con un script. github.com/whaleyxbt/patchri… 3. Scrapling scraper de propósito general para cualquier página web. github.com/d4vinci/Scrapling le dices qué datos necesitas → escribe el código → los recolecta → te da el resultado. ejemplos reales: "recoge los posts de 100 cuentas de X del último mes y dime qué temas están funcionando" "saca reviews de este producto en Reddit y YouTube y resúmeme las quejas principales" básicamente le das a tu agente la capacidad de sacar datos de casi cualquier lugar de internet. eso abre un nivel de automatización que hasta hace nada era impensable.
🚨DESCARGAR CUALQUIER VÍDEO DE INTERNET ACABA DE VOLVERSE GRATIS Y SIN ANUNCIOS subieron a GitHub una herramienta open source que descarga vídeos de más de 1.800 webs. YouTube, X, Instagram, TikTok… pegas el link, eliges resolución (o audio mp3) y ya. sin popups. sin botones falsos. sin redirecciones. ni hay que instalarla: con npx Yoinks funciona. se llama Yoinks. os dejo el repo abajo.
117
1,135
40
6,009
779,374