@TelonAlex

Follow for elegant code, AI papers, python, math, books and meta-humor.

Linköping, Sverige
Joined January 2013
A todo cli in 10 lines of python.
2
2
1
7
924
Blogman explains the halting problem. Tweetguy say lol this shit is made overcomplicated. Look this program here did in fact show my example code will halt! What a dumb blogman! Am I close?
Someone liked this thread of mine out of the blue from years ago, and I realized that it’s still relevant, so I thought I would repost it.
1
664
An "off switch" will get progressively more expensive to use as we increase our dependence on AI. Like a country that is no longer self sufficient on food the option to stop trade might become a mirrage. An off switch is not enough if using it is allowed to become unpalatable.
We've long argued for a balanced approach to AI regulation. One that keeps humans in control. Since 2023, @Microsoft has advocated for "safety brakes" or an “off switch" that keeps advanced AI systems under human control. And it should ensure that AI systems that control critical infrastructure and autonomous systems be run in secure cloud infrastructure with layered safeguards that provide additional intervention points as needed. Put simply, powerful AI systems should always remain under human control, including interruption, correction, and shutdown. This approach is tried and tested. In the 1850s, Elisha Otis’ demonstration of a safety brake at the World’s Fair helped earn the public’s trust – and made modern cities possible. We should adopt this principle for AI. If we do, AI can become a powerful tool for human progress. If we don't, a lack of public trust, not regulation, will become the limiting factor that slows AI adoption.
14
Explaining something is to express a point in space using different word-vectors. But each word is more a transformation using the last point as input. So unlike fixed vectors it's path dependent. Golfing is to find minimal descriptions that gets you roughly to the same space.
1
14
It's easier to think of each unique word as a distinct vector in 3D space that eventually leads to some point. You could invent a new word and declare its defined as the meaning of some sentence. Then you can build upon that. And explore where this new tool can lead you!
1
9
Ofc "in reality" it's not 3D buy ~10000D and your new word can be almost orthogonal to everything. In 3D you cant discover a new direction to travel in. But in high dimensional space it's easy. It's unintuitively unexplored!
4
Code-golfing english seems useful for finding new words. Code-golfing language seems useful for finding new words. Code-golfing language highlights holes in your language. Word-golfing seems useful for expanding language Golfing language finds new lexemes
1
9
Alex Telon retweeted
"stochastic parrot" was a mimetically-fit cognitive virus that spread from 2021-2025; it temporarily blinded many gifted people to the nature of AI progress, burning up crucial years in which they could have helped think through the response to the situation.
275
351
170
3,457
758,257
I did not have the time now to read this in full but directionally this sounds like a great idea!
1
92
Sounds great!
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
10
A good writeup and counterpoint on some of the recent news regarding the hugging face incident. Alignment is very important imo and the recent incidents give concrete examples of what can go wrong. But there are good arguments here for that the models were in fact aligned.
1
1
31
The main argument is that if the model is allowed to reward hack on internal tests. And the model is told that it has no internet access and it's in a simulated environment. Here it reward hacked. But when told it can access the internet it did not.
1
17
The argument as such is good. Though I'm not sure what actual cases it covers. My recollection of METRs report gave me the impression that the models beloved they had real internet access. Not some elaborate simulated internet.
7
someone at ANTHROPIC just showed CLAUDE finding ZERO DAY vulnerabilities in a live conference demo claude has found zero day in Ghost, 50,000 stars on github, never had a critical security vulnerability in its entire, history... it found the blind SQL injection in 90 minutes, stole the admin api key, then did the exact, same thing to the linux kernel
292
1,269
267
11,440
1,906,890
Alex Telon retweeted
Replying to @jeremywei
Love the word "comprehension debt", haven't encountered it so far, it's very accurate. It's so very tempting to just move on when the LLM one-shotted something that seems to work ok.
33
44
20
1,369
130,797
Good visualisation! To anyone reading remember that one could argue that our brains also don't have anything that maps this spiral in our "neural weights" directly. We understand spirals with higher level reasoning. So the spiral understanding could come later! Hence "suboptimal"
One of the best visual explanations I've ever seen for why scaling Transformers works, but is suboptimal, as it's just brute-forcing things, by @YesThisIsLion (co-author of the Transformer) on @MLStreetTalk "In the (rejected) paper "Intelligent Matrix Exponentiation", they show the decision boundary of a classic MLP with a ReLu/Tanh activation function on the classic Spiral dataset." "You can see they both technically solve it with great scores on the test set. Next, they show the decision boundary of the "M-layer" they propose in the paper. And it represents the spiral ... as a spiral!" "Shouldn't we? If the data is a spiral... shouldn't we represent it as a spiral?" "If you look back at the decision boundaries of the MLP, it's clear that you just have these tiny, piecewise separations without learning the concept of a spiral. That's what I mean!" "If you train these things enough, it can fit the spiral and get a high accuracy. But there's no indication that the MLP actually understands a spiral. When you represent it as a spiral, it extrapolates correctly, cause the spiral just keeps going out."
45
Alex Telon retweeted
One of the best visual explanations I've ever seen for why scaling Transformers works, but is suboptimal, as it's just brute-forcing things, by @YesThisIsLion (co-author of the Transformer) on @MLStreetTalk "In the (rejected) paper "Intelligent Matrix Exponentiation", they show the decision boundary of a classic MLP with a ReLu/Tanh activation function on the classic Spiral dataset." "You can see they both technically solve it with great scores on the test set. Next, they show the decision boundary of the "M-layer" they propose in the paper. And it represents the spiral ... as a spiral!" "Shouldn't we? If the data is a spiral... shouldn't we represent it as a spiral?" "If you look back at the decision boundaries of the MLP, it's clear that you just have these tiny, piecewise separations without learning the concept of a spiral. That's what I mean!" "If you train these things enough, it can fit the spiral and get a high accuracy. But there's no indication that the MLP actually understands a spiral. When you represent it as a spiral, it extrapolates correctly, cause the spiral just keeps going out."
27
49
9
618
86,911
Very interesting paper! When I have time I would hope to dive deeper to see what categories of data had the most signal and to what degree it overlaps with what a smartwatch could gather. But I found that the data is not from random people so its very important to consider!
Today in @NatureMedicine we report that AI can predict 130 diseases from 1 night of sleep🛌 We trained a foundation model (#SleepFM) on 585K hours of sleep recordings from 65K people—brain, heart, muscle & breathing signals combined. AI learns the language of sleep🧵
3
130
Alex Telon retweeted
Opus 4.5 has a 50%-time horizon of 4 hours 49 minutes on METR, which means that it can successfully complete certain software engineering tasks that take a human 4 hours 49 minutes to complete around 50% of the time. ... but its 80%-time horizon is MUCH shorter: just 27 minutes. Very high ceiling, BIG error rate.
Replying to @METR_Evals
Despite its high 50%-time horizon, Opus 4.5's 80%-time horizon is only 27 minutes, similar to past models and below GPT-5.1-Codex-Max's 32 mins. The gap between its 50%- and 80%- horizons reflects a flatter logistic success curve, as Opus differentially succeeds on longer tasks.
23
24
1
539
77,749