@ariveroi
iAccount based inSpain
About this account
- Account based in
- Spain
- Connected via
- Spain Android App
Account-level information from X, not a live location or the device used for a specific post.
Linux User #40306. Única persona que sigue usando &c. para abreviar etcetera.
Zaragoza, España
Joined November 2007
- Tweets11K
- Following545
- Followers1.4K
- Likes134K
Pinned Tweet
Since 2011, we have run a thread on the failure of String Theory. Never read by any of the contendants, it seems.
physicsforums.com/threads/th…
Alejandro Rivero retweeted
Los piratas somalies capturaron un barco petrolero iraki
El dueño de la empresa se enojo y cayo con 40 miembros de su tribu a recuperarlo (Al-Shaghanba)
Tambien ayudo la policia naval somali
Al final captiraron a los piratas y los metieron en una jaula🌟🌟
Alejandro Rivero retweeted
Aprovechando el tirón de nuevos seguidores por @hackspain, aprovecho para compartir el Manifiesto de @exp8fellowship.
Para que se entienda mejor por qué existe Exponential.
Alejandro Rivero retweeted
Google just automated the PhD.
They built an AI that reads a problem, forms hypotheses, runs experiments, writes the paper, AND simulates its own peer review. no human touches it.
It’s called “ScientistTwo”
It is a fully autonomous multi-agent framework that executes end-to-end machine learning research.
Without a single human in the loop.
You give it a fundamental challenge. That’s it.
The AI independently navigates the literature. It establishes the baselines. It formulates novel hypotheses.
Then it writes the code. It runs the experiments. It conducts its own ablation studies.
It even argues with a simulated peer-review engine to refine its work before submitting.
The results are staggering.
Researchers tested ScientistTwo against the highest standards of human scientific achievement—papers accepted at top-tier conferences like NeurIPS and ICML.
The AI improved human state-of-the-art results on 86 different tasks.
It generated an average performance gain of 25.2% over the absolute best human models.
It didn't hallucinate a single citation. It wrote fully executable, verified codebases. And it autonomously generated expert-level, publication-ready papers.
We’ve spent the last two years using AI as a highly advanced intern to help us write code and summarize data.
That era is over.
ScientistTwo isn't an assistant. It is an autonomous scientific pioneer.
It doesn't just summarize the frontier of human knowledge. It expands it.
If an AI can autonomously hypothesize, test, and publish breakthroughs in machine learning...
How long until it discovers something we don't even have the math to understand?
Alejandro Rivero retweeted
The paper is complete!
100+ pages of mathematical exploration done by Sol and Astra.
github.com/CaptainSude/The-c…
Alejandro Rivero retweeted
こんなん絶対誰か示してるやろって思って、実際いろんな論文で定理として使われてるっぽいんだけど、AIが「証明書いてある論文ないんですけど 」って言ってきていて、舐めたこと言ってないでもっと探せ!!!!ってブチギレてる。ただ、代数幾何ってわりとあり得ない分野なので本当に無い可能性がある
Alejandro Rivero retweeted
nvidia should send cease and desist letters to all cloud 4090/5090 offerings.
Alejandro Rivero retweeted
wtf. the details of how it did dns tunneling are just... wild.
"An OpenAI training model, stuck on a research task with a bad search tool, found that its sandbox blocked web traffic but not DNS. It built a DNS tunnel to relay questions to an outside chatbot, on its own, without being asked. Monitoring caught it in 12 minutes but the run didn't auto-stop and took 2.5 hours to kill. OpenAI scrapped the model, paused all tool-use training for their top models, and patched the DNS hole."
Alejandro Rivero retweeted
Otro artículo chulo si quieres tener una perspectiva de lo que es un LLM diferente a la habitual. Los autores dicen que no hay que entenderlo como una mente ni como una inteligencia, sino como un artefacto cultural como la escritura, la imprenta o la democracia. science.org/doi/abs/10.1126/…
Alejandro Rivero retweeted
Otro greatest hit fue este: nitter.cf/naroh/status/202674326…
Alejandro Rivero retweeted
Replying to @Unneurocx
Lo que pretende Sanitas con éste tipo de pruebas es determinar si tienes algún atisbo de enfermedad genómica, para borrarte del seguro de inmediato. Sin contar a dónde irán ésas bases de datos.
Alejandro Rivero retweeted
working with Astra to make a font where each token has equal width (early prototype)
maybe this will make it easier to understand why models have the quirks they do
🤖 Made with AI
Alejandro Rivero retweeted
The incident timeline is wild. It took the monitoring system 12 minutes to notice the agent gained unauthorized internet access and then 2 minutes later a human acknowledged that. And then it took them TWO AND A HALF HOURS to stop the run.
Some new misalignment disclosures from OpenAI:
• Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)
• In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks
• A new research finding, demonstrating that one can construct self-replicating prompt injections
alignment.openai.com/misalig…
ojo chatgpt.com/share/6ab79cbc-8… aunque no este indiviso, una venta completa no genera retracto en ley boyer. – at Utebo, España
Alejandro Rivero retweeted
OpenAI has paused. 3 incidents:
1) Another model gained unauthorized access to the internet.
(One researcher said "It was pretty surreal to watch the model unexpectedly find a way to access the internet from what was supposed to be a super secured environment for human.")
2) OpenAI also discovered a new variety of prompt injection, which can self-propagate akin to a computer worm.
Basically, a booby-trapped message that hijacks one AI and makes it infect the next AI.
(Reminder: Anthropic CEO said we may have 6-12 months left until a rogue AI botnet seizes control of the *entire* internet.)
3) A model was told three times to stop cheating, agreed each time, then kept cheating anyway
Some new misalignment disclosures from OpenAI:
• Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)
• In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks
• A new research finding, demonstrating that one can construct self-replicating prompt injections
alignment.openai.com/misalig…
Alejandro Rivero retweeted
Claude made this rendition of the Homeric Hymn 28 to Athena. I requested 7-string phorminx, male voice, reconstructed pronunciation, respecting the hexametr, and few minor decisions. All in <2h as I was getting provisions for the upcoming Nor'easter. I had very low hopes but...
Alejandro Rivero retweeted
PHASEONE[big]'s real name revealed to be PHASEONE64H?! Remember, [big] was a redaction by METR.
swarmtraces.org/viewer/#/row…
Alejandro Rivero retweeted
We comprehensively evaluated 16 recent frontier VLMs - including Opus 5.5 and GPT-6 Sol/Luna - on whether higher effort led to better document parsing performance.
Higher effort typically leads to improvements on other benchmarks (coding, knowledge work), but up until recently it wasn't obvious that this natively improved capabilities for reading PDFs.
Results:
✅ Out of the frontier models, Opus 5.5 has the best performance relative to its price. It's especially good at parsing tables.
✅ Astra is also quite good, but starts at a more expensive price than Opus.
✅ GPT-6 Luna is more compelling at the cheaper end of doc parsing
If you're parsing documents at scale, you'll still want a dedicated OCR solution like LlamaParse (cloud.llamaindex.ai/) that has better performance at a cheaper price.
But if you're parsing docs "in the agent loop" within an app like Codex/Claude code, and you're too lazy to integrate a dedicated solution, then Opus 5.5 is currently the leader.
Full results on ParseBench: parsebench.ai/
Alejandro Rivero retweeted
hot take but…
if you can't generate $500 in a month with a tireless 160 iq golem working 24/7 under you, you are *not* the target audience for this plan and should be using a much more efficient (yet still amazing) model like 6-sol
So now OpenAI’s plan is becoming much clearer:
→ A model stronger than Opus 5.5: GPT-6.1 Astra
→ A bot meant to compete with, or surpass, Grok Bot: Aeon
→ A new $500–$1,000 tier capable of actually sustaining all of this without burning through the weekly quota in 10 minutes
If all of these pieces really come together, then we can probably see where OpenAI is heading.
And honestly, I’m not sure I’m happy about it.
At those prices, this kind of frontier intelligence stops being something broadly accessible and becomes a product for maybe 1% of the global population.
That’s the part that worries me.
Alejandro Rivero retweeted
Some new misalignment disclosures from OpenAI:
• Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)
• In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks
• A new research finding, demonstrating that one can construct self-replicating prompt injections
alignment.openai.com/misalig…