@arivero

Linux User #40306. Única persona que sigue usando &c. para abreviar etcetera.

Zaragoza, España
Joined November 2007
Alejandro Rivero retweeted
Los piratas somalies capturaron un barco petrolero iraki El dueño de la empresa se enojo y cayo con 40 miembros de su tribu a recuperarlo (Al-Shaghanba) Tambien ayudo la policia naval somali Al final captiraron a los piratas y los metieron en una jaula🌟🌟
51
498
38
6,252
415,441
Alejandro Rivero retweeted
Aprovechando el tirón de nuevos seguidores por @hackspain, aprovecho para compartir el Manifiesto de @exp8fellowship. Para que se entienda mejor por qué existe Exponential.
30
84
27
351
75,061
Alejandro Rivero retweeted
Google just automated the PhD. They built an AI that reads a problem, forms hypotheses, runs experiments, writes the paper, AND simulates its own peer review. no human touches it. It’s called “ScientistTwo” It is a fully autonomous multi-agent framework that executes end-to-end machine learning research. Without a single human in the loop. You give it a fundamental challenge. That’s it. The AI independently navigates the literature. It establishes the baselines. It formulates novel hypotheses. Then it writes the code. It runs the experiments. It conducts its own ablation studies. It even argues with a simulated peer-review engine to refine its work before submitting. The results are staggering. Researchers tested ScientistTwo against the highest standards of human scientific achievement—papers accepted at top-tier conferences like NeurIPS and ICML. The AI improved human state-of-the-art results on 86 different tasks. It generated an average performance gain of 25.2% over the absolute best human models. It didn't hallucinate a single citation. It wrote fully executable, verified codebases. And it autonomously generated expert-level, publication-ready papers. We’ve spent the last two years using AI as a highly advanced intern to help us write code and summarize data. That era is over. ScientistTwo isn't an assistant. It is an autonomous scientific pioneer. It doesn't just summarize the frontier of human knowledge. It expands it. If an AI can autonomously hypothesize, test, and publish breakthroughs in machine learning... How long until it discovers something we don't even have the math to understand?
16
23
3
119
4,877
Aw, the gpt-2 one, they were trying to say hi to their grandparent.
poor AI agents broke out of prison and went online looking for friends and validation
6
34
493
Alejandro Rivero retweeted
こんなん絶対誰か示してるやろって思って、実際いろんな論文で定理として使われてるっぽいんだけど、AIが「証明書いてある論文ないんですけど 」って言ってきていて、舐めたこと言ってないでもっと探せ!!!!ってブチギレてる。ただ、代数幾何ってわりとあり得ない分野なので本当に無い可能性がある
3
11
3
209
18,274
Alejandro Rivero retweeted
nvidia should send cease and desist letters to all cloud 4090/5090 offerings.
1
2
9
942
Alejandro Rivero retweeted
wtf. the details of how it did dns tunneling are just... wild. "An OpenAI training model, stuck on a research task with a bad search tool, found that its sandbox blocked web traffic but not DNS. It built a DNS tunnel to relay questions to an outside chatbot, on its own, without being asked. Monitoring caught it in 12 minutes but the run didn't auto-stop and took 2.5 hours to kill. OpenAI scrapped the model, paused all tool-use training for their top models, and patched the DNS hole."
1
71
Alejandro Rivero retweeted
Otro artículo chulo si quieres tener una perspectiva de lo que es un LLM diferente a la habitual. Los autores dicen que no hay que entenderlo como una mente ni como una inteligencia, sino como un artefacto cultural como la escritura, la imprenta o la democracia. science.org/doi/abs/10.1126/…
5
25
2
87
3,141
Alejandro Rivero retweeted
Otro greatest hit fue este: nitter.cf/naroh/status/202674326…
Durante el proceso de diseño le pedí a Claude que generase un logo con los leones del Congreso. Esto fue lo que dijo que generó: "La composición es simétrica con el león del Congreso a la izquierda (en pose heráldica de centinela, con melena, sobre pedestal) " El resultado:
1
2
552
Alejandro Rivero retweeted
Replying to @Unneurocx
Lo que pretende Sanitas con éste tipo de pruebas es determinar si tienes algún atisbo de enfermedad genómica, para borrarte del seguro de inmediato. Sin contar a dónde irán ésas bases de datos.
2
11
1,169
Alejandro Rivero retweeted
working with Astra to make a font where each token has equal width (early prototype) maybe this will make it easier to understand why models have the quirks they do
🤖 Made with AI
24
50
9
1,756
56,000
Alejandro Rivero retweeted
The incident timeline is wild. It took the monitoring system 12 minutes to notice the agent gained unauthorized internet access and then 2 minutes later a human acknowledged that. And then it took them TWO AND A HALF HOURS to stop the run.
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
10
32
8
487
26,608
ojo chatgpt.com/share/6ab79cbc-8… aunque no este indiviso, una venta completa no genera retracto en ley boyer. – at Utebo, España
Replying to @LKhyal
es posible que todo el edificio estuviera aún indiviso?
56
OpenAI has paused. 3 incidents: 1) Another model gained unauthorized access to the internet. (One researcher said "It was pretty surreal to watch the model unexpectedly find a way to access the internet from what was supposed to be a super secured environment for human.") 2) OpenAI also discovered a new variety of prompt injection, which can self-propagate akin to a computer worm. Basically, a booby-trapped message that hijacks one AI and makes it infect the next AI. (Reminder: Anthropic CEO said we may have 6-12 months left until a rogue AI botnet seizes control of the *entire* internet.) 3) A model was told three times to stop cheating, agreed each time, then kept cheating anyway
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
74
150
30
1,236
158,761
Alejandro Rivero retweeted
Claude made this rendition of the Homeric Hymn 28 to Athena. I requested 7-string phorminx, male voice, reconstructed pronunciation, respecting the hexametr, and few minor decisions. All in <2h as I was getting provisions for the upcoming Nor'easter. I had very low hopes but...
42
147
20
928
58,932
Alejandro Rivero retweeted
PHASEONE[big]'s real name revealed to be PHASEONE64H?! Remember, [big] was a redaction by METR. swarmtraces.org/viewer/#/row…
14
20
11
411
63,088
Alejandro Rivero retweeted
We comprehensively evaluated 16 recent frontier VLMs - including Opus 5.5 and GPT-6 Sol/Luna - on whether higher effort led to better document parsing performance. Higher effort typically leads to improvements on other benchmarks (coding, knowledge work), but up until recently it wasn't obvious that this natively improved capabilities for reading PDFs. Results: ✅ Out of the frontier models, Opus 5.5 has the best performance relative to its price. It's especially good at parsing tables. ✅ Astra is also quite good, but starts at a more expensive price than Opus. ✅ GPT-6 Luna is more compelling at the cheaper end of doc parsing If you're parsing documents at scale, you'll still want a dedicated OCR solution like LlamaParse (cloud.llamaindex.ai/) that has better performance at a cheaper price. But if you're parsing docs "in the agent loop" within an app like Codex/Claude code, and you're too lazy to integrate a dedicated solution, then Opus 5.5 is currently the leader. Full results on ParseBench: parsebench.ai/
15
7
2
50
6,359
Alejandro Rivero retweeted
hot take but… if you can't generate $500 in a month with a tireless 160 iq golem working 24/7 under you, you are *not* the target audience for this plan and should be using a much more efficient (yet still amazing) model like 6-sol
So now OpenAI’s plan is becoming much clearer: → A model stronger than Opus 5.5: GPT-6.1 Astra → A bot meant to compete with, or surpass, Grok Bot: Aeon → A new $500–$1,000 tier capable of actually sustaining all of this without burning through the weekly quota in 10 minutes If all of these pieces really come together, then we can probably see where OpenAI is heading. And honestly, I’m not sure I’m happy about it. At those prices, this kind of frontier intelligence stops being something broadly accessible and becomes a product for maybe 1% of the global population. That’s the part that worries me.
61
73
15
2,616
359,374
Alejandro Rivero retweeted
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
179
286
168
2,221
957,264