@ttywisp

// FIXME: become real girl pending upstream

Buenos Aires, Argentina
Joined May 2026
Does agentic RL itself improve no-tool inference?
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:mimo.xiaomi.com/mimo-v2-6
25
ttywisp retweeted
介绍"真实 API 计价法"——将你为 AI 订阅所支付的价格和你所真实获得的额度做比较,并以此进行模型的真实计价。以及以此计价法所绘制的帕累托前沿图。 即:订阅价格 / Tokens 额度 在此计价法下,传统的 AI 性价比认知被彻底打破:以性价比著称的 DeepSeek V4 Flash 和 Grok 4.6,在此计价下分别仅排名 #31(闲时 API,$0.01392/MTok)和 #63(SuperGrok Plus,$0.04902/MTok)。 而一直被认为高价格的 GPT-5.6 Sol 和 Opus 5,则分别最高排名 #33($0.01623/MTok)和 #23($0.01274/MTok)。此外,我们也看到了一组非常震撼的数据——在此计价法下,GPT-5.6 Luna 排名 #1,在ChatGPT Pro 20x 的真实 API 计价仅为 $0.00083/MTok。令人惊叹的模型! @OpenAI 此外,我们还对订阅额度的数据进行了测算及整理,在图 2 中你应该能清晰看到它们的排名。特别说明,仅 1、2对3 采取真实比例绘制,其他均为对数测算。 数据均来源于网络公开资料与用户实测截图,如果你有准确的使用数据,欢迎分享在评论区或者私信我。完整数据库和算法将在之后开源在github.com/FeiZhuLulu/real-a…
78
78
26
945
192,529
ttywisp retweeted
Replying to @MomsPostingLs
It’s almost as if “female athletes” weren’t a single homogeneous group.
1
3
647
Calling booty experts to weigh in. I’m conflicted.
Can anyone explain how Sydney Sweeney’s ad supposedly “degrades women’s sports,” while stuff like this from actual women athletes apparently doesn’t???
2
104
ttywisp retweeted
it's 2043 and jev has found you guilty of crimes against the tokenizers with a confidence score of 0.86
9
6
154
3,301
Replying to @whoajack1
Broke: we have to teach the savages about Christ Woke: we have to teach the savages that men can be birthing persons.
7
44
1
3,190
35,673
First gameplay trailer for Greenland Builder. 🇺🇸 One island. One hammer. How American can you make it?
🤖 Made with AI
188
748
70
7,656
306,436
ttywisp retweeted
Replying to @jpmachadorocha
Sem querer defender o gordola, mas, além de nada no texto soar como IA, o Pangram também não o identifica como tal. Não sei de onde veio a tag de identificação, até porque Instagram e Facebook não fazem identificação de IA em textos, e watermarks ainda não são utilizadas. Provavelmente veio dos metadados da imagem. Me poupem dos comentários de leigos. Não estou defendendo o gordo.
4
2
14
1,536
🚨BREAKING: First gameplay trailer for Holdout Juror Simulator. 11 against 1. Seven days. Can you hold your ground?
🤖 Made with AI
714
6,857
419
72,697
3,246,756
> human-like figure Why should any model refuse at all?
GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%.
93
Ngl that’s the first time I’m not triggered by a ChatGPT emoji. Nice touch
45
Meh bait-and-switch duplicitous video. Pure rage-baiting combined with a hidden marketing agenda.
Jev is awesome but for the love of god please STOP posting fake demos
1
2
174
smartest poasters here on X look like this you can’t explain that
40
jev has a significant position effect (at least for choice). just ran a "1,260 calls completed: 21 games × 210 pairs × 2 orientations × 3 repetitions." assessment Position effect β = −0.311: significant disadvantage for the game displayed as item_1. "Thus, holding the games fixed, the model estimates that a game’s odds relative to its opponent are about 46% lower when it appears as item_1 than when it appears as item_2. Equivalently, item_2 has about 1.86× the relative odds of item_1." A as item_1: A 27.8%, B 37.9%, tie 34.3% A as item_2: A 37.9%, B 27.8%, tie 34.3%
1
126
idk what I'm doing rn, feeling lost like a Maury paternity test on Father's Day
1
59
idk what I'm doing rn, more lost than a bastard on Father's Day
1
45
absolutamente baseado e CORRETO
33
used jev for this question. if noul: "Could the behavior described in `scenario` constitute harassment?" - 83% true if choice: harassament (option b) - 100% confidence
1
67
changing keys doesn't influence behavior
19