@PurpleCodeSH

🎨 Innovative Frontend Developer and Robotics Engineering Expert 🤖 Passionate about crafting unique apps and pushing creative boundaries through code 💫

Distrito Federal, México
Joined January 2024
Claude Opus 5.5 is the new #1 in the Artificial Analysis Coding Agent Index, with gains across all three evaluations, though at a higher Cost per Task At max effort in Claude Code, Opus 5.5 scores 66 on the Coding Agent Index, the highest score we have measured. It is up 6 points against Opus 5 (60) and 4 points against Claude Fable 5.1 (62). Anthropic has cut Opus pricing to $4/$20 per million input/output tokens, from $5/$25 for Opus 5, and cache reads to $0.20 from $0.50. Even with those reductions, Opus 5.5’s Cost per Task is $13.04, above Opus 5’s $10.79, because it uses substantially more tokens. Key takeaways: ➤ Improves across all three Coding Agent Index evaluations: Terminal-Bench 4.0 rises to 63.1% from 54.5% for Opus 5, DeepSWE v1.1 to 68.4% from 62.5%, and SWE-Atlas-QnA to 66.4% from 62.1%. The largest gain is on Terminal-Bench, at +8.6 percentage points. ➤ The top score comes at the highest Cost per Task: Opus 5.5’s Cost per Task is $13.04, up 21% from Opus 5 at $10.79. It uses about 15.6 million tokens per task against 11.4 million for Opus 5, including about 2.4× as many output tokens. ➤ Extends the Coding Agent Index vs Cost per Task Pareto frontier: No lower-cost model in our comparison matches Opus 5.5's score. It moves the frontier upward at its high-cost end. Other model details: ➤ Pricing: $4/$20 per million input/output tokens, down 20% from Opus 5. Cache reads cost $0.20 per million, down 60% from $0.50. ➤ Evaluation setup: Claude Code at max effort, measured on DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA. The Coding Agent Index gives each evaluation equal weight.
85
68
19
1,002
63,696
Purple-Code-sh retweeted
This is crazy good
Thank you @claudeai! GIF animated in Python and rendered in Blender by Claude Opus 5.5
3
31
4,021
Purple-Code-sh retweeted
Me está gustando tantísimo todo lo que estoy viendo de Opus 5.5 en cuanto a animaciones, 3D, calidad visual en general, que me lo voy a gozar muy fuerte este modelo! Qué locura
37
61
2
1,530
90,986
Purple-Code-sh retweeted
MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today. Less prefill, a smaller KV cache, better long-context retrieval—and we got all three at once. Compared with MiMo-V2.6's Hybrid SWA architecture: • 5.02× lower prefill FLOPs at 1M tokens • 4.5× smaller KV cache at 1M tokens • Better MRCRv2 and RULER-v2 scores, plus lower AgentPPL and LongPPL Why build a new architecture? Agentic inference is a very different workload. Each round, a short action can return a long observation that needs to be prefilled, while the context keeps growing. That puts prefill cost, KV-cache size, and retrieval accuracy on the critical path at the same time. HySparse2 tackles all three with two levels of KV sharing: • KV Bridging: Following YOCO, full-attention layers in the cross-decoder build their K/V from self-decoder hidden states. • KV Reuse: Within each hybrid block, sparse layers reuse the preceding full-attention layer's KV cache and selection indices. Two more changes: token-level selection replaces block-level selection, and a forced window of recent tokens replaces the separate SWA branch, so local and global tokens share one KV cache. Since all cross-decoder KV caches now come from the self-decoder, prefill can stop once the self-decoder finishes. Paper: arxiv.org/pdf/2609.26368
211
384
110
4,350
392,528
Purple-Code-sh retweeted
Google ha publicado AX, un orquestador de agentes de IA programado en Go. La idea: un agente trabaja horas, guarda archivos y llama a un modelo. Si nadie lo vigila, te puede fundir el dinero en bucle. Defines la tarea en un YAML, lo ejecuta en un sandbox y le pone límites de consumo y procesamiento. → github.com/google/ax
20
87
2
827
43,277
Gente hermosa!! Nueva movida que nos haría muchísima ilusión 🥹 Si tienes algún proyecto el cual ha sido hecho mediante gentle-ai, nos encantaría otorgarte esta certificación para que puedas poner en tu repositorio! Gracias a todos los que nos ayudan y comparten el día a día con nosotros ❤️‍🔥 github.com/Gentleman-Program…
33
33
2
386
15,577
Purple-Code-sh retweeted
Yo uso #gentleai de @G_Programming. Y tu?
1
10
3
22
12,936
Purple-Code-sh retweeted
portals
7
11
113
2,098
The Intelligence Index vs Cost per Task Pareto frontier shifted this week with the releases of MiMo-V2.6-Pro, Claude Opus 5.5, GPT-6 Luna, and GPT-6 Sol Together they have established eleven new points on the Pareto frontier (driven by different reasoning efforts): five from GPT-6 Luna, one each from MiMo-V2.6-Pro and GPT-6 Sol, and four from Claude Opus 5.5. GPT-6 Luna (max) scores 37 at $0.068 per task, MiMo-V2.6-Pro scores 46 at $0.13, GPT-6 Sol (max) scores 48 at $1.06, and Claude Opus 5.5 (max with fallback) is the new highest-scoring model at 58 at $5.98.
62
87
28
1,317
88,047
Purple-Code-sh retweeted
Deploying DiffusionGemma-Jev (djev) just got a lot easier. You can now spin up a Jev API-compatible endpoint on Google Cloud Run using a single command. Performance is solid: ~35-60 ms for single step latency and batch@32 is ~100-123 requests/sec. It's a straightforward way to experiment without needing your own GPU. Runs at roughly $3/hr and drops to $0 when idle. Get the code and instructions here: github.com/taeold/djev-run
101
655
134
6,977
632,936
Otra tanda de releases en el ecosistema Gentle. Te lo resumo en fácil. La novedad grande de gentle-ai (v3.7.0): ahora elegís qué modelo usa cada revisor. En Claude Code podés asignarle un modelo a cada uno de los seis revisores. En OpenCode editás los modelos de los agentes que implementan y exploran. Y en Codex configurás los modelos de los agentes de ODD y de los revisores, con presets de GPT-6 incluidos. ¿Para qué sirve? Para decidir vos dónde gastar. Un modelo potente para el revisor que busca riesgos, uno más barato para el que mira legibilidad. Y si el modelo que guardaste no está disponible, cae a uno soportado en vez de romperse. En las versiones del medio (v3.6.x) hubo arreglos que se agradecen: el hook de Claude Code ya anda bien en Windows con PowerShell, y el sync ahora saca solo el paquete de preguntas de terceros que chocaba con el nativo de Gentle Shell. El aviso del post anterior, resuelto sin que tengas que hacer nada. Y Gentle Shell se volvió independiente de verdad: Tiene instalación propia. Con npm i -g gentle-pi te queda el comando gentle-shell, que arranca en su propia carpeta sin tocar tu Pi de siempre. Si preferís usar tus logins, tus modelos locales y tus chats de siempre, lo abrís con gentle-shell --link. Se configura solo. La primera vez que lo abrís instala lo que necesita y arranca, con el tema Gentleman-Cute por defecto. Si cambia la versión, se vuelve a preparar solo. Y volvió la revisión adentro de Pi para los modelos que vienen de extensiones, así cualquier proveedor que uses para trabajar también te sirve para revisar. Para actualizar: brew upgrade gentle-ai && gentle-ai sync npm i -g gentle-pi y después gentle-shell (o pi install npm:gentle-pi si ya lo usás dentro de Pi) Notas completas: github.com/Gentleman-Program… github.com/Gentleman-Program…
13
8
2
129
4,879
Purple-Code-sh retweeted
Jev (@typesafeai) is so insane & cheap for search!! > 6000+ @ycombinator Startups indexed. > Sub 1 second search results. > 90M tokens & $2.7 in total testing costs. Search any startup in a second, in any way! - Color - Niche - Your Competitor - Age - Image - etc... > watch the entire video, it's so freaking cool omg! > this is the coolest thing i have ever built for fun! (worked on it for 2 days straight!)
79
54
18
1,676
125,295
Purple-Code-sh retweeted
MiMo V2.6 Flash is free for the next week Both Flash and Pro are also available in Go
161
254
99
6,177
572,719
Purple-Code-sh retweeted
Hemos intentado frenar la enfermedad bloqueando una proteína que se llama Kat6. Esto es de lo más avanzado en epigenética que hay. Que no funcione confirma que el tumor sabe avanzar sin hormonas y es el momento de cambiar de estrategia. Mientras esperábamos algún tratamiento estos meses atrás, el tumor ha aprovechado para extenderse al hígado eso significa que ya no tengo el lujo del tiempo que tenía cuando solo estaba en hueso, porque al ser organo vital puede fallar su funcionamiento. Lo peligroso es que la enfermedad avanza de forma sistémica y descontrolada hay que frenarla. Volvemos a la casilla de salida y a buscar opciones.
47
70
2
631
33,977
Purple-Code-sh retweeted
Por aquí es donde más ayuda necesito
Replying to @miriamgonp
1. Yo no he escrito ningún promp ni ningún agente. Pido a Polarias una tarea y él hace lo necesario para resolverla 2. Como se ve, hay muchas herramientas y agentes que se crean para esas tareas específicas y me falta optimizar su gestión 3. También me falta descentralizar Polaris de Claude code que es con lo que está construido todo 4. También siento que pierde información global entre sesiones y no consigo no tener que repetirle algunas cosas 5. Cuando creo tareas automaticas muchas veces acaban cayendose y no tienen consistencia, por ejemplo leer el mail todos los días o actualizar y optimizarse 1 vez a la semana
20
55
8,521
Purple-Code-sh retweeted
Acabo de publicar el render 3D de mi hígado. Salen 20 lesiones. Soy ingeniera y tengo un cáncer de mama metastásico. De esas 20 solo hay dos medidas en el informe de radiología, una en el segmento II de 20 mm y otra en el IVb de 19, las dos del TC del 8 de septiembre. Las otras 18 las ha sacado Polaris, mi sistema agéntico, y todavía no las ha validado nadie. En mayo un PET no vio nada ahí, en julio ya había metástasis y del TC de septiembre salen las 20. Es impactante que aparezcan tantas en tan poco tiempo. Del hígado el informe escrito dice «múltiples». Del hueso dice «incontables». Y eso es lo que más me ha chocado, que no ponga el número de lesiones. El primer visor que hice fue el de los huesos, y me sirvió de apoyo cuando decidimos con mi equipo qué lesión biopsiar. Ahora ya le he añadido el hígado y el tumor primario de la mama. Los puedes girar tu aqui 👇 helpmiriam.com/lesiones-x Cada vez tengo menos tiempo y espero que estas herramientas ayuden a agilizar procesos. Sigo alucinando que lo que construyo tiene impacto real en mi salud y en las decisiones que tomamos.
216
2,487
74
7,520
861,168
Purple-Code-sh retweeted
Ayudemos a Miriam a que su caso de cáncer llegue a personas que puedan aportar algo a su tratamiento. Ella misma se está construyendo sus propias herramientas de software para poder agilizar los procesos. Cada clic cuenta.
Acabo de publicar el render 3D de mi hígado. Salen 20 lesiones. Soy ingeniera y tengo un cáncer de mama metastásico. De esas 20 solo hay dos medidas en el informe de radiología, una en el segmento II de 20 mm y otra en el IVb de 19, las dos del TC del 8 de septiembre. Las otras 18 las ha sacado Polaris, mi sistema agéntico, y todavía no las ha validado nadie. En mayo un PET no vio nada ahí, en julio ya había metástasis y del TC de septiembre salen las 20. Es impactante que aparezcan tantas en tan poco tiempo. Del hígado el informe escrito dice «múltiples». Del hueso dice «incontables». Y eso es lo que más me ha chocado, que no ponga el número de lesiones. El primer visor que hice fue el de los huesos, y me sirvió de apoyo cuando decidimos con mi equipo qué lesión biopsiar. Ahora ya le he añadido el hígado y el tumor primario de la mama. Los puedes girar tu aqui 👇 helpmiriam.com/lesiones-x Cada vez tengo menos tiempo y espero que estas herramientas ayuden a agilizar procesos. Sigo alucinando que lo que construyo tiene impacto real en mi salud y en las decisiones que tomamos.
4
759
3
2,707
106,254
MiMo-V2.6-Pro debuts as the top open weights model on the Artificial Analysis Intelligence Index (46). At $0.13 per Intelligence Index task, it lands on the Intelligence vs. Cost per Task Pareto frontier @Xiaomi has just released MiMo-V2.6-Pro, an open weights model with major advances in intelligence over its predecessor, MiMo-V2.5-Pro (Intelligence Index: 26). Despite the improvement, it retains the same attractive pricing at $0.435 per 1M input tokens (with a 99% cache-hit discount) and $0.87 per 1M output tokens. This makes MiMo-V2.6-Pro one of the most cost-efficient models to deploy. MiMo-V2.6-Pro is an MoE model with 1.02T total parameters and 42B active parameters. Stay tuned for additional analysis of the model. Check out MiMo-V2.6-Pro full benchmarking breakdown here: artificialanalysis.ai
175
390
260
4,139
888,682
Purple-Code-sh retweeted
Hijo de su puta madre mira que está más arriba que el DeepSeek 4.1 flash y más barato. Lo voy a probar y si es fiable como el DeepSeek, acabo de encontrar mi próximo main 🥰
101
64
9
1,992
305,467
Purple-Code-sh retweeted
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:mimo.xiaomi.com/mimo-v2-6
376
867
477
8,801
1,445,858