@icools

Android / Kotlin / JAVA ,主要分享以上或者相關資訊領域技術或發發牢騷使用,目前對於 AI / LLM 應用也有高度興趣

北部
Joined September 2007
J LIN retweeted
the reality of vibe coding agree?
534
464
208
4,646
1,096,763
LM Studio 的Splash版本讓Qwen3.8 27B可以在原本28 tokens/sec的設備,跑到 65.78 tokens/sec
1
1
109
但有時候又跑到48左右 (on Mac M3 MAX)
2
GPT-6 Luna (Max) at #24
Real-world results are in for GPT-6 Luna (Max). It just landed #24 in the Code Arena: WebDev with 1593 pts! GPT-6 Luna (Max) by @OpenAI is a significant +74pt improvement from GPT-5.6 Luna (xHigh) which is now ranked at #41. The latest Luna delivers on par with lower-cost models like: Gemini 3.7 Flash High and Qwen 3.8-27B, but at a fraction of the price (blended $.40/Mtoken). Congrats to the @OpenAI team!
42
大多公司管理的本质
120
237
25
3,108
197,802
Used ChatGPT Voice with plugins for a couple of hours. It's remarkably good. I have a feeling it's going to fundamentally change my usage of ChatGPT. First, it works with any plugin I've tested. That includes a custom MCP I made that lets me access my Mac from ChatGPT. I asked to pull a photo from my desktop and my Reminders, and it all just worked. Voice mode with tools can also preview images inline. (I'd love to see MCP UIs rendered inline, though.) Second: live bidirectional voice with tools feels great to use. When I asked to search for a file in Superhuman, I was able to interrupt the model, say "no actually it's in Slack", and it pivoted on the fly. Combined with Fast mode, it's even nicer. The result: I can realistically use ChatGPT Voice to drive background tasks that a) just don't work with Siri AI and b) go beyond simple questions and kick off workflows that involve multiple plugins – all without ever typing. I've been waiting for OpenAI to ship this for a while, and they pretty much nailed this.
35
26
3
512
50,973
ChatGPT Voice mode is now supporting Plugins and can be used in ChatGPT Work! Besides that, underlying tasks are now powered by GPT 6 model family. This means that you can finally do a real work without a need to use Remote session via a desktop app. A huge enablement! 🔥
We heard you loud and clear. ChatGPT Voice can now: - Use plugins like your email, calendar, and Slack. - Be powered by GPT-6 Astra, Sol, and Luna. - Be used in ChatGPT Work on web and mobile, so you can create docs, decks, sites, and spreadsheets or tackle complex tasks in the browser, just by talking. Rolling out globally today in the latest version of the app.
11
24
3
471
36,686
Claude Opus 5.5 is the new #1 in the Artificial Analysis Coding Agent Index, with gains across all three evaluations, though at a higher Cost per Task At max effort in Claude Code, Opus 5.5 scores 66 on the Coding Agent Index, the highest score we have measured. It is up 6 points against Opus 5 (60) and 4 points against Claude Fable 5.1 (62). Anthropic has cut Opus pricing to $4/$20 per million input/output tokens, from $5/$25 for Opus 5, and cache reads to $0.20 from $0.50. Even with those reductions, Opus 5.5’s Cost per Task is $13.04, above Opus 5’s $10.79, because it uses substantially more tokens. Key takeaways: ➤ Improves across all three Coding Agent Index evaluations: Terminal-Bench 4.0 rises to 63.1% from 54.5% for Opus 5, DeepSWE v1.1 to 68.4% from 62.5%, and SWE-Atlas-QnA to 66.4% from 62.1%. The largest gain is on Terminal-Bench, at +8.6 percentage points. ➤ The top score comes at the highest Cost per Task: Opus 5.5’s Cost per Task is $13.04, up 21% from Opus 5 at $10.79. It uses about 15.6 million tokens per task against 11.4 million for Opus 5, including about 2.4× as many output tokens. ➤ Extends the Coding Agent Index vs Cost per Task Pareto frontier: No lower-cost model in our comparison matches Opus 5.5's score. It moves the frontier upward at its high-cost end. Other model details: ➤ Pricing: $4/$20 per million input/output tokens, down 20% from Opus 5. Cache reads cost $0.20 per million, down 60% from $0.50. ➤ Evaluation setup: Claude Code at max effort, measured on DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA. The Coding Agent Index gives each evaluation equal weight.
86
68
19
1,003
64,640
The Intelligence Index vs Cost per Task Pareto frontier shifted this week with the releases of MiMo-V2.6-Pro, Claude Opus 5.5, GPT-6 Luna, and GPT-6 Sol Together they have established eleven new points on the Pareto frontier (driven by different reasoning efforts): five from GPT-6 Luna, one each from MiMo-V2.6-Pro and GPT-6 Sol, and four from Claude Opus 5.5. GPT-6 Luna (max) scores 37 at $0.068 per task, MiMo-V2.6-Pro scores 46 at $0.13, GPT-6 Sol (max) scores 48 at $1.06, and Claude Opus 5.5 (max with fallback) is the new highest-scoring model at 58 at $5.98.
62
88
28
1,319
89,056
60Hz gives you ~16.6ms per frame, 120Hz drops that to ~8.3ms. A single main-thread disk read or heavy layout pass during VSYNC-app will trigger dropped frames instantly. Use the Display track in Android Studio Profiler to isolate exact jank causes before your users leave a 1-star review. ⚡️ developer.android.com/studio… #AndroidDev #AppPerformance #Kotlin
1
3
71
3,194
Breaking: Luna and Sol are INSANE for Browser Use 👀 Luna is 22x cheaper than Opus 5.5 with similar performance. Mind that these are extremely long and very hard browser agent tasks, even for humans. GPT 6 models live in a league of its own.
Browser Use Bench v2 Pareto frontier got completely redrawn today > Claude Opus 5.5: 59.4 > GPT‑6 Sol medium: 66.9 (3.5x cheaper) > GPT‑6 Luna xhigh: 57.6 (22x cheaper than Opus) OpenAI is in its own league 🔥 All models available to try on our cloud. [x-axis is log scale]
43
34
8
622
59,269
J LIN retweeted
DiffusionGemma-Jev now runs on vLLM 🚀 Ask yes/no, multiple-choice, or scored questions and get confidence with every answer. vLLM seeds a canvas with the response template, leaves only the answer slots noisy, then reads a probability distribution from every slot in a single denoising step. Huge thanks to @mmastrac for driving this upstream! 🙏 github.com/vllm-project/vllm…
Deploying DiffusionGemma-Jev (djev) just got a lot easier. You can now spin up a Jev API-compatible endpoint on Google Cloud Run using a single command. Performance is solid: ~35-60 ms for single step latency and batch@32 is ~100-123 requests/sec. It's a straightforward way to experiment without needing your own GPU. Runs at roughly $3/hr and drops to $0 when idle. Get the code and instructions here: github.com/taeold/djev-run
24
135
7
1,279
105,539
J LIN retweeted
Jev-omni open source no HF Video, audio, imagens e texto como contexto
Introducing Jev-Omni, the first multimodal system one model. (Other OSS versions miss atleast a modality) Supports all modalities: text, images, audio and video ! On Par with Jev on Typed-benchmarks. Scaled -> 30k examples on 8xH200 (data mix matters a lot) < 100ms on 1 H100 huggingface.co/akhilaaa3/Jev… More work is coming, so follow along !
2
4
99
8,668
J LIN retweeted
Tada!!一覺醒來人類文明正式進入 GPT-6 家族世代啦!是說毫不意外雞肋般的 terra 正式被踢出門外,從此就是 astra/sol/luna 高中低三階配置打天下!🎉 ( 話說這下子今天得重新 eval 一次 #agentflow 準備來發佈最新操作指引嚕~🤓
2
1
45
1,523
谷歌Gemini 躲在角落里哭…
62
113
30
635
76,752
Qwen-Image-2.1 can now run locally on 12GB VRAM with Unsloth GGUFs! 🖼️ The 7B model performs on par with Nano Banana 2.0. For higher quality, you can also run Dynamic FP8 on just 6GB of VRAM via offloading. GGUF: huggingface.co/unsloth/Qwen-… Guide: unsloth.ai/docs/models/qwen-…
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: 👀 - Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs. - Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images. - Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products. - Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography. Start to create your next masterpiece with Qwen-Image-2.1! 🖼️ - Blog: qwen.ai/blog?id=qwen-image-2… - GitHub: github.com/QwenLM/Qwen-Image… - Model Scope: modelscope.cn/models/Qwen/Qw… - Hugging Face: huggingface.co/Qwen/Qwen-Ima…
75
271
34
2,914
351,546
Deploying DiffusionGemma-Jev (djev) just got a lot easier. You can now spin up a Jev API-compatible endpoint on Google Cloud Run using a single command. Performance is solid: ~35-60 ms for single step latency and batch@32 is ~100-123 requests/sec. It's a straightforward way to experiment without needing your own GPU. Runs at roughly $3/hr and drops to $0 when idle. Get the code and instructions here: github.com/taeold/djev-run
103
656
134
6,985
635,695
Q: Which models have been most impacted by the arrival of Jev? A: The flash versions of models across many labs. Also: nearly half of Jev users on OpenRouter hadn’t touched any models the week before. This launch sparked enough excitement to pull them off the sidelines.
18
6
5
212
16,172