@ivanalog_comi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Focus on 1) Technology and Dividend Blended 2) Drawdown protection. Check highlights and articles. No investment advice. Now building ModelDock
Iceland
Joined October 2011
- Tweets12.8K
- Following753
- Followers16.5K
- Likes15.9K
发现大家对用谷歌云GPU比较感兴趣,
Colab本身几乎是没有前端的,所以Collabosm作了:
一键开 GPU,目前有Qwen Flash Next和27B两种预设
自带聊天页面,可以发图片
隧道到本地固定地址,方便转发进Codex
状态一目了然:显存、隧道、prefill/decode 速度、CU 余额
可能是目前使用Colab的最佳方式
调试了一会,目前谷歌A100运行Qwen Flash的速度(3000+ prefill, 90+ tps)已经比一般的商用API还要快了,偶尔短时重度并发使用可以放到Collab上去。
为了节约大家的时间,我把在Collab上重建这个Qwen Flash的Recipe放到Github了,需要可以自取。
github.com/architectds/colla…
智力,长期以来,在人类社会都是极度稀缺的。
仅仅在一百多年以前的中国以及全世界,连传递家书这样的需求,大部分人都要通过请人帮助才能满足。
普鲁士依靠领先的义务教育和大量供应的受教育阶层,在一战之前实现了智力立国,然后是经济立国,遗憾的最终因为实力刺激了野心,多次挑起战争。
但这不改变智力对人类的意义:人类活动中,存在大量的未被优化的空间。智力可以帮助我们更高效,更优雅的生产生活。
而今天,智力终于成为了水和电一样自然的事情。全人类,迎来了我们的普鲁士时代。
仅凭这一点,我们就必将效率升级,让自然出产更多,让每个人的生活更加优雅(如果认为有效率就是繁忙,繁忙也许不可避免),向宇宙进发,成为星际文明。
AI 智能正在以人类历史上几乎没见过的速度变便宜。
Epoch AI 最新的数据:过去三年,获得同等 AI 能力的成本平均每季度下降 47%,一年大概便宜 13 倍。比 DNA 测序降本快 4 倍、计算快 6 倍、锂电池快 18 倍、电力快 54 倍。
2025 年初 o3 在 GPQA 做到 75% 左右,一道题成本大概 $0.30;不到 18 个月,GPT-5.6 Luna 做到差不多的成绩,只要 $0.0004,直接便宜了 725 倍。
AI 正在从“昂贵的软件工具”,变成一种近乎无限供应的生产要素。
但便宜不等于算力需求下降,反而很可能是 Jevons paradox:以前舍不得让 AI 干的活,现在全让它干;以前问一句,现在 Agent 可以自己跑几个小时、并行十几个任务。单次 intelligence 越便宜,人类消耗的 intelligence 反而越多。
未来被快速商品化的,可能恰恰是“ llm”本身。真正值钱的东西会越来越往数据、workflow、分发、用户关系,以及能把廉价智能转化成实际收入的产品的能力
AJ is one of the best, most diligent AI content creators.
Here's another one from @XiaomiMiMo Mimo-v2.6-flash!
A three.js animated mechanical Koi fish.
Same way as before, no assets are used, all in code, three.js and without any iterations i.e. no harness.
I changed the subject in prompt and this again delivered.
中年男人的消费力曾经不如狗,广告商看了都摇头。
如果一个中年人一不玩车,二不玩表,三不滑雪,那么作为一个中年男人,在这个消费主义时代,他就是隐形的。
人工智能硬件让中年男在消费市场上站起来了。
能跑前沿AI模型的硬件,让人心痒的睡不着觉。几万的硬件,在几千的订阅面前,是理直气壮的投资。
问题在于,本地AI,和大部分个人消费品一样,是有高闲置时间的。每天至少闲置8-16小时,即产能利用率30-60%。
叠加小硬件本身的配置劣势,大部分本地的优化不足。
同样的硬件,利用率比云端低一个数量级。
那么,同样的推理需求,就对应十倍的硬件需求。
所以一年前我就说,等NAS佬入场,AI硬件才会真正的快速涨价。
云商怎么算折旧,还算理性主义范畴。
末日危机时你我有人工智能相伴,这是消费主义冲动。
需求变化了一个数量级
Qwen flash is THE model for local inference right now.
I don’t know how, but I think Qwen should’ve a business model for releasing the best usable local models.
そういえば、Qwen3.8 Flash Nextはwevdevベンチでは
fabele 5の上の11位に居たりします。このサイズではかなり賢い!!!!
ちなみにQwen3.8 27Bは25位です!!!!
arena.ai/leaderboard/code/we…
One of the most helpful community note:
No the photo is not from that human zoo show,
It was from other human zoo shows around the similar time.
So back then there are multiple human zoo shows?
The Last Human Zoo exhibition, Belgium (1958).
Readers added context they thought people might want to know
The 1958 Brussels World's Fair included a Congolese village exhibit referred to as the last human zoo. The image of a child in a birdcage is from the Belgian Congo in the 1950s and not from that exhibition.
africamuseum.be/en/discover/hi… europeannewsroom.com/photo-taken-du… snopes.com/fact-check/hum…
我单方面宣布
Deepseek v4.1 Flash 重获我最喜爱的模型第一名称号。
Operation Cheepseek Phase 2:
$60 of usage for DeepSeek v4.1 Flash is now permanent
enjoy :)
大约三四十年前,彭丽媛是最红的歌唱家,习近平还没什么名气。
彭丽媛:“我只能嫁给一个能驾驭我的思想,我的爱情的人。”
主持人:“你已经这么成功了,你老公怎么可能比你更成功?”
Dario said that he will support a sweep game GPU ban to prevent this grieve future and urge President Trump to discuss a joint US-China GPU ban with President Xi to benefit the world at large.
Still developing
BREAKING: anthropic ceo dario amodei is reportedly concerned that an rtx 3060 12gb now runs a 27 billion parameter model fully locally at 50 tok/s fresh, and has called it a grave danger that gpus this powerful are allowed in homes. sources say the new essay's biggest concern is that a gpu this powerful can burn a house down if it isn't power limited, which requires advanced technical safety skills.
more as it develops.
事到如今,最对不起的就是家人。
A good day for updates ⚡️
Update for Qwen3.8 Flash for a solo DGX Spark
- Decode on prose is now 49 tok/s single stream.
- Decode on code is now 62 tok/s single stream.
- Faster follow-up replies.
- Optional vLLM 0.30 path with official Nvidia NVFP4.
This is still the BEST model to run a single DGX Spark.
The love-hate sentimeter for M7 today
🤖 Made with AI