攻城狮/业余投机/右侧交易/游戏开发/Haskell/Rust/C++/Unity3D/C#/对java有偏见/RL/智商欠费/浅尝辄止故平庸/反乌托邦/竹林中/地狱变/手撕菠萝蜜/胸口碎榴莲/单机推特中/乐视一生黑(乐视已阵亡)/华为一生黑(迟早会阵亡 )/中华跪族/器材党/预防式B台支

上海, 中华人民共和国
Joined June 2012
CXMT (688825.SH) announced plans to invest RMB 24.1bn in a new technology R&D project, including RMB 13bn in excess IPO proceeds. It also plans to invest RMB 10.8bn in Phase II of its memory wafer back-end testing base, including RMB 5bn in excess IPO proceeds. The new facility will provide DRAM chip testing and module assembly. Total planned investment: RMB 34.9bn (~US$4.9bn). #CXMT #DRAM
2
103
Muse is basically a data harvester that goes out of its way to collect your personal information, while offering very little practical value. Straight to the trash.
1
3
163
my timeline after opus 5.5 release
1
2
134
Qwen 4 is coming soon. Qwen 4.5 and Qwen 5 are targeting a parameter size of 4–10T.
14
823
The previously missing data from MiMo’s RL training has now been added.
MIMO’s RL training has stopped at step 30. We can observe the following: The Pro model achieved a DeepSWE score of 72.57, but this score was reported at step 28; data for steps 29 and 30 are missing. The number of active environments for Pro began increasing at step 23, then dropped rapidly after step 26. The duration of each training step also rose sharply.
5
624
MIMO’s RL training has stopped at step 30. We can observe the following: The Pro model achieved a DeepSWE score of 72.57, but this score was reported at step 28; data for steps 29 and 30 are missing. The number of active environments for Pro began increasing at step 23, then dropped rapidly after step 26. The duration of each training step also rose sharply.
2
1
5
989
Replying to @ZixuanLi_
此时,智谱公关团队的心情:
1
48
1,835
Replying to @mranti @Ion_Mio_
DGX 三台可以三角高速互联,四台需要 200G 交换机,通讯带宽减半,不算是个好方案。
217
Kimi K2.8 Preview offers a 1M context window, and its performance is officially claimed to be close to K3. kimi.com/code/docs/en/kimi-c…
11
1,058
DeepSeek-V4.1-Flash delivers beastly performance at a fraction of Kimi K3’s price. ↓
DeepSeek-V4.1-Flash: 552B parameters, with just 8B active during prefill and 16B during decoding. Kimi K3: 2.8T total, 104B active. Those benchmark numbers make V4.1-Flash look like an absolute beast! And judging by the release cadence, Kimi’s next model—whether it’s K3 Pro or K3.1—should be just around the corner.
5
298
DeepSeek-V4.1-Flash: 552B parameters, with just 8B active during prefill and 16B during decoding. Kimi K3: 2.8T total, 104B active. Those benchmark numbers make V4.1-Flash look like an absolute beast! And judging by the release cadence, Kimi’s next model—whether it’s K3 Pro or K3.1—should be just around the corner.
1
1
12
940
DS has a very interesting question in the V4.1 Flash survey: "Do you think this model can fully replace the online (production) version of DeepSeek V4 Pro?" It seems that high expectations are being placed on this multimodal V4.1 Flash.
"DeepSeek V4.1 Flash is now in internal beta as an intermediate version — welcome to try it out. It features a new model architecture with native multimodal support, stronger capabilities, faster speed, and lower cost. Keep your base_url unchanged and set the model name to deepseek-v4.1-flash-expires-on-0910 to start using it. Pricing is currently the same as deepseek-v4-flash, with a rate limit of 20 concurrent requests per account."
2
2
33
38,683