CXMT (688825.SH) announced plans to invest RMB 24.1bn in a new technology R&D project, including RMB 13bn in excess IPO proceeds.
It also plans to invest RMB 10.8bn in Phase II of its memory wafer back-end testing base, including RMB 5bn in excess IPO proceeds. The new facility will provide DRAM chip testing and module assembly.
Total planned investment: RMB 34.9bn (~US$4.9bn).
#CXMT #DRAM
The previously missing data from MiMo’s RL training has now been added.
MIMO’s RL training has stopped at step 30. We can observe the following:
The Pro model achieved a DeepSWE score of 72.57, but this score was reported at step 28; data for steps 29 and 30 are missing.
The number of active environments for Pro began increasing at step 23, then dropped rapidly after step 26. The duration of each training step also rose sharply.
MIMO’s RL training has stopped at step 30. We can observe the following:
The Pro model achieved a DeepSWE score of 72.57, but this score was reported at step 28; data for steps 29 and 30 are missing.
The number of active environments for Pro began increasing at step 23, then dropped rapidly after step 26. The duration of each training step also rose sharply.
Kimi K2.8 Preview offers a 1M context window, and its performance is officially claimed to be close to K3.
kimi.com/code/docs/en/kimi-c…
DeepSeek-V4.1-Flash delivers beastly performance at a fraction of Kimi K3’s price. ↓
DeepSeek-V4.1-Flash: 552B parameters, with just 8B active during prefill and 16B during decoding. Kimi K3: 2.8T total, 104B active. Those benchmark numbers make V4.1-Flash look like an absolute beast!
And judging by the release cadence, Kimi’s next model—whether it’s K3 Pro or K3.1—should be just around the corner.
DeepSeek-V4.1-Flash: 552B parameters, with just 8B active during prefill and 16B during decoding. Kimi K3: 2.8T total, 104B active. Those benchmark numbers make V4.1-Flash look like an absolute beast!
And judging by the release cadence, Kimi’s next model—whether it’s K3 Pro or K3.1—should be just around the corner.
DS has a very interesting question in the V4.1 Flash survey:
"Do you think this model can fully replace the online (production) version of DeepSeek V4 Pro?"
It seems that high expectations are being placed on this multimodal V4.1 Flash.
"DeepSeek V4.1 Flash is now in internal beta as an intermediate version — welcome to try it out. It features a new model architecture with native multimodal support, stronger capabilities, faster speed, and lower cost.
Keep your base_url unchanged and set the model name to deepseek-v4.1-flash-expires-on-0910 to start using it. Pricing is currently the same as deepseek-v4-flash, with a rate limit of 20 concurrent requests per account."