@vimalyohen

Verification Engineer #semicon-professional.

Trichy, tamilnadu, India
Joined March 2013
Vimal Manivannan retweeted
MS: CXL TAM expansion is going to be steep. As memory demand outpaces supply consistently, CXL is emerging as viable solution. Memory conundrum: - Intensity of memory shortage may fluctuate but all memory supply will be used by AI in immediate future - $NVDA CEO Jensen just said industry needs to workaround these limitations - De-speccing, rack scale central memory reduction, HBM reduction are already being tried upon - Disaggregation ( $CBRS ) and CXL are two such workarounds - MS overweight on $SNDK and $MU - Memory names like $SKHY Samsung have had lean period recent What is CXL? - Compute Express link is a high speed interconnect to give processors flexible access to memory - Using special memory & cache protocols, CXL make external memory to behave like processor memory - Key benefit being that memory need not be rigidly tied to specific CPU - Traditionally used for higher memory than lower latency / higher bandwidth (that AI needs) - Due to "memory wall" this is not emerging as a partial solution MS sees 3 main use cases: - Allowing servers to add more memory - Enables multiple processors to access same memory - Creates common memory pool that can dynamically be allocated to multiple servers / processors MS picks $ALAB and $MRVL as 2 key beneficiaries
2
7
24
1,755
Vimal Manivannan retweeted
A complete advanced deep dive of the architecture of High-Bandwidth Memory (HBM). The current HBM architecture is running into several scaling challenges from thermal performance, interconnect shoreline limits, and manufacturing yield. In my opinion, I think the #1 problem with HBM is thermal gradients in the stack. The temperature can vary greatly at different layers in the stack, requiring sophisticated sensors and circuits to compensate for thermal mismatches. In this advanced deep dive (linked below), I discuss the challenges of scaling HBM performance. 👇👇👇
6
21
150
5,779
Vimal Manivannan retweeted
GPU-Initiated Communication: Dissecting Down to the Bone arxiv.org/abs/2610.01380 github.com/ParCoreLab/Dissec… Javid Baydamirli, Ismayil Ismayilov, Kaan Oktay, Didem Unat
21
106
5,066
Vimal Manivannan retweeted
SPI2HDMI の話題が流れてきたので便乗。LCD や OLED に表示するノリでモニタやテレビに画像を出せます。BOOTH で基板も販売中。 shapoco.github.io/lcdtap/
4
58
263
10,889
Vimal Manivannan retweeted
The biggest bottleneck as we know as industry is: ✨The cost of data movement in processor-centric architectures 
@SKhynix latest AI Ecosystem piece is worth a read which is written by Dr. Onur Mutlu @_onurmutlu_ Professor at @ETH, one of PIM (processing-in-memory's) most prominent advocates. His core argument chimes with our view that AI performance now depends on how little data moves, not how fast compute runs. 
▪️One DRAM access can cost 150–2,000x the energy of a simple arithmetic operation. ▪️ Memory access and data movement can account for more than 90% of system energy for large ML models according to some studies. ▪️ This is where architecture matters and high-speed interconnect approaches such as CXL & NVLink also matters ▪️ CXL and NVLink/NVSwitch are solving different problems: ▫️CXL = capacity. It expands, pools and shares memory across hosts.
 ▫️NVLink = bandwidth. It moves weights, activations and KV cache between GPUs. ▪️Today, GPU memory expansion runs mostly through Nvidia's own paths, such as NVLink-C2C to Grace's LPDDR5X, not through CXL. CXL pooling at scale is still more roadmap than deployment. ▪️The KV cache is what's really driving this. Longer contexts and agentic reasoning make the KV cache grow faster than HBM capacity can keep up. ▪️ That's where Dr. Mutlu argues for disaggregation: moving the KV cache and attention work into separate memory pools or near-memory compute. e.g PIM could be ultimate goal but maybe not for every workload. 
▪️ So, every architectural approach raises memory's share of system value. That includes HBM next to the GPU, CXL pools, near-memory accelerators and PIM. This all points to we need more and more memory in different forms. And @SKhynix is not shy of emphasizing it 😅 $SKHY $MRVL
3
5
10
1,226
Vimal Manivannan retweeted
So something I have been working on, and just sent the initial rough draft over to my editor. Circa 300 pages (A4 pages) and 28 labs all based on the up coming explorer board.
22
99
2
973
22,500
Vimal Manivannan retweeted
Roadmap: Understanding GPU Architecture This roadmap is intended for those who are relatively new to GPUs or who would just like to learn more about the computer technology that goes into them. No particular parallel programming experience is assumed, and the exercises are based on standard NVIDIA sample programs that are included with the CUDA Toolkit. cvw.cac.cornell.edu/gpu-arch… More Roadmaps cvw.cac.cornell.edu/roadmaps
2
61
1
378
13,449
Vimal Manivannan retweeted
At Intel, we spent millions throwing hardware fixes at bad software. @luminal_ai exists to make the compiler take on more complexity so hardware does less.
8
8
2
153
10,128
Vimal Manivannan retweeted
"Glass Packaging for Chiplets Heterogeneous Integration", John Lau, Unimicron, J of Microele & Ele Pkg, 26-09-15 imapsjmep.org/article/169986… <= IEEE SCV EPS, 09-24 nitter.cf/ogawa_tter/status/2103… CPO, 03-26 Slides URL (80 pp) r6.ieee.org/scv-eps/wp-conte… 电子与封装, 07-06 nitter.cf/ogawa_tter/status/2092…
Replying to @ogawa_tter
=> 🇨🇳 "Research on the Underlying Core Technology of τ  - Law and Innovative Glass-Based Packaging", 北京清大博联科技中心 & 上海澈芯科技有限公司, 电子与封装, Jul 6, 2026 ep.org.cn/CN/10.16257/j.cnki… τ - Law, Tingbo He, Huawei, Science China Info Sci, Jul 21 nitter.cf/ogawa_tter/status/2084…
4
18
2
68
10,314
Vimal Manivannan retweeted
Chip Scale Review The Future of Semiconductor Packaging You can download PDF directly chipscalereview.com/read-the… 2026, Volume 30 2025, Volume 29 2024, Volume 28 2023, Volume 27 2022, Volume 26 2021, Volume 25 2020, Volume 24
2
44
1
301
12,663
Vimal Manivannan retweeted
If you're looking for a complete overview of glass packaging, this is it. 80-page slide deck just published by John Lau, IEEE fellow. r6.ieee.org/scv-eps/wp-conte…
12
103
10
655
45,955
Vimal Manivannan retweeted
其实,英伟达也是 RISC-V 架构的重度使用者 它的每块 GPU 内部通常包含 10 到 40 个 RISC-V 核心 英伟达在转向 RISC-V 之前 长期使用自研的 32 位专有架构 Falcon(快速逻辑控制器) 该架构最早于 2005 年左右随 G98 这一代产品首次亮相 截止2016十年间使用的 Falcon 核心总数约 30 亿个 但Falcon 为 32 位核心,缺乏 64 位寻址空间 既没有数据缓存,也无法运行真正的操作系统 英伟达从 2016 年起逐步将其替换为更具扩展性和标准化的 RISC-V 架构 (当时的备选方案包括 ARM 的 A 系列和 R 系列核心,以及 Synopsys 的 ARC、MIPS 和 Cadence 的相关技术,也曾评估过加州大学伯克利分校开发的参考核心 Rocket) xda-developers.com/your-nvid…
2
18
1
76
9,002
Vimal Manivannan retweeted
It’s raining 2nm chips!! 2nm is finally reaching flagship phones.
But the next #chipset battle will not be won by process node alone. @Qualcomm , @MediaTek , @Apple and @Samsung are pushing flagship platforms to 2nm. @Xiaomi XRING O3 and @Google Tensor G6 show 3nm can still compete through architecture and software. @Huawei is taking a different path with the Kirin 9050 Pro unfortunately due to US restrictions but still holding share in China market The real shift to monitor is the transition from GenAI to Agentic AI. #GenAI was mostly about generating images, summaries or answers. Agentic AI has to understand context, reason, talk to apps and take actions. That changes what smartphone SoCs must deliver. - Qualcomm is combining custom Oryon CPU, Hexagon NPU and Neural Fusion for more persistent AI workloads. - MediaTek is going AI-first efficiency: AI Compute Fusion, NPU upgrades and KV-cache optimization. - Apple is doubling down on vertical integration via A20 Pro’s 2nm design, higher memory bandwidth, Neural Engine and tightly coupled software. The side by side DRAM with SoC takes the thermal management to next level. - Samsung Semiconductor is using Exynos 2600 + 2nm GAA to own more of the Galaxy on-device AI experience. - Xiaomi is building its own silicon with XRING O3 for tighter hardware–software control. - Google is designing Tensor G6 around Gemini and on-device AI, not just peak CPU/GPU scores. - Huawei Hisilicon is betting on vertical stacking and shorter interconnects instead of pure node scaling. Peak performance still matters. For Agentic #AI, sustained performance per watt matters more. That’s where the real competition beckons. Great slide form @CounterPointTR team.
7
14
3
59
27,041
Vimal Manivannan retweeted
Hardware voltage glitching with success in seconds: AES key recovery, authentication bypass, and more! 📟⚡🤖⏱️🏆 More details on: LinkedIn: lnkd.in/p/dUuwS5X9 Substack: it4sec.substack.com/p/hardwa… Telegram: t.me/it4sec/54
5
42
259
8,965
Vimal Manivannan retweeted
Lightning, d-Matrix’s 3rd gen inference accelerator, will have a four high-stack DRAM directly on top of the compute die. Straight out of AI Infra Summit 2026. Raptor, the 2nd gen accelerator which is currently in the works, was presented at Hot Chips 2 weeks ago.
4
2
3
23
10,064
Vimal Manivannan retweeted
Hack IoT and in-car modems from a SIM: RUN AT proactive commands & how to use them. 🚗🏭၊၊||၊🎫😈 More details on: LinkedIn: lnkd.in/p/dwgjB-tb Substack: it4sec.substack.com/p/hack-i… Telegram: t.me/it4sec/52
1
29
172
6,349
Vimal Manivannan retweeted
Went back through NVIDIA’s CUDA Refresher series this week with some new folks. developer.nvidia.com/blog/ta… Origins of GPU computing: why we ended up here. CPUs stopped getting faster the easy way, and parallel hardware picked up the workloads that could split. Getting started with CUDA: what the toolkit is and how you get a first program running on the GPU. The GPU computing ecosystem: the libraries and tools sitting on top of CUDA. Useful for knowing what you should NOT be writing yourself. The CUDA programming model: kernels, grids, blocks, threads, how work actually maps onto the GPU. Note: it’s from 2020, so a lot of the tooling and setup steps have moved on and may not apply anymore. But the core GPU primitives still hold up.
GPU architecture | LLM Inference Handbook handbook.modular.com/kernel-… Add this to your LLM learning resource bundle. "Before writing or tuning GPU kernels, you need a working model of how a GPU runs code. Without it, suggestions like “increase occupancy” or "reduce shared memory bank conflicts" are just a set of rules to memorize. You don't fully understand when they apply and when they don't. This section explains modern GPU architecture at the level needed for kernel work. The details lean toward NVIDIA hardware because CUDA dominates much of the LLM inference ecosystem today. However, the core concepts apply broadly to AMD GPUs and other parallel accelerators as well."
3
38
1
278
31,815