iAccount based inIndia
About this account
- Account based in
- India
- Connected via
- India Android App
Account-level information from X, not a live location or the device used for a specific post.
Verification Engineer #semicon-professional.
- Tweets1.2K
- Following4.9K
- Followers301
- Likes2.1K
ALT This paper dissects GPU-initiated communication at the GPU–NIC boundary. We first detail the GPU-side network path: queue placement, work-request construction, doorbell ordering, and completion semantics. A minimal GPU path issues an operation in 0.7 µs and completes in 4.0 µs; libraries add up to 4.6 µs of issue time through queue management, memory ordering, and completion scope, and issue time scales with the SM clock. A tuned CPU proxy matches or beats the GPU path at idle, at the cost of a dedicated core whose operating state sets its latency and throughput. Reaching the 260 M msg/s ceiling of our InfiniBand platform requires doorbell batching and queue parallelism, and both have resource costs: communication code can reduce GPU block residency even when unused, and all-to-all traffic loses 59% of its NIC message rate at about 3,000 active connections.
ALT Figure 1: The GPU-submitted path with GPU-resident queues and in-kernel WQE construction. In a CPU-proxy path the host performs (1)–(3) and (6), and the GPU instead enqueues a request descriptor. Figure 2: mlx5 send WQEs for an RDMA write, built from 16 B segments (ctrl: control, AV: address vector, raddr: remote address) and fetched in 64 B basic blocks. Only DC WQEs carry the AV (§3.4). An inline payload replaces the data segment’s pointer with a length word and the payload. The bottom row expands the control segment; the doorbell store carries its first 8 B. Correct GPU-side RDMA delivery hinges on two ordering requirements: (1) WQE and source-payload stores must be visible to the NIC before the doorbell that announces them; and (2) when one thread rings the doorbell for WQEs that other threads wrote, each writer’s stores must be visible before the ringing thread sees its slot as ready.
ALT Table 3: One 8 B operation. Completion is one CQE for mini-gda, a host-resident counter for the baseline mini-proxy, and a GPU-resident counter written through GDRCopy for the tuned one. Figure 6: RTT under load on (a) P-IB and (b) P-GB200, with background CTAs loading both endpoints. Dashed curves share one ring or context with the bulk traffic; solid curves reserve a probe queue. IBGDA uses a fixed 16-QP pool and GDAKI a private context. Labels give the background M msg/s carried at 64 CTAs. We can tune the proxy to the workload by allocating one producer per ring, caching progress, and encoding validity in the descriptor, which together reduce enqueue to a single posted 16 B store taking 0.13 µs. Idle latency reflects per-operation software costs, while loaded latency depends on queue isolation and carried traffic. Separate rings reduce interference, while workers and chaining raise proxy capacity.
ALT Figure 9: Useful-work throughput loss relative to the kernel without communication code, for (a) a compute caller with eight live values (2.1 ms) and (b) a streaming caller with 16 live values (2.6 ms). The three bars of an arm are not additive. Figure 10: Connection scaling on P-H100. (a) 8 B rate through one NIC against its active QPs per direction. Long reuse increases writes per connection visit from 32 to 8,192. Receive-only uses total incoming QPs and aggregate rate over a common duration; other curves use the median sender. (b) 32-PE RC/DC and dense 128-PE RC sweeps, normalized to their two-peer controls within each pass. The x axis counts visited QPs or DCI–peer pairs. Medians over passes with their range; hollow markers are single-pass points; (b) uses the mean over NICs. In the streaming caller, both NVSHMEM builds and GDAKI lose about 37% before sending a message, consistent with lower block residency (Figure 9b). All-to-all loses 59% by 3,000 active connections.