@hgpui
iAccount based inGermany
About this account
- Account based in
- Germany
- Connected via
- Germany Android App
Account-level information from X, not a live location or the device used for a specific post.
High performance computing on graphics processing units (GPU): AMD/ATI, nVidia, Intel Xeon Phi, CUDA, OpenCL, OpenGL, GPGPU, HPC
Joined May 2011
- Tweets10.7K
- Following119
- Followers3.9K
- Likes104
DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance Engineering
#AMD #HIP #LLM #Performance #DeepSeek
hgpu.org/?p=31229
Automated Instruction Encoding Synthesis for Modern GPU ISA Compression
#CUDA #ISA #HardwareArchitecture #Package
hgpu.org/?p=31226
PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving
#CUDA #LLM #Performance #Package
hgpu.org/?p=31227
Accelerating the Solving of Many Tiny General Linear Systems on GPUs: Application to Constitutive Laws
#CUDA #LinearAlgebra
hgpu.org/?p=31228
AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines
#CUDA #LLM #Performance #Package
hgpu.org/?p=31225
Taming Bitwise Behavior in GPU Kernels with Tensor Core: Black-Box Reconstruction, Compiler Enforcement, and Static Verification
#CUDA #PTX #Triton
hgpu.org/?p=31199
PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
#PTX #LLM #Package
hgpu.org/?p=31135
Benchmarking Confidential Computing Performance on NVIDIA Blackwell GPUs
#CUDA #Security
hgpu.org/?p=31134
Concurrency Response of Plain Global Loads on the NVIDIA H100
#CUDA #Performance #HardwareArchitecture
hgpu.org/?p=31133
RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
#Triton #CUDA #LLM #Package
hgpu.org/?p=31131
Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4
#PTX #Package
hgpu.org/?p=31108
A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family
#Triton #LLM
hgpu.org/?p=31107
Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell Ultra
#PTX #Triton #ISA
hgpu.org/?p=31106
CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution
#PTX #CUDA #Compilers
hgpu.org/?p=31105
Validation-Centric AI-Assisted GPU Porting of a 250,000+ Line Legacy Weather Simulation Code
#OpenACC #Fortran #OpenMP #CodeGeneration
hgpu.org/?p=31104