@SpectralCom

Nerds taming the green dragon with SCALE, our framework for compiling CUDA codebases for AMD GPUs, with support for more accelerated platforms coming soon.

483 Green Lanes London England
Joined September 2025
You shouldn't have to pick a GPU vendor up front. And your codebase shouldn't pick one for you. That was our CEO @SpectralMichael's point on stage at the The Xcelerated Compute Show in London, in a fireside chat with @dcdnews's Vlad - Gabriel Anghel. He calls it team rainbow 🌈: rather than signing up with team green, red or blue, you mix and match. He compared the CUDA ecosystem to the UK plug: you can argue it's the better socket, but nobody rips out the wiring to switch. And rewriting code for each vendor is a maintenance bill that grows with every vendor you add. So keep your CUDA code. scale-lang.com compiles it for NVIDIA and AMD GPUs. You get more performance per dollar from the hardware you already own, and you buy your next GPU because it's the best one for the workload, not because it's the only one your code runs on. @BerryWrites wrote it up for @sdxcentral: sdxcentral.com/news/can-team…
1
104
The hardest test for a CUDA compiler is code that was never written with it in mind. So on top of our own tests, we build 48 open-source CUDA projects - FFmpeg, OpenCV, GROMACS, hashcat and more. We follow each project's own build instructions as closely as we can, then run its own test suite on AMD and NVIDIA GPUs. That is a deliberately unforgiving test, and it steers our roadmap. API coverage is driven by what real projects actually reach for, discovered by watching real builds fail - not by working alphabetically through a specification. The scripts and the results for each release are public, failures included. Read them and pick holes in them - repo link: github.com/spectral-compute/… 𝗔 𝗰𝗼𝗺𝗽𝗶𝗹𝗲𝗿 𝘁𝗵𝗮𝘁 𝗼𝗻𝗹𝘆 𝗰𝗼𝗺𝗽𝗶𝗹𝗲𝘀 𝘁𝗵𝗲 𝗲𝘅𝗮𝗺𝗽𝗹𝗲𝘀 𝗶𝗻 𝗶𝘁𝘀 𝗼𝘄𝗻 𝗱𝗼𝗰𝘂𝗺𝗲𝗻𝘁𝗮𝘁𝗶𝗼𝗻 𝗶𝘀 𝗮 𝗱𝗲𝗺𝗼. Come tell us what to test: scale-lang.com/s/discord?utm…
91
𝗦𝗖𝗔𝗟𝗘 𝟭.𝟳.𝟯 𝗶𝘀 𝗼𝘂𝘁. This release adds a Rocky Linux / RHEL 8 package, and supports older glibc in general. More of the compiler now ships statically linked, so it is less likely to clash with other LLVM-based compilers. Texture API support expands substantially. And there is the usual long list of nvcc-compatibility, runtime and CUDA-X library fixes. Grab the new release from scale-lang.com. 𝗗𝗿𝗼𝗽-𝗶𝗻 𝘀𝗵𝗼𝘂𝗹𝗱 𝗰𝗼𝘃𝗲𝗿 𝘁𝗵𝗲 𝘀𝘆𝘀𝘁𝗲𝗺 𝗮𝘀 𝘄𝗲𝗹𝗹 𝗮𝘀 𝘁𝗵𝗲 𝗰𝗼𝗱𝗲. Full release notes: scale-lang.com/s/scale173tw
1
2
101
Spectral Compute retweeted
Replying to @SpectralCom
llama.cpp carrying a hip tree that trails cuda is the quiet tax. cutlass having no amd door at all is the sharper split
1
1
39
There are two kinds of well-known CUDA project. The first kind ships two codebases: a CUDA one and a hand-written HIP port. Hashcat, GROMACS, Cycles, llama.cpp all carry both. Every change has to land twice, and the port usually trails the original. The second kind has no working AMD route at all. CUTLASS, AMGX and others simply do not run on AMD hardware - unless you compile them with scale-lang.com. For the first group we are a way to delete a codebase. For the second we are the only door in. Best part? SCALE is free for non-commercial use. Try it against your codebase, and drop into our Discord to let us know how it fared! scale-lang.com/s/scale-disco…
1
1
3
259
"Portable code is slower than hand-tuned code" is one of those things everyone knows and nobody rechecks. CPU compilers disproved it decades ago. They exploit diverse hardware features better than most people manage by hand. The GPU version of that argument is very concrete. Moving data between lanes can go through a general permute instruction, which bounces the data off the shared memory cache at roughly ten times the cost of a register move. Or it can go through a hardwired permutation pattern that is 𝗹𝗶𝘁𝗲𝗿𝗮𝗹𝗹𝘆 𝘄𝗶𝗿𝗲𝘀 - close to free, and expressed as a flag on an operand rather than an instruction of its own. The catch is that the available patterns change between hardware generations, so almost nobody reaches for them by hand. A compiler can. scale-lang.com takes a whole reduction loop, works out what is actually being computed across which lanes, and rebuilds it from the cheap patterns - chaining two of them where no single one matches. The same analysis then found something on NVIDIA hardware: a whole-warp reduction instruction most people never write the metaprogramming to reach. SCALE now uses it automatically wherever it applies. Read our blog post "Optimizing CUDA Shuffles with SCALE": tinyurl.com/5y5vhtxy
1
4
229
𝗡𝗼𝘁 𝗲𝘃𝗲𝗿𝘆 𝗸𝗲𝗿𝗻𝗲𝗹 𝗴𝗲𝘁𝘀 𝗮 𝟭𝟬𝟬𝘅. A Monte Carlo run splits into millions of independent jobs. GPUs eat that. Hundreds of times a CPU core, easily. An FDTD stencil is limited by memory bandwidth, not arithmetic. It runs as fast as the memory allows and no faster. A sparse implicit solve is the awkward one. On the open-source OPM Flow reservoir simulator, a published comparison[1] put one GPU at about 5.6x a dual-threaded MPI process. All three get sold as "GPU accelerated". So when someone quotes you a speed-up, ask which kernel it came from. Borrowed from another shape, it tells you nothing about yours. Ours, same treatment: on Rodinia, against hand-written HIP ports on an MI300X, scale-lang.com averaged 𝟲.𝟭𝟵𝘅. Up to 33.8x. Faster on 10 of the 14. Benchmarks: tinyurl.com/4ku2yx47
1
1
1
6
327
Most real CUDA codebases have inline assembly in them somewhere. Automated porting gives up on it. There is no source-level translation of inline PTX, so the code someone wrote specifically to be fast is the code that does not come across. scale-lang.com parses it instead. Inline PTX is raised into the compiler's own intermediate representation, optimised like any other code, then lowered back onto whatever instructions the target hardware actually has. A three-input bitwise operation written as inline assembly comes out the other side as two native AMD instructions, because once it is real code the optimiser can fold it. This pays off on NVIDIA hardware too. Plenty of inline PTX in the wild was written around toolchain limitations that no longer exist, and assembly a compiler cannot read blocks it from optimising everything around it. 𝗣𝗼𝗿𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝘀𝗵𝗼𝘂𝗹𝗱𝗻'𝘁 𝗰𝗼𝘀𝘁 𝘆𝗼𝘂 𝘀𝗽𝗲𝗲𝗱.
1
3
560
𝗖𝗼𝗺𝗽𝘂𝘁𝗲 𝗮𝘃𝗮𝗶𝗹𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗶𝘀 𝗮 𝗯𝗼𝘁𝘁𝗹𝗲𝗻𝗲𝗰𝗸. New accelerators reach a handful of clouds first, and capacity stays tight for a while after that. If your code only builds for one vendor, you're stuck waiting on that vendor's capacity. scale-lang.com compiles accelerated HPC codebases to run natively on both AMD and NVIDIA GPUs. That widens the pool of capacity you can actually use. 𝗣𝗿𝗶𝗰𝗲 𝗶𝘀 𝗼𝗻𝗲 𝗯𝗲𝗻𝗲𝗳𝗶𝘁. 𝗚𝗲𝘁𝘁𝗶𝗻𝗴 𝗰𝗼𝗺𝗽𝘂𝘁𝗲 𝘄𝗵𝗲𝗻 𝘆𝗼𝘂 𝗻𝗲𝗲𝗱 𝗶𝘁 𝗶𝘀 𝘁𝗵𝗲 𝗯𝗶𝗴𝗴𝗲𝗿 𝗼𝗻𝗲.
1
1
88
A recent scale-lang.com release included a significant atomics optimisation that only benefits AMD GFX11 cards. Hardware that is around five years old. Silicon vendors rarely fund that work, and for a completely rational reason. Making last year's card faster does not sell this year's card. Their incentive is the next chip.
1
1
7
1,696
A compiler company's incentive runs the other way. We are paid for code running fast on whatever you already own. Older hardware, mixed fleets, the cards you bought two procurement cycles ago - 𝘁𝗵𝗮𝘁 𝗶𝘀 𝘁𝗵𝗲 𝗷𝗼𝗯.
1
1
2
102
The largest body of GPU code in the world is CUDA. Models write it better than anything else. You can see people working around that. 🧵
1
1
2
144
Something still has to compile that code then, and many of the optimisations that matter aren't expressible in source at all. Agents speed up plenty of GPU work. They don't remove the compiler from it. Try scale-lang.com out and tell us what you think!
1
2
20
Every GPU strategy is secretly a hiring strategy. That now includes the models your team codes with. 5M+ developers have CUDA experience. And every coding agent learned from the same corpus - by far the largest body of GPU code in the world - so it writes CUDA code better than it writes anything else. Ask for an alternative and it will often hand back CUDA anyway. An alternative stack means retraining twice. That should be a line in the hardware comparison sheet - it's often bigger than the hardware delta. scale-lang.com removes it. The code your team and your agents already write compiles natively for AMD and NVIDIA GPUs. 𝗕𝘂𝘆 𝘁𝗵𝗲 𝘀𝗶𝗹𝗶𝗰𝗼𝗻 𝗼𝗻 𝗶𝘁𝘀 𝗺𝗲𝗿𝗶𝘁𝘀. Come say hi on Discord: tinyurl.com/bdfpy33p
2
93
𝗧𝗵𝗲 𝘃𝗼𝗹𝘂𝗺𝗲 𝗼𝗳 𝗺𝗮𝗰𝗵𝗶𝗻𝗲-𝘄𝗿𝗶𝘁𝘁𝗲𝗻 𝗚𝗣𝗨 𝗰𝗼𝗱𝗲 𝘄𝗶𝗹𝗹 𝗴𝗿𝗼𝘄 𝗮𝘁 𝗲𝘃𝗲𝗿-𝗶𝗻𝗰𝗿𝗲𝗮𝘀𝗶𝗻𝗴 𝗿𝗮𝘁𝗲𝘀. A lot of research effort is going into getting models to write GPU kernels, and building benchmarks for scoring them. The yardstick is almost always CUDA code, since that is what the models were trained on most. Generating a kernel is cheap now. Knowing whether it is correct is not: you will only find out after compiling. Don't waste tokens, use SCALE-clangd. It brings a modern development workflow to your CUDA project - autocomplete, go-to-definition, diagnostics - so a bad type or a missing include shows up in the editor, not after the build and not in review. 𝗬𝗼𝘂𝗿 𝘁𝗶𝗺𝗲 𝗺𝗮𝘁𝘁𝗲𝗿𝘀, 𝗮𝘀 𝘄𝗲𝗹𝗹 𝗮𝘀 𝘆𝗼𝘂𝗿 𝘁𝗼𝗸𝗲𝗻𝘀. 𝗗𝗼𝗻'𝘁 𝘄𝗮𝘀𝘁𝗲 𝘁𝗵𝗲𝗺. Integrate SCALE-clangd in your workflow 👉 tinyurl.com/z76kw6wp
1
85