How to Practice CUDA Locally Without an NVIDIA GPU
July 14, 2026
If you're learning CUDA without access to an NVIDIA GPU, you're no longer blocked. Compiler Explorer can compile and even execute CUDA programs on remote hardware. Just head over to godbolt.org, choose a CUDA compiler, and you're ready to experiment.
This post, however, aims to set up a local CLI development environment. With it, you'll be able to code in the comfort of your favorite terminal and editor. The only catch is that there's a few milliseconds of delay between when you run the command and get your outputs, due to the network latency.
This setup is good for practicing CUDA syntax, checking that kernels and memory management compile correctly, and confirming small programs behave as expected, all without installing the CUDA toolkit or owning NVIDIA hardware. It's not a good fit for serious performance benchmarking, since you're sharing remote hardware and timing won't be reliable, or for large, stateful projects, since each run starts fresh.
To set this up locally, we'll use a small CLI that talks to Compiler Explorer and, if needed, a tool to flatten
Since Compiler Explorer receives a single source file, projects that rely on local headers need to be flattened first. multi-file projects into a single source file.
I picked cexpl and quom, respectively
Set up a virtual environment and install cexpl. If your project spans multiple files, install quom as well.
$ python -m venv venv
$ source venv/bin/activate # on Windows: venv\Scripts\activate
$ pip install cexpl quom
Next, write your CUDA code in a file, say main.cu.
#include <cstdio>
#include <cuda_runtime.h>
__global__ void hello() {
printf("Hello from thread %d\n", threadIdx.x);
}
int main() {
hello<<<1, 8>>>();
cudaDeviceSynchronize();
return 0;
}
This kernel just prints a message from 8 GPU threads.
Also, in practice you should check the return value of cudaDeviceSynchronize.
If your code includes local header files, make a combined file using
quom.$ quom main.cu combined.cu
Finally, compile and run it remotely with
$ cexpl --exec --skip-asm --lang cuda main.cu
--exec runs the compiled program, not just compile it;
--skip-asm hides the generated assembly output, so you only see program results;
--lang cuda tells Compiler Explorer to treat the file as CUDA code.
After a short delay, you should see something like:
STDOUT:
Hello from thread 0
Hello from thread 1
Hello from thread 2
Hello from thread 3
Hello from thread 4
Hello from thread 5
Hello from thread 6
Hello from thread 7
That output was produced by an NVIDIA GPU running on Compiler Explorer's servers.
Keep in mind that the execution happens in a fresh environment every run.
--execcan sometimes crash instead of showing errors. If your code has a compilation error,cexpl --execshould normally print the compiler errors to stderr and stop there, but depending on your system or network, it sometimes crashes outright instead of showing you anything useful. If that happens, just drop--execand run:$ cexpl --skip-asm --lang cuda main.cuThis only compiles the code and prints the compiler output, including any errors, without trying to execute it. Once your code compiles cleanly, add
--execback to actually run it.
That's all there is to it. With that, you can practice and learn CUDA a bit more easily.