Hello author, I want to change nattenqkrpb_cuda_forward_kernel to achieve the desired

how to debug cuda kernel? about neighborhood-attention-transformer HOT 3 CLOSED

shi-labs commented on June 27, 2024

how to debug cuda kernel?

from neighborhood-attention-transformer.

Comments (3)

alihassanijr commented on June 27, 2024

Hello and thank you for your interest.
It really depends on what you mean by debugging.

When it's a compilation issue, you may find using setup.py better as the logs are a bit more concise.
If you want to check your implementation, you probably have to either come up with a pure python/pytorch implementation, and compare outputs of the two. That's what we typically do, which is send random tensors with different shapes through the CUDA version and the "torch version" with the same weights, compute the outputs, check if they allclose, and then do the same for the backward pass.
If you're sure of your forward pass being correct, gradcheck is probably a better way to check if your backward pass kernel is correct.
It's also not a bad idea to have assertions in place, even in the kernel when debugging, but I'd recommend leaving in as few assertions in the device code (the kernel) as possible and do assertions mostly before the kernel call.

If you're getting into optimization, you may find pytorch's profiler useful for measuring latency, and you'd probably need NVIDIA Nsight to profile in more detail.

I hope you find these useful, but if you need more details, please let us know.

from neighborhood-attention-transformer.