- CUDA
- MLsys
- Technology
•
•
-
How the “Correct” Way to Align CUDA Pointers Can Slow Down Shared-Memory Loads
A pointer-to-integer round trip can erase CUDA shared-memory provenance, turn LDS/STS into generic LD.E/ST.E instructions, and slow down an otherwise correct kernel.
-
The Catastrophe of
#pragma unrollin CUDA ProgrammingThis post records a CUDA performance issue I previously overlooked, as a reminder to myself and to friends who might encounter similar situations in the future.