A transpose does no arithmetic. Every nanosecond it spends goes to moving bytes, so whatever it loses against a plain copy is lost to the order in which it touches memory. That makes it the cleanest kernel for watching the memory system work.

Each episode starts from the kernel the previous one left behind and changes one idea. Every version is checked against the same exact reference and reported as the same number, effective bandwidth, next to a copy kernel that moves the same bytes.

The series builds on the CUDA C++ Course. If you have not taken it, the essay on how the course works is the place to start.