The CUDA Handbook: Now Available Exclusively On The Website
At long last, The CUDA Handbook text is available for perusal with my permission!
A few years ago, I secured IP rights to The CUDA Handbook back from the publisher, and have been intending to refresh and update the material. TensorCores were added to the platform with the V100, almost a decade ago! The platforms have shifted; data center GPUs have become the flagship product offerings, with many GPU ASICs designed exclusively for deployment in data centers and, relatedly, PCI Express has been supplanted by faster and lower-latency interconnects, often suitable for enabling cache coherency between the CPU and GPUs.
The technology has evolved in eye-popping fashion. The first CUDA-capable GPU was a then-whopping 684M transistors – the biggest chip TSMC feasibly could manufacture at the time – and its main market goal was to run 3D games. Today, a single Rubin chiplet is 168B transistors, 245x bigger than G80, for a growth rate of 30% per year across 20 years. Rubin chiplets also are designed to collaborate with their peers, with packaging innovations that enable multiple chiplets to emulate a single huge GPU (“two GPUs in a trenchcoat”) with performance and seamlessness that GPU architects of yore could only imagine. The addition of TensorCores requires special attention, especially in light that starting with Blackwell, they now have their own on-chip memory.
The state of the art has advanced in other areas. Duane Merrill, the NVIDIA researcher whose Ph.D. thesis (cited in The CUDA Handbook) was on optimized GPU scan algorithms, improved on that state of the art with this work, done in collaboration with Michael Garland. Merrill and Garland’s implementation is nicely packaged for use by developers in the CUB and Thrust libraries, which are not mentioned anywhere in The CUDA Handbook. In light of these developments, the chapter on Scan serves more as a historical reference than resource that can be brought to bear on practical applications.
Another area where the state of the art has evolved: The C++11 standard was still young when the first edition of The CUDA Handbook was published. Modern C++ is more mature, NVIDIA has built several libraries that leverage its features to make available CUDA capabilities in a more ergonomic fashion, and std::thread has withstood the test of time and enjoys widespread toolchain and platform support; so some of the CUDA Handbook library code could stand to be refactored.
Finally, although most of the benchmarking code is as relevant today as when it was written, the benchmark numbers themselves are laughably out of date. I have some ideas on how to proceed, but won’t share further details at this time.
So, there’s no shortage of work to do.
Here are the steps taken so far:
The full text of The CUDA Handbook, including listings and figures, is now live on the website, and some measures have been taken to bring it onto the Internet and further into the 21st Century: for example, the errata fixes have been applied to the source material, and listing links now reference the GitHub repository.
All the ParallelProgrammer Substack posts have been moved to the Blog section, so the website is the canonical reference for the CUDA Handbook blog; and
I worked with the good folks at ethicalads.ai to transition the website to be ad-supported, with the option to remove ads for paid members. Tentatively, membership is priced the same as a subscription to my Substack ($10/mo), with an annual fee of $99 and a lifetime fee of $199.
Paying Substack subscribers will be “grandfathered” in to the ad-free experience on cudahandbook.com, and we’ll be emailing you all soon.
For those wondering Why is there a revenue model at all?:
By making the text of The CUDA Handbook available in this way – easily accessible, and more dynamic/easily updated, but copyrighted with no authorization to present it elsewhere without express written permission – I am reclaiming ownership of the material in a way that we’ve never enjoyed before. When the book was just a one-time byproduct of a person-year in nights and weekends, the content was more passive, like a big thick business card, and it is no longer even in print1.
With this new model, I’m hoping it can become something more.
Additionally, besides the obvious quid pro quo between authors and their readers, a revenue model enables copyright claims to be asserted that allege real economic damages. There is at least one GitHub repository with an unauthorized PDF scan of The CUDA Handbook available for download, and Microsoft has refused to take it down despite my filing a specific complaint and ticket with a link to the file. If I threaten to sue for damages, they may rethink their strategic neglect.
Every author has to navigate the tension between notoriety and being fairly compensated for their work. In software, we often bias toward the former, with open source initiatives supported by indirect economic models. For the sample source code, making The CUDA Handbook’s source code freely available was never in question. For the text and the presentation of material, I’m hoping this avenue will enable more activity and for the material to be updated and to remain current.
If you have thoughts on the matter, don’t hesitate to email me. You may be surprised at how few people do!
I reserve the option of self-publishing a print edition, and if I do, it’ll likely be slimmer and denser in content - more presentation and less reference type material.

