cuBLAS-level performance from Python using CuTe DSL kernels
A GPU programmer walks through six kernels and profiling bottlenecks showing that Python-hosted CuTe DSL can approach vendor-library performance on Ada Lovelace
i hit cuBLAS level performance on Ada Lovelace using CuTe DSLs kernels in python i walk through a series of 6 kernels, show how to profile them and understand profiling bottlenecks, tons of visuals and how to write understand CuTe DSL kern