Optimizing the GPU kernels that inference workloads spend their time in.

The pages in this section are listed below in the recommended reading order.