Techniques that increase throughput and reduce latency for a given model and hardware.

The pages in this section are listed below in the recommended reading order.