Where is LLM inference run?
Survey the deployment targets for inference — cloud, on-premises, hybrid and edge — and their trade-offs.
Draft
Foundations
Draft page
This page is an outline. It describes what will be covered and is not yet complete technical documentation.
Where inference runs determines what you can control and what you must accept. The decision is rarely purely technical: data residency, procurement, existing hardware and the shape of demand all constrain it.
This page lays out the options and the criteria for choosing between them.
What you will learn
- The main deployment targets and what distinguishes them.
- How data residency and compliance constrain the choice.
- When on-premises accelerators are worth the operational burden.
- Why demand shape matters as much as demand volume.
Recommended outline
This page is an outline. The subsections below are the planned structure; they are filled in as the handbook is written.