Inference
The process of running a trained AI model to generate outputs, such as text, images or predictions, in response to new inputs.
Also called: Model inference, serving
Once a model has been trained, inference is how it is used: each prompt, query or agent step requires computation to process inputs and produce outputs, often measured in tokens. Inference runs on GPUs and other accelerators in data centers, and increasingly on custom chips and devices. Techniques such as reasoning at test time, where models spend more computation thinking before answering, raise the compute required per request.
Training is a large but periodic cost, while inference scales with usage, so it drives ongoing revenue and capacity needs for AI providers and cloud companies. Inference economics depend on hardware efficiency, model size, utilization and pricing per token. Falling cost per token tends to expand usage, so total inference spending can rise even as unit costs decline.
For allocators, inference demand is a key indicator of whether AI investment is converting into recurring revenue. Example: a company deploys a customer service model that handles millions of conversations a month, paying a cloud provider based on the tokens processed.
Related terms
Part of the BlockWest Glossary, plain-language definitions for markets, AI and digital assets. Educational content, not investment advice.
