Growth Marketing Glossary

Inference (AI)

in·fer·encenoun

Putting a trained model to work - generating an answer from a prompt. Where the cost, speed, and scale of every live AI feature are decided.

promptoutputrunning a trained model to produce an answerthe cost & speed of every AI request
Schematic — a trained model producing an output
Term
Inference (AI)
Is
Running a trained model on new input
Versus training
The 'using' phase, not the 'learning' phase
Determines
Per-request cost, latency, scale

Forms & parts of speech

inference · noun
Running a model to get output.
"At scale, inference cost dominated - every customer request meant another model run to pay for."

Definition in plain terms

Inference is the act of using a trained AI model: you give it an input, like a prompt or an image, and it produces an output, like a response or a classification. It contrasts with training, which is the resource-intensive process of building the model in the first place.

Once a model is trained, inference is what happens every time it's actually used - so for a deployed AI feature, inference happens on every single request. Inference has real costs: it consumes computing power, takes time (latency), and must scale with usage.

While training a large model is enormously expensive up front, inference is the ongoing, per-use cost that accumulates as a product is used. Understanding inference as the 'using' phase clarifies where the day-to-day economics and performance of AI features come from.

Why it matters to growth leaders

For a growth leader running AI in production, inference is where the ongoing economics live. Every time a customer interacts with an AI feature - a chat response, a generated recommendation, a piece of content - that's an inference, with a cost in compute and a delay in latency.

At scale, across many users and requests, inference cost and speed become defining constraints on an AI product's unit economics and user experience. A feature that's cheap to demo can become expensive to run when usage grows, because inference cost scales with usage.

Understanding inference helps a growth leader forecast the true cost of an AI feature, design for acceptable latency, and make sound build decisions

because the question isn't just whether an AI capability works, but whether its per-inference cost and speed make it viable at the scale the business needs. It connects AI capability to the unit economics that determine whether it's worth deploying.

Worked example. A growth leader launches an AI recommendation feature that performs brilliantly in the pilot, then watches its costs and response times balloon as usage scales - and inference is where the problem lives.

Every customer interaction triggers an inference: the trained model runs on that user's input to produce a recommendation, consuming compute and adding latency on each request.

In the small pilot, the per-inference cost was negligible; across millions of live requests, inference cost became a defining constraint on the feature's unit economics, and latency began to hurt the user experience.

The growth leader, understanding inference as the ongoing 'using' phase distinct from one-time training, treats it as the real economic question: not just whether the feature works, but whether its per-inference cost and speed are viable at scale.

The leader optimizes - caching common results, using a smaller model where quality allows, capping where appropriate - to bring inference cost and latency into a sustainable range.

Understanding inference, the growth leader forecasts the true cost of AI features and designs for the per-request economics and speed that determine whether an AI capability is actually worth running at the scale the business needs.
Failure modes to watch. Forecasting AI cost from training alone while ignoring per-request inference cost that scales with usage; designing AI features without accounting for inference latency; assuming a cheap demo means a cheap product at scale; and overlooking inference economics in build decisions.

Synonyms & antonyms

Synonyms

inferenceAI inferencemodel inference

Antonyms

trainingmodel training

Origin & history

Inference is the use phase of a trained AI model - generating outputs from new inputs - distinct from training; as the per-request cost and latency that scale with usage, it governs the real economics and performance of deployed AI features.

Etymology: source.

Usage trends

Search interest for this term over the last five years:

View interest-over-time on Google Trends →

Common questions

What is inference in AI?
The process of running a trained AI model on new input to generate an output — the 'using' phase, as opposed to training; it determines the per-request cost, latency, and scalability of an AI feature.
How is inference different from training?
Training is the up-front, resource-intensive process of building the model; inference is using the trained model on each new request, so inference cost accumulates with usage.
Why does inference matter for AI economics?
Every use of a deployed AI feature is an inference with a compute cost and latency, so at scale inference cost and speed define the feature's unit economics and user experience.

Related tools & calculators

Resources & people to follow

Curated, non-competitor resources verified per term.

Related training

Disciplines

Areas of marketing where inference (ai) is a core concern:

Sources

  1. trendsGoogle Trends — "ai inference"