Back to News & insightsGuides

Understand the cost behind a prompt.

Input, output, context, and repeated requests all affect an AI application’s running costs.

Editorial guide · Updated September 17, 2026 · 1 min read
Floating metal and glass layers connected to a central core, illustrating application architecture.

A prompt is only one part of an AI request. System instructions, attached material, conversation history, and the generated response can all affect usage. A small chat message does not necessarily imply a small bill.

Tokens are units of processing

Models break information into tokens. A token is not always a word, and tokenization varies across models and input types. Use your provider's usage metadata instead of relying on a fixed words-to-tokens assumption.

Estimate the complete workflow

Read the current rate for the exact model and service you use. Check input and output charges, caching conditions, tool charges, and any separate storage or infrastructure costs. A workflow that calls a model several times should include every call in its estimate.

Set limits before scaling

For your own application, define input and output bounds, concurrency limits, and a spending budget. Include failed or repeated attempts when reviewing actual usage. Revisit the estimate whenever you change the model or workflow.

Further reading

Google AI for Developers: Technical documentation

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.