Back to News & insightsIndustry

AI energy claims: measure a useful task before comparing footprints

Distinguish power from energy, define the measurement boundary, and include retries and quality when comparing AI inference workloads.

Editorial guide · Updated September 27, 2026 · 4 min read
A computing tile sits beside a glass energy reservoir and a metal heat radiator.

A claim about the energy used by one AI query sounds simple until you ask what counts as a query. A short classification, a long generated report, and a multi-step agent run can involve very different amounts of work.

Useful comparisons need a defined task, a measurement boundary, and an accepted result. Without those details, a precise-looking number can hide more than it explains.

Power and energy answer different questions

Power is a rate of energy use. Energy accumulates over time. A device drawing more power can still use less energy for a task if it finishes sufficiently sooner; a low-power device can run long enough to consume more.

That distinction is why peak device specifications are not a measurement of one completed request. You need observations over the relevant execution period and a clear account of what equipment was included.

The Power Hungry Processing research compares inference energy across defined model categories and tasks. It provides evidence about those experiments, not a permanent universal footprint for every current AI service. Hardware, software, workload, and utilization change the comparison.

Define the useful unit of work

Imagine a hypothetical publisher tagging incoming articles. Candidate systems include a task-specific classifier and a general language model. The useful unit is an accepted set of tags for an article under the same quality requirements, not an arbitrary token or API call.

Fix the input distribution and the acceptance criteria. Include long articles, ambiguous topics, and cases where no existing tag fits. A system that saves energy by producing inadequate output has not completed equivalent work.

Count retries and manual correction where relevant. If one approach needs several generation attempts, the energy for those attempts belongs to the workflow. Report failures instead of measuring only the successful subset.

State the measurement boundary

Decide whether you are measuring accelerator energy, whole-server energy, or a broader service estimate. Include model loading, idle time, preprocessing, retrieval, and networking when they fall within the stated boundary.

If only accelerator telemetry is available, say so. Do not label that measurement as a complete data-center footprint. Shared infrastructure and cooling may require additional information that a client-side experiment cannot observe directly.

For the publisher, report both a cold run that includes startup and a steady-state run if both occur in practice. Amortizing a large startup cost across an assumed huge batch can be misleading when the real workload is small and intermittent.

Compare under realistic utilization

A continuously busy server and an on-demand service can have different operating profiles. Batching can change throughput, waiting time, and energy per accepted item. Measure the arrival pattern your product is likely to produce.

Run enough repeated trials to see variability, and keep environmental conditions and software settings visible. A single measurement can be dominated by background activity or startup behavior.

Measure latency beside energy. A highly efficient batch process may be unsuitable for an interactive tool. A decision should acknowledge that tradeoff rather than presenting one resource metric as the complete definition of efficiency.

Do not turn energy into emissions without assumptions

Energy use and associated emissions are related but distinct. Estimating emissions requires information about electricity generation and the accounting boundary. Time, location, and methodology can affect the result.

If those inputs are unavailable, report energy or a clearly labeled estimate with its assumptions. Avoid using an unrelated average to imply a precise footprint for a specific remote service.

Likewise, operational measurements do not automatically include manufacturing or other lifecycle impacts. A narrower measurement can still be useful as long as its scope is stated accurately.

Optimize the workflow you actually control

For the tagging system, first eliminate unnecessary repeated work and evaluate whether a smaller task-specific approach meets the quality contract. Then test batching, shorter outputs, and appropriate reuse. Each change should preserve the accepted outcome.

Keep a record of energy, quality, latency, and volume before and after the change. If a lower per-item cost encourages much more usage, total consumption can still rise. Report both per-task efficiency and overall workload where available.

A responsible energy comparison is a reproducible account of a defined task. Its value comes from helping people make a better decision, not from attaching a dramatic universal number to every interaction with AI.

Research background

Read the original research paper on arXiv

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.