2. Scope and Applicability
What is in scope, out of scope, and the evaluation boundary.
2.1 In Scope
The methodology applies to the following:
- Cloud AI inference workloads. Energy consumption, carbon emissions, and embodied emissions arising from the processing of input tokens in the prefill phase and the generation of output tokens in the decode phase, by large language models and embedding models hosted on major cloud providers.
- Water consumption. The water consumed in cooling the hardware that serves inference, and the water consumed in generating the electricity drawn by inference, estimated from the same energy values on the basis set out in the cloud methodology.
- Supported cloud providers. Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP).
- Software as a Service token usage. The per-token outputs apply to any usage report that exposes a per-model token count broken down by token type, which includes vendor API logs and per-seat AI tooling as well as cloud provider billing data.
- Model coverage. The methodology covers large language models (LLMs) and embedding models available through the supported cloud providers' AI service catalogues. At the time of publication, the methodology's prediction models are applied across over 260 models in Greenpixie's database.
- Token types. Input tokens, output tokens, cached input tokens, and embedding tokens.
- Billing configurations. The methodology distinguishes between configurations that affect energy profiles, including the provider's processing tier, being real-time, batched, priority, and flexible, and whether prefix caching applies to the request.
- Training energy. The energy consumed in training a model is accounted for as an uplift on the operational energy of inference, amortising it across the tokens the model serves, as described in Section 7.4.
2.2 Out of Scope
The following are not covered by this methodology:
- Usage-phase emissions per training token. Training energy is included as an uplift on inference energy rather than modelled per training token.
- Image, video, and audio inference. This is a text inference methodology, and the energy of non-text modalities is not represented.
- Prompt optimisation cloud AI products.
- The additional energy of agentic tool usage, such as web search, code execution, and retrieval, beyond the token counts those tools generate.
- Locally hosted or on-premises AI inference, unless the customer provides sufficient deployment information, including token counts, model identity, and infrastructure configuration, to enable Greenpixie to apply its predictive models. Where such information is provided, on-premises inference may be assessed on a case-by-case basis.
2.3 Evaluation Boundary
The evaluation boundary encompasses server-level energy consumption attributable to AI inference. GPU board power is measured directly via the NVIDIA Management Library (NVML); the remaining server components are accounted for via wall-time-based estimation rather than direct instrumentation. Facility-level overhead such as cooling and power distribution is accounted for through Power Usage Effectiveness (PUE) factors. An uplift is applied to operational energy to account for the training of the model being served. Embodied emissions are calculated for server hardware used during inference, amortised over an assumed operational lifespan. Water consumption covers the IT cooling water consumption from cooling the inference hardware and the electricity generation water consumption from the electricity drawn by inference.
Energy is attributed on a request-scoped basis, under which a request accounts for the server time it causes and not for idle capacity between requests. Section 7.1 describes this attribution.