2. Scope and Applicability
What is in scope, out of scope, and the evaluation boundary.
2.1 In Scope
The methodology applies to the following:
- Cloud AI inference workloads. Energy consumption and embodied emissions arising from the processing of input tokens in the prefill phase and the generation of output tokens in the decode phase, by large language models and embedding models hosted on major cloud providers.
- Water consumption. The water consumed in cooling the hardware that serves inference, and the water consumed in generating the electricity drawn by inference, estimated from the same energy values on the basis set out in the cloud methodology.
- Supported cloud providers. Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP).
- Model coverage. The methodology covers large language models (LLMs) and embedding models available through the supported cloud providers' AI service catalogues. At the time of publication, the methodology's prediction models are applied across over 220 models in Greenpixie's database.
- Token types. Input tokens, output tokens, cached input tokens, and embedding tokens.
- Billing configurations. The methodology distinguishes between configurations that affect energy profiles, including batch versus real-time inference, prompt-length tiers where the cloud provider differentiates them, and prefix caching.
2.2 Out of Scope
The following are not covered by this methodology:
- Model training and fine-tuning energy on a per-token basis. Training energy is accounted for as an uplift on amortised embodied emissions, as described in Section 7.4, but the methodology does not model usage-phase emissions per training token.
- Image generation, video generation, and prompt optimisation cloud AI products.
- Locally hosted or on-premises AI inference, unless the customer provides sufficient deployment information, including token counts, model identity, and infrastructure configuration, to enable Greenpixie to apply its predictive models. Where such information is provided, on-premises inference may be assessed on a case-by-case basis.
2.3 Evaluation Boundary
The evaluation boundary encompasses server-level energy consumption attributable to AI inference. GPU power is measured directly via NVML; the remaining server components are accounted for via wall-time-based estimation rather than direct instrumentation. Facility-level overhead such as cooling and power distribution is accounted for through standard data centre efficiency factors. Embodied emissions are calculated for server hardware used during inference, amortised over an assumed operational lifespan. Water consumption covers the IT cooling water consumption from cooling the inference hardware and the electricity generation water consumption from the electricity drawn by inference.