AI

9. Quality Assurance and Oversight

Standards alignment, validation, and review controls.

9.1 Standards Alignment

The carbon intensity, water, and embodied emissions factors applied by this methodology are provided by the core Greenpixie Methodology for Cloud Emission Measurement, which is aligned to the GHG Protocol, prepared with reference to ISO 14064-1, and independently verified in accordance with ISO 14064-3 at a limited level of assurance. Operational emissions are calculated in accordance with the GHG Protocol Scope 2 Guidance.

The AI-specific per-token estimation layer described in this document is being progressed towards the same standard of assurance. It is not independently verified at the time of publication and should not be presented as such.

9.2 Benchmarking Controls

Benchmarking is performed using standardised workload suites, consistent hardware configurations, and repeatable measurement procedures. The inference engine and its configuration are held constant across a benchmarking campaign, so that measurements taken at different points within it remain comparable. Energy measurements are captured using NVML to ensure consistency across benchmark runs.

9.3 Regression Model Validation

Predictive models are validated against held-out benchmark data to assess accuracy. Outlying measurements and outlying models are excluded from the fit and the model is refitted, against criteria defined in advance. Model performance is reviewed when new benchmark data becomes available or when material changes to the prediction database are introduced.

9.4 Cross-Referencing and External Comparison

Where available, Greenpixie's estimates are cross-referenced against published research, open-source tools such as EcoLogits, and vendor-published environmental data to identify material discrepancies. The principal reference points are the peer-reviewed inference energy benchmarking of Jegham et al., which measures energy through provider APIs, a model vendor's third-party-reviewed life cycle assessment, and a hyperscaler's published disclosure of the per-query energy, carbon, and water footprint of its consumer AI product.

Agreement across the majority of these comparisons falls within the stated uncertainty of the methodology. The comparison against the vendor life cycle assessment is materially higher than Greenpixie's prediction. That assessment amortises a model's cumulative life cycle impact, dominated by training and by hardware manufacturing, across the usage served to date, and presents inference as a marginal share of the total, so the two figures do not describe the same quantity. Material discrepancies are documented and investigated.

9.5 Price-Ratio Validation

Greenpixie applies a price-ratio validation as an independent check on predicted energy values for models that cannot be benchmarked directly. This tests whether the implied cost of the electricity consumed per token remains within plausible bounds relative to the price the provider charges for that token. If the implied energy cost exceeds a defined proportion of the token price, the prediction is flagged for review and the underlying band assignment is revisited.

The check depends on a margin between the price of a token and the cost of serving it. Where a provider prices inference at or near its cost of service, that margin is compressed and the check loses sensitivity. The validation is therefore applied as a plausibility bound alongside the external comparisons described in Section 9.4, and not as a definitive check in isolation.

9.6 Independent Production Validation

Greenpixie has validated its predicted energy values against independent measurements from enterprise clients running AI inference on their own infrastructure. These exercises compare modelled per-token energy estimates with metered energy data from real infrastructure. Discrepancies are investigated and used to inform model refinement.

9.7 Internal Review

Methodology changes, new model additions, band assignments that depart from the usual naming convention of a provider, and benchmark updates are subject to internal review before publication or deployment.

9.8 Academic and Technical Collaboration

The water consumption factors applied by this methodology are developed in collaboration with Dr Shaolei Ren at the University of California, Riverside, as set out in the cloud methodology.

On this page