AI

8. Monitoring, Updates, and Review Timing

How the methodology is kept current.

8.1 Ongoing Maintenance

The model database is updated on an ongoing basis as cloud providers add new models, which are assessed against the eligibility criteria in Section 6 and assigned predicted values using the existing regression models. Cloud provider pricing structures and billing configurations are monitored continuously, and changes that affect the mapping to usage types are incorporated as they are identified. Benchmarking is refreshed when material changes occur, such as new GPU hardware architectures, significant shifts in inference engine performance, or model architectures that fall outside the existing benchmark set. Parameter bands are recalibrated periodically, as the parameter count implied by a given provider tier label moves as the frontier advances. Regional carbon intensity, water, and data centre efficiency factors are updated in line with the cadence described in the core Greenpixie Methodology for Cloud Emission Measurement.

8.2 Extensions in Development

The following extensions to the methodology are in development:

  • Wider hardware coverage, extending benchmarking beyond the current accelerator set, and measuring at rack and cluster level in addition to the individual card, so that at-scale interconnect and utilisation effects are captured directly.
  • Larger models, benchmarking open-weights models above one trillion total parameters to improve extrapolation to the largest proprietary models, where measured reference points are currently sparsest.
  • Long context and agentic workloads, extending the benchmarked prompt and response range to represent coding, document summarisation, and multi-step agent workloads.
  • Multimodal inference, extending the methodology to image, video, and audio tokens.
  • Measured cache behaviour, replacing the assumed prefix-caching adjustment with directly measured hit rates across models, hardware, batching regimes, and prompt lengths.
  • Tokenizer normalisation, expressing energy per character or per byte alongside energy per token, so that comparison between models with differing tokenizers is exact.
  • Automated benchmark ingestion, admitting newly released open-weights models from a curated set of providers to the benchmark set within days of release, so that the regression models track the frontier without manual intervention.
  • Narrower parameter bands, reducing band widths and the associated reported interquartile range as further public information on proprietary model scale becomes available.
  • Agentic tool energy, quantifying the energy of tool usage such as web search, code execution, and retrieval, and amortising it onto a per-token basis.

On this page