Region Scorecard

4. Methodology Overview

How the methodology works at a high level.

Two obstacles prevent a region from being scored by averaging its prices. Prices arrive in units that cannot be added together, since an hour of compute, a gigabyte-month of storage, and a million API requests are not commensurable. Regions also sell different mixes of products, so a region selling only inexpensive products would appear inexpensive whatever its prices were.

Price list entries are first grouped into product categories, so that entries for the same item of purchase are compared with each other. Entries inconsistent with the same product category elsewhere are removed as data errors, and a product category whose prices remain inconsistent is excluded in full.

A comparison basket is then fixed for the provider, holding the product categories priced widely enough across its regions for a comparison to mean something. Every region is scored on that basket, whatever else it does or does not sell.

Each price becomes a price ratio against the median price for its product category, which removes the unit problem, because a ratio carries no unit. A region's relative cost percentage is the median of its ratios. Where a region prices too little of the basket, no cost score is published for it.

The service coverage percentage is a separate count, comparing the product categories a region prices against the count priced by the provider's most complete region. Both percentages are placed in five bands using Jenks natural breaks, and the resulting grades are published alongside them.