Load Greenpixie data into CloudHealth (Azure)
How to integrate your enriched Azure cloud usage report data into CloudHealth.
This guide takes you from an enriched Azure cost file in your Azure Blob Storage container through to a fused cost and carbon dataset inside CloudHealth. It is written for a platform, FinOps or data engineer who already has the enriched file. That file carries the standard Azure cost-export columns plus Greenpixie's numeric sustainability columns for energy in kWh, carbon in tonnes CO₂e and water in litres.
The integration uses CloudHealth Data Connect, the Bring Your Own Data (BYOD) feature, to load the file as a custom dataset. Custom Datasets then joins it onto the native Azure cost model on shared keys such as resource ID and meter.
For an AWS estate, follow the AWS version of this guide.
Before you start
You need an enriched Azure cost file landing in a Blob Storage container you control, carrying both the standard Azure cost-export columns and Greenpixie's numeric sustainability columns.
- Data Connect is available only at the top-level organisational unit (TLOU) of the CloudHealth tenant. Create the connection there. A FlexOrg administrator without TLOU access will not find the feature.
- The schema cannot be edited after creation. Getting the sample file right the first time is critical.
- Maximum 100 MB per file, as CSV or Parquet, and only one file per subscription per billing month. Together these mean the export has to be pre-aggregated down to one file per subscription-month.
- Column names must match
^[a-zA-Z][a-zA-Z0-9_-]*$, so letters, numbers, underscores and hyphens only, starting with a letter. Standard Azure export names such asResourceId,MeterIdandSubscriptionIdalready comply, as do Greenpixie's numeric sustainability columns. - Loaded data is queryable within about 24 hours of creation. After that, Data Connect collects on a 12-hour schedule.
- Azure cost exports differ by billing-account type. Enterprise Agreement (EA) and Microsoft Customer Agreement (MCA) use different account and subscription column names, and exports come as actual or amortized cost. Mapping your identifier columns onto CloudHealth's FOCUS fields should let the join work regardless of which your tenant uses.
Prepare the enriched file
Before touching CloudHealth, make the file join-ready.
- Confirm the identifier columns are present and clean. The join back to the cost model depends on columns that also exist in CloudHealth's native Azure dataset. Include at minimum:
ResourceId, the full Azure Resource Manager (ARM) resource URI, in the form/subscriptions/.../resourceGroups/.../providers/..., and the strongest join key.MeterIdand/orMeterName, the meter-level equivalent of an AWS SKU or usage type, for meter-level joins and where resource IDs are blank, such as some marketplace, data-transfer or tax lines.SubscriptionId, also writtenSubscriptionGuid, the subscription that owns the resource. On EA you may also carryAccountNameandAccountOwnerId; on MCA,billingProfileIdandinvoiceSectionName.- A period key,
BillingPeriodStartDateorDatetruncated to the day or month.
- Convert the period columns to
YYYY-MM-DD. CloudHealth infers a column as a Date only inYYYY-MM-DDform, and as a Timestamp only inYYYY-MM-DDTHH:MM:SSform. Anything else is stored as a string, and a string period key cannot be used as a date in reports or matched cleanly against CloudHealth's own period fields. Check the format your export actually writes forDate,BillingPeriodStartDateandBillingPeriodEndDate, and convert where it differs. - Fix the grain. Azure usage detail is granular to the meter, resource and day, while Greenpixie metrics are usually per-resource and often daily. Decide the join grain now, where resource ID, meter and day or month is the natural meeting point, and pre-aggregate the coarser side so you have exactly one row per key. A one-to-many match will multiply rows and inflate both cost and emissions.
- Keep sustainability columns strictly numeric so CloudHealth classifies them as measures, not dimensions, for example
usage_electricity_consumption_kwh,total_tonnes_co2e,total_water_litres. - Get each subscription-month under 100 MB. Because the folder structure allows one file per subscription per month and nothing finer, a single subscription-month that exceeds 100 MB in Parquet cannot be split. Reduce it instead, either by pre-aggregating to a coarser grain, so resource ID and month where you currently have resource ID and day, or by dropping columns that serve neither the join nor the report. Parquet is preferred for size and typed columns.
- Produce a representative sample file containing every column with realistic values, which is what defines the immutable schema.
Lay out the files in Blob Storage
Data Connect reads one fixed folder structure, and Broadcom's documentation states it must be followed exactly.
BYOD/<datasetName>/<accountId>/<YYYY-MM>/<version>/file.parquetBYODis a literal prefix.<datasetName>is the Data Connect name you enter under Metadata, for examplegreenpixie_sustainability.<accountId>is the Azure subscription ID the data belongs to.<YYYY-MM>is the billing month, in exactly that format.<version>is an incrementing folder, used when you re-upload the same month.
Accepted extensions are .csv, .csv.gz and .parquet, and a worked example path is below.
BYOD/greenpixie_sustainability/00000000-0000-0000-0000-000000000000/2026-07/1/gpx-enriched.parquetData Connect does not collect a file written to a free-form prefix.
Grant CloudHealth access to your Blob Storage container
Where the AWS integration uses an IAM (Identity and Access Management) assume-role with an External ID, Azure authenticates through a service principal, which is an Entra ID or Azure AD app registration. You create the service principal, give it read access to the container, and hand CloudHealth its credentials.
-
In the target storage account, note the storage account name, the container name, and the blob prefix where the enriched files land.
-
Register an app or create a service principal for this connection and assign it a least-privilege blob read role scoped to the container, or the storage account.
Storage Blob Data Readeris the correct data-plane read role.az ad sp create-for-rbac \ --name "cloudhealth-greenpixie-reader" \ --role "Storage Blob Data Reader" \ --scopes "/subscriptions/<SUB_ID>/resourceGroups/<RG>/providers/Microsoft.Storage/storageAccounts/<ACCOUNT>/blobServices/default/containers/<CONTAINER>"This returns
appId,passwordandtenant, which CloudHealth calls Client ID, Client Secret and Tenant ID. Capture the secret now, as it is not shown again. -
If the storage account restricts network access through firewall rules, private endpoints, or the "allow trusted Microsoft services" setting, permit CloudHealth's access. A blocked path shows up later, as a failure when CloudHealth tries to reach the storage location.
Server-side encryption with customer-managed keys is transparent to a data-plane reader, so the service principal does not need Key Vault permissions to read encrypted blobs.
Set up Data Connect
- In CloudHealth, go to Setup & Configuration > Data Connect and choose Add New Connection.
- In Select Connection Type, choose Custom File Upload. The automated connector option currently only covers Databricks, so custom files from Blob Storage go through this path.
- Under Metadata, enter a dataset name and description, for example
greenpixie_sustainability, then choose Next. The name you enter here is the<datasetName>folder in the Blob path. - Under Authentication, set Data Location to Azure Blob, then enter:
- Tenant ID
- Client ID
- Client Secret
- Storage Account Name
- Container Name
- Blob Prefix
- Choose Next, then enter an Account Identifier and a Usage Month. Any unique identifier works for the first, and the subscription ID is the natural choice. Give the second in
YYYY-MMformat, matching the month folder you wrote. - Choose Connect, then Next. CloudHealth verifies it can reach the storage location and only opens the Schema section once that succeeds. A storage firewall or private endpoint blocking the path shows up here.
- In the Schema section, enter the data location file path to your sample file and choose Get Schema.
Define the schema
The schema is set once and locked, so confirm it carefully.
- CloudHealth reads the sample file and auto-classifies each column as a measure or a dimension, and assigns a data type.
- Correct the classification before saving:
usage_electricity_consumption_kwh,total_tonnes_co2e,total_water_litresbecome numeric measures.ResourceId,MeterId/MeterName, subscription and period become dimensions.
- Confirm the period column's inferred type is Date. If it shows as a string, the file's date format is wrong and the sample needs regenerating before you save.
- Set Mapping Identifiers. These are CloudHealth's cloud-agnostic FOCUS fields, being ResourceId, ServiceName, Billingaccount and ChargePeriod. Map your identifier columns onto them. This is what lets the dataset line up with the native Azure data, and is the foundation for the join. Because the fields are cloud-agnostic, the join should hold whether your tenant's native dataset uses EA or MCA column names. Confirm the join holds with your tenant's own EA or MCA column names when you build the Custom Dataset.
- Review and finish, after which the schema is locked. To change it, you create a new connection with a corrected sample file.
- The dataset imports and, within about 24 hours, appears as a Data Connect dataset in the Reports section and in classic FlexReports.
Load each new month
- Write each new month's file into a new
<YYYY-MM>folder under the same dataset and subscription path, using the same schema and column names. - Data Connect collects on a 12-hour schedule. To pull a new file immediately, without waiting for the next scheduled run, use Trigger Collection on the connection.
- To correct a month already uploaded, write a complete replacement file containing every row for that month into a new
<version>folder. CloudHealth does not merge partial updates, so the latest file must include all previously uploaded data for that month. - Monitor row counts per period as a smoke test that new files are loading correctly.
Join onto the cost model
- Go to Custom Datasets and create a new dataset.
- Add two source datasets, the native CloudHealth Azure cost dataset and your
greenpixie_sustainabilityData Connect dataset. - Add a JOIN operation, and choose the type deliberately:
- A LEFT join with the Azure cost dataset as the left dataset keeps every cost line and attaches Greenpixie measures wherever the keys match. Use this by default, to avoid dropping cost rows.
- Use INNER only if you want to restrict to resources Greenpixie has measured.
- Select the join columns, which must match on both sides at the same grain. Prefer the FOCUS mapping identifiers so the match does not depend on EA or MCA column naming:
- The primary key is resource ID,
ResourceId, which maps to the FOCUSResourceIdfield. - Add meter,
MeterIdorMeterName, for meter-level precision and to help on lines without a resource ID. - Add subscription and period, the FOCUS
BillingaccountandChargePeriodfields, to keep the match in scope and stop one row matching every period on the other side. More keys make for a tighter, safer match.
- The primary key is resource ID,
- Validate immediately. Compare the joined dataset's total cost against the unjoined Azure cost total; for a LEFT join they should be identical. If cost has inflated, a one-to-many match is fanning rows out, so revisit the grain or add join keys.
- Add calculated columns for the metrics that matter, for example
carbon_intensity_per_dollar = total_tonnes_co2e / cost,emissions_by_service, orkwh_per_service.
Report and dashboard
- Build a FlexReport or Report on the joined custom dataset, grouping cost and emissions by service, subscription, team tag or meter on the same rows.
- Embed one or more of these reports into a CloudHealth dashboard for a unified cost-and-carbon view.
- For programmatic extraction, the GraphQL reporting API can pull the joined dataset into a BI tool.
References
- Bring Your Own Data (BYOD) with Data Connect
- Custom Datasets, covering JOIN, MERGE and calculated columns
- FlexReports
- GraphQL API guides
- CloudHealth API docs
- Azure, register an app and create a service principal
- Azure, Storage Blob Data Reader role
- Azure, Cost Management exports