Load Greenpixie data into CloudHealth (AWS)

How to integrate your enriched AWS cloud usage report data into CloudHealth.

This guide takes you from an enriched Cost and Usage Report (CUR) file in your Amazon S3 bucket through to a fused cost and carbon dataset inside CloudHealth. It is written for a platform, FinOps or data engineer who already has the enriched file. That file carries the standard AWS CUR columns plus Greenpixie's numeric sustainability columns for energy in kWh, carbon in tonnes CO₂e and water in litres.

The integration uses CloudHealth Data Connect, the Bring Your Own Data feature, to load the file as a custom dataset. Custom Datasets then joins it onto the native cost model on shared keys such as resource ID and SKU.

For an Azure estate, follow the Azure version of this guide.

Before you start

You need an enriched CUR file landing in an S3 bucket you control, carrying both the standard AWS CUR columns and Greenpixie's numeric sustainability columns.

Several constraints shape the whole design, so read them first:

  • Data Connect is available only at the top-level organisational unit (TLOU) of the CloudHealth tenant. Create the connection there. A FlexOrg administrator without TLOU access will not find the feature.
  • The schema cannot be edited after creation. Getting the sample file right the first time is critical.
  • Maximum 100 MB per file, as CSV or Parquet, and only one file per account per billing month. Together these mean the CUR has to be pre-aggregated down to one file per account-month.
  • Column names must match ^[a-zA-Z][a-zA-Z0-9_-]*$, so letters, numbers, underscores and hyphens only, starting with a letter. Legacy CUR column names contain a slash and fail this rule, so they have to be renamed.
  • Loaded data is queryable within about 24 hours of creation. After that, Data Connect collects on a 12-hour schedule.
  • Native CUR loading requires Legacy CUR, not CUR 2.0, so the join keys line up with CloudHealth's own CUR dataset.

Prepare the enriched file

Before touching CloudHealth, make the file join-ready.

  1. Confirm the identifier columns are present and clean. The join back to the cost model depends on columns that also exist in CloudHealth's native CUR dataset. Include at minimum:

    • lineItem/ResourceId, the strongest join key.
    • lineItem/UsageType and/or product/sku, for SKU-level joins and where resource IDs are blank, such as some data-transfer or tax lines.
    • lineItem/UsageAccountId, the linked account.
    • A period key, bill/BillingPeriodStartDate or lineItem/UsageStartDate truncated to the day or month.
  2. Rename the CUR columns to pass the naming rule. Legacy CUR uses a slash in every prefixed column name, which CloudHealth rejects at schema inference. Replace the slash with an underscore across the whole file.

    CUR columnRename to
    lineItem/ResourceIdlineItem_ResourceId
    lineItem/UsageTypelineItem_UsageType
    lineItem/UsageAccountIdlineItem_UsageAccountId
    product/skuproduct_sku
    bill/BillingPeriodStartDatebill_BillingPeriodStartDate

    Greenpixie's own enrichment columns contain only letters and underscores, so they need no change.

  3. Convert the period columns to YYYY-MM-DD. CloudHealth infers a column as a Date only in YYYY-MM-DD form, and as a Timestamp only in YYYY-MM-DDTHH:MM:SS form. Anything else is stored as a string, and a string period key cannot be used as a date in reports or matched cleanly against CloudHealth's own period fields.

  4. Fix the grain. CUR is granular to line item, resource and hour, while Greenpixie metrics are usually per-resource and often daily. Decide the join grain now, where resource ID, SKU and day or month is the natural meeting point, and pre-aggregate the coarser side so you have exactly one row per key. A one-to-many match will multiply rows and inflate both cost and emissions.

  5. Keep sustainability columns strictly numeric so CloudHealth classifies them as measures, not dimensions, for example usage_electricity_consumption_kwh, total_tonnes_co2e, total_water_litres.

  6. Get each account-month under 100 MB. Because the folder structure allows one file per account per month and nothing finer, a single account-month that exceeds 100 MB in Parquet cannot be split. Reduce it instead, either by pre-aggregating to a coarser grain, so resource ID and month where you currently have resource ID and day, or by dropping columns that serve neither the join nor the report. Parquet is preferred for size and typed columns.

  7. Produce a representative sample file containing every column with realistic values, which is what defines the immutable schema.

Lay out the files in S3

Data Connect reads one fixed folder structure, and Broadcom's documentation states it must be followed exactly.

BYOD/<datasetName>/<accountId>/<YYYY-MM>/<version>/file.parquet
  • BYOD is a literal prefix.
  • <datasetName> is the Data Connect name you enter under Metadata, for example greenpixie_sustainability.
  • <accountId> is the AWS account ID the data belongs to.
  • <YYYY-MM> is the billing month, in exactly that format.
  • <version> is an incrementing folder, used when you re-upload the same month.

Accepted extensions are .csv, .csv.gz and .parquet, and a worked example path is below.

BYOD/greenpixie_sustainability/123456789012/2026-07/1/gpx-enriched.parquet

A file written to a free-form prefix is not collected.

Grant CloudHealth access to your S3 bucket

Data Connect authenticates to S3 through an IAM assume-role, using an External ID that CloudHealth generates. Collect two values from CloudHealth first.

  1. In CloudHealth, go to Setup > Accounts > AWS, open the target account, and copy the External ID, a roughly 30-character hex string, and the exact Account Name.

  2. In the AWS IAM console, create an IAM role for this connection with a trust policy that lets CloudHealth's account assume it, conditioned on that External ID. Verify the CloudHealth principal against current CloudHealth docs before applying.

    {
      "Version": "2012-10-17",
      "Statement": [{
        "Effect": "Allow",
        "Principal": { "AWS": "arn:aws:iam::<CLOUDHEALTH_AWS_ACCOUNT_ID>:root" },
        "Action": "sts:AssumeRole",
        "Condition": { "StringEquals": { "sts:ExternalId": "<EXTERNAL_ID_FROM_CLOUDHEALTH>" } }
      }]
    }
  3. Attach a least-privilege read policy scoped to the BYOD prefix in the bucket.

    {
      "Version": "2012-10-17",
      "Statement": [
        {
          "Effect": "Allow",
          "Action": ["s3:GetObject"],
          "Resource": "arn:aws:s3:::<BUCKET_NAME>/BYOD/*"
        },
        {
          "Effect": "Allow",
          "Action": ["s3:ListBucket"],
          "Resource": "arn:aws:s3:::<BUCKET_NAME>",
          "Condition": { "StringLike": { "s3:prefix": ["BYOD/*"] } }
        }
      ]
    }
  4. Copy the new role's ARN, which Data Connect asks for as the Assume Role ARN.

  5. If the bucket uses SSE-KMS, grant the role kms:Decrypt on the relevant key.

  6. Note the bucket's name, the region it is in, and the path to the BYOD prefix. Data Connect asks for all three alongside the role details.

Set up Data Connect

  1. In CloudHealth, go to Setup & Configuration > Data Connect and choose Add New Connection.
  2. In Select Connection Type, choose Custom File Upload. The automated connector option currently only covers Databricks, so custom files from S3 go through this path.
  3. Under Metadata, enter a dataset name and description, for example greenpixie_sustainability, then choose Next. The name you enter here is the <datasetName> folder in the S3 path.
  4. Under Authentication, set Data Location to AWS S3, then enter:
    • Assume Role ARN
    • Assume Role External ID
    • Assume Role Account Name
    • S3 Bucket Name
    • S3 Bucket Region
    • S3 Bucket Path
  5. Choose Next, then enter an Account Identifier and a Usage Month. Any unique identifier works for the first, and the AWS account ID is the natural choice. Give the second in YYYY-MM format, matching the month folder you wrote.
  6. Choose Connect, then Next. CloudHealth verifies it can reach the storage location and only opens the Schema section once that succeeds. A bucket policy, KMS or path error surfaces here.
  7. In the Schema section, enter the data location file path to your sample file and choose Get Schema.

Define the schema

The schema is set once and locked, so confirm it carefully.

  1. CloudHealth reads the sample file and auto-classifies each column as a measure or a dimension, and assigns a data type.
  2. Correct the classification before saving:
    • usage_electricity_consumption_kwh, total_tonnes_co2e, total_water_litres become numeric measures.
    • lineItem_ResourceId, product_sku, lineItem_UsageType, account and period become dimensions.
  3. Confirm the period column's inferred type is Date. If it shows as a string, the file's date format is wrong and the sample needs regenerating before you save.
  4. Set Mapping Identifiers. These are CloudHealth's cloud-agnostic FOCUS fields (the FinOps Open Cost and Usage Specification), being ResourceId, ServiceName, Billingaccount and ChargePeriod. Map your identifier columns onto them. This is what lets the dataset line up with other CloudHealth data, and is what the join depends on.
  5. Review and finish, after which the schema is locked. To change it, you create a new connection with a corrected sample file.
  6. The dataset imports and, within about 24 hours, appears as a Data Connect dataset in the Reports section and in classic FlexReports.

Load each new month

  1. Write each new month's file into a new <YYYY-MM> folder under the same dataset and account path, using the same schema and column names.
  2. Data Connect collects on a 12-hour schedule. To pull a new file immediately, without waiting for the next scheduled run, use Trigger Collection on the connection.
  3. To correct a month already uploaded, write a complete replacement file containing every row for that month into a new <version> folder. CloudHealth does not merge partial updates, so the latest file must include all previously uploaded data for that month.
  4. Monitor row counts per period as a smoke test confirming the load is working.

Join onto the cost model

  1. Go to Custom Datasets and create a new dataset.
  2. Add two source datasets, the native CloudHealth CUR / cost dataset and your greenpixie_sustainability Data Connect dataset.
  3. Add a JOIN operation, and choose the type deliberately:
    • A LEFT join with the CUR as the left dataset keeps every cost line and attaches Greenpixie measures wherever the keys match. Use this by default, to avoid dropping cost rows.
    • Use INNER only if you want to restrict to resources Greenpixie has measured.
  4. Select the join columns, which must match on both sides at the same grain:
    • The primary key is resource ID. Your renamed lineItem_ResourceId does not share a name with CloudHealth's native lineItem/ResourceId, so match the two through the ResourceId mapping identifier you set on the schema, since the column names no longer agree. Confirm this behaves as expected on your tenant before building reports on top of it, since the join across renamed columns is not covered by Broadcom's documentation.
    • Add SKU / usage type, product_sku or lineItem_UsageType, for SKU-level precision and to help on lines without a resource ID.
    • Add account and period to keep the match in scope, avoid cross-period fan-out, and produce a tighter, safer match overall.
  5. Compare the joined dataset's total cost against the unjoined CUR total as soon as the join runs; for a LEFT join, they should be identical. If cost has inflated, a one-to-many match is fanning rows out, so revisit the grain or add join keys.
  6. Add calculated columns for the metrics that matter, for example carbon_intensity_per_dollar = total_tonnes_co2e / cost, emissions_by_service, or kwh_per_service.

Report and dashboard

  1. Build a FlexReport or Report on the joined custom dataset, grouping cost and emissions by service, account, team tag or SKU on the same rows.
  2. Embed one or more of these reports into a CloudHealth dashboard for a unified cost-and-carbon view.
  3. For programmatic extraction, the GraphQL reporting API can pull the joined dataset into a BI tool.

References

Next steps

On this page