Partner onboarding
How partners send usage data for multiple customers and get it back enriched.
This guide is for partners who send Greenpixie usage data on behalf of several customers. It covers the setup in order for Amazon S3, Azure Blob Storage and Google Cloud Storage, and links to the reference pages for detail.
How it works
- You upload each customer's usage data into your bucket or container, one directory per upload.
- When an upload is complete, you write an empty
.upload_finishedfile into its directory. - Greenpixie starts a run for each marked upload:
- Amazon S3: in a nightly check between 01:00 and 04:00 UTC, or on demand as soon as the marker is written if you add the optional event notification.
- Azure Blob Storage: in a nightly check between 01:00 and 04:00 UTC.
- Google Cloud Storage: in a nightly check between 01:00 and 04:00 UTC.
- The enriched files are written back to the same bucket or container, under
outputat the same path as the input.
1. Prepare the usage data
Greenpixie reads each customer's billing data in the provider's own export schema, with all of its columns. If you already collect your customers' billing data, write it out unchanged in the layout from step 3. If a customer needs to set up an export first, the platform guides show how:
- AWS: CUR 2.0 export
- Google Cloud: detailed billing export via BigQuery
- Azure: Cost Management export
A customer's export lands in their own storage, so copy the data files from there into your bucket or container.
Your setup reads one file type and, on Azure, one dataset (actual, amortized or FOCUS). Agree both with your Greenpixie contact during onboarding, and send every upload in them. Parquet is preferred, with no file larger than 500 MB. Keep the file extension the export writes, such as .snappy.parquet, and keep the same columns and data types across every file in an upload. See Data format.
2. Connect your storage
Use one bucket or container for both input and output. Greenpixie reads your uploads from it and writes the enriched files back to it, so it needs read and write access: on Azure, both Storage Blob Data Reader and Storage Blob Data Contributor; on Google Cloud, Storage Object User. One bucket or container can hold every customer and every cloud.
Follow the guide for your platform, then send your Greenpixie contact the details listed at the end of it:
On Azure, the client secret expires on the date you chose when creating it. Send a new one before then, or Greenpixie can no longer read your uploads.
3. Lay out each upload
Every upload goes in its own directory:
<bucket_container_name>/input/<org_id>/<platform>/<account_id>/<timestamp>/{partition files}<org_id>is your own ID for the customer. Keep it stable, and use the same one for a customer on every platform. You don't need to tell us about a new customer first: the first upload under a new ID sets it up as a new customer.<platform>isaws,gcporazure.<account_id>is the account the export covers: the payer account on AWS, the billing account on Google Cloud, or the billing account or subscription on Azure.<timestamp>identifies the upload, in UTC ISO 8601 basic format:YYYYMMDDTHHMMSSZ, for example20260722T120000Z. Use a new timestamp for every upload.
Provider default export layouts are not read for partners, so every upload must follow this layout.
For example, an Azure upload for customer org_12345, made at 12:00 UTC on 22 July 2026:
input/org_12345/azure/<billing_account_id>/20260722T120000Z/billing-part-0000.snappy.parquetSee Directory structure for the full specification.
4. Mark the upload complete
Once every file has finished uploading, write an empty file named .upload_finished into the same <timestamp> directory:
input/org_12345/azure/<billing_account_id>/20260722T120000Z/.upload_finishedA directory without the marker is never enriched. A marker written before the upload finishes means the enrichment runs on whatever files are there when it starts.
5. When it runs
On every platform, Greenpixie checks for marked uploads each night between 01:00 and 04:00 UTC and starts a run for every one not yet enriched. A marker written before 01:00 UTC is picked up that night; one written later may wait until the following night.
Each upload is enriched once, including older ones you add later, such as a history backfill. The first check after you connect picks up every marked upload already in your bucket or container, and a large backfill can take more than one night. To enrich an upload again, upload it under a new timestamp.
Optional: on-demand runs on S3
On S3 you can also have each upload start as soon as its marker is written, at any time of day. Without this, S3 uploads run in the nightly check like the other platforms. Azure Blob Storage and Google Cloud Storage don't support on-demand runs.
To set it up for a bucket:
- Send your Greenpixie contact your AWS account ID, your bucket name and the bucket's region. Your contact confirms the region is supported, creates a notification topic that accepts events from that bucket only, and sends you its SNS topic ARN.
- In the bucket, create an event notification with event type
s3:ObjectCreated:*, suffix filter.upload_finished, and the SNS topic ARN as its destination.
The nightly check skips any upload the notification has already started.
6. Collect the output
Enriched files land under output, at the same path as the input:
output/org_12345/azure/<billing_account_id>/20260722T120000Z/billing-part-0000.snappy.parquetThey keep every original column and row, with sustainability columns added. See Output format for the columns.
If output does not appear
Check that the upload's path follows step 3, and that .upload_finished sits in its <timestamp> directory and was written after every file finished uploading. If you set up on-demand runs on S3, also check that the bucket's event notification targets the SNS topic ARN your contact sent. Troubleshooting covers permission, file size and format problems. If none of these explain it, send your Greenpixie contact the path of the upload.