How large companies manage Scope 3 data collection across 5,000 or more suppliers

Supplier Engagement
Marc Munier
,

CEO

6 min read
Table of contents

Howden manages Scope 3 PG&S emissions across 55 countries with DitchCarbon.

See what the platform could do for you.
Book a demo

A supplier list in the thousands breaks any process built for hundreds. Most Category 1 programmes start life on a spreadsheet designed for the pilot: a few dozen strategic suppliers, a questionnaire, a few weeks of chasing responses. That approach doesn't fail gradually as the list grows, it fails at a threshold. Somewhere between 200 and 500 suppliers, response-chasing stops being a task a person can do alongside their other work, and by 5,000 it isn't a task at all, it's a second full-time job with a response rate as its only output metric.

Why a bigger list needs a different starting point, not just more effort

The GHG Protocol Scope 3 Standard names four calculation methods for Category 1, purchased goods and services: supplier-specific, hybrid, average-data and spend-based. A programme built around outreach is implicitly betting on the supplier-specific method for the whole list, asking every organisation to report before anything gets calculated. At 50 suppliers that bet is survivable. At 5,000, the response rate becomes the ceiling on the entire inventory, whoever is running it, because the total can't be more complete than what's been sent back.

The alternative the Standard itself allows is to model first and replace second. Start with a spend-based or average-data estimate across the full list so the total is complete from day one, then improve individual lines as supplier-specific figures arrive. Nothing in the Standard requires one method applied uniformly across a category, and a hybrid inventory, mixing reported figures where they exist with modelled ones elsewhere, is compliant provided the method and source are documented line by line.

Start from what's already known, not from zero

DitchCarbon provides verified emissions data for over 2 million organisations. For an enterprise supplier list, a meaningful share of names on it already have a company-specific figure on record before a single outreach email goes out, matched through entity resolution against DUNS, LEI and ISIN identifiers rather than by name alone. Name-matching is where large lists usually lose accuracy: a spend export with "Acme Ltd", "ACME Limited" and "Acme (UK) Ltd" as three separate rows either gets treated as three different suppliers or merged by hand, and neither scales past a few hundred names.

This changes what "5,000 suppliers" actually means as a project. It isn't 5,000 unknowns, it's however many of the 2 million already have a documented figure, plus a smaller remaining list that genuinely needs work. The size of that remaining list, not the size of the original one, is what determines how much outreach a programme actually has to run.

Bring your own list in without a separate integration project

Organisations can be added in bulk by uploading a CSV matched to a provided template, rather than one supplier at a time. Once a list is in, organisations can be sorted by embodied emissions and by DitchCarbon Score, a composite rating built from more than 30 underlying data points with a benchmark percentile attached, so the suppliers carrying the largest share of the total, and the ones furthest behind their peers, are visible rather than buried in a spreadsheet of thousands of rows.

That sorting step is what turns a flat list into a worklist. A supplier at the top of the embodied-emissions ranking with no reported figure is a priority for direct engagement. A supplier near the bottom with a low DitchCarbon Score isn't worth the same attention, whatever its position in the procurement system.

What "large" actually changes about the calculation itself

Scale doesn't just add volume, it changes which of the four Category 1 methods is realistic for which part of the list. At 50 suppliers, a team can reasonably aim for supplier-specific data across most of the list within a reporting cycle. At 5,000, that aim has to be scoped to a segment, the top few hundred by spend or by modelled emissions, with the rest of the list carried on hybrid or average-data methods deliberately rather than by default. That's not a lower standard, it's the standard the GHG Protocol itself describes: methods matched to where the data effort produces the most improvement in the total, not applied uniformly because uniform effort is easier to explain.

The same logic applies to how often data gets refreshed. A 50-supplier list can be reviewed in full every reporting cycle. A 5,000-supplier list needs a refresh strategy that revisits the top of the ranking most often and the long tail less frequently, because revisiting every line at the same cadence spends as much effort on suppliers that don't move the total as on the ones that do.

Close the gap that's left, not the whole list

Once a list is matched against existing data, what remains to chase is a smaller pile than the original 5,000, and it's ranked rather than flat. Effort goes where it moves the total: the top 200 suppliers by modelled contribution first, not the alphabetically first 200. This is the same principle procurement teams already apply to spend, an 80/20 concentration in most Category 1 lists, applied to data collection instead of sourcing.

Requests that do need to go out can be made through the Survey Responder, which routes a buyer's questionnaire into a tracked queue answered from a supplier's own profile data rather than a blank form started from scratch each time. A supplier who has already answered a similar request from another customer isn't starting from zero on this one either.

Documentation at this scale, not just data at this scale

ISO 14064-1:2018 requires an organisation to document its quantification methodology and the reason it was selected, not just report a total. At 5,000 line items, each produced by a different method depending on data availability, that documentation burden is itself a scale problem: which method produced which line, which factor set and vintage, where a reported figure came from and what changed since the last cycle. Every DitchCarbon figure carries its source and change history for this reason, so the record exists at the point the number is generated rather than being reconstructed at year end.

Where these programmes actually stall

The common failure mode isn't a bad first quarter, it's a plateau after it. A programme reaches 40 or 50% coverage relatively quickly, because that's the share of the list where a figure was already available or the easiest suppliers to reach responded first. Progress then slows sharply, because what's left is the suppliers who don't respond, don't have a public disclosure, or sit several tiers down a subcontracting chain where nobody on the buying side has a direct relationship to lean on.

Two things tend to keep that plateau from becoming permanent. The first is ranking the remaining gap by modelled contribution rather than treating every unanswered supplier as equally worth chasing, since a non-responder near the bottom of the embodied-emissions ranking is costing the total far less than one near the top. The second is accepting that some share of a 5,000-supplier list will stay on modelled data indefinitely, and building the audit trail to say so honestly rather than treating an unclosed gap as a reporting failure.

How this fits an existing procurement system rather than replacing it

A programme at this scale rarely gets to start from a blank sheet. There's usually an existing spend file, a supplier master in an ERP or procurement platform, and a sustainability team that doesn't own either. The practical entry point is the CSV export that already exists, brought in as a batch rather than requiring a live systems integration before any data collection can start. Sorting and prioritisation happen on top of that import, so the first useful output, a ranked list of the suppliers worth asking first, can exist before any decision gets made about deeper system integration.

What this costs to run

Numbers you can defend within 2 weeks. A recent deployment reached about 60% coverage on a large supplier base within that window, with the coverage gap shown rather than hidden, so a programme starting at 5,000 suppliers begins from an honest picture of what's known and what genuinely still needs to be asked for, rather than a spreadsheet that looks complete because every row has a number in it.

Send a sample of your own supplier list and see the coverage split before committing to a full onboarding.

See the coverage on your own category register

Send us your supplier list and we will show you the coverage and the data quality behind each figure, so you can see which of your significant categories can be upgraded off spend-based data.

Join the industry leaders and solve your Scope 3 emissions data challenge

See how DitchCarbon can transform your sustainability journey with auditable insights and verified data.