What the data holds, what it derives, and how it is built

Capabilities first, plumbing as the evidence

Most descriptions of a data platform open with its architecture. This one does not, because architecture rarely conveys what becomes possible. The pipeline is set out in full further down, but it appears as evidence beneath each capability rather than as the opening argument. The question the page is arranged around is a research one: which problems become tractable once a supply-chain model can see corporate structure, and not only country totals?

What is distinctive here?

An entity-resolution layer, rather than a minerals dataset.

sections 02, 05

What can be derived from it?

Ten analytical surfaces, and one measure not yet in common use.

sections 07, 08

What could it support scientifically?

Three candidate directions, ranked by what the measurements allow.

sections 13

The value is in the join, rather than the rows

Minerals data is, for the most part, public - USGS, national surveys, study groups - and any research group can obtain it. The scarce ingredient is the layer that makes those sources composable against corporate reality: a legal-entity backbone able to establish that a Peruvian cadastre record, an SEC filer, a Dutch disclosure and a Brazilian tax number all describe one economic actor. Two hundred registers are, in effect, two hundred directories written in different naming conventions. This is the concordance between them.

Sources converge through entity resolution into one object898,749facility records6,895,243entities2,747,104ownership + control edgescatalogued sourcesentity resolutionone queryable object3,434,524 same_entity_as edges

208

catalogued sources

206 of them public

3,434,524

same_entity_as edges

the resolution layer itself

6,895,243

entity records

1,708,611 resolved clusters

brightquery~470M documents

US-strong legal-entity and corporate-structure corpus. CIK + LEI + EIN + ticker + ISIN + FIGI resolved onto one row.

brightquery_global~124M rows · 95 jurisdictions

Rest-of-world legal-entity backbone. Contributes LEI and (jurisdiction, company number) registration numbers.

Two of the 208 sources are proprietary. The other 206 are public, and any team could fetch every one of them. What is harder to assemble is the identifier hub that makes those 206 composable - the reason a Peruvian cadastre record, an SEC filer, a Dutch disclosure and a Brazilian tax number can be shown to describe one economic actor.

The headline count is accurate, and easily misread

898,749 facility records is a correct number and an easily misread one. An atlas marks every place a mineral has ever been found; a census counts what is actually in production today. This corpus is largely the former, with a smaller, well-evidenced core of the latter - so the narrowing below carries more meaning than the headline, and it is what determines whether a claim built on this data will survive review.

The corpus narrows from 898,749 records to 2,197 corroborated898,749Facility records in storeany spatial asset record - all geocoded, 200 countries482,552Commodity-attributedrecord carries at least one commodity attribution16,133Operating facilitiesclassified as an operating production asset2,197Corroboratedtwo or more independent sources agree
Occurrence306,25163.5%
Ambiguous160,16833.2%
Operating16,1333.3%

The three largest contributing registries are cadastral, not operational: USGS MRDS (237,087), Brazil ANM SIGMINE (267,423) and Peru INGEMMET (66,263) record that something was found or claimed, not that anything is being produced. We hold a very large asset gazetteer with a small, well-evidenced operating core. The gazetteer gives global reach; the operating core is what supports quantitative claims. They are not interchangeable.

catalogued160,149
crosswalk154,984
occurrence151,267
observed13,955
corroborated2,197

Only the bottom two tiers carry independent confirmation. The top three are catalogue presence, not observation - which is exactly why the pyramid narrows so hard.

Global in reach, uneven in publication

The corpus spans 200 countries, yet 92% of records sit in four of them. That distribution reflects which governments publish bulk cadastral data rather than where minerals happen to be, which is why a country ranking built on raw record counts should be read with care.

92.3% in the top four

USUnited States441,05049.39%
BRBrazil269,53130.18%
PEPeru68,3537.65%
CACanada44,8405.02%
DEGermany6,3260.71%
CLChile4,7470.53%
GBUnited Kingdom4,6360.52%
AUAustralia4,4140.49%
CNChina4,0860.46%
MXMexico3,5710.4%
MNMongolia3,2840.37%
NONorway3,0300.34%
-All other40,8814.5%

China produces more mined output than any other country and appears here ninth, at 0.46% of records. Brazil sits second at 30%, on the strength of one open cadastre. The ranking measures publication regimes, not geology - much as a global rainfall map drawn only from countries that happen to operate weather stations would say more about instruments than about rain.

Six million edges, of which fewer than half are ownership

The ownership table is the largest object in the store and the easiest figure to misquote. More than half of its rows record that two names refer to the same company, not that one company holds another - roughly the difference between a person's aliases and their relatives. Those rows are therefore drawn outside the ownership group rather than as its largest member.

6,181,628 rows total

same_entity_as3,434,524

identity resolution - the join itself, not a holding

not ownership
controls1,017,820

control assertion without a disclosed percentage

owns_share869,464

a shareholding, from 13 disclosure-register lanes

47,869 with exact %
ultimate_parent_via736,659

resolved chain to the top of the control tree

15,573 with exact %
consolidates123,161

accounting consolidation from filed statements

16,198 with exact %

2,747,104

genuine ownership edges

79,640

carry an exact percentage

2.9%

of ownership edges

The percentage layer is thin, and we describe it as thin. It comes from 13 disclosure registers - SEC 13D/G, Korea DART, ASX, AFM, BaFin, HKEX, China CNINFO, Brazil CVM and others - each of which only captures holdings above a statutory threshold. Sub-threshold ownership is structurally unobservable in every jurisdiction, so no amount of ingestion fixes it.

Panels, rather than a snapshot

Whether the data carries history determines what kind of research it can support: a current-state snapshot permits description and little more. The answer here is encouraging. The mine-safety assets form a genuine three-decade annual panel, and the corporate-control graph is dated back to 2003.

Panel spans by asset1970198019902000201020202026MSHA workforce19942026MSHA control history19702026controls edges20032026IBAMA enforcement19822026

MSHA workforce 1994-2026

employment, hours, injuries, fatalities, citations, penalties

354,524 mine-years · 28,258 mines · 30 years

MSHA control history 1970-2026

dated controller spells with start and end

111,488 spells · 47,101 mines · 33,175 controllers

controls edges 2003-2026

annual corporate-control graph, 20k-130k edges/yr

1,014,692 dated edges

IBAMA enforcement 1982-2026

Brazilian environmental sanctions, CNPJ-keyed

15,517 acts · 5,015 companies · R$4.53bn

A natural experiment already present in the data

Dated changes of controller on one side; thirty years of annual mine outcomes on the other; both keyed on the same identifier. That pairing supports a familiar research design - when a mine changes hands, what happens to employment, injury rate and citation rate, and does the answer depend on who the incoming owner turns out to be?

Controller change as a treatment event on a mine outcome panel19942026controller changeobserved outcomepost-change pathcounterfactual

the treatment

facility_control_history - 111,488 dated controller spells across 33,175 controllers, each with a start and an end date.

the outcome panel

facility_workforce - 354,524 mine-years over 30 years: employment, hours, injuries, fatalities, citations, penalties.

Both are keyed on the same msha_mine_id, so the join is exact rather than fuzzy. The ownership graph then identifies who the incoming controller is and what else they hold worldwide - the step that needs both a data layer and a modelling one. This is a design rather than a result; it has not been run.

What comes out, and one measure that does not yet exist

Ten registered analytical surfaces render into 65 commodity dossiers. The most interesting candidate is not yet among them: published supply-risk measures almost always express concentration by country, while the layer needed to express it by owner is already in place.

by country - computed today

0.137

diversified

Top country Chile at 26.48% of mine production, 2025. Every published supply-risk metric works this way.

by beneficial owner - not yet computed

?

unknown

We have not found this published at global scale, which is unsurprising: it requires the cross-registry entity resolution described in section 02. The testable claim is that a country-diversified supply base can still be owner-concentrated, in which case diversification measured on geography would overstate the resilience actually present.

65 commodities x 25 sections

commodity_dossierfacility_dossiermatrix_criticality_concentrationmatrix_recycling_gapmatrix_demand_uncertaintydeposit_atlascommodity_economicssupply_cost_curveslifecycle_flowsindustry_chain

Rendering into production, reserves, capacity, concentration, midstream facilities, demand outlook, deposits, geochemistry, trade, pricing, recycling, material flows, environmental LCA, facility operations, ownership, risk, permitting, components, lifecycle, events and export restrictions - plus an explicit per-commodity gaps section.

Operating records59%
All records33.8%

44,416 of 75,315 operating records carry a named operator. Enough to demonstrate owner concentration on well-covered commodities; not enough for a global ranking until production weighting and the attribution gap are handled explicitly.

Four stages, each of them gated

Fetch, resolve, validate, adapt - across four persisted tiers and a single command-line entry point. The topology is deliberately unremarkable. What matters is that every transition is gated, and that each record carries a structured provenance chain back to a primary source, much as a laboratory sample carries its chain of custody.

Four-stage pipeline from raw fetch to adapter outputfetchraw01resolveresolved02validateready03adaptadapter-output04
fetch

132 source-specific fetchers + 75 curated seeds, each with a metadata card

resolve

cross-source identity resolution onto one entity spine

validate

at-rest invariants + the dossier-invariant defect-class suite

adapt

emit for one downstream consumer - GCAM, PROMMIS, Hector, Xanthos

152,624

source lines

1,556 files

83,179

test lines

889 files

132

fetchers

+ 75 curated seeds

24

GB ready tier

84 GB all tiers

Enforced at commit: lint and format, type checking, architectural import boundaries, a file-length cap, and a dossier-invariant suite that encodes known defect classes as executable rules - for instance forbidding any sum of component parts against a world total, an error this domain produces repeatedly. Every record carries a structured provenance chain: source id, fetch date, URL, hash, page or section, citation, licence, and a nested chain for derived records.

208 sources

official survey97
primary government46
reputable compilation43
curated14
industry association4
commercial2
primary company2

Every one of the 208 carries a metadata card - publisher, homepage, coverage, cadence, access method, format, licence, access URL, and an explicit list of known traps in that source.

Delivered in GCAM's own input format

Strata does not run GCAM and makes no claim to. It emits GCAM inputs in GCAM's own file shapes, including ready-to-run 7.2 supply-constraint XML - delivered in the receiving model's gauge rather than requiring a translation step. The adapter also records, per commodity, which layer rests on measurement and which on a screening assumption.

vintage 2026-08-05

11

GCAM 7.2 XMLs

ready-to-run supply constraints

431

A11 curve rows

865 in vintage form

13

commodities

26 GCAM regions

54

A32 intensity rows

confidence-graded

aluminiumcobaltcoppergraphitelithiummanganesenickelplatinumrare earthssteeltelluriumuraniumvanadium

Strata does not run GCAM and does not claim to. It emits GCAM inputs, in GCAM's own file shape.

Quantity + regionUSGS MCS reserves, unit-normalised, rolled to GCAM-32measured
Cost - uraniumNEA/IAEA Red Book cost-of-recovery pyramidmeasured
Cost - all othersbenchmark C1 cash-cost screening model, 1975 USD/kgscreening
Grade split - copperUSGS grade-tonnage model (measured deposit-grade distribution)measured
Grade split - all othersfixed screening proportion - reserves carry no measured grade pyramidscreening

223

region-commodity pairs hold reserves but no cost model

The adapter ships its own drop ledger - 229 rollups dropped in total, 223 of them for this one reason - so a country missing from a supply curve is visible rather than silent. That number is the joint work item in a single line: the reserves exist, an empirically grounded cost does not. Replacing the screening cost model with asset-level economics improves their model rather than selling them ours.

The research-grade core is clear to publish

Collaborative work can stall late if a central input turns out to be encumbered, so it is worth establishing early. Six of 208 sources carry any restriction at all, and the single share-alike licence among them can be isolated.

6 of 208

sources carry any restriction. The research-grade core is publishable.

brightqueryhard

internal commercial corpus

cannot be redistributed

brightquery_globalshaped

commercial - only canonical IDs + derived analytics may egress

derived results publishable, raw rows are not

cvm_frehard

ODbL-1.0 share-alike

a derived database inherits the licence - isolate it

baci_v202601_hs22shaped

CC-BY-NC-4.0

fine for academic publication, blocks commercial use

co_dian_customsshaped

derived aggregates only

aggregate outputs only

cl_aduana_customshandled

ODbL vintages 2009-2018

already excluded from default pulls

The other 202 are public-domain, CC-BY, open-government or free-with-attribution. The one to design around early is brightquery_global: derived analytics may egress, raw entity rows may not - which happens to suit a joint paper exactly, since the contribution is the derivation rather than the rows.

Stated plainly, rather than left to be discovered

Each item below is a known limitation with a measured size. None is fatal to the directions in the next section, but each one shapes what can honestly be claimed.

01

Occurrence versus operating

306,251 of 482,552 commodity-attributed records are geological occurrences; 16,133 are operating and 2,197 corroborated.

02

Geographic skew

87% of records sit in four countries - entirely a publication-regime artifact of open cadastral data, not a statement about geology.

03

Operator attribution

59% on operating records, 33.8% overall. This is what bounds the owner-concentration work today.

04

Exact ownership percentages

79,640 edges - 2.9% of genuine ownership edges. Statutory disclosure thresholds make sub-threshold holdings structurally unobservable in every jurisdiction.

05

Commissioning year

4.1% populated (37,140 of 898,749). Not usable as a vintage variable.

06

Cost models

223 region-commodity pairs have reserves but no cost model; outside uranium, every cost is a screening estimate rather than an observation.

07

Heuristic attribution emitted as fact

Process role, status and commodity attribution are partly heuristic in some paths. Row-level derivation labels are the known fix, not yet complete everywhere.

08

Dated material-flow data

The copper flow matrix latest year is 2004 (Yale STAF vintage).

Three directions, ranked by what the measurements allow

Each of these would draw on a modelling capability as well as a data one. The ranking reflects not how interesting they sound, but what the attribution rates, panel spans and drop ledgers above will actually support today.

The adapter emits A11 and 11 GCAM 7.2 supply-constraint XMLs today. The gap is named by its own drop ledger: 223 region-commodity pairs hold reserves but no cost model. Replacing the screening cost model with asset-level economics improves the receiving model rather than merely supplying it with data.

boundLowest technical risk. Most credible December deliverable.

The three axes are an editorial ranking, not a measurement. The feasibility evidence behind them is the measured material above - attribution rates, panel spans, and the adapter drop ledger.

Every figure was measured against the shipped data/ready tier on 2026-08-20, and the queries behind them are kept alongside the source brief. Nothing here is projected. Where the data cannot support a number, the page says so rather than showing an estimate - the same rule the platform itself follows, described at /methodology.