OpenData · Strata
What the data holds, what it derives, and how it is built
01 · Orientation
Capabilities first, plumbing as the evidence
Most descriptions of a data platform open with its architecture. This one does not, because architecture rarely conveys what becomes possible. The pipeline is set out in full further down, but it appears as evidence beneath each capability rather than as the opening argument. The question the page is arranged around is a research one: which problems become tractable once a supply-chain model can see corporate structure, and not only country totals?
What is distinctive here?
An entity-resolution layer, rather than a minerals dataset.
sections 02, 05
What can be derived from it?
Ten analytical surfaces, and one measure not yet in common use.
sections 07, 08
What could it support scientifically?
Three candidate directions, ranked by what the measurements allow.
sections 13
02 · What is distinctive
The value is in the join, rather than the rows
Minerals data is, for the most part, public - USGS, national surveys, study groups - and any research group can obtain it. The scarce ingredient is the layer that makes those sources composable against corporate reality: a legal-entity backbone able to establish that a Peruvian cadastre record, an SEC filer, a Dutch disclosure and a Brazilian tax number all describe one economic actor. Two hundred registers are, in effect, two hundred directories written in different naming conventions. This is the concordance between them.
208
catalogued sources
206 of them public
3,434,524
same_entity_as edges
the resolution layer itself
6,895,243
entity records
1,708,611 resolved clusters
brightquery~470M documentsUS-strong legal-entity and corporate-structure corpus. CIK + LEI + EIN + ticker + ISIN + FIGI resolved onto one row.
brightquery_global~124M rows · 95 jurisdictionsRest-of-world legal-entity backbone. Contributes LEI and (jurisdiction, company number) registration numbers.
Two of the 208 sources are proprietary. The other 206 are public, and any team could fetch every one of them. What is harder to assemble is the identifier hub that makes those 206 composable - the reason a Peruvian cadastre record, an SEC filer, a Dutch disclosure and a Brazilian tax number can be shown to describe one economic actor.
03 · Evidence
The headline count is accurate, and easily misread
898,749 facility records is a correct number and an easily misread one. An atlas marks every place a mineral has ever been found; a census counts what is actually in production today. This corpus is largely the former, with a smaller, well-evidenced core of the latter - so the narrowing below carries more meaning than the headline, and it is what determines whether a claim built on this data will survive review.
What the 482,552 attributed records actually are
The three largest contributing registries are cadastral, not operational: USGS MRDS (237,087), Brazil ANM SIGMINE (267,423) and Peru INGEMMET (66,263) record that something was found or claimed, not that anything is being produced. We hold a very large asset gazetteer with a small, well-evidenced operating core. The gazetteer gives global reach; the operating core is what supports quantitative claims. They are not interchangeable.
Evidence tier, all commodities
Only the bottom two tiers carry independent confirmation. The top three are catalogue presence, not observation - which is exactly why the pyramid narrows so hard.
04 · Reach
Global in reach, uneven in publication
The corpus spans 200 countries, yet 92% of records sit in four of them. That distribution reflects which governments publish bulk cadastral data rather than where minerals happen to be, which is why a country ranking built on raw record counts should be read with care.
Records by country
92.3% in the top four
China produces more mined output than any other country and appears here ninth, at 0.46% of records. Brazil sits second at 30%, on the strength of one open cadastre. The ranking measures publication regimes, not geology - much as a global rainfall map drawn only from countries that happen to operate weather stations would say more about instruments than about rain.
05 · Ownership
Six million edges, of which fewer than half are ownership
The ownership table is the largest object in the store and the easiest figure to misquote. More than half of its rows record that two names refer to the same company, not that one company holds another - roughly the difference between a person's aliases and their relatives. Those rows are therefore drawn outside the ownership group rather than as its largest member.
ownership_relationship, decomposed
6,181,628 rows total
same_entity_as3,434,524identity resolution - the join itself, not a holding
not ownershipcontrols1,017,820control assertion without a disclosed percentage
owns_share869,464a shareholding, from 13 disclosure-register lanes
47,869 with exact %ultimate_parent_via736,659resolved chain to the top of the control tree
15,573 with exact %consolidates123,161accounting consolidation from filed statements
16,198 with exact %2,747,104
genuine ownership edges
79,640
carry an exact percentage
2.9%
of ownership edges
The percentage layer is thin, and we describe it as thin. It comes from 13 disclosure registers - SEC 13D/G, Korea DART, ASX, AFM, BaFin, HKEX, China CNINFO, Brazil CVM and others - each of which only captures holdings above a statutory threshold. Sub-threshold ownership is structurally unobservable in every jurisdiction, so no amount of ingestion fixes it.
06 · Time depth
Panels, rather than a snapshot
Whether the data carries history determines what kind of research it can support: a current-state snapshot permits description and little more. The answer here is encouraging. The mine-safety assets form a genuine three-decade annual panel, and the corporate-control graph is dated back to 2003.
Observed span, by asset
MSHA workforce 1994-2026
employment, hours, injuries, fatalities, citations, penalties
354,524 mine-years · 28,258 mines · 30 years
MSHA control history 1970-2026
dated controller spells with start and end
111,488 spells · 47,101 mines · 33,175 controllers
controls edges 2003-2026
annual corporate-control graph, 20k-130k edges/yr
1,014,692 dated edges
IBAMA enforcement 1982-2026
Brazilian environmental sanctions, CNPJ-keyed
15,517 acts · 5,015 companies · R$4.53bn
07 · Experiment
A natural experiment already present in the data
Dated changes of controller on one side; thirty years of annual mine outcomes on the other; both keyed on the same identifier. That pairing supports a familiar research design - when a mine changes hands, what happens to employment, injury rate and citation rate, and does the answer depend on who the incoming owner turns out to be?
the treatment
facility_control_history - 111,488 dated controller spells across 33,175 controllers, each with a start and an end date.
the outcome panel
facility_workforce - 354,524 mine-years over 30 years: employment, hours, injuries, fatalities, citations, penalties.
Both are keyed on the same msha_mine_id, so the join is exact rather than fuzzy. The ownership graph then identifies who the incoming controller is and what else they hold worldwide - the step that needs both a data layer and a modelling one. This is a design rather than a result; it has not been run.
08 · Derivations
What comes out, and one measure that does not yet exist
Ten registered analytical surfaces render into 65 commodity dossiers. The most interesting candidate is not yet among them: published supply-risk measures almost always express concentration by country, while the layer needed to express it by owner is already in place.
Copper concentration, two ways of asking
by country - computed today
0.137
diversified
Top country Chile at 26.48% of mine production, 2025. Every published supply-risk metric works this way.
by beneficial owner - not yet computed
?
unknown
We have not found this published at global scale, which is unsurprising: it requires the cross-registry entity resolution described in section 02. The testable claim is that a country-diversified supply base can still be owner-concentrated, in which case diversification measured on geography would overstate the resilience actually present.
Registered analytical surfaces
65 commodities x 25 sections
commodity_dossierfacility_dossiermatrix_criticality_concentrationmatrix_recycling_gapmatrix_demand_uncertaintydeposit_atlascommodity_economicssupply_cost_curveslifecycle_flowsindustry_chainRendering into production, reserves, capacity, concentration, midstream facilities, demand outlook, deposits, geochemistry, trade, pricing, recycling, material flows, environmental LCA, facility operations, ownership, risk, permitting, components, lifecycle, events and export restrictions - plus an explicit per-commodity gaps section.
Operator attribution, commodity-facility join
44,416 of 75,315 operating records carry a named operator. Enough to demonstrate owner concentration on well-covered commodities; not enough for a global ranking until production weighting and the attribution gap are handled explicitly.
09 · Pipeline
Four stages, each of them gated
Fetch, resolve, validate, adapt - across four persisted tiers and a single command-line entry point. The topology is deliberately unremarkable. What matters is that every transition is gated, and that each record carries a structured provenance chain back to a primary source, much as a laboratory sample carries its chain of custody.
fetch132 source-specific fetchers + 75 curated seeds, each with a metadata card
resolvecross-source identity resolution onto one entity spine
validateat-rest invariants + the dossier-invariant defect-class suite
adaptemit for one downstream consumer - GCAM, PROMMIS, Hector, Xanthos
152,624
source lines
1,556 files
83,179
test lines
889 files
132
fetchers
+ 75 curated seeds
24
GB ready tier
84 GB all tiers
Enforced at commit: lint and format, type checking, architectural import boundaries, a file-length cap, and a dossier-invariant suite that encodes known defect classes as executable rules - for instance forbidding any sum of component parts against a world total, an error this domain produces repeatedly. Every record carries a structured provenance chain: source id, fetch date, URL, hash, page or section, citation, licence, and a nested chain for derived records.
Catalogue by authority tier
208 sources
Every one of the 208 carries a metadata card - publisher, homepage, coverage, cadence, access method, format, licence, access URL, and an explicit list of known traps in that source.
10 · GCAM handoff
Delivered in GCAM's own input format
Strata does not run GCAM and makes no claim to. It emits GCAM inputs in GCAM's own file shapes, including ready-to-run 7.2 supply-constraint XML - delivered in the receiving model's gauge rather than requiring a translation step. The adapter also records, per commodity, which layer rests on measurement and which on a screening assumption.
Emitted bundle
vintage 2026-08-05
11
GCAM 7.2 XMLs
ready-to-run supply constraints
431
A11 curve rows
865 in vintage form
13
commodities
26 GCAM regions
54
A32 intensity rows
confidence-graded
Strata does not run GCAM and does not claim to. It emits GCAM inputs, in GCAM's own file shape.
What is measured, and what is a screening model
223
region-commodity pairs hold reserves but no cost model
The adapter ships its own drop ledger - 229 rollups dropped in total, 223 of them for this one reason - so a country missing from a supply curve is visible rather than silent. That number is the joint work item in a single line: the reserves exist, an empirically grounded cost does not. Replacing the screening cost model with asset-level economics improves their model rather than selling them ours.
11 · Publishable
The research-grade core is clear to publish
Collaborative work can stall late if a central input turns out to be encumbered, so it is worth establishing early. Six of 208 sources carry any restriction at all, and the single share-alike licence among them can be isolated.
6 of 208
sources carry any restriction. The research-grade core is publishable.
brightqueryhardinternal commercial corpus
cannot be redistributed
brightquery_globalshapedcommercial - only canonical IDs + derived analytics may egress
derived results publishable, raw rows are not
cvm_frehardODbL-1.0 share-alike
a derived database inherits the licence - isolate it
baci_v202601_hs22shapedCC-BY-NC-4.0
fine for academic publication, blocks commercial use
co_dian_customsshapedderived aggregates only
aggregate outputs only
cl_aduana_customshandledODbL vintages 2009-2018
already excluded from default pulls
The other 202 are public-domain, CC-BY, open-government or free-with-attribution. The one to design around early is brightquery_global: derived analytics may egress, raw entity rows may not - which happens to suit a joint paper exactly, since the contribution is the derivation rather than the rows.
12 · Thin spots
Stated plainly, rather than left to be discovered
Each item below is a known limitation with a measured size. None is fatal to the directions in the next section, but each one shapes what can honestly be claimed.
Occurrence versus operating
306,251 of 482,552 commodity-attributed records are geological occurrences; 16,133 are operating and 2,197 corroborated.
Geographic skew
87% of records sit in four countries - entirely a publication-regime artifact of open cadastral data, not a statement about geology.
Operator attribution
59% on operating records, 33.8% overall. This is what bounds the owner-concentration work today.
Exact ownership percentages
79,640 edges - 2.9% of genuine ownership edges. Statutory disclosure thresholds make sub-threshold holdings structurally unobservable in every jurisdiction.
Commissioning year
4.1% populated (37,140 of 898,749). Not usable as a vintage variable.
Cost models
223 region-commodity pairs have reserves but no cost model; outside uranium, every cost is a screening estimate rather than an observation.
Heuristic attribution emitted as fact
Process role, status and commodity attribution are partly heuristic in some paths. Row-level derivation labels are the known fix, not yet complete everywhere.
Dated material-flow data
The copper flow matrix latest year is 2004 (Yale STAF vintage).
13 · Candidates
Three directions, ranked by what the measurements allow
Each of these would draw on a modelling capability as well as a data one. The ranking reflects not how interesting they sound, but what the attribution rates, panel spans and drop ledgers above will actually support today.
The adapter emits A11 and 11 GCAM 7.2 supply-constraint XMLs today. The gap is named by its own drop ledger: 223 region-commodity pairs hold reserves but no cost model. Replacing the screening cost model with asset-level economics improves the receiving model rather than merely supplying it with data.
boundLowest technical risk. Most credible December deliverable.
The three axes are an editorial ranking, not a measurement. The feasibility evidence behind them is the measured material above - attribution rates, panel spans, and the adapter drop ledger.
On this page
Every figure was measured against the shipped data/ready tier on 2026-08-20, and the queries behind them are kept alongside the source brief. Nothing here is projected. Where the data cannot support a number, the page says so rather than showing an estimate - the same rule the platform itself follows, described at /methodology.