How many sources become one record

Eleven registers describe the same world in their own words. Reconciliation is the work of deciding when two records are one asset - and, more often, proving that they are not.

Every source describes the world in its own words

No register is wrong; they simply disagree about names, boundaries and identity. Reconciliation resolves that - and these are the four ways it goes wrong when done carelessly, each guarded by a sweep over every shipped page.

False empty

An empty state denying what a sibling field already holds.

A commodity page said it had no export controls while the same payload carried three.

Asserted as measured

A placeholder or hardcoded value presented with the authority of a measurement.

A model card read status “phase 2” while a live health probe sat unused two lines above it.

Resolution key

A grouping key missing a discriminator silently splits or fuses records.

One antimony producer ranked twice at 25 kt because its name arrived in two spellings.

Population mismatch

Two statistics compared across different denominators.

A pair count of 57,438 was really 40,073 - the filter compared sets whose element order was not stable.

Four of these classes now have permanent guards that run over every shipped commodity page on each build. A rule only earns a place there once a real instance has been found live - the sweep is a record of things that happened, not a checklist of things that might.

The same mine, described four times

Each register mints its own record for a site it observes, which is correct behaviour for each of them. Reconciliation turns four true records into one asset - without merging things that only look alike.

before · 6 facility records

ICMMChuquicamata
USGS MRDSChuquicamata
USGS MASChuquicamata
WikidataChuquicamata
USGSChuquicamata Copper Smelter
USGSChuquicamata Copper Refinery

after · 3 facilities

Chuquicamata
one pit · 4 records merged
Chuquicamata Copper Smelter
kept separate
Chuquicamata Copper Refinery
kept separate

The smelter and refinery sit at the pit's exact coordinates and score as strong a match as the duplicates do. Proximity alone cannot tell them apart - the names can.

What a merge has to prove

Name similarity is the weakest possible evidence for identity, so it is never enough on its own. Each gate removes a population the previous one could not tell apart.

Same name, same country328,705

raw candidates

Within 2 km133,150

proximity gate

Not a placeholder name81,362

"Unknown" is not a name

Different sources40,073

one source's own ids are its answer

Merged25,450

facilities actually collapsed

Each gate removes candidates that look identical to the previous one. The largest single cut is proximity: Escondida is a common Chilean place name carried by 33 records from latitude −19.7 to −30.2, which name matching alone would have fused into one mine.

A reconciled count is smaller, and truer

A reconciled census is smaller than an unreconciled one. Nothing is lost: every source record survives, attached to the single asset it was always describing.

CommodityFacility censusChange
Copper
88,766 → 76,566
-13.7%
Lead
76,768 → 70,053
-8.7%
Gold
67,377 → 63,254
-6.1%
Zinc
68,896 → 64,184
-6.8%
Silver
41,473 → 36,771
-11.3%

Across the platform the facility corpus moved from 898,749 to 873,299 records. No source was dropped and no site was removed; 25,450 duplicate records were attached to the asset they had always described.

The automation we measured, then declined

The stronger claim a data platform can make is not what it automated, but what it declined to automate after measuring the damage.

Match company names automatically

5 lanes measured

Best coverage was 5.4%, and the specificity needed to do better fused governments, unrelated firms and private individuals.

Merge records inside one source

48.8k pairs examined

98.6% carried distinct source identifiers - the source had already decided they were different sites. About 179 were real duplicates.

Rank on how a name matched

regression set of 6 queries

Ranking by match shape alone put shell companies above the producer they were named after.

Where automation was refused, identity is asserted by hand instead, and every assertion carries evidence a reviewer can re-check from the register or filing it came from - never from the fact that two names look alike. Records that share a name but carry two separate registrations stay separate, however similar they read.

Measured twice, because once is not enough

A reconciliation decision is only as good as the measurement behind it, and a first measurement is often wrong in ways that read as right.

  • Unit tests pass on invented strings. A name-cleaning rule was green across its whole test file while it was truncating real company names in production. String rules are now checked against the actual corpus, not against examples chosen to illustrate them.
  • A fix can make the symptom worse. Consolidating one producer's duplicate records gave its cluster a new display name that no longer contained the ticker people search for - so the company it had just repaired dropped further down the results.
  • Two hundred-odd numbers are easy to get wrong. A candidate count reported here as 40,073 was first computed as 57,438, because a set comparison depended on element order that was never guaranteed.

Where the record lives

Findings are tracked individually rather than summarised, including the ones closed by deciding not to act.

Each defect carries its instance, its measurement and its outcome. Two of the most recent were closed as refusals: one because its premise did not survive measurement, another because the missing link it asked for was a deliberate guard against fusing companies that held the same concession in different decades.