# Methods and reproducibility

This is a descriptive analysis of published modeled immigration composition, with independent flow weights. It is not a re-estimation of the authors' model or a causal study. Source bytes were acquired on 8 September 2026.

## Source acquisition and version identity

Composition: Zenodo record 20814220, version 2, published 23 June 2026. Exact CSV: `Immigration_composition_by_RCSL_20260520.csv`, MD5 `525374940bbffcdfcb9418de77fc1d8c`. All nine supporting R scripts are preserved, together with the record metadata. Every file's MD5 matches that metadata. The paper is Yildiz and Abel (2026), Scientific Data, DOI 10.1038/s41597-026-07969-8. The linked census supplement is included.

Independent flows: Abel–Cohen pseudo-Bayesian closed demographic-accounting estimates (`da_pb_closed`). Frozen author repository tree: `8ffca255dc15d149d1fc804963d519d29d47d693`. The flow files were last changed at `d0c74c860b5654a69c92255fcdee7b118ec0ddde` on 28 April 2023. The newer Figshare 2025 total and sex-type downloads returned an access error during this exercise. The accessible authors' matrix mirror is used consistently instead. Its exact upstream CSV version is not independently recoverable; it is not described as the latest UN vintage. The dataset authors' recommendation to use the separately estimated total series for magnitude is followed. Female and male estimates supply only their relative shares.

The JSON matrices include 231 country/territory nodes and 11 duplicate regional aggregate nodes. Remove all regional nodes before aggregating country matrices. Rows are origins and columns destinations. Diagonals are zero. All flow cells are nonnegative. Column sums reconcile to separately rounded country inflows within 15 persons in the preferred total series. Use matrix column sums consistently. Do not sum region and country nodes together. Region diagrams sum only country-to-country moves.

The original explorer visualization source code and README were inspected for provenance but are not redistributed. The D3 visualization here is newly authored. Published underlying flow data carry the Abel–Cohen CC BY 4.0 attribution. Natural Earth geometry is public domain.

## Unit of observation and cleaning

Keys retained: destination `iso3c`, five-year `period`, `sex`, `age`, `education`, estimated `prop`. The enriched retained file preserves `prop_raw`, cleaned `prop`, external `sex_weight`, combined `share`, total flow and implied modeled counts. These implied counts are products of independent estimates, not observed microdata.

The country-period-SEX distribution has 55 structural cells: three child ages with `Under 15`, plus thirteen adult/young-adult age bands crossed with four education labels. It is not a 16 × 5 complete Cartesian grid. Structural combinations such as a 5–9-year-old in a completed education group are inapplicable, not missing observations. Each present 55-cell distribution sums to one; summing both sexes without weights would produce two.

All 123,090 source rows are present, with no missing values or duplicate keys. All 2,238 present groups have 55 cells. Maximum absolute sum error is 3.1086244689504383e-15. There are 274 negative floating-point residues, minimum −7.133359717722099e-17, all in the 15–19 post-secondary cell. Clip values below zero to zero and renormalize within country-period-sex. The largest cell adjustment is 4.996e-16. Any negative value below −1e-12 or sum error above 1e-10 would fail the analysis assertions.

Fifteen destinations appear only in 2015–20: ATG, BRB, BRN, DJI, ERI, ESH, GRD, GUM, LBY, MRT, MYT, PNG, SYC, UZB, VIR. This creates 150 absent country-period-sex groups and 8,250 absent expected joint rows. There are no partial groups. Additional unusable early flow matches include CUW, MNE, SDN, SRB and SSD. Historical Sudan and current Sudan remain distinct; no composition is invented for split states. The source metadata contain an ambiguous duplicate Chile label; Chile is explicitly mapped to CHL, not the Channel Islands code CHI.

## Combined composition and measures

Let p(c,t,s,a,e) be the cleaned sex-conditional share and w(c,t,s) the independently estimated sex flow divided by female plus male flow. Combined q = p × w. For every usable destination-period, sum q = 1. The main total is the independent total-flow series, not female plus male flows.

- Child share: sum q for ages 0–14.
- Prime working age: sum q for ages 25–54. This does not establish employment.
- Older-migrant share: sum q for ages 55+. This is not a retirement category.
- Tertiary proxy: sum q in the source `Post Secondary` category at ages 20+, divided by all arrivals. The paper maps this category from completed university education; the label is not assumed to be identical to all OECD ISCED definitions. Tiny positive post-secondary cells at 15–19 are excluded from this measure, following the paper's age-20 interpretation. Maximum combined country-period effect: 0.0342617 percentage points. The source-faithful partition Sankey retains those source cells.
- Tertiary 25+: post-secondary arrivals aged 25+ divided by all arrivals aged 25+.
- Female share: w(female). It is not identifiable from the composition file alone.
- Age diversity: normalized Shannon entropy −Σa qa log(qa) / log(16), using the full 16 age groups, including open-ended 75+.
- Education diversity: normalize four attainment shares among ages 25+, then compute −Σe qe log(qe) / log(4). Exclude the structural `Under 15` category. Source `No Education` means less than completed primary schooling, not necessarily no schooling.

All differences are endpoint minus baseline. Shares are reported in percentage points; entropy differences are reported as index units multiplied by 100 when compared in a heatmap. Entropy depends on the published category definition. A larger value means more even category shares, not higher attainment, better integration or welfare.

## Comparison samples, rankings and global aggregation

The balanced sample has 179 countries with all six periods and usable weights. Global shares pool implied modeled counts and divide by the independent total for this same country set. Global entropy is computed from the pooled composition, not an average of national entropy. The sample covers 97.20% of independent world flow in 1990–95 and 97.37% in 2015–20.

The default absolute-change rankings require at least 100,000 estimated arrivals in both 1990–95 and 2015–20: 87 countries qualify. Ranks retain signed changes and separately sort absolute magnitudes. Sensitivity files repeat all eight metrics using 50,000, 250,000 and 500,000 floors. The 30-country heatmap is selected independently from the largest final-period flow totals among all 231 flow destinations. Missing composition remains NA. No country size is inferred from proportions.

GCC pyramids combine BHR, KWT, OMN, QAT, SAU and ARE. The selected Southern Europe aggregate combines ITA, ESP, GRC and PRT. Both use independent total flows times external sex fractions to pool male and female age distributions. Each pyramid side sums to 100% within sex, so its shape is not a direct male/female count comparison.

The symmetric decomposition is exact: ΔP = Σ(mean(w) × Δp) + Σ(mean(p) × Δw). Endpoint destination weights each sum to one. The two terms are called within-destination composition and destination mix. They reconcile within 1e-9 percentage points. The tertiary result is 6.650778 + 1.396473 = 8.047251 points. This is descriptive accounting, not a causal policy attribution.

## CLR–PCA

Use 48 age–education cells: twelve age groups from 20–24 to 75+, times four attainment categories. Structural child schooling zeros and the near-zero 15–19 post-secondary cells are excluded by this adult scope. Pool the sexes first. All retained cells are strictly positive. Close each adult profile to one, take logs, subtract each row's mean log share, and run full-SVD PCA with four components. No pseudocount, no additional variance standardization, and equal country-period weights. There are 1,074 observations. The first two components explain 81.91% and 7.80% of variance. Scores and loadings are included. Arbitrary axis signs carry no normative meaning. Similarities may partly reflect the common source prediction model.

## Sensitivity and independent benchmark

The published census inventory contains 143 samples and 74 countries. Eight excluded samples mentioned in the paper are already absent from this inventory; the code checks their keys and does not subtract them again. Mark destinations with a listed census as contributors and others as prediction without a listed local census. This is a coverage indicator, not a prediction interval; all profiles are modeled.

Exclude the sparsest last period by ending in 2010–15. Calculate sign agreement and Spearman correlation of absolute changes on the original 87-country eligible set. Tertiary: 97.70% sign agreement and 0.922 rank correlation. Older share: 81.61% and 0.784. The atlas also allows the endpoint to be changed directly and recalculates its eligibility threshold for that endpoint.

Reweight with the authors' minimum closed demographic-accounting sex flow series. The maximum effect on tertiary-share change is 1.318 pp among default eligible countries. Female-share change is substantially more sensitive (see the exact sensitivity CSV); avoid presenting that ordering as settled evidence. No fabricated uncertainty intervals are supplied.

Eurostat migr_imm8 snapshots contain annual immigration by age and sex for DE, IT, ES and DK, 2015–19. Decode JSON-stat indices; prefer completed age (`COMPLET`), using age reached (`REACH`) for Denmark where the completed-age cells are unavailable. Sum single ages including less than one and 100+. Divide pooled age-group counts by the pooled sex-specific total over all five years. Unknown age remains in the total denominator and is documented. A mean-of-annual-shares sensitivity is also supplied. Compare to the source's sex-conditional 2015–20 composition, so independent sex weights do not enter this benchmark. Time framing, age definitions and recorded events versus estimated five-year transitions differ. Germany female 55+: model 10.15%, Eurostat 5.14%; Denmark: 11.52% versus 3.83%. This is limited external evidence, not comprehensive model validation. No education-series benchmark is claimed.

## Visualization integrity

The world map contains two independent evidence layers. Corridor width uses a square-root scale of total moves, with an explicit legend; destination circle area is proportional to inflow. Destination shading uses a diverging ±15-point scale, with exact values in the selected profile and saturated extremes. Neither corridor colour nor width is assigned an educational composition. Geodesic connectors are not observed travel paths. The interface includes all six periods, all available destination profiles, independent flow floors, census coverage, route limits, incoming/world views, geographic scale, and light/dark/system themes. Country clicking has a native selector equivalent. Query parameters restore the evidence state. JSON, CSV and SVG exports carry selected evidence.

For offline size, the interactive edge subset keeps every corridor of at least 10,000 moves plus each destination's strongest 30 incoming corridors. The map renders only the selected strongest 25/75/150 available edges with coordinates and reports their share of the full independent flow denominator. The exact bilateral CSV contains every positive country-to-country cell, including edges omitted by the browser subset. No filtering changes totals or region matrices.

The ordinary Sankey partitions one cohort by sex, age and education; it does not show changes in individual people. The circular Sankey places six regions in separate origin and destination roles. Africa combines both African regions; Europe/Central Asia combines Europe and East Europe & Central Asia; other Asia combines South, East and South-East Asia; the Americas combine both American regions. West Asia and Oceania remain separate. The chord uses the original eleven regions and directional arrowheads. Neither diagram depicts age–education transitions across periods.

## Rebuild

Python dependencies: pandas, numpy, scipy, scikit-learn, beautifulsoup4, Pillow, reportlab. Run from the package root:

```
python 04_research/code/acquire.py       # optional; frozen raw files are included
python 04_research/code/analyze.py
python 04_research/code/build_article.py
node 04_research/code/render_assets.cjs
python 04_research/code/build_article.py
node 04_research/code/build_carousel.mjs
python 04_research/code/finish_pdf.py
```

The renderer requires Node.js, Playwright and Chromium. Set `SCHYM_CHROMIUM_PATH` to an installed Chromium executable. D3 7.9.0 and fonts are included; no browser network calls are required by the article. The carousel source uses `@oai/artifact-tool` in the Codex primary runtime, with editable native text and vector shapes and embedded exact D3 chart assets. Its page PNGs are rendered from that canonical source; the PDF is built from those same page images. PowerPoint text and layout remain editable; complex charts are embedded images with SVG and JavaScript source provided separately. Exact source and build hashes are in the package manifest and checksums.

## Measured performance and package size

The offline release deliberately includes its source snapshot, exact downloads and static fallbacks. Executable code is about 103 KiB gzipped; non-executable evidence JSON is about 996 KiB gzipped. In the local headless delivery checks, switching to Bahrain took 143–191 ms after per-figure caching. Full page loading took 3.1–3.3 seconds; the automated paint observer recorded 3.4–3.9 seconds over the screenshot procedure. The 2.5-second paint target is not consistently met by that observer in this environment. These automated screenshot timings are not field Core Web Vitals certification. No external runtime request is made.

The pyramid CSV contains exact sex-conditional age marginals for all countries, plus weighted GCC and selected Southern European aggregates. The HTML download is explicitly limited to the six reference panels to avoid duplicating the full dataset in the page; the country JSON export includes any selected individual age profile. All 2,262 country/group-period-sex pyramid profiles close to one within 5.33e-15.

The separate combined-share audit checks the denominator after external sex weighting: 1,099 country-periods close to one within 6.44e-15, while 20 country-periods are explicitly flagged for missing usable sex weights. These missing weights are never treated as zero demographic shares.
