# Methods and Identification Boundary

Edition: September 11, 2026. Research question: does data-center investment expand total activity, redirect scarce resources, or do both in different locations?

## Source and Vintage Discipline

The source receipt binds every acquired input to its original URL, retrieval time, byte count and SHA-256. Inputs are kept outside the deployable tree under `.local-data/datacenter-opportunity/raw/2026-09-11`. The public source inventory includes failed acquisitions. The ERCOT December 2025 planning document returned HTTP 403 and contributes no numerical observations.

A source's latest date is read from its contents, not its filename. The first EIA sales workbook ends in 2020 and is not used for current analysis. The replacement EIA-861 workbook covers 2010-2024. Table 2 changed from GWh to MWh in 2023; its current header is authoritative. The CSO housing API includes revisions: the frozen sum for 2025 is 36,215 dwellings, compared with 36,284 in the initial release. Both vintages are described without mixing their values.

Raw recovery is permitted only when the returned bytes match the existing receipt. Revised upstream bytes are rejected. A public research ZIP carries the cleaned observations, code, source ledger and methods; it does not redistribute copyrighted source pages or guarantee that mutable upstream endpoints still serve the frozen vintage. Exact raw-source reproduction therefore requires the retained local snapshot if an upstream source has changed.

## National Construction Account

Census C30 data-center buildings are a private subcategory of Office. They include building construction and integral systems, not all owner equipment, land acquisition or announced campus investment. Servers and racks are outside the building measure.

For month t:

`NR_EX_DC[t] = TOTAL_NONRESIDENTIAL[t] - PRIVATE_DATA_CENTER[t]`

`TOTAL_EX_DC[t] = TOTAL_CONSTRUCTION[t] - PRIVATE_DATA_CENTER[t]`

`DC_SHARE_ALL_NR[t] = 100 * PRIVATE_DATA_CENTER[t] / TOTAL_NONRESIDENTIAL[t]`

`DC_SHARE_PRIVATE_NR[t] = 100 * PRIVATE_DATA_CENTER[t] / PRIVATE_NONRESIDENTIAL[t]`

Both ownership denominators are retained. The first numerator is private even when the denominator includes government construction. Office excluding data centers is computed once, so data centers are not double counted among categories. Civil infrastructure combines highways/streets, sewage/waste disposal, water supply and conservation/development across ownership; it is not labeled all public infrastructure.

Monthly values are seasonally adjusted annual rates, in nominal USD million. Annual flows sum unadjusted monthly estimates, with the month count and completeness flag preserved. January-July 2026 is a partial-year total, never a full-year estimate. Individual annual sums can differ by several million dollars from separately rounded published annual tables. The 2025 private data-center annual sum reconciles to the published $49,737 million.

Each indexed series has its own 2019 monthly mean as denominator. Nominal indexes do not adjust construction inflation. Removing the observed component is an accounting residual, not an estimate of the economy in which that investment never occurred.

## BEA Investment Account

Use the August 26, 2026 vintage of Tables 5.3.5 and 5.4.5U. The latter separately identifies data-center structures from 2020 onward in this workbook. Earlier dots are missing, not zero.

Private fixed investment is decomposed into data-center structures, manufacturing structures, remaining nonresidential structures, equipment, intellectual property, and residential investment. For component j:

`CONTRIBUTION[j,t] = 100 * (I[j,t] - I[j,t-1]) / TOTAL_PFI[t-1]`

These are percentage-point contributions to nominal private investment growth, not real GDP growth contributions. The component sum matches the total within published rounding. Quarterly levels are SAAR; quarter-on-quarter changes of those rates are not silently annualized growth rates.

The nominal residual is `PFI - BEA_DC_STRUCTURES`. Census and BEA vintages are not forced to agree. Chained 2017-dollar component values are never subtracted to construct a real residual. The data-center nominal, real-volume and price series are independently rebased to 2020=100 for the price-versus-volume exhibit. Chain volume is not building area.

No matched data-center/non-data-center floor-area panel was obtained. Midlothian's 260,000 square feet is converted using exactly 0.09290304 square meters per square foot, but not divided into the broader $600m site-development announcement. Expected permanent jobs are not reported as realized employment, and no realized jobs-per-dollar ratio is estimated.

## County Panel

The eligibility frame is 694 counties with private all-industry employment of at least 25,000 in 2015. It uses pre-period size rather than later outcomes. All 11 annual geographic rows must exist; individual outcomes may still be missing. Connecticut is excluded because its county geography changed; Puerto Rico and other territories are outside this U.S. county frame. No claim is made that the resulting sample represents all rural counties.

QCEW filters are private ownership (`own_code=5`), all establishment sizes (`size_code=0`), a valid stable county FIPS, and the requested annual industry. Industries are all employment (10), construction (23), specialty trades (238), electrical contractors (23821), plumbing/heating/air-conditioning contractors (23822), utility systems construction (2371), architectural/engineering services (5413), and data processing/hosting (518210).

Electrical 23821 is a consistent parent category, rather than a guessed six-digit code. Its pay is the average across the industry's workers, not a wage observation for the electrician occupation. Suppression flag N turns employment, payroll and wages into missing values even if the raw numeric field is zero. Establishment counts are retained as separately published counts. Entirely absent industry rows remain missing, not inferred zeros.

Annual building permits are authorized housing units, not building counts, starts or completions. All permit columns are used rather than the reported-only duplicate set. Five 2015 geographic name variants have identical FIPS and numeric data; only exact numeric/geographic aliases are collapsed. Conflicting duplicates raise an error.

## Fixed Effects and Robustness

For outcome y in county i and year t:

`y_it = county_FE_i + year_FE_t + beta * log(1 + hosting_establishments_it) + gamma * log(all_private_employment_it) + error_it`

Permits use log(1+units); positive wages and employment use natural logs. Wages are nominal. Year effects absorb common national price changes but not local inflation or composition shifts. The exposure is a broad industry proxy, not data-center MW or AI adoption.

For each of five outcomes, show five specifications: county/year effects only; plus contemporaneous employment; employment-weighted; exclusion of Loudoun; and a sample ending in 2022. Weighted estimates use fixed 2019 employment, not population weights. Control inclusion is a sensitivity, not proof of identification: contemporaneous employment may be a channel of investment.

The implementation uses statsmodels WLS after a converged weighted alternating projection of county and year effects. Cluster scores are aggregated by county. The CR1 finite-sample correction counts both absorbed fixed-effect dimensions and estimated regressors. Intervals use a t distribution with G-1 degrees of freedom, where G is the number of county clusters. Synthetic unbalanced-panel tests compare coefficients and covariance with an explicit statsmodels dummy regression.

These intervals allow within-county dependence but not arbitrary spatial dependence across counties. The specifications and outcomes are multiple comparisons; no unadjusted significant coefficient is presented as a discovered causal mechanism. Each model records its actual observations and clusters, because suppression changes the sample.

The dynamic exposure profile uses the 2019 ratio of hosting establishments to all establishments, scaled to a difference of 10 per 1,000. It interacts this fixed exposure with calendar years 2019-2025, omitting 2022. A joint cluster-robust F test examines 2019-2021 coefficients. The 2023 marker is a national era, not a local construction date. Several construction outcomes reject a flat pre-period differential path; failure to reject for another outcome does not prove parallel trends. This is not a Bartik IV design.

## Dated Project Cases

Two 2019 groundbreakings are recorded: Midlothian/Ellis County from Google's contemporaneous post and a project engineer, and Papillion/Sarpy County from contemporaneous local reporting. The date can mark a ceremony rather than first site work. Later phases are not counted again as separate initial projects. The ledger is not an exhaustive national treatment inventory.

For each case, select ten nearest comparison counties from its predeclared Census division using only 2015-2018 mean log employment, mean log pay, mean log(1+permits), and log-permit trend. Standardize features within the candidate pool; use Euclidean distance and deterministic FIPS tie-breaking. Exclude the two project counties and named major clusters. Compare to the equally weighted mean of the same counties; industry suppression can reduce available comparisons for an outcome/year and its count is displayed.

The event window is -4 to +5 years, with the previous year indexed to 100. A second clock moves the anchor to the following calendar year. Missing values remain gaps. The comparison range is its 25th-75th percentile, not a confidence or prediction interval. The pre-period RMS index gap is descriptive balance evidence, not a significance test. A test verifies that modifying all post-2018 observations cannot alter the selected comparisons.

Sites and comparison counties are not randomly assigned; comparison counties can contain other data centers. With two case events, cluster-asymptotic causal intervals would be misleading. No pooled causal project effect is estimated. County manufacturing-construction and county data-center electricity outcomes remain unavailable. The case view must not be described as the requested fully identified natural experiment.

## International and Grid Evidence

Ireland: use the separately rounded annual CSO table for annual electricity; MEC02 provides quarterly context. Total and components can differ by 1 GWh from rounding. NDQ01 dwelling completions are original counts with their own current revision vintage. No dual axis equates dwellings and electricity.

Great Britain: DESNZ's June 2026 workbook provides counts, TWh and area-specific consumption shares for 2020-2024. Shares are stored as fractions and multiplied by 100. External-serving centers are included; enterprise centers and Northern Ireland are excluded. Overlapping national/regional/local areas are never stacked. Sources of undercounting and mixed-use meters are acknowledged in the report.

Netherlands: CBS dedicated-activity electricity connections exclude institutional embedded facilities. 2024 is provisional. Published small/large components sum to 5,093 GWh; the 5,100 GWh headline is rounded. The 4.58% share is used as published, not recomputed from an incompatible total.

Australia: separate original current-price building approvals, work done, equipment and selected imports in the February 2026 ABS article. Building categories and SD59 are broader than data centers. Suppressed `np` cells remain null. Annual building activity workbooks are retained for context, not substituted for floor area.

Canada: AESO's June 2025 policy is an Alberta load-request and connection-limit snapshot, not national consumption. More than 16 GW is a lower bound; 1.2 GW is a prospective interim limit through 2028. Neither is generating capacity or a generation interconnection queue. No conversion probability or refused-project count is inferred.

Germany/EU: Frankfurt's adopted 2022 planning policy documents recognized competition for commercial land. EU operator reporting is not treated as a long national time series. No harmonized German electricity share is imputed from reporting coverage. The fixed-order country comparison leaves missing values explicitly absent and retains each source's year, denominator and observed/model status.

## Reproduction

Use the repository's locked Node dependencies (`npm ci`) and a Python environment with NumPy 2.3.5, pandas 2.3.3, SciPy 1.16.3, statsmodels 0.14.6 and openpyxl 3.1.5. The analysis was executed with the editor-selected repository environment; no dependency or interpreter settings were changed.

Run from the repository root, in order:

```text
node scripts/datacenter-opportunity/acquire.mjs
node scripts/datacenter-opportunity/build-data.mjs
python scripts/datacenter-opportunity/analyze.py
python scripts/datacenter-opportunity/electricity.py
node scripts/datacenter-opportunity/international.mjs
node scripts/datacenter-opportunity/geography.mjs
node scripts/datacenter-opportunity/design.mjs
node scripts/datacenter-opportunity/figures.mjs
node scripts/datacenter-opportunity/publish.mjs
node scripts/datacenter-opportunity/assets.mjs
```

The design step requires the current personal Schym chart/design skills; its generated contracts are included for inspection. Each data generator with `--check` compares its output without rewriting. Source and model tests run separately. The research package also includes the processed county panel, so its observational models can be reproduced without re-downloading BLS inputs. The website itself requires the native repository and its Next.js shell; it is not represented as a self-contained offline HTML file.