2.4 KiB

Methodology and limitations

Inclusion

The seed selects names with sustained or count-based evidence in US or UK birth-registration series from 1900 onward, then retains every required hand-curated canonical, diminutive, nickname and variant from the prototype. The result is intentionally broad enough to include established names of many linguistic origins used in English-speaking societies.

Usage profiles

  • en-US: derived from US Social Security Administration annual national name counts.
  • en-GB: derived from the unified ONS, National Records of Scotland and NISRA series. Historical ranked rows lacking counts receive a decreasing rank-based proxy so the period remains usable without implying exact births.
  • en-IE: Northern Ireland evidence with conservative UK/US smoothing; labelled as an estimate rather than Republic of Ireland direct statistics.
  • en-CA: a documented US/UK evidence blend.
  • en-AU and en-NZ: documented UK/US evidence blends.

Within each locale and band, gender weights use observed counts with light smoothing. Usage weights use a logarithmic scale relative to the strongest name in the same evidence pool, preventing a few extremely common names from flattening the remainder. Zero-count bands retain a low usage weight and a broader all-era gender prior with reduced confidence.

Sources

Limitations

Registration spelling does not prove etymological canonicity, and official privacy thresholds omit very rare names. Locale estimates are deliberately labelled. This package is a deterministic knowledge layer, not an onomastic authority or demographic classifier. Future releases may replace blended profiles with direct official observations without changing stable name keys or consumer semantics.