31 lines
2.4 KiB
Markdown

# Methodology and limitations
## Inclusion
The seed selects names with sustained or count-based evidence in US or UK birth-registration series from 1900 onward, then retains every required hand-curated canonical, diminutive, nickname and variant from the prototype. The result is intentionally broad enough to include established names of many linguistic origins used in English-speaking societies.
## Usage profiles
- `en-US`: derived from US Social Security Administration annual national name counts.
- `en-GB`: derived from the unified ONS, National Records of Scotland and NISRA series. Historical ranked rows lacking counts receive a decreasing rank-based proxy so the period remains usable without implying exact births.
- `en-IE`: Northern Ireland evidence with conservative UK/US smoothing; labelled as an estimate rather than Republic of Ireland direct statistics.
- `en-CA`: a documented US/UK evidence blend.
- `en-AU` and `en-NZ`: documented UK/US evidence blends.
Within each locale and band, gender weights use observed counts with light smoothing. Usage weights use a logarithmic scale relative to the strongest name in the same evidence pool, preventing a few extremely common names from flattening the remainder. Zero-count bands retain a low usage weight and a broader all-era gender prior with reduced confidence.
## Sources
- SSA Popular Baby Names: https://www.ssa.gov/oact/babynames/limits.html
- SSA qualifications: https://www.ssa.gov/OACT/babynames/background.html
- ONS baby names: https://www.ons.gov.uk/peoplepopulationandcommunity/birthsdeathsandmarriages/livebirths/datasets/babynamesinenglandandwalesfrom1996/1996tocurrent
- UK unified processing and upstream source links: https://github.com/rossbowen/babynames
- SSA mirror and processing documentation: https://github.com/hackerb9/ssa-baby-names
- Curated nickname mappings: https://github.com/dxdc/babynames
- Irish CSO reference for future direct enrichment: https://www.cso.ie/en/statistics/birthsdeathsandmarriages/irishbabiesnames/
## Limitations
Registration spelling does not prove etymological canonicity, and official privacy thresholds omit very rare names. Locale estimates are deliberately labelled. This package is a deterministic knowledge layer, not an onomastic authority or demographic classifier. Future releases may replace blended profiles with direct official observations without changing stable name keys or consumer semantics.