💬 Language & Media

Wiki Articles' First Named Entity's Language Origin Drift

Tracks the language origin of the first named entity in Wikipedia articles over time.

This subject is not publishing figures yet — its collector is still being built.

No figures yet

This subject has been created but its first collection run hasn't produced data yet.

About this data

The page measures the proportion of first named entities (people, places or things) in Wikipedia articles that belong to three language families: Indo‑European, Sino‑Tibetan and Afro‑Asiatic. For each article, the earliest named entity is identified, its language of origin is determined from Wikidata or language links, and the counts are aggregated into yearly or decadal buckets.

Data are sourced directly from the public Wikipedia dump and the Wikidata knowledge base, using automated scripts that extract the first entity label and map it to its language family. No external surveys or proprietary data are used; the calculation is transparent and reproducible.

Observing these proportions reveals subtle shifts in what the world chooses to highlight in its encyclopaedic record. Changes in the dominance of certain language families can reflect evolving cultural, economic or geopolitical interests, offering a unique lens on global attention that is not tracked by any other known source.

Why this isn't published anywhere else

No live dashboard, recurring report, or public dataset was found that tracks the language origin drift of the first named entity in Wikipedia articles; only general linguistic concept pages appear in search results.

Uniqueness score 0.98 — assessed against live web search results when this subject was created.