Datasets:
Harness the power of Ethnologue datasets for Analytics, AI and Academia
Better data plus better analysis leads to better outcomes, with licensed access to the full depth & breadth of the world’s most comprehensive source of language information. Harness our datasets as raw data files for maximum flexibility, or leverage our applied research & language consulting team for special projects. 70 years of experience combined with the world’s ‘gold standard’ in language datasets, language analytics and computational linguistics.
Ethnologue Global Dataset: the raw data files behind our website in tab-delimited format for virtually any form of advanced language analysis, including longitudinal studies. Collated and curated by hundreds of linguists, experts, and field contributors from around the world, with 7,467 languages in 242 countries and 11,292 Language-in-Country data points (28th edition) and historical datasets (1996 onwards). Users include Governments (e.g., USITC) and Top Ranking Universities.
Training Datasets: audio & text content in 1,000s languages, created over decades around a central theme, and suitable for training Artificial Intelligence & Machine Learning engines and models (including parallel data). The datasets can be licensed for academic and commercial purposes, and include ‘hard to source’ indigenous and low resource languages. Users include Big Tech (e.g. Meta’s ‘MMS’ engine) and Top Ranking Universities (e.g. Carnegie Mellon’s ‘XEUS’ engine).
Language GIS Dataset: the most comprehensive, up-to-date geographic dataset of the locations of the world’s 7,100+ living languages (aka, ‘polygons’, ‘shape files’). Designed for digital language mapping, including centroid coordinates for each language, and over seventy-five percent are also represented by boundary polygons that display the traditional homeland of each indigenous language (geodatabase GIS or Esri shapefiles). Users include Big Tech (e.g. Ancestry.com) and Economics Universities.
Digital Language Support: the world’s first and only dataset dedicated to the measurement of digital language capabilities for the world’s living languages. Each language is categorised on a 5-point scale with component scores across a range of factors. The world’s most advanced dataset to understand the Digital Language Divide. Users include Global Industry Associations (e.g., GSMA, SOMIC report and website) and Global Nonprofits (e.g. Clear Global).
Applied Research: Want a ‘slice’ of data for a project or a customized digital map? Need to source some specific language data or create an infographic? Want to conduct language research or undertake product prototyping? Just let us know what you’re aiming to achieve, and we will gladly offer free no-strings 1 hour consultation on the best options available. Users include global NGOs (e.g., WorldBank) and Nonprofits (e.g. YouVersion).
Contact Sales
Datasets are not included in subscriptions to Ethnologue.com and must be purchased separately. Please Contact Us for licensing and pricing information.