|
||
|
||
Co-authored by David Castillo Parra1 and Christian Dawson.
There are approximately 7,100 known living languages in the world, but AI’s most popular tools support fewer than 1% of them. ChatGPT, one of the most widely used AI systems, officially supports 59 of them. That gap is not a technical footnote. It is the defining challenge of the AI era.
AI is rapidly becoming the primary way billions of people access the Internet. It shapes how people search for information, access services, and communicate across borders. But these systems are only as inclusive as the data and infrastructure they are built on. In Sub-Saharan Africa, less than 12% of the population currently has access to AI. For many communities, the barrier is not only connectivity or affordability, but whether AI systems can understand, process, and serve users in the languages they’re most comfortable speaking. The communities left out today risk becoming permanently invisible to the next generation of AI systems.
The problem is not just about translation. It runs deeper, into the Internet’s foundations. True language enablement requires integrated digital infrastructure that recognizes every script, identifiers such as domain names and email addresses, locally relevant content, representative datasets, and applications designed to serve communities in their own languages. Without these foundational layers working together, translation alone cannot ensure that people can fully participate in the digital economy or be represented in the AI systems increasingly shaping it.
Today, millions of people using non-Latin scripts encounter websites and platforms that fail to recognize their names, email addresses, or domain names. The Internet was built on the assumption that users would communicate in a limited set of Latin characters and domain endings like .com or .org. That assumption no longer reflects reality.
Fixing this requires a technical standard called Universal Acceptance (UA). UA is the principle that all valid domain names and email addresses work equally across the Internet, regardless of language, script, or character length. Without it, whole communities are locked out of basic digital participation. And as AI systems are trained on existing infrastructure, these gaps do not just persist. They get encoded and amplified at global scale.
There are signs of progress. Governments and international organizations are beginning to treat linguistic diversity as a core requirement for digital inclusion. Recent examples include:
Declarations alone will not close the gap. What is needed is coordinated investment in three specific areas:
Language inclusion cannot be a late-stage add-on. If a language is absent at the infrastructure level, it becomes invisible to AI. The communities who speak it lose access to the economic, educational, and social benefits AI is delivering everywhere else.
AI has the potential to be the most powerful force for linguistic and cultural inclusion the world has ever seen. It also has the potential to be the most efficient engine of exclusion. The decisions being made right now about infrastructure, datasets, and language support will determine which it becomes. If AI is going to rewrite the Internet, it cannot do so in only a handful of voices.
Sponsored byRadix
Sponsored byCSC
Sponsored byDNIB.com
Sponsored byVerisign
Sponsored byWhoisXML API
Sponsored byVerisign
Sponsored byIPv4.Global