Monolingual LLMs in the Age of Multilingual Chatbots
地理
历史
作者
Xuenan Cao
出处
期刊:boundary 2 [Duke University Press] 日期:2025-05-01卷期号:52 (2): 3-27被引量:2
标识
DOI:10.1215/01903659-11636583
摘要
Abstract Among the many exciting features of commercial-grade chatbots such as ChatGPT is their ability to converse in multiple languages. The chatbots are families of large language models (LLMs) built on multilingual textual data, the majority of which is licensed through third-party providers that scraped from the Western internet. In an unevenly distributed internet governed by various data sovereigns, LLMs are fed content accessible from wherever their developers are based: mainly the Anglophone West. The overrepresentation of one region of the world through LLMs is an issue largely disguised by their apparently inclusive multilingual fronts. The versatile chatbots that translate texts also produce and disseminate ideological content, constituting a digital hegemony yet to be named. This essay is a step toward a understanding of the mathematics behind a new kind of language politics. Expanding the Marxist, feminist, and critical race study canons in the critique of knowledge dissemination, this essay shows that LLMs shape culture not only at narrative level but also at the deeper level of parameters and word embeddings.