The Gates Foundation has launched a coalition of 60 partners, including Anthropic, Google and the OpenAI Foundation, to build more representative language data for AI systems, aiming to reach more than 3 billion people within five years. The effort follows the foundation's earlier $1 billion pledge to bring AI tools to health workers, teachers and farmers, and reflects a shift in focus from simply deploying AI in low-resource settings to fixing the underlying data gaps that make those deployments risky in the first place.
Bill Gates said "the smart use of AI," paired with sustained global aid, could help fight inequality, but was unusually blunt about what he sees as a foundational failure in how the industry has built these systems so far: "We failed to build in the right human values and ways of keeping our moral code controlling what these AIs do." The admission is notable coming from one of the technology sector's most prominent philanthropic voices rather than from an AI safety researcher or critic.
The coalition's own research supplies a concrete example of why that failure matters in practice, not just in theory. AI systems trained on limited language data have produced dangerous mistranslations in medical contexts, including one case where a pregnant woman's statement that her "water has broken" was rendered by an AI translation system as her having "thrown away water" — an error that, in a real clinical setting, could delay urgent care during labour. The example is the kind of concrete harm the coalition is using to justify building out representative language data across the roughly 3 billion people it hopes to eventually reach, most of whom speak languages current AI systems handle poorly.
Foundation CEO Mark Suzman said the work has to continue "full speed ahead," a position that puts the Gates Foundation somewhat at odds with AI labs and researchers, including Anthropic's own leadership, who have separately called for slowing the pace of frontier AI capability development. The foundation's stance treats those as two different problems: capability speed is one debate, and whether the AI already being deployed serves the roughly two-thirds of the world's population whose languages are poorly represented in training data is a separate, more urgent one that Suzman argues cannot wait for the capability debate to resolve.

