LLMs Mirror Country-Specific Gender Patterns If Asked, but Skew Male When Generating Media in Local Languages
Sharif Kazemi, Tanya Popli, Neil K. R. Sehgal, Sunny Rai, Niyati Malhotra, Victor Orozco-Olvera, Ana María Muñoz Boudet, Samuel P. Fraiberger, Sharath Chandra Guntuku, Manuel Tonneau
Abstract
Large language models (LLMs) are increasingly used to generate media, but whether their content perpetuates gender stereotypes is unknown: standard benchmarks rely on selection-based formats rather than long-form generation, and surveyed baselines for local gender associations are scarce outside the West. We collect gender associations for 22 occupational and domestic roles from 695 respondents across the United States, India, Kenya, and Nigeria, and evaluate eight LLMs under two regimes: direct questioning and media generation. Models track the surveyed associations under direct questioning but skew substantially more male under media generation in major local-language cells, consistent with the male bias documented in human-produced media. Outside the US, the shift is much smaller and non-significant under English prompting, so English-only or country-agnostic evaluation would miss this bias in the languages where these models are most deployed. Instruction prompting reduces the shift directionally, but trades off against alignment with the surveyed associations. Evaluating LLM gender bias for global deployment therefore requires generation-format testing, local-language prompting, and locally-collected human baselines.