PaperScope
LIVE · 2026-09-03 05:40 UTC

CHASE: How Content Ecosystems Are Reshaped When Ranking Is the Only Target

Qianwen Gao, Zichang Su, Yiwen Hou, Arlen Kumar, Leanid Palkhouski

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.30466 v1
Category
Submitted
2026-08-31

Abstract

Generative Engine Optimization (GEO) is increasingly used to improve content visibility in LLM-based retrieval systems, yet its population-level effects under repeated optimization remain poorly understood. We introduce Content Homogenization under rAnking Signal Exploitation (CHASE), a controlled simulation framework for studying how content ecosystems are reshaped when creators repeatedly adapt documents to an LLM ranking signal. We use ranking as a proxy for source visibility and validate this abstraction against citations in grounded generated responses, obtaining a rank-citation AUC of 0.853 $\pm$ 0.093 across six domains. CHASE then iterates ranking, feature discrimination, rewriting, and evaluation over 20 rounds across different domains. Quality-ranking alignment decreases in all six domains: from R0 to R20, the change in Spearman's rho ranges from -0.107 to -0.018, with a mean change of -0.068, which means documents closer to the ranking feature profile become less aligned with independently judged document quality over the simulation horizon. A random-target control has shown that it is associated with adaptation toward ranking-derived incentives rather than iterative rewriting alone. The resulting ecosystem dynamics are strongly domain-dependent. Together, these findings show how repeated optimization against a fixed LLM ranking signal can reshape both content populations and the incentives faced by content creators.

Comment: Accepted to the Conference on Language Modeling (COLM) 2026

arXiv abs page · PDF · same-day batch