PaperScope
LIVE · 2026-09-15 05:40 UTC

Not All Duplicates Are Coordination: Generic vs. Non-Generic Duplicate Campaigns in Information Operations

Ashfaq Ali Shafin, Khandaker Mamun Ahmed

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.13671 v1
Category
Submitted
2026-09-12

Abstract

Duplicate content is widely used to study coordinated behavior in social media information operations (IOs), but not all repetition provides equally meaningful evidence of coordination. Generic, reusable, or low-information posts may create noisy account-account links when projected into coordination graphs. We study this problem using 187,000 English-language tweets from six Russian Twitter Information Operations datasets. We introduce a generic/non-generic distinction for duplicate campaigns, label tweets using an LLM-assisted protocol with independent human validation, and train supervised classifiers over sentence embeddings to scale the labels. We construct duplicate campaigns using lexical similarity and two embedding-based methods. Generic campaigns are rare under lexical matching but account for nearly 39% of campaigns detected by embedding-based methods. Restricting graphs to non-generic campaigns reduces graph size and the largest connected component while increasing density, suggesting a smaller but more focused coordination structure. These findings show that duplicate-based coordination analysis should consider both textual similarity and semantic specificity.

Comment: Accepted in the 11th Workshop on Natural User-generated Text (W-NUT collocated with EMNLP 2026)

arXiv abs page · PDF · same-day batch