PaperScope
LIVE · 2026-09-03 05:40 UTC

failure 14 papers this week · -53% WoW

papers mentioning "failure" in title/abstract · 30d window

Latestcs.CLcs.LGcs.AIcs.CV

Mentions per Day (30d)

Latest Papers

#TitleAuthorsCatDate
1EarlyEval: Cheaper Agent Evaluation via Early Outcome PredictionYuling Shi, Zhensu Sun, Junsen Dong +3cs.CL2026-09-02
2DKL: Decoupled Knowledge Learning for Instruction-Tuned Language ModelsKushagra Bhushan, Meghanadh Pulivarthi, Sai Krishna Reddy Sathi +7cs.CL2026-09-02
3From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMsUrja Pawar, Rajitha Ramanayake, Owen O'Neill +4cs.CL2026-09-02
#4When Persona Attributes Improve Population Alignment in Large Language ModelsLeon Fröhling, Jens Rupprecht, Markus Strohmaier +1cs.CL2026-09-02
#5PragAlign: Feedback-Guided Pragmatic Alignment for Controlled Synthetic Dialogue GenerationSmitha Muthya Sudheendra, Jaideep Srivastavacs.CL2026-09-02
#6Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral AbstractionsJiayi Bi, Yanjie Gao, Yuanmin Xie +4cs.AI2026-09-02
#7Towards a Foundational Ontology for Identifying and Resolving Contradictions in Dialogue-based Human-Robot InteractionsMaitreyee Tewari, Michele Persianics.HC2026-09-02
#8LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic GuardrailsVansh Wahics.AI2026-09-02
#9PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic TasksYuyao Zheng, Haipeng Sun, Junwei Bao +4cs.AI2026-09-02
#10Beyond Context Windows: Persistent Discovery Context for Data-Centric AgentsJalal Mahmudcs.AI2026-09-02
#11Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization GapNirajan Kunwor, Sanjaya Poudel, Quoc-Huy Trinh +2cs.CV2026-09-02
#12Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step SupervisionSitong Pan, Yipeng Shen, Yilin Lu +3cs.AI2026-09-02
#13The Dynamics of Continuous Mixture Collapse in Language ModelsAli Backourcs.LG2026-09-02
#14Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM AgentsYanting Yang, Can Jin, Jinman Zhao +6cs.LG2026-09-02
#15Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM OversightXinyu Fu, Narayan Ramasubbu, Dennis Gallettacs.HC2026-09-02
#16The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory AgentsJundong Hu, Shekar Ramachandrancs.AI2026-09-01
#17Agent Memory Is a Surface for Endogenous Authorization LaunderingTommaso Cerruti, Mika Okamoto, Ansel Kaplan Erolcs.CR2026-09-01
#18VakyArth: Evaluating Pragmatic Competence in LLMs across Indic LanguagesUsneek Singh, Poorvaja Veera Balaji Kumar, Parth Nanda +4cs.CL2026-09-01
#19Ten Architectures, One Error: Shared Failure Modes in Hyperspectral Classification under Spatially Disjoint EvaluationEhsan Faghih, Fatemeh Ashrafi, Marguerite Moore +1cs.CV2026-09-01
#20MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal ModelsTawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir +5cs.CL2026-09-01
#21Evidential Deep Learning for Multi-Modal Anti-UAV DetectionDmitry Golovchits, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahagcs.CV2026-09-01
#22SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech RecognitionBiraj Subedics.CL2026-09-01
#23Adaptive Critical Token-Aware Retrieval for Repository-Level Code GenerationKefeng Duan, Dewu Zheng, Yanlin Wang +8cs.SE2026-09-01
#24Facet-0: A Robotic Foundation Model for Contact-Rich Precise ManipulationHaoyuan Deng, Haichao Liu, Wenkai Guo +6cs.RO2026-09-01
#25From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request MixOlga Tsymboi, Dmitrii Stoianov, Ramil Latypov +11cs.CL2026-09-01
#26Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectoriesNabira Rashid, Manolis Kelliscs.LG2026-09-01
#27When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce EvaluationPeiying Zhu, Sidi Changcs.AI2026-09-01
#28FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-MakingVahid Reza Khazaie, Ahmed Y. Radwan, Shaina Razacs.CV2026-09-01
#29Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their ObserversEgor Pakhomov, Erik Nijkampcs.AI2026-09-01
#30When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-TuningYitong Guo, Xiaoyi Chen, Siyuan Zhang +2cs.CR2026-09-01
#31Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution SpeedsClinton Enwerem, John S. Baras, Calin Beltacs.RO2026-09-01
#32EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM SystemsJun Hou, Priya Pitre, Yi Fang +1cs.AI2026-09-01
#33Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVREsther Xincs.CL2026-09-01
#34CHARM: Character Hallucination for Multicultural Role Play BenchmarkSunkyung Han, Nahyeon Park, Gaeun Seo +2cs.CL2026-09-01
#35ExBind: A Controlled Diagnostic Benchmark for Visual-to-Executable CorrespondenceZiqian Wang, Yuxiao Cheng, Tingxiong Xiao +1cs.CV2026-09-01
#36Explore Before Committing: Hypothesis-Guided Search for Deep Research AgentsRuochen Zhou, Zhengyu Chen, Luan Zhang +3cs.CL2026-09-01
#37HiLRP: Toward One Trustworthy Explanation for Vision Transformer: Conservation-Valid Attribution via Attention PrimitivesSathiyamohan Nishankar, Pubudu Sanjeewani, Asanka Perera +1cs.CV2026-09-01
#38What Does an Agentic Software Engineering Benchmark Measure? Profiling Task Demands and Agent Behaviour Beyond What Category Labels RevealRadin Shayanfar, Keheliya Gallaba, Ahmed E. Hassancs.SE2026-09-01
#39MeRoPE: Metric Rotary Position Embedding for Camera-Controlled Video GenerationZhijian Qiao, Xinjiang Wang, Jiajie Chen +5cs.CV2026-09-01
#40One Prompt Is Enough: Watermark Laundering Through Foundation Image ModelsJidong Yang, Qi Li, Wei Zong +5cs.CV2026-09-01
#41Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive AgentsLiming Pu, Xiaoxia Li, Yifu Liu +2cs.LG2026-09-01
#42MutMem-V2: Cryptographically Authorized Mutation in Persistent Agent Memory Portable Verification and Reproducible EvidenceWalid Saidics.CR2026-09-01
#43Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screeningZhilong Song, Lixue Chengcond-mat.mtrl-sci2026-09-01
#44Hints Help But Do They Teach? Evaluating Skills Transfer in Code GenerationWill Badrcs.SE2026-09-01
#45When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIPShota Sato, Hajime Kiyama, Tosho Hirasawa +1cs.CL2026-09-01
#46Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language ModelsJiayu Ding, Zhuodong Liu, Lei Zhang +6cs.CV2026-09-01
#47Disclosure-Gated User Simulation for Companion-Agent EvaluationYao Liu, Yu Hecs.CL2026-09-01
#48Candidate-Expanding Routing with Permutation-Stabilized Experts for Mixed-Format Medical VQAHai-Dang Nguyen, Huy-Hieu Phamcs.CV2026-09-01
#49Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-CallingKangjia Zhao, Jiajun Li, Haozhan Shen +8cs.CL2026-09-01
#50A Dataset for Modeling Iterative Problem-SolvingFagun Patel, Sang T. Truong, Duc Q. Nguyen +4cs.CL2026-09-01

all trends · matching is case-insensitive substring after tokenization