| 1 | GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay | Boning Li, Longbo Huang | cs.DC | 2026-09-10 |
| 2 | The widening evaluation gap in medical large language model research 2023 to 2026 | Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif | cs.CL | 2026-09-10 |
| 3 | MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions | Antoine Saillenfest | cs.LG | 2026-09-10 |
| #4 | Legible Failures: Detecting and Repairing In-Context Binding Errors | Manas Venkata Sai Ravulapalli, Samrath Singh Chadha, Abhinav M. Hari | cs.LG | 2026-09-10 |
| #5 | Breaking Predictions Is Not Enough: Specified-Foil Counterfactuals for Temporal Graphs | Minwoo Yu, Young-guk Ha | cs.AI | 2026-09-10 |
| #6 | Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation | Anqi Peter Li, Kaden Kim | cs.LG | 2026-09-10 |
| #7 | scDEFT: A deep learning framework for drug-effect prediction and counterfactual reasoning | Murthy Devarakonda | q-bio.QM | 2026-09-09 |
| #8 | Counterfactual Marginalisation: Framework for Evaluating Robustness to Nuisance Variables | Yasin Ibrahim, Hermione Warr, Robin J. Evans +1 | cs.LG | 2026-09-09 |
| #9 | CARRE: Counterfactual Action Retrieval and Reason Evaluation for Explainable Churn Prescription | MinJoo Kim, SanJin Park, SeungHwan Cho | cs.CL | 2026-09-09 |
| #10 | Do LLMs Make More Mistakes If They Do Not Believe the Input Data? | Peter Kochelka, Aleš Manuel Papáček, Vojtěch Dvořák +1 | cs.CL | 2026-09-08 |
| #11 | NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting | Tobias Susetzky, Raphael Rehms, Dmitrii Seletkov +5 | cs.LG | 2026-09-08 |
| #12 | Evaluating and Improving Evidence-Grounded Fact-Checking in LLMs via Multi-Round Evidence Ablation | Xingyu Deng, Mingzi Cao, Nikolaos Aletras +2 | cs.CL | 2026-09-08 |
| #13 | What Fixed-Rollout pass@k Evaluations Can Identify | Pranav Singh, Prashant Singh | stat.ML | 2026-09-08 |
| #14 | CAR-MIL: Counterfactual Attention Regularization for Multiple Instance Learning | Imane Chraki, Pierre Marza, Stergios Christodoulidis +1 | cs.CV | 2026-09-08 |
| #15 | What Eviction Destroys: A Restore-Counterfactual Audit of Forgetting in Agent Memory | Chen Shen | cs.CL | 2026-09-08 |
| #16 | Three Types of Negation of Triple and its Elements and an Extension of Triple | Zhenghua Pan | cs.AI | 2026-09-08 |
| #17 | Evidence-Aligned Entity Verification for Hallucination Detection in Retrieval-Augmented Generation | Runsong Jia, Zhen Fang, Mengjia Wu +2 | cs.AI | 2026-09-08 |
| #18 | ActionSplice: In-Flight Action Editing for Interactive World Models | Pardis Taghavi, Tingyu Guo, Jonas Lossner +2 | cs.CV | 2026-09-08 |
| #19 | Vision: Data-Centric Anchoring for Robust and Interpretable Agentic AI | Arun Vignesh Malarkkan, Xinyuan Wang, Yanjie Fu | cs.AI | 2026-09-08 |
| #20 | Less Is Personal: Learning Minimal Sufficient User Profiles for Personalized Language Models | Minghang Liu, Qiang Qiu, Yuanzhuo Wang +2 | cs.AI | 2026-09-08 |
| #21 | TaskGuard: Task-Conditioned Restoration Utility for Risk-Aware Object Detection | Vung Pham | cs.CV | 2026-09-07 |
| #22 | AVCG: A Generalized Variational Framework for Counterfactual Generation under Hypothesis Distributions | Jamie Duell, Alejandro Jimenez Rodriguez, Mahault Albarracin | cs.LG | 2026-09-07 |
| #23 | InfluenceField: A Differentiable Field with Interventionally Identifiable Causal Structure for Multimodal World Modeling | Zihao Yang, Zijia Wang, Zhiqiu Huang | cs.LG | 2026-09-07 |
| #24 | Modus Tollens and Counterfactuals and Counterfactual Reasoning Based on Three Types of Negation | Zhenghua Pan | cs.AI | 2026-09-07 |
| #25 | SkillAlign: Aligning Skill Interfaces for LLM-based Agents | Shuo Ren, Xiaomian Kang, Jiajun Zhang | cs.AI | 2026-09-07 |
| #26 | Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise | Mika Okamoto, Gabriele Sarti | cs.AI | 2026-09-07 |
| #27 | Generalist Open-World Temporal Perception | Cristian Sminchisescu | cs.CV | 2026-09-06 |
| #28 | Counterfactual Tests for Measuring Chain-of-Thought Faithfulness in Visual Language Models | Bayar Menzat, Maximilian Süss, Ruizhi Wang +3 | cs.CV | 2026-09-06 |
| #29 | MemCorr-DP: Counterfactual Correspondence Conditioning for a Diffusion Policy Guided by a Reference | Tan Su, Haoxiang Yang, Ruxin Wang +1 | cs.RO | 2026-09-06 |
| #30 | A Statistical and Machine Learning Framework for Quantifying Offensive Impact in Professional Box Lacrosse | Robert Jimerson | cs.LG | 2026-09-06 |
| #31 | Reading Decoder Trajectories: Training-Free Counterfactual Query-Trajectory Reliability for Small-Object Detection | Zhaoning Shi, Bo Ma | cs.CV | 2026-09-06 |
| #32 | Separating Capability from Confidence: Grounded Dual-State Calibration for GRPO-Trained Medical Vision-Language Models | Yangyang Xie, Ke Hao, Jiaqi Liu +2 | cs.CV | 2026-09-06 |
| #33 | CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review | Jicheng Zhou, Kemou Li, Kahim Wong +5 | cs.AI | 2026-09-04 |
| #34 | A Comparative Study of Counterfactual Explainers for Graph Neural Networks Enabling Multiple Types of Graph Edit | Maria Myrto Villia, Filippos Gouidis, Theodore Patkos +1 | cs.LG | 2026-09-04 |
| #35 | Confounding-Valid Conformal Inference for Counterfactual KPIs in Wireless Networks | Abdessamed Qchohi, Jessica Moysen Cortes, Matteo Zecchin | cs.LG | 2026-09-04 |
| #36 | Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents | Chao Yao, Yangbo Wei, Zhen Huang +5 | cs.CR | 2026-09-04 |
| #37 | How Faithful Is Attribution for Sales Forecasting? A Counterfactual Study | Glib Kechyn | cs.LG | 2026-09-04 |
| #38 | DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems | Zehao Wang, Lanjun Wang, Shilong Jin +2 | cs.AI | 2026-09-04 |
| #39 | A Computationally Feasible Framework for Causal Probabilistic Explanation | Rafal Urbaniak, Sam Witty, Daniel Waxman +7 | cs.AI | 2026-09-03 |
| #40 | The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs | Yumeng Shi, Quanyu Long, Yin Wu +1 | cs.CV | 2026-09-03 |
| #41 | Common-Witness Certificates and Sharp Feature Bounds for Counterfactual Image Auditing | Usef Faghihi, Amir Saki | cs.AI | 2026-09-03 |
| #42 | Pushing the (Decision) Boundaries: Dynamically Calibrating Differentially Private Noise to Explainability in Federated Learning | Michael Khavkin, Kichang Lee, Jaeho Jin +2 | cs.LG | 2026-09-03 |
| #43 | Rethinking World Models for Safety-Critical Embodied Systems | Kailang Ma, Heye Huang, Inhi Kim +1 | cs.AI | 2026-09-03 |
| #44 | Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation | Yan Tang, Tingyu Cao, Yuanbo Tang +2 | cs.AI | 2026-09-03 |
| #45 | Counterfactual Routing Using Integer Programming with Constraint Generation | Daniël Vos, Sterre Lutz | cs.AI | 2026-09-03 |
| #46 | When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents | Wen-Yu Chang, Yun-Nung Chen | cs.CL | 2026-09-03 |
| #47 | Guide, Not Bind: Why Defeasible Priors Fail in Augmented Lagrangian Causal Discovery | Sairam Sundararaman, Sara Girdhar, Manit Narasimha Murthy +2 | cs.LG | 2026-09-03 |
| #48 | It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories | Yigit Utku Bulut | cs.LG | 2026-09-03 |
| #49 | Lngram v2: Latent N-Gram Memory with Interpretable Discrete Representations | Yunao Zheng, Bin Wen, Xiaojie Wang | cs.CL | 2026-09-03 |
| #50 | Accountable AI with Grounded, Faithful, Consistent, Actionable Rationales: A Case Study in Clinical Trial Matching with VERDICT | Zikai Zhou, Yufei Jin, Yilin Xu +3 | cs.CL | 2026-09-03 |