| 1 | Optimal Low-Rank Quantum State Tomography with Bounded-Sample Joint Measurements | Ashwin Nayak, Xingyu Zhou | quant-ph | 2026-09-09 |
| 2 | Cross-Model Agreement as a Deployment-Time Reliability Signal for Automatic Polyp Segmentation | Siddharth Gupta, Jitin Singla | cs.CV | 2026-09-09 |
| 3 | IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier | Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan | cs.CL | 2026-09-09 |
| #4 | Maverick: Private and Verifiable LLM Inference Made Practical via Matrix-Vector Multiplication Delegation | Ben Merbaum, Mohammad Amin Raeisi, Wenhao Wang +3 | cs.CR | 2026-09-09 |
| #5 | DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs | Bhuvan Arora, Devesh Saraogi, Sravya Varada +1 | cs.CL | 2026-09-09 |
| #6 | Forward-Free LLM Depth Pruning via Weight Redundancy | Vincent-Daniel Yun, Woosang Lim | cs.LG | 2026-09-09 |
| #7 | Oracle Complexity of Stochastic Fixed-Point Equations with Nonexpansive Maps | Jelena Diakonikolas, Cristóbal Guzmán, David Martínez-Rubio | math.OC | 2026-09-08 |
| #8 | VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models | Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain +6 | cs.CV | 2026-09-08 |
| #9 | Measuring LLM Sycophancy under Sustained Multi-Turn Pressure | Leyuan Tang, Kangda Wei, Tianyu Jiang +1 | cs.CL | 2026-09-08 |
| #10 | Evaluation of Contextual Understanding in Large Language Models | Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe +4 | cs.CL | 2026-09-08 |
| #11 | Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers | Louis Yiven Zhu | cs.CY | 2026-09-08 |
| #12 | Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models | Luka Debevc, Nishan Chatterjee, Antoine Doucet +2 | cs.CL | 2026-09-08 |
| #13 | Leveraging contextual events on structure-aware next activity prediction | Alessandro Mele, Claudia Diamantini, Domenico Potena | cs.LG | 2026-09-08 |
| #14 | In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning | Iliano Fasolino | cs.CR | 2026-09-08 |
| #15 | A Quantitative Evaluation Framework for Temporal Explainability in Echocardiographic Video Segmentation | Jiyoo Noh, Jonathan H. Chan | cs.CV | 2026-09-07 |
| #16 | A Black-Box Adversarial Attack on Human Pose Estimation and Keypoint-Based Action Recognition Models | Kacper Mroczek, Michal Kepski | cs.CV | 2026-09-07 |
| #17 | Scoring Without the Engine: Validating a Deterministic, Manipulation-Resistant Content Score for Generative Engines, End to End | Elisha Bajemon, Andre-Louis Rochet | cs.AI | 2026-09-07 |
| #18 | Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy | Ayoub Kirouane, Georgios Giaples, Christos Petrocheilos | cs.RO | 2026-09-07 |
| #19 | FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect | Hazel H. Kim, Andrew M. Bean, Guilherme Affonso Ferreira de Camargo +10 | cs.CL | 2026-09-07 |
| #20 | The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs | Ziyue Feng, Hongbo Fang, James A. Evans | cs.CL | 2026-09-07 |
| #21 | BEFORE THE FLIP: Measuring Hidden Score Shifts In Quantized Vision Language Models Before The Answer Changes for Visual Question Answering | Sourajit Saha, Shubhashis Roy Dipta, Shaswati Saha +2 | cs.CV | 2026-09-07 |
| #22 | Measuring GEO Visibility: Prompt Corpora Define the Answer Market | Olivier Martinez | cs.IR | 2026-09-06 |
| #23 | Counterfactual Tests for Measuring Chain-of-Thought Faithfulness in Visual Language Models | Bayar Menzat, Maximilian Süss, Ruizhi Wang +3 | cs.CV | 2026-09-06 |
| #24 | Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence | Urja Pawar, Rajitha Ramanayake, Nabeel Kemal +4 | cs.AI | 2026-09-04 |
| #25 | Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models | José Luciano Verçosa Marques, Frederico Jorge Heitmann, Daniel Omar Perez +2 | cs.AI | 2026-09-04 |
| #26 | Measuring the Novelty of Biomedical Papers Using the Latent Distances between Knowledge Units | Yi Zhao, Heng Zhang, Yuzhuo Wang +3 | cs.DL | 2026-09-04 |
| #27 | Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny? | Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert | cs.AI | 2026-09-04 |
| #28 | On Epistemic Diversity in Large Language Models | Elisabeth Kirsten, Nicole Krämer, Muhammad Bilal Zafar | cs.CL | 2026-09-04 |
| #29 | Catalogue Photography as a Cold Start: Toward Deployable Carbide Burr Recognition | Abilash Philip Madavath, Chandra Yuvesh Aubeeluck, Augustin Raju +3 | cs.CV | 2026-09-03 |
| #30 | WorldReward: Reward Modeling for Camera-Conditioned World Models | Yibin Wang, Zehan Wang, Junshu Tang +13 | cs.CV | 2026-09-03 |
| #31 | Pushing the (Decision) Boundaries: Dynamically Calibrating Differentially Private Noise to Explainability in Federated Learning | Michael Khavkin, Kichang Lee, Jaeho Jin +2 | cs.LG | 2026-09-03 |
| #32 | SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation | Marco Cipriano, Leonardo Zini, Alexandra Schild +5 | cs.AI | 2026-09-03 |
| #33 | KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents | Yaxing Lyu, Shengjie Zhou, Binbin Toh +2 | cs.AI | 2026-09-03 |
| #34 | Modern Transformers Are Implicit Hybrids: From Functional Differentiation to Principled Hybrid Architecture Design | Runlin Shi, Bojian Yin, Guoqi Li | cs.LG | 2026-09-02 |
| #35 | Dimension Dependent Correlation Gap Bounds under Restricted Independence | Arjun Ramachandra | math.PR | 2026-09-02 |
| #36 | Doppio: A Dataset for Contactless Weight Estimation of Falling Particles | Simon Kiefhaber, Jan-Martin O. Steitz, Julia Grabinski +5 | cs.CV | 2026-09-02 |
| #37 | Do Large Language Models Capture the Diversity in their Training Data? | Youqi Wu, Farzan Farnia | cs.CL | 2026-09-02 |
| #38 | Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial Imaging | Khawaja Murad ul Hassan, Mehran Ebrahimi | cs.CV | 2026-09-02 |
| #39 | ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval | Aaryan Kapoor, Md Abdullah Al Hafiz Khan | cs.SE | 2026-09-01 |
| #40 | FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making | Vahid Reza Khazaie, Ahmed Y. Radwan, Shaina Raza | cs.CV | 2026-09-01 |
| #41 | Contribution-Aware Bandwidth Allocation for Multimodal Split Learning | Iason Ofeidis, Leandros Tassiulas | cs.LG | 2026-09-01 |
| #42 | Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity | Sinjini Banerjee, Tim Marrinan, Anand D. Sarwate | stat.ML | 2026-09-01 |
| #43 | How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation | Elitsa Yotkova, Violeta Kastreva, Petar Velkov +4 | cs.CL | 2026-09-01 |
| #44 | Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades | Dushyant Rajput | cs.AI | 2026-09-01 |
| #45 | Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations | Yi Fei Cheng, Fan Yang, Iremsu Bas +3 | cs.AI | 2026-09-01 |
| #46 | CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs | Maryam Alshehyari, Dushyant Singh Chauhan, Samuele Poppi +3 | cs.LG | 2026-09-01 |
| #47 | Lightweight Interpretable RGB-Guided Hyperspectral Super-Resolution under Real Cross-resolution Misalignment | Mohamad Jouni, Aurélien Godet, Mauro Dalla Mura | eess.IV | 2026-09-01 |
| #48 | Sim2Signal: Sim-to-Real Benchmarks for Traffic Signal Control | Ferdous Al Rafi, Susrik Mukherjee, Latika Liladhar Dekate +6 | cs.LG | 2026-09-01 |
| #49 | Beyond the Clock: Measuring the Value of Adaptive Revision | Ayushi Chadha | cs.AI | 2026-09-01 |
| #50 | Measuring Optimal Transport in Transformer Depth | Alexandre Quemy | cs.CL | 2026-09-01 |