PaperScope
LIVE · 2026-09-10 05:40 UTC

measuring 13 papers this week · +17% WoW

papers mentioning "measuring" in title/abstract · 30d window

Latestcs.CLcs.LGcs.AIcs.CV

Mentions per Day (30d)

Latest Papers

#TitleAuthorsCatDate
1Optimal Low-Rank Quantum State Tomography with Bounded-Sample Joint MeasurementsAshwin Nayak, Xingyu Zhouquant-ph2026-09-09
2Cross-Model Agreement as a Deployment-Time Reliability Signal for Automatic Polyp SegmentationSiddharth Gupta, Jitin Singlacs.CV2026-09-09
3IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model IdentifierBlake Stenstrom, Charangan Vasantharajan, Brian Sathianathancs.CL2026-09-09
#4Maverick: Private and Verifiable LLM Inference Made Practical via Matrix-Vector Multiplication DelegationBen Merbaum, Mohammad Amin Raeisi, Wenhao Wang +3cs.CR2026-09-09
#5DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMsBhuvan Arora, Devesh Saraogi, Sravya Varada +1cs.CL2026-09-09
#6Forward-Free LLM Depth Pruning via Weight RedundancyVincent-Daniel Yun, Woosang Limcs.LG2026-09-09
#7Oracle Complexity of Stochastic Fixed-Point Equations with Nonexpansive MapsJelena Diakonikolas, Cristóbal Guzmán, David Martínez-Rubiomath.OC2026-09-08
#8VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language ModelsZaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain +6cs.CV2026-09-08
#9Measuring LLM Sycophancy under Sustained Multi-Turn PressureLeyuan Tang, Kangda Wei, Tianyu Jiang +1cs.CL2026-09-08
#10Evaluation of Contextual Understanding in Large Language ModelsSubavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe +4cs.CL2026-09-08
#11Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI DevelopersLouis Yiven Zhucs.CY2026-09-08
#12Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language ModelsLuka Debevc, Nishan Chatterjee, Antoine Doucet +2cs.CL2026-09-08
#13Leveraging contextual events on structure-aware next activity predictionAlessandro Mele, Claudia Diamantini, Domenico Potenacs.LG2026-09-08
#14In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document PoisoningIliano Fasolinocs.CR2026-09-08
#15A Quantitative Evaluation Framework for Temporal Explainability in Echocardiographic Video SegmentationJiyoo Noh, Jonathan H. Chancs.CV2026-09-07
#16A Black-Box Adversarial Attack on Human Pose Estimation and Keypoint-Based Action Recognition ModelsKacper Mroczek, Michal Kepskics.CV2026-09-07
#17Scoring Without the Engine: Validating a Deterministic, Manipulation-Resistant Content Score for Generative Engines, End to EndElisha Bajemon, Andre-Louis Rochetcs.AI2026-09-07
#18Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action PolicyAyoub Kirouane, Georgios Giaples, Christos Petrocheiloscs.RO2026-09-07
#19FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing EffectHazel H. Kim, Andrew M. Bean, Guilherme Affonso Ferreira de Camargo +10cs.CL2026-09-07
#20The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMsZiyue Feng, Hongbo Fang, James A. Evanscs.CL2026-09-07
#21BEFORE THE FLIP: Measuring Hidden Score Shifts In Quantized Vision Language Models Before The Answer Changes for Visual Question AnsweringSourajit Saha, Shubhashis Roy Dipta, Shaswati Saha +2cs.CV2026-09-07
#22Measuring GEO Visibility: Prompt Corpora Define the Answer MarketOlivier Martinezcs.IR2026-09-06
#23Counterfactual Tests for Measuring Chain-of-Thought Faithfulness in Visual Language ModelsBayar Menzat, Maximilian Süss, Ruizhi Wang +3cs.CV2026-09-06
#24Necessary or Sufficient? Evaluating LLM Explanations With Behavioural EvidenceUrja Pawar, Rajitha Ramanayake, Nabeel Kemal +4cs.AI2026-09-04
#25Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language ModelsJosé Luciano Verçosa Marques, Frederico Jorge Heitmann, Daniel Omar Perez +2cs.AI2026-09-04
#26Measuring the Novelty of Biomedical Papers Using the Latent Distances between Knowledge UnitsYi Zhao, Heng Zhang, Yuzhuo Wang +3cs.DL2026-09-04
#27Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libertcs.AI2026-09-04
#28On Epistemic Diversity in Large Language ModelsElisabeth Kirsten, Nicole Krämer, Muhammad Bilal Zafarcs.CL2026-09-04
#29Catalogue Photography as a Cold Start: Toward Deployable Carbide Burr RecognitionAbilash Philip Madavath, Chandra Yuvesh Aubeeluck, Augustin Raju +3cs.CV2026-09-03
#30WorldReward: Reward Modeling for Camera-Conditioned World ModelsYibin Wang, Zehan Wang, Junshu Tang +13cs.CV2026-09-03
#31Pushing the (Decision) Boundaries: Dynamically Calibrating Differentially Private Noise to Explainability in Federated LearningMichael Khavkin, Kichang Lee, Jaeho Jin +2cs.LG2026-09-03
#32SVG-Score: Human-Aligned Evaluation of Text-to-SVG GenerationMarco Cipriano, Leonardo Zini, Alexandra Schild +5cs.AI2026-09-03
#33KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM AgentsYaxing Lyu, Shengjie Zhou, Binbin Toh +2cs.AI2026-09-03
#34Modern Transformers Are Implicit Hybrids: From Functional Differentiation to Principled Hybrid Architecture DesignRunlin Shi, Bojian Yin, Guoqi Lics.LG2026-09-02
#35Dimension Dependent Correlation Gap Bounds under Restricted IndependenceArjun Ramachandramath.PR2026-09-02
#36Doppio: A Dataset for Contactless Weight Estimation of Falling ParticlesSimon Kiefhaber, Jan-Martin O. Steitz, Julia Grabinski +5cs.CV2026-09-02
#37Do Large Language Models Capture the Diversity in their Training Data?Youqi Wu, Farzan Farniacs.CL2026-09-02
#38Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial ImagingKhawaja Murad ul Hassan, Mehran Ebrahimics.CV2026-09-02
#39ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding RetrievalAaryan Kapoor, Md Abdullah Al Hafiz Khancs.SE2026-09-01
#40FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-MakingVahid Reza Khazaie, Ahmed Y. Radwan, Shaina Razacs.CV2026-09-01
#41Contribution-Aware Bandwidth Allocation for Multimodal Split LearningIason Ofeidis, Leandros Tassiulascs.LG2026-09-01
#42Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicitySinjini Banerjee, Tim Marrinan, Anand D. Sarwatestat.ML2026-09-01
#43How Correct Is Your Answer? A Semantic Correctness Framework for Open QA EvaluationElitsa Yotkova, Violeta Kastreva, Petar Velkov +4cs.CL2026-09-01
#44Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving CascadesDushyant Rajputcs.AI2026-09-01
#45Measuring the Behavioral Fidelity of Long-Horizon Human Activity SimulationsYi Fei Cheng, Fan Yang, Iremsu Bas +3cs.AI2026-09-01
#46CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMsMaryam Alshehyari, Dushyant Singh Chauhan, Samuele Poppi +3cs.LG2026-09-01
#47Lightweight Interpretable RGB-Guided Hyperspectral Super-Resolution under Real Cross-resolution MisalignmentMohamad Jouni, Aurélien Godet, Mauro Dalla Muraeess.IV2026-09-01
#48Sim2Signal: Sim-to-Real Benchmarks for Traffic Signal ControlFerdous Al Rafi, Susrik Mukherjee, Latika Liladhar Dekate +6cs.LG2026-09-01
#49Beyond the Clock: Measuring the Value of Adaptive RevisionAyushi Chadhacs.AI2026-09-01
#50Measuring Optimal Transport in Transformer DepthAlexandre Quemycs.CL2026-09-01

all trends · matching is case-insensitive substring after tokenization