PaperScope
LIVE · 2026-09-03 05:40 UTC

verification 15 papers this week · -20% WoW

papers mentioning "verification" in title/abstract · 30d window

Latestcs.CLcs.LGcs.AIcs.CV

Mentions per Day (30d)

Latest Papers

#TitleAuthorsCatDate
1Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision FrameworkCagri Temelcs.RO2026-09-02
2Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM SystemsYihang Chen, Yuxiang Chen, Yuxuan Huang +3cs.AI2026-09-02
3Repo-To-Skill: Distilling GitHub Repositories Into AI4AI SkillsJianlyu Chen, Yuyang Hu, Hongjin Qian +8cs.AI2026-09-02
#4Dimension Dependent Correlation Gap Bounds under Restricted IndependenceArjun Ramachandramath.PR2026-09-02
#5TrajMind: Chaining Role-Specialized LoRAs for Fast-and-Slow Collective Trajectory Anomaly DiagnosisJiahao Wu, Zhenqun Yang, Chen Jason Zhang +1cs.LG2026-09-02
#6Fine-Grained Anomaly Perception in Wild UGC-Enhanced Images: A Comprehensive Dataset and Difference-Fusion FrameworkYan Zhong, Gefei Chen, Qiufang Ma +4cs.CV2026-09-02
#7When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language ModelsSmitha Muthya Sudheendra, Jaideep Srivastavacs.CL2026-09-02
#8Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit AssignmentChenyu Zhou, Qiliang Jiang, Shuning Wu +1cs.LG2026-09-02
#9YesTrack: Referring Multi-Object Tracking via MLLM-based Yes/No VerificationQuansheng Hu, Qin Sun, Qiansen Dai +4cs.CV2026-09-02
#10Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime OptimizationBizhe Bai, Jiakang Yuan, Hongming Wu +8cs.CL2026-09-02
#11Retrosynthesis of Synthetic Media for Explainable AI Provenance ForensicsYijie Lin, Ching-Chun Chang, Isao Echizen +2cs.CR2026-09-02
#12T2LSC-Bench: Benchmarking Localized Semantic Control in Text-to-Image GenerationYan Wang, Xinyi Hou, Weiguo Lin +2cs.CV2026-09-02
#13LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic GuardrailsVansh Wahics.AI2026-09-02
#14PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview AssessmentFan Yuxuan, Huang Miaojun, Zhang Haimei +2cs.AI2026-09-02
#15Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image UnderstandingShanji Liu, Kelu Yao, Junxiao Xue +5cs.CV2026-09-02
#16World-Coherent Decoding: Self-Verifying Test-Time Planning for World Action ModelsChuhan Zhang, Seiji Ito, Kenta Hoshino +2cs.CV2026-09-02
#17OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human DemonstrationsYixiong Xiao, Lang An, Hucheng Yang +9cs.HC2026-09-02
#18Rendering-in-the-Loop: An Execution-Driven Agent for Interactive Web DevelopmentYilong Guo, Hanqi Chen, Zixiao Ye +3cs.CV2026-09-02
#19ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark ConstructionKe Zhang, Yankang Liu, Roya Zandi +1cs.AI2026-09-02
#20HyGRAIL: Cost-Aware and Evidence-Grounded Scientific Hypothesis Discovery over Knowledge GraphsYihang Sun, Zhihan Zhu, Zhiyuan Jiang +3cs.CL2026-09-02
#21Privacy Washing: Detecting Internal Contradictions in Privacy PoliciesThomas Brackincs.CY2026-09-02
#22Detecting Object Hallucinations in Large Vision-Language Models via Cross-Modal Attention Drifts and Mask-Based VerificationXuanbing Wen, Boxu Chen, Le Yang +4cs.CV2026-09-02
#23GeoStore: Finding Small Storefronts in Large Scenes -- A Fine-Grained POI Localization Benchmark with Global-to-Local Asymmetric MatchingLu Han, Xiting Sun, Hao Wang +4cs.CV2026-09-02
#24ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent EvaluationsPeiying Zhu, Sidi Changcs.AI2026-09-02
#25Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM OversightXinyu Fu, Narayan Ramasubbu, Dennis Gallettacs.HC2026-09-02
#26Candidate Generation and Definition-Guided Verification for Sentence-Level Depression Symptom RecognitionWeiming Li, Catarina Barata, Miguel Constante +1cs.CL2026-09-01
#27InSight: A Benchmark for Agentic Claim Verification in Interactive VisualizationsMaeve Hutchinson, Syed Mahbubul Huq, Mohammad Albinhassan +3cs.CL2026-09-01
#28IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective SignalsMd. Atabuzzaman, Christian Alexander, Chris Thomascs.CV2026-09-01
#29EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA AgentsWei Wang, Wenqiao Zhang, Yutong Lin +14cs.RO2026-09-01
#30Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive AgentsLiming Pu, Xiaoxia Li, Yifu Liu +2cs.LG2026-09-01
#31MutMem-V2: Cryptographically Authorized Mutation in Persistent Agent Memory Portable Verification and Reproducible EvidenceWalid Saidics.CR2026-09-01
#32Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety JudgesRui Yang, Shuang Huang, Junhua Liu +7cs.CR2026-09-01
#33Revisiting Face Recognition for Monozygotic Twins: The Celeb Twins Test SetMichael Zang, Haiyu Wu, Mrinal Sharma +1cs.CV2026-09-01
#34EDRAC: Benchmarking Arabic Dialect Reading ComprehensionNoor Abo Mokh, Kirill Chirkunov, Teresa Lynn +15cs.CL2026-09-01
#35Causal Evidentiary Governance for High-Risk Machine Learning SystemsSamah Kareem, Barış Çeliktaşcs.CY2026-09-01
#36Low-Quality Face Recognition using Center Aligned Representations and Local Margin ConstraintsVedat Can Dilaver, Benjamin S. Riggancs.CV2026-09-01
#37VerNav: Verifier-First Low-Latency Vision-and-Language NavigationZhixin Wang, Chengzheyi Yao, Leyuan Liu +2cs.RO2026-09-01
#38Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste SortingUdo Schlegel, Shubhangi, Gabriel Dax +3cs.CV2026-09-01
#39Agentic programs: an emerging form of scientific software in computational materials scienceYunsung Lim, Haekwan Jeon, Jaesun Kim +2cond-mat.mtrl-sci2026-09-01
#40SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT VerificationSwapnil Bhattacharyya, Mayank Baranwalcs.AI2026-09-01
#41Controllable Image Captioning with Prompt-Conditioned Scene RewardsJongyeop Hyun, Taeyoung Kim, Hyounghun Kimcs.CV2026-09-01
#42Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary SearchEnrong Pan, Ryan Zhou, Ting Hucs.AI2026-09-01
#43A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLossSuryaansh Jain, Rahasya Barkur, Vishal G +8cs.CV2026-09-01
#44Enoki: Efficient Multi-Level Hallucination DetectionElisei Rykov, Timur Ionov, Nikolay Ivanov +5cs.CL2026-09-01
#45DeSyR: A Decoupled Symbolic Recovery Framework with PINN-Guided Structure Search and Physics-Informed Coefficient RefinementPancheng Niu, Jun Guo, Qiaolin He +2cs.LG2026-09-01
#46CoVer: Conflict-Aware Claim VerificationShuning Zhang, Dai Shi, Bohao Chu +7cs.AI2026-09-01
#47Validity-Aware Jailbreak Evaluation for Large Language ModelsQilong Wu, Sahil Wadhwa, Pranab Mohanty +2cs.AI2026-08-31
#48MemeBridge: A Dataset for Benchmarking and Mitigating the Bidirectional Cultural Gap in Meme InterpretationHangxiao Zhu, Suliu Qin, Zhuoyan Li +3cs.CL2026-08-31
#49TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge PoisoningMuhaimin Bin Munir, Akib Jawad Ononto, Nazia Shehnaz Joynab +2cs.CL2026-08-31
#50From Tool Use to Technological Agency: LoopCAT as a Local-First, Open-Source Tool for Translation Technology EducationGokhan Dogru, Adrià Martín Morcs.CL2026-08-31

all trends · matching is case-insensitive substring after tokenization