PaperScope
LIVE · 2026-09-25 05:40 UTC

post-training 15 papers this week · +78% WoW

papers mentioning "post-training" in title/abstract · 30d window

Latestcs.CLcs.LGcs.AIcs.CV

Mentions per Day (30d)

Latest Papers

#TitleAuthorsCatDate
1AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive ControlJiabin Qiu, Zixuan Chen, Hongye Cao +3cs.AI2026-09-24
2PoEM: Predicting RL Outcomes from Existing PoliciesKimia Hamidieh, Giannis Daras, Antonio Torralbacs.LG2026-09-24
3Self-Adaptive VLA for Robust Robot DeploymentHongxin Zhang, Chunru Lin, Tsun-Hsuan Wang +2cs.RO2026-09-24
#4GHOST-Q: Towards Studying Grounding Hallucinations Overlooked Under Same-score TradeOffs in Quantized VLMSSaim Rehman, Muhammad Shafiquecs.CV2026-09-24
#5AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video GenerationZhiyu Xu, Weilong Yan, Yufei Shi +4cs.CV2026-09-24
#6SEEK: Skill-Routed Evaluation with Evolvable Knowledge for Industrial SearchZhongxin Huang, Songyang Li, Renzhe Zhou +5cs.IR2026-09-24
#7Rufus-Air: An Open LLM Post-Training RecipeChia-Yuan Chang, Renyuan Cheng, Rui Feng +19cs.CL2026-09-24
#8Transcript-Supervised Post-Training of Generative Speech Enhancement on Real Recordings via Reinforce Adjoint MatchingJulius Richter, Christoph Boeddeker, Yoshiki Masuyama +4eess.AS2026-09-24
#9Post-Training Leaves Behavioral Shadows on Unrelated DecisionsZiyang Zhang, Yubin Jing, Yuanhao Zeng +3cs.CL2026-09-24
#10From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn AgentsXingyu Su, Abhishek Kumar, Qing Ping +5cs.AI2026-09-24
#11AlphaDiverse: Post-Training Local Quantitative Research Agents for Diverse Exploration in Alpha Factor MiningQingzhuo Wang, Zikun Wei, Zhihua Wei +1cs.AI2026-09-24
#12Same Bit Width, Different Outcomes: Post-Training Quantization of Text-to-Speech Across ArchitecturesSe Un Park, Yutae Kim, Junyoung Parkeess.AS2026-09-24
#13ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video GenerationZichong Meng, Chongjian Ge, Chun-Hao P. Huang +2cs.CV2026-09-24
#14Small yet Assistive: Spatially-Aware Post-Training for Low VisionRishabh Choudhary, Shreyansh Raj, Umesh Goyal +6cs.CV2026-09-23
#15Temporal Taxation Compounds Under Post-Training Compression of Whisper ModelsSrishti Ginjala, Eric Fosler-Lussier, Christopher W. Myers +1cs.CL2026-09-23
#16Thinking Leakage: A Causal Audit of NoThink Post-Training in Hybrid Reasoning ModelsZehao Liu, Vasant G. Honavarcs.LG2026-09-23
#17Mizar: A 159M-Parameter Audio-Language Model for Audio UnderstandingKaiyang Li, Shaobo Han, Yue Tian +1cs.SD2026-09-23
#18Predicting Quantization Price for Selecting PTQ Configurations Before DeploymentJunbin Qiu, Jian Mu, Weitong Zhang +1cs.CL2026-09-23
#19Exact Feedback Is Not Control: Evaluating Text-based Closed-Loop Revision in LLMsHaitong Jiang, Chunlin Liu, Yile Wang +1cs.CL2026-09-23
#20InternW0: A Foundational Physical World Model for Efficient Real-World InteractionsJisong Cai, Yao Mu, Ganlin Yang +23cs.RO2026-09-23
#21When Context Misleads: In-context Learning with Jurisdiction in Large Language ModelsPei-lin Li, Qingle Liu, Junyang Feng +4cs.CL2026-09-23
#22The Capability Manifold and ML Scaling LawsSyed Ali Raza Zaidi, Maryam Hafeezcs.LG2026-09-23
#23Pistis Technical ReportHeyun Chen, Xiaohan Lan, Jiaxi Li +17cs.AI2026-09-23
#24Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes InteractionJialu Wang, Jianing Deng, Shuqing Luo +6cs.LG2026-09-23
#25CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned EnvironmentsYuxuan Li, Will Epperson, Wesley Deng +1cs.AI2026-09-23
#26WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow MapsAbbas Mammadov, Jerry Y. Huang, Justin Lin +5cs.LG2026-09-22
#27OMatG-flash: An All-Atom Flow Map with Reinforce Adjoint Matching for Scalable Materials DiscoveryThomas Egg, Harry Winston Sullivan, Ellad B. Tadmor +1cs.LG2026-09-22
#28PACT: From Credit Assignment to Critic AlignmentJiayan Fu, Hang Xu, Yong Zhang +4cs.LG2026-09-22
#29Disaggregated Quantization: Specializing LLM Prefill and DecodeAndrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi +2cs.LG2026-09-22
#30COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-EvaluationRuike Cao, Fugen Yao, Liang Dong +3cs.LG2026-09-22
#31GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer CompressionBaher Mohammad, Ammar Ali, Stamatios Lefkimmiatiscs.LG2026-09-22
#32Visual Jev: Accurate and Efficient Decisions from Shared Visual ContextGuanxu Yu, Yuhang Yaocs.CV2026-09-22
#33The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural VarianceRojin Ziaeics.AI2026-09-22
#34Toolcompass: Guiding Tool Trialing, Not Suppressing ItJunlin Fang, Chong Zhang, Do Nguyen-Thanh +3cs.AI2026-09-22
#35Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRAWenzhi Fang, Nicholas Tzou, Lazar Valkov +1cs.AI2026-09-22
#36VLAQuantBench: Closed-Loop Evaluation of Post-Training Quantization for Vision-Language-Action ModelsJiuyi Xu, Qing Jin, Meida Chen +3cs.RO2026-09-21
#37TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom TasksBohao Wang, Chenwei Wu, Hang Zou +7cs.CL2026-09-21
#38Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed EngineersWeihang Ding, Junfei Zhancs.LG2026-09-21
#39SPHQuant: Efficient extreme low bit weight quantization for Vision-Language ModelsKewei Zhang, Zheng Chen, Haotong Qin +1cs.CV2026-09-21
#40When Quantization Preserves Accuracy but Not Evidence: Explanation-Aware Post-Training Quantization for Medical LLMsYeji Kim, Mi-Young Kim, Randy Goebelcs.CL2026-09-21
#41Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech EnhancementQing Yao, Lijian Gao, Qirong Maocs.LG2026-09-21
#42FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent FoldingHung T. Ho, Khanh D. Nguyen, Quang D. Nguyen +5cs.RO2026-09-21
#43End-to-end Jordanian dialect speech-to-text self-supervised learning frameworkAli A. Safieh, Ibrahim Abu Alhaol, Rawan Ghnematcs.CL2026-09-21
#44SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and CorroborationLinhan Luo, Lequan Lin, Dai Shi +3cs.LG2026-09-21
#45MemCalib: Benchmarking and Optimizing Memory Use in LLM AgentsRuike Cao, Fanyu Zhao, Fugen Yao +6cs.LG2026-09-21
#46Memory vs. Context? Influential Factors of Factual Recall in Language ModelsGuilhem Fouilhé, Nicholas Asher, Philippe Mullercs.CL2026-09-21
#47P2Flow: Phoneme-aware Progressive Flow Matching for Extreme Speech Super-ResolutionNingyuan Yang, Yize Li, Pu Zhao +5eess.AS2026-09-21
#48Cost-Accuracy Trade-offs: Neural Operator vs Classical Numerical SolverDaniel Zhengyu Huang, Andrew M. Stuartmath.NA2026-09-21
#49ACLArena: Agent Continue Learning in Multi-stage Post-trainingHaixin Wang, Xiaoxuan Wang, Junkai Zhang +9cs.AI2026-09-21
#50Measuring the Assistant's Harmlessness Preferences on the User TurnJord Nguyencs.CL2026-09-20

all trends · matching is case-insensitive substring after tokenization