Daily AI Papers — September 18, 2026
Published:
1. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
Authors: DeepSeek-AI, :, Xu, Anyi, Li, B., Lin, Bangcai, Xue, Bing, Xian, BingCheng, Xu, Bingzheng, Wu, Bochao, Zhang, Bowei, Deng, Boyi, Yu, C. C., Jin, Chao, Lin, Chaofan, Dong, Chen, Wang, Chenbing, Feng, Chenfan, Lu, Chengda, Zhao, Chenggang, Deng, Chengqi, Zhang, Chengyuan, Xu, Chenhao, Zhao, Chenqi, Shao, Chenze, Wang, Chuhao, Zhang, Chuqi, Dai, Damai, Yang, Dejian, Chen, Deli, Huang, Di, Wu, Di, Li, Donghao, Li, Erhang, Fu, Eric, Zhou, F., Zhou, Fangwei, Lin, Fangyun, Yuan, Fangzhou, Xia, Feiyu, Dai, Fucong, Hao, Guangbo, Li, Guanglin, Chen, Guanting, Cao, Guoai, Fan, Guofan, Meng, Guolai, Li, Guowei, Zhang, Haichuan, Ma, Haiyang, Shen, Haiyang, Li, Han, Yu, Han, Zhang, Han, Deng, Hangyuan, Xu, Hanwei, Xu, Hanxiang, Zhong, Hanxun, Guo, Hao, Jiang, Hao, Li, Hao, Qin, Hao, Wen, Haodong, Liang, Haofen, Huang, Haofeng, Liu, Haohua, Zhang, Haoling, Luo, Haoming, Yang, Haoran, Xu, Haotian, Yuan, Haotian, Huang, Haoting, Luo, Haowen, Cai, Haoyang, Chen, Haoyu, Ji, Haozhe, Zhang, Hengran, Wang, Hengrui, Wu, Hengxu, Ding, Honghui, Tang, Hongxuan, Wang, Huadong, Cao, Huanqi, Gao, Huazuo, Qu, Hui, Zeng, Hui, Yang, J., Jin, J. H., Zhang, J. H., Zou, J. X., Yu, Jia, Zhou, Jiahui, Chen, Jiajun, Huang, Jialiang, Zhao, Jialin, Tang, Jiamin, Zhou, Jian, Tong, Jianan, Li, Jianwen, Zhu, Jiaqi, Wang, Jiarui, Ye, Jiasheng, Li, Jiashi, Xu, Jiaxin, Ding, Jiaying, Lu, Jibai, Hu, Jiewen, Yan, Jin, Zhai, Jincheng, Chen, Jingchang, Hu, Jingcheng, Zhou, Jingli, Xu, Jingsheng, Xiang, Jingting, Yun, Jingyan, Yuan, Jingyang, Cheng, Jingyuan, Zhu, Jinhua, Wang, Jinpeng, Chen, Jinyi, Hu, Jinyi, Yu, Jiping, Guo, Jueliang, Pei, Junbo, Sun, Junbo, Jiang, Junguang, Qiu, Junjie, Zhou, Junkang, Liu, Junqi, Li, Junren, Li, Junxian, Song, Junxiao, Guo, Junyi, Dong, Kai, Chen, Kaifeng, Gao, Kaige, Guan, Kang, Yuan, Kangdong, Hong, Ke, Xu, Ke, Zhao, Kefan, Ji, Kexin, Zhang, Kexin, Zhou, Kexing, Yu, Kuai, Zhang, Lan, Wang, Lean, Zhang, Lecong, Wang, Lei, Gao, Letian, Zhao, Liang, Xu, Liansheng, Guo, Lihua, Luo, Lingxiao, Fu, Lingyue, Deng, Litao, Wang, Litong, Zhang, Liyue, Chen, Longhao, Chen, Lu, Huang, Luotian, Ma, Luyao, Wang, Luyao, Di, M. S., Mei, Max, Ye, Menghao, Cui, Miao, Zhang, Mingchuan, Zhang, Minghua, Tang, Minghui, Zhang, Mingjing, Wei, Mingqi, Chen, Mingshu, Liu, Mingxing, Zhou, Mingxu, Xu, Mingyu, Yang, Mingyu, Wang, Mingze, Chen, Muyang, Shentu, Ni, Wang, Ning, Ning, Niufang, Huang, Panpan, Cong, Peixin, Wang, Peiyi, Xin, Peiyuan, Ren, Pengfei, Yan, Pengfei, Zhang, Pengle, Kang, Qi, Tang, Qi, Wang, Qiancheng, Li, Qiang, Zhu, Qihao, Li, Qingyang, Chen, Qinyu, Du, Qiushi, Guo, Qizhou, Xu, Rongxian, Ding, Rui, Hu, Rui, Tian, Rui, Yu, Rui, Zhu, Ruidong, Xu, Ruifan, Yang, Ruihan, Xia, Ruihang, Lu, Ruijie, Geng, Ruilin, Hong, Ruipeng, Ge, Ruiqi, Zhang, Ruisong, Sun, Ruize, Pan, Ruizhe, Wang, Runji, Chen, Runqian, Xu, Runxin, Tian, Ruohong, Shen, Ruomeng, Zhang, Ruoyu, X., Ryan, Liu, S. H., Lu, Shanghao, Zhou, Shangyan, Chen, Shanhuang, Cai, Shaofei, Nie, Shaoheng, Chen, Shaoyuan, Hu, Shengding, Lin, Shengkai, Ran, Shengwen, Liu, Shengyu, Jia, Shengyuan, Bai, Shi, Feng, Shi, Xu, Shicheng, Liu, Shichun, Hu, Shiqiang, Ma, Shirong, Wang, Shiyu, Feng, Shiyuan, Gong, Shufan, Lin, Shuhan, Yu, Shuiping, Zhou, Shunfeng, Yang, Shuo, Wang, Shuomeng, Guo, Shuting, Pan, Shuting, Yu, Shuying, Cao, Sinuo, Lin, Siyi, Chen, Sizhe, Chen, Songyang, Zhou, Songyang, Ni, Tao, Yun, Tao, Jin, Tian, Pei, Tian, Ye, Tian, Lin, Tianle, Ji, Tianran, Cui, Tianyi, Yue, Tianyuan, Yu, Tingting, Xiong, Tongrui, Zeng, Wangding, Liu, Wei, Zhang, Wei, Xu, Weibin, Zeng, Weihao, Zhao, Weilin, Liu, Wen, Liang, Wenfeng, Pang, Wenjie, Luo, Wenjing, Yao, Wenjing, Gao, Wenjun, Shao, Wenkai, Yang, Wenkai, Zhang, Wenli, Wang, Wenlu, Huang, Wenlve, Yan, Wenqian, Zhang, Wentao, Gao, Xi, He, Xiang, Li, Xiang, Li, Xiangli, Wang, Xiangwen, Zhang, Xiangying, Wei, Xiankui, Bi, Xiao, Liu, Xiaodong, Wang, Xiaohan, Qu, Xiaojian, Chen, Xiaokang, Zhang, Xiaokang, Nie, Xiaotao, Zou, Xiaoyao, Li, Xiaoyuan, Guo, Xicheng, Chu, Xieting, Cheng, Xin, Liu, Xin, Xie, Xin, Xu, Xinbo, Liu, Xingchao, Liu, Xingchen, Yu, Xingkai, Li, Xingyou, Yao, Xintong, Chen, Xinyang, Jiang, Xinyong, Yang, Xinyu, Yang, Xinyu, Chen, Xu, Wang, Xuanyu, Zhong, Xubei, Su, Xuecheng, Liu, Xuejie, Lin, Xuheng, Fan, Xujie, Zhao, Xuncheng, Fu, Xuwei, Yan, Y. C., Jiang, Y. H., Wu, Y. T., M., Y. W., Wang, Y. Z., Gao, Yafei, Yang, Yang, Zhang, Yang, Ma, Yanru, Huang, Yanwen, Li, Yao, Li, Yao, Meng, Yao, Zhao, Yao, Sun, Yaofeng, Wang, Yaohui, Ye, Yaoyang, Yin, Yehang, Wu, Yexinrui, Qian, Yi, Tao, Yi, Yu, Yi, Zhang, Yichao, Jiang, Yichen, Wang, Yicheng, Ding, Yifan, Shi, Yifan, Peng, Yifeng, Zhai, Yifeng, Wu, Yijia, Xiong, Yiliang, Wang, Yilun, He, Ying, Zhou, Ying, Luo, Yingjia, Zhong, Yinmin, Wang, Yiping, Wang, Yisong, Zhang, Yixiang, Chen, Yixiao, Tan, Yixuan, Wei, Yixuan, Ma, Yiyang, Yang, Yiyao, Liu, Yiyuan, Cai, Yizai, Wei, Yizhen, Wang, Yizhi, Yang, Yonglun, Zhuo, Yongqi, Guo, Yongqiang, Wu, Yongtong, Wu, Yu, Zhang, Yu, Bian, Yuan, Cheng, Yuan, Ou, Yuan, Sun, Yuan, Xu, Yuanfan, Sun, Yuanhang, Li, Yuanhao, Liu, Yuchen, Yao, Yuchen, Han, Yudong, Wang, Yuduan, Wu, Yuhan, Meng, Yuhao, Zou, Yuheng, Li, YuKun, Wang, Yunchuan, Xiao, Yunfan, Xiong, Yunfan, Chen, Yupeng, Cao, Yuqian, Wang, Yuqian, Chen, Yuqing, Zhang, Yushun, Lin, Yutong, Xiao, Yuwei, Gu, Yuxian, Chen, Yuxiang, Huang, Yuxiang, Luo, Yuxiang, You, Yuxiang, Chen, Yuxin, Xiang, Yuxin, Liu, Yuxuan, Zhou, Yuxuan, Zhou, Yuyang, Guo, Yuzhe, Huang, Yuzhen, Bai, Yuzhuo, Z., Z. Y., Ni, Zanlin, Wang, Zehao, Zhao, Zehua, Ren, Zehui, Zhao, Zejun, Sha, Zhangli, Wang, Zhanying, Zhang, Zhaochen, Du, Zhaoshuai, Fu, Zhe, Xu, Zhean, Xie, Zhenda, Liu, Zheng, Zhang, Zhengyan, Dong, Zhenhua, Hao, Zhewen, Wang, Zhibang, Gou, Zhibin, Ma, Zhicheng, Li, Zhihao, Shao, Zhihong, Huang, Zhihuan, Li, Zhijie, Lu, Zhirui, Huang, Zhixian, Chen, Zhixuan, Chen, Zhixuan, Pan, Zhixuan, Wu, Zhiyu, Ren, Zhizhou, He, Zhu, Li, Zhuoshu, Zhang, Zhuping, Xu, Zian, Wang, Zihao, Gu, Zihui, Zhu, Zijia, Zhang, Zili, Li, Zilin, Hou, Zilong, Lyu, Zilong, Wang, Ziqiao, Xie, Ziwei, Zhang, Ziya, Gao, Ziyi, Pan, Zizheng, Li, Zonglin, Yao, Zongqing, Chen, Zui, Wu, Zuofan, Ling, Chenchen, Hou, Chengyu, Chen, Chong, Li, D., Qi, Di, Ji, Dongjie, Wei, Fang, Xia, Fanyi, Xie, Fei, Tan, Feiyi, Guo, Hailong, Zhai, Haiyan, Zhou, Hui, Tan, Huihui, Li, Huijie, Luo, Jia, Song, Jia, Cai, Jialu, Liang, Jian, Zhou, Jiangting, Gao, Jiaqi, Shao, Jiayi, Chen, Jie, Yang, Jieyu, Chen, Jin, Zhang, Jingde, Zhou, Jingzi, Wang, Jinqian, Liu, Jinyang, Sun, JinZhao, Ling, Junhua, Zheng, Junmin, Yang, Kaicheng, Xu, Ke, Su, Le, Xia, Leyi, Ding, Liangfeng, Zhuo, Lin, Ma, Linwang, Zhu, Linyan, Cai, Liyu, Yao, Luqi, Zhang, M. K., Li, Meng, Lin, Miao, Wang, Miaojun, Zhang, Min, Li, Mingming, Wang, Mingming, Yin, Mingze, Han, Minmin, Cao, Nan, Wang, Ning, Ma, Ningxin, Wang, Panpan, Lin, Peihan, Sun, Peng, Zhang, Peng, Ying, Qian, Xiang, Qiang, Wang, Qiao, Mao, Qingmiao, Jiang, Qiwei, Jin, Rongli, Chen, Ruyi, Tao, Sha, Sun, Shangmian, Wu, Shaoqing, Zou, Shichao, Lei, Si, Zhang, Tianyang, Sun, Tianyu, Yin, Tingting, Xiao, W. L., An, Wei, Li, Wei, Wang, Wei, Lin, Weiwei, Hou, Wenqing, Lin, X., Meng, Xiangfei, Huang, Xianzhu, Peng, Xiao, Li, Xiaoqian, Zhang, Xiaoting, Sun, Xiaowen, Wang, Xiaoxiang, Ye, Xiaoyu, Zhang, Xinrou, Zhang, Xinyu, Cao, Xue, Chen, Xueyin, Zhou, Yanan, Xu, Yanhong, Xia, Yao, Xu, Yao, Shao, Yi, Zhang, Yihong, Ma, Yiling, Tang, Ying, Lou, Yining, Chen, Yiru, Piao, Yishi, Chen, Yixuan, Xiong, Yong, Xuan, Yuchen, Yang, Yuehan, Xu, Yuer, Zha, Yukun, Ma, Yunxian, Lin, Yuping, Yan, Yuting, Xie, Yutong, Sheng, Yuwen, Zhu, Yuxuan, Zhang, Zekai, Ju, Zhe, Lin, Zhenzhen, Gao, Zheren, Sun, Zheyang, Yan, Zhigang, Wu, Zhongyu, Wang, Zi, Qu, Zihua, Yan, Ziling, Wan, Ziyi arXiv: arxiv.org/abs/2609.19969 Summary: To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. With its Causal Encoder-Decoder (CED) architecture, the model activates 16B parameters per token during decode but only 8B parameters during prefill, substantially improving cost efficiency for agentic workloads. Trending because: 51 HuggingFace upvotes + pairs a one-million-token multimodal MoE with aggressive KV-cache and prefill efficiency for agentic workloads
2. SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
Authors: Liu, Haozhe, Ye, Tian, Gao, Sensen, Cao, Qihang, Li, Yitong, Zhuge, Mingchen, Wang, Duomin, Zhang, Ruihua, Luo, Ping, Bian, Jiawang, Zhu, Lei, Zhu, Ligeng, Xie, Enze, Han, Song arXiv: arxiv.org/abs/2609.20519 Summary: We take an RSI-inspired approach at the harness layer, scaling auto-research loops across increasingly numerous and diverse environments for harness rollouts. At this scale, the process yields reusable improvements that transfer beyond their development setting, moving automated harness discovery toward production-level outcomes. Trending because: 40 HuggingFace upvotes + scales recursive auto-research loops across diverse environments to discover transferable harness improvements
3. An Empirical Study of Harness Design for Coding Agents
Authors: Fan, Run-Ze, Zhang, Zihao, Ma, Simin, Hu, Yebowen, Wang, Shouju, Song, Kaiqiang, Liu, Fei, Zamani, Hamed, Wang, Xiaoyang arXiv: arxiv.org/abs/2609.20804 Summary: Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management. Trending because: 33 HuggingFace upvotes + isolates how planning, action space, and context management affect coding-agent accuracy and cost
4. JEPA-Anything: Learning Predictive Models across Different Worlds
Authors: Cui, Taoyong, Wang, Zhongyao, Xu, Xinyue, Liu, Weiyang, Yu, Zhaochen, Zhang, Yuying, Gao, Qiang, Yang, Mengyue, Ouyang, Wanli, Heng, Pheng Ann, Wu, Yingcheng, Yin, Zhenfei, Yang, Ling arXiv: arxiv.org/abs/2609.20800 Summary: We introduce JEPA-Anything, a domain-agnostic framework based on orthogonal predictive factorization (OPF). Against matched JEPA baselines, JEPA-Anything improves reported metrics on all 10 dynamics tasks and reduces single-intervention prediction error on Interventional Pong by 34.8%. Trending because: 22 HuggingFace upvotes + tests a shared factorized predictive principle across seven radically different domains
5. Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model
Authors: Zhao, Haoyu, Zhao, Zihao, Deng, Tianyu, Xu, Ziqin, Zhang, Zihao, Wang, Xudong, Guo, Jinxiang, Gao, Chen, Ye, Ziyi, Jin, Yeying, Gu, Jiaxi, Wu, Zuxuan, Yan, Shuicheng arXiv: arxiv.org/abs/2609.18323 Summary: To investigate this question, this work introduces a comprehensive evaluation framework organized around four complementary dimensions of physical world reasoning. Across 517 evaluation instances, MiniMax-H3 achieves an overall success rate of 41.97%. Trending because: 21 HuggingFace upvotes + evaluates omni-modal physical-world reasoning with tasks that require complementary cross-modal evidence
6. RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation
Authors: Liu, ZhuoXin, Ma, Zhiming, Zhang, Ying, Yang, Mengzheng, Wang, Yifan, Huang, Zhengqi, Zhou, Yanhan, Lin, Zekun, Zhang, Jun, Zhang, Shun, Chen, Yue, Zhao, Qiao, Chen, Peng arXiv: arxiv.org/abs/2609.16900 Summary: We introduce RiskChainBench, pairing 3,600 synthetic token-text restoration inputs from 600 source sessions with 600 corresponding human-labeled local web environments. A model first restores the message, operational intent, and destination; the same underlying model then acts as a VLM-driven web agent that investigates the correctly associated website and produces a frozen, evidence-cited risk report without message-side semantics or domain-reputation cues. Trending because: 17 HuggingFace upvotes + connects obfuscated-message restoration to evidence-grounded investigation of risky web destinations
7. When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
Authors: Yang, Yuxiao, Yu, Tianrun, Li, Shangzhe, Zhao, Kaixiang, Zhang, Xuchao, Bansal, Chetan, Yao, Huaxiu, Killian, Taylor W., Zhang, Weitong arXiv: arxiv.org/abs/2609.20511 Summary: We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We show that aligning the decoding stopping set alone is insufficient, while treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigates mismatch-induced length inflation across all three model families. Trending because: 16 HuggingFace upvotes + traces runaway OPD generations to mismatched termination tokens across major model families
8. RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Authors: Yu, Yan, Lu, Zhengxi, Liu, Yizhou, Pan, Yichen, Wang, Aozhe, Chen, Qipeng, Yang, Hua, Zhang, Wenqi, Lu, Weiming, Chen, Qianglong, Shen, Yongliang arXiv: arxiv.org/abs/2609.20784 Summary: We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpasses its own skill-conditioned teacher in every setting. Trending because: 14 HuggingFace upvotes + lets an RL student retire its distillation teacher automatically after closing the performance gap
9. Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents
Authors: Chen, Bofan, Zhang, Boxuan, Tang, Fei, Lu, Zhengxi, Du, Yong, Chen, Tongbo, Lu, Weiming, Xiao, Jun, Zhuang, Yueting, Shen, Yongliang arXiv: arxiv.org/abs/2609.17653 Summary: We propose EvoSkill-GUI, a training-free framework in which each skill is a structured multi-file package containing retrieval metadata, executable plans, backup localization, failure-recovery rules, accessibility utilities, and failure cases. EvoSkill-GUI operates through a reflect-revise-reuse loop: the executor performs instant in-rollout revisions, an isolated critic diagnoses failed trajectories under strict information isolation, and the executor edits specific skill files through a restricted tool interface. Trending because: 10 HuggingFace upvotes + allows GUI-agent skills to evolve from deployment feedback without additional training
10. WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing
Authors: Yu, Hao, Liu, Kang, Zhao, Linnan, Zhan, Jiabo, Sun, Chong, Li, Chen, Lyu, Jing arXiv: arxiv.org/abs/2609.20423 Summary: We present WeVisDoc, a two-stage data-centric framework for robust end-to-end document parsing. Stage II uses a held-out probe to measure the Stage I parser’s residual errors within fixed visual-structural clusters. Trending because: 9 HuggingFace upvotes + uses residual-error probes to target document-parser weaknesses rather than merely adding coverage
11. VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control
Authors: Zhang, Zhongbo, Jin, Jiayi, Wang, Yifan, Zhang, Zaibin, Diao, Haiwen, Wang, Lijun, Lu, Huchuan arXiv: arxiv.org/abs/2609.19554 Summary: We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedural context from RGB-only demonstrations, actively select camera viewpoints, issue metric Cartesian commands, and revise them from execution feedback. Trending because: 7 HuggingFace upvotes + measures the full observe-reason-act-revise loop for embodied spatial intelligence
12. Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
Authors: Xi, Haocheng, Xie, Yiming, Zhao, Hexu, Zhang, Yiwen, Liu, Michael, Creavin, Thomas, Keutzer, Kurt, Li, Xiuyu, Lv, Zhaoyang, Xu, Chenfeng, Feng, Haiwen arXiv: arxiv.org/abs/2609.20744 Summary: We present Video DeltaNet (VDN), which combines local Softmax attention with bidirectional linear memory for long-range video context. Separate output projections and learnable gates calibrate the two branches, while a staged teacher-alignment recipe progressively introduces the new pathway into pretrained models. Trending because: 6 HuggingFace upvotes + combines local Softmax attention with bidirectional linear memory for efficient video generation
13. UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation
Authors: Zhang, Danning, Lin, Yijing, Zhuang, Shuhan, Huang, Mengqi, Wu, Shaojin, Fang, Shancheng, Mao, Zhendong arXiv: arxiv.org/abs/2609.12397 Summary: To address this challenge, we propose UFO, the first unified framework for omni-condition alignment simultaneous evaluation. Specifically, UFO introduces a novel Atomized Chain-of-Evaluation paradigm, i.e., it first decomposes omni-condition alignment into a sequential chain of fine-grained, disentangled Atomic Evaluation Units (AEUs), categorizes them into distinct modality-relevance classes, and then employs general or dedicated functional calls for accurate verification of different AEU types. Trending because: 4 HuggingFace upvotes + evaluates simultaneous multimodal condition alignment through an atomized chain of checks
14. FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations
Authors: Qu, Kevin, Sun, Tao, Viola, Massimiliano, Zhu, Liyuan, Zhou, Zhizhuo, Sarkar, Sayan Deb, Schindler, Konrad, Armeni, Iro arXiv: arxiv.org/abs/2609.20817 Summary: We present FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds. To aggregate articulation cues across observations, we introduce a Multi-state Articulation Transformer with alternating state-wise and global attention. Trending because: 4 HuggingFace upvotes + infers articulated parts and joints from sparse unordered partial point clouds
15. When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
Authors: Shim, Jaejun, Kim, HyunJin, Kim, Young Jin, Bak, JinYeong arXiv: arxiv.org/abs/2609.19671 Summary: We formulate efficient reasoning as an instance-adaptive computation allocation problem and propose When2Think, a post-training framework for hybrid reasoning that dynamically allocates computation based on problem difficulty. Our method introduces Instance-level Difficulty-Aware Control (IDAC), a reward-shaping mechanism that leverages pre-computed reference statistics (accuracy and token usage) to regulate reasoning depth. Trending because: 3 HuggingFace upvotes + allocates reasoning computation per instance instead of imposing uniform length controls
16. Region-Level Policy Optimization for Fine-grained MLLM Perception
Authors: Shi, Yuheng, Pei, Xiaohuan, Dong, Minjing, Xu, Chang arXiv: arxiv.org/abs/2609.19745 Summary: We therefore optimize the proposal network with region-level reinforcement learning, which we call Vision-RL2. Across six fine-grained benchmarks and four MLLM backbones, Vision-RL2 improves accuracy over the base model at every token budget and surpasses its largest-budget accuracy with about 4 times fewer visual tokens. Trending because: 3 HuggingFace upvotes + uses region-level reinforcement learning to preserve fine-grained perception with far fewer visual tokens
17. What Does Privileged Information Add to On-Policy Self-Distillation?
Authors: Zhang, XiuYu, Chow, Wei, Fang, Junfeng, Liang, Zhenkai, Chua, Tat-Seng arXiv: arxiv.org/abs/2609.20612 Summary: To isolate that contribution, we construct AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, and compare each view with matched reference-free distillation. With a thinking-enabled teacher supervising direct-response rollouts, reference-free distillation accounts for much of Qwen3-1.7B’s improvement under thinking-enabled evaluation, both in domain and on external benchmarks. Trending because: 3 HuggingFace upvotes + separates the benefit of privileged teacher information from self-distillation itself
18. PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
Authors: Okamoto, Mika, Erol, Ansel Kaplan arXiv: arxiv.org/abs/2609.18605 Summary: We introduce PACT (Pressure-Applied Compliance Testing), a benchmark for rule-following under pressure in AI agents assisting employees in daily tasks across twelve regulated enterprise domains and forty-eight scenarios, each set in a realistic multi-turn conversation. Each benchmark item pairs a standing rule against a rule-violating shortcut, and applies a battery of pressures across different wordings and system-prompt modes. Trending because: 2 HuggingFace upvotes + benchmarks enterprise-agent rule compliance under realistic pressure across regulated domains
19. Don’t Mask the Environment: Observation Supervision Changes How Agents Explore Under RL
Authors: Zhang, Juzheng, Makhija, Disha, Arivazhagan, Manoj Ghuhan, Kumar, Vinayshekhar Bannihatti, Gangadharaiah, Rashmi arXiv: arxiv.org/abs/2609.20715 Summary: We introduce ActObs, which also supervises the observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy to model action consequences without adding data, parameters, sequence tokens, or forward passes. Trending because: 2 HuggingFace upvotes + adds observation supervision so agent policies learn the consequences of their actions before RL
20. Self-Evolving Search Index
Authors: Lee, Sangam, Lee, Wonjae, Kim, Sunghwan, Kim, Deogyong, Kim, Jaehoon, Nam, Daye, Kang, SeongKu, Lee, Dongha arXiv: arxiv.org/abs/2609.19656 Summary: We propose SELF-INDEX, a framework that enables an index to self-evolve without human intervention. Its Optimizer autonomously diagnoses retrieval shortfalls, selectively revises the responsible index keys, and validates each revision before updating the index. Trending because: 1 HuggingFace upvotes + lets retrieval indexes diagnose failures, revise keys, and simulate new query demands autonomously
