Daily AI Papers — August 14, 2026

12 minute read

Published:

1. Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

Authors: Yuanyang Yin, Gongxuan Wang, Yifan Zhan, Chuanhao Li, Kaipeng Zhang, Feng Zhao arXiv: arxiv.org/abs/2608.13546 Summary: Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory, while low-latency interaction relies on few-step generation whose capabilities are bounded by its teacher. Trending because: 81 HuggingFace upvotes + one of the most-upvoted papers in today’s feed.


2. DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

Authors: DreamX Team, Rui Chen, Xiangxiang Chu, Geng Li, Jifan Li, Qingfeng Shi, Datao Tang, Jing Tang, Jun Wang, Pengfei Zhang arXiv: arxiv.org/abs/2608.13489 Summary: We present DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. Trending because: 79 HuggingFace upvotes + one of the most-upvoted papers in today’s feed.


3. LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

Authors: Tao Feng, Fangxu Yu, Haozhen Zhang, Zhongjie Dai, Liangqi Yuan, Zijie Lei, Weizhi Zhang, Kunlun Zhu, Haodong Yue, Keyang Xuan, Ge Liu, Jiaxuan You arXiv: arxiv.org/abs/2608.06867 Summary: No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair comparison and extension difficult. Trending because: 76 HuggingFace upvotes + one of the most-upvoted papers in today’s feed.


4. DarwinX: Evolving Agent Harnesses Through Natural Selection

Authors: Yifan Zhang, Yutong Dai, Juntao Tan, Luyu Yang, Rishi Mullur, Thai Hoang, Zhiyuan Hu, James Zhu, Phil Mui, Silvio Savarese, Ran Xu, Zeyuan Chen arXiv: arxiv.org/abs/2608.07545 Summary: An LLM agent’s capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. Trending because: 52 HuggingFace upvotes + strong engagement in today’s HuggingFace feed.


5. Intern-S2-Preview: Scientific Agentic Foundation Model

Authors: Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng, Erfei Cui, Xuanlang Dai, Shengyuan Ding, Shangheng Du, Yanhui Duan, Yue Fan, Youqing Fang, Quan Gan, Yuanyuan Gao, Jiaye Ge, Lixin Gu, Yuzhe Gu, Qipeng Guo, Junjun He, Xin Hong, Ming Hu, Zhouqi Hua, Haian Huang, Junhao Huang, Zixian Huang, Minxi Jin, Lingkai Kong, Alexander Lam, Zehao Li, Zonglin Li, Tianhao Liang, Dahua Lin, Junyao Lin, Tianyang Lin, Zhouhan Lin, Jiangning Liu, Jin Liu, Kuikun Liu, Wenran Liu, Yifei Liu, Yuhong Liu, Yuhong Liu, Zhoumianze Liu, Ziyan Liu, Ziyu Liu, Haijun Lv, Han Lv, Chengqi Lyu, Le Ma, Ningsheng Ma, Zerun Ma, Haoyang Peng, Runyu Peng, Jifei Shan, Zixin Shang, Kou Shi, Xiang Shi, Qisheng Su, Xuerui Su, Hao Sun, Xiao Sun, Yanan Sun, Yu Sun, Huanze Tang, Yinghao Tang, Wenhui Tian, Zhongbo Tian, Bingli Wang, Haomin Wang, Jiarui Wang, Jingzhi Wang, Rui Wang, Xiquan Wang, Yi Wang, Zhecan Wang, Ziyi Wang, Zun Wang, Rubin Wei, Lianyi Wu, Wen Wu, Yue Wu, Yuhan Wu, Zhenyu Wu, Zijian Wu, Shuhao Xing, Jun Xu, Xingle Xu, Xuenan Xu, Xiangchao Yan, Ziang Yan, Bowen Yang, Danni Yang, Lin Yang, Zhiqi Yang, Qian Yao, Haochen Ye, Peng Ye, Jinhui Yin, Jiashuo Yu, Dingbo Yuan, Fei Yuan, Yuhang Zang, Bo Zhang, Chao Zhang, Chen Zhang, Hongjie Zhang, Junming Zhang, Wenlong Zhang, Wenwei Zhang, Yiming Zhang, Zhuo Zhang, Ziyang Zhang, Haiteng Zhao, Penghao Zhao, Yibo Zhao, Zhonghan Zhao, Zhihang Zhong, Bowen Zhou, Peiheng Zhou, Xin Zhou, Xinyu Zhou, Yunhua Zhou, Dongsheng Zhu, Yicheng Zou arXiv: arxiv.org/abs/2608.13505 Summary: Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. Trending because: 41 HuggingFace upvotes + strong engagement in today’s HuggingFace feed.


6. How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

Authors: Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou arXiv: arxiv.org/abs/2608.08975 Summary: As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Trending because: 37 HuggingFace upvotes + strong engagement in today’s HuggingFace feed.


7. AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

Authors: Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li, Zhengrong Yue, Jing Li, Xiaofu Chen, Xiaohan Zhao, Jiacheng Liu, Jiacheng Cui, Zhiqiang Shen, Xiaotong Li arXiv: arxiv.org/abs/2608.13560 Summary: Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to drive recursive self-improvement, existing paradigms remain static and fall short of this capability. Trending because: 31 HuggingFace upvotes + strong engagement in today’s HuggingFace feed.


8. PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

Authors: Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao arXiv: arxiv.org/abs/2608.13552 Summary: Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. Trending because: 31 HuggingFace upvotes + strong engagement in today’s HuggingFace feed.


9. Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence

Authors: Haokai Zhang, Yuhang Ding, Yunshu Zhou, Xinze Du, Shengtao Zhang, Zhiyue Zhao, Yuling Xi, Hao Chen arXiv: arxiv.org/abs/2608.12743 Summary: Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. Trending because: 26 HuggingFace upvotes + strong engagement in today’s HuggingFace feed.


10. Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

Authors: Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Jing Xiong, Hengyuan Zhang, Zhongzhu Zhou, Tiantian Zhang, Ngai Wong, Chuan-Wei Kuo arXiv: arxiv.org/abs/2608.12149 Summary: We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attention spikes (PAS), and can persist through intervening linear attention layers, giving rise to inter-spike plateaus (ISP). As full attention becomes denser, successive PAS become increasingly connected through ISP, ultimately recovering the stable MA morphology of full attention LLMs. Trending because: 16 HuggingFace upvotes + solid engagement in today’s HuggingFace feed.


11. AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

Authors: Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri arXiv: arxiv.org/abs/2608.11216 Summary: World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous researchers–a setting in which the improvement direction is not specified in advance, unlike the engineering-to-spec tasks that dominate current agent benchmarks. Trending because: 11 HuggingFace upvotes + solid engagement in today’s HuggingFace feed.


12. UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

Authors: Yuxuan Zhang, Haozhong Xiong, Jiayi Song, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang, Liwei Wang arXiv: arxiv.org/abs/2608.11752 Summary: Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. Trending because: 11 HuggingFace upvotes + solid engagement in today’s HuggingFace feed.


13. Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning

Authors: Haodong Li, Shaoteng Liu, Tianyu Wang, Chongjian Ge, Sihui Ji, Jiahan Zhang, Xin Lin, Haolin Lu, Zhe Lin, Manmohan Chandraker arXiv: arxiv.org/abs/2608.09926 Summary: The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how the pixels transit over time. Trending because: 10 HuggingFace upvotes + solid engagement in today’s HuggingFace feed.


14. LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time

Authors: Yuxuan Zhang, Haozhong Xiong, Yubo Huang, Jiayi Song, Jinpeng Yu, Haofan Wang, Jiaming Liu, Ruihua Huang, Liwei Wang arXiv: arxiv.org/abs/2608.11745 Summary: Pose-driven human animation synthesizes a video of a target person from a single reference image and a driving pose stream. Real-time generation is essential for interactive applications such as live streaming, telepresence, and virtual avatars, yet diffusion-based systems require minutes to hours per clip, precluding responsive interaction. Trending because: 10 HuggingFace upvotes + solid engagement in today’s HuggingFace feed.


15. From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection

Authors: Zepeng Wang, Jiagao Hu, Fuhao Li, Yuxuan Chen, Fei Wang, Daiguo Zhou arXiv: arxiv.org/abs/2608.11562 Summary: Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, video reflection removal remains largely underexplored due to the lack of paired video data, temporally coherent removal models, and dedicated evaluation benchmarks. Trending because: 9 HuggingFace upvotes + solid engagement in today’s HuggingFace feed.


16. CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

Authors: Gyuwan Kim, Cheoneum Park, Tao Yang arXiv: arxiv.org/abs/2608.07458 Summary: Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the coarse-grained chunks. This paper optimizes the Pareto frontier under low prefill latency constraints while maximizing accuracy by proposing CoinRAG (Contextualized Information Nugget KV Cache Reuse for Long-Context RAG). Trending because: 8 HuggingFace upvotes + solid engagement in today’s HuggingFace feed.


17. ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

Authors: Yutao Mou, Pengfei Yang, Zhe Yin, Zhangchi Xue, Xiaotian Luan, Dingyao Yu, Tong Zhang, Shikun Zhang, Wei Ye arXiv: arxiv.org/abs/2608.11878 Summary: Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. Trending because: 8 HuggingFace upvotes + solid engagement in today’s HuggingFace feed.


18. JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Authors: Shawn Li, Wei Yang, Jike Zhong, Jiate Li, Jiawei Yang, You Qin, Ryan Rossi, Franck Dernoncourt, Roger Zimmermann, Yue Wang, Zhengzhong Tu, Vicente Ordonez, Mohit Bansal, Yue Zhao arXiv: arxiv.org/abs/2607.27670 Summary: Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \ours{}, a benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth. Trending because: 7 HuggingFace upvotes + fresh entry gaining traction in today’s HuggingFace feed.


19. Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

Authors: Burc Gokden arXiv: arxiv.org/abs/2608.10288 Summary: The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator G_{LM}, built from a positive tensor A_{LM} by elementwise power laws. The architecture is fully specified, verified against pinned reference releases; claims are labeled theorem, conditional theorem, measurement, or conjecture. Trending because: 7 HuggingFace upvotes + fresh entry gaining traction in today’s HuggingFace feed.


20. The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images

Authors: Zhiheng Wang, Bo Peng, Lai Wei, Chaochao Lu arXiv: arxiv.org/abs/2608.06270 Summary: The “thinking-with-images” paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. Trending because: 7 HuggingFace upvotes + fresh entry gaining traction in today’s HuggingFace feed.