Daily AI Papers — August 09, 2026
Published:
1. FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
Authors: Kapil Wanaskar, Gaytri Jena, Aman Chadha, Vinija Jain, Vasu Sharma, Amitava Das arXiv: arxiv.org/abs/2608.01049 Summary: World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) offer a particularly compelling direction. Trending because: 10 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
2. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
Authors: Gaytri Jena, Kapil Wanaskar, Vinija Jain, Aman Chadha, Vasu Sharma, Amitava Das arXiv: arxiv.org/abs/2608.01851 Summary: Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Trending because: 10 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
3. GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization
Authors: Baihan Yang, Tiexin Li, Yuheng Liu, Xin Lin, Xinke Li, Xiaohui Xie, Truong Nguyen arXiv: arxiv.org/abs/2608.01492 Summary: Selecting a complete 3D object from a reconstructed scene with minimal user effort is essential for practical scene editing and embodied interaction. Existing 3DGS-based methods either retrain the Gaussian representation to embed per-object labels, or build dense multi-view SAM observations, both requiring heavy computation and dense viewpoint coverage that is rarely available in practice. Trending because: 9 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
4. Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation
Authors: Tirth Bhatt, Naren Kumar S, Mayank Singh arXiv: arxiv.org/abs/2608.05785 Summary: Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional Flow Matching (TCFM), a multilingual embedding adaptation framework that selectively applies Flow Matching to translation tasks while optimizing retrieval, classification, and pair-classification tasks with objectives better aligned to their learning dynamics. Trending because: 9 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
5. ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
Authors: Mingxin Wang, Bin Hu, Bin Qian, Kaitao Jiang, Haoning Wu, Feng Yan, Bowen Jing, Ruiyang Hao, Enyi Wang, Kangning Niu, Yandan Yang, Mu Xu, Yan Wang, Houde Liu, Tianlun Li arXiv: arxiv.org/abs/2607.28993 Summary: World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions with task-irrelevant visual content, limiting robustness under visual distribution shifts. Trending because: 7 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
6. ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads
Authors: Şuayp Talha Kocabay, Talha Rüzgar Akkuş, Kamer Ali Yuksel arXiv: arxiv.org/abs/2608.02703 Summary: Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the final language-modeling head (LM-head) in BF16 or FP16. Quantizing this projection naively can strongly perturb the vocabulary-logit distribution. Trending because: 6 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
7. Self-Evolving Coding Agents
Authors: Hao Zhou, Haichuan Hu, Ye Shang, Quanjun Zhang arXiv: arxiv.org/abs/2608.03392 Summary: Large language models are increasingly embedded in software engineering workflows as coding agents that can inspect repositories, invoke tools, execute tests, debug failures, and generate patches. Yet most existing agents remain largely static after deployment, even though software development is a dynamic, feedback-rich process in which repositories evolve, dependencies change, tests fail, and repair attempts leave reusable experience. Trending because: 6 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
8. When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings
Authors: Christopher Schröder, Lukas Gienapp, Ferdinand Schlatt, Martin Potthast, Gerhard Heyer arXiv: arxiv.org/abs/2608.03994 Summary: We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows floating-point precision, which zeroes out a large fraction of attention weights and renders the affected attention heads partially blind. We analyze this failure mode, characterize its impact, and examine four mitigation strategies. Trending because: 6 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
9. Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming
Authors: Yanting Wang, Chenlong Yin, Runpeng Geng, Jinyuan Jia arXiv: arxiv.org/abs/2608.05108 Summary: Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Trending because: 6 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
10. TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex
Authors: Yuliang Yan, Shuo Yan, Haochun Tang, Yiqin Sun, Enyan Dai arXiv: arxiv.org/abs/2607.22143 Summary: Molecular glue degraders have emerged as a promising strategy for targeted protein degradation by inducing ternary complex formation between an E3 ubiquitin ligase and a target protein. Despite their therapeutic potential, computational design of molecular glues remains largely unexplored. Trending because: 5 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
11. Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
Authors: Jiazhen Liu, Mingkuan Feng, Long Chen arXiv: arxiv.org/abs/2608.02791 Summary: MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability, and fast inference. Embedding-prediction methods may disrupt language modeling through pixel-level objectives, whereas next-token generation is inefficient for dense masks. Trending because: 4 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
12. CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
Authors: Xuehang Guo, Pingyue Zhang, Ruiyi Zhang, Zhenhailong Wang, Hanrui Lyu, Heng Ji, Tong Sun, Qingyun Wang, Manling Li arXiv: arxiv.org/abs/2608.02833 Summary: Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains. While extrinsic chain-of-thought prompting and visual cues significantly improve performance, current MLLMs lack intrinsic visual grounded reasoning capabilities, leading to inaccurate perception and reasoning disconnected from visual evidence. Trending because: 4 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
13. ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
Authors: Gagan Bhatia, Julian Schlenker, Simone Paolo Ponzetto, Steffen Eger arXiv: arxiv.org/abs/2608.03507 Summary: Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve together across languages. We address this problem by asking how the magnitude and direction of change vary across linguistic levels, languages, and historical periods within a single analytical space. Trending because: 4 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
14. LegalPincite: Multi-level Legal Information Retrieval Dataset
Authors: Theresia Veronika Rampisela, Henrik Palmer Olsen, Giovanni Colavizza arXiv: arxiv.org/abs/2608.03756 Summary: A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While legal practice often requires pinpoint citations (pincites) to specific case paragraphs, most existing public legal IR datasets lack paragraph-level citation annotations. Trending because: 4 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
15. ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
Authors: Jinhe Bi, Chennan Zhou, Zengjie Jin, Aniri, Shuo Lu, Wenke Huang, Hu Cao, Xun Xiao, Zhihong Zhu, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua arXiv: arxiv.org/abs/2608.03972 Summary: On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-guided methods lose their main source of supervision, and these failed trajectories are typically discarded as negative samples. Trending because: 4 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
16. What AI Red-Team Evaluations Can and Cannot Prove
Authors: Bandana Kaur arXiv: arxiv.org/abs/2607.21735 Summary: Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. Trending because: 3 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
17. Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents
Authors: Shuai Wang, Haodong Chen, Yu Yin, Shengyao Zhuang, Bevan Koopman, Guido Zuccon arXiv: arxiv.org/abs/2608.02751 Summary: Existing deep-research agents use a search-visit workflow that retrieves and reads whole pages, without considering the addressable structure that web sources expose through titles, headings, sections, and metadata. This prevents agents from directly constraining retrieval to document fields and often carries irrelevant page content into their context. Trending because: 3 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
18. DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack
Authors: Hoseong Tae, Jong-Seok Lee arXiv: arxiv.org/abs/2608.03207 Summary: Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. We show that this robustness is largely illusory: it stems from prior attacks ignoring the multi-step denoising ODE. Trending because: 3 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
19. Multi-Task Multi-Frame Visual Piano Transcription
Authors: Yonghyun Kim, Hoyeol Sohn, Juhan Nam, Alexander Lerch arXiv: arxiv.org/abs/2608.03419 Summary: Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release. Yet existing Visual Piano Transcription (VPT) systems focus on onset detection from short video windows, offset accuracy lags onset by a wide margin, and note-level velocity has not been reported. Trending because: 3 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
20. When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs
Authors: Omatharv Bharat Vaidya, Connor Thomas Jerzak, Zayne Rea Sprague, Fangcong Yin, Nhat Ho arXiv: arxiv.org/abs/2608.03506 Summary: Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a valid minority trace. We introduce CALVER (Causal Axiom-Level VERification), a training-free symbolic verifier that scores structured traces against Pearl’s causal criteria, including -separation, backdoor adjustment, and intervention, and selects the highest-scoring candidate without consulting a reference answer. Trending because: 3 HuggingFace upvotes; among the most-upvoted fresh papers in the current feed.
