Daily AI Papers — August 19, 2026
Published:
1. StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
Authors: Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang, Kai Wang arXiv: arxiv.org/abs/2608.15089 Summary: Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track of mutable state, fail to reactivate lessons from earlier executions, skip known procedures, or stop prematurely. Trending because: 284 HuggingFace upvotes + a timely benchmark drawing evaluation-focused attention
