Blog Post number 1
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

AI / ML Engineer @ PyTorch-Meta
less than 1 minute read
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
less than 1 minute read
Published:
This post will show up by default. To disable scheduling of future posts, edit config.yml and set future: false.
10 minute read
Published:
Authors: Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo arXiv: arxiv.org/abs/2609.33757 Summary: YuE2 uses a single autoregressive/non-autoregressive Mixture-of-Transformers to plan a readable score, expand it into semantic music tokens, and render full-song audio. It beats evaluated public baselines on WildSongBench, approaches proprietary systems in expert listening, and supports score-guided edits and zero-shot covers from the same checkpoint.
11 minute read
Published:
Authors: Hongyang Du, Yunfei Xie, Junjie Ye, Jiawei Yang, Xiaoyan Cong, Haodong Zhang, Yongchao Huang, Haiyu Wu, Zongxia Li, Shihang Gui, Dawei Liu, Runhao Li, Jingcheng Ni, Chen Wei, Randall Balestriero, Yue Wang arXiv: arxiv.org/abs/2609.31620 Summary: FuseReg trains representation autoencoders over random subsets of visual-encoder layers, penalizing sensitivity to cross-layer disagreement and allowing one decoder to reconstruct from full, sparse, or single-layer fusions. On ImageNet-256, it improves reconstruction flexibility and cuts unguided gFID by 27% with decoder replacement alone and 29% when regularizing both decoder and diffusion training.
9 minute read
Published:
Authors: Zhongwen Xu, Zihan Ding arXiv: arxiv.org/abs/2509.13232 Summary: Single-stream Policy Optimization replaces group-based baselines with a persistent KL-adaptive value tracker and globally normalized advantages, avoiding degenerate groups and synchronization barriers in LLM reinforcement learning. On five hard mathematics benchmarks with Qwen3-8B, it improves average maj@32 by 3.4 percentage points over GRPO while converging more smoothly and wasting less computation.