Yuxuan Zhang

Yuxuan Zhang

Yuxuan Zhang
I build post-training methods, environments, and benchmarks toward recursive self-improvement.
PhD @ UBC · Vector Institute
Open to internship opportunities in AI research. Email me

Hi there! I'm Yuxuan Zhang, a PhD student at the University of British Columbia, advised by Prof. Kelsey R. Allen (Senior Research Scientist, DeepMind). I'm also affiliated with the Vector Institute. I work closely with Prof. Wenhu Chen, Dongfu Jiang, Yubo Wang, and Ping Nie. My research interests span:

Agentic AI Large Language Models Reinforcement Learning Recursive Self-Improvement

I make AI agents work in the real world: ClawBench for live web-agent evaluation (adopted by ByteDance Seed, EMNLP 2026 Findings), RewardHarness for self-evolving agentic post-training (COLM 2026), and VidGround for grounded video understanding.


Let's collaborate

I'm open to collaborations and research chats. Reach out at reacher.zh [at] gmail.com ( WeChat works too).
Actively looking for research internships. Let's grab a coffee and have a chat!

Research Highlights

Cover image for ClawBench: Can AI Agents Complete Everyday Online Tasks?
EMNLP 2026 Findings

ClawBench: Can AI Agents Complete Everyday Online Tasks?

ClawBench evaluates AI agents on 153 real-world tasks across 144 live platforms, from booking appointments to filing job applications. It runs on production websites and intercepts only the final submission, keeping evaluation safe without losing real-world complexity.

Cover image for VidGround: Watch Before You Answer
Under Review

VidGround: Watch Before You Answer

Language models can answer video questions from text priors alone, without watching the video. We filter out such spurious training samples, producing a cleaner dataset that teaches models to ground answers in what they see.

Cover image for RewardHarness: Learning Human Preferences for Image Editing with Only 100 Demonstrations
COLM 2026

RewardHarness: Learning Human Preferences for Image Editing with Only 100 Demonstrations

RewardHarness beats GPT-5 by 5.3 points on image-editing evaluation while using 0.05% of the EditReward data. It reframes reward modeling as context evolution: from 100 preference demonstrations it evolves a library of tools and skills.

  • University of British Columbia
    PhD in Computer Science
    Jan. 2025 - Dec. 2028
  • University of Toronto
    MSc in Applied Computing
    Sept. 2022 - June 2024
  • Peking University
    Bachelor in Economics
    Sept. 2019 - June 2022
  • SOTI
    Data Scientist
    Jan. 2024 - Dec. 2024
    Research Intern
    May 2023 - Dec. 2023
  • 2026 Tinker Research Grant ($5,000), Thinking Machines Lab
  • 2026 Compute Transparency Champion Award, CVPR 2026
  • 2025 President's Academic Excellence Award
  • 2023 Mitacs Accelerate Fellowship ($30,000)

News

2026
Self-Developing Agents: Aspire, S3Gym, and HarnessDev study self-improvement through goals, experience, and agent harnesses.
Five papers accepted at EMNLP 2026: WebWorld, OpenSkill, and VGI-BENCH; ClawBench in Findings; and Dr. Claw in System Demonstrations.
We are organizing ARRSI 2026, the first workshop on Autonomous Research and Recursive Self-Improvement, co-located with AACL-IJCNLP 2026. Call for papers is open!
ClawBench has been accepted at the COLM 2026 Workshop on Agent Behavior. See the paper.
ClawBench has been adopted by ByteDance Seed (Doubao).
Function-Aware Fill-in-the-Middle (FIM) mid-training for coding agent foundation models is out. See the paper, data, and models.
earlier news (8)
2026
OpenSkill public preprint is available. See the paper and code.
RewardHarness has been accepted to COLM 2026. See the paper and code.
ClawBench public preprint is available. See the paper, code, and data.
Watch Before You Answer public preprint is available. Read the paper.
2025
ScholarCopilot has been accepted at COLM 2025! Play with the live demo.
Retri3D has been accepted at ICLR 2025 as a Spotlight 🌟!
WikiGap is released! Try the Chrome extension.
StructEval is released! Check out the dataset.

Selected Publications (view all )

2026 15 papers
Cover image for RewardHarness: Learning Human Preferences for Image Editing with Only 100 Demonstrations

RewardHarness: Learning Human Preferences for Image Editing with Only 100 Demonstrations

Yuxuan Zhang, Cong Wei, Penghui Du, Bo Li, Huaisong Zhang, Songcheng Cai, Yubo Wang, Dongfu Jiang, Yuyu Zhang, Changqian Yu, Ping Nie, Wenhu Chen, Kelsey R Allen

COLM 2026

Cover image for ClawBench: Can AI Agents Complete Everyday Online Tasks?

ClawBench: Can AI Agents Complete Everyday Online Tasks?

Yuxuan Zhang, Yubo Wang, Yipeng Zhu, Penghui Du, Junwen Miao, Xuan Lu, Zhuofeng Li, Xingwei Qu, Zhengkang Guo, Yuanzhe Shen, Dingjie Song, Han Zhou, Tuney Zheng, Xian Wu, Hao Yu, Songcheng Cai, Yi Lu, Yunzhuo Hao, Minyi Lei, Liang Chen, Kai Zou, Huifeng Yin, Wendong Xu, Dongfu Jiang, Ping Nie, Jiaheng Liu, Wenhu Chen, Kelsey R. Allen

EMNLP 2026 Findings

Adopted by ByteDance Seed

Cover image for Watch Before You Answer: Learning from Visually Grounded Post-Training

Watch Before You Answer: Learning from Visually Grounded Post-Training

Yuxuan Zhang, EunJeong Hwang, Huaisong Zhang, Penghui Du, Yiming Jia, Dongfu Jiang, Xuan He, Shenhui Zhang, Ping Nie, Peter West, Kelsey R. Allen

Technical Report

Cover image for HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Xinping Lei, Qingshui Gu, Yuxuan Zhang, Zexuan Wang, Chen He, Chen Huang, Maojia Song, Zhiyuan Zeng, Shaowen Wang, Jinkai Liu, Yunfeng Shi, Jiaheng Liu, Shen Yan, Wenhao Huang, Ge Zhang, Wenxuan Zhang

Technical Report

Cover image for Dr. Claw: An AI Scientist Workspace for Vibe Research

Dr. Claw: An AI Scientist Workspace for Vibe Research

Dingjie Song, Hanrong Zhang, Dawei Liu, Yixin Liu, Zongxia Li, Zhengqing Yuan, Siqi Zhang, Henry Peng Zou, Zhiling Yan, Yuxuan Zhang, Yanfang Ye, Philip S. Yu, Lichao Sun

EMNLP 2026 System Demonstrations

Cover image for Aspire: Can Models Self-Evolve from Vague Goals?

Aspire: Can Models Self-Evolve from Vague Goals?

Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Yuxuan Zhang, Xinping Lei, Junting Zhou, Zexuan Wang, Yuchen Wu, Huan Zhou, Duo Wang, Yinzhu Piao, Yongchang Peng, Yunfeng Shi, Jin Chen, Zuo Wang, Jinkai Liu, Jiaheng Liu, Wenxuan Zhang, Shen Yan, Wenhao Huang, Ge Zhang

Technical Report

Cover image for S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

Jiajun Shi, Siyuan Tao, Yuhao Wu, Zexuan Wang, Jingyuan Zhang, Jiaheng Liu, Xinping Lei, Xinrong Zhang, Siyuan Fang, Zhewen Tan, Tianle Cai, Junhao Fang, Jiameng Huang, Yueyang Wang, Jinkai Liu, Yuxuan Zhang, Jian Yang, Zhoujun Li, Shen Yan, Wenhao Huang, Ge Zhang

Technical Report

Cover image for WebWorld: The Browser as a World Model for Self-Improving Web Code

WebWorld: The Browser as a World Model for Self-Improving Web Code

Jiajun Wu, Jian Yang, Yaxin Du, Wei Zhang, Haowen Wang, Junhang Cheng, Yuxuan Zhang, Tuney Zheng, Xianglong Liu, Ming Zhou

EMNLP 2026

Cover image for ModularRSI: Toward Generalizable Harness RSI

ModularRSI: Toward Generalizable Harness RSI

Siwei Wu*, Jincheng Ren*, Yizhi Li*, Haau-Sing Li, Chengran Yang, Weicheng Gu, Yuxuan Zhang, Jian Yang, Riza Batista-Navarro, Chuanyi Zhang, Xianglong Liu, Ming Zhou, Bryan Dai, Chenghua Lin (* equal contribution)

Technical Report

Cover image for VGI-BENCH: Probing Visual Intelligence in Video Generation Models

VGI-BENCH: Probing Visual Intelligence in Video Generation Models

Xuan He§, Cong Wei, Yuhao Cheng, Linrui Ma, Yuxuan Zhang, Zuojun Li, Yuhao Wen, Jize Jiang, Zeyi Liu, Yuren Hao, Songcheng Cai, Keming Wu, Penghui Du, Kai Zou, Rui Yang, Chenkai Sun, Ke Yang, Ping Nie, Kelsey R Allen, Chenglong Wang, Michel Galley, Jianfeng Gao, ChengXiang Zhai ( Main Contributor§ Project Lead.)

EMNLP 2026

Cover image for MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning

MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning

Yingying Fan, Penghui Du, Leyan Zhu, Runze He, Zimeng Wu, Yuxuan Zhang, Liang Chen, Jiahao Xie, Jiangtang Wang, Shuai Shao, Anchao Yang, Yutong Bai, Yan Wang

Technical Report

Cover image for Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

Yubo Wang, Jiarong Liang, Yuxuan Zhang, Xuye Liu, Cong Wei, Yuyu Zhang, Ping Nie, Wenhu Chen

Technical Report

Cover image for Learning from the Self-future: On-policy Self-distillation for dLLMs

Learning from the Self-future: On-policy Self-distillation for dLLMs

Yifu Luo, Zeyu Chen, Haoyu Wang, Xinhao Hu, Yuxuan Zhang, Zhizhou Sha, Shiwei Liu

Technical Report

Cover image for OpenSkill: Open-World Self-Evolution for LLM Agents

OpenSkill: Open-World Self-Evolution for LLM Agents

Zhiling Yan, Dingjie Song, Hanrong Zhang, Wei Liang, Yuxuan Zhang, Yutong Dai, Lifang He, Philip S. Yu, Ran Xu, Xiang Li, Lichao Sun

EMNLP 2026

Cover image for MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection

MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection

Haowen Wang, Yaxin Du, Jian Yang, Jiajun Wu, Shukai Liu, Yuxuan Zhang, Pingjie Wang, Siheng Chen, Tuney Zheng, Ming Zhou, Xianglong Liu, Bryan Dai

Technical Report

2025 5 papers
Cover image for Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports

Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports

Yang Yao, Yixu Wang, Yuxuan Zhang, Yi Lu, Tianle Gu, Lingyu Li, Dingyi Zhao, Keming Wu, Haozhe Wang, Ping Nie, Yan Teng, Yingchun Wang

Technical Report

Cover image for ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations

ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations

Yubo Wang, Xueguang Ma, Ping Nie, Huaye Zeng, Zhiheng Lyu, Yuxuan Zhang, Benjamin Schneider, Yi Lu, Xiang Yue, Wenhu Chen

COLM 2025

Cover image for Retri3D: 3D Neural Graphics Representation Retrieval

Retri3D: 3D Neural Graphics Representation Retrieval

Yushi Guan, Daniel Kwan, Jean Sebastien Dandurand, Xi Yan, Ruofan Liang, Yuxuan Zhang, Nilesh Jain, Nilesh Ahuja, Selvakumar Panneer, Nandita Vijaykumar

ICLR 2025 Spotlight

Cover image for VideoScore2: Think before You Score in Generative Video Evaluation

VideoScore2: Think before You Score in Generative Video Evaluation

Xuan He, Dongfu Jiang, Ping Nie, Minghao Liu, Zhengxuan Jiang, Mingyi Su, Wentao Ma, Junru Lin, Chun Ye, Yi Lu, Keming Wu, Benjamin Schneider, Quy Duc Do, Zhuofeng Li, Yiming Jia, Yuxuan Zhang, Guo Cheng, Haozhe Wang, Wangchunshu Zhou, Qunshu Lin, Yuanxing Zhang, Ge Zhang, Wenhao Huang, Wenhu Chen

TMLR 2026

Cover image for StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs

StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs

Jialin Yang, Dongfu Jiang, Lipeng He, Sherman Siu, Yuxuan Zhang, Disen Liao, Zhuofeng Li, Huaye Zeng, Yiming Jia, Haozhe Wang, Benjamin Schneider, Chi Ruan, Wentao Ma, Zhiheng Lyu, Yifei Wang, Yi Lu, Quy Duc Do, Ziyan Jiang, Ping Nie, Wenhu Chen

TMLR 2026

Impact

ClawBench
AdoptedClawBench is used as a headline agent benchmark in frontier model technical reports. Li Auto's Mach-Mind-4-Flash reports its ClawBench score in the abstract and benchmarks eight model families on it, including Kimi K2.5, Qwen3.5, GLM-4.7 and Nemotron-3. 2026
CitedClawBench has been cited by agent-benchmark work from Shanghai AI Laboratory (WildClawBench, π-Bench), Carnegie Mellon (MyPCBench), HKU MMLab and Meituan (UniClawBench), and Peking University (Harness-Bench). 2026
LeaderboardClawBench is tracked as a standing leaderboard by the browser-agent platform Steel, alongside WebVoyager and BrowseComp. 2026
VideoScore2
IntegratedVideoScore2 ships as a built-in evaluation metric in FastVideo, UC San Diego Hao AI Lab's video generation framework. 2026
IntegratedVideoScore2 is used as a reward model by Google's VQQA video-evaluation framework, which runs it as a Best-of-N selector across its main results tables. 2026
StructEval
IntegratedStructEval is integrated as a scored benchmark in llm-jp-eval, the evaluation suite of Japan's national LLM project (LLM-jp / NII). 2026
CitedStructEval has been cited by structured-generation work from Apple, Microsoft Research and Amazon Science. 2025 - 2026
ScholarCopilot
CitedScholarCopilot is surveyed in "Transforming Science with Large Language Models," a widely cited review of AI-assisted scientific discovery. 2025 - 2026

Talks

Academic Service

Organizer
AACL 2026 Workshop on Autonomous Research and Recursive Self-Improvement (ARRSI)
Reviewer
ACL ARR 2026 ICMLW 2026 ICLRW 2026 NeurIPS 2026 COLM 2026 CVPR 2026 ECCV 2026

Advisors & Collaborators