Yuxuan Zhang

Yuxuan Zhang

Yuxuan Zhang
I build post-training methods, environments, and benchmarks toward recursive self-improvement.
PhD @ UBC · Vector Institute
Open to internship opportunities in AI research. Email me

Hi there! I'm Yuxuan Zhang, a PhD student at the University of British Columbia, advised by Prof. Kelsey R. Allen (Senior Research Scientist, DeepMind). I'm also affiliated with the Vector Institute. I work closely with Prof. Wenhu Chen, Dongfu Jiang, Yubo Wang, and Ping Nie. My research interests span:

Agentic AI Large Language Models Reinforcement Learning Recursive Self-Improvement

I make AI agents work in the real world: ClawBench for live web-agent evaluation (adopted by ByteDance Seed · Doubao; COLM 2026 WAB), RewardHarness for self-evolving agentic post-training (COLM 2026), and VidGround for grounded video understanding.


Let's collaborate

I'm open to collaborations and research chats. Reach out at reacher [at] cs.ubc.ca ( WeChat works too).
Actively looking for research internships. Let's grab a coffee and have a chat!

Research Highlights

Cover image for ClawBench: Can AI Agents Complete Everyday Online Tasks?
COLM 2026 WAB

ClawBench: Can AI Agents Complete Everyday Online Tasks?

ClawBench evaluates AI agents on 153 real-world tasks across 144 live platforms, from booking appointments to filing job applications. It runs on production websites and intercepts only the final submission, keeping evaluation safe without losing real-world complexity.

Cover image for VidGround: Watch Before You Answer
Under Review

VidGround: Watch Before You Answer

Language models can answer video questions from text priors alone, without watching the video. We filter out such spurious training samples, producing a cleaner dataset that teaches models to ground answers in what they see.

Cover image for RewardHarness: Self-Evolving Agentic Post-Training
COLM 2026

RewardHarness: Self-Evolving Agentic Post-Training

RewardHarness beats GPT-5 by 5.3 points on image-editing evaluation while using 0.05% of the EditReward data. It reframes reward modeling as context evolution: from 100 preference demonstrations it evolves a library of tools and skills.

  • University of British Columbia
    PhD in Computer Science
    Jan. 2025 - Dec. 2028
  • University of Toronto
    MSc in Applied Computing
    Sept. 2022 - June 2024
  • Peking University
    Bachelor in Economics
    Sept. 2019 - June 2022
  • SOTI
    Data Scientist
    Jan. 2024 - Dec. 2024
    Research Intern
    May 2023 - Dec. 2023
  • 2026 Tinker Research Grant ($5,000), Thinking Machines Lab
  • 2026 Compute Transparency Champion Award, CVPR 2026
  • 2025 President's Academic Excellence Award
  • 2023 Mitacs Accelerate Fellowship ($30,000)

News

2026
We are organizing ARRSI 2026, the first workshop on Autonomous Research and Recursive Self-Improvement, co-located with AACL-IJCNLP 2026. Call for papers is open!
ClawBench has been accepted at the COLM 2026 Workshop on Agent Behavior. See the paper.
ClawBench has been adopted by ByteDance Seed (Doubao).
Function-Aware Fill-in-the-Middle (FIM) mid-training for coding agent foundation models is out. See the paper, data, and models.
OpenSkill public preprint is available. See the paper and code.
RewardHarness has been accepted to COLM 2026. See the paper and code.
ClawBench public preprint is available. See the paper, code, and data.
earlier news (5)
2026
Watch Before You Answer public preprint is available. Read the paper.
2025
ScholarCopilot has been accepted at COLM 2025! Play with the live demo.
Retri3D has been accepted at ICLR 2025 as a Spotlight 🌟!
WikiGap is released! Try the Chrome extension.
StructEval is released! Check out the dataset.

Selected Publications (view all )

2026 8 papers
Cover image for RewardHarness: Self-Evolving Agentic Post-Training

RewardHarness: Self-Evolving Agentic Post-Training

Yuxuan Zhang, Cong Wei, Penghui Du, Bo Li, Junwen Miao, Huaisong Zhang, Songcheng Cai, Yubo Wang, Dongfu Jiang, Yuyu Zhang, Ping Nie, Wenhu Chen, Changqian Yu, Kelsey R. Allen

COLM 2026

Cover image for ClawBench: Can AI Agents Complete Everyday Online Tasks?

ClawBench: Can AI Agents Complete Everyday Online Tasks?

Yuxuan Zhang, Yubo Wang, Yipeng Zhu, Penghui Du, Junwen Miao, Xuan Lu, Zhuofeng Li, Xingwei Qu, Zhengkang Guo, Yuanzhe Shen, Dingjie Song, Han Zhou, Tuney Zheng, Xian Wu, Hao Yu, Songcheng Cai, Yi Lu, Yunzhuo Hao, Minyi Lei, Liang Chen, Kai Zou, Huifeng Yin, Wendong Xu, Dongfu Jiang, Ping Nie, Jiaheng Liu, Wenhu Chen, Kelsey R. Allen

COLM 2026 WAB

Adopted by ByteDance Seed · Doubao

Cover image for Watch Before You Answer: Learning from Visually Grounded Post-Training

Watch Before You Answer: Learning from Visually Grounded Post-Training

Yuxuan Zhang, EunJeong Hwang, Huaisong Zhang, Penghui Du, Yiming Jia, Dongfu Jiang, Xuan He, Shenhui Zhang, Ping Nie, Peter West, Kelsey R. Allen

Under Review

Cover image for Terminal-Bench-Science: Evaluating AI Agents on Complex Real-World Scientific Workflows in the Terminal

Terminal-Bench-Science: Evaluating AI Agents on Complex Real-World Scientific Workflows in the Terminal

Open Source Benchmark · Contributor

Cover image for Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

Yubo Wang, Jiarong Liang, Yuxuan Zhang, Xuye Liu, Cong Wei, Yuyu Zhang, Ping Nie, Wenhu Chen

Under Review

Cover image for OpenSkill: Open-World Self-Evolution for LLM Agents

OpenSkill: Open-World Self-Evolution for LLM Agents

Zhiling Yan, Dingjie Song, Hanrong Zhang, Wei Liang, Yuxuan Zhang, Yutong Dai, Lifang He, Philip S. Yu, Ran Xu, Xiang Li, Lichao Sun

Under Review

Cover image for MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection

MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection

Haowen Wang, Yaxin Du, Jian Yang, Jiajun Wu, Shukai Liu, Yuxuan Zhang, Pingjie Wang, Siheng Chen, Tuney Zheng, Ming Zhou, Xianglong Liu, Bryan Dai

Under Review

Cover image for Dr. Claw: An AI Research Assistant and Full-Stack Research Workspace

Dr. Claw: An AI Research Assistant and Full-Stack Research Workspace

Open Source · 1k+ stars

2025 5 papers
Cover image for Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports

Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports

Yang Yao, Yixu Wang, Yuxuan Zhang, Yi Lu, Tianle Gu, Lingyu Li, Dingyi Zhao, Keming Wu, Haozhe Wang, Ping Nie, Yan Teng, Yingchun Wang

Under Review

Cover image for ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations

ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations

Yubo Wang, Xueguang Ma, Ping Nie, Huaye Zeng, Zhiheng Lyu, Yuxuan Zhang, Benjamin Schneider, Yi Lu, Xiang Yue, Wenhu Chen

COLM 2025

Cover image for Retri3D: 3D Neural Graphics Representation Retrieval

Retri3D: 3D Neural Graphics Representation Retrieval

Yushi Guan, Daniel Kwan, Jean Sebastien Dandurand, Xi Yan, Ruofan Liang, Yuxuan Zhang, Nilesh Jain, Nilesh Ahuja, Selvakumar Panneer, Nandita Vijaykumar

ICLR 2025 Spotlight

Cover image for VideoScore2: Think before You Score in Generative Video Evaluation

VideoScore2: Think before You Score in Generative Video Evaluation

TMLR 2026

Cover image for StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs

StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs

Jialin Yang, Dongfu Jiang, Lipeng He, Sherman Siu, Yuxuan Zhang, Disen Liao, Zhuofeng Li, Huaye Zeng, Yiming Jia, Haozhe Wang, Benjamin Schneider, Chi Ruan, Wentao Ma, Zhiheng Lyu, Yifei Wang, Yi Lu, Quy Duc Do, Ziyan Jiang, Ping Nie, Wenhu Chen

TMLR 2025

Impact

ClawBench
AdoptedClawBench is used as a headline agent benchmark in frontier model technical reports. Li Auto's Mach-Mind-4-Flash reports its ClawBench score in the abstract and benchmarks eight model families on it, including Kimi K2.5, Qwen3.5, GLM-4.7 and Nemotron-3. 2026
CitedClawBench has been cited by agent-benchmark work from Shanghai AI Laboratory (WildClawBench, π-Bench), Carnegie Mellon (MyPCBench), HKU MMLab and Meituan (UniClawBench), and Peking University (Harness-Bench). 2026
LeaderboardClawBench is tracked as a standing leaderboard by the browser-agent platform Steel, alongside WebVoyager and BrowseComp. 2026
VideoScore2
IntegratedVideoScore2 ships as a built-in evaluation metric in FastVideo, UC San Diego Hao AI Lab's video generation framework. 2026
IntegratedVideoScore2 is used as a reward model by Google's VQQA video-evaluation framework, which runs it as a Best-of-N selector across its main results tables. 2026
StructEval
IntegratedStructEval is integrated as a scored benchmark in llm-jp-eval, the evaluation suite of Japan's national LLM project (LLM-jp / NII). 2026
CitedStructEval has been cited by structured-generation work from Apple, Microsoft Research and Amazon Science. 2025 - 2026
ScholarCopilot
CitedScholarCopilot is surveyed in "Transforming Science with Large Language Models," a widely cited review of AI-assisted scientific discovery. 2025 - 2026

Talks

Academic Service

Organizer
AACL 2026 Workshop on Autonomous Research and Recursive Self-Improvement (ARRSI)
Reviewer
ACL ARR 2026 ICMLW 2026 ICLRW 2026 NeurIPS 2026 COLM 2026 CVPR 2026 ECCV 2026

Advisors & Collaborators