Yuxuan Zhang
Hi there! I'm Yuxuan Zhang, a PhD student at the University of British Columbia, advised by Prof. Kelsey Allen. I'm also affiliated with the Vector Institute. My research interests span:
I build benchmarks and post-training methods that make AI agents work in the real world: ClawBench for live web-agent evaluation (adopted by ByteDance Seed · Doubao), RewardHarness for self-evolving agentic post-training (COLM 2026), and VidGround for grounded video understanding.
I'm open to collaborations and always excited to chat about research ideas. Feel free to reach out at
reacher [at] cs.ubc.ca ( WeChat works too).
I am actively looking for research internships. Please reach out if my background looks like a great fit for your organization.
Let's grab a coffee and have a chat!
Research Highlights
Can AI Agents Complete Everyday Online Tasks?
ClawBench evaluates AI agents on 153 real-world tasks across 144 live platforms — from booking appointments to filing job applications. Unlike sandboxed benchmarks, ClawBench runs on production websites, intercepting only the final submission to keep evaluation safe without losing real-world complexity.
Watch Before You Answer
Language models can answer video questions using text priors alone, without truly watching the video. We identify and filter such spurious training samples, producing a cleaner dataset that teaches models to genuinely ground their answers in visual content.
RewardHarness: Self-Evolving Agentic Post-Training
Evaluating instruction-guided image edits requires rewards that capture subtle human preferences — yet existing reward models rely on hundreds of thousands of annotations. RewardHarness reframes reward modeling as context evolution: from only 100 preference demonstrations it iteratively evolves a library of tools and skills, surpassing GPT-5 by 5.3 points on image-editing evaluation benchmarks while using just 0.05% of the EditReward data.
-
University of British ColumbiaPhD in Computer ScienceJan. 2025 - Present
-
University of TorontoMSc in Applied ComputingSept. 2022 - June 2024 -
Peking UniversityBachelor in EconomicsSept. 2019 - June 2022
-
SOTIData ScientistJan. 2024 - Dec. 2024Research InternMay 2023 - Dec. 2023
-
2026 Compute Transparency Champion Award, CVPR 2026
-
2025 President's Academic Excellence Award
-
2023 Mitacs Accelerate Fellowship ($30,000)
News
Selected Publications (view all )
2026 10 papers
RewardHarness: Self-Evolving Agentic Post-Training
Conference on Language Modeling (COLM) 2026
Project arXiv Code HuggingFace BibTeX
@article{zhang2026rewardharness,
title={RewardHarness: Self-Evolving Agentic Post-Training},
author={Zhang, Yuxuan and Du, Penghui and Li, Bo and Wei, Cong and Miao, Junwen and Zhang, Huaisong and Cai, Songcheng and Wang, Yubo and Jiang, Dongfu and Zhang, Yuyu and others},
journal={arXiv preprint arXiv:2605.08703},
year={2026}
}
RewardHarness: Self-Evolving Agentic Post-Training
Conference on Language Modeling (COLM) 2026
Project arXiv Code HuggingFace BibTeX
@article{zhang2026rewardharness,
title={RewardHarness: Self-Evolving Agentic Post-Training},
author={Zhang, Yuxuan and Du, Penghui and Li, Bo and Wei, Cong and Miao, Junwen and Zhang, Huaisong and Cai, Songcheng and Wang, Yubo and Jiang, Dongfu and Zhang, Yuyu and others},
journal={arXiv preprint arXiv:2605.08703},
year={2026}
}
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Under Review
Project arXiv Code Dataset HuggingFace BibTeX
@article{zhang2026clawbench,
title={ClawBench: Can AI Agents Complete Everyday Online Tasks?},
author={Zhang, Yuxuan and Wang, Yubo and Zhu, Yipeng and Du, Penghui and Miao, Junwen and Lu, Xuan and Xu, Wendong and Hao, Yunzhuo and Cai, Songcheng and Wang, Xiaochen and others},
journal={arXiv preprint arXiv:2604.08523},
year={2026}
}
ClawBench: Can AI Agents Complete Everyday Online Tasks?
Under Review
Project arXiv Code Dataset HuggingFace BibTeX
@article{zhang2026clawbench,
title={ClawBench: Can AI Agents Complete Everyday Online Tasks?},
author={Zhang, Yuxuan and Wang, Yubo and Zhu, Yipeng and Du, Penghui and Miao, Junwen and Lu, Xuan and Xu, Wendong and Hao, Yunzhuo and Cai, Songcheng and Wang, Xiaochen and others},
journal={arXiv preprint arXiv:2604.08523},
year={2026}
}
Watch Before You Answer: Learning from Visually Grounded Post-Training
Under Review
Project arXiv Code HuggingFace BibTeX
@article{zhang2026watch,
title={Watch before you answer: Learning from visually grounded post-training},
author={Zhang, Yuxuan and Hwang, EunJeong and Zhang, Huaisong and Du, Penghui and Jia, Yiming and Jiang, Dongfu and He, Xuan and Zhang, Shenhui and Nie, Ping and West, Peter and others},
journal={arXiv preprint arXiv:2604.05117},
year={2026}
}
Watch Before You Answer: Learning from Visually Grounded Post-Training
Under Review
Project arXiv Code HuggingFace BibTeX
@article{zhang2026watch,
title={Watch before you answer: Learning from visually grounded post-training},
author={Zhang, Yuxuan and Hwang, EunJeong and Zhang, Huaisong and Du, Penghui and Jia, Yiming and Jiang, Dongfu and He, Xuan and Zhang, Shenhui and Nie, Ping and West, Peter and others},
journal={arXiv preprint arXiv:2604.05117},
year={2026}
}
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
Under Review
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
Under Review
Learning from the Self-future: On-policy Self-distillation for dLLMs
Under Review
@article{luo2026learning,
title={Learning from the Self-future: On-policy Self-distillation for dLLMs},
author={Luo, Yifu and Chen, Zeyu and Wang, Haoyu and Hu, Xinhao and Zhang, Yuxuan and Sha, Zhizhou and Liu, Shiwei},
journal={arXiv preprint arXiv:2606.18195},
year={2026}
}
Learning from the Self-future: On-policy Self-distillation for dLLMs
Under Review
@article{luo2026learning,
title={Learning from the Self-future: On-policy Self-distillation for dLLMs},
author={Luo, Yifu and Chen, Zeyu and Wang, Haoyu and Hu, Xinhao and Zhang, Yuxuan and Sha, Zhizhou and Liu, Shiwei},
journal={arXiv preprint arXiv:2606.18195},
year={2026}
}
CompRank: Efficient LLM Reranking via Token-Level Compression and Decoding-Free Scoring
Under Review
@article{lu2026comprank,
title={CompRank: Efficient LLM Reranking via Token-Level Compression and Decoding-Free Scoring},
author={Lu, Xuan and Huang, Haohang and Fan, Yingqi and Tong, Junlong and Zhang, Yuxuan and Nie, Ping and Meng, Rui and Shen, Xiaoyu},
journal={arXiv preprint arXiv:2606.11700},
year={2026}
}
CompRank: Efficient LLM Reranking via Token-Level Compression and Decoding-Free Scoring
Under Review
@article{lu2026comprank,
title={CompRank: Efficient LLM Reranking via Token-Level Compression and Decoding-Free Scoring},
author={Lu, Xuan and Huang, Haohang and Fan, Yingqi and Tong, Junlong and Zhang, Yuxuan and Nie, Ping and Meng, Rui and Shen, Xiaoyu},
journal={arXiv preprint arXiv:2606.11700},
year={2026}
}
OpenSkill: Open-World Self-Evolution for LLM Agents
Under Review
@article{yan2026openskill,
title={OpenSkill: Open-World Self-Evolution for LLM Agents},
author={Yan, Zhiling and Song, Dingjie and Zhang, Hanrong and Liang, Wei and Zhang, Yuxuan and Dai, Yutong and He, Lifang and Yu, Philip S and Xu, Ran and Li, Xiang and others},
journal={arXiv preprint arXiv:2606.06741},
year={2026}
}
OpenSkill: Open-World Self-Evolution for LLM Agents
Under Review
@article{yan2026openskill,
title={OpenSkill: Open-World Self-Evolution for LLM Agents},
author={Yan, Zhiling and Song, Dingjie and Zhang, Hanrong and Liang, Wei and Zhang, Yuxuan and Dai, Yutong and He, Lifang and Yu, Philip S and Xu, Ran and Li, Xiang and others},
journal={arXiv preprint arXiv:2606.06741},
year={2026}
}
Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback
Under Review
@article{zhang2026and,
title={Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback},
author={Zhang, Huaisong and Yu, Hao and Zhang, Yuxuan and Wang, Jiahe and Chen, Xinrui and Cao, Haoxiang and Lu, Feng and Zhang, Wendong and Yu, Changqian and Yuan, Chun},
journal={arXiv preprint arXiv:2606.06113},
year={2026}
}
Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback
Under Review
@article{zhang2026and,
title={Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback},
author={Zhang, Huaisong and Yu, Hao and Zhang, Yuxuan and Wang, Jiahe and Chen, Xinrui and Cao, Haoxiang and Lu, Feng and Zhang, Wendong and Yu, Changqian and Yuan, Chun},
journal={arXiv preprint arXiv:2606.06113},
year={2026}
}
MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection
Under Review
@article{wang2026mira,
title={MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection},
author={Wang, Haowen and Du, Yaxin and Yang, Jian and Wu, Jiajun and Liu, Shukai and Zhang, Yuxuan and Wang, Pingjie and Chen, Siheng and Zheng, Tuney and Zhou, Ming and others},
journal={arXiv preprint arXiv:2605.30288},
year={2026}
}
MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection
Under Review
@article{wang2026mira,
title={MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection},
author={Wang, Haowen and Du, Yaxin and Yang, Jian and Wu, Jiajun and Liu, Shukai and Zhang, Yuxuan and Wang, Pingjie and Chen, Siheng and Zheng, Tuney and Zhou, Ming and others},
journal={arXiv preprint arXiv:2605.30288},
year={2026}
}
2025 8 papers
Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
Under Review
@article{yao2026dr,
title={Dr. bench: A multidimensional evaluation for deep research agents, from answers to reports},
author={Yao, Yang and Wang, Yixu and Zhang, Yuxuan and Lu, Yi and Gu, Tianle and Li, Lingyu and Zhao, Dingyi and Wu, Keming and Wang, Haozhe and Nie, Ping and others},
year={2026}
}
Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
Under Review
@article{yao2026dr,
title={Dr. bench: A multidimensional evaluation for deep research agents, from answers to reports},
author={Yao, Yang and Wang, Yixu and Zhang, Yuxuan and Lu, Yi and Gu, Tianle and Li, Lingyu and Zhao, Dingyi and Wu, Keming and Wang, Haozhe and Nie, Ping and others},
year={2026}
}
ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations
Conference on Language Modeling (COLM) 2025
Project Paper Code Dataset Model Demo BibTeX
@article{wang2025scholarcopilot,
title={Scholarcopilot: Training large language models for academic writing with accurate citations},
author={Wang, Yubo and Ma, Xueguang and Nie, Ping and Zeng, Huaye and Lyu, Zhiheng and Zhang, Yuxuan and Schneider, Benjamin and Lu, Yi and Yue, Xiang and Chen, Wenhu},
journal={arXiv preprint arXiv:2504.00824},
year={2025}
}
ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations
Conference on Language Modeling (COLM) 2025
Project Paper Code Dataset Model Demo BibTeX
@article{wang2025scholarcopilot,
title={Scholarcopilot: Training large language models for academic writing with accurate citations},
author={Wang, Yubo and Ma, Xueguang and Nie, Ping and Zeng, Huaye and Lyu, Zhiheng and Zhang, Yuxuan and Schneider, Benjamin and Lu, Yi and Yue, Xiang and Chen, Wenhu},
journal={arXiv preprint arXiv:2504.00824},
year={2025}
}
Retri3D: 3D Neural Graphics Representation Retrieval
International Conference on Learning Representations (ICLR) 2025 Spotlight
@inproceedings{guan2025retri3d,
title={Retri3D: 3D Neural Graphics Representation Retrieval},
author={Guan, Yushi and Kwan, Daniel and Dandurand, Jean and Yan, Xi and Liang, Ruofan and Zhang, Yuxuan and Jain, Nilesh and Ahuja, Nilesh and Panneer, Selvakumar and Vijaykumar, Nandita},
booktitle={International Conference on Learning Representations},
volume={2025},
pages={31771--31813},
year={2025}
}
Retri3D: 3D Neural Graphics Representation Retrieval
International Conference on Learning Representations (ICLR) 2025 Spotlight
@inproceedings{guan2025retri3d,
title={Retri3D: 3D Neural Graphics Representation Retrieval},
author={Guan, Yushi and Kwan, Daniel and Dandurand, Jean and Yan, Xi and Liang, Ruofan and Zhang, Yuxuan and Jain, Nilesh and Ahuja, Nilesh and Panneer, Selvakumar and Vijaykumar, Nandita},
booktitle={International Conference on Learning Representations},
volume={2025},
pages={31771--31813},
year={2025}
}
VideoScore2: Think before You Score in Generative Video Evaluation
Transactions on Machine Learning Research (TMLR) 2026
Project Paper Code Dataset Model HF Space Twitter BibTeX
@article{he2025videoscore2,
title={Videoscore2: Think before you score in generative video evaluation},
author={He, Xuan and Jiang, Dongfu and Nie, Ping and Liu, Minghao and Jiang, Zhengxuan and Su, Mingyi and Ma, Wentao and Lin, Junru and Ye, Chun and Lu, Yi and others},
journal={arXiv preprint arXiv:2509.22799},
year={2025}
}
VideoScore2: Think before You Score in Generative Video Evaluation
Transactions on Machine Learning Research (TMLR) 2026
Project Paper Code Dataset Model HF Space Twitter BibTeX
@article{he2025videoscore2,
title={Videoscore2: Think before you score in generative video evaluation},
author={He, Xuan and Jiang, Dongfu and Nie, Ping and Liu, Minghao and Jiang, Zhengxuan and Su, Mingyi and Ma, Wentao and Lin, Junru and Ye, Chun and Lu, Yi and others},
journal={arXiv preprint arXiv:2509.22799},
year={2025}
}
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
Transactions on Machine Learning Research (TMLR) 2025
Project Paper Code Dataset BibTeX
@article{yang2025structeval,
title={StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs},
author={Yang, Jialin and Jiang, Dongfu and He, Lipeng and Siu, Sherman and Zhang, Yuxuan and Liao, Disen and Li, Zhuofeng and Zeng, Huaye and Jia, Yiming and Wang, Haozhe and others},
journal={arXiv preprint arXiv:2505.20139},
year={2025}
}
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
Transactions on Machine Learning Research (TMLR) 2025
Project Paper Code Dataset BibTeX
@article{yang2025structeval,
title={StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs},
author={Yang, Jialin and Jiang, Dongfu and He, Lipeng and Siu, Sherman and Zhang, Yuxuan and Liao, Disen and Li, Zhuofeng and Zeng, Huaye and Jia, Yiming and Wang, Haozhe and others},
journal={arXiv preprint arXiv:2505.20139},
year={2025}
}
WikiGap: Promoting Epistemic Equity by Surfacing Knowledge Gaps Between English Wikipedia and other Language Editions
Under Review
@article{wang2025wikigap,
title={WikiGap: Promoting epistemic equity by surfacing knowledge gaps between English Wikipedia and other language editions},
author={Wang, Zining and Zhang, Yuxuan and Yoon, Dongwook and Vincent, Nicholas and Samir, Farhan and Shwartz, Vered},
journal={arXiv preprint arXiv:2505.24195},
year={2025}
}
WikiGap: Promoting Epistemic Equity by Surfacing Knowledge Gaps Between English Wikipedia and other Language Editions
Under Review
@article{wang2025wikigap,
title={WikiGap: Promoting epistemic equity by surfacing knowledge gaps between English Wikipedia and other language editions},
author={Wang, Zining and Zhang, Yuxuan and Yoon, Dongwook and Vincent, Nicholas and Samir, Farhan and Shwartz, Vered},
journal={arXiv preprint arXiv:2505.24195},
year={2025}
}
PLAICraft: Large-Scale Time-Aligned Vision-Speech-Action Dataset for Embodied AI
Under Review
@article{he2025plaicraft,
title={Plaicraft: Large-scale time-aligned vision-speech-action dataset for embodied ai},
author={He, Yingchen and Weilbach, Christian D and Wojciechowska, Martyna E and Zhang, Yuxuan and Wood, Frank},
journal={arXiv preprint arXiv:2505.12707},
year={2025}
}
PLAICraft: Large-Scale Time-Aligned Vision-Speech-Action Dataset for Embodied AI
Under Review
@article{he2025plaicraft,
title={Plaicraft: Large-scale time-aligned vision-speech-action dataset for embodied ai},
author={He, Yingchen and Weilbach, Christian D and Wojciechowska, Martyna E and Zhang, Yuxuan and Wood, Frank},
journal={arXiv preprint arXiv:2505.12707},
year={2025}
}
Enhancing Vector Quantization with Distributional Matching: A Theoretical and Empirical Study
Under Review
@article{fang2025enhancing,
title={Enhancing vector quantization with distributional matching: A theoretical and empirical study},
author={Fang, Xianghong and Guo, Litao and Chen, Hengchao and Zhang, Yuxuan and Song, Dingjie and Liu, Yexin and Wang, Hao and Yang, Harry and Yuan, Yuan and Sun, Qiang and others},
journal={arXiv preprint arXiv:2506.15078},
year={2025}
}
Enhancing Vector Quantization with Distributional Matching: A Theoretical and Empirical Study
Under Review
@article{fang2025enhancing,
title={Enhancing vector quantization with distributional matching: A theoretical and empirical study},
author={Fang, Xianghong and Guo, Litao and Chen, Hengchao and Zhang, Yuxuan and Song, Dingjie and Liu, Yexin and Wang, Hao and Yang, Harry and Yuan, Yuan and Sun, Qiang and others},
journal={arXiv preprint arXiv:2506.15078},
year={2025}
}
2024 1 paper
Edge-Enhanced Dilated Residual Attention Network for Multimodal Medical Image Fusion
IEEE International Conference on Bioinformatics and Biomedicine (BIBM) 2024
@inproceedings{zhou2024edge,
title={Edge-enhanced dilated residual attention network for multimodal medical image fusion},
author={Zhou, Meng and Zhang, Yuxuan and Xu, Xiaolan and Wang, Jiayi and Khalvati, Farzad},
booktitle={2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)},
pages={4108--4111},
year={2024},
organization={IEEE}
}
Edge-Enhanced Dilated Residual Attention Network for Multimodal Medical Image Fusion
IEEE International Conference on Bioinformatics and Biomedicine (BIBM) 2024
@inproceedings{zhou2024edge,
title={Edge-enhanced dilated residual attention network for multimodal medical image fusion},
author={Zhou, Meng and Zhang, Yuxuan and Xu, Xiaolan and Wang, Jiayi and Khalvati, Farzad},
booktitle={2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)},
pages={4108--4111},
year={2024},
organization={IEEE}
}
2021 1 paper
CelebHair: A New Large-Scale Dataset for Hairstyle Recommendation Based on CelebA
Knowledge Science, Engineering and Management (KSEM) 2021
@inproceedings{chen2021celebhair,
title={Celebhair: A new large-scale dataset for hairstyle recommendation based on celeba},
author={Chen, Yutao and Zhang, Yuxuan and Huang, Zhongrui and Luo, Zhenyao and Chen, Jinpeng},
booktitle={International Conference on Knowledge Science, Engineering and Management},
pages={323--336},
year={2021},
organization={Springer}
}
CelebHair: A New Large-Scale Dataset for Hairstyle Recommendation Based on CelebA
Knowledge Science, Engineering and Management (KSEM) 2021
@inproceedings{chen2021celebhair,
title={Celebhair: A new large-scale dataset for hairstyle recommendation based on celeba},
author={Chen, Yutao and Zhang, Yuxuan and Huang, Zhongrui and Luo, Zhenyao and Chen, Jinpeng},
booktitle={International Conference on Knowledge Science, Engineering and Management},
pages={323--336},
year={2021},
organization={Springer}
}
Academic Service
I have served as a reviewer for the following venues: