Publications — Yaoyao (Freax) Qian
14 publications by Yaoyao Qian across LLM agents, agent evaluation, dialogue systems, robotic manipulation, and AI ethics.
DeployBench: Benchmarking LLM Agents for Research Artifact Deployment
Yuanli Wang, Yaoyao Qian, Yue Zhang, Hanhan Zhou, Jindan Huang, Tianfu Fu, Qiuyang Mang, Huanzhi Mao, Wenhao Chai, Wendong Fan, Liqiang Jing. “DeployBench: Benchmarking LLM Agents for Research Artifact Deployment.” EMNLP Findings 2026.
LLM agents have made rapid progress on software engineering and ML research tasks, but these advances often assume access to a working runnable environment. For research artifacts released alongside published papers, setting up such an environment from a fresh machine remains a major bottleneck. Existing environment setup benchmarks do not cover the full scope of research artifact deployment, which involves multi-language toolchains, system-level dependencies beyond containers (e.g. GPU/CUDA and kernel configurations), and legacy artifact compatibility. We introduce DeployBench, a multi-domain benchmark of 51 research-artifact deployment tasks spanning AI/ML, computer systems, and scientific computing, covering all these dimensions. Each task is verified by a hidden pipeline that executes the paper's designated experiment and checks its outputs. Evaluating four state-of-the-art LLMs with OpenHands yields pass-rates from 7.8% - 51.0%. Failures are dominated by a completion-judgment problem: 97 of 154 are agent-terminated self-stops, where the agent's pre-finish checks validate a different or weaker target than the paper-specific task requires. DeployBench highlights the gap between current agents and autonomous deployment, and offers a realistic testbed for scientific research agents.
Keywords: LLM Agents, Benchmark, Research Artifact Deployment, Environment Setup, Software Engineering, Evaluation.
arXiv · Project page
Inside the Agentic Classroom: Simulating Instructional Dynamics with AI Student Agents
Zhicheng Guo, Xiaoying Zheng, Yueru Yan, Eryclis Rodrigues Bezerra Silva, Yaoyao Qian, Kyrie Zhixuan Zhou, Ece Gumusel. “Inside the Agentic Classroom: Simulating Instructional Dynamics with AI Student Agents.” ASIS&T 2026.
Keywords: AI Student Agents, Instructional Dynamics, Agent Simulation, Education, LLM Agents.
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Xiangyi Li, Yimin Liu, Wenbo Chen, Bingran You, Zonglin Di, Yifeng He, Shenghan Zheng, Kyoung Whan Choe, Jiankai Sun, Shuyi Wang, Chujun Tao, Binxu Li, Xuandong Zhao, Hejia Geng, Xiaojun Wu, Junwei Zhou, Xiaokun Chen, Hanwen Xing, Yubo Li, Qunhong Zeng, Di Wang, Yuanli Wang, Roey Ben Chaim, Penghao Jiang, Haotian Shen, Luyang Kong, Xinyi Liu, Runhui Wang, Xuanqing Liu, Jiachen Li, Xin Lan, Yueqian Lin, Wengao Ye, Junwei He, Songlin Li, Yue Zhang, Yipeng Gao, Yijiang Li, Ze Ma, Liqiang Jing, Tianyu Wang, Kaixin Li, Yiqi Xue, Haoran Lyu, Yizhuo He, Yuchen Tian, Shutong Wu, Bowei Wang, Yixuan Gao, Bo Chen, Litong Liu, Sikai Cheng, Jiajun Bao, Shuaicheng Tong, Shuwen Xu, Terry Yue Zhuo, Tinghan Ye, Qi Qi, Miao Li, Longtai Liao, Zelin Tan, Chang Shi, Xilin Tang, Srinath Tankasala, Boqin Yuan, Yaoyao Qian, Jianhong Tu, Chenguang Wang, Yizhou Sun, Wei Wang, Aaron Taylor, Ziyue Yang, Changkun Guan, Zhikang Dong, Xinyu Zhang, Steven Dillmann, Han-chung Lee, Dawn Song. “SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.” Preprint.
Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark whose current inventory contains 87 tasks across 8 domains paired with curated Skills and deterministic verifiers. Our latest aggregate evaluation runs the 87-task benchmark under matched no-Skills and curated-Skills conditions for 18 model-harness configurations. Curated Skills raise the average pass rate from 33.9% to 50.5% (+16.6 percentage points; 25.5% normalized gain), with configuration-level gains ranging from +4.1 to +25.7 pp. Focused Skills with at most three modules outperform larger or exhaustive bundles, and smaller models with Skills can match larger models without them. SkillsBench establishes paired evaluation as the foundation for rigorous measurement of Skill efficacy on agentic, expertise-heavy work.
Keywords: Agent Skills, LLM Agents, Benchmark, Procedural Knowledge, Evaluation, Paired Evaluation.
arXiv · Project page · Dataset
A Very Big Video Reasoning Suite
Maijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji, Thaddäus Wiedemer, Qingying Gao, Dezhi Luo, Yaoyao Qian, Lianyu Huang, Zelong Hong, Jiahui Ge, Qianli Ma, Hang He, Yifan Zhou, Lingzi Guo, Lantao Mei, Jiachen Li, Hanwen Xing, Tianqi Zhao, Fengyuan Yu, Weihang Xiao, Yizheng Jiao, Jianheng Hou, Danyang Zhang, Pengcheng Xu, Boyang Zhong, Zehong Zhao, Gaoyun Fang, John Kitaoka, Yile Xu, Hua Xu, Kenton Blacutt, Tin Nguyen, Siyuan Song, Haoran Sun, Shaoyue Wen, Linyang He, Runming Wang, Yanzhi Wang, Mengyue Yang, Ziqiao Ma, Raphaël Millière, Freda Shi, Nuno Vasconcelos, Daniel Khashabi, Alan Yuille, Yilun Du, Ziming Liu, Bo Li, Dahua Lin, Ziwei Liu, Vikash Kumar, Yijiang Li, Lei Yang, Zhongang Cai, Hokin Deng. “A Very Big Video Reasoning Suite.” ICML 2026.
This work addresses the underexploration of reasoning capabilities in video models by introducing the Very Big Video Reasoning (VBVR) Dataset, encompassing 200 curated reasoning tasks and over one million video clips. The authors present VBVR-Bench, an evaluation framework using rule-based, human-aligned scorers. Through scaling studies, they identify early signs of emergent generalization to unseen reasoning tasks, establishing a foundation for advancing research in spatiotemporal reasoning over video content.
Keywords: Video Reasoning, Spatiotemporal Reasoning, Video Understanding, Benchmark, Evaluation, Large-Scale Dataset.
arXiv · Project page · Dataset · Eval dataset · Leaderboard · Code · Infrastructure · Model
"Everyone Else Does It": The Rise of Preprinting Culture in Computing Disciplines
Kyrie Zhixuan Zhou, Justin Eric Chen, Xiang Zheng, Yaoyao Qian, Yunpeng Xiao, Kai Shu. “"Everyone Else Does It": The Rise of Preprinting Culture in Computing Disciplines.” ASIS&T 2026.
Preprinting has become a norm in fast-paced computing fields such as artificial intelligence (AI) and human-computer interaction (HCI). In this paper, we conducted semi-structured interviews with 15 academics in these fields to reveal their motivations and perceptions of preprinting. The results found a close relationship between preprinting and characteristics of the fields, including the huge number of papers, competitiveness in career advancement, prevalence of scooping, and imperfect peer review system -- preprinting comes to the rescue in one way or another for the participants. Based on the results, we reflect on the role of preprinting in subverting the traditional publication mode and outline possibilities of a better publication ecosystem. Our study contributes by inspecting the community aspects of preprinting practices through talking to academics.
Keywords: Preprints, Scholarly Communication, Computing Disciplines.
arXiv · Project page
WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph Representation
Yaoyao Qian, Yuanli Wang, Jinda Zhang, Yun Zong, Meixu Chen, Hanhan Zhou, Jindan Huang, Yifan Zeng, Xinyu Hu, Chan Hee Song, Danqing Zhang. “WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph Representation.” MTI-LLM @ NeurIPS 2025.
Current evaluation of web agents largely reduces to binary success metrics or conformity to a single reference trajectory, ignoring the structural diversity present in benchmark datasets. We present WebGraphEval, a framework that abstracts trajectories from multiple agents into a unified, weighted action graph. This graph representation is directly compatible with existing benchmarks such as WebArena, using both leaderboard trajectories and newly collected runs, and provides a principled basis for analyzing solution spaces without modifying environments. The framework canonically encodes actions, merges recurring behaviors, and applies structural analyses including reward propagation and success-weighted edge statistics. Evaluations across thousands of trajectories from six web agents demonstrate that the graph abstraction captures cross-model regularities, highlights redundancy and inefficiency, and identifies critical decision points overlooked by outcome-based metrics. By framing web interaction as graph-structured data, WebGraphEval establishes a general methodology for multi-path, cross-agent, and efficiency-aware evaluation of web agents.
Keywords: Web agent, Evaluation, Large Language Models.
Demo · Project page
Clebsch-Gordan Transformer: Fast and Global Equivariant Attention
Owen Lewis Howell, Linfeng Zhao, Xupeng Zhu, Yaoyao Qian, Haojie Huang, Lingfeng Sun, Wil Thomason, Robert Platt, Robin Walters. “Clebsch-Gordan Transformer: Fast and Global Equivariant Attention.” Preprint.
The global attention mechanism is one of the keys to the success of transformer architecture, but it incurs quadratic computational costs in relation to the number of tokens. On the other hand, equivariant models, which leverage the underlying geometric structures of problem instance, often achieve superior accuracy in physical, biochemical, computer vision, and robotic tasks, at the cost of additional compute requirements. As a result, existing equivariant transformers only support low-order equivariant features and local context windows, limiting their expressiveness and performance. This work proposes Clebsch-Gordan Transformer, achieving efficient global attention by a novel Clebsch-Gordon Convolution on SO(3) irreducible representations. Our method enables equivariant modeling of features at all orders while achieving O(N log N) input token complexity. Additionally, the proposed method scales well with high-order irreducible features, by exploiting the sparsity of the Clebsch-Gordon matrix. Lastly, we also incorporate optional token permutation equivariance through either weight sharing or data augmentation. We benchmark our method on a diverse set of benchmarks including n-body simulation, QM9, ModelNet point cloud classification and a robotic grasping dataset, showing clear gains over existing equivariant transformers in GPU memory size, speed, and accuracy.
Keywords: Equivariance, Transformer, SO(3), Clebsch-Gordan, Global Attention, Fast Attention, Point Cloud, Robotics.
arXiv
The Ranking Blind Spot: Decision Hijacking in LLM-based Text Ranking
Yaoyao Qian*, Yifan Zeng*, Yuchao Jiang, Chelsi Jain, Huazheng Wang. “The Ranking Blind Spot: Decision Hijacking in LLM-based Text Ranking.” EMNLP 2025.
Large Language Models (LLMs) show strong performance in information retrieval tasks like passage ranking. We examine how instruction-following capabilities interact with multi-document comparison and identify a Ranking Blind Spot—a characteristic of LLM decision processes during comparative evaluation. We study its impact on LLM-based evaluation systems via two attack families: Decision Objective Hijacking (altering the ranking objective in pairwise ranking systems) and Decision Criteria Hijacking (shifting relevance criteria across ranking schemes). These attacks push the ranker to prefer a specific passage, placing it at the top. Malicious content providers can exploit this weakness to gain exposure by attacking the ranker. Empirically, we show the attacks are effective across multiple LLMs and generalize across ranking schemes on realistic examples, with stronger LLMs often more vulnerable. Code: https://github.com/blindspotorg/RankingBlindSpot
Keywords: LLM, Text Ranking, Information Retrieval, Adversarial Attack, Decision Making, NLP.
Paper · arXiv · Project page · Code
WHEN TO ACT, WHEN TO WAIT: Modeling the Intent-Action Alignment Problem in Dialogue
Yaoyao Qian, Jindan Huang, Yuanli Wang, Simon Yu, Kyrie Zhixuan Zhou, Jiayuan Mao, Mingfu Liang, Hanhan Zhou. “WHEN TO ACT, WHEN TO WAIT: Modeling the Intent-Action Alignment Problem in Dialogue.” Social Sim @ COLM 2025.
Dialogue systems often fail when user utterances are semantically complete yet lack the clarity and completeness required for appropriate system action. This mismatch arises because users frequently do not fully understand their own needs, while systems require precise intent definitions. This highlights the critical Intent-Action Alignment Problem: determining when an expression is not just understood, but truly ready for a system to act upon. We present STORM, a framework modeling asymmetric information dynamics through conversations between UserLLM (full internal access) and AgentLLM (observable behavior only). STORM produces annotated corpora capturing trajectories of expression phrasing and latent cognitive transitions, enabling systematic analysis of how collaborative understanding develops. Our contributions include: (1) formalizing asymmetric information processing in dialogue systems; (2) modeling intent formation tracking collaborative understanding evolution; and (3) evaluation metrics measuring internal cognitive improvements alongside task performance. Experiments across four language models reveal that moderate uncertainty (40-60%) can outperform complete transparency in certain scenarios, with model-specific patterns suggesting reconsideration of optimal information completeness in human-AI collaboration. These findings contribute to understanding asymmetric reasoning dynamics and inform uncertainty-calibrated dialogue system design.
Keywords: Dialogue Systems, Intent Recognition, Task-Oriented Dialogue, NLP, Asymmetric Information, Human-AI Collaboration.
Paper · Project page · Code · Dataset · Dashboard
Hierarchical Equivariant Policy via Frame Transfer
Haibo Zhao*, Dian Wang*, Yizhe Zhu, Xupeng Zhu, Owen Howell, Linfeng Zhao, Yaoyao Qian, Robin Walters, Robert Platt. “Hierarchical Equivariant Policy via Frame Transfer.” ICML 2025.
Recent advances in hierarchical policy learning highlight the advantages of decomposing systems into high-level and low-level agents, enabling efficient long-horizon reasoning and precise fine-grained control. However, the interface between these hierarchy levels remains underexplored, and existing hierarchical methods often ignore domain symmetry, resulting in the need for extensive demonstrations to achieve robust performance. To address these issues, we propose Hierarchical Equivariant Policy (HEP), a novel hierarchical policy framework. We propose a frame transfer interface for hierarchical policy learning, which uses the high-level agent's output as a coordinate frame for the low-level agent, providing a strong inductive bias while retaining flexibility. Additionally, we integrate domain symmetries into both levels and theoretically demonstrate the system's overall equivariance. HEP achieves state-of-the-art performance in complex robotic manipulation tasks, demonstrating significant improvements in both simulation and real-world settings.
Keywords: Robotics, Equivariant Learning, Policy Learning, Hierarchical Learning, Machine Learning.
arXiv
VisualTreeSearch: Understanding Web Agent Test-time Scaling
Danqing Zhang, Yaoyao Qian, Shiying He, Yuanli Wang, Jingyi Ni, Junyu Cao. “VisualTreeSearch: Understanding Web Agent Test-time Scaling.” ECML-PKDD 2025.
We present VisualTreeSearch, a fully-deployed system for visualizing and understanding web agent test-time scaling. While test-time search algorithms substantially improve web agent success rates, they remain confined to research contexts with limited practical deployment. Our system bridges this gap with three key contributions: (1) a production-ready solution with cloud-based architecture, (2) an efficient API-based state reset mechanism that reduces state reset time from 50 to 2 seconds, and (3) an interactive web UI that transparently demonstrates the agent's decision-making process. VisualTreeSearch provides an intuitive framework for both researchers and users to understand tree search execution in web agents.
Keywords: Web Agents, Tree Search, Vision-Language Models, Test-time Scaling, Interactive Visualization.
Paper
ThinkGrasp: A Vision-Language System for Strategic Part Grasping in Clutter
Yaoyao Qian, Xupeng Zhu, Ondrej Biza, Shuo Jiang, Linfeng Zhao, Haojie Huang, Yu Qi, Robert Platt. “ThinkGrasp: A Vision-Language System for Strategic Part Grasping in Clutter.” CoRL 2024.
Robotic grasping in cluttered environments remains a significant challenge due to occlusions and complex object arrangements. We have developed ThinkGrasp, a plug-and-play vision-language grasping system that makes use of GPT-4o's advanced contextual reasoning for heavy clutter environment grasping strategies. ThinkGrasp can effectively identify and generate grasp poses for target objects, even when they are heavily obstructed or nearly invisible, by using goal-oriented language to guide the removal of obstructing objects. This approach progressively uncovers the target object and ultimately grasps it with a few steps and a high success rate. In both simulated and real experiments, ThinkGrasp achieved a high success rate and significantly outperformed state-of-the-art methods in heavily cluttered environments or with diverse unseen objects, demonstrating strong generalization capabilities.
Keywords: Robotic Grasping, Vision-Language Models, Manipulation, Robotics, GPT-4o, Cluttered Environments.
arXiv · Project page · Code · Video
Imagination Policy: Using Generative Point Cloud Models for Learning Manipulation Policies
Haojie Huang, Karl Schmeckpeper, Dian Wang, Ondrej Biza, Yaoyao Qian, Haotian Liu, Mingxi Jia, Robert Platt, Robin Walters. “Imagination Policy: Using Generative Point Cloud Models for Learning Manipulation Policies.” CoRL 2024.
Humans can imagine goal states during planning and perform actions to match those goals. In this work, we propose Imagination Policy, a novel multi-task key-frame policy network for solving high-precision pick and place tasks. Instead of learning actions directly, Imagination Policy generates point clouds to imagine desired states which are then translated to actions using rigid action estimation. This transforms action inference into a local generative task. We leverage pick and place symmetries underlying the tasks in the generation process and achieve extremely high sample efficiency and generalizability to unseen configurations. Finally, we demonstrate state-of-the-art performance across various tasks on the RLBench benchmark compared with several strong baselines.
Keywords: Point Clouds, Generative Models, Manipulation, Robotics, Policy Learning.
Project page
Analysis of football game performance based on social network
Yaoyao Qian, Xianming Wang. “Analysis of football game performance based on social network.” ICAID 2022.
With the advent of big data and network era, the interpersonal relationship is getting closer and closer, and the advantage of teamwork is becoming more and more prominent. By analyzing the team competition rules, combining the team members' abilities, characteristics, and interactions between team members in previous competitions, team cooperation, and coordination strategy can be further optimized effectively. Based on the data samples recorded in 38 football games, this paper takes the 14th huskies vs. O14 game with a 4-0 result as an example. After data preprocessing, DBSCAN, a clustering algorithm, is introduced to simplify the data set by presenting a scatter graph. Each player's resident coordinates in each game are obtained, which provides data for social network analysis. We used a Python program to get the overall network analysis. The D6, M1, and F2 players on the Huskies form a triad configuration, the players D3 and D4, the players D1 and M2 in the O14 team form multiple dyadic configurations. To better understand the characteristics of the passing network, the individual network analysis is carried out. Through the study of particular net research degree center degree, it can be found that the Huskies coach's tactical arrangement was 4-4-2 formation, which was changed into 4-3-3 attacking formation in an actual combat operation. The tactical arrangement of the O14 coach was the same as the essential combat operation, which was 2-3-5 formation and more defensive. The in-depth analysis found that the Huskies' core was a midfielder, while the O14's core was not a midfielder. The tempo of the game and the fluidity of the passing network of the O14 team were significantly reduced, which was also the main factor leading to the failure.
Keywords: Football Team Performance Evaluation, Social Network Analysis, Betweenness, Clustering Algorithm, DBSCAN, Sports Analytics.
Paper