I am a Ph.D. student in the Data Systems Group at the Cheriton School of Computer Science, University of Waterloo, advised by Prof. M. Tamer Özsu. I work on multimodal and agentic data management: natural-language interfaces to databases, complex analytical tasks over heterogeneous federated systems, computer-use agents, and multimodal learning.
I was a research intern at Salesforce AI Research in 2026 and a visiting researcher at ServiceNow Research from 2024 to 2025. Before Waterloo, I received my M.Sc. in Data Science from City University of Hong Kong, advised by Prof. Yu Yang, working on data mining and graph learning. Off the clock, I follow Manchester City closely.
xiangru=#SELECT position, status
xiangru-#FROM job_market
xiangru-#WHERE available_from = 'May 2027';
position
status
Tenure-track Assistant Professor
open
Postdoctoral Researcher
open
Industry Research Scientist
open
(3 rows)
xiangru=#
I expect to complete my Ph.D. in May 2027 and would be glad to talk about faculty, postdoc, and research scientist roles. Email me for the most recent version of my CV.
I build systems that let LLM agents answer hard analytical questions over data that is scattered, heterogeneous, and multimodal: relational and graph databases, documents, videos, and the desktop software where much of that data actually lives.
Thesis (tentative)A Multi-Agent Architecture for Complex Analytical Tasks over Heterogeneous Federated Systems
Agentic data management
Natural-language interfaces to relational and graph databases, multi-agent analytics over federated data lakes, and the interplay between LLMs and data management systems.
RA-SQL: Relational Algebra as Deterministic Chain-of-Thought for Natural Language to SQL accepted by SIGMOD 2027.
Sep 2026paper
TRL-Bench, a cross-paradigm benchmark for tabular encoders, accepted by the NeurIPS 2026 Datasets & Benchmarks Track. [project]
Sep 2026paper
FORGE, fine-grained multimodal evaluation for manufacturing, accepted by Findings of EMNLP 2026. [project]
Sep 2026talk
Talk at Peking University: Answering Complex Analytical Questions over Heterogeneous Federated Systems.
Aug 2026career
Wrapped up my research internship at Salesforce AI Research (Singapore) on computer-use and MCP tool-use agents, mentored by Junnan Li.
Aug 2026release
Preprints released: MCP-Universe RL (RL for MCP tool-use agents, [project]) and TEXAS (task-expert-aware supervision for MoE LLM adaptation).
Jul 2026release
Preprint released: StateAct, program state before pixels for long-horizon computer-use agents.
2026paper
MolEmb appears at the ICML 2026 Workshop on Multi-modal Foundation Models and LLMs for Life Sciences; CUA-Suite appears at the ICLR 2026 Workshop on Lifelong Agents.
CUA-Suite released and ranked #1 Paper of the Day on Hugging Face. It unifies UI-Vision, GroundCUA and VideoCUA (55 hours of expert video across 87 professional apps); the series has 500K+ downloads on Hugging Face.
Mar 2026talk
Invited talk at Tencent Hunyuan: Recent Efforts on CUAs' Data, Training, Evaluation, and (Potentially) Beyond.
RA-SQL: Relational Algebra as Deterministic Chain-of-Thought for Natural Language to SQL
Xiangru Jian, Wei Pang, Chao Zhang, Xi He, M. Tamer Özsu
ACM SIGMOD International Conference on Management of Data (SIGMOD)
Serializes queries as relational-algebra operator trees and uses them as a deterministic chain of thought, so the model reasons in RA before it writes SQL.
NeurIPS 2026 Datasets and Benchmarks Track · * core contributors
Probes row, column, and table embeddings from 20 tabular encoders across 16 tasks and 87 datasets, and finds that encoder quality is capability-specific rather than a single leaderboard.
Xiangru Jian*, Hao Xu*✉︎, Wei Pang*, Xinjian Zhao, Chengyu Tao, Qixin Zhang, Xikun Zhang✉︎, Chao Zhang, Guanzhi Deng, Alex Xue, Juan Du, Tianshu Yu, Garth Tarr, Linqi Song, Qiuzhuang Sun, Dacheng Tao
Findings of EMNLP 2026
A benchmark of real 2D images and 3D point clouds annotated down to exact model numbers; across 18 MLLMs, missing domain knowledge, not visual grounding, is the main bottleneck.
Yan Yang, Xiangru Jian, Ziyang Luo, Zirui Zhao, Yutong Dai, Ziji Shi, Hanshu Yan, Jun Hao Liew, Silvio Savarese, Junnan Li
arXiv preprint
A code-first multi-agent harness that grounds computer-use agents in program state and calls a GUI agent only when needed, raising OSWorld 2.0 binary success from 20.6% to 26.9% (Claude Opus 4.8) at about 9× lower cost per task.
Xiangru Jian*✉︎, Shravan Nayak*, Kevin Qinghong Lin, Aarash Feizi, Kaixin Li, Patrice Bechard, Spandana Gella, Sai Rajeswar
arXiv preprint · short version at the ICLR 2026 Workshop on Lifelong Agents
About 10K expert desktop tasks across 87 professional apps (55 hours of 30 fps video with cursor traces and reasoning), unified with UI-Vision and GroundCUA into one ecosystem for training and evaluating CUAs.
A two-way view of LLMs for databases and databases for LLMs, from LLM-synthesized database engines to data systems that become agentic data environments.
A training-free refinement loop that explains generated SPARQL in natural language and lets users and LLMs fix it, improving accuracy on QALD-9 and QALD-10.
International Conference on Learning Representations (ICLR)
Benchmarks LLM reasoning on graph problems across task, graph type, serialization, and prompt scheme, with an RL-inspired selector that finds strong configurations cheaply.
Aarash Feizi*, Shravan Nayak*, Xiangru Jian, Kevin Qinghong Lin, Kaixin Li, Rabiul Awal, Xing Han Lù, Johan Obando-Ceron, Juan A. Rodriguez, Nicolas Chapados, David Vazquez, Adriana Romero-Soriano, Reihaneh Rabbany, Perouz Taslakian, Christopher Pal, Spandana Gella, Sai Rajeswar
International Conference on Learning Representations (ICLR)
GroundCUA: 3.56M human-verified UI element annotations across 87 desktop apps. GroundNext models trained on it reach state of the art at the 3B and 7B scales on five grounding benchmarks with a tenth of the usual data.
IEEE International Conference on Data Engineering (ICDE), Demo Track
Decomposes multi-frame video queries into fine-grained operations and pushes most of the work into relational query execution and vector search, so video analytics scales beyond what VLMs alone can do.
Wei Pang*, Kevin Qinghong Lin*✉︎, Xiangru Jian*, Xi He✉︎, Philip Torr
NeurIPS 2025 Datasets and Benchmarks Track · Oral at the ICML 2025 MAS Workshop
The first benchmark and metric suite for academic poster generation, plus PosterAgent, a visual-in-the-loop multi-agent pipeline that turns papers into editable posters.
Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu, Nicolas Chapados, Yoshua Bengio, Enamul Hoque, Christopher Pal, Issam H. Laradji, David Vazquez, Perouz Taslakian, Spandana Gella, Sai Rajeswar
Advances in Neural Information Processing Systems (NeurIPS)
Shravan Nayak*✉︎, Xiangru Jian*✉︎, Kevin Qinghong Lin, Juan A. Rodriguez, Montek Kalsi, Rabiul Awal, Nicolas Chapados, M. Tamer Özsu, Aishwarya Agrawal, David Vazquez, Christopher Pal, Perouz Taslakian, Spandana Gella, Sai Rajeswar
International Conference on Machine Learning (ICML)
A license-permissive desktop GUI benchmark over 83 applications with dense human annotations for element grounding, layout grounding, and action prediction.
Juan A. Rodriguez*, Xiangru Jian*†, Siba Smarak Panigrahi, Tianyu Zhang, Aarash Feizi, Abhay Puri, Akshay Kalkunte, François Savard, Ahmed Masry, Shravan Nayak, Rabiul Awal, Mahsa Massoud, Amirhossein Abaskohi, Zichao Li, Suyuchen Wang, Pierre-André Noël, Chao Wang, Mats Leon Richter, Saverio Vadacchino, Shubham Agarwal, Sanket Biswas, Sara Shanian, Ying Zhang, Noah Bolger, Kurt MacDonald, Simon Fauvel, Sathwik Tejaswi, Srinivas Sunkara, Joao Monteiro, Krishnamurthy DJ Dvijotham, Torsten Scholak, Nicolas Chapados, Sepideh Kharagani, Sean Hughes, M. Tamer Özsu, Siva Reddy, Marco Pedersoli, Yoshua Bengio, Christopher Pal, Issam Laradji, Spandana Gella, Perouz Taslakian, David Vazquez, Sai Rajeswar
International Conference on Learning Representations (ICLR)
A 7.5M-sample, license-permissive dataset across 30 document and code tasks (e.g., screenshot-to-HTML, table-to-LaTeX), released with the BigDocs-Bench benchmark.
Yu Yin*, Xiangru Jian*, Haoxiang Liu, Wei Gong, Yu Yang
Under review
No papers in this view. Try “All”.
Datasets & benchmarks
10 total
Much of my work builds the data that agents and multimodal models learn from and are measured on. Every number below comes from the paper or the official release.
ICLR 2026 Workshop on Lifelong AgentsCo-first author · Corresponding
A unified ecosystem of expert video demonstrations and dense annotations for training and evaluating desktop computer-use agents across 87 professional applications. The series has 500K+ downloads on Hugging Face.
Pixel-precise desktop grounding data built entirely by human annotators; it trains the GroundNext 3B/7B models, state of the art at their scale on five grounding benchmarks.
An open, license-permissive dataset for document understanding and code generation, plus BigDocs-Bench, 10 new tasks such as Screenshot2HTML, Table2LaTeX, and Image2SVG.
The first benchmark and metric suite for academic poster generation, with PosterAgent, a multi-agent pipeline that turns papers into editable posters. Oral at the ICML 2025 MAS Workshop.
100
paper–poster pairs from NeurIPS, ICML, and ICLR
4
evaluation dimensions, including PaperQuiz
87%
fewer tokens than GPT-4o multi-agent systems (open-source PosterAgent)
ICLR 2026Co-first author · Project lead · Corresponding
A benchmark of LLM reasoning on graph-theoretic problems that varies graph type, serialization format, and prompt scheme across three difficulty levels.
A representation-level benchmark that probes frozen row, column, and table embeddings from tabular encoders with shared lightweight heads, across three suites: TRL-CTbench, TRL-Rbench, and TRL-DLTE.
A manufacturing benchmark of real 2D images and 3D point clouds annotated down to exact model numbers, covering workpiece verification, surface inspection, and assembly verification.
A molecular context-aware retrieval benchmark from the MolEmb paper: it tests whether task instructions let molecular embeddings retrieve the right context when molecule identity alone is not enough.
Computer-use and MCP tool-use agents. Contributed to StateAct and MCP-Universe RL; responsible for data development and reinforcement-learning training of the internal SFR-CUA model. Mentor: Junnan Li