On the job market · Ph.D. expected May 2027

Xiangru (Edward) Jian

Ph.D. student in Computer Science · Data Systems Group · University of Waterloo

I am a Ph.D. student in the Data Systems Group at the Cheriton School of Computer Science, University of Waterloo, advised by Prof. M. Tamer Özsu. I work on multimodal and agentic data management: natural-language interfaces to databases, complex analytical tasks over heterogeneous federated systems, computer-use agents, and multimodal learning.

I was a research intern at Salesforce AI Research in 2026 and a visiting researcher at ServiceNow Research from 2024 to 2025. Before Waterloo, I received my M.Sc. in Data Science from City University of Hong Kong, advised by Prof. Yu Yang, working on data mining and graph learning. Off the clock, I follow Manchester City closely.

Portrait of Xiangru Jian
Waterloo, Canada
psql · xiangru@uwaterloo

On the job market, 2026–27

xiangru=# SELECT position, status
xiangru-#   FROM job_market
xiangru-#  WHERE available_from = 'May 2027';
positionstatus
Tenure-track Assistant Professoropen
Postdoctoral Researcheropen
Industry Research Scientistopen

(3 rows)

I expect to complete my Ph.D. in May 2027 and would be glad to talk about faculty, postdoc, and research scientist roles. Email me for the most recent version of my CV.

Research

I build systems that let LLM agents answer hard analytical questions over data that is scattered, heterogeneous, and multimodal: relational and graph databases, documents, videos, and the desktop software where much of that data actually lives.

Thesis (tentative) A Multi-Agent Architecture for Complex Analytical Tasks over Heterogeneous Federated Systems

500K+
Hugging Face downloads of the UI-Vision / GroundCUA / CUA-Suite series
4.2k+ ★
GitHub stars across the code releases of my papers, led by Paper2Poster
12
Published papers as first, co-first, or corresponding author (4 as corresponding author)
3 Orals
EMNLP 2023 · IEEE BigData 2023 · ICML 2025 MAS Workshop

News

  1. Sep 2026 paper

    RA-SQL: Relational Algebra as Deterministic Chain-of-Thought for Natural Language to SQL accepted by SIGMOD 2027.

  2. Sep 2026 paper

    TRL-Bench, a cross-paradigm benchmark for tabular encoders, accepted by the NeurIPS 2026 Datasets & Benchmarks Track. [project]

  3. Sep 2026 paper

    FORGE, fine-grained multimodal evaluation for manufacturing, accepted by Findings of EMNLP 2026. [project]

  4. Sep 2026 talk

    Talk at Peking University: Answering Complex Analytical Questions over Heterogeneous Federated Systems.

  5. Aug 2026 career

    Wrapped up my research internship at Salesforce AI Research (Singapore) on computer-use and MCP tool-use agents, mentored by Junnan Li.

  6. Aug 2026 release

    Preprints released: MCP-Universe RL (RL for MCP tool-use agents, [project]) and TEXAS (task-expert-aware supervision for MoE LLM adaptation).

  7. Jul 2026 release

    Preprint released: StateAct, program state before pixels for long-horizon computer-use agents.

  8. 2026 paper

    MolEmb appears at the ICML 2026 Workshop on Multi-modal Foundation Models and LLMs for Life Sciences; CUA-Suite appears at the ICLR 2026 Workshop on Lifelong Agents.

  9. 2026 paper

    The Duality between Large Language Models and Data Management Systems published in IEEE Data Engineering Bulletin 50(1).

  10. Apr 2026 paper

    InteracSPARQL accepted by Findings of ACL 2026.

  11. Mar 2026 release

    CUA-Suite released and ranked #1 Paper of the Day on Hugging Face. It unifies UI-Vision, GroundCUA and VideoCUA (55 hours of expert video across 87 professional apps); the series has 500K+ downloads on Hugging Face.

  12. Mar 2026 talk

    Invited talk at Tencent Hunyuan: Recent Efforts on CUAs' Data, Training, Evaluation, and (Potentially) Beyond.

  13. Feb 2026 paper

    Spatio-temporal traffic accident detection via graph-based GAN published in Engineering Applications of Artificial Intelligence.

  14. Jan 2026 paper

    Two papers accepted by ICLR 2026: GraphOmni and GroundCUA (Grounding Computer Use Agents on Human Demonstrations).

  15. Jan 2026 talk

    Invited talks at Snowflake (Dedicated Multi-Agent System over Heterogeneous Data Lakes) and LlamaIndex (Multimodal Learning on Documents and Beyond).

  16. Dec 2025 paper

    LazyVLM, a neuro-symbolic approach to video analytics, accepted by ICDE 2026 (Demo).

  17. Dec 2025 talk

    Talk at ONDBD 2025: An Interactive Tool for SPARQL Query Refinement Using Natural Language Explanations.

  18. Sep 2025 paper

    Three papers accepted by NeurIPS 2025: Paper2Poster (3.9k ★ on GitHub), vision models for graph structural understanding, and AlignVLM.

  19. May 2025 paper

    UI-Vision, a desktop-centric GUI benchmark, accepted by ICML 2025.

  20. Jan 2025 paper

    BigDocs accepted by ICLR 2025; spectral augmentation for graph SSL accepted by TMLR.

  21. Apr 2024 career

    Started a visiting researcher internship at ServiceNow Research (Montreal), working on multimodal learning and GUI understanding (through May 2025).

  22. Oct 2023 paper

    Two papers accepted by EMNLP 2023: Balance Act (Oral) and InvGC (Findings).

  23. Aug 2023 paper

    Paper on loss landscapes of PDE-solving networks accepted by IEEE BigData 2023 (Oral).

  24. Jul 2023 paper

    Decentralized online DR-submodular maximization accepted by CIKM 2023.

Publications

Google Scholar

* equal contribution† project lead✉︎ corresponding author

  1. SIGMOD 2027

    RA-SQL: Relational Algebra as Deterministic Chain-of-Thought for Natural Language to SQL

    Xiangru Jian, Wei Pang, Chao Zhang, Xi He, M. Tamer Özsu

    ACM SIGMOD International Conference on Management of Data (SIGMOD)

    Serializes queries as relational-algebra operator trees and uses them as a deterministic chain of thought, so the model reasons in RA before it writes SQL.

  2. NeurIPS 2026

    TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

    Wei Pang*, Xiangru Jian*, Hehan Li*, Zhixuan Yu*, Alex Xue*

    NeurIPS 2026 Datasets and Benchmarks Track · * core contributors

    Probes row, column, and table embeddings from 20 tabular encoders across 16 tasks and 87 datasets, and finds that encoder quality is capability-specific rather than a single leaderboard.

  3. EMNLP F 2026

    FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios

    Xiangru Jian*, Hao Xu*✉︎, Wei Pang*, Xinjian Zhao

    Findings of EMNLP 2026

    A benchmark of real 2D images and 3D point clouds annotated down to exact model numbers; across 18 MLLMs, missing domain knowledge, not visual grounding, is the main bottleneck.

  4. arXiv 2026

    MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning

    Ziyang Luo, Yan Yang, Xiangru Jian

    arXiv preprint

  5. arXiv 2026

    StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents

    Yan Yang, Xiangru Jian, Ziyang Luo

    arXiv preprint

    A code-first multi-agent harness that grounds computer-use agents in program state and calls a GUI agent only when needed, raising OSWorld 2.0 binary success from 20.6% to 26.9% (Claude Opus 4.8) at about 9× lower cost per task.

  6. arXiv 2026

    TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation

    Guanzhi Deng, Haibo Wang, Kuan Wu, Xiangru Jian

    arXiv preprint

  7. ICLR-W 2026

    CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents

    Xiangru Jian*✉︎, Shravan Nayak*, Kevin Qinghong Lin, Aarash Feizi, Kaixin Li, Patrice Bechard

    arXiv preprint · short version at the ICLR 2026 Workshop on Lifelong Agents

    About 10K expert desktop tasks across 87 professional apps (55 hours of 30 fps video with cursor traces and reasoning), unified with UI-Vision and GroundCUA into one ecosystem for training and evaluating CUAs.

  8. ICML-W 2026

    MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models

    Xinjian Zhao, Xiangru Jian, Yaoyao Xu, Xiaozhuang Song, Wei Pang, Lei Bai, Tianshu Yu

    ICML 2026 Workshop on Multi-modal Foundation Models and LLMs for Life Sciences

  9. IEEE DEB 2026

    The Duality between Large Language Models and Data Management Systems

    M. Tamer Özsu, Kerem Akillioglu, Xiangru Jian

    IEEE Data Engineering Bulletin, 50(1):40–58

    A two-way view of LLMs for databases and databases for LLMs, from LLM-synthesized database engines to data systems that become agentic data environments.

  10. ACL F 2026

    InteracSPARQL: An Interactive System for SPARQL Query Refinement Using Natural Language Explanations

    Xiangru Jian, Zhengyuan Dong, M. Tamer Özsu

    Findings of ACL 2026

    A training-free refinement loop that explains generated SPARQL in natural language and lets users and LLMs fix it, improving accuracy on QALD-9 and QALD-10.

  11. ICLR 2026

    GraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic Tasks

    Hao Xu*, Xiangru Jian*†✉︎, Xinjian Zhao*, Wei Pang*

    International Conference on Learning Representations (ICLR)

    Benchmarks LLM reasoning on graph problems across task, graph type, serialization, and prompt scheme, with an RL-inspired selector that finds strong configurations cheaply.

  12. ICLR 2026

    Grounding Computer Use Agents on Human Demonstrations

    Aarash Feizi*, Shravan Nayak*, Xiangru Jian, Kevin Qinghong Lin

    International Conference on Learning Representations (ICLR)

    GroundCUA: 3.56M human-verified UI element annotations across 87 desktop apps. GroundNext models trained on it reach state of the art at the 3B and 7B scales on five grounding benchmarks with a tenth of the usual data.

  13. ICDE 2026

    LazyVLM: Neuro-Symbolic Approach to Video Analytics

    Xiangru Jian*†✉︎, Wei Pang*, Zhengyuan Dong*, Chao Zhang*, M. Tamer Özsu

    IEEE International Conference on Data Engineering (ICDE), Demo Track

    Decomposes multi-frame video queries into fine-grained operations and pushes most of the work into relational query execution and vector search, so video analytics scales beyond what VLMs alone can do.

  14. EAAI 2026

    Spatio-temporal Traffic Accidents Detection via Graph-based Generative Adversarial Network

    Lyuyi Zhu, Qixin Zhang, Xiangru Jian, Yu Yang, Lishuai Li

    Engineering Applications of Artificial Intelligence, vol. 165

  15. arXiv 2026

    When Vision Meets Graphs: A Survey on Graph Reasoning and Learning

    Xinjian Zhao, Wei Pang, Zhixuan Yu, Xiangru Jian

    arXiv preprint

  16. NeurIPS 2025

    Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers

    Wei Pang*, Kevin Qinghong Lin*✉︎, Xiangru Jian*, Xi He✉︎, Philip Torr

    NeurIPS 2025 Datasets and Benchmarks Track · Oral at the ICML 2025 MAS Workshop

    The first benchmark and metric suite for academic poster generation, plus PosterAgent, a visual-in-the-loop multi-agent pipeline that turns papers into editable posters.

  17. NeurIPS 2025

    The Underappreciated Power of Vision Models for Graph Structural Understanding

    Xinjian Zhao*, Wei Pang*, Zhongkai Xue*, Xiangru Jian*

    Advances in Neural Information Processing Systems (NeurIPS)

  18. NeurIPS 2025

    AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding

    Ahmed Masry, Juan A. Rodriguez, Tianyu Zhang, Suyuchen Wang, Chao Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian

    Advances in Neural Information Processing Systems (NeurIPS)

  19. ICML 2025

    UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction

    Shravan Nayak*✉︎, Xiangru Jian*✉︎, Kevin Qinghong Lin, Juan A. Rodriguez

    International Conference on Machine Learning (ICML)

    A license-permissive desktop GUI benchmark over 83 applications with dense human annotations for element grounding, layout grounding, and action prediction.

  20. ICLR 2025

    BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks

    Juan A. Rodriguez*, Xiangru Jian*†, Siba Smarak Panigrahi, Tianyu Zhang

    International Conference on Learning Representations (ICLR)

    A 7.5M-sample, license-permissive dataset across 30 document and code tasks (e.g., screenshot-to-HTML, table-to-LaTeX), released with the BigDocs-Bench benchmark.

  21. TMLR 2025

    Rethinking Spectral Augmentation for Contrast-based Graph Self-Supervised Learning

    Xiangru Jian*, Xinjian Zhao*, Wei Pang*, Chaolong Ying, Yimu Wang, Yaoyao Xu, Tianshu Yu

    Transactions on Machine Learning Research (TMLR)

  22. NAACL 2025

    DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models

    Yimu Wang, Shuai Yuan, Bo Xue, Xiangru Jian

    NAACL 2025 (Long Papers)

  23. EAAI 2025

    Graph Convolutional Network for Traffic Incidents Duration Classification

    Lyuyi Zhu, Qixin Zhang, Xiangru Jian, Yu Yang

    Engineering Applications of Artificial Intelligence, vol. 151

  24. arXiv 2025

    Scope: Selective Cross-modal Orchestration of Visual Perception Experts

    Tianyu Zhang, Suyuchen Wang, Chao Wang, Juan Rodriguez, Ahmed Masry, Xiangru Jian

    arXiv preprint

  25. EMNLP 2023

    Balance Act: Mitigating Hubness in Cross-Modal Retrieval with Query and Gallery Banks

    Yimu Wang, Xiangru Jian, Bo Xue

    Conference on Empirical Methods in Natural Language Processing (EMNLP)

  26. EMNLP F 2023

    InvGC: Robust Cross-Modal Retrieval by Inverse Graph Convolution

    Xiangru Jian, Yimu Wang

    Findings of EMNLP 2023

  27. BigData 2023

    Roughness Index for Loss Landscapes of Neural Network Models of Partial Differential Equations

    Keke Wu, Xiangru Jian, Rui Du, Jingrun Chen, Xiang Zhou

    IEEE International Conference on Big Data (BigData)

  28. CIKM 2023

    Communication-Efficient Decentralized Online Continuous DR-Submodular Maximization

    Qixin Zhang, Zengde Deng, Xiangru Jian, Zaiyi Chen, Haoyuan Hu, Yu Yang

    ACM International Conference on Information and Knowledge Management (CIKM)

  29. Preprint

    A Pattern-based Subset Choice Model

    Yu Yin*, Xiangru Jian*, Haoxiang Liu, Wei Gong, Yu Yang

    Under review

Datasets & benchmarks

10 total

Much of my work builds the data that agents and multimodal models learn from and are measured on. Every number below comes from the paper or the official release.

Computer-use agents

CUA-Suite

ICLR 2026 Workshop on Lifelong AgentsCo-first author · Corresponding

A unified ecosystem of expert video demonstrations and dense annotations for training and evaluating desktop computer-use agents across 87 professional applications. The series has 500K+ downloads on Hugging Face.

87
open-source desktop apps in 12 categories
~10K
expert-designed tasks
500K+
Hugging Face downloads (series)

VideoCUA

CUA-Suite · 2026Co-first author · Corresponding

The largest open expert video corpus for desktop computer use, with cursor traces and multi-layered reasoning annotations.

~55 h
continuous 30 fps expert video
6M
frames
~10K
human-demonstrated tasks
87
desktop apps

GroundCUA

ICLR 2026

Pixel-precise desktop grounding data built entirely by human annotators; it trains the GroundNext 3B/7B models, state of the art at their scale on five grounding benchmarks.

3.56M+
human-verified UI element annotations
56K
annotated screenshots
700K
instruction-tuning samples for GroundNext
87
desktop apps

UI-Vision

ICML 2025Co-first author · Corresponding

A license-permissive desktop benchmark for element grounding, layout grounding, and action prediction.

450
human demonstrations
8.2K+
query–label pairs
83
apps across 6 domains
3
evaluation tasks
Documents & code

BigDocs

ICLR 2025Co-first author · Project lead

An open, license-permissive dataset for document understanding and code generation, plus BigDocs-Bench, 10 new tasks such as Screenshot2HTML, Table2LaTeX, and Image2SVG.

7.5M
training image–text pairs
4M
unique training images
30
tasks
10
new benchmark tasks (BigDocs-Bench)
Documents & agents

Paper2Poster

NeurIPS 2025 Datasets & BenchmarksCo-first author

The first benchmark and metric suite for academic poster generation, with PosterAgent, a multi-agent pipeline that turns papers into editable posters. Oral at the ICML 2025 MAS Workshop.

100
paper–poster pairs from NeurIPS, ICML, and ICLR
4
evaluation dimensions, including PaperQuiz
87%
fewer tokens than GPT-4o multi-agent systems (open-source PosterAgent)
3.9k ★
GitHub stars
Graphs & LLMs

GraphOmni

ICLR 2026Co-first author · Project lead · Corresponding

A benchmark of LLM reasoning on graph-theoretic problems that varies graph type, serialization format, and prompt scheme across three difficulty levels.

241K
question–answer samples
6
graph-theoretic tasks
7 × 7
graph types × serialization formats
9
prompt schemes
Graphs & vision

GraphAbstract

NeurIPS 2025Co-first author

A benchmark of holistic graph-structure perception (archetypes, symmetry, connectivity strength, critical elements), tested on progressively larger out-of-distribution graphs.

4
structural-understanding tasks
3
test scales: in-distribution, near-OOD, far-OOD
150
nodes per graph at the far-OOD scale (max)
Tables

TRL-Bench

NeurIPS 2026 Datasets & BenchmarksCore contributor

A representation-level benchmark that probes frozen row, column, and table embeddings from tabular encoders with shared lightweight heads, across three suites: TRL-CTbench, TRL-Rbench, and TRL-DLTE.

20
tabular encoders evaluated
16
tasks in three suites
87
datasets
47,772
tables in the TRL-DLTE data lake
Manufacturing

FORGE

Findings of EMNLP 2026Co-first author

A manufacturing benchmark of real 2D images and 3D point clouds annotated down to exact model numbers, covering workpiece verification, surface inspection, and assembly verification.

3,320
evaluation cases
90
distinct model numbers
14
workpiece categories
18
MLLMs evaluated
Molecules

MolCAR

MolEmb · ICML 2026 Workshop (FM4LS)

A molecular context-aware retrieval benchmark from the MolEmb paper: it tests whether task instructions let molecular embeddings retrieve the right context when molecule identity alone is not enough.

4,092
molecule–context evaluation pairs
10,844
training pairs (MolCAR-Train)
1,653
held-out molecules
8
property-prediction task families

Experience

Research positions

  1. Salesforce AI Research May – Aug 2026

    Research Intern · Singapore

    Computer-use and MCP tool-use agents. Contributed to StateAct and MCP-Universe RL; responsible for data development and reinforcement-learning training of the internal SFR-CUA model. Mentor: Junnan Li

  2. ServiceNow Research Apr 2024 – May 2025

    Visiting Researcher (Internship) · Montreal, Canada

    Multimodal learning and GUI understanding: BigDocs, UI-Vision, GroundCUA, and CUA-Suite. Mentor: Sai Rajeswar

  3. Hong Kong Institute of Data Science, CityU Jul 2021 – Jun 2022

    Research Assistant · Hong Kong SAR

    Order prediction for industrial partners including Alibaba and HKTV Mall.

  4. City University of Hong Kong Jun – Aug 2021

    Research Assistant · Hong Kong SAR

    Loss landscapes of deep neural networks.

Education

  1. Ph.D., Computer Science

    2022 – 2027 (expected)

    Advisor: M. Tamer Özsu

  2. M.Sc., Data Science

    2020 – 2021

    Advisor: Yu Yang · Graduated with distinction, ranked 1st

  3. B.Eng., Civil Engineering; minor in Mathematics

    2014 – 2019

    Outstanding Graduate (top 5%)

Talks

  1. Sep 2026
    Peking University Answering Complex Analytical Questions over Heterogeneous Federated Systems
  2. Mar 2026
    Tencent Hunyuan Recent Efforts on CUAs' Data, Training, Evaluation, and (Potentially) Beyond
  3. Jan 2026
    Snowflake Dedicated Multi-Agent System over Heterogeneous Data Lakes
  4. Jan 2026
    LlamaIndex Multimodal Learning on Documents and Beyond
  5. Dec 2025
    ONDBD 2025 An Interactive Tool for SPARQL Query Refinement Using Natural Language Explanations slides

Teaching

  • CS 338 Computer Applications in Business: DatabasesTeaching assistant · Spring & Fall 2023
  • CS 106 Introduction to Computer Science 2Teaching assistant · Winter 2023
  • CS 114 Principles of Computing for ScienceTeaching assistant · Fall 2022

Service

  • Area ChairACL ARR
  • ReviewerICLR · ICML · NeurIPS · AAAI · ECCV · ACL ARR · TMLR · TKDD · IEEE TBD · ACI Materials Journal
  • External reviewerVLDB 2023, 2024 · SIGMOD 2024 · CIKM 2022, 2023 · SIGIR 2022 · KDD 2022

Honors

  • Graduated with distinction, ranked 1st in the M.Sc. programCity University of Hong Kong · 2021
  • Outstanding Graduate (highest honor, top 5%)Tongji University · 2019
  • National Scholarship (top 2%)2015