Undergraduate Researcher · University of Wisconsin–Madison

Zelong Xu

I study how to evaluate AI agents, align them with what we intend, and understand how they reason.

Applying to CS PhD programs for Fall 2027

Fig. 1Gradient descent on a drifting loss landscape.

About me

Zelong Xu

I am an undergraduate at the University of Wisconsin–Madison, where I study Statistics with certificates in Computer Science and Mathematics (B.S. expected January 2027), and work as an undergraduate researcher on machine learning and AI safety.

My research asks a simple question with difficult answers: how do we know whether an AI system is actually doing what we want? I approach it from three directions — building evaluations that faithfully measure what agents can and cannot do, developing alignment methods that are robust and data-efficient, and studying how models reason internally so that their failures can be predicted and corrected.

Most recently I contributed to DOG-DPO, a training-free data-selection framework for safety alignment that treats preference pairs as geometric signals in a model’s representation space (Findings of EMNLP 2026).

Before Madison I studied Finance at Shandong University. I am applying to CS PhD programs for Fall 2027 — if my interests overlap with yours, I would love to hear from you.

Research interests

  1. AI agent evaluation

    Benchmarks tell us whether an agent succeeded; they rarely tell us why, or whether the result will hold outside the benchmark. I am interested in evaluations that are faithful to real tasks, hard to game, and diagnostic, so that a score actually predicts behavior.

  2. Alignment

    Alignment methods are only as good as the data and objectives behind them. I am interested in making preference-based alignment more robust and data-efficient, and in understanding when and why safety training fails to generalize.

  3. Reasoning & interpretability

    When a model “reasons”, what is actually happening inside it? I am interested in connecting internal representations to reasoning behavior, and in using that understanding to make models more reliable and easier to evaluate.

Publications

2026

  1. DOG-DPO: Dynamic Optimization in Geometry for Safety Alignment

    Yi Nian, Tiankai Yang, Yudi Zhang, Qi Pan, Zelong Xu, Shenzhe Zhu, Qingqing Luan, Yue Huang, Xiangliang Zhang, Yue Zhao

    Findings of EMNLP 2026 arXiv:2606.07678

    arXivPDF
    Abstract

    Safety alignment for large language models relies on preference data, but current pipelines often train on large, redundant datasets. Existing data selection methods typically score each preference pair independently, collapsing directional preference information into scalar quality or diversity scores. This sample-centric view is especially limiting in multi-dataset settings, where shared safety directions coexist with dataset-specific residual risks. We propose DOG-DPO, a training-free data selection framework that treats preference pairs as structured geometric signals. DOG-DPO first represents each preference pair as a direction in model representation space. It then decomposes multi-dataset preference geometry into a global anchor subspace and dataset-specific residual subspaces. Finally, it selects subsets by maximizing diversity-based coverage, encouraging broad, non-redundant coverage of alignment directions before DPO training. Across six safety benchmarks and two model backbones, DOG-DPO achieves a strong utility-robustness trade-off using only 11% of the preference pairs. It recovers most of the safety gains of full-data training while remaining entirely teacher-free, training-free, and substantially faster than representative selection baselines.

    BibTeX
    @inproceedings{nian2026dog,
      title        = {DOG-DPO: Dynamic Optimization in Geometry for Safety Alignment},
      author       = {Yi Nian and Tiankai Yang and Yudi Zhang and Qi Pan and Zelong Xu and Shenzhe Zhu and Qingqing Luan and Yue Huang and Xiangliang Zhang and Yue Zhao},
      booktitle    = {Findings of EMNLP 2026},
      year         = {2026},
      url          = {https://arxiv.org/abs/2606.07678},
    }

Education

  1. Sep 2024 — Jan 2027 (expected)

    University of Wisconsin–Madison

    B.S. in Statistics

    Certificates in Computer Science and Mathematics

    GPA 3.90 / 4.00

  2. Sep 2022 — Jun 2024

    Shandong University

    Finance

    GPA 88.1 / 100

Get in touch

The best way to reach me is email. I am always happy to talk about research, PhD opportunities, or anything on this page.