Hi, I am a 4th year PhD candidate at Machine Learning Department, Carnegie Mellon University. I am advised by Aarti Singh and Drew Bagnell. I am also closely working with Wen Sun. I am partially supported by Two Sigma PhD fellowship.

In the past, I have interned at FAIR Paris (with Remi Munos), Amazon NYC (with Udaya Ghai and Dean Foster), and Microsoft Research NYC (with Akshay Krishnamurthy and Dylan Foster). I finished my master degree in MLD, advised by Kris Kitani. I completed my undergraduate at UC San Diego with CS and Math majors and I was advised by Sicun Gao.

Research

I study interactive decision making via reinforcement learning, where agents act and learn in environments with complex dynamics and rich observations. My goal is to make such agents learn as effectively as possible, in two complementary senses. First, an agent should be able to learn from all sources of data. This calls for a single learning objective that accommodates any data distribution, keeping the system streamlined and scalable. Second, an agent should extract maximum learning signal from every sample, making it highly data-efficient. Currently I focus on the continual learning problem (e.g., recursive self-improvement, frontier discovery), where both properties are crucial under limited and heterogeneous data.

/

Hybrid Reinforcement Learning from Offline Observation Alone
Yuda Song, J. Andrew Bagnell, Aarti Singh
ICML, 2024
We consider a practical setting of hybrid RL where the agent only has access to offline observation data without action labels (e.g., videos of human demonstrations), and we show that it is possible to achieve efficient learning in this setting with a practical algorithm.
Provable Benefits of Representational Transfer in Reinforcement Learning
(alphabetical order) Alekh Agarwal, Yuda Song, Wen Sun, Kaiwen Wang, Mengdi Wang, Xuezhou Zhang
COLT, 2023
We prove the benefit of representation learning on diverse source environments which allows efficient learning on the source environment with the learned representation under the low-rank MDPs setting.

Talks

Reinforcement Learning beyond Reward Maximization
[Slides]
  • Google Research NYC, May 2026.
  • Frontiers in Online Reinforcement Learning Workshop, March 2026.
  • Stanford, March 2026.
  • Harvard ML Foundations Group, February 2026.
To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning
[Slides]
  • New Directions in Reinforcement Learning and Control, May 2026.
  • RL Theory Seminar, March 2026. [Recording]
Harnessing Additional Feedback in LLM Post-Training
[Slides]
  • CMU Goomba Lab, October 2025.
  • FAIR Paris, July 2025.
Hybrid RL: Efficient RL with Both Online and Offline Data
[Slides]
  • Amazon NYC, July 2024.
  • ISAIM Special Session on Deep Reinforcement Learning: Bridging Theory and Practice, January 2024.
  • RL Theory Seminar, November 2023. [Recording]

Teaching

Lecturer
  • CMU 10734: Foundations of Autonomous Decision Making under Uncertainty (Fall 2024, Fall 2025)
Guest Lecturer
  • Cornell CS6789: Foundations of Reinforcement Learning (Fall 2024)
  • CMU 17740: Algorithmic Foundations of Interactive Learning (Fall 2024)
Teaching Assistant
  • UCSD CSE291: Topics in Search and Optimization (Winter 2020)
  • UCSD CSE154: Deep Learning (Fall 2019)
  • UCSD CSE150: Introduction to AI: Search and Reasoning (Winter 2019, Spring 2020)
  • UCSD CSE30: Computer Organization and Systems Programming (Spring 2019, Winter 2018)
  • UCSD CSE11: Introduction to CS & OOP (Fall 2018)