Mehul Damani

Hello! I am a fifth year Ph.D. student at MIT advised by Jacob Andreas.

My research interests lie at the intersection of RL and LLMs, with a focus on how learning objectives shape model behavior and reliability. Currently, I am thinking about multi-agent RL and swarm alignment.

My recent work has focused on improving LLM reliability through RL for calibration and hallucination reduction (RLCR), self-distillation for continual learning (SDFT), and adversarial learning for alignment (VARL).

I am currently on the job market looking for roles starting in early 2027. If you are recruiting or know of opportunities, please reach out!

Previously, I worked with Lerrel Pinto at NYU on developing automatic curriculum learning methods for RL agents. Before that, I was a part of the MARMot Lab at NUS, where I worked with Guillaume Sartoretti on applying multi-agent reinforcement learning to traffic signal control and multi-agent pathfinding.

Selected Publications

  1. Thumbnail for Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
    Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
    ICLR, 2026
    Mehul Damani, Isha Puri, Stewart Slocum, Idan Shenfeld, Leshem Choshen, Yoon Kim, and Jacob Andreas
  2. Thumbnail for Self-Distillation Enables Continual Learning
    Self-Distillation Enables Continual Learning
    ICML, 2026 (Spotlight)
    Idan Shenfeld, Mehul Damani, Jonas Hubotter, and Pulkit Agrawal
  3. Thumbnail for Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations
    Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations
    ICML 2026 Workshop on RL from World Feedback
    Mehul Damani, Isha Puri, Idan Shenfeld, and Jacob Andreas
  4. Thumbnail for Vector policy optimization: Training for diversity improves test-time search
    Vector policy optimization: Training for diversity improves test-time search
    NeurIPS, 2026
    Ryan Bahlous-Boldi, Isha Puri, Idan Shenfeld, Akarsh Kumar, Mehul Damani, Sebastian Risi, Omar Khattab, Zhang-Wei Hong, and Pulkit Agrawal
  5. Thumbnail for Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
    Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
    ICML, 2026
    Isha Puri, Mehul Damani, Idan Shenfeld, Marzyeh Ghassemi, Jacob Andreas, and Yoon Kim
  6. Thumbnail for Position: It's Time to Optimize for Self-Consistency
    Position: It's Time to Optimize for Self-Consistency
    Itamar Pres, Belinda Z. Li, Laura Ruis, Zifan Carl Guo, Keya Hu, Mehul Damani, Isha Puri, Ekdeep Singh Lubana, and Jacob Andreas
  7. Thumbnail for The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
    The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
    ICML, 2025
    Ekin Akyürek, Mehul Damani, Adam Zweiger, Linlu Qiu, Han Guo, Jyo Pari, Yoon Kim, and Jacob Andreas
  8. Thumbnail for Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
    Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
    ICLR, 2025
    Mehul Damani, Idan Shenfeld, Andi Peng, Andreea Bobu, and Jacob Andreas