Mehul Damani
Hello! I am a fifth year Ph.D. student at MIT advised by Jacob Andreas.
My research interests lie at the intersection of RL and LLMs, with a focus on how learning objectives shape model behavior and reliability. Currently, I am thinking about multi-agent RL and swarm alignment.
My recent work has focused on improving LLM reliability through RL for calibration and hallucination reduction (RLCR), self-distillation for continual learning (SDFT), and adversarial learning for alignment (VARL).
Previously, I worked with Lerrel Pinto at NYU on developing automatic curriculum learning methods for RL agents. Before that, I was a part of the MARMot Lab at NUS, where I worked with Guillaume Sartoretti on applying multi-agent reinforcement learning to traffic signal control and multi-agent pathfinding.
Selected Publications
Right in the Right Way: LM Training with Verifiable Rewards and Human DemonstrationsICML 2026 Workshop on RL from World Feedback