I am a senior algorithm engineer with Qwen Unit, Alibaba Group, where I focus on reinforcement learning for language models with a strong drive toward real-world applications.
Before this, I worked at Ant Group, where I focused on doing LLM-RL research and its application in the safety post-training stage of the company’s foundation models.
Before that, I did my Ph.D. in Dr. Tianyi Chen’s group which I was fortunate to join as the first Ph.D. student. My Ph.D. research focused on optimization and reinforcement learning. Meanwhile, I was a research intern with IBM Research AI where I collaborated with Pin-Yu Chen, Payel Das, Songtao Lu, Xiaodong Cui and many other talented researchers. My research at IBM focused on LLM alignment and RL.
News and highlights
- [Jul. 2026] I joined the Qwen Bussiness Unit of Alibaba Group, where I will be working on the LLM-RL foundational algorithms.
- [May. 2026] One co-authored paper accepted in ICML 2026.
- [Jan. 2026] Our paper on entropy regularization of LLM-RL is accepted in ICLR 2026.
- [Mar. 2025] I am excited to join Ant Group via its research talent program Ant Star.
- [Feb. 2025] The extended study of our ICML 2024 paper is accepted in JMLR.
- [Jan. 2025] Our paper is accepted in ICLR 2025:
- [Dec. 2024] An extended study of our ICML 2023 paper has been accepted in Mathematical Programming.
- [Oct. 2024] New paper on improved LLM alignment framwork:
Industry experiences
Qwen Unit, Alibaba Group. Present
- Senior foundational algorithm engineer.
Ant Group. 03.2025 - 07.2026
- Senior research engineer, joined via Ant Star talent program.
IBM Research AI, T. J. Watson Research Center. 05.2024 - 08.2024
- Research intern, mentored by Dr. Pin-Yu Chen and Dr. Payel Das.
IBM Research AI, T. J. Watson Research Center. 05.2021 - 08.2021
- Research intern, mentored by Dr. Songtao Lu and Dr. Xiaodong Cui.
Selected research
On Entropy Control in LLM-RL Algorithms
Han Shen
ICLR 2026. [paper]SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection
Han Shen, Pin-Yu Chen, Payel Das, Tianyi Chen
ICLR 2025. [paper]Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF
Han Shen, Zhuoran Yang, Tianyi Chen
ICML 2024, extended work in JMLR [paper]Mitigating Gradient Bias in Multi-objective Learning: A Provably Convergent Approach
Heshan D. Fernando, Han Shen, Miao Liu, Subhajit Chaudhury, Keerthiram Murugesan, Tianyi Chen
ICLR 2023 Oral. [paper]On Penalty-based Bilevel Gradient Descent Method
Han Shen, Quan Xiao, Tianyi Chen
ICML 2023, extended work in Mathematical Programming. [paper]A Single-timescale Analysis for Stochastic Approximation with Multiple Coupled Sequences
Han Shen, Tianyi Chen
NeurIPS 2022 Oral. [paper]
Services
Reviewer for NeurIPS, ICML, ICLR, AISTATS, AAAI, IEEE Transactions on Signal Processing, etc.
