Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
01 · Profile
About me
I am 2nd year PhD student at The Chinese University of Hong Kong (CUHK), advised by Prof. Viet Anh Nguyen. Prior to my PhD, I spent 2 wonderful years at VinAI Research as a Research Resident.
While I have broad experience in recommender systems, graph neural networks, and continual learning, my current research focuses on reasoning optimization for Transformer-based Language Models. I work to improve LLM reasoning performance, diversity, and efficiency using techniques like LLM Post-Training (GRPO, self-distillation), KV Cache Compression, model pruning, and routing.
Please feel free to reach out via email (hilljun.2000@gmail.com) or WeChat (ID: junhill9961).
02 · Research
Publications
2026
2025
-
Reasoning Planning for Language Models
-
Mixture-of-Personas Language Models for Population Simulation
-
Structured Pruning for Diverse Best-of-N Reasoning Optimization
-
Task-driven Layerwise Additive Activation Intervention
2024
-
Cold-start Recommendation by Personalized Embedding Region Elicitation
-
Explaining Graph Neural Networks via Structure-aware Interaction Index
-
Generative Conditional Distributions by Neural (Entropic) Optimal Transport
2022
-
Combining Soft-Actor Critic with Cross-Entropy Method for Policy Search in Continuous Control
Journals
-
Retrospective Feature Estimation for Continual Learning
03 · Community
Academic Services
- ICML 20262025
- ICLR 20262025
- WWW 2025
- NeurIPS 2024
04 · Recognition
Honors and Awards
- 2026: ICML 2026 Golden Reviewer.
- July 2024: ICLR2026 Travel Grant! (Rio de Janeiro, Brazil)
- July 2024: ICML2024 Travel Grant! (Vienna, Austria)
- April 2022: Bachelors Thesis with highest score.
- September 2018 - April 2022: Honor Student Scholarship for all Academic Years - UIT