Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization
MAQ learns reusable macro actions from human demonstrations and integrates them with off-the-shelf reinforcement learning algorithms. The method improves human-likeness on D4RL Adroit tasks while preserving task-oriented behavior.
Jian-Ting Guo, Yu-Cheng Chen, Ping-Chun Hsieh, Kuo-Hao Ho, Po-Wei Huang, Ti-Rong Wu, I-Chen Wu. NeurIPS 2025 Main Track.