Career Profile
I am currently a Senior Algorithm Engineer at Tencent, where I lead an algorithm team of ~15 people on frontier game-AI research and development. Our work spans several directions:
- evaluation systems for LLM-based game generation and high-quality game-data synthesis;
- training point-cloud foundation models for 3D scene generation;
- exploring a new paradigm of video-model-based AI-generated games;
- and the core algorithms behind online AI game generation, such as asset retrieval and 2D/3D asset processing.
I received my Bachelor's and Master's degrees from Harbin Institute of Technology. My bachelor thesis on model-based MARL was completed during an exchange at INSA Lyon in France, supervised by Prof. Dibangoye. During my master's, I was advised by Prof. Lei Feng and Prof. Bing Liu, and worked on cooperative multi-agent reinforcement learning, including multi-level credit assignment and correcting biased value estimation in mixing-based MARL. After graduation, I joined Tencent as a reinforcement learning engineer and shipped production AI bots with RL and imitation learning. I later built LLM-driven NPCs for AAA games and applied SFT/ORPO/RLHF to 3D scene generation, before leading the current game-AI team. My research interests span generative game AI, agentic 3D scene generation, and reinforcement learning.
News
Publications
Education
Experiences
- Lead an algorithm team of ~15 people on frontier game-AI R&D, spanning: LLM-based game-generation evaluation systems and high-quality game-data synthesis; training a point-cloud foundation model for scene generation; exploring novel AI game generation with video models; and core algorithms for online AI game generation such as asset retrieval and 2D/3D asset processing.
- Optimized Qwen with SFT/ORPO/RLHF for 3D scene generation.
- Built LLM-driven high-quality NPCs for AAA games, with human-like dialogue, actions, and expressions/emotions.
- Developed AI bots with reinforcement learning and imitation learning, successfully shipped to production and improved game retention.
- Studied how game-content output affects player retention to guide numerical design.
- Proposed belief occupancy state as a summary to recast Dec-POMDPs under one-sideness sharing to boMDP which is MDP actually.
- Implemented belief occupancy state Heuristic search and value iteration algorithm to solve boMDP.
- Applied linear programming and tabular method to improve the scalability.
- Quantized floating point data of DL Networks into 16 or 8 bits on Caffe.
- Applied KL Divergence to decrease the loss caused by quantization of 8 bits to just 1.5 for MobileNet-SSD.
- Verified the Quantization Scheme for 16 bits on FPGA.
- Designed and trained DL Network for diagnosis of pneumonia on Caffe and implemented it on Zynq.
- Won 2nd Place in the 16th Challenge Cup and Silver Award in the 9th Zuguang Cup.
- Manufactured and debugged control module for a wireless charging system.
