I am currently a Technical Staff Member at DeepSeek AI, where I lead the LLM Alignment Team. My team focuses on advancing model capabilities in core areas including Writing, QA, AI Search, General Reasoning, General Agent, and Safety. My primary research goal is to build generic artificial general intelligence (AGI) through scaling laws and reinforcement learning.
At DeepSeek, I am deeply involved in the development of the DeepSeek model family, including DeepSeek-V1/2/3/4, DeepSeek-R1, and DeepSeek Math. Most notably, my team pioneered the Group Relative Policy Optimization (GRPO) algorithm and DeepSeek-R1-Zero. These technical breakthroughs laid the foundation for DeepSeek-R1, our work published in Nature, which demonstrates the feasibility of incentivizing complex reasoning capabilities in LLMs via pure reinforcement learning without supervised fine-tuning.
Before joining DeepSeek, I was a Senior Researcher at the Natural Language Computing Group, Microsoft Research Asia (MSRA). During my tenure there, I developed foundational models including WavLM, which was recognized with the IEEE Signal Processing Society (SPS) Best Paper Award. I received my Ph.D. and B.S. degrees from Beihang University, advised by Prof. Ming Zhou and Prof. Zhoujun Li.
To date, I have published over 100 papers in top-tier conferences and journals such as Nature, NeurIPS, ICML, ICLR, ACL, and EMNLP, with more than 50,000 citations on Google Scholar.
Email: wumark at 126 dot com
DeepSeek AI
Technical Staff Member and Head of the LLM Alignment Team
September 2023 – Present
Microsoft Research Asia
Senior Researcher, Natural Language Computing Group
March 2021 – September 2023
Microsoft Research Asia
Researcher, Natural Language Computing Group
June 2019 – March 2021
My recent research explores how large language models can acquire more reliable reasoning, planning, tool use, and general problem-solving abilities through large-scale post-training and reinforcement learning.
Selected contributions include:
DeepSeek-R1 — Reinforcement learning for complex reasoning in large language models
GRPO — A group-relative optimization method for reinforcement learning
DeepSeekMath — Mathematical reasoning in open language models
DeepSeek-V3 — Efficient scaling and mixture-of-experts language modeling
WavLM — Self-supervised pretraining for full-stack speech processing
BEATs — Audio pretraining with acoustic tokenizers
For a complete list of publications, please visit my Google Scholar profile.
IEEE Signal Processing Society (SPS) Best Paper Award
Top 10 most significant Innovations, Netexplo, 2023
Global Top 50 Chinese Young Scholars in NLP, Baidu, 2022
InterSpeech Best Student Paper Nomination 2021
AdeptMind ScholarShip 2018
Microsoft Research Asia Ph.D. Fellowship, 2017
National Scholarship, Beihang University, 2015, 2018