MTS, Microsoft AI Superintelligence
Rengan
Xu
LLM infrastructure — model and systems optimization.
About
Research shaped by
systems thinking.
I am a Member of Technical Staff at Microsoft AI Superintelligence (MSI). Previously, I was a Research Scientist at Meta SuperIntelligence Labs. I focus on scaling reinforcement learning post-training for high throughput—spanning attention, MoE, grouped GEMM, parallelism strategies, and E2E system efficiency.
Education
Ph.D., Computer Science
University of Houston, 2016
Advisor: Dr. Barbara Chapman
B.S., Computer Science
Hefei University of Technology, 2009
Research focus
LLM
Reinforcement Learning (RL)
Inference, Post Training
GPU Kernels
Publications
Selected work
Research across ads delivery, learned representations, accelerator programming, and high-performance systems.
-
2026
Article ↗
In-Kernel Broadcast Optimization: Co-Designing Kernels for RecSys Inference
-
2024
PDF ↗
Enhancing Performance and Scalability of Large-Scale Recommendation Systems with Jagged Flash Attention
-
2024
PDF ↗
Async Learned User Embeddings for Ads Delivery Optimization
-
2017
PDF ↗
Implementing the OpenACC Data Model
-
2015
PDF ↗
Compiler Transformation of Nested Loops for General Purpose GPUs
-
2014
PDF ↗
SPEC ACCEL: A Standard Application Suite for Measuring Hardware Accelerator Performance
- 2019PDF ↗
Densifying Assumed-Sparse Tensors
- 2016PDF ↗
An Analytical Model-Based Auto-Tuning Framework for Locality-Aware Loop Scheduling
- 2015PDF ↗
Multi-GPU Support on Single Node Using Directive-Based Programming Model
- 2014PDF ↗
NAS Parallel Benchmarks for GPGPUs Using a Directive-Based Programming Model
- 2014PDF ↗
Reduction Operations in Parallel Loops for GPGPUs
- 2018PDF ↗
The OpenACC Data Model: Preliminary Study on Its Major Challenges and Implementations
- 2016PDF ↗
ACC-SVM: Accelerating SVM on GPUs Using OpenACC
- 2016PDF ↗
Optimizing GPU Register Usage: Extensions to OpenACC and Compiler Optimizations
- 2014PDF ↗
Accelerating Kirchhoff Migration on GPU Using Directives
- 2014PDF ↗
A Validation Testsuite for OpenACC 1.0
- 2014PDF ↗
Compiling a High-Level Directive-Based Programming Model for GPGPUs
- 2014
OpenUH — An Open Source OpenACC Compiler
- 2013
Exploring Programming Multi-GPUs Using OpenMP and OpenACC-Based Hybrid Model
- 2013
Filesystem Aware Scalable I/O Framework for Data-Intensive Parallel Applications
- 2013PDF ↗
OpenACC Programming Experiences Using Scientific Applications
- 2012PDF ↗
Directive-Based Programming Models for Scientific Applications — A Comparison
- 2012PDF ↗
Parallel I/O Framework for Data-Intensive Parallel Applications
Contact
Let’s connect.
For research conversations, collaborations, or professional inquiries, reach me by email or find my work and profile online.