Chaoyu Wang
王超宇
AI Researcher LLM Alignment × Systems
About
Hi, I'm Chaoyu 👋, Founder at MAL.LAB and independent ML researcher — interested in building AI systems that are more capable, more reliable, and genuinely beneficial to humanity.
I received an M.S. from
Northwestern University and a B.S. from
UC San Diego, both in Applied Mathematics — two places I am deeply grateful for. I was fortunate to be mentored by Prof. Zhaoran Wang at Northwestern and Prof. Ioana Dumitriu at UCSD, both of whom shaped how I think about research. My previous work spans LLM fine-tuning and alignment, retrieval-augmented generation, and synthetic data construction.
You can find more about my background in my CV.
🔬 Interests: Agentic Reinforcement Learning, Long-Horizon Decision-Making & Credit Assignment, Self-Evolving Agents & Environments, Scalable Agent Data Synthesis, Efficient Training & Inference Systems.
I am applying to CS PhD programs for Fall 2027, and actively looking for long-term Research Assistant positions in the meantime.
More than a position, I am looking for the right fit — a lab that takes its time with ideas, a collaborator who wants to build something meaningful over the long run, or a mentor genuinely invested in helping someone learn to think independently. I believe this kind of match has to go both ways.
I am available to work fully onsite for six months or more, and I take that commitment seriously — good research takes time, and I am not looking to pass through.
If any of this resonates, I would love to talk: email · calendly
Latest News
Launchpad S1 successfully held in Shanghai
Launchpad S1 successfully held in Shanghai
🚀 Co-organized Launchpad S1 with MAL.LAB — a product launch and go-to-market event in Minhang, Shanghai. Brought together founders and builders, sponsored by DeepTech. From idea to execution in two months.
Graduated from Northwestern University
Graduated from Northwestern University
🎓 Completed M.S. in Engineering Science & Applied Mathematics at Northwestern University.
Graduated from UC San Diego
Graduated from UC San Diego
🎓 Completed B.S. in Applied Mathematics at UC San Diego.
Check out my latest work

Syncopate_Async_AgenticRL
Syncopate_Async_AgenticRL
Head-to-head study of synchronous vs fully-async agentic RL training on verl, featuring multi-turn tool-calling GRPO, long-tail rollout profiling, and staleness/partial-rollout ablations. Quantifies when asynchronous training pays off — and at what scale the crossover arrives.

Darkroom_VeRL-Omni
Darkroom_VeRL-Omni
Diffusion RL post-training on verl: Flow-GRPO on Qwen-Image where rollouts are denoising trajectories, evidence-based upstream reconnaissance, and a CPU OCR reward replacing the VLM judge with 6.9× headroom — the full generate-then-grade loop validated on one RTX 5090.

Switchboard_MoE-DeepEP
Switchboard_MoE-DeepEP
Hand-built MoE expert parallelism: routing-skew capture from DeepSeek-V2-Lite, fused grouped-GEMM Triton kernels (1.28×), 2-all2all dispatch/combine, and overlap ablations on 2×H100 — proving by counter-example why DeepEP exists.

Halftone_QuantizedKVCache
Halftone_QuantizedKVCache
Data-driven INT8 KV-cache quantization: SQNR-guided granularity, near-lossless PPL (+0.2%), 0.50× memory, and a fused Triton int8 decode-attention kernel at 1.56× — with the gap to the 2× ceiling fully explained.

Laminar_GPU-Bubble-Lab
Laminar_GPU-Bubble-Lab
Profile-first CUDA microbenchmark lab for killing GPU pipeline bubbles: multi-stream overlap (1.69×), CUDA Graphs vs torch.compile, and speculative-decoding overlap (1.44×) — every mechanism Nsight-verified on sm_120.








