I build and study infrastructure for large language models: inference serving, speculative decoding, and distributed training. I am completing a B.Eng. (Hons) in Software Engineering at the University of Sydney. Most recently I was an Applied Scientist Intern at Microsoft, working on foundation-model serving; before that, an Agent Engineering Intern at Xiaohongshu (RedNote). My research is on allocating inference compute: consequence-aware reasoning budgets and visual token compression, budget-robust speculative decoding, reward-aware execution gating for agents, and market-aware routing across inference providers.
Open to LLM Engineer / LLM Infrastructure roles.
-
Not All Errors Are Equal: Consequence-Aware Reasoning Compute Allocation
Liang He, Jingbo Wen, Haoyu Wang, Ziqi He, Yixiong Chen, Kangning Cui, Xilu Wang
arXiv:2606.04402, 2026. 22–33% lower cost-weighted loss than difficulty-aware compute routing on SWE-bench Lite. -
BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding
Liang He, Jingbo Wen, Qishi Zhan, Yixiong Chen, Kangning Cui, Qizhen Lan, Xilu Wang
arXiv:2606.00144, 2026. One budget-robust drafter for sparse KV Cache; 6.55× end-to-end speedup at 4K context. [code] -
Not All Visual Tokens Are Equally Safe to Remove: Consequence-Sensitive Visual Token Compression
Jingbo Wen, Liang He, Mingyu Cao, Haoyu Wang, Minxuan Hu, Kangning Cui, Xilu Wang
arXiv:2608.09176, 2026. High-stakes VLM errors 0.300 → 0.133 at fixed token budget; 38% lower cost-weighted error. -
From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents
Liang He, Jingbo Wen, Hongyu Gu, Hao Li, Haoyu Wang, Yixiong Chen, Kangning Cui, Xilu Wang
arXiv:2608.09168, 2026. RADEG skips 68% of agent calls while retaining 61% of total reward. -
You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference
Liang He, Jingbo Wen, Yixiong Chen, Yue Yang, Qizhen Lan, Kangning Cui, Xilu Wang
arXiv:2609.37902, 2026. Price does not predict quality; measured routing saves ~50% at matched quality.
Microsoft — Applied Scientist Intern, Azure community infra
Jun 2026 – Sep 2026
- Built a dynamic-batching request scheduler in Go for Foundation Model serving (routing, admission control, deadlines), sustaining 2K+ concurrent streams; cut p99 TTFT by 28% under bursty traffic.
- Integrated EAGLE-3 speculative decoding into the serving runtime, with per-request KV-cache rollback so batches with unequal acceptance lengths verify in one pass; 1.8× tokens/s at mean acceptance length 3.6.
- Profiled the scheduler-to-GPU path with pprof and Nsight Systems; traced a TPOT regression to per-token scheduler–runtime RPC overhead and removed it with batched streaming, lowering TPOT by 17%.
- Built benchmarking and observability for TTFT, TPOT, tokens/s, and KV-cache usage across batch sizes and speculation configs; its load sweeps set the concurrency threshold above which speculation is disabled.
Xiaohongshu (RedNote) — Agent Engineering Intern, Social Networking Engineering
Dec 2025 – Mar 2026
- Designed and developed an internal Coding Agent from 0→1, enabling autonomous codebase understanding, multi-file editing, tool use, code execution, and iterative debugging from natural-language instructions.
- Built a 0→1 OCR verification pipeline (extraction, validation, exception handling) for a production mobile app and took it from prototype to launch: 98% verification accuracy, 21K+ users verified on day one.
- Optimized LLM serving (vLLM, INT4 quantization, continuous batching): +45% QPS, −40% cost.
- Built and maintained internal LLM training and evaluation pipelines supporting SFT / DPO / RLHF workflows.
Distributed LLM Training & Inference System — personal project, LLM Systems / Distributed Computing
Jan 2026 – Present
- Implemented a 350M–1.3B GPT-style Transformer from scratch (GQA, RoPE, SwiGLU, RMSNorm).
- Built Megatron-style Tensor Parallelism (Column/RowParallelLinear, VocabParallelEmbedding) and Pipeline Parallelism (GPipe / 1F1B scheduling, P2P activation transfer, pipeline bubble analysis).
- Integrated DeepSpeed ZeRO-1/2/3 with CPU Offload into a 3D parallel training stack (TP × PP × DP), pre-training a 1.3B model on 4×A100 (TP=2, PP=2) over 1.5B tokens.
- Built a KV Cache inference engine (Prefill/Decode separation, dynamic batching, O(n) per-step attention).
- Wrote custom CUDA C++ and Triton kernels (elementwise fusion, reduction, softmax/LayerNorm, tiled matmul) and profiled them against PyTorch native ops with Nsight Compute.
| LLM Training | LoRA · SFT / DPO / RLHF · DDP · Tensor / Pipeline Parallelism · DeepSpeed ZeRO-1/2/3 · NCCL |
| LLM Inference | KV Cache · Paged Attention · Speculative Decoding · Quantization · vLLM · TensorRT-LLM |
| GPU & Kernel | CUDA C++ · Triton · Nsight Compute · CUDA Streams / Events |
| ML Engineering | PyTorch · Go · FAISS · ChromaDB · RAG · FastAPI · Docker · Linux · Git |
University of Sydney — B.Eng. (Hons), Software Engineering (ECE)
Mar 2023 – Mar 2027 (expected)
GPA 3.8 / 4.0 · 2025 Dean’s List · TOEFL iBT 110 / 120 (5.5 / 6)
I write 拆解大模型 (“Taking LLMs Apart”), a series in Chinese on Juejin about how large language models work: language modeling, Transformer internals, attention, and training.
