Tsinghua University · AI Systems

Huajun Bai

Ph.D. Student in Computer Science

I study Tokenomics: making every token an LLM computes do useful work.

I am advised by Prof. Jiwu Shu in Tsinghua University's Storage Research Group.

My work spans speculative decoding, efficient reasoning, and agentic inference, connecting algorithm design with real serving systems.

Research agenda

Systems that spend inference time well

I study where autoregressive systems wait, repeat work, or compute tokens they do not need, then design mechanisms that remove those costs without changing model quality.

  1. 01

    Agentic inference

    Overlap model reasoning with tool execution to reduce agent latency.

  2. 02

    Speculative decoding

    Make drafting cheaper and more useful under real serving constraints.

  3. 03

    Efficient reasoning

    Reduce unnecessary generation while preserving answer quality.

Selected systems work

From inference mechanism to measured system result

First author arXiv 2026 Tsinghua

SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference

Huajun Bai, Weiwei Lv, Huichuan Zheng, Youyou Lu, Jiwu Shu

SPORK forks a lightweight probe while an agent is still reasoning, predicts the upcoming tool call, and dispatches it early so tool execution overlaps with the remaining decode. It requires no retraining, auxiliary model, or offline trace.

  • 18% lower GAIA P95
  • 4B–32B model coverage
  • Open source controller
Co-first author ICML 2026 Tencent

SpecExit: Accelerating Large Reasoning Model via Speculative Exit

Rubing Yang*, Huajun Bai*, Song Liu, et al.

SpecExit predicts future tokens and an early-exit signal from a lightweight draft model, reducing overthinking without the separate probing overhead used by prior early-exit systems.

  • 66% shorter generation
  • 2.5× end-to-end speedup
  • No accuracy loss
Co-first author ICLR 2026 AMD

PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation

Zihao An*, Huajun Bai*, Ziqiong Liu, Dong Li, Emad Barsoum

PARD adapts one target-independent draft model across a family of target models and predicts multiple future tokens in parallel, reducing both adaptation and serving cost.

  • training efficiency
  • 3.67× LLaMA3.1-8B speedup
  • 264.88 tokens/s

Additional publications

  • Hierarchical N-Gram speculative decoding for AI PCs
    First author · MobiCom 2025 EdgeFM Workshop
    Paper
  • FACE: Evaluating Natural Language Generation with Fourier Analysis of Cross-Entropy
    NeurIPS 2023
    Paper
  • Expressive User Embedding from Churn and Recommendation Multi-Task Learning
    First author · WWW 2023
    Paper
  • A Corpus for Reasoning about Natural Language Grounded in Photographs
    ACL 2019
    Paper

* Equal contribution

Background

Research and engineering

  1. Meituan · AI SystemsResearch Intern, agentic inference and end-to-end latency · SPORK
  2. Tencent · AI InfraResearch Intern, inference acceleration
  3. AMD · LLM InferenceResearch Intern, speculative decoding and AI PC inference
  4. Early-stage startupFounder and Forward Deployed Engineer
  5. News feed systemsMachine Learning Engineer

Education

Tsinghua and Cornell

  1. Tsinghua UniversityPh.D., Computer Science and Technology
  2. Cornell TechM.Eng., Computer Science
  3. Cornell UniversityB.A., Mathematics and Computer Science
Academic service

Reviewer for ICML, NeurIPS, COLM, and EMNLP; ICML 2026 Gold Reviewer.

MLSYS Feed cover

Research communication

MLSYS Feed

I host a Chinese-language, AI-native podcast that retrieves, ranks, and narrates fresh MLSYS papers through an agentic pipeline. It is also my way of continuously mapping the systems research landscape beyond my immediate projects.

Contact

Let’s talk systems.

F5, Ziqiang Science & Technology Building
Tsinghua University · Beijing, China