Bingyang Wu
Bingyang Wu
Light
Dark
Automatic
3
ExpertPlex: A High-Goodput Disaggregated Serving System for MoE LLMs with Adaptive Persistent Kernels
LLMs scale Mixture-of-Experts (MoE) parameters for superior intelligence, but massive weights and dynamic computation impede efficient …
Bingyang Wu
,
Chao Jin
,
Zili Zhang
,
Xinming Wei
,
Yinmin Zhong
,
Ruidong Zhu
,
Chengxu Yang
,
Xin Jin
,
Yuliang Liu
PDF
Cite
DOI
UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing
Large-scale expert parallelism (EP) is becoming pivotal for training and serving frontier MoE models, but it also amplifies …
Xinming Wei
,
Chao Jin
,
Tuo Dai
,
Yinmin Zhong
,
Shan Yu
,
Chengxu Yang
,
Bingyang Wu
,
Zili Zhang
,
Jing Mai
,
Qianchao Zhu
,
Zhouyang Li
,
Yuliang Liu
,
Guojie Luo
PDF
Cite
DOI
ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning
Load imbalance is a long-standing challenge in Mixture-of-Experts (MoE) training and is exacerbated in reinforcement learning (RL) for …
Chao Jin
,
Xinming Wei
,
Yinmin Zhong
,
Chengxu Yang
,
Bingyang Wu
,
Ruidong Zhu
,
Zili Zhang
,
Yuliang Liu
,
Xin Jin
PDF
Cite
DOI
Heddle: A Distributed Orchestration System for Agentic RL Rollout
Agentic Reinforcement Learning (RL) enables LLMs to solve complex tasks by alternating between a data-collection rollout phase and a …
Zili Zhang
,
Yinmin Zhong
,
Chengxu Yang
,
Chao Jin
,
Bingyang Wu
,
Xinming Wei
,
Yuliang Liu
,
Xin Jin
PDF
Cite
DOI
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
Prefix caching is crucial to accelerate multi-turn interactions and requests with shared prefixes. At the cluster level, existing …
Bingyang Wu
,
Zili Zhang
,
Yinmin Zhong
,
Guanzhe Huang
,
Yibo Zhu
,
Xuanzhe Liu
,
Xin Jin
PDF
Cite
DOI
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
Reinforcement learning (RL) has become the core post-training technique for large language models (LLMs). RL for LLMs involves two …
Yinmin Zhong
,
Zili Zhang
,
Xiaoniu Song
,
Hanpeng Hu
,
Chao Jin
,
Bingyang Wu
,
Nuo Chen
,
Yukun Chen
,
Yu Zhou
,
Changyi Wan
,
Hongyu Zhou
,
Yimin Jiang
,
Yibo Zhu
,
Daxin Jiang
PDF
Cite
DOI
A Survey of Resource-efficient LLM and Multimodal Foundation Models
Large foundation models, including large language models (LLMs), vision transformers (ViTs), diffusion, and LLM-based multimodal …
Mengwei Xu
,
Wangsong Yin
,
Dongqi Cai
,
Rongjie Yi
,
Daliang Xu
,
Qipeng Wang
,
Bingyang Wu
,
Yihao Zhao
,
Chen Yang
,
Shihe Wang
,
Qiyang Zhang
,
Zhenyan Lu
,
Li Zhang
,
Shangguang Wang
,
Yuanchun Li
,
Yunxin Liu
,
Xin Jin
,
Xuanzhe Liu
PDF
Cite
DOI
Cite
×