Bingyang Wu
Bingyang Wu
Light
Dark
Automatic
Zili Zhang
Latest
ExpertPlex: A High-Goodput Disaggregated Serving System for MoE LLMs with Adaptive Persistent Kernels
UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing
ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning
FastServe: Iteration-Level Preemptive Scheduling for Large Language Model Inference
Heddle: A Distributed Orchestration System for Agentic RL Rollout
Epiphron: Resource-Efficient Distributed Key-Value Storage
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
Optimizing RLHF Training for Large Language Models with Stage Fusion
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM Serving
Transparent GPU Sharing in Container Clouds for Deep Learning Workloads
Cite
×