跳到正文
HuggingFace Daily Papers(社區熱門論文)·· 1 天前AI 評分46

TokenRouter:面向 token 級 LLM 路由的高效服務系統

TokenRouter: Efficient Serving System for Token-Level LLM Routing

AI 導讀

TokenRouter 是一套面向 token 級 LLM 路由推理的高效服務系統,針對現有系統的步驟失同步和批次接納延遲問題而設計。它在不同路由算法、工作負載及模型配對下,解碼吞吐量較現有系統提升 2.01-64.15x。系統採用「以請求為中心的程式設計、以模型為中心的執行」,由各 LLM 子伺服器非同步分派請求,並使用延遲批次排程器。

正文

Abstract

Large language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving. While coarse-grained routing at the session or query level has been widely adopted in production systems, recent algorithmic work shows that fine-grained token-level routing can yield substantial efficiency and quality gains. However, efficiently serving token-level routed inference poses significant challenges to existing systems. Built on single-LLM assumptions, current systems suffer from severe step desynchronization and frequent batch admission delays under token-level routing, and they also impose high implementation complexity on developers. To address these challenges, we design TokenRouter, an efficient and developer-friendly serving system for token-level routed LLM inference. TokenRouter follows the principle of request-centric programming, model-centric execution: developers describe routing logic from the perspective of a single request, while the runtime launches a subserver for each LLM and dispatches requests asynchronously. Each subserver employs a delayed-batching scheduler, whose optimal hyperparameters are derived from a mathematical throughput model of the system. Across diverse routing algorithms, workloads, and model pairs, TokenRouter achieves 2.01-64.15x higher decoding throughput than existing systems, substantially advancing the serving efficiency of token-level LLM routing. Our code is available at https://github.com/thu-nics/TokenRouter.

View arXiv page View PDF Project page GitHub 4 Add to collection

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.12242 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.12242 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.12242 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co