Hugging Face Blog·· 2025-12-11精選AI 評分68
llama.cpp server 新增路由模式,支援無需重啟切換多個模型
New in llama.cpp: Model Management
AI 導讀
llama.cpp server 現已提供 router mode,可在不重啟的情況下動態載入、卸載及切換多個模型。它會從 llama.cpp 快取或指定的 --models-dir 尋找 GGUF,並在收到請求時載入模型;達到 --models-max 上限時按 LRU 卸載,預設最多同時載入 4 個模型。各模型以獨立程序運行,一個模型崩潰不會影響其他模型。
推薦理由
多進程隔離配合按需載入和 LRU 卸載,展示 llama.cpp 如何在同一服務中管理多個模型並限制同時載入數量。
來源:Hugging Face Blog · huggingface.co