Cerebras:Blog·· 3 小時前AI 評分57
Cerebras 開放預覽 Gemma 4 31B,多模態推理速度達每秒 1,851 個 token
Gemma 4 on Cerebras—The Fastest Inference is Now Multimodal
AI 導讀
Cerebras 已在 Cerebras Inference Cloud 限時公開預覽 Gemma 4 31B,支援以圖片輸入進行多模態推理。Artificial Analysis 測得其輸出速度為每秒 1,851 個 token,是典型 GPU 端點的 35 倍;首次回應 token(包括推理)在 1.5 秒內返回。
來源:Cerebras:Blog · cerebras.ai