跳到正文
原文
Cerebras:Blog·· 3 小時前AI 評分57

Cerebras 開放預覽 Gemma 4 31B,多模態推理速度達每秒 1,851 個 token

Gemma 4 on Cerebras—The Fastest Inference is Now Multimodal

AI 導讀

Cerebras 已在 Cerebras Inference Cloud 限時公開預覽 Gemma 4 31B,支援以圖片輸入進行多模態推理。Artificial Analysis 測得其輸出速度為每秒 1,851 個 token,是典型 GPU 端點的 35 倍;首次回應 token(包括推理)在 1.5 秒內返回。

來源:Cerebras:Blog · cerebras.ai