跳到正文
原文
Google DeepMind·· 2026-06-09精選AI 評分72

Google DeepMind 發佈 Gemma 4 12B 統一免編碼器多模態模型

Introducing Gemma 4 12B: a unified, encoder-free multimodal model

AI 導讀

Google DeepMind 發佈 Gemma 4 12B,定位為可在筆記型電腦本機運行的多模態模型,並為中型模型加入原生音訊輸入。它採用免多模態編碼器架構,將視覺和音訊輸入直接整合至 LLM backbone;標準基準表現接近 26B MoE,總記憶體佔用不足其一半。Gemma 4 12B 以 Apache 2.0 授權釋出,可透過 LM Studio、Ollama 等工具試用。

推薦理由

Gemma 4 12B 的免編碼器設計、接近 26B MoE 的基準表現與較低記憶體佔用,提供比較中型多模態模型架構和本機部署定位的具體依據。

來源:Google DeepMind · deepmind.google