跳到正文
Hugging Face Blog·· 1 小時前精選AI 評分61

Liquid AI 發佈端側決策模型 d1-3B 與實驗版 d1-omni-600M

Multimodal open d1 decision models for the edge

AI 導讀

Liquid AI 發佈兩款開放權重端側決策模型 d1-3B 和實驗版 d1-omni-600M。d1-3B 支援文字與圖像輸入,在 Decision Index 0.2.1 得分 48.57,並在七個公開數據集取得 82.9 的平均分;d1-omni-600M 支援文字與圖像或文字與音訊輸入。

推薦理由

文章把 d1-3B 的公開基準成績與多款邊緣設備的延遲數據並列,讓讀者瞭解其決策表現及端側推理速度。

正文 · 繁體中文

今日,我們推出 d1 決策模型系列中的兩款開放式決策模型:d1-3B 和 d1-omni-600M(實驗版)。

  • Decision Index 0.2.1 上 10B 以下最佳的決策模型:d1-3B 得分 48.57,領先所有 4B 和 9B 模型,亦高於 Decider 35B-A3B(47.11)。
  • 多模態:d1-3B 支援文字和圖片,d1-omni-600M 則支援文字和圖片,或文字和音訊。
  • 快速:d1-3B 在 NVIDIA Jetson AGX Thor 上回答問題需時 16 ms,在 Jetson AGX Orin 上需時 26 ms,在 Jetson Orin Nano 上需時 50ms。

我們如何為邊緣裝置打造決策模型

這些開放式 d1 決策模型以我們的 Liquid Foundation Models(LFMs)為基礎。與生成模型不同,決策模型不會產生 token,而是在單次前向傳播中作答。

d1-3B 和 d1-omni-600M 分別以兩個截然不同的主幹模型訓練:

  • d1-3B 以我們最新的 VLM LFM2.5-VL-3B 訓練,該模型採用僅解碼器架構,接收文字和圖片作為輸入。
  • d1-omni-600M 以 LFM2.5-Encoder-350M 訓練,這是一個雙向編碼器。它加入視覺和音訊編碼器,以處理全部三種模態。它可接收文字和圖片,或文字和音訊作為輸入。這個模型目前處於早期研究發布階段,仍在進一步開發。

基準測試結果

我們在七個公開數據集上對 d1-3B 和 d1-omni-600M 進行基準測試,涵蓋閲讀理解、毒性偵測、意圖分類、醫療問答和跨語言理解。d1-3B 的平均得分為 82.9,是表中最高分,亦高於 Decider 4B。d1-omni-600M 得分 78.4,以僅四分之一的參數量超越 Decider 2B(77.1)。

基準測試 d1-omni-600M d1-3B Decider 2B Decider 4B
SQuAD 2.0 74.0 83.3 67.7 76.0
Civil Comments 95.8 93.3 93.6 92.8
MASSIVE intent 86.1 86.9 81.1 88.3
PubMedQA 61.3 68.3 65.7 63.3
BoolQ 77.7 86.3 87.3 89.0
XNLI 74.7 85.6 85.0 88.6
PAWS-X 79.5 76.4 59.5 69.8
平均值 78.4 82.9 77.1 81.1

我們已驗證,d1-3B 在標準視覺基準測試中保留了其主幹模型 LFM2.5-VL-3B 的視覺能力,而 d1-omni-600M 能處理全部三種模態。由於 Decision Index v0.3 只包含一個私有視覺子集,而音訊決策基準測試目前仍是尚待解決的問題,因此我們沒有公佈任何視覺或音訊基準測試結果。

速度

我們與 NVIDIA 合作,使用 NVIDIA 技術堆疊,在 NVIDIA GeForce RTX 4090、NVIDIA Jetson AGX Thor、Jetson AGX Orin 64 GB 和 Jetson Orin Nano 上評估 d1-3B。由於 d1-omni-600M 仍處於早期研究發布階段,本次發布不會提供其速度數據。

邊緣推理。d1-3B 在每部已測試裝置上回答單一問題均少於 50 ms。回答三個問題所需時間僅為回答一個問題的 1.3x;AGX Thor 的時間則由 16 ms 增至 20 ms。

一個問題 3 個問題 3.4K-token 狀態 384px 圖片 64 個打包狀態
Apple M5 Pro 30 ms 41 ms 640 ms 62 ms 78 / s
Jetson AGX Thor 16 ms 20 ms 220 ms 35 ms 262 / s
Jetson AGX Orin 64 GB 26 ms 35 ms 560 ms 83 ms 110 / s
Jetson Orin Nano 50 ms 73 ms 1,640 ms 202 ms 38 / s

GPU 推理。在 GPU 上,d1-3B 在兩個平台上回答問題均少於 10 ms,處理 384px 圖片則少於 18 ms。

一個問題 3 個問題 3.4K-token 狀態 384px 圖片 64 個打包狀態
NVIDIA RTX 4090 8 ms 21 ms 102 ms 17 ms 475 / s
AMD MI325X 9 ms 14 ms 44 ms 18 ms 1,106 / s

如何使用開放式 d1 決策模型

當你需要快速、結構化的決策(包括多模態輸入)時,可選用 d1 決策模型。以其規模而言,d1-3B 的決策品質最高;若重視模型佔用空間,d1-omni-600M 則更合適。

安裝相依套件(需要 transformers>=5.14):

pip install "transformers>=5.14" torch torchvision pillow

這些模型隨附專屬程式碼,因此請使用 trust_remote_code=True 載入:

import io
import urllib.request

import torch
from PIL import Image
from transformers import AutoModel

device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True,
                                  dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)

# Several named questions over one text state, answered in one pass
questions = {
    "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
                          "fraud": "Suspected unauthorised use"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["Can wait", "Today", "Blocking the customer now"]},
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))

# An image as the whole state
url = "http://images.cocodataset.org/val2017/000000039769.jpg"  # two cats on a sofa
photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))
print(model.system_one(None, {"cats": {"type": "choice", "instructions": "How many cats are there?",
                                       "criteria": {"one": "One", "two": "Two", "more": "Three or more"}}},
                       images=[photo]))

# Many requests, packed together with no padding
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))

為求簡潔,我們只列出 d1-3B 的示例。請參閲 d1-omni-600M 模型卡,瞭解如何執行該模型。

開始使用開放權重 d1 決策模型

這兩個決策模型均採用開放權重,現已在 Hugging Face 上提供:

期待看看你會打造出甚麼。

引用

如果你使用這項成果,請引用發布網誌文章:

@article{liquidAI2026opend1,
  author  = {Liquid AI},
  title   = {Open d1: Edge decision models for text, vision, and audio},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {www.liquid.ai/blog/open-d1},
}

來源:Hugging Face Blog · huggingface.co