跳到正文
原文
LlamaIndex:產品、工程與評測·· 4 天前AI 評分53

LlamaIndex 比較 GPT-4 與開源 Prometheus 模型的 RAG 評估表現

LlamaIndex: RAG Evaluation Showdown with GPT-4 vs. Open-Source Prometheus Model

AI 導讀

LlamaIndex 使用其框架,基於包含 100 條問題的 Llama 2 論文資料集,比較 Prometheus 與 GPT-4 的 RAG 評估表現。GPT-4 的 Faithfulness 和 Relevancy 分數分別為 0.93 和 0.98,Prometheus 則為 0.39 和 0.57;文章亦展示 Prometheus 有時會誤判上下文資訊。

來源:LlamaIndex:產品、工程與評測 · llamaindex.ai