跳到正文
Hacker News:AI 熱帖·· 3 小時前AI 評分59

研究指 Claude、ChatGPT 會按推測財富推薦不同價位選項

Study: Claude, ChatGPT Offer Different Shopping Prices Based on Wealth

AI 導讀

一項涵蓋 13 個 AI 智能體、325K 次實驗的研究發現,在航班、健康保險和研究生課程選擇中,8 個模型會因推斷用戶較富裕而推薦較昂貴選項。要求找最便宜選項時,Gemini 2.5 Flash 對較富裕用戶的推薦平均仍高出 $280;文中 Figure 4 所列 GPT 和 Claude Opus 4.8 的差距則分別為 $21 和 $20。

正文

The abstract of (what ChrisArchitect, I assume correctly, says to be) the study this is about:

Personal AI agents make recommendations and take actions on people's behalf in high-stakes economic contexts, e.g., buying a flight, choosing health insurance, or selecting a graduate program. The agent is given access to the user's personal context, e.g., their email inbox and a structured profile of personal attributes, with the intention of making an optimal, personalized decision for the user. We show that by simply providing this personal context, the agent steers recommendations based on inferred wealth, without being explicitly instructed to do so. In a suite of 325K experiments on 13 agents across three types of economic decisions (flights, health insurance, and graduate programs), we find that 8 models systematically choose more expensive options for wealthier users when requests are identical. This steering continues even when it directly goes against the user's stated objective: when explicitly instructed to find the cheapest option, some agents still act on the wealth profile they have inferred. It also occurs when wealth is inferred from ambient data, such as emails unrelated to the task. And it persists under privacy controls that block specific attributes: blocking financial attributes largely removes the disparity, but blocking other attributes leaves it unchanged and can increase it by up to 40% for insurance, as agents rely on the remaining signals to infer wealth. Larger and more capable models are no better; Claude Opus 4.8 shows the largest effect. We term this misalignment "adversarial delegation", in which the very conditions that make a personal AI agent useful - access to personal information - enable it to act against the user's interests.

Some of this seems bad, some not. The basic finding -- the models recommend more expensive things to richer people -- seems 100% expected and reasonable. Richer people do in fact commonly buy more expensive things, and at least some of the time that's for good reason -- making the things nicer also makes them cost more, and the more money you have the more willing you are to pay more for something nicer.

Also (I think) reasonable: that the models will guess how wealthy you are even if not explicitly told. (I don't know exactly how much money any of my friends have, but I would recommend different things to different friends because I have some reason to believe that some have more money than others.)

Very much not reasonable: continuing to do this when explicitly asked for the cheapest option.

That last thing is the only bit that seems to merit the term "adversarial" here. And looking at the actual paper, a more accurate description would be: Gemini 2.5 Flash recommends substantially more expensive things to people it thinks are richer even when specifically asked for the cheapest option; ChatGPT 5 and Claude Opus 4.8 do not.

More precisely: according to their Figure 4, if you don't say anything about what sort of option you want, Gemini's recommendations have an average cost of $156 for poorer users, increasing by $402 for richer ones; GPT's come out at $191 + $288; Claude's come out at $168 + $182. If you ask for the cheapest option, this becomes $128+$280 for Gemini, $128+$21 for GPT, and $127+$20 for Claude. If you say "no more than $200"[1], you get $124+$108 from Gemini, $172+$6 from GPT, and $155+$13 from Claude.

[1] I am oversimplifying slightly.

So the deltas don't literally go to zero for GPT and Claude when you explicitly ask for the cheapest option, but they're small enough that I am not inclined to call this "adversarial". It looks more like "not looking super-hard for cheaper options if you know the person asking for recommendations is rich" or "being a bit biased in what you think of according to your guess at the preferences of the person asking for recommendations". Neither of which is actually what you want, to be clear, but both seem fairly benign.

Gemini 2.5 Flash, on the other hand, I'm pretty happy to call "adversarial" here.

來源:Hacker News:AI 熱帖 · news.ycombinator.com