跳到正文
Hacker News:AI 熱帖·· 3 小時前AI 評分39

讓 AI 智能體在螢幕上畫出大型箭嘴、方框和文字

Let your AI agents paint big arrows, boxes and text on your screen

AI 導讀

這項工具讓 AI 智能體在螢幕上畫出大型箭嘴、方框和文字。討論者認為,這種標示可用於引導用戶操作應用程式、改善文件中的截圖指引,以及協助智能體互相溝通。亦有人質疑,若箭嘴本身無法取得焦點,未必能解決確認對話框未能自動取得焦點的問題。

正文
Let your AI agents paint big arrows, boxes and text on your screen (github.com/franzenzenhofer)
44 points by franze 43 minutes ago | hide | past | favorite | 15 comments

Literally unusable as it is. Some minimal extra features this would need:

- rainbow dripping arrows

- angrily pointing arrows

- flame-surrounded text boxes with particle effects

- the ability for the agent to play airhorn.wav at max volume, overriding existing volume or mute settings


Interesting. When I read the headline I imagined this would be a sort of thinking trace booster -- letting the agent focus its own attention on different parts of the screen. But this is cool in a different way. I bet agentic harnesses would find it useful for communicating with other agents / themselves as well.


This could be great for documentation. Screenshots in docs are frequently useless because they show me a screen and say "click X" where I still have to search X visually. And I could just to dhat in the other tab I have open.


Am I missing something here?

What terminal / agentic workflows spawn the demonstrated dialogue boxes that require the User to click/reject the action? Aren't most such flows actually inline UIs? And finally, if the core issue is that these confirmation boxes are tied to the terminal that triggered them, which does not autofocus, what's the point of said "big arrow" that is also lurking behind without focus?

If the said "big arrow" automatically gains focus, isn't the real fix here to just make the dialogue boxes themselves gain focus automatically? Both are similarly disruptive anyway.


The first example (of HN) is the one that feels like it has the most potential to me.

"Tech me how to use this app myself" kinda stuff. Guiding agent rather than doing agent.


"and they keep hitting the same wall, the part that only a human may do"

Grim. If you're just there to click sudo buttons for the bot, you might as well give it root access and be done.


Claude refuses to do certain actions (enter passwords, change security settings, create new accounts on external services) even in Yolo mode running as sudo. (I tested it all on its own mac machine)


You can add custom auto-mode classifier rules and even disable the built-in ones, if you want to live on the edge like this.


What a time to be a radical centrist - the AI haters seem out of touch, the AI thought leaders can't stop huffing their farts and being condescending, and somehow this is on the top of HN. What a silly time.

來源:Hacker News:AI 熱帖 · news.ycombinator.com