▶
ExecutionTrigger
Automation/LLM/Vision
Uses vision LLM to comprehensively observe and describe the current screen
Scores range from 0 to 10. Higher values mean more impact, exposure, or operational weight.
Trigger
Vision-capable LLM model
Screenshot as base64 PNG, JPEG, WebP or GIF (a data URL is fine). Ignored when Image is connected
Screenshot as an image, e.g. from the Screenshot node. Takes precedence over the base64 Screenshot
Screen frame of the screenshot, from the capture node. Coordinates are desktop input coordinates (ready for the mouse nodes) when Frame is connected, otherwise pixels of the original screenshot
Specific area or aspect to focus on (optional)
Continue
Complete screen observation
Text description of the screen
Observed elements. Optional x/y is the element's center in desktop input coordinates (ready for the mouse nodes) when Frame is connected, otherwise pixels of the original screenshot
Center x: desktop input coordinates when a frame was connected, otherwise screenshot pixels. Absent when the model gave no point on the screenshot.
Center y, in the same space as `x`.