▶
ExecutionTrigger
Automation/LLM/Vision
Uses a vision LLM to locate UI elements based on natural language description
Scores range from 0 to 10. Higher values mean more impact, exposure, or operational weight.
Trigger
Vision-capable LLM model
Screenshot of the screen as base64 PNG, JPEG, WebP or GIF (a data URL is fine). Ignored when Image is connected
Screenshot of the screen as an image, e.g. from the Screenshot node. Takes precedence over the base64 Screenshot
Screen frame of the screenshot, from the capture node. Coordinates are desktop input coordinates (ready for the mouse nodes) when Frame is connected, otherwise pixels of the original screenshot
Natural language description of the element to find (e.g., 'the blue submit button')
Optional context about the application or page
Continue
Element not found, or the model did not give a point on the screenshot
Element location. x/y is the element's center and width/height its size, in desktop input coordinates (ready for the mouse nodes) when Frame is connected, otherwise pixels of the original screenshot