▶
ExecutionTrigger
Automation/LLM/Planning
Uses LLM to suggest the most appropriate next action given current screen and goal
Scores range from 0 to 10. Higher values mean more impact, exposure, or operational weight.
Trigger
Vision-capable LLM model
Current screenshot as base64 PNG, JPEG, WebP or GIF (a data URL is fine). Ignored when Image is connected
Current screenshot as an image, e.g. from the Screenshot node. Takes precedence over the base64 Screenshot
Screen frame of the screenshot, from the capture node. Coordinates are desktop input coordinates (ready for the mouse nodes) when Frame is connected, otherwise pixels of the original screenshot
Ultimate goal we're trying to achieve
JSON array of actions already taken
Result/outcome of the last action
Continue
Goal appears to be achieved
Next step suggestion; target_coordinates are in desktop input coordinates (ready for the mouse nodes) when Frame is connected, otherwise pixels of the original screenshot
Target point: desktop input coordinates when a frame was connected, otherwise screenshot pixels.
Type of suggested action
Target description