▶
ExecutionTrigger
Automation/LLM/Vision
Uses LLM to rank multiple element candidates based on match quality
Scores range from 0 to 10. Higher values mean more impact, exposure, or operational weight.
Trigger
Vision-capable LLM model
Screenshot as base64 PNG, JPEG, WebP or GIF (a data URL is fine). Ignored when Image is connected
Screenshot as an image, e.g. from the Screenshot node. Takes precedence over the base64 Screenshot
Screen frame of the screenshot, from the capture node. Coordinates are desktop input coordinates (ready for the mouse nodes) when Frame is connected, otherwise pixels of the original screenshot
Candidate elements to rank, each with a unique id. Optional x/y are in desktop input coordinates (ready for the mouse nodes) when Frame is connected, otherwise pixels of the original screenshot
What the target element should match (description/intent)
Additional context for ranking
Continue
Full ranking result; only given candidate ids appear
ID of the best matching candidate
Candidates sorted by rank