▶
ExecutionTrigger
Automation/Computer/Agent
Lets a vision model operate the desktop until a goal is reached: it looks at a screenshot (optionally with numbered accessibility marks), calls mouse and keyboard tools, waits for the screen to settle, checks the result and repeats. Works with any vision model that supports tool calling. Ends when the model reports done or asks the user, or when it is stuck or out of steps or time
Scores range from 0 to 10. Higher values mean more impact, exposure, or operational weight.
Trigger
Computer session handle
Browser context info if browser is attached
Current page info if a page is open
Visible text contained in the element; script, style and head content is ignored.
ARIA role, optionally filtered by accessible name: `button`, `button|Sign in` (case-insensitive substring), `button|=Sign in` (exact) or `button|/^sign/i` (regex).
Element ref from the latest Browser Snapshot, such as `e12`.
Vision model with tool calling that operates the desktop
The task in plain language, e.g. 'Rename report.txt on the Desktop to final.txt'
Display the agent sees and acts on; -1 is the primary display
Optional window title (or part of it): the agent sees the display showing this window and its elements are the numbered ones. Empty uses the focused window outside Flow-Like
screenshot: the plain screenshot; screenshot_marks: numbered boxes on accessibility elements plus an element list; screenshot_marks_ocr: also number recognized text
Most model turns before the agent stops with max_steps (1–500)
Wall-clock budget in seconds before the agent stops with timeout (1–86400)
Most tool calls executed per model turn (1–20)
Longest wait after actions for the screen to stop changing before the next screenshot (0–10000)
Regular expressions the agent may never type; a matching type call is refused
Optional guidance added to the task, e.g. which app to use or what to avoid
The model reported the goal as achieved
The model reported failure, the agent was stuck, ran out of steps or time, or hit an error
The model asked the user a question; it is in Answer
Computer session handle (pass-through)
Browser context info if browser is attached
Current page info if a page is open
Visible text contained in the element; script, style and head content is ignored.
ARIA role, optionally filtered by accessible name: `button`, `button|Sign in` (case-insensitive substring), `button|=Sign in` (exact) or `button|/^sign/i` (regex).
Element ref from the latest Browser Snapshot, such as `e12`.
done, failed, stuck, max_steps, timeout or needs_input
The model's result or summary, the question for the user, or why the agent stopped
Trajectory: per turn the screenshot seen, the model's text, each action with its arguments, desktop coordinates and result, and the duration
1-based turn number.
The screenshot the model looked at.
Text the model wrote next to its tool calls.
One tool call of a turn.
Arguments as the model sent them; coordinates are pixels of its screenshot.
Desktop input coordinates the action used.
Not run: an earlier call of the turn failed, the turn limit was reached or the task ended.
The last screenshot, without marks
Desktop rectangle and pixel size of the final screenshot