▶
ExecutionTrigger
Automation/Computer/Vision
Recognizes text in an image or on the screen with the operating system's OCR engine (Apple Vision, Windows OCR, Tesseract on Linux). Lines carry pixel boxes and, when the screen frame is known, desktop coordinates for the mouse nodes
Scores range from 0 to 10. Higher values mean more impact, exposure, or operational weight.
Trigger
Computer session handle
Browser context info if browser is attached
Current page info if a page is open
Visible text contained in the element; script, style and head content is ignored.
ARIA role, optionally filtered by accessible name: `button`, `button|Sign in` (case-insensitive substring), `button|=Sign in` (exact) or `button|/^sign/i` (regex).
Element ref from the latest Browser Snapshot, such as `e12`.
Image to read; leave unconnected to capture the window, region or display below
Screen frame of the connected image; maps recognized boxes to desktop coordinates
Look only inside the window with this title; empty uses the region or display
Display to capture when no window or region is given
Left edge of the region in desktop input coordinates
Top edge of the region in desktop input coordinates
Region width; 0 captures the whole display
Region height; 0 captures the whole display
Comma-separated languages such as en-US, de; empty lets the OS engine choose
Let the engine correct words with its language model (Apple Vision); turn off for codes and identifiers
Drop lines below this confidence (0–1); engines without confidence keep all lines
Continue
Computer session handle (pass-through)
Browser context info if browser is attached
Current page info if a page is open
Visible text contained in the element; script, style and head content is ignored.
ARIA role, optionally filtered by accessible name: `button`, `button|Sign in` (case-insensitive substring), `button|=Sign in` (exact) or `button|/^sign/i` (regex).
Element ref from the latest Browser Snapshot, such as `e12`.
All recognized lines in reading order, separated by newlines
Recognized lines with confidence, pixel boxes, desktop boxes and word boxes
A rectangle in image pixels, origin top-left.
A rectangle in desktop input coordinates, the space the mouse nodes use.
One recognized word. `bbox` is filled when the image's screen frame is known.
0–1; absent when the engine does not report confidence (Windows).
A rectangle in image pixels, origin top-left.
A rectangle in desktop input coordinates, the space the mouse nodes use.
Screen frame of the recognized image; empty when an image without frame was read