Skip to content

LLM Observe Screen Node

Automation/LLM/Vision

Uses vision LLM to comprehensively observe and describe the current screen

llm_observe_screenautomationLong running
Inputs6
Outputs4
Security exposure4/10
Packageautomation

Ratings

Scores range from 0 to 10. Higher values mean more impact, exposure, or operational weight.

SecurityAttack surface and exposure impact.
4/10Medium
PrivacyPotential sensitivity of processed data.
3/10Low
PerformanceRuntime or resource pressure.
4/10Medium
GovernancePolicy, audit, or compliance impact.
5/10Medium
ReliabilityOperational stability considerations.
7/10High
CostExternal or compute cost impact.
5/10Medium

Input Pins

6

▶

Execution
exec_in

Trigger

Model

Struct
model

Vision-capable LLM model

BitBit19 fields
idstring
default ""
typeBitTypes
enum "Llm", "Vlm", "Tts", "Stt"...default "Other"
metaMap<string, Metadata>
default {}
*Metadatamap value
namestringrequired
descriptionstringrequired
long_descriptionstring | null
release_notesstring | null
tagsArray<string>required
itemsstringarray item
+11 more fields
authorsArray<string>
default []
itemsstringarray item
repositorystring | null
default null
download_linkstring | null
default null
file_namestring | null
default null
hashstring
default ""
sizeinteger | null
format uint64default nullmin 0
hubstring
default ""
parametersvalue
default null
versionstring | null
default null
licensestring | null
default null
dependenciesArray<string>
default []
itemsstringarray item
dependency_tree_hashstring
default ""
createdstring
default ""
updatedstring
default ""
model_slugstring | null
default null
+1 more fields
Schema enforced

Screenshot

String
screenshot

Screenshot as base64 PNG, JPEG, WebP or GIF (a data URL is fine). Ignored when Image is connected

Image

Struct
image

Screenshot as an image, e.g. from the Screenshot node. Takes precedence over the base64 Screenshot

NodeImageNodeImage1 fields
image_refstringrequired

Frame

Struct
frame

Screen frame of the screenshot, from the capture node. Coordinates are desktop input coordinates (ready for the mouse nodes) when Frame is connected, otherwise pixels of the original screenshot

ScreenFrameScreenFrame7 fields
display_indexinteger | null
format uint32default nullmin 0
xinteger:int32required
format int32
yinteger:int32required
format int32
widthinteger:uint32required
format uint32min 0
heightinteger:uint32required
format uint32min 0
pixel_widthinteger:uint32required
format uint32min 0
pixel_heightinteger:uint32required
format uint32min 0

Focus Area

String
focus_area

Specific area or aspect to focus on (optional)

Output Pins

4

▶

Execution
exec_out

Continue

Observation

Struct
observation

Complete screen observation

ScreenObservationScreenObservation6 fields
descriptionstringrequired
app_contextstringrequired
interactive_elementsArray<ObservedElement>required
itemsObservedElementarray item
element_typestringrequired
descriptionstringrequired
approximate_locationstringrequired
is_interactivebooleanrequired
current_statestring | null
+2 more fields
text_contentArray<string>
default []
itemsstringarray item
notable_featuresArray<string>
default []
itemsstringarray item
possible_actionsArray<string>
default []
itemsstringarray item

Description

String
description

Text description of the screen

Elements

Struct Array
elements

Observed elements. Optional x/y is the element's center in desktop input coordinates (ready for the mouse nodes) when Frame is connected, otherwise pixels of the original screenshot

ObservedElementObservedElement7 fields
element_typestringrequired
descriptionstringrequired
approximate_locationstringrequired
is_interactivebooleanrequired
current_statestring | null
xinteger | null

Center x: desktop input coordinates when a frame was connected, otherwise screenshot pixels. Absent when the model gave no point on the screenshot.

format int32default null
yinteger | null

Center y, in the same space as `x`.

format int32default null

Node Info

Internal name
llm_observe_screen
Category
Automation/LLM/Vision
Version
5