Use GLM-4.6V vision model to control browser and complete tasks
- Grounding Capability: Model can precisely identify bounding box coordinates [x1,y1,x2,y2] of elements in screenshots
- Vision Understanding: Understand page content and layout through base64 screenshots
- Automated Operations: Execute clicks, inputs, scrolls and other actions based on visual analysis results
- Screenshot current page → Convert to base64
- Send to GLM-4.6V for analysis
- Model uses Grounding capability to locate elements, returns bbox coordinates
- Calculate bbox center point, execute precise click
- Loop until task completion
-
Make sure you have installed the latest version of Google Chrome.
-
Node.js 20.x or higher is required.
-
Edit
.envfile to set your API key and other configurations.
# LLM model
LLM_MODEL=glm-4.6v
# BigModel API base URL
LLM_API_BASE=https://open.bigmodel.cn/api/paas/v4
# BigModel API key, Go to https://bigmodel.cn/usercenter/proj-mgmt/apikeys to get your API key
LLM_API_KEY=your_bigmodel_api_key
# Your browser platform: windows / macOS / linux
BROWSER_PLATFORM=windows
# Your browser version
BROWSER_VERSION=142.0.7444.162# Install dependencies and build the project
yarn install
yarn build
# Start a browser agent task, set max steps to 100, you can adjust it as needed
yarn start -t "Please play a game of online Minesweeper by yourself until you win the game." -s 100