Vinlic/glm-agent-browser

Use GLM-4.6V vision model to control browser and complete tasks

★ 11Forks 0TypeScriptGitHub ↗Compare

README

GLM Agent Browser

Use GLM-4.6V vision model to control browser and complete tasks

GLM Agent Browser

Features

  • Grounding Capability: Model can precisely identify bounding box coordinates [x1,y1,x2,y2] of elements in screenshots
  • Vision Understanding: Understand page content and layout through base64 screenshots
  • Automated Operations: Execute clicks, inputs, scrolls and other actions based on visual analysis results

Workflow

  1. Screenshot current page → Convert to base64
  2. Send to GLM-4.6V for analysis
  3. Model uses Grounding capability to locate elements, returns bbox coordinates
  4. Calculate bbox center point, execute precise click
  5. Loop until task completion

Configuration

  1. Make sure you have installed the latest version of Google Chrome.

  2. Node.js 20.x or higher is required.

  3. Edit .env file to set your API key and other configurations.

# LLM model
LLM_MODEL=glm-4.6v
# BigModel API base URL
LLM_API_BASE=https://open.bigmodel.cn/api/paas/v4
# BigModel API key, Go to https://bigmodel.cn/usercenter/proj-mgmt/apikeys to get your API key
LLM_API_KEY=your_bigmodel_api_key
# Your browser platform: windows / macOS / linux
BROWSER_PLATFORM=windows
# Your browser version
BROWSER_VERSION=142.0.7444.162

Usage

# Install dependencies and build the project
yarn install
yarn build
# Start a browser agent task, set max steps to 100, you can adjust it as needed
yarn start -t "Please play a game of online Minesweeper by yourself until you win the game." -s 100

Contributors

Vinlic

Issues