Headless browser automation for terminals and AI agents.
w3txt converts live web pages into structured text — Markdown with indexed action tokens — so that shell scripts, LLM agents, and humans in a terminal can read and interact with any website without a graphical browser.
$ w3txt open "https://example.com"
$ w3txt state
# Example Domain
URL: https://example.com
This domain is for use in illustrative examples. More information... ⟦0⟧
## Actions
Links:
[0] More information... (link)
$ w3txt click 0 # clicks the link at index 0
$ w3txt input 2 "hello" # types into the input at index 2
$ w3txt keys "Enter" # presses Enter
- Single binary, no runtime. Rust. No Python, no Node.js.
- Text-first output. Every page becomes Markdown + a numbered action
list. Pipe
w3txt state --jsoninto any language. - Session server. The first command starts a background server; subsequent commands are fast IPC round-trips. Browser state (cookies, login, tabs) persists across commands.
- TUI browser.
w3txt tuiopens a BBS-style terminal UI for interactive browsing. - AI-agent friendly. The indexed action tokens (
⟦0⟧ ⟦1⟧ ...) let an LLM read a page and choose actions by number, no DOM knowledge required.
# Auto-download (recommended):
w3txt setup
# Or manually: install Google Chrome / Chromium and chromedriver,
# then ensure `chromedriver` is on your PATH.git clone <repo-url>
cd w3txt
cargo build --release
# binary: target/release/w3txt# Open a page
w3txt open "https://news.ycombinator.com"
# See the page as text + actions
w3txt state
# Click link #5
w3txt click 5
# JSON output for scripting
w3txt state --jsonw3txt tuiKey bindings: numbers to click, g to go to URL, b/f for back/forward,
v to toggle action tokens, / to search, : for commands, ? for help,
q to quit.
| Command | Description |
|---|---|
open <url> |
Navigate to a URL |
state |
Get current page as Markdown + actions |
click <index> |
Click an element |
input <index> <text> |
Clear a field and type text |
type <text> |
Type into the focused element |
keys <combo> |
Send key combination (e.g. Enter, Control+a) |
back / forward |
Browser navigation |
scroll <up|down> |
Scroll the page |
hover <index> |
Hover over an element |
dblclick <index> |
Double-click |
rightclick <index> |
Right-click |
select <index> <value> |
Select a dropdown option |
screenshot [path] |
Capture PNG (file or base64) |
wait text <text> |
Wait for text to appear |
wait selector <css> |
Wait for a CSS selector |
get title|html|text|value|attributes|bbox |
Extract page data |
eval <script> |
Execute JavaScript |
switch <tab> |
Switch browser tab |
close-tab [tab] |
Close a tab |
cookies get|set|clear|export|import |
Cookie management |
server status |
Show server info and sessions |
server stop |
Stop the background server |
setup |
Download Chrome + chromedriver |
tui |
Launch the terminal UI |
Full details: docs/commands.md
| Flag | Description |
|---|---|
--json |
Machine-readable JSON output |
--session <name> |
Use a named session (default: default) |
--headed |
Show the browser window |
--log-level <level> |
Log verbosity (error, warn, info, debug, trace) |
--idle-timeout <secs> |
Auto-stop server after N seconds of inactivity |
Optional config file at ~/.w3txt/config.toml:
chromedriver_path = "/usr/local/bin/chromedriver"
chrome_path = "/usr/bin/google-chrome"
window_width = 1280
window_height = 720
default_timeout_ms = 30000
idle_timeout_secs = 0
auto_screenshot_on_error = false
log_level = "info"Environment variables (W3TXT_CHROMEDRIVER_PATH, W3TXT_CHROME_PATH, etc.)
override the config file. CLI flags override everything.
w3txt CLI ──TCP──▶ Session Server ──HTTP──▶ chromedriver ──CDP──▶ Chrome
│
PageSnapshot
(markdown + actions)
- The CLI sends JSON-RPC requests over a local TCP socket.
- The server manages browser sessions via W3C WebDriver.
- On
state, JavaScript is injected to collect visible interactive elements and convert the DOM to Markdown with⟦index⟧action tokens. - Actions (
click,input, etc.) target elements by their snapshot index.
See docs/architecture.md for details.
The core abstraction: every web page becomes a PageSnapshot containing Markdown text and a flat list of indexed interactive elements.
# Page Title
Some paragraph text with a login link ⟦0⟧ and a search box ⟦1⟧.
## Actions
Links:
[0] Login (link)
Inputs:
[1] Search (input)
Buttons:
[2] Submit (button)
The ⟦N⟧ tokens (U+27E6 / U+27E7) appear inline where the element exists
in the document flow. The Actions section at the bottom provides a
grouped summary.
See docs/snapshot-format.md for the full specification.
w3txt/
├── Cargo.toml # Workspace
├── crates/
│ ├── w3txt-core/ # Core library
│ │ └── src/
│ │ ├── chromedriver.rs # ChromeDriver process management
│ │ ├── webdriver.rs # W3C WebDriver HTTP client
│ │ ├── session.rs # Browser session + snapshot extraction
│ │ ├── server.rs # JSON-RPC TCP server
│ │ ├── snapshot.rs # PageSnapshot / IndexedElement types
│ │ ├── protocol.rs # JSON-RPC message types
│ │ ├── config.rs # Configuration loading
│ │ ├── setup.rs # Auto-download Chrome + chromedriver
│ │ ├── error.rs # Error types + exit codes
│ │ └── paths.rs # Data/runtime directory paths
│ └── w3txt/ # CLI + TUI binary
│ └── src/
│ ├── main.rs # Entry point + command dispatch
│ ├── cli.rs # clap CLI definition
│ ├── client.rs # TCP client (auto-starts server)
│ ├── output.rs # Human-readable formatters
│ └── tui/ # Terminal UI (ratatui)
│ ├── mod.rs # Event loop + state machine
│ ├── render.rs # Screen rendering
│ └── worker.rs # Background server communication
└── docs/
├── architecture.md
├── snapshot-format.md
└── commands.md
- Rust 1.75+ (for building)
- chromedriver + Chrome (or chrome-headless-shell)
- Run
w3txt setupto auto-install both
- Run
MIT