This repository holds research notes for a planned open-source project: an agent that operates a Linux machine by understanding the full UI. The research answers one question: what would state-of-the-art Linux computer use look like if we started building it today?
The research is done. No code exists yet. Read REPORT.md for the final answer. Read INDEX.md for the map of evidence.
An agent that:
- Sees the screen (pixels) and, when the desktop exposes it, the accessibility tree (AT-SPI2).
- Acts with mouse, keyboard, and accessibility actions.
- Works on live Wayland sessions, X11 sessions, and headless CI containers.
- Verifies that its actions worked before it claims success.
- Runs from open components under permissive licenses.
Date of research: 2026-09-18. Two passes ran with parallel subagents:
- Eight researchers: three read local projects in
~/(jev,smelt,decision-model-benchmark,uhm,fly, and others). Five searched the web: model landscape, Linux control stack, OSS frameworks, training methods, hybrid perception. - Three architects designed the project from the evidence. One red team attacked their proposals. One critic audited the notes for gaps.
Every number in the notes carries its source. Where a fact could not be verified from a primary source, the notes say so. Leaderboard rows were read on 2026-09-18 and will drift.
REPORT.md Final report: what SOTA Linux computer use is, and the plan
INDEX.md Map of every note with one-line summaries
notes/ 00-08: the evidence base (details, numbers, sources)
panel/ Raw panel outputs: three architects, red team, critic (JSON)
experiments/ Validation experiments, starting with the AT-SPI truth audit
report-html/ Standalone HTML export of REPORT.md (published to AXINBOX)
tools/ Helper scripts used to build this repository
~/jev— measured study of the Jev (TypeSafe System One) classifier for session-tree labeling. Cost floors, confidence gates, cache rules.~/smelt— teacher-student distillation for local UI detection. Machine-checked size and latency budgets, honest status reporting.~/decision-model-benchmark— Jev versus eight LLMs on typed decisions. Tells us when a cheap classifier beats a language model.~/fly,~/fly-vision-experiment— pointer control from pixels, with an honest negative result (0.0077 correct-click rate against a 0.95 gate).