Midscene.js

Project lead · Open source · July 2024 – October 2025

Midscene brings multimodal models into UI automation. Developers describe what to do, what to extract and what to verify in natural language, while retaining the integrations and debugging tools needed for an engineering workflow.

Midscene product illustration
Product illustration from the official Midscene website. It illustrates the project; my contributions below refer to my 2024–2025 tenure.

From intent to UI interaction

Traditional E2E scripts often depend on selectors that need maintenance as an interface evolves. Visual and semantic checks can also be difficult to express through DOM state alone.

I led the planning-and-execution architecture, connecting multimodal task understanding with concrete UI actions and feedback. YAML, JavaScript SDK and Playwright integrations made the framework usable within existing test projects.

Midscene planning and execution workflow
A simplified view of the planning, execution and inspection workflow.

Making automation observable

I developed Chrome tooling for recording, replay and script generation, and worked on reports that make actions, page state and extraction results inspectable. Caching reduces repeated inference work, while replayable reports help explain where a task diverged from the intended behavior.

The framework was adopted in ByteDance business workflows and by external companies for UI automation and regression testing.

GUI agents and model engineering

Working with the algorithm team, I contributed to UI-TARS engineering and the GUI Agent implementation. I also built rollout data pipelines to collect successful and failed interaction traces and connect them to annotation workflows.

I am a co-author of UI-TARS: Pioneering Automated GUI Interaction with Native Agents, listed as Xiao Zhou. My Feday talk covers the UI automation architecture and the engineering work behind it.

GitHub · Documentation · UI-TARS paper