Page Agent: Alibaba's Open-Source In-Page GUI Agent — One Line of Script Lets AI Live Inside Any Web Page

·Toolin Editorial Team

Alibaba's open-source in-page GUI Agent: one CDN script tag or one npm install gives any web page AI natural-language operation. 26.8k stars on GitHub, MIT licensed.

Page Agent: Alibaba's Open-Source In-Page GUI Agent — One Line of Script Lets AI Live Inside Any Web Page

Page Agent Banner

Ever run into this: a creaky internal admin panel, a government services site, a supplier's order system — the UI is stuck a decade in the past, and every operation means a dozen clicks and a stack of forms. How much easier would it be to just say "sort the items in the cart by price" and let the AI do it?

Alibaba's open-source Page Agent does exactly that. The official tagline captures it well: "A GUI agent that lives in your web page — one line of script gives any page its own AI agent." It has 26.8k stars and 2.4k forks on GitHub, MIT licensed, latest version v1.12.2 (released 2026-07-16, 37 releases so far).

Repository: https://github.com/alibaba/page-agent

What Page Agent Is

It's an in-page GUI Agent built on JavaScript (TypeScript is 82.6% of the codebase) that runs in the browser client and manipulates the current page's DOM directly. The official positioning is unambiguous: designed for client-side web enhancement, not server-side automation.

That draws a key distinction:

TypeRepresentative ProjectsControl ScopeDeployment
Browser automationbrowser-use, PlaywrightThe entire browser (multi-tab, cross-origin)Server-side or extension
In-page agentPage AgentThe current page's DOMClient-side JS

Page Agent's DOM-handling components and prompt engineering are derived from the well-known browser-use project (acknowledged explicitly in the official Acknowledgments), but the positioning differs — it solves "giving one specific web page AI operation capability," not "letting AI drive my whole browser."

How It Works

Four core traits (all from the official README):

  • Pure client-side operation: no browser extension, no Python, no headless browser — everything runs inside your web page.
  • Text-based DOM manipulation: no screenshots, no reliance on multimodal models, no special permissions — it converts the DOM into structured text for the model, then translates the model's output back into DOM operations.
  • Bring Your Own LLM: compatible with mainstream models, including locally deployed ones.
  • Optional Chrome extension + MCP Server (Beta): use the extension for multi-tab tasks; the MCP Server lets external agent clients control the browser in reverse.

Typical execution flow:

  1. Parse the DOM: read the current page structure and identify interactive elements (buttons, inputs, links, tables)
  2. Understand semantics: convert the DOM into a structured description the model can understand
  3. Receive natural-language instructions: like "sort the table by date, descending" or "fill in this form, use the default shipping address"
  4. Execute within the page: trigger clicks, typing, navigation, and other operations directly on the DOM

Because it doesn't rely on screenshots, deployment is extremely lightweight, and it sidesteps the latency and cost that multimodal models bring.

Quick Integration: Two Ways

Method 1: One-Line CDN Script (Fastest to Try)

The officially recommended fastest path, running against their free demo LLM:

<script
    src="https://cdn.jsdelivr.net/npm/page-agent@1.12.2/dist/iife/page-agent.demo.js"
    crossorigin="anonymous"
></script>

<!-- If jsDelivr is unreachable in your region, use the mirror -->
<!-- https://registry.npmmirror.com/page-agent/1.12.2/files/dist/iife/page-agent.demo.js -->

⚠️ The official docs state plainly: this demo CDN uses their free test LLM API, for technical evaluation only. To use your own model, add ?autoInit=false so the script doesn't auto-create the demo agent, then instantiate with new window.PageAgent(...) passing in your own LLM configuration.

Method 2: npm Integration (For Building Into Products)

Frontend/full-stack developers who want to integrate it natively into their own product (say, adding an AI operations assistant to a SaaS) go through npm:

npm install page-agent
import { PageAgent } from 'page-agent'

const agent = new PageAgent({
    model: 'qwen3.5-plus',
    baseURL: 'https://dashscope.aliyuncs.com/compatible-mode/v1',
    apiKey: 'YOUR_API_KEY',
    language: 'en-US',
})

await agent.execute('Click the login button')

Configuration supports model / baseURL / apiKey / language, and agent.execute(...) takes a single natural-language task. Page Agent is a monorepo (npm workspaces) with the core under packages/page-agent/, and it ships with a UI Panel.

💡 Production integration tip: when embedding into your own product, watch out for your site's CSP (Content Security Policy) restrictions, and never hardcode LLM API keys in the frontend — for enterprise scenarios, run your own LLM gateway or route through a proxy. Validate in staging before going live.

Officially Listed Use Cases

The README names five concrete use cases, more specific than just "adding AI to old websites":

  • SaaS AI Copilot: embed an AI copilot into your product with a few lines of code, no backend rewrite needed
  • Smart Form Filling: compress a 20-click workflow into one sentence — a fit for ERP / CRM / admin systems
  • Accessibility: make any web page operable through natural language, with voice commands and screen-reader support, at zero barrier
  • Multi-page Agent: extend the in-page agent's reach across multiple browser tabs via the Chrome extension
  • MCP: let your agent client control the browser in reverse

Where it fits poorly: critical business processes that demand 100% determinism (agents are probabilistic; for critical paths, the official API is still the safer route).

The Relationship with browser-use

The most frequently asked question. The official Acknowledgments is clear: Page Agent's DOM-handling components and prompts are derived from browser-use (Gregor Zunic's MIT-licensed project). But the two have different missions — browser-use controls the entire browser, while Page Agent works only within the current page, deployed as client-side JS with no Python, headless browser, or extension required.

Final Thoughts

Page Agent drives the barrier for "AI operating web pages" down to a single line of script, and it's especially friendly to frontend developers — it's also a codebase worth dissecting: how to build a screenshot-free in-page agent with client-side JS + an LLM. Head to the GitHub repository for the README and source; that's where you'll find the most accurate integration guide and demo videos.

Repository: https://github.com/alibaba/page-agent (MIT licensed, npm package name page-agent, latest version v1.12.2)

Honesty note: this article is based on public information from the Page Agent official GitHub repository README (stars / version numbers / code snippets / use cases all come from the repository page, as of 2026-07-18). For specifics on the API and configuration, defer to the repository's latest documentation.