Docs
Snapshots and refs
How the agent reads a page and picks elements to act on.
Instead of reading raw HTML, the agent asks for a snapshot: a compact, indented view of the page's accessibility tree. Interactive elements carry a ref such as @21 that the agent can act on directly.
Reading a page
const page = task.page("p1");
console.log(await page.snapshot()); // current viewport
console.log(await page.snapshot({ scope: "full_page" })); // whole page
Acting on refs
await page.fill("@21", "user@example.com");
await page.click("loc=role:button[name='Sign in']");
await page.waitForSelector("loc=css:#account-home", { state: "visible" });
console.log(await page.snapshot());
Element actions accept snapshot refs (@21), text=…, loc=css: / loc=role: / loc=href: locators, xpath=… and plain CSS selectors. A selector must match exactly one element.
After the page changes, refs from the old snapshot may no longer be valid, so the agent takes a new snapshot before the next step.
When a snapshot isn't enough
Canvas apps, maps and some editors don't expose useful page structure. There the agent uses page.screenshot() with mouse and keyboard actions, then checks the result with another screenshot.