> ## Documentation Index
> Fetch the complete documentation index at: https://tbd-6fc993ce-hypeship-docs-ia-v2.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# How You Drive the Browser

> Choose a control surface and where your agent loop runs

You make two choices before you write any automation. They're independent, but the first constrains the second:

1. **How you drive the browser** — the control surface your code or model uses to act on the page.
2. **Where the loop runs** — the machine your decision-making code runs on, relative to the browser.

## 1. How you drive the browser

Kernel browsers accept four control surfaces. Pick by what's driving the page, not by what you already know.

| Surface | Use it when | Trade-off |
| - | - | - |
| [Playwright execution](/browsers/playwright-execution) | **Default.** You know what to do on the page — navigate, fill, extract, upload. | Needs a selector or DOM path that exists. |
| [Computer controls](/browsers/computer-controls) | **Recommended fallback.** A model is looking at pixels, or the page can't be driven programmatically. | Slower per step, and the model has to see the state to act. |
| CDP | You have an existing Playwright, Puppeteer, or CDP codebase to point at Kernel. | Adds a protocol fingerprint and a network hop. |
| WebDriver BiDi | You need the W3C standard protocol. | Smaller client ecosystem. |

For agents, start with [playwright execution with a computer use fallback](/browsers/playwright-computer-use-fallback): script the deterministic steps, and hand the page to a computer use model when a step doesn't respond to a selector.

### Why the choice matters on hardened sites

CDP is what Playwright and Puppeteer speak, and anti-bot vendors scan for its signatures. Computer controls carry no CDP connection, so there's no protocol fingerprint to leak. That makes them the stronger option on sites with aggressive detection, and it's why [managed auth](/auth/managed-auth) drives logins with coordinate-based input rather than CDP. How much this matters is site-specific, so test before you commit — see [bot anti-detection](/browsers/bot-detection/overview).

### Control surface examples

<Tabs>
  <Tab title="Computer Use">
    Kernel's [Computer Controls](/browsers/computer-controls) API exposes OS-level mouse, keyboard, and screen primitives — the surface a computer-use model already knows how to drive (screenshot, click, type, key, scroll, drag). No CDP or WebDriver connection required, so there's no protocol fingerprint to leak. Ideal for [Claude](/integrations/computer-use/anthropic), [OpenAI](/integrations/computer-use/openai), or [Gemini](/integrations/computer-use/gemini) computer-use loops.

    <CodeGroup>
      ```typescript Typescript/Javascript theme={null}
      import Kernel from '@onkernel/sdk';

      const kernel = new Kernel();
      const kernelBrowser = await kernel.browsers.create();

      const screenshot = await kernel.browsers.computer.captureScreenshot(kernelBrowser.session_id);

      await kernel.browsers.computer.clickMouse(kernelBrowser.session_id, {
        x: 420,
        y: 280,
      });

      await kernel.browsers.computer.typeText(kernelBrowser.session_id, {
        text: 'kernel cloud browsers',
      });
      ```

      ```python Python theme={null}
      from kernel import Kernel

      kernel = Kernel()
      kernel_browser = kernel.browsers.create()

      screenshot = kernel.browsers.computer.capture_screenshot(id=kernel_browser.session_id)

      kernel.browsers.computer.click_mouse(
          id=kernel_browser.session_id,
          x=420,
          y=280,
      )

      kernel.browsers.computer.type_text(
          id=kernel_browser.session_id,
          text="kernel cloud browsers",
      )
      ```

      ```go Go theme={null}
      package main

      import (
      "context"

      "github.com/kernel/kernel-go-sdk"
      )

      func main() {
      ctx := context.Background()
      client := kernel.NewClient()

      kernelBrowser, err := client.Browsers.New(ctx, kernel.BrowserNewParams{})
      if err != nil {
      	panic(err)
      }

      screenshot, err := client.Browsers.Computer.CaptureScreenshot(
      	ctx,
      	kernelBrowser.SessionID,
      	kernel.BrowserComputerCaptureScreenshotParams{},
      )
      if err != nil {
      	panic(err)
      }
      defer screenshot.Body.Close()

      if err := client.Browsers.Computer.ClickMouse(
      	ctx,
      	kernelBrowser.SessionID,
      	kernel.BrowserComputerClickMouseParams{
      		X: 420,
      		Y: 280,
      	},
      ); err != nil {
      	panic(err)
      }

      if err := client.Browsers.Computer.TypeText(
      	ctx,
      	kernelBrowser.SessionID,
      	kernel.BrowserComputerTypeTextParams{
      		Text: "kernel cloud browsers",
      	},
      ); err != nil {
      	panic(err)
      }
      }
      ```
    </CodeGroup>
  </Tab>

  <Tab title="Playwright Execution">
    Run any Playwright code from anywhere — no local Playwright install, no Chromium download, no CDP connection to manage. Your code executes inside the browser's VM with the full Playwright API in scope and returns structured data back to your agent. Ships with [Patchright](/browsers/bot-detection/stealth) by default.

    <CodeGroup>
      ```typescript Typescript/Javascript theme={null}
      const response = await kernel.browsers.playwright.execute(
        kernelBrowser.session_id,
        {
          code: `
            await page.goto('https://example.com');
            return await page.title();
          `,
        },
      );

      console.log(response.result);
      ```

      ```python Python theme={null}
      response = kernel.browsers.playwright.execute(
          id=kernel_browser.session_id,
          code="""
            await page.goto('https://example.com')
            return await page.title()
          """,
      )

      print(response.result)
      ```

      ```go Go theme={null}
      response, err := client.Browsers.Playwright.Execute(
      ctx,
      kernelBrowser.SessionID,
      kernel.BrowserPlaywrightExecuteParams{
      	Code: `
            await page.goto('https://example.com');
            return await page.title();
          `,
      },
      )
      if err != nil {
      panic(err)
      }

      fmt.Println(response.Result)
      ```
    </CodeGroup>
  </Tab>

  <Tab title="CDP">
    Chrome DevTools Protocol — the wire format Playwright, Puppeteer, and most browser frameworks speak. Use `cdp_ws_url` from the created browser session for deterministic, scripted automation driven from your own infra.

    <CodeGroup>
      ```typescript Typescript/Javascript theme={null}
      import { chromium } from 'playwright';

      const browser = await chromium.connectOverCDP(kernelBrowser.cdp_ws_url);
      const context = browser.contexts()[0];
      const page = context.pages()[0];

      await page.goto('https://example.com');
      const title = await page.title();
      console.log(title);
      ```

      ```python Python theme={null}
      from playwright.async_api import async_playwright

      async with async_playwright() as playwright:
          browser = await playwright.chromium.connect_over_cdp(kernel_browser.cdp_ws_url)
          context = browser.contexts[0]
          page = context.pages[0]

          await page.goto('https://example.com')
          title = await page.title()
          print(title)
      ```
    </CodeGroup>
  </Tab>

  <Tab title="WebDriver BiDi">
    W3C-standard browser control. Use `webdriver_ws_url` with [Vibium](/integrations/vibium) or any other BiDi client.

    <CodeGroup>
      ```typescript Typescript/Javascript theme={null}
      import { browser } from 'vibium';

      const bro = await browser.start(kernelBrowser.webdriver_ws_url);
      const page = await bro.page();

      await page.goto('https://example.com');
      const title = await page.title();
      console.log(title);
      ```

      ```python Python theme={null}
      from vibium.sync_api import browser

      bro = browser.start(kernel_browser.webdriver_ws_url)
      page = bro.page()

      page.goto('https://example.com')
      title = page.title()
      print(title)
      ```
    </CodeGroup>
  </Tab>
</Tabs>

## 2. Where the loop runs

Your loop is whatever decides the next action: a script, an agent, or a model. It can run in three places.

<Tabs>
  <Tab title="Your own infrastructure">
    Connect to `cdp_ws_url` or `webdriver_ws_url` from wherever your code already runs. Any CDP client works, and there's no lock-in.

    **Costs:** a network round trip per action, disconnects to handle, screenshot and DOM bandwidth, and the CDP fingerprint above. It's fine for low-frequency or deterministic work, and it hurts most in a vision loop.
  </Tab>

  <Tab title="Playwright execution API">
    Send code, not commands. Each call runs in the browser's VM against the live session, so state carries across calls and an agent can drive the page turn by turn — one tool call per step, structured data back.

    **Costs:** the code you send has to be self-contained per call. There's nothing to install and no connection to manage.
  </Tab>

  <Tab title="Code execution platform">
    Deploy the whole agent next to the browser with the [code execution platform](/apps/develop), invoked on demand or on a schedule, with no infrastructure of your own.

    **Costs:** your agent has to be deployable as a Kernel app. It's worth it once the automation is long-running, stateful, or triggered by events rather than by a person.
  </Tab>
</Tabs>

### Where computer use fits

A computer use agent answers the first question, not the second — it still has to run its loop somewhere. Because every turn ships a screenshot instead of a small script, running that loop off-platform costs far more than it does for a Playwright-driven agent: you pay image bandwidth and a round trip on every step. That makes computer use the strongest case for running your loop next to the browser. Model inference stays with the model vendor either way.

### Putting it together

| Your automation | Control surface | Where the loop runs |
| - | - | - |
| Scheduled scrape of a known page | Playwright execution | Anywhere — one call, one result |
| Agent doing multi-step work on a normal site | Playwright execution, computer use fallback | Playwright execution API, or the code execution platform once it's long-running |
| Agent on a site with aggressive detection | Computer controls | Code execution platform |
| Existing Playwright suite you're migrating | CDP | Your own CI, then move hot paths to playwright execution |

## Why computer use for agents

Kernel's computer controls are built to match how computer-use models were trained — the same primitives the model emits (screenshot, click at coords, type, key, scroll, drag) map 1:1 onto the API. There's no harness translating model output into framework calls.

* **Native fit.** Screenshot, click, type, key, scroll, drag — the primitives the model already speaks.
* **Faster screenshots.** Captures bypass CDP, which removes the largest source of latency in a vision loop.
* **Better against bot detection.** No CDP connection means no CDP fingerprint to leak. Pairs naturally with [stealth mode](/browsers/bot-detection/stealth) and [residential proxies](/proxies/residential).
* **Human-like input.** OS-level events with Bézier-curve mouse paths, variable typing speed, and configurable mistype rate.
* **Not DOM-limited.** Screenshots capture the full VM, so the agent can see and interact with native dialogs, canvas elements, iframes, and PDFs — not just things you can address with a selector.

## Why playwright execution over a direct CDP connection

If you're reaching for Playwright, prefer the execution API over `connectOverCDP`. Same Playwright API you already know, none of the setup.

* **Run from anywhere.** No `playwright` package to version-pin, no Chromium download, no CDP connection to manage. Send the code, get the result.
* **Co-located with the browser.** Code runs in the same VM as the browser — no network hop between your script and the page, fewer flakes.
* **Patchright by default.** Hardened against bot detection out of the box.
* **Full Playwright API.** `page`, `context`, and `browser` are all in scope. Anything Playwright can do — DOM queries, file uploads, full-page screenshots — works here.
* **Returns values.** `return` from your code and the result comes back in the response. Easy to use as an agent tool.

## Computer use + playwright execution

Computer controls drive the browser the way a person would — they don't speak the programmatic API surface. Anything you'd reach for the DOM or Playwright client for (reading text and attributes, `page.goto`, file uploads, cookie or storage access, switching tabs) belongs on the [playwright execution](/browsers/playwright-execution) side. When computer use is driving, expose playwright execution to the agent as a tool it can call for structured data or a programmatic action. For the full pattern in the other direction — playwright execution first, computer use when a step doesn't respond to a selector — see [playwright with computer use fallback](/browsers/playwright-computer-use-fallback).

<CodeGroup>
  ```typescript Typescript/Javascript theme={null}
  const response = await kernel.browsers.playwright.execute(
    kernelBrowser.session_id,
    {
      code: `
        const rows = await page.$$eval('table tr', (trs) =>
          trs.map((tr) => Array.from(tr.querySelectorAll('td')).map((td) => td.textContent))
        );
        return rows;
      `,
    },
  );

  console.log(response.result);
  ```

  ```python Python theme={null}
  response = kernel.browsers.playwright.execute(
      id=kernel_browser.session_id,
      code="""
        const rows = await page.$$eval('table tr', (trs) =>
          trs.map((tr) => Array.from(tr.querySelectorAll('td')).map((td) => td.textContent))
        );
        return rows;
      """,
  )

  print(response.result)
  ```

  ```go Go theme={null}
  response, err := client.Browsers.Playwright.Execute(
  	ctx,
  	kernelBrowser.SessionID,
  	kernel.BrowserPlaywrightExecuteParams{
  		Code: `
  			const rows = await page.$$eval('table tr', (trs) =>
  				trs.map((tr) => Array.from(tr.querySelectorAll('td')).map((td) => td.textContent))
  			);
  			return rows;
  		`,
  	},
  )
  if err != nil {
  	panic(err)
  }

  fmt.Println(response.Result)
  ```
</CodeGroup>

## Going deeper

* [Computer Controls reference](/browsers/computer-controls) — every mouse, keyboard, and screen primitive.
* [Playwright Execution reference](/browsers/playwright-execution) — the full execution surface, return values, and timeouts.
* [Computer use integrations](/integrations/computer-use/anthropic) — drop-in examples for Anthropic, Gemini, OpenAI, and more.
* [Network access](/info/network-access) lists the domains and ports to allow for API, CDP, and WebDriver BiDi connections.
