Skip to Content
View all events

Harness Engineering for Browser Agents

Published by O'Reilly Media, Inc.

Intermediate content levelIntermediate

Give agents the capability to see, click, and act safely on the web

What you’ll learn and how you can apply it

  • Understand the browser harness end to end
  • Build and operate a browser harness
  • Reason about safety, reliability, and cost

Course description

The web is a high-leverage environment for agents, yet tasks like booking, buying, or researching with a browser often require reliable, secure execution. Sajal Sharma demystifies a browser agent’s “harness,” the code that manages page perception, executes actions, and ensures security. You’ll build a browser harness yourself and construct a perception-to-action tool set using the Chrome DevTools Protocol. You’ll also implement production-grade features, including verification, recovery, guardrails, and session isolation. You’ll finish by mounting a production browser tool into the same harness, running a task at capability, and mapping each step back to the tool set you built.

This live event is for you because...

  • You’re an AI engineer who’s building agents, and you want to understand how the browser layer works under the hood.
  • You’re a software engineer integrating browser automation or web agents into a product and want to build, debug, and extend the harness with confidence.
  • You’re a technical builder evaluating browser-agent tools, or running a site that agents increasingly visit, and you want enough internals to choose well, operate safely, and see your pages the way an agent does.

Prerequisites

  • A Python environment—Python 3.12+ with pip/uv available to install packages
  • Node.js, to install the browser tools
  • An active Anthropic API key with access to Claude models
  • A code editor (VS Code recommended) and terminal access
  • Comfort with Python and the command line
  • A basic understanding of LLMs and AI agents (the agent loop, tools, and tool calls)
  • Familiarity with how web pages are structured (HTML and the DOM)
  • Familiarity with API keys and JSON configuration

Recommended preparation:

  • The prewired agent loop, tool stubs, and all demo code will be shared before the course

Recommended follow-up:

  • Watch Building AI Agents with LangGraph (on-demand course)
  • Read AI Agents: The Definitive Guide (book)
  • Take Advanced Harness Engineering (live online course with Richmond Alake)
  • Take Getting Started with Claude Agent SDK (live online course with Sajal Sharma)

Schedule

The time frames are only estimates and may vary according to how the class is progressing.

The machine and its interaction surfaces (45 minutes)

  • Presentation: How a computer-use agent works as a loop (perceives the screen, grounds an action, executes it, verifies the result, and recovers when the GUI does not provide a clean success signal); how screenshots, coordinates, resolution, and context management affect reliability and cost; why the desktop should be treated as one environment with multiple surfaces (GUI, shell, and files)
  • Hands-on exercises: Wire the computer tool into a prebuilt agent harness; capture screenshots, dispatch click/type/key actions
  • Q&A
  • Break

The shell, the files, and routing (40 minutes)

  • Presentation: Routing work across the shell, filesystem, and GUI; when scripts are cheaper and more reliable than pixels; how command output and file checks become verification signals; how cross-surface recovery catches failures the screen alone can hide
  • Hands-on exercises: Add shell and filesystem tools and implement routing capabilities between different interaction surfaces
  • Q&A
  • Break

Production, safety, and evaluation (35 minutes)

  • Presentation: What changes when a model can drive a whole machine (blast radius, prompt injection through pixels and files, network boundaries, human approval gates, trace logging, and deterministic verification); evaluating computer-use agents with small task suites and public benchmarks without mistaking leaderboard scores for production readiness
  • Hands-on exercise: Run an end-to-end receipts-to-report workflow in a cloud computer
  • Q&A

Your Instructor

Sajal Sharma

Sajal Sharma is an AI engineer and technology leader with over eight years of experience in AI/ML, specializing in natural language processing. He works at Liminal, a venture studio in Singapore, where he focuses on building AI-first products and shaping technology strategies. Previously, he led AI initiatives at various consulting and product companies, developing innovative AI solutions across industries. His on-demand course, Building AI Agents with LangGraph, is available on the O’Reilly learning platform. Sajal has delivered a guest lecture at Yale University and has been a mentor for Udacity and the University of Melbourne, where he guided students in machine learning and AI. He holds a master’s degree in information technology from the University of Melbourne.

Skill covered

Security Engineering