I have been working with our customers @browserbase for over a year now. It feels like living through a cambrian explosion, and I have collected a whole book worth of stories. There’s unbounded creativity in how the world adopts new technology, and it is in full effect with computer-use. Every week, I’m surprised by something that our customers have built with @browserbase , and I’m excited to share some of those stories with y’all!
We are seeing computer-use cross the chasm in real time, as more of the early majority is exposed to this voodoo magic. But before we get carried away too far into the future, I want to recap and share about how this all started.
The Innovators
October '24: Anthropic published about their newest ability: computer-use in Sonnet 3.5
Claude can now use computers. The latest version of Claude 3.5 Sonnet can, when run through the appropriate software setup, follow a user’s commands to move a cursor around their computer’s screen, click on relevant locations, and input information via a virtual keyboard, emulating the way people interact with their own computer. We think this skill—which is currently in public beta—represents a significant breakthrough in AI progress.
I remember seeing this release, and wondering how this will spread like wildfire in 3 months. But, when I started tinkering with it, I realized it was terrible at doing anything meaningful yet.
Fast forward, Jan 2025: OpenAI releases a research preview of Operator.
Operator got people taking attention at computer-use, and I was so excited when I got access to this on my 200$ Pro plan. I told Operator to join a Google Meet with my friends, and it started commenting in the meeting. Black magic.
Two days later, Browserbase shipped Open Operator.
This is when I started following Browserbase closely, and I remember thinking how this is the beginning of a new modality of AI interaction. In my mind, it was inevitable that this technology will have lasting implications on knowledge work, where almost all my daily work is done inside a browser.
In early '25, coding agents had gone exponential and developers started to realize how this would translate to better computer-use. Websites are basically code (HTML & Javascript), and you could let a coding agent manipulate a website by writing more code.
A lot of signs were pointing me towards this being the the most fun next big thing, and so I joined Browserbase in June '25.
The Early Adopters
Mid '25: Browser Agents started to become a category of its own. Customers would come up to us in zoom calls and show off their latest demos. After a day of customer conversations, I would leave brimming with ideas on how our infrastructure will power this new wave of applications.
I saw enterprise customers using @Stagehanddev to power their "agentic crawling" efforts for KYC.
I saw healthcare companies automating insurance claims using browser automation.
I saw developers building full-scale behavior simulation and UI-testing suites using different "personas".
I saw an agent that would login to critical software and ensure that all configurations were compliant.
Everyday, it was something new.
Meanwhile, claude-code was absolutely ripping, and the models were getting much better at computer-use. All the labs started publishing their computer-use benchmarks on Mind2Web, WebVoyager and OSWorld.
^^Gemini 2.5's model card shows its benchmark performance, as measured on Browserbase.
2026 came around, and Openclaw took everyone by surprise. We all realized how important a good harness is to building an agent.
And then a few people started to notice, if you took this harness and gave it access to a browser, it could start doing "work" for you.
I setup my Openclaw on a VM, and started using it as my personal trainer. It would use a browser session to pull relevant YouTube videos, and prepare my workout for the next day. I started planning my trips and date nights with openclaw. I was obsessed, and so were many others!
Somewhere in mid-2026, models and harnesses started to get really really good with computer-use. In fact, people (👀 @danshipper) started giving Codex their email access.
Developers stopped testing their apps, and let their Codex run wild testing everything in its browser.
The Chasm
Something very interesting started happening in the last couple months.
xAI shipped GrokBot in August
Instinct, which went live in August, has been the most viral invite-only product launch in recent history.
Meta went live with Muse in September. It is currently on a faster growth curve than ChatGPT in 2022.
Now millions of people across the world can feel that black magic I felt with OpenAI Operator. They could all see the power of an agent planning out their travels, using a browser to book their tickets. These agents are now making e-commerce transactions with credit cards, which is a strong sign of the user's trust in these systems.
At Browserbase, we have seen many tailwinds to our business, but this has been the strongest by far. It seems like the average internet user can just download an app from the App Store, and use state-of-the-art computer-use out of the box. The packaging of these personal agents is extremely appealing and will spread very far.
However, even after all these improvements, I'd argue that we haven't yet crossed the computer-use chasm. Most people are willing to let an agent run some errands for them, but haven't yet offloaded mission-critical workloads.
My taxes haven't been filed by an agent. My investments aren't being managed by an agent. My insurance is not supervised by an agent. Payroll for my company is not being run by an agent.
Yet.
Why?
Geoffrey Moore’s central argument in Crossing the Chasm is that early adopters will tolerate an incomplete product because they want an advantage. The early majority will not. They want something reliable, proven, and safe enough to become part of normal work.
One of the biggest challenges with agents is that of identity. How do you delegate your identity and permissions to your agent. How do you configure your agent to have all your passwords, and what about 2FAs.
Even if we solve the problem of delegating identity, the bigger question arises: what about permissions? Should an agent have the same permissions as the human, and if not, where do you draw the line. For critical decisions, agents ask their human for explicit approval, but if not designed well, this can lead to approval fatigue.
One of the biggest friction points with the broader adoption of computer-use is security. For agents to perform economically valuable work, they are going to have access to sensitive information like health records, bank accounts, and personally identifiable information. The big question for our industry is how we ensure these agents browse the web without leaking any information to honeypots or vulnerabilities. Additionally, how do we ensure the world wide web doesn't inject my agent's context.
In addition to the challenges of agent identity & security, you have the challenge of working within the current incentive structure of the internet. Antibots and CAPTCHAs, which originally existed to prevent bad bots from doing bad things, are now blocking good-intentioned agents.
Crossing the Chasm
All of these are challenges that get solved at an infrastructure and ecosystem level. We should enable our agents to have emails, phone lines, IP addresses, credentials, credit cards, securely governed within isolated sandboxes and browsers.
As the agents build more distribution leverage, it forces the ecosystem to change its incentive structure. The internet economy will start to open itself to agents, and WebMCP is a promising direction to the future.
At Browserbase, we say that we're in the diffusion business, and these are all problems that we're wrestling with everyday. Just a matter of time till we cross into the majority.
Slowly, then all at once!