Blog · · 12 min read

Phones are the last big body of software only people can use

In the 1970s and 80s, a lot of the world's important software lived on IBM mainframes and could only be used through 3270 terminals, the green screens that bank clerks and airline agents typed into all day. Those applications had no way for another program to talk to them, so when companies wanted their new PCs to work with them, IBM shipped an interface called HLLAPI that let a PC program read what was on the terminal screen and type into it, exactly as a clerk would. 3270 screens were made of labelled fields rather than pixels, so these programs read the fields, and for decades a large share of business integration ran on software operating screens that had been designed for people.

I think the same thing happens whenever a large body of software is built only for humans to operate: a layer appears that operates those screens on behalf of programs, and it does the job until proper interfaces arrive, which can take a very long time. Phone apps are the biggest collection of software like that today, and the one most people use every day. My bank, the taxi, the parking and the parcel locker are all phone apps with no interface an agent can use, and every one of them sends its login codes to the same phone. Coding agents can find their way around a website they've never seen, but they couldn't touch any of that. Opposable is our attempt at that layer for phones. It lets an agent use a real phone more or less the way a person does, opening apps, reading the screen, typing and dealing with whatever pops up in between.

On Android, the equivalent of those fields comes from accessibility. In 2009, Android 1.6 added a small framework so that a screen reader called TalkBack could describe other apps' screens to people who couldn't see them and press buttons on their behalf, and that framework is how our app finds and taps things today. A screen reader and an agent need nearly the same thing from a phone, a description of the screen in words and a way to press a button without aiming at pixels. What follows is what we learned building on that, and some guesses about where this goes that I'm less sure of.

Why phones are harder than computers

On a laptop, computer use is conceptually simple. The agent takes a screenshot, works out where to click and moves the mouse, and the operating system has no particular opinion about who is moving it. Phones are designed so that every app runs in its own sandbox and the system goes to a lot of trouble to stop one app from seeing or touching another, which is a big part of why phones never got the kind of malware that Windows PCs did. I think that design is right, but it does mean there's no obvious place for an agent to plug in.

Android's one opening is the accessibility service, the descendant of that 2009 framework, which today also serves people who operate their phones with switches or their voice. An app holding that permission can read the contents of other apps and perform taps and gestures in them, which is also what an agent needs.

The iPhone has no equivalent, so we went around it. A Mac can present itself to an iPhone as a Bluetooth keyboard and mouse, and with AssistiveTouch switched on, iOS treats a mouse click as a tap, while the screen comes back to the Mac over a USB cable. I wrote up the details in a separate post; in our tests on an iPhone 17 Pro, every tap landed where it was aimed.

What the agent looks at

The first decision that turned out to matter was what to show the model on each step. A screenshot is the obvious choice, since it's what a person sees and what desktop computer use is built around, and we started there. It works, but it's slow, and models are better at describing a screenshot than at saying precisely where a small toggle sits in it, so they tap a few pixels to the side, nothing happens, and they spend another step taking a screenshot to find out why.

The accessibility service gives you something better at no extra cost, because it already knows every element on the screen along with its label, its type and where it is, which is the same information TalkBack reads aloud. So we hand the agent a numbered list, and it replies with something like "tap 12" instead of a pair of coordinates. A screen as text is a few dozen lines, against roughly a thousand image tokens for a screenshot, so each step costs less as well as missing less often. Screenshots are still available for the cases where the agent needs to see a photo or a map, and in our test runs it rarely needs one.

Most of the speed came from two changes to how the app talks to the phone. The first version opened a fresh connection to the phone for every action and then waited a fixed half second for the screen to settle, so we kept one connection open and waited until the screen had actually stopped changing, which is sometimes 80 milliseconds and sometimes two seconds. On the three tasks we use as a benchmark, the total time dropped from 140 seconds to 56 and the number of round trips with the model from 62 to 18, and most of what's left is the model thinking rather than the phone.

Interruptions

If you watch yourself use a phone for ten minutes, you'll see that a lot of it has nothing to do with what you were trying to do. The camera asks for a permission the first time you open it, an app wants a rating, a website wants you to accept cookies, and a six-digit code arrives by SMS while you're in the middle of something else. People handle all of this without noticing, but an agent that hasn't been told about it will look at the same screen over and over, trying to work out why the button it wanted isn't there.

A lot of our time went into making the phone describe these situations clearly. When something is covering the app, the element list now says so, the harmless kinds like rating prompts and update reminders are dismissed automatically, and one-time codes are picked up from messages and notifications so the agent can carry on with what it was doing. Permission dialogs are deliberately left for the agent to decide, because granting one is a real choice.

Doing more of the work on the phone

Most of what people ask an agent to do on a phone doesn't need a large model at all. Setting an alarm, texting a contact, switching on Wi-Fi or opening a particular settings page takes the same few steps every time, so the app now does about forty of these by itself without any model, and checks the screen afterwards to confirm the thing actually happened.

When an agent does something new on your phone through Opposable, the steps are recorded against the text on the screen rather than against pixel positions, and you can save them as a routine. Next time the phone replays them on its own in a few seconds, so the large model only has to work a task out once.

For anything else there's a small screen model, about two billion parameters, that runs on the phone itself and decides where to tap by looking at the screen. It's nowhere near as capable as Claude and on unfamiliar tasks it's slower than I'd like, but it works without a connection, the screen never leaves the phone, and it keeps going after you close the laptop.

When we timed a single step of that model on a laptop CPU, almost half of it went into turning the screenshot into something the language part of the model could read, and generating the answer, which element to tap, took the least time.

That stage, the vision encoder, is very close to the work phone chips already do all day for the camera. Every recent phone has a neural processor whose main job is making photos look better, and on a Galaxy S25 it ran our encoder in about a quarter of a second, roughly sixteen times faster than a laptop CPU. We haven't yet run the full loop end to end on a real phone, and the rest of each step still has to happen, so I don't know yet what a whole step will cost there. If the rest of the step can be brought down as well, running the screen model on the phone could end up cheaper and faster than sending screenshots to a server, but we don't have those numbers yet.

Payments, and three bugs in our own code

People are right to get nervous about an AI agent and a banking app on the same device. From the beginning, anything that looked like a payment paused and asked the owner on the phone, with Approve and Deny buttons, and everything the agent did went into a log on the phone. Last week someone asked how exactly we decide that a screen is a payment screen, and when I went back to check, the check had gaps.

The first version only looked at the label on the button being tapped. It caught "Pay", "Buy", "Place order" and "Send money", but a plain "Confirm" underneath a €25 transfer went through, and so did an unlabelled button or a slide-to-pay control. We changed it to read the whole screen, so that a confirm-type tap on a screen showing an amount next to words like total, amount or recipient asks first, and any confirm inside a known payment app asks regardless.

While testing that fix I found three other problems. The approval request is an ordinary Android notification, and since the agent can pull down the notification shade like anyone else, nothing was stopping it from tapping Approve on its own payment. While I was closing that, our automatic pop-up handler saw the words "update available" somewhere on the screen and tapped the first dismissal button it found, which happened to be Deny on a pending approval. And when the phone was plugged into a computer over USB with Opposable's control switched off, the tool on the computer fell back to driving the phone over the debugging connection, which skips the check entirely, so a transfer on our test page went straight through.

All three are fixed, and none of them involved a model misbehaving. Each was a piece of ordinary plumbing written on the assumption that only a person would ever touch the notification shade or that a fallback path would never matter. I suspect most security work on phone agents will look like this: walking every route an action can take and making sure the owner's confirmation can't be reached from any of them. It gets more important once you consider that a phone is mostly text written by other people. Every SMS, notification and web page the agent reads is input, and eventually one of them will tell it to ignore its instructions and send money somewhere, so I would much rather rely on a check the agent can't reach than on the model noticing.

Most real money movement also ends in authentication an agent can't perform, like a fingerprint, the bank app's own PIN, or on an iPhone, Face ID plus a double press of the side button. An agent can do all twenty steps up to a payment and still be unable to make the last one, so for most payments the final step stays with the owner even if our own check misses something.

Where I think this goes

The mainframe story had a more recent sequel in personal finance around 2010. Banks didn't offer APIs, so Mint and Plaid connected to people's accounts by logging in with their passwords and reading the account pages. Banks disliked this and tried to block it, it kept working because customers wanted it, and in Europe regulators eventually required banks to provide proper interfaces. Reading the screen was how those apps got the data until a better way existed, the same role HLLAPI played for the mainframes, and agents using phone screens are in a similar position now.

Most apps offer agents no way in except the interface designed for people, and many of them earn their money through that interface, from ads or from the offer that appears just before checkout. Some will push back, and the early signs are already there: some banking apps refuse to run on virtual phones, and Google asks you to register a virtual phone before the Play Store will work on it. At the same time, Android and iOS both have early mechanisms for apps to expose actions directly to the system's assistant. My guess is that the two approaches will exist side by side for a long time, with proper interfaces for apps that want agents and the screen as the fallback for everything else, much as the web kept its pages long after it got APIs.

Apple and Google will also build agents into their operating systems, and those will probably be good. Our bet is that people already have an agent they like, and it isn't necessarily the one that came with their phone. If you do your work with Claude or ChatGPT or Codex, you want that agent to be able to use your phone too. An agent running on your computer also needs a phone for the same reasons you do, because that's where the codes arrive and where half the services live.

The phone has also become the thing that proves who you are online. Two-factor codes and bank authenticator apps rest on the assumption that whoever holds the phone is its owner. Once an agent holds the phone that assumption gets weaker, and I'd expect services to respond by moving the decisive step to something that needs a body, a fingerprint or a face. That seems healthy to me, since it separates the work from the decision in the same way the payment approvals do.

The permission that makes all of this possible on Android exists for disabled people, and Google and app makers both have reasons to restrict it. Android already makes it harder to grant to apps installed from outside the Play Store, and Google Play limits which apps may use it. It would be easy to tighten further, and I don't think we should pretend that risk away. Our case rests on the same idea accessibility itself rests on, which is that it's your device and you get to decide what operates it, but I can't promise that argument will always win.

Agents may also change app design, once a real share of users are agents. A lot of mobile interface design is about steering people, whether that's a cancel button that is slightly too small or the six taps it takes to end a subscription. An agent doesn't get tired or embarrassed, so if enough people start delegating those flows, the steering stops working and some leverage shifts back to the person who owns the phone. We don't know yet how app makers will respond.

How we're approaching it

Put together, these are the principles we're building around. The phone should do as much of the work as it can by itself, because that keeps your screen private and keeps working without a connection, and because phone chips are getting very good at the expensive part. Any agent should be able to drive it, because which agent you use is your choice. You keep the last tap on anything that moves money, and the agent must not be able to reach that tap by any route. Everything the agent does is written down, so you never have to take its word for what happened.

All of this is early: apart from the iPhone and the one neural processor measurement, all of our timings come from an emulator, the on-device model is slower than I want, and we keep finding assumptions that only held because a person was holding the phone. Computer use got attention because it let AI use software built for people instead of for other programs, and for most people, especially in countries where a phone is the only computer a household has, that software now lives on a phone.

If you try Opposable and it gets stuck somewhere, I'd like to hear where. The Android app is free at getopposable.com, the iPhone route needs a Mac, and my email is hello@getopposable.com.