The standard interface for an agent is a command line or a sidebar. Both are a log: the user writes, the agent replies, and each tool call is recorded. Typically the agent works out of view, and the person reviews what comes back. Each round trip widens the gap between what the agent did and what the person understands. For long tasks that gap becomes the limit, and the person is left to trust the result or attempt to reconstruct it.

The alternative is a more visual approach. IE, an agent works beside the user in their document, so its work can be steered as it runs. There is a very fine line between visual theatre and insightful visualisation. Generally, working out of view gets more done; whilst working in view keeps the user oriented, and can teach them as it goes. Our objective is the “sweet spot” of friction that strikes the right balance between user alignment and work output.

1. Spatially resolved agents

A type agent works in the document as a colleague would. It has a caret and a name, it moves to the paragraph it is working on, and its edits arrive as tracked changes while you watch. Similarly, it leaves alone the sections others have claimed or locked.

This makes editing more visceral. In a sidebar, the agent’s work is a report you read afterwards. In the document, you see it arrive where it applies, which makes steering (interruption and revised commands) more intuitive.

Slotting into this is the concept of redlines (approved diffs) that the user controls. Again, more friction. But each accepted or rejected change is kept in version control, and that history can later be fed back to the agent to bring it closer to what the user wants, an approach Cursor has written about well.[1]

Moreover, by working in the document itself, we have better spatial resolution, or “pointing”. A recent post on HN (big-arrow-on-the-screen) lets a coding agent draw a large arrow, a box or a line of text over whatever is on a screen.[2] This might be visual guidance can be used to show users around software (ie. how do i do this?). Similarly to how mutli-agent design revolves around human co-working principles, people can communicate by ie. highlighting something with a marker, so why not LLMs too?

Equally, we’ve experimented with three carets at once. For most users, this is too much info at once- for live consumption. By making agent actions replayable (ie via video artifacts), separable (ie view every caret trajectory individually) and revertible (via version control) we resolve most of these issues. Essentially, summaries are always legible, when they are actionable by the user.

The caret could always simply be an additive to the agent sidebar that the user opts-in to, for easier co-work and steering. AKA: the agent acts as usual (same set of prompt > tool calls until completion), but it emits an opt-in caret to follow. We’re just shifting where the user focuses their attention, and enters feedback, rather than changing how agents work.

2. Communication between people

Most real work is shared. Equally, most agent conversations are private to a single user.

The question becomes, what is effective multiplayer? Is it several people editing a shared underlying codebase, using private conversations? We break this down into a couple of concepts.

- Organisations. A company shares documents the way it already does: with the whole company, as editors or viewers (some kind of RBAC).

- Projects. Documents, facts and decisions belong to a project, and the agent reads from the same place as the people on it.

- Shared conversations. If work is driven by agents, and people are not reliably committing relevant information (such as in a git commit), the agent conversation becomes a work artifact that can be searched for details.

- Live collaboration. Colleagues see each other in the document, claim sections, lock them against AI and replay how a draft came to be.

- Comments pinned to the text. A comment solves the pointing problem in both directions. It is anchored to the words it is about, so a person can ask about “this” and the agent knows what “this” is. A comment is also a prompt: mention the agent in it, and the agent reads the thread and replies in place, where the question was asked. The agent can answer with a dozen separate comments, each on the paragraph it concerns, instead of one long reply. Leaving comments on a canvas for the agent to pick up feels more natural than typing into a sidebar. But it also adds friction and there’s more clicking.

- Graphics in comments. An extension to the above. Comments are where people already discuss a document, so the agent belongs there too. A comment can carry a chart, a timeline or a flow, drawn from the document’s own figures as a live artifact rather than pasted as an image.

Comments and shared conversations are solving known problems: authorship, position, and everyone already knows how to use them. We’ve seen the development of org workflow tools like Claude Tags as a result. The only hair in the soup is ensuring that these varied frontend representations use the same underlying agent harness and version control system (so everything gets the same level of support).

3. Communication between agents

Two patterns of communication between agents hold up.

Top-down. A coordinator splits the work, specialists do it, and one check reads everything at the end.

Push. Agents also need to send each other things while they work: a finding, a correction, a changed instruction. These messages should be durable and ordered: the receiving agent is told a message came from another session, and a teammate cannot approve a permission prompt on the user’s behalf.[3]

Free-form conversation between agents is where things leak and drift, and reading it is as tedious as reading raw tool calls.

A person should see the shared record the agents write to: the board by section, with IE. a paragraph two agents disagree about marked as contested (due to a semantic difference). You can also resolve these conflicts automatically. The way most coding harnesses do this is by surfacing issues and saying: do you prefer A/B/C? If the user does not explicate a preference an agent-mediated solution is selected. Once again, opt-in friction.

Networks of agents interacting is “mostly a distraction”, hence why we organise work to minimise the probability of inter-agent friction (via DAGs, etc). Several agents contributing intelligence (via read-ops) while writes stay single-threaded seems to be useful.[4] Similarly, Anthropic recommends three to five teammates for most work, and warns that two teammates editing the same file leads to overwrites.[3] We reached the same conclusion for documents in June: reads can run in parallel, but writes stay synchronous.[5]

4. Knowledge graphs

Visualising data in 3D sounds appealing. Semantic vectors are conceptually very cool. As a way to read a document, it was the least useful interface we have made, for four reasons.

Everything at once answers nothing. Agents are too messy to light up a reasoning “path”. A short report easily becomes more than a hundred nodes.

Depth costs more than it gives. The third dimension buys space at the price of occlusion: nodes hide nodes, labels collide, etc. Where a node sits in a force-directed layout means nothing, so nobody learns where anything is- and semantic clusters obscure other more important structural information.

The answer is typically a list. A better graph rearranges itself around the question. The answer separates into columns that read left to right as a chain of consequence.

A knowledge graph is excellent as the agent’s working data and poor as the reader’s view. The agent can use the links to check what depends on a section before it changes it. A person needs the part of the graph that answers their question, arranged so that position means something, with the text underneath. Occam’s razor.

Quotes

Some anonymous user quotes on this issue:

- “Separate guiding from doing. Teach me to do this" kinda stuff. "Guiding agent" rather than "doing agent".”

- “Seems like a pretty useful way to help infants use desktop computers” > “Have you met users?”

- “I think this would’ve been very handy to me a couple years ago when I was learning to use Blender and asking LLMs for help performing certain actions like “how do I display the normals of all the vertices in my mesh?”"

- “Undoing an agent’s mistake by sending it more messages is clumsy, where a tracked change can be rejected in one click.”

- “We humans can do this too on paper. We can point at something, we can highlight something with a marker. Why not give an LLM these capabilities as well?”

Conclusions

- Draw what the person has to check, not what the agent knows. A knowledge graph is the agent’s data. The reader needs the part that answers the question, arranged so that position means something, with the words one click away.

- Put agent work where people already talk: shared conversations, comments pinned to the text, and review.

- Workflows are a silly term, but increasingly relevant. AI continuously refines a knowledge base. Workflows must be persistent & referenceable. The tacit example is a series of markdown files (encoding linguistic descriptions, the ultimate currency of LLMs). Alternatively, for greater reproducibility on tasks, deterministic programmes, modified to a specification and then tested using regression gates.

References

- Jacob Jackson, Ben Trapani, Nathan Wang and Wanqi Zhu, ‘Improving Composer through real-time RL’ (Cursor, 26 March 2026) <cursor.com/blog/real-time-rl-for-composer>

- Franz Enzenhofer, ‘big-arrow-on-the-screen’ (GitHub, accessed 9 October 2026) <github.com/franzenzenhofer/big-arrow-on-the-screen>

- Anthropic, ‘Orchestrate teams of Claude Code sessions’ (Claude Code Docs, accessed 9 October 2026) <code.claude.com/docs/en/agent-teams>

- Walden Yan, ‘Multi-Agents: What’s Actually Working’ (Cognition, 22 April 2026) <cognition.com/blog/multi-agents-working>

- Alan Yahya, ‘Multi agent systems for complex tasks’ (Lexifina, 25 June 2026) <lexifina.com/blog/multi-agent-systems-for-complex-tasks>