The talk about AI is the talk of the town. If you are part of a team that develops software, it’s likely that which AI works and which doesn’t is a topic of some interest to you. We often hear colleagues discuss whether this AI agent is better than that or how good some AI is.
In such discussions, however, people may use the word ‘good’ to mean different things. For example, by ‘good’ AI, one could imply one which is reliable in most conditions, or one that is safe or one that is socially beneficial or even morally desirable. These are different claims. This matters because two people arguing about whether an AI system is “good” may be judging it by different standards.
Recently, I read a book titled “Morality” by the philosopher Bernard Williams. In one of the chapters titled “Good”, Williams asks: what does good mean?
The word “good” looks simple, but it does not work like an ordinary adjective such as “red.” If you say, “This is a red car. A car is a vehicle,” then we can infer, “This is a red vehicle.” What makes that car red is the same thing that makes that vehicle red, its color. But if you say, “He is a good soccer player. A soccer player is a human being,” it does not follow that “He is a good human being.” What makes someone a good soccer player is not the same as what makes someone a good human being.
We cannot always detach the word “good” from the kind of thing being judged and expect the same standards to apply to other objects. The properties that make someone a good doctor are not the same properties that make something a good knife or someone a good soccer player.
Williams also points out that comparison alone does not define goodness. If we say, “Smith is better than most soccer players,” we are still not sure how good a soccer player Smith is since we are still relying on the hidden definition of what makes a good soccer player. But I think once we have clear criteria and baselines, comparisons can be useful.
Reading this chapter made me, as someone interested in AI as a user, think a bit. What if someone says, “ABC is a good AI agent. It can write software, essays, business reports, medical summaries, and legal memos.” Can we then infer that ABC is a good programmer, good writer, good business analyst, and a good lawyer? At most, we can infer that it performs some tasks associated with those roles well, but only in specific situations, possibly not in all. That is not the same as being good at the role itself.
One could also wonder what a good AI coding agent is. An AI coding agent can generate a lot of code quickly. But programming is not just producing code quickly. Programming can involve clarifying requirements, understanding an existing system, coordinating with people in different roles across a company, reasoning about trade-offs and correctness, and thinking about maintainability and reliability. Programmers may also train people, transfer knowledge, interact with customers, communicate ideas, bring their past experience to their work, disagree when they think something is wrong, and persuade colleagues when they believe there is a better approach.
But even saying what an AI is good at may not be enough. A good outcome could depend on the nature of the task, its purpose, why it is being carried out, the user carrying it out, and the environment. The kinds of outcomes we measure and the failures we consider acceptable are also important. A simple script written with little verification by an engineer for a quick throwaway prototype may be fine at a startup, but an engineer shipping software relied upon by millions with no verification inspires little confidence.
When you hear someone say, this thing is good, for a meaningful conversation you can ask: good at what? An AI coding agent good at creating prototypes fast is different from one good at helping security researchers find problems in untrusted code. An AI coding agent that creates patches quickly may be good for solving an immediate problem, but not for a team that has to maintain those patches over the long run, especially if it is used without the team’s oversight.
The word “good” can work well when people share enough context to understand what is meant. Friends who have dined together for years do not need specifications when one of them calls a restaurant “good”; their shared experience supplies that context. The problem arises when claims carry from one context to another; for example, when a vendor claims that their product is good at x, which could definitely be good in that context, but a team working in a completely different context, with different tasks, a different environment, and different success and failure criteria assumes that it would also be good for them.
Towards the end of the chapter, Williams suggests that to judge whether something is a good x, we need to understand what kind of thing x is, what it does, and the relevant facts about this particular x. That can often be enough, at least broadly, to judge whether the claim is true or false. When evaluating whether technologies and tools are “good,” teams therefore need enough understanding of both the work and the technology to choose meaningful tasks, baselines, success criteria, and acceptable failure conditions.
And importantly, before calling a technology or tool good, we can ask: good at what, for whom, compared with what, and under which conditions? What outcomes are we measuring, and which failures are unacceptable?