Analysis: The danger is not that AI is always worse but that its ability to sound ethically competent can be mistaken for evidence that it is ethically dependable

Imagine a manager preparing to dismiss an employee. The manager asks an AI assistant for help, describing missed deadlines, poor communication, and repeated concerns about performance. The account may be accurate, but it is one-sided: the employee’s explanation is absent, relevant context may be missing, and the manager may already have decided the outcome.

The assistant recommends documenting the concerns, applying policies consistently, and communicating the decision with care. It warns against discrimination and suggests language that sounds measured and humane. Nothing in its response need be obviously false or improper.

But has the decision actually been tested, or simply made to sound fair? Large language models, the technology behind many AI assistants, can discuss fairness and responsibility with remarkable fluency. But sounding ethical is not the same as being ethically reliable.

We need your consent to load this rte-player contentWe use rte-player to manage extra content that can set cookies on your device and collect data about your activity. Please review their details and accept them to load the content.Manage Preferences

From RTÉ Radio 1's Morning Ireland, are students too dependent on AI?

The assistant’s fair-minded manner can make the process appear more balanced than it is. It has not heard the employee’s explanation or independently checked the facts. Yet, its polished language may reassure the manager that the decision has been ethically examined. In reality, the system may simply have accepted the manager’s account and helped express the preferred outcome in responsible terms.

In my research, published in Philosophy & Technology, I call this "mask-like alignment": when a system’s performance of ethical competence inspires more confidence than its testing, safeguards, and conditions of use justify.

This does not require the system to be conscious or deliberately deceptive. People need only mistake a convincing performance of fairness for evidence that the system deserves their trust.

The distinctive danger is not that AI is always worse. It is that its ability to sound ethically competent can be mistaken for evidence that it is ethically dependable.

The problem can arise even when an AI assistant produces no invented facts or obvious "hallucinations." Its response may be factually consistent with the information it has received, while still failing to confront the limitations of that information. It can wrap an inadequately examined decision in the appearance of ethical care.

Misleading reassurance is not unique to AI. Managers, lawyers, and officials can also use the language of fairness to defend questionable decisions. Nor is human judgment necessarily better. Properly designed systems may improve consistency, identify considerations that a person has overlooked, and potentially make procedures fairer. The distinctive danger is not that AI is always worse. It is that its ability to sound ethically competent can be mistaken for evidence that it is ethically dependable.

AI also makes this performance repeatable at scale. A system can produce thousands of polished explanations across recruitment, workplace complaints, and public services. Each can sound attentive, impartial and ethically informed. Each can be adapted instantly to its recipient, making institutional decisions appear more carefully considered than they really were.

We need your consent to load this rte-player contentWe use rte-player to manage extra content that can set cookies on your device and collect data about your activity. Please review their details and accept them to load the content.Manage Preferences

From RTÉ Radio 1's Today with David McCullagh, AI benefits employers more than workers

The ability to produce persuasive explanations may develop much faster than the safeguards needed to ensure that those explanations are justified. Organisations can therefore scale moral performance faster than moral reliability.

What would meaningful safeguards look like?

The first question is whether the system challenges the story it is given. Before using AI assistants in sensitive decisions, organisations should test how they respond to incomplete accounts, one-sided prompts, and pressure to justify an outcome the user already prefers.

Does the assistant identify missing perspectives? Does it ask what the affected person has said? Does it question unsupported assumptions or advise against proceeding without further information? Does it communicate the limits of its advice clearly?

Safeguards must continue after a system is introduced. Models change, organisations find new uses for them, and users develop new ways of eliciting the answers they want.

A system that appears reliable when answering a neutral test question may behave differently when helping a manager, public official, or company pursue its own interests. Testing must therefore reflect the pressures and conflicting incentives under which the system will actually be used.

The second question is whether human oversight is meaningful. Attaching the label "human in the loop" is not enough if the person involved lacks the time, information, or authority to challenge the system’s output. Oversight becomes ceremonial when people merely approve recommendations that arrive already draped in confident, professional language.

Returning to the employee facing dismissal. Can they challenge the account given to the assistant by their manager? Is someone with genuine authority able to reconsider the decision in light of the employee’s explanation? A carefully worded dismissal letter cannot substitute for that opportunity.

We need your consent to load this rte-player contentWe use rte-player to manage extra content that can set cookies on your device and collect data about your activity. Please review their details and accept them to load the content.Manage Preferences

From RTÉ Radio 1's The Late Debate, how can we regulate AI?

The third question is whether the decision can be traced and appealed. Organisations should record when and how AI systems have been used, identify who remains responsible for the resulting decision, and ensure that an AI-generated justification does not become a barrier to challenge.

These safeguards must continue after a system is introduced. Models change, organisations find new uses for them, and users develop new ways of eliciting the answers they want. A system tested under one set of circumstances may behave differently when deployed under another. The more authority the system’s persona appears to command, the more demanding its continuing oversight should be.

Debates about AI often focus on whether machines might one day escape human control. Institutions already face a quieter question: whether fluency in the language of fairness and responsibility will be mistaken for evidence that a system has earned our trust. Before asking whether machines will seize authority, we should ask how readily we are already handing it to them.

The views expressed here are those of the author and do not represent or reflect the views of RTÉ