The Human-Likeness Penalty

The More Human AI Feels, the More Human Its Failures Become
Enhancing technology doesn't just lead to greater connection; it also increases the number of ways it can let us down and alters the experience of disappointment.
The brief usually states the same point: make it feel more human. Give the bot a name and a personality, and use empathetic language and a conversational, warm tone. This should make the customer feel as if they are speaking to a person, not a machine.
The brief is not incorrect; anthropomorphic design does improve the quality of interaction in ordinary situations because users tend to engage more, gain trust more quickly, and report greater satisfaction when the automation they are dealing with acts in socially coherent ways.
However, the brief is incomplete because it considers only the good conditions and ignores the bad ones. Moreover, no matter how well designed the automation is, it will eventually fail. The question the brief has not asked is this:
What effect does it have on the user's experience of failure when the system they had been interacting with is made to seem human?
The question has now been answered, and the answer changes the brief.
What the Research Found
A study examining customers' reactions to service failures in AI chatbots tested under both high- and low-anthropomorphic conditions. The chatbot failed in two ways: in process failures, it responded inappropriately during the interaction; in outcome failures, the service result was incorrect, regardless of how the interaction had seemed.
Customer aggression was high when the outcome was incorrect, regardless of how human-like the chatbot seemed. The outcome was incorrect, and the bot's anthropomorphism had no effect.
Regarding process failures, the results were both specific and counterintuitive: customers showed considerably greater aggression toward highly anthropomorphic chatbots than toward less anthropomorphic ones when both types of chatbots suffered the same failure. This was because users' expectations had been disconfirmed: they had formed higher expectations of socially competent behavior from the bot because it resembled a human, and when these expectations were not met, the experience was not seen as a limitation of the software but rather as a social violation. Moreover, a reaction to a social violation is different from a reaction to a software error.
The importance of that distinction lies in the fact that when an AI system fails in a way that annoys users, most organizations tend to rely on disclosure as their first response. However, the research shows that this reflex is incorrect; the right course of action isn't to tell the user that they were speaking to software, but rather to regard the failure as the user has experienced it, that is, as a social situation calling for a social reply.
A bot that acts like software can manage software-related failures. If you give it a name, a personality, and empathetic language, the organization has gently altered the standard used to judge every failure.
The Human-Likeness Penalty
Whenever automation is given any human characteristic, a corresponding liability arises regarding expectations. This is known as the Human-Likeness Penalty, that is, the gap between the level of expectation raised by an anthropomorphically designed system and the level the system is in fact able to deliver when things go wrong.
The penalty is not uniform; it is focused specifically on the process level, that is, on the quality of the interaction from moment to moment, rather than on the outcome. A system that is highly human-like and produces an incorrect outcome is no more likely to generate aggression than a machine-like system that produces the same incorrect outcome. However, a highly human-like system that responds awkwardly, misunderstands the social tone of the conversation, or is unable to handle the emotional aspect of a difficult interaction will cause considerably more aggression than a machine-like system that makes the same mistake.
This presents a particular design risk because the existing literature on anthropomorphism has not adequately identified it. The more convincingly human an automated system seems, the more its failures must be perceived as human-like; in a system intended to appear like a person, awkward phrasing that would be entirely normal in a clearly mechanical system is instead treated as a social failure. A missed cue that would simply be regarded as a software limitation becomes a relational breach when the system has suggested that it understands how conversations work.
Three Decisions This Changes
The Human-Likeness Penalty is not an objection to anthropomorphism; it is, in fact, an argument for intentionally adjusting anthropomorphism, based on a clear explanation of what the system can and cannot do at the process level, rather than automatically adopting a human-like design simply because it yields better engagement metrics under ordinary conditions.
01 The naming and personality decision
It is common for designers to give an automated system a name and endow it with a conversational personality, which sets the earliest expectations. When a system is called Alex and has a friendly, conversational style, a social agreement forms before any interaction takes place. This agreement defines what constitutes socially appropriate behavior. Companies that make this choice without first considering whether the system can provide socially competent responses, not merely factually correct ones, are setting themselves up for expectation liability they might not be able to meet in all cases.
02 The recovering script
Almost all chatbot failure-recovery scripts fall into one of two types: explaining the technical limitation or escalating the issue to a human agent. The research points to a third type that outperforms the other two in situations with a high degree of anthropomorphism—namely, a sincere apology that frames the failure as a social incident rather than a system error. In cases where the system is designed to appear human, 'I'm sorry, I didn't handle that well' is better than 'I'm sorry, as an AI I have limitations'. The recovery procedure must match the persona and not contradict it.
03 The failure mode audit
When organizations use systems that are highly anthropomorphic, they should specifically examine the possible ways the process might fail, that is, the manner in which the interaction itself could go wrong, rather than concentrating mainly on the accuracy of the outcome. A chatbot that achieves the correct result through an awkward, misinterpreted, or socially inconsistent interaction process accumulates greater expectation liability than a less human-like system with the same level of outcome accuracy. The questions that should be asked are not just 'when does this fail?' but also 'when it fails, what kind of failure is it, and does the design set up the expectation that the failure will be violated?'
Where This Shows Up Most
The consequences of the penalty are most significant when the emotional stakes are already high. For example, healthcare scheduling robots that appear warm and empathetic interact with patients who are anxious, in pain, or concerned about a condition. If such a process fails—say, by misinterpreting the emotional tone or responding awkwardly to a patient who has expressed distress—this is not considered a software malfunction. Instead, it is perceived as a breach of the care relationship implicitly created by the robot's anthropomorphic design. The aggression that follows is proportional to the level of care expected, not to the seriousness of the technical error.
AI receptionists, virtual assistants used in customer service, and AI-generated communications that use a human name and a friendly tone are all affected by the same dynamic; the more convincingly the system appears to have social competence, the greater the damage each instance of social incompetence inflicts on the relationship and the associated brand.
Compared with disclosures, the study of apologies also has particular consequences for organizations that have invested heavily in highly anthropomorphic AI for customer communications. In such cases, when the communication leads to frustration, there is no point in reminding the audience that it was produced by AI; doing so is not a remedy but an admission that could erode trust rather than increase it. It is better to offer a well-constructed apology in which the failure is taken seriously as a relational event.
The recovery procedure must be in keeping with the persona rather than contradict it; stating that the system is an AI when a human-like AI fails does not reduce agression; it only confirms that the user's frustration was unfounded.
The statement that aims to give it a more human feel hits on a genuine point; designs with a human touch improve the quality of interaction, lead to greater engagement, and establish rapport more quickly than machine-like designs. All of that remains true.
The contribution of the research is the other half of the equation. Each degree of human-likeness achieved by the design also means the system is obliged to meet that expectation, not only in favorable conditions but also in unfavorable ones. The penalty is not incurred when the system is working well; it is incurred exactly when the system fails, which is precisely when the organization's relationship with the user is most vulnerable.
The decision to deliberately calibrate anthropomorphism—adjusting the degree of human-like characteristics in the design to match the level of human-likeness the system can actually maintain in the face of failure—remains one most organizations have not made. The research indicates that the cost of not making this decision exceeds what most expectation-liability models have accounted for.
Some ideas are worth discussing in the context of your organization.



