What Do We Really Mean When We Say AI ‘Thinks?’
Carnegie Mellon research traces today's AI rhetoric to decades-old debates about language, technology and human intelligence
By Stefanie Johndrow Email Stefanie Johndrow
When organizations describe artificial intelligence as “thinking,” “writing” or “learning,” the words can sound familiar. But according to Carnegie Mellon University historian Christopher Phillips, those terms may tell us as much about how humans talk about technology as they do about the technology itself.
In a new paper published in IEEE Annals of the History of Computing, Phillips and co-author Alison Langmead of the University of Pittsburgh examine how language surrounding artificial intelligence has evolved since the 1950s. Their research argues that descriptions of AI have long relied on what they call “strategic ambiguity” — using words that carry precise technical meanings for computer scientists while suggesting something much broader to the public.
According to Phillips, the result is a growing tendency to compare computers with people instead of describing what each does best.
“We use words like ‘smart,’ ‘read,’ ‘write’ and ‘think’ very differently when we're talking about humans than when we're talking about computers,” said Phillips, professor and head of the Department of History in CMU’s Dietrich College of Humanities and Social Sciences. “The question we wanted to ask was: Why are we using the same words at all?”
"The work Chris and I are doing focuses on how human beings have allowed computers — tools of our own invention — to become an integral part of our daily life,” Langmead said. “Because they are now enmeshed in our social world, how we talk about them and imagine them matters a great deal. We would like for there to be a larger, ongoing conversation that transcends the hype cycle about precisely what computers can and cannot do."
How language shapes the AI conversation
Rather than debating whether machines can think like humans, Phillips and Langmead looked backward, examining conversations about computing during the 1950s and 1960s — before artificial intelligence became a household term and before computers became widespread.
They found that today's debates are far from new.
Some early computing pioneers embraced human-centered language to describe computers, while others deliberately chose more precise, if less elegant, descriptions of what machines actually accomplished. Computer scientist Norbert Wiener, for example, described computers as “learning” when they implemented rules that resulted in more successful outcomes. To fellow researchers, the way he used the term had a well-defined technical meaning. To broader audiences, however, it could evoke the much richer human experience of learning.
For Phillips and Langmead, that distinction matters.
Strategic ambiguity allows technical language to be easily understood across audiences while sometimes making technologies appear more humanlike than they are. The practice isn't necessarily intentional, Phillips said, but it can shape how society understands AI.
“Why can't we say the computer is executing a particular set of instructions?” Phillips said. “Why do we have to call it thinking?”
The researchers stress that this does not diminish the accomplishments of modern AI. Phillips described today's large language models as amazing technological achievements capable of producing outputs that humans immediately recognize as meaningful.
Instead, Phillips argues that their accomplishments are remarkable precisely because they differ from human cognition.
“When you ask an image generator for a dog riding a pony in a New York Mets parade, and it produces exactly what you imagined, that's amazing,” Phillips said. “But let's not call that creativity. Let's call it the technical achievement that it is.”
Lessons from computing's early history
The paper also revisits modern benchmarks used to evaluate AI systems. Tests such as Massive Multitask Language Understanding, or MMLU, and the recently introduced “Humanity's Last Exam” are frequently described as measuring machine knowledge or reasoning. Phillips and Langmead argue that those benchmarks more accurately measure classification accuracy — how well systems identify correct answers on standardized evaluations — rather than demonstrating human knowledge or understanding.
When AI is described as thinking or writing, people can begin viewing machines as competitors rather than tools. That framing, Phillips argues, risks reducing complex human activities such as creativity, learning and reading to computational outputs while overlooking the relationships, lived experiences and judgment that shape those processes.
“Most of us read poetry because we're interested in the person who wrote it,” Phillips said. “We're interested in the emotion, the lived experience and the beauty that comes from having an actual human being produce something.”
One of the paper's more surprising discoveries was how interdisciplinary conversations about computing once were.
Phillips and Langmead found that during the mid-20th century, historians, literary scholars, psychologists, engineers and computer scientists all participated in discussions about what computers should do and where they belonged in society.
“It wasn't obvious who should control computing or what role computers should play,” Phillips said. “People across disciplines were asking those questions together.”
Why precision matters today
Today, Phillips believes those broader conversations are just as important. Across Dietrich College, colleagues in the humanities and social sciences bring different disciplinary perspectives to questions about AI, language and human knowledge
“This timely study reminds us that language does not simply deliver scientific or technological ideas: it changes these ideas and it transforms our understanding of them,”said Andreea Ritivoi, William S. Dietrich Professor of English and associate dean of research in Dietrich College. “The authors' compelling plea for precision is a great opportunity to underscore the need for productive collaborations across the ‘two cultures’ of science and humanities, as C.P. Snow, one of the heroes in this article, insisted decades ago. Too often the humanities and STEM are seen as islands of sorts, but the debates around AI happen in the troubled waters between them. We remain on one island at our own peril.”
Ritivoi’s perspective on language and interdisciplinary collaboration is complemented by Rob Kass’ focus on the importance of precision in the technical and statistical dimensions of AI.
“Computational algorithms, and mathematical derivations, rely on precision. Everyday language, however, makes heavy use of individual words that have multiple context-dependent meanings: they are often understood in a particular way by small groups of people within a limited setting, but others may easily attribute to those words very different meanings,”said Kass, Maurice Falk University Professor of Statistics & Computational Neuroscience. “This is a big issue in statistical reasoning from data. Phillips and Langmead convincingly document both the early appearance of ambiguous terminology in AI research and the ways it continues to be used for strategic advantage, much to the detriment of society as a whole.”
Ultimately, Phillips hopes the paper encourages a more thoughtful conversation about AI — one grounded less in sweeping claims about machine intelligence and more in an honest assessment of what these systems actually do.
“We're not arguing that these technologies aren't impressive,” Phillips said. “We're asking people to be very clear about what the machines do and don't do.”