(Credit: Gage Skidmore, CC BY-SA 3.0)
About 2,000 years ago, the ancient Greek philosopher Aristotelēs devised a way of constructing arguments. He called this “rhetoric,” describing how logic, an understanding of the audience's needs and comprehension, and the speaker's authority could be strategically used within an argument or speech to persuade others.
Beyond relying solely on logic or trust in the speaker, politicians and actors have long recognized that using emotion is often the most effective way to capture an audience's heart—and, as a result, move it.
Last week's announcement of GPT-4o may have shown us a machine ideally suited to this task. While many view this as a remarkable advance with the potential to benefit countless people, some are approaching it with more caution.
Actress Scarlett Johansson, who had previously declined to provide a voice sample to OpenAI, said she was “shocked” and “angered” when she heard the new GPT-4o speak.
One of the five voices used by GPT-4o, named “Sky,” bore a striking resemblance to Samantha, the AI character Johansson voiced in the 2013 film Her. On the very day GPT-4o was announced, OpenAI founder and CEO Sam Altman tweeted “her”, seemingly underscoring the comparison between Sky and Samantha/Johansson.
OpenAI later posted on X that it was “working on pausing the use of Sky,” and on May 19 created a web page explaining that a different actress had been used for the voice. The company also detailed how the voices were selected.
The near-instant references to the film Her when GPT-4o was announced likely helped raise general public awareness of the technology, and perhaps made its capabilities seem less alarming.
This is fortunate, because iOS 18 is set to launch next month, and rumors of a partnership with Apple have already sparked privacy concerns. Similarly, OpenAI has partnered with Microsoft on its new generation of AI-powered Windows systems, “Copilot+ PC.”
Unlike other large language models (LLMs), GPT-4o (or “omni”) was built from the ground up to understand text, vision, and audio in a unified way. This represents true multimodality that goes far beyond the capabilities of “conventional” LLMs.
GPT-4o can recognize nuances of speech—emotion, breathing, ambient sounds, birdsong—and integrate this with visual information.
As a unified multimodal model (capable of processing both images and text), it responds at roughly the same speed as human speech (an average of 320 milliseconds) and can be interrupted. The result feels remarkably natural, and it can appropriately vary its tone and emotional intensity. It can even sing. Some have complained that GPT-4o is “flirtatious.” It's no wonder actors are concerned.
This represents a new way of interacting with AI, subtly shifting our relationship with technology and offering a new kind of “natural” interface—sometimes referred to as EAI (emotional AI).
The pace of this progress is unsettling many government agencies and police organizations. It remains unclear how to respond if this technology is weaponized by malicious states or criminals. With the rise of audio deepfakes, distinguishing what's real from what's fake is becoming increasingly difficult—even Johansson's own friends reportedly mistook the AI voice for her.
In a year with elections involving more than 4 billion potential voters worldwide, and amid a rise in scams centered on targeted deepfake audio, the dangers of weaponized AI should not be underestimated.
As Aristotelēs discovered, persuasive power often lies not in what is said, but in how it is said. We all carry unconscious biases, as highlighted by an intriguing UK report on accent bias. Some accents are perceived as more trustworthy, authoritative, or credible than others. For precisely this reason, call center workers have used AI to “westernize” their voices. In the case of GPT-4o, how something is said may matter just as much as what is said.
If AI can understand an audience's needs and reason logically, then—just as Aristotelēs pointed out 2,000 years ago—what may be needed is simply the right way to deliver a message. If so, we could end up creating an AI that becomes a superhuman master of rhetoric, wielding a persuasive power audiences find impossible to resist.
