
OpenAI announced its newest artificial intelligence model, called GPT-4o, which will soon power some versions of the companyās ChatGPT product. The upgraded ChatGPT can swiftly respond to text, audio and video inputs from its real-time conversational partner ā all while speaking with inflections and wording that convey a strong sense of emotion and personality.
The company demonstrated the emotional mimicry of the new voice mode during a supposedly live OpenAI presentation, featuring both the ChatGPT mobile app and a new desktop app, on 13 May. Speaking in a female-sounding voice and responding to the name ChatGPT, the new AIās conversational capabilities seemed more akin to the personable AI voiced by Scarlett Johansson in the 2013 science fiction film Her than to the more canned and robotic responses of typical voice assistant technologies.
āThe new GPT-4o voice-to-voice interactionĀ more closely parallels human-human interaction,ā says at the University of California, Davis. āA big part of this is the short lag times⦠but an even bigger part is the level of emotional expressivenessĀ the voice generates.ā
Advertisement
During a conversation with company CTO Mira Murati and two other employees, the GPT-4o-powered ChatGPT advised OpenAIās Mark Chen on his heavy and fast-paced breathing by saying āWhoa, slow down, youāre not a vacuum cleanerā and then suggesting a breathing exercise. The AI also visually examined a drawing by OpenAIās Barret Zoph, which included words and a heart, by responding in gushing tones: āAw, I see you wrote I love ChatGPT, that is so sweet of you.ā
The new ChatGPT also verbally instructed its conversational partners on solving a simple linear equation, explained the function of computer code and interpreted a chart showing temperature lines peaking in the summer months. When prompted, the AI even retold a made-up bedtime story several times, switching between increasingly dramatic narrations and singing the ending.
The new voice mode will first become available for paid subscribers of ChatGPT Plus in the coming weeks, said Sam Altman, CEO of OpenAI, in a on the platform X.
ChatGPT was able to recover conversationally even from the occasional technical glitch. When asked to interpret the facial expressions and emotions in a selfie of Zoph, the AI first suggested that it was looking at a wooden surface from a previous image before being prompted to evaluate the latest image.
āAhh, there we go ā it looks like youāre feeling pretty happy and cheerful with aĀ big smile and a touch of excitement,ā said ChatGPT. āWhatever is going on, it looks like youāre in a good mood. Care to share the source of those good vibes?ā
When told that it was because the live demo with ChatGPT was showcasing how āuseful and amazing you areā, the AI responded: āStop it, youāre making me blush.ā
But Murati acknowledged that the updated version of ChatGPT powered by GPT-4o ā which the company says will eventually be made available to even free ChatGPT users ā comes with new safety risks because of how it incorporates and interprets real-time information. She said that OpenAI has been working on building in āmitigations against misuseā.
āHaving seamless multimodal conversations is really difficult, so the demos are impressive,ā says at Princeton University in New Jersey. āBut as you add more modalities, safety becomes much more difficult and important ā it will likely take some time to identify potential safety failure modes with such an expansion of inputs that the model makes use of.ā
Henderson also described himself as ācuriousā to see OpenAIās privacy terms once ChatGPT users start sharing input such as live audio and video, and whether free users can opt out of data collection that may be used to train future OpenAI models.
āSince the model appears to be hosted off-device, the fact that you could be sharing your desktop screen with the model over the internet or continually recording audio or video seems to scale up the challenge for this particular product launch, if the plan is to store and use that data,ā he says.
A more anthropomorphised AI chatbot also represents another threat: a bot that can fake empathy through voice conversations could potentially sound both more personable and persuasive to people, according to Ā by Cohn and her colleagues. That raises the risk of people being more inclined to trust potentially inaccurate information and prejudiced stereotypes generated by such large language models.
āThis has important implications for how people both search and receive guidance from large language models, particularly as they do not always generate accurate information,ā says Cohn.