
Artificial intelligence models that generate audio are being trained on datasets plagued with bias, offensive language and potential copyright infringement, sparking concerns about their use.
Generative audio products, such as song generators, voice cloning tools and transcription services, are increasingly popular, but while text and image generators have been subject to much scrutiny, audio has received less attention.
To help rectify this, at Carnegie Mellon University in Pennsylvania and his colleagues combed a yearâs worth of audio-modelling studies to spot common datasets used to train AI. In total, they found 175 datasets containing more than 680,000 hours of audio between them.
Advertisement
The team then conducted a more detailed audit on seven of the most commonly used datasets, which were mostly in English. These included voice recordings, such as sentences read by volunteers for the Mozilla Common Voice project, music recordings such as the Free Music Archive, and AudioSet, a set of 2 million 10-second YouTube clips.
To investigate biases, the team obtained transcripts of audio that contained words, such as speech or song lyrics, and identified keywords. With the exception of Mozilla Common Voice, which was more balanced, the datasets used the words âmanâ or âmenâ more than three times as frequently as âwomanâ or âwomenâ. Across datasets, âmanâ was strongly associated with words like âwarâ, âkillâ and âhistoryâ, while words associated with âwomanâ included âstoreâ, âmomâ and âbitchâ.
Many other words linked with identity appeared infrequently, suggesting a lack of representation. âMuslimâ appeared five to 10 times less than âChristianâ, while ânon-binaryâ only occurred around 10 times in total.
The researchers also tallied instances of profanity in the text and found that two datasets in particular â Free Music Archive and LibriVox â contained many thousands of occurrences of racist and queerphobic terms.
LibriVox comprises audio readings of public domain texts, meaning that the vast majority are over 70 years old and are likely to reflect outdated views and language. âCultural norms have shifted a lot,â says team member , also at Carnegie Mellon.
The team also found that phrases such as âall rights reservedâ featured frequently in the datasets, suggesting they may contain significant amounts of copyrighted material.
at the University of Cambridge says the work fills an important gap in ethical discussions about audio AI, which are currently dominated by fears over deepfakes being used to imitate a personâs voice.
A risk of training AI models on audio datasets like those studied, she says, is that they could end up producing offensive content â repeating a slur, for instance. Context is key: a musician may reclaim a racial slur in a song, for example, but would AI systems understand that the same word isnât acceptable in other contexts, or perhaps for an AI model to produce at all? âI think there is a risk that they start using these kinds of terms unprompted or uncritically,â says McInerney.
arxiv