Norwegian Researchers Create Artificial Children's Voice

Getting a computer to understand the voice of a child is a great challenge. It is also difficult to get a computer to synthesise speech in a child’s voice. Norwegian researchers have found simple, effective solutions to both challenges.

“Synthesised speech has grown more and more similar to human speech. Yet children communicating via a speech device are still forced to use a synthetic adult voice,” explains Magne Lunde, Managing Director of Media LT, a company developing tools to assist disabled persons.

This drawback was the driver behind a collaborative research project involving MedialT and Lingit, a software company. Together they are developing Norway’s first synthesised childlike voice.

“We start with what is known as a master voice, which is the product of three or four adult speakers recording several thousands of phrases. Then we record a single child reading a smaller number of phrases aloud. We use this recording to modify the master voice, making it sound like a child’s voice,” relates Torbjørn Nordgård from Lingit. Dr Nordgård is also a professor of linguistics at the University of Nordland.

The phrases recorded by the child have been selected to include a number of the most essential sounds found in Norwegian.

“The master voice still carries the intonation, i.e. a phrase’s melody. The result sounds rather like a child with unusual elocution skills, but it’s still much better than the voice of an adult,” says Mr Nordgård.

Very little research has been carried out on this subject internationally. MediaLT and Linget’s innovative method of synthesising a child’s voice is bringing them to the forefront of their field.

Everything is now in place to start testing trial versions of the child’s voice. “We hope to have a beta version in place this summer,” says Magne Lunde.

Mr Lunde and his colleagues are also researching voice control such as use of verbal commands to operate a PC.

In order to operate a computer by means of speech, the machine must successfully decipher what is being said. Interpreting the speech of individuals on both the young and the older end of the scale is especially challenging since the distance from their vocal cords to their lips is shorter than that of the average adult.

“Teaching a speech recognition program to understand the pronunciation of the various sounds of a language requires a relatively large amount of recorded speech. Unfortunately, insufficient data exist today in terms of actual children’s speech,” states Professor Torbjørn Svendsen from the Norwegian University of Science and Technology.

“The converted adult speech resembles the way children speak in terms of sound as well. Thus, we could apply our conversion technique to a large database of adult speech and generate a functional database of artificial childlike voices. We then used this to train a separate speech recognition program for children,” explains Professor Svendsen.

“This greatly improved the recognition fidelity of children’s speech. The error rate was reduced by 50 to 70 per cent,” he states.

The Norwegian language poses a number of especially steep challenges to speech recognition experts.

The Research Council of Norway:

Norwegian Researchers Create Artificial Children's Voice

Tuesday, February 21, 2012

0 comments:

Post a Comment

Popular Stories