bg
News
18:36, 06 September 2026
views
6

St. Petersburg State University Researchers Teach Neural Network to Speak With Proper Intonation

Researchers at St. Petersburg State University have completed the first phase of a study aimed at improving the intonation of speech generated by neural network models.

Photo: Alexey Danichev / RIA Novosti

The approach was presented by Ulyana Kochetkova, Associate Professor at the Department of Phonetics and Foreign Language Teaching Methodology at SPbU. The work was carried out as part of a master's thesis and focused not on creating a new architecture, but on preparing high-quality training data for the existing Tacotron 2 model.

A key element of the study was a corpus created at the SPbU Department of Phonetics in the 2000s in collaboration with the Center for Speech Technologies. Its distinguishing feature is expert intonation annotation rather than automatic labeling.

"The beauty of our corpus lies in the fact that the intonation annotation was done by experts. That is, it was not network-generated, not automatic annotation, but expert annotation — which is, essentially, spot on," Kochetkova noted.

Training the model on this material noticeably improved the intonation of synthesized speech. On the MOS metric, the researchers achieved scores of 4.5 to 4.7 — a strong result given the limited volume of data.

According to the researchers, large datasets handle speech sounds well but struggle with subtle intonational nuances. In the future, the team plans to expand the experiment and add emotional coloring to the speech.

like
heart
fun
wow
sad
angry
Latest news
Important
Recommended
previous
next