OKI and Chuo University Develop AI Self-Supervised Learning Method; Paper Published in IEEE TASLP|電経新聞

OKI and Chuo University Develop AI Self-Supervised Learning Method; Paper Published in IEEE TASLP

OKI and the Faculty of Science and Engineering at Chuo University have jointly developed a self-supervised learning method that addresses the distortion caused by short-term temporal variations when segmenting audio data into units suitable for AI processing. This achievement was published in the peer-reviewed international journal *IEEE Transactions on Audio, Speech and Language Processing* (TASLP).
Evaluations using multiple publicly available datasets demonstrated that the method achieves recognition performance equal to or better than leading existing techniques and offers improved stability against slight temporal shifts in the audio.
Building on this result, the team aims to advance AI technology capable of more stably learning the characteristics of sound and vibration—even in environments where large amounts of labeled training data are difficult to obtain—to support the maintenance and management of social infrastructure such as bridges, roads, and tunnels. This advancement will facilitate AI technologies that help streamline inspections and enable the early detection of anomalies in the field.
Unlike methods that rely solely on features obtained by dividing spectrograms into a grid, this new approach incorporates short-term variation patterns—which are often lost during segmentation—into the learning process. A key feature is the integration of a mechanism within the self-supervised learning framework that utilizes this distortion as useful information for sound recognition, rather than simply eliminating it. The method’s effectiveness was verified using public datasets covering environmental sounds, speech, and music.
A spectrogram is a visual representation of waveforms (such as sound or vibration) that displays frequency components and their intensities over time. It makes it easier to identify the characteristics of speech and mechanical sounds and is utilized in applications such as anomaly sound detection and acoustic analysis.

※Translating Japanese articles into English with AI