cs.SD语音【1】Enhancing the analysis of murine neonatal ultrasonic vocalizations: Development, evaluation, and application of different mathematical models标题:加强对小鼠新生儿超声发声的分析:不同数学模型的开发、评估和应用链接:https://arxiv.org/abs/2405.12957作者:Rudolf Herdt,Louisa Kinzel,Johann Georg Maaß,Marvin Walther,Henning Fröhlich,Tim Schubert,Peter Maass,Christian Patrick Schaaf摘要:啮齿动物使用超声波发声(USV)进行社会交流。由于这些发声为动物的情感状态、社会互动和发育阶段提供了有价值的见解,各种深度学习方法的目标是自动化USV的定量(检测)和定性(分类)分析。在这里,我们首次对不同类型的神经网络进行系统评估,用于USV分类。我们评估了各种前馈网络,包括定制的全连接网络和卷积神经网络,不同的残差神经网络(ResNets),EfficientNet和Vision Transformer(ViT)。与一个精炼的基于熵的检测算法(实现94.9%的召回率和99.3%的精度)相结合,最佳架构(实现86.79%的准确率)被集成到一个完全自动化的管道中,能够以高可靠性分析广泛的USV数据集。此外,用户可以根据自己的研究需求指定个人最低准确度阈值。在这种半自动化设置中,管道选择性地对具有高伪概率的调用进行分类,将其余部分留给手动检查。我们的研究仅关注新生儿USV。作为正在进行的表型研究的一部分,我们的管道已被证明是一个有价值的工具,用于识别具有自闭症样行为的小鼠产生的USV的关键差异。摘要:Rodents employ a broad spectrum of ultrasonic vocalizations (USVs) for social communication. As these vocalizations offer valuable insights into affective states, social interactions, and developmental stages of animals, various deep learning approaches have aimed to automate both the quantitative (detection) and qualitative (classification) analysis of USVs. Here, we present the first systematic evaluation of different types of neural networks for USV classification. We assessed various feedforward networks, including a custom-built, fully-connected network and convolutional neural network, different residual neural networks (ResNets), an EfficientNet, and a Vision Transformer (ViT). Paired with a refined, entropy-based detection algorithm (achieving recall of 94.9% and precision of 99.3%), the best architecture (achieving 86.79% accuracy) was integrated into a fully automated pipeline capable of analyzing extensive USV datasets with high reliability. Additionally, users can specify an individual minimum accuracy threshold based on their research needs. In this semi-automated setup, the pipeline selectively classifies calls with high pseudo-probability, leaving the rest for manual inspection. Our study focuses exclusively on neonatal USVs. As part of an ongoing phenotyping study, our pipeline has proven to be a valuable tool for identifying key differences in USVs produced by mice with autism-like behaviors.
【2】 A Dataset and Baselines for Measuring and Predicting the Music Piece Memorability标题:用于测量和预测音乐作品互换性的数据集和基线链接:https://arxiv.org/abs/2405.12847作者:Li-Yang Tseng,Tzu-Ling Lin,Hong-Han Shuai,Jen-Wei Huang,Wen-Whei Chang备注:None摘要:如今,人们不断地接触音乐,无论是通过自愿的流媒体服务还是在商业休息期间的偶然相遇。尽管音乐丰富,但某些作品仍然更令人难忘,而且往往更受欢迎。受这一现象的启发,我们专注于测量和预测音乐记忆力。为了实现这一点,我们收集了一个新的音乐作品数据集与可靠的记忆标签使用一种新的交互式实验程序。然后,我们训练基线来预测和分析音乐记忆性,利用可解释的特征和音频梅尔频谱图作为输入。据我们所知,我们是第一个使用基于数据驱动的深度学习方法探索音乐记忆力的人。通过一系列的实验和消融研究,我们证明,虽然有改进的余地,预测音乐记忆有限的数据是可能的。某些内在的元素,如更高的效价,唤醒和更快的节奏,有助于令人难忘的音乐。随着预测技术的不断发展,音乐推荐系统和音乐风格转移等实际应用无疑将受益于这一新的研究领域。摘要:Nowadays, humans are constantly exposed to music, whether through voluntary streaming services or incidental encounters during commercial breaks. Despite the abundance of music, certain pieces remain more memorable and often gain greater popularity. Inspired by this phenomenon, we focus on measuring and predicting music memorability. To achieve this, we collect a new music piece dataset with reliable memorability labels using a novel interactive experimental procedure. We then train baselines to predict and analyze music memorability, leveraging both interpretable features and audio mel-spectrograms as inputs. To the best of our knowledge, we are the first to explore music memorability using data-driven deep learning-based methods. Through a series of experiments and ablation studies, we demonstrate that while there is room for improvement, predicting music memorability with limited data is possible. Certain intrinsic elements, such as higher valence, arousal, and faster tempo, contribute to memorable music. As prediction techniques continue to evolve, real-life applications like music recommendation systems and music style transfer will undoubtedly benefit from this new area of research.
【3】 SYMPLEX: Controllable Symbolic Music Generation using Simplex Diffusion with Vocabulary Priors标题:SYMPFLEX:使用具有词汇先验的单纯形扩散的可控符号音乐生成链接:https://arxiv.org/abs/2405.12666作者:Nicolas Jonason,Luca Casini,Bob L. T. Sturm摘要:我们提出了一种新的方法,快速和可控生成的符号音乐的基础上的单纯形扩散,这本质上是一个扩散过程中的概率,而不是信号空间。这个目标已经应用于自然语言处理等领域,但在这里,我们将其应用于使用无序表示生成4小节多乐器音乐循环。我们表明,我们的模型可以与词汇先验,这提供了相当大的水平控制音乐生成过程中,例如,填充时间和音高和乐器的选择-所有没有特定任务的模型适应或应用外部控制。摘要:We present a new approach for fast and controllable generation of symbolic music based on the simplex diffusion, which is essentially a diffusion process operating on probabilities rather than the signal space. This objective has been applied in domains such as natural language processing but here we apply it to generating 4-bar multi-instrument music loops using an orderless representation. We show that our model can be steered with vocabulary priors, which affords a considerable level control over the music generation process, for instance, infilling in time and pitch and choice of instrumentation -- all without task-specific model adaptation or applying extrinsic control. 【4】 On a time-frequency blurring operator with applications in data augmentation标题:关于时频模糊运算符及其在数据增强中的应用链接:https://arxiv.org/abs/2405.12899作者:Simon Halvdansson备注:22 pages, 4 figures摘要:受最近的数据增强方法的成功的启发,作用于时间-频率表示的信号,我们引入了一个操作员卷积的信号的短时傅立叶变换与指定的内核。分析性质,包括有界性,紧性和积极性的时频分析的角度进行了研究。训练卷积神经网络和Vision Transformer以使用具有不同增强设置的频谱图对音频信号进行分类,包括上述时间-频率模糊算子,结果表明该算子可以显著提高测试性能,特别是在数据匮乏的情况下。摘要:Inspired by the success of recent data augmentation methods for signals which act on time-frequency representations, we introduce an operator which convolves the short-time Fourier transform of a signal with a specified kernel. Analytical properties including boundedness, compactness and positivity are investigated from the perspective of time-frequency analysis. A convolutional neural network and a vision transformer are trained to classify audio signals using spectrograms with different augmentation setups, including the above mentioned time-frequency blurring operator, with results indicating that the operator can significantly improve test performance, especially in the data-starved regime.
【5】 Mamba in Speech: Towards an Alternative to Self-Attention标题:言语中的曼巴舞:走向自我注意力的替代方案链接:https://arxiv.org/abs/2405.12609作者:Xiangyu Zhang,Qiquan Zhang,Hexin Liu,Tianyi Xiao,Xinyuan Qian,Beena Ahmed,Eliathamby Ambikairajah,Haizhou Li,Julien Epps摘要:Transformer及其衍生产品已在计算机视觉、自然语言处理和语音处理的各种任务中取得成功。为了降低在Transformer中的多头自注意机制内的计算的复杂性,选择性状态空间模型(即,Mamba)被提议作为替代方案。Mamba在自然语言处理和计算机视觉任务中表现出了其有效性,但其优越性在语音信号处理中很少被研究。本文探讨了将Mamba应用于语音处理的解决方案,使用两个典型的语音处理任务:语音识别,需要语义和顺序信息,语音增强,主要集中在顺序模式。结果表明,双向曼巴(BiMamba)的语音处理优于香草曼巴。此外,实验证明了BiMamba作为Transformer及其衍生物中自我注意模块的替代品的有效性,特别是对于语义感知任务。然后在消融研究和讨论部分总结了将曼巴转移到语音的关键技术,为未来的研究提供见解。摘要:Transformer and its derivatives have achieved success in diverse tasks across computer vision, natural language processing, and speech processing. To reduce the complexity of computations within the multi-head self-attention mechanism in Transformer, Selective State Space Models (i.e., Mamba) were proposed as an alternative. Mamba exhibited its effectiveness in natural language processing and computer vision tasks, but its superiority has rarely been investigated in speech signal processing. This paper explores solutions for applying Mamba to speech processing using two typical speech processing tasks: speech recognition, which requires semantic and sequential information, and speech enhancement, which focuses primarily on sequential patterns. The results exhibit the superiority of bidirectional Mamba (BiMamba) for speech processing to vanilla Mamba. Moreover, experiments demonstrate the effectiveness of BiMamba as an alternative to the self-attention module in Transformer and its derivates, particularly for the semantic-aware task. The crucial technologies for transferring Mamba to speech are then summarized in ablation studies and the discussion section to offer insights for future research.
【6】 A Survey of Integrating Wireless Technology into Active Noise Control标题:无线技术融入主动噪音控制的综述链接:https://arxiv.org/abs/2405.12496作者:Xiaoyi Shen,Dongyuan Shi,Zhengding Luo,Junwei Ji,Woon-Seng Gan摘要:主动噪声控制(ANC)是一种广泛采用的技术,用于在各种场景中降低环境噪声。本文着重于提高降噪性能,特别是通过细化信号质量馈入ANC系统。我们讨论的主要无线技术集成到ANC系统,配备了一些创新的算法,在不同的环境。无线技术的应用避免了额外的计算需求,而不是使用麦克风阵列来隔离多个噪声源以提高降噪性能,这增加了ANC系统的计算复杂度。参考信号、误差信号和控制信号的无线传输也被应用于改善ANC系统的收敛性能。此外,本文还列出了一些无线ANC应用,如耳塞、耳机、窗户和头枕,强调了它们在各种环境中的适应性和效率。摘要:Active Noise Control (ANC) is a widely adopted technology for reducing environmental noise across various scenarios. This paper focuses on enhancing noise reduction performance, particularly through the refinement of signal quality fed into ANC systems. We discuss the main wireless technique integrated into the ANC system, equipped with some innovative algorithms, in diverse environments. Instead of using microphone arrays, which increase the computation complexity of the ANC system, to isolate multiple noise sources to improve noise reduction performance, the application of the wireless technique avoids extra computation demand. Wireless transmissions of reference, error, and control signals are also applied to improve the convergence performance of the ANC system. Furthermore, this paper lists some wireless ANC applications, such as earbuds, headphones, windows, and headrests, underscoring their adaptability and efficiency in various settings. eess.AS音频处理【1】 Mamba in Speech: Towards an Alternative to Self-Attention标题:言语中的曼巴舞:走向自我注意力的替代方案链接:https://arxiv.org/abs/2405.12609作者:Xiangyu Zhang,Qiquan Zhang,Hexin Liu,Tianyi Xiao,Xinyuan Qian,Beena Ahmed,Eliathamby Ambikairajah,Haizhou Li,Julien Epps摘要:Transformer及其衍生产品已在计算机视觉、自然语言处理和语音处理的各种任务中取得成功。为了降低在Transformer中的多头自注意机制内的计算的复杂性,选择性状态空间模型(即,Mamba)被提议作为替代方案。Mamba在自然语言处理和计算机视觉任务中表现出了其有效性,但其优越性在语音信号处理中很少被研究。本文探讨了将Mamba应用于语音处理的解决方案,使用两个典型的语音处理任务:语音识别,需要语义和顺序信息,语音增强,主要集中在顺序模式。结果表明,双向曼巴(BiMamba)的语音处理优于香草曼巴。此外,实验证明了BiMamba作为Transformer及其衍生物中自我注意模块的替代品的有效性,特别是对于语义感知任务。然后在消融研究和讨论部分总结了将曼巴转移到语音的关键技术,为未来的研究提供见解。摘要:Transformer and its derivatives have achieved success in diverse tasks across computer vision, natural language processing, and speech processing. To reduce the complexity of computations within the multi-head self-attention mechanism in Transformer, Selective State Space Models (i.e., Mamba) were proposed as an alternative. Mamba exhibited its effectiveness in natural language processing and computer vision tasks, but its superiority has rarely been investigated in speech signal processing. This paper explores solutions for applying Mamba to speech processing using two typical speech processing tasks: speech recognition, which requires semantic and sequential information, and speech enhancement, which focuses primarily on sequential patterns. The results exhibit the superiority of bidirectional Mamba (BiMamba) for speech processing to vanilla Mamba. Moreover, experiments demonstrate the effectiveness of BiMamba as an alternative to the self-attention module in Transformer and its derivates, particularly for the semantic-aware task. The crucial technologies for transferring Mamba to speech are then summarized in ablation studies and the discussion section to offer insights for future research. 【2】 A Survey of Integrating Wireless Technology into Active Noise Control标题:无线技术融入主动噪音控制的综述链接:https://arxiv.org/abs/2405.12496作者:Xiaoyi Shen,Dongyuan Shi,Zhengding Luo,Junwei Ji,Woon-Seng Gan摘要:主动噪声控制(ANC)是一种广泛采用的技术,用于在各种场景中降低环境噪声。本文着重于提高降噪性能,特别是通过细化信号质量馈入ANC系统。我们讨论的主要无线技术集成到ANC系统,配备了一些创新的算法,在不同的环境。无线技术的应用避免了额外的计算需求,而不是使用麦克风阵列来隔离多个噪声源以提高降噪性能,这增加了ANC系统的计算复杂度。参考信号、误差信号和控制信号的无线传输也被应用于改善ANC系统的收敛性能。此外,本文还列出了一些无线ANC应用,如耳塞、耳机、窗户和头枕,强调了它们在各种环境中的适应性和效率。摘要:Active Noise Control (ANC) is a widely adopted technology for reducing environmental noise across various scenarios. This paper focuses on enhancing noise reduction performance, particularly through the refinement of signal quality fed into ANC systems. We discuss the main wireless technique integrated into the ANC system, equipped with some innovative algorithms, in diverse environments. Instead of using microphone arrays, which increase the computation complexity of the ANC system, to isolate multiple noise sources to improve noise reduction performance, the application of the wireless technique avoids extra computation demand. Wireless transmissions of reference, error, and control signals are also applied to improve the convergence performance of the ANC system. Furthermore, this paper lists some wireless ANC applications, such as earbuds, headphones, windows, and headrests, underscoring their adaptability and efficiency in various settings.
【3】 Enhancing the analysis of murine neonatal ultrasonic vocalizations: Development, evaluation, and application of different mathematical models标题:加强对小鼠新生儿超声发声的分析:不同数学模型的开发、评估和应用链接:https://arxiv.org/abs/2405.12957作者:Rudolf Herdt,Louisa Kinzel,Johann Georg Maaß,Marvin Walther,Henning Fröhlich,Tim Schubert,Peter Maass,Christian Patrick Schaaf摘要:啮齿动物使用超声波发声(USV)进行社会交流。由于这些发声为动物的情感状态、社会互动和发育阶段提供了有价值的见解,各种深度学习方法的目标是自动化USV的定量(检测)和定性(分类)分析。在这里,我们首次对不同类型的神经网络进行系统评估,用于USV分类。我们评估了各种前馈网络,包括定制的全连接网络和卷积神经网络,不同的残差神经网络(ResNets),EfficientNet和Vision Transformer(ViT)。与一个精炼的基于熵的检测算法(实现94.9%的召回率和99.3%的精度)相结合,最佳架构(实现86.79%的准确率)被集成到一个完全自动化的管道中,能够以高可靠性分析广泛的USV数据集。此外,用户可以根据自己的研究需求指定个人最低准确度阈值。在这种半自动化设置中,管道选择性地对具有高伪概率的调用进行分类,将其余部分留给手动检查。我们的研究仅关注新生儿USV。作为正在进行的表型研究的一部分,我们的管道已被证明是一个有价值的工具,用于识别具有自闭症样行为的小鼠产生的USV的关键差异。摘要:Rodents employ a broad spectrum of ultrasonic vocalizations (USVs) for social communication. As these vocalizations offer valuable insights into affective states, social interactions, and developmental stages of animals, various deep learning approaches have aimed to automate both the quantitative (detection) and qualitative (classification) analysis of USVs. Here, we present the first systematic evaluation of different types of neural networks for USV classification. We assessed various feedforward networks, including a custom-built, fully-connected network and convolutional neural network, different residual neural networks (ResNets), an EfficientNet, and a Vision Transformer (ViT). Paired with a refined, entropy-based detection algorithm (achieving recall of 94.9% and precision of 99.3%), the best architecture (achieving 86.79% accuracy) was integrated into a fully automated pipeline capable of analyzing extensive USV datasets with high reliability. Additionally, users can specify an individual minimum accuracy threshold based on their research needs. In this semi-automated setup, the pipeline selectively classifies calls with high pseudo-probability, leaving the rest for manual inspection. Our study focuses exclusively on neonatal USVs. As part of an ongoing phenotyping study, our pipeline has proven to be a valuable tool for identifying key differences in USVs produced by mice with autism-like behaviors. 【4】 On a time-frequency blurring operator with applications in data augmentation标题:关于时频模糊运算符及其在数据增强中的应用链接:https://arxiv.org/abs/2405.12899作者:Simon Halvdansson备注:22 pages, 4 figures摘要:受最近的数据增强方法的成功的启发,作用于时间—频率表示的信号,我们引入了一个操作员卷积的信号的短时傅立叶变换与指定的内核。分析性质,包括有界性,紧性和积极性的时频分析的角度进行了研究。训练卷积神经网络和Vision Transformer以使用具有不同增强设置的频谱图对音频信号进行分类,包括上述时间—频率模糊算子,结果表明该算子可以显著提高测试性能,特别是在数据匮乏的情况下。摘要:Inspired by the success of recent data augmentation methods for signals which act on time-frequency representations, we introduce an operator which convolves the short-time Fourier transform of a signal with a specified kernel. Analytical properties including boundedness, compactness and positivity are investigated from the perspective of time-frequency analysis. A convolutional neural network and a vision transformer are trained to classify audio signals using spectrograms with different augmentation setups, including the above mentioned time-frequency blurring operator, with results indicating that the operator can significantly improve test performance, especially in the data-starved regime.
【5】 A Dataset and Baselines for Measuring and Predicting the Music Piece Memorability标题:用于测量和预测音乐作品互换性的数据集和基线链接:https://arxiv.org/abs/2405.12847作者:Li-Yang Tseng,Tzu-Ling Lin,Hong-Han Shuai,Jen-Wei Huang,Wen-Whei Chang备注:None摘要:如今,人们不断地接触音乐,无论是通过自愿的流媒体服务还是在商业休息期间的偶然相遇。尽管音乐丰富,但某些作品仍然更令人难忘,而且往往更受欢迎。受这一现象的启发,我们专注于测量和预测音乐记忆力。为了实现这一点,我们收集了一个新的音乐作品数据集与可靠的记忆标签使用一种新的交互式实验程序。然后,我们训练基线来预测和分析音乐记忆性,利用可解释的特征和音频梅尔频谱图作为输入。据我们所知,我们是第一个使用基于数据驱动的深度学习方法探索音乐记忆力的人。通过一系列的实验和消融研究,我们证明,虽然有改进的余地,预测音乐记忆有限的数据是可能的。某些内在的元素,如更高的效价,唤醒和更快的节奏,有助于令人难忘的音乐。随着预测技术的不断发展,音乐推荐系统和音乐风格转移等实际应用无疑将受益于这一新的研究领域。摘要:Nowadays, humans are constantly exposed to music, whether through voluntary streaming services or incidental encounters during commercial breaks. Despite the abundance of music, certain pieces remain more memorable and often gain greater popularity. Inspired by this phenomenon, we focus on measuring and predicting music memorability. To achieve this, we collect a new music piece dataset with reliable memorability labels using a novel interactive experimental procedure. We then train baselines to predict and analyze music memorability, leveraging both interpretable features and audio mel-spectrograms as inputs. To the best of our knowledge, we are the first to explore music memorability using data-driven deep learning-based methods. Through a series of experiments and ablation studies, we demonstrate that while there is room for improvement, predicting music memorability with limited data is possible. Certain intrinsic elements, such as higher valence, arousal, and faster tempo, contribute to memorable music. As prediction techniques continue to evolve, real-life applications like music recommendation systems and music style transfer will undoubtedly benefit from this new area of research.
【6】 Blind Separation of Vibration Sources using Deep Learning and Deconvolution标题:使用深度学习和反卷积进行振动源盲分离链接:https://arxiv.org/abs/2405.12774作者:Igor Makienko,Michael Grebshtein,Eli Gildish备注:20 pages, 13 figures摘要:旋转机械的振动主要来源于两个来源,这两个来源在到达传感器的过程中都会被机器的传递函数所扭曲:与齿轮相关的主要振动和与轴承故障相关的低能量信号。所提出的方法有利于振动源的盲分离,消除了对被监测设备或外部测量的任何信息的需要。该方法分两个阶段估计两个源:首先,使用扩张的CNN隔离齿轮信号,然后使用残差的平方对数包络估计轴承故障信号。使用一种新的基于白化的反卷积方法(WBD)从两个源中去除传递函数的影响。仿真和实验结果表明,该方法的能力,早期检测轴承故障时,没有额外的信息。本研究考虑了局部和分布式轴承故障,假设振动记录在稳定的操作条件下。摘要:Vibrations of rotating machinery primarily originate from two sources, both of which are distorted by the machine's transfer function on their way to the sensor: the dominant gear-related vibrations and a low-energy signal linked to bearing faults. The proposed method facilitates the blind separation of vibration sources, eliminating the need for any information about the monitored equipment or external measurements. This method estimates both sources in two stages: initially, the gear signal is isolated using a dilated CNN, followed by the estimation of the bearing fault signal using the squared log envelope of the residual. The effect of the transfer function is removed from both sources using a novel whitening-based deconvolution method (WBD). Both simulation and experimental results demonstrate the method's ability to detect bearing failures early when no additional information is available. This study considers both local and distributed bearing faults, assuming that the vibrations are recorded under stable operating conditions.
【7】 SYMPLEX: Controllable Symbolic Music Generation using Simplex Diffusion with Vocabulary Priors标题:SYMPFLEX:使用具有词汇先验的单纯形扩散的可控符号音乐生成链接:https://arxiv.org/abs/2405.12666作者:Nicolas Jonason,Luca Casini,Bob L. T. Sturm摘要:我们提出了一种新的方法,快速和可控生成的符号音乐的基础上的单纯形扩散,这本质上是一个扩散过程中的概率,而不是信号空间。这个目标已经应用于自然语言处理等领域,但在这里,我们将其应用于使用无序表示生成4小节多乐器音乐循环。我们表明,我们的模型可以与词汇先验,这提供了相当大的水平控制音乐生成过程中,例如,填充时间和音高和乐器的选择-所有没有特定任务的模型适应或应用外部控制。摘要:We present a new approach for fast and controllable generation of symbolic music based on the simplex diffusion, which is essentially a diffusion process operating on probabilities rather than the signal space. This objective has been applied in domains such as natural language processing but here we apply it to generating 4-bar multi-instrument music loops using an orderless representation. We show that our model can be steered with vocabulary priors, which affords a considerable level control over the music generation process, for instance, infilling in time and pitch and choice of instrumentation -- all without task-specific model adaptation or applying extrinsic control.