今天跟大家分享一篇语音相关的论文合集:cs.SD语音4篇,eess.AS音频处理4篇。

cs.SD语音

【1】 End-To-End Audiovisual Feature Fusion for Active Speaker Detection

标题:基于端到端音视频特征融合的主动说话人检测

链接:https://arxiv.org/abs/2207.13434

作者:Fiseha B. Tesema,Zheyuan Lin,Shiqiang Zhu,Wei Song,Jason Gu,Hong Wu
备注:To appear on the proceeding of the Fourteenth International Conference on Digital Image Processing (ICDIP 2022), May 20-23, Wuhan, China, 8 pages, 3 figures
摘要:主动说话人检测在人机交互中起着至关重要的作用.近年来出现了一些端到端的视听系统框架.然而,这些模型的复杂性和输入数据量大,其推理时间没有得到充分的研究,不适用于实时应用.另外,他们探索了一种类似的特征提取策略,该策略将ConvNet应用于音频和视频输入。该网络将通过VGG-M从图像中提取的特征与从音频波形中提取的原始Mel频率倒谱系数特征进行融合,在AVA-300上的实验结果表明,该模型具有较好的动态响应特性. ActiveSpeaker数据集的实验结果表明,与ConvNet模型相比,本文提出的特征提取方法具有更好的鲁棒性和更快的推理速度,该模型的检测精度达到了88.929%,与现有的检测方法的检测精度基本一致。
摘要:Active speaker detection plays a vital role in human-machine interaction. Recently, a few end-to-end audiovisual frameworks emerged. However, these models' inference time was not explored and are not applicable for real-time applications due to their complexity and large input size. In addition, they explored a similar feature extraction strategy that employs the ConvNet on audio and visual inputs. This work presents a novel two-stream end-to-end framework fusing features extracted from images via VGG-M with raw Mel Frequency Cepstrum Coefficients features extracted from the audio waveform. The network has two BiGRU layers attached to each stream to handle each stream's temporal dynamic before fusion. After fusion, one BiGRU layer is attached to model the joint temporal dynamics. The experiment result on the AVA-ActiveSpeaker dataset indicates that our new feature extraction strategy shows more robustness to noisy signals and better inference time than models that employed ConvNet on both modalities. The proposed model predicts within 44.41 ms, which is fast enough for real-time applications. Our best-performing model attained 88.929% accuracy, nearly the same detection result as state-of-the-art -work.


【2】 Perception-Aware Attack: Creating Adversarial Music via  Reverse-Engineering Human Perception

标题:感知感知攻击:通过逆向工程人类感知创建对抗性音乐

链接:https://arxiv.org/abs/2207.13192

作者:Rui Duan,Zhe Qu,Shangqing Zhao,Leah Ding,Yao Liu,Zhuo Lu
备注:ACM CCS 2022
摘要:近来,对抗性机器学习攻击已经对包括语音识别、说话人识别先前的研究主要集中在通过在原始信号上创建小的类似噪声的扰动来确保攻击音频信号分类器的有效性。目前还不清楚攻击者是否能够在攻击有效性的基础上产生人类能够很好感知的音频信号扰动。这对于音乐信号尤其重要,因为它们是精心制作的,具有人类喜欢的音频特征。  在本文中,我们将对抗性攻击的概念表述为一种新的感知感知攻击框架,它将人类研究与对抗性攻击的设计相结合.具体来说,我们进行了一项人类研究,以量化人类对音乐信号变化的感知.我们邀请人类参与者基于原始和扰动音乐信号对他们感知的偏差进行评级,然后,通过回归分析对人类感知过程进行逆向工程,以预测给定扰动信号下的人类感知偏差.感知感知感知攻击被表述为一个优化问题,即寻找最优扰动信号,以最小化由回归的人类感知模型预测的感知偏差.我们使用感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知实验结果表明,感知感知感知攻击产生的对抗性音乐的感知质量明显优于已有的攻击。
摘要:Recently, adversarial machine learning attacks have posed serious security threats against practical audio signal classification systems, including speech recognition, speaker recognition, and music copyright detection. Previous studies have mainly focused on ensuring the effectiveness of attacking an audio signal classifier via creating a small noise-like perturbation on the original signal. It is still unclear if an attacker is able to create audio signal perturbations that can be well perceived by human beings in addition to its attack effectiveness. This is particularly important for music signals as they are carefully crafted with human-enjoyable audio characteristics.  In this work, we formulate the adversarial attack against music signals as a new perception-aware attack framework, which integrates human study into adversarial attack design. Specifically, we conduct a human study to quantify the human perception with respect to a change of a music signal. We invite human participants to rate their perceived deviation based on pairs of original and perturbed music signals, and reverse-engineer the human perception process by regression analysis to predict the human-perceived deviation given a perturbed signal. The perception-aware attack is then formulated as an optimization problem that finds an optimal perturbation signal to minimize the prediction of perceived deviation from the regressed human perception model. We use the perception-aware framework to design a realistic adversarial music attack against YouTube's copyright detector. Experiments show that the perception-aware attack produces adversarial music with significantly better perceptual quality than prior work.


【3】 Knowledge-driven Subword Grammar Modeling for Automatic Speech  Recognition in Tamil and Kannada

标题:泰米尔语和卡纳达语自动语音识别中知识驱动的子词语法建模

链接:https://arxiv.org/abs/2207.13333

作者:Madhavaraj A,Bharathi Pilar,Ramakrishnan A G
摘要:本文提出了一种新的方法,我们提出了专门设计的自动语音识别(ASR)系统,它可以识别无限词汇量的单词,我们使用子词作为基本词汇单元进行识别,并构造了子词语法加权有限状态转换器(SG-WFST)图形的分词,它捕捉了语言中大多数复杂的构词规则。(i)动词,(ii)名词,(ii)代名词,以及(iv)号码。本文提出了一种启发式分词算法,该算法可以对不符合SG-WFST图规则的异常词进行分词,并利用该算法对SG-WFST图中的大部分数据进行了分词处理.驱动子词词典创建算法是计算驱动的,因此不能保证语素类单元,因此我们使用语言的语言学知识并手动创建子词词典和图。最后,我们训练深度神经网络声学模型,并将其与子词词典的发音词典和SG-WFST图相结合,以构建子词-由于子词-ASR产生子词序列作为给定测试语音的输出,我们对其输出进行后处理以获得最终的词序列,从而实际可识别的词的数量要高得多。在使用IISc-MILE泰米尔语和卡纳达语ASR语料库对子词-ASR系统进行实验时,我们观察到,与泰米尔语和卡纳达语的基线基于单词的ASR系统相比,绝对单词错误率分别减少了12.39%和13.56%。
摘要:In this paper, we present specially designed automatic speech recognition (ASR) systems for the highly agglutinative and inflective languages of Tamil and Kannada that can recognize unlimited vocabulary of words. We use subwords as the basic lexical units for recognition and construct subword grammar weighted finite state transducer (SG-WFST) graphs for word segmentation that captures most of the complex word formation rules of the languages. We have identified the following category of words (i) verbs, (ii) nouns, (ii) pronouns, and (iv) numbers. The prefix, infix and suffix lists of subwords are created for each of these categories and are used to design the SG-WFST graphs. We also present a heuristic segmentation algorithm that can even segment exceptional words that do not follow the rules encapsulated in the SG-WFST graph. Most of the data-driven subword dictionary creation algorithms are computation driven, and hence do not guarantee morpheme-like units and so we have used the linguistic knowledge of the languages and manually created the subword dictionaries and the graphs. Finally, we train a deep neural network acoustic model and combine it with the pronunciation lexicon of the subword dictionary and the SG-WFST graph to build the subword-ASR systems. Since the subword-ASR produces subword sequences as output for a given test speech, we post-process its output to get the final word sequence, so that the actual number of words that can be recognized is much higher. Upon experimenting the subword-ASR system with the IISc-MILE Tamil and Kannada ASR corpora, we observe an absolute word error rate reduction of 12.39% and 13.56% over the baseline word-based ASR systems for Tamil and Kannada, respectively.


【4】 Subword Dictionary Learning and Segmentation Techniques for Automatic  Speech Recognition in Tamil and Kannada

标题:用于泰米尔语和卡纳达语自动语音识别的子词词典学习和分割技术

链接:https://arxiv.org/abs/2207.13331

作者:Madhavaraj A,Bharathi Pilar,Ramakrishnan A G
摘要:提出了一种自动语音识别方法(ASR)系统,用于泰米尔语和卡纳达语,基于子词建模,以有效地处理无限的词汇量,由于语言的高度粘着性。(BPE)算法,并提出了一种改进的扩展BPE算法,和Morfessor工具将每个单词分割为子词。我们有效地将最大似然法使用加权有限状态转换器的(ML)和Viterbi估计技术(WFST)框架在这些算法中用于从大型文本语料库中学习子词词典,训练数据转录中的词被分割成子词,并且我们训练深度神经网络ASR系统,该系统识别子词序列,对于泰米尔语ASR,我们使用152个小时的数据用于训练,65个小时用于测试,而对于卡纳达语ASR,我们使用了275小时的训练时间和72小时的测试时间,通过实验不同的切分和估计技术的组合,我们发现单词错误率(WER)与基线单词水平ASR相比大幅降低,泰米尔语和卡纳达语的最大绝对WER降低分别为6.24%和6.63%。
摘要:We present automatic speech recognition (ASR) systems for Tamil and Kannada based on subword modeling to effectively handle unlimited vocabulary due to the highly agglutinative nature of the languages. We explore byte pair encoding (BPE), and proposed a variant of this algorithm named extended-BPE, and Morfessor tool to segment each word as subwords. We have effectively incorporated maximum likelihood (ML) and Viterbi estimation techniques with weighted finite state transducers (WFST) framework in these algorithms to learn the subword dictionary from a large text corpus. Using the learnt subword dictionary, the words in training data transcriptions are segmented to subwords and we train deep neural network ASR systems which recognize subword sequence for any given test speech utterance. The output subword sequence is then post-processed using deterministic rules to get the final word sequence such that the actual number of words that can be recognized is much larger. For Tamil ASR, We use 152 hours of data for training and 65 hours for testing, whereas for Kannada ASR, we use 275 hours for training and 72 hours for testing. Upon experimenting with different combination of segmentation and estimation techniques, we find that the word error rate (WER) reduces drastically when compared to the baseline word-level ASR, achieving a maximum absolute WER reduction of 6.24% and 6.63% for Tamil and Kannada respectively.


eess.AS音频处理

【1】 Knowledge-driven Subword Grammar Modeling for Automatic Speech  Recognition in Tamil and Kannada

标题:泰米尔语和卡纳达语自动语音识别中知识驱动的子词语法建模

链接:https://arxiv.org/abs/2207.13333

* 与cs.SD语音【3】为同一篇

作者:Madhavaraj A,Bharathi Pilar,Ramakrishnan A G
摘要:本文提出了一种新的方法,我们提出了专门设计的自动语音识别(ASR)系统,它可以识别无限词汇量的单词,我们使用子词作为基本词汇单元进行识别,并构造了子词语法加权有限状态转换器(SG-WFST)图形的分词,它捕捉了语言中大多数复杂的构词规则。(i)动词,(ii)名词,(ii)代名词,以及(iv)号码。本文提出了一种启发式分词算法,该算法可以对不符合SG-WFST图规则的异常词进行分词,并利用该算法对SG-WFST图中的大部分数据进行了分词处理.驱动子词词典创建算法是计算驱动的,因此不能保证语素类单元,因此我们使用语言的语言学知识并手动创建子词词典和图。最后,我们训练深度神经网络声学模型,并将其与子词词典的发音词典和SG-WFST图相结合,以构建子词-由于子词-ASR产生子词序列作为给定测试语音的输出,我们对其输出进行后处理以获得最终的词序列,从而实际可识别的词的数量要高得多。在使用IISc-MILE泰米尔语和卡纳达语ASR语料库对子词-ASR系统进行实验时,我们观察到,与泰米尔语和卡纳达语的基线基于单词的ASR系统相比,绝对单词错误率分别减少了12.39%和13.56%。
摘要:In this paper, we present specially designed automatic speech recognition (ASR) systems for the highly agglutinative and inflective languages of Tamil and Kannada that can recognize unlimited vocabulary of words. We use subwords as the basic lexical units for recognition and construct subword grammar weighted finite state transducer (SG-WFST) graphs for word segmentation that captures most of the complex word formation rules of the languages. We have identified the following category of words (i) verbs, (ii) nouns, (ii) pronouns, and (iv) numbers. The prefix, infix and suffix lists of subwords are created for each of these categories and are used to design the SG-WFST graphs. We also present a heuristic segmentation algorithm that can even segment exceptional words that do not follow the rules encapsulated in the SG-WFST graph. Most of the data-driven subword dictionary creation algorithms are computation driven, and hence do not guarantee morpheme-like units and so we have used the linguistic knowledge of the languages and manually created the subword dictionaries and the graphs. Finally, we train a deep neural network acoustic model and combine it with the pronunciation lexicon of the subword dictionary and the SG-WFST graph to build the subword-ASR systems. Since the subword-ASR produces subword sequences as output for a given test speech, we post-process its output to get the final word sequence, so that the actual number of words that can be recognized is much higher. Upon experimenting the subword-ASR system with the IISc-MILE Tamil and Kannada ASR corpora, we observe an absolute word error rate reduction of 12.39% and 13.56% over the baseline word-based ASR systems for Tamil and Kannada, respectively.


【2】 Subword Dictionary Learning and Segmentation Techniques for Automatic  Speech Recognition in Tamil and Kannada

标题:用于泰米尔语和卡纳达语自动语音识别的子词词典学习和分割技术

链接:https://arxiv.org/abs/2207.13331

* 与cs.SD语音【4】为同一篇

作者:Madhavaraj A,Bharathi Pilar,Ramakrishnan A G
摘要:提出了一种自动语音识别方法(ASR)系统,用于泰米尔语和卡纳达语,基于子词建模,以有效地处理无限的词汇量,由于语言的高度粘着性。(BPE)算法,并提出了一种改进的扩展BPE算法,和Morfessor工具将每个单词分割为子词。我们有效地将最大似然法使用加权有限状态转换器的(ML)和Viterbi估计技术(WFST)框架在这些算法中用于从大型文本语料库中学习子词词典,训练数据转录中的词被分割成子词,并且我们训练深度神经网络ASR系统,该系统识别子词序列,对于泰米尔语ASR,我们使用152个小时的数据用于训练,65个小时用于测试,而对于卡纳达语ASR,我们使用了275小时的训练时间和72小时的测试时间,通过实验不同的切分和估计技术的组合,我们发现单词错误率(WER)与基线单词水平ASR相比大幅降低,泰米尔语和卡纳达语的最大绝对WER降低分别为6.24%和6.63%。
摘要:We present automatic speech recognition (ASR) systems for Tamil and Kannada based on subword modeling to effectively handle unlimited vocabulary due to the highly agglutinative nature of the languages. We explore byte pair encoding (BPE), and proposed a variant of this algorithm named extended-BPE, and Morfessor tool to segment each word as subwords. We have effectively incorporated maximum likelihood (ML) and Viterbi estimation techniques with weighted finite state transducers (WFST) framework in these algorithms to learn the subword dictionary from a large text corpus. Using the learnt subword dictionary, the words in training data transcriptions are segmented to subwords and we train deep neural network ASR systems which recognize subword sequence for any given test speech utterance. The output subword sequence is then post-processed using deterministic rules to get the final word sequence such that the actual number of words that can be recognized is much larger. For Tamil ASR, We use 152 hours of data for training and 65 hours for testing, whereas for Kannada ASR, we use 275 hours for training and 72 hours for testing. Upon experimenting with different combination of segmentation and estimation techniques, we find that the word error rate (WER) reduces drastically when compared to the baseline word-level ASR, achieving a maximum absolute WER reduction of 6.24% and 6.63% for Tamil and Kannada respectively.


【3】 End-To-End Audiovisual Feature Fusion for Active Speaker Detection

标题:基于端到端音视频特征融合的主动说话人检测

链接:https://arxiv.org/abs/2207.13434

* 与cs.SD语音【1】为同一篇

作者:Fiseha B. Tesema,Zheyuan Lin,Shiqiang Zhu,Wei Song,Jason Gu,Hong Wu
备注:To appear on the proceeding of the Fourteenth International Conference on Digital Image Processing (ICDIP 2022), May 20-23, Wuhan, China, 8 pages, 3 figures
摘要:主动说话人检测在人机交互中起着至关重要的作用.近年来出现了一些端到端的视听系统框架.然而,这些模型的复杂性和输入数据量大,其推理时间没有得到充分的研究,不适用于实时应用.另外,他们探索了一种类似的特征提取策略,该策略将ConvNet应用于音频和视频输入。该网络将通过VGG-M从图像中提取的特征与从音频波形中提取的原始Mel频率倒谱系数特征进行融合,在AVA-300上的实验结果表明,该模型具有较好的动态响应特性. ActiveSpeaker数据集的实验结果表明,与ConvNet模型相比,本文提出的特征提取方法具有更好的鲁棒性和更快的推理速度,该模型的检测精度达到了88.929%,与现有的检测方法的检测精度基本一致。
摘要:Active speaker detection plays a vital role in human-machine interaction. Recently, a few end-to-end audiovisual frameworks emerged. However, these models' inference time was not explored and are not applicable for real-time applications due to their complexity and large input size. In addition, they explored a similar feature extraction strategy that employs the ConvNet on audio and visual inputs. This work presents a novel two-stream end-to-end framework fusing features extracted from images via VGG-M with raw Mel Frequency Cepstrum Coefficients features extracted from the audio waveform. The network has two BiGRU layers attached to each stream to handle each stream's temporal dynamic before fusion. After fusion, one BiGRU layer is attached to model the joint temporal dynamics. The experiment result on the AVA-ActiveSpeaker dataset indicates that our new feature extraction strategy shows more robustness to noisy signals and better inference time than models that employed ConvNet on both modalities. The proposed model predicts within 44.41 ms, which is fast enough for real-time applications. Our best-performing model attained 88.929% accuracy, nearly the same detection result as state-of-the-art -work.


【4】 Perception-Aware Attack: Creating Adversarial Music via  Reverse-Engineering Human Perception

标题:感知感知攻击:通过逆向工程人类感知创建对抗性音乐

链接:https://arxiv.org/abs/2207.13192

* 与cs.SD语音【2】为同一篇

作者:Rui Duan,Zhe Qu,Shangqing Zhao,Leah Ding,Yao Liu,Zhuo Lu
备注:ACM CCS 2022
摘要:近来,对抗性机器学习攻击已经对包括语音识别、说话人识别先前的研究主要集中在通过在原始信号上创建小的类似噪声的扰动来确保攻击音频信号分类器的有效性。目前还不清楚攻击者是否能够在攻击有效性的基础上产生人类能够很好感知的音频信号扰动。这对于音乐信号尤其重要,因为它们是精心制作的,具有人类喜欢的音频特征。  在本文中,我们将对抗性攻击的概念表述为一种新的感知感知攻击框架,它将人类研究与对抗性攻击的设计相结合.具体来说,我们进行了一项人类研究,以量化人类对音乐信号变化的感知.我们邀请人类参与者基于原始和扰动音乐信号对他们感知的偏差进行评级,然后,通过回归分析对人类感知过程进行逆向工程,以预测给定扰动信号下的人类感知偏差.感知感知感知攻击被表述为一个优化问题,即寻找最优扰动信号,以最小化由回归的人类感知模型预测的感知偏差.我们使用感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知—感知实验结果表明,感知感知感知攻击产生的对抗性音乐的感知质量明显优于已有的攻击。
摘要:Recently, adversarial machine learning attacks have posed serious security threats against practical audio signal classification systems, including speech recognition, speaker recognition, and music copyright detection. Previous studies have mainly focused on ensuring the effectiveness of attacking an audio signal classifier via creating a small noise-like perturbation on the original signal. It is still unclear if an attacker is able to create audio signal perturbations that can be well perceived by human beings in addition to its attack effectiveness. This is particularly important for music signals as they are carefully crafted with human-enjoyable audio characteristics.  In this work, we formulate the adversarial attack against music signals as a new perception-aware attack framework, which integrates human study into adversarial attack design. Specifically, we conduct a human study to quantify the human perception with respect to a change of a music signal. We invite human participants to rate their perceived deviation based on pairs of original and perturbed music signals, and reverse-engineer the human perception process by regression analysis to predict the human-perceived deviation given a perturbed signal. The perception-aware attack is then formulated as an optimization problem that finds an optimal perturbation signal to minimize the prediction of perceived deviation from the regressed human perception model. We use the perception-aware framework to design a realistic adversarial music attack against YouTube's copyright detector. Experiments show that the perception-aware attack produces adversarial music with significantly better perceptual quality than prior work.