今日论文合集:cs.SD语音7篇,eess.AS音频处理5篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】 Balancing physical modeling and musical requirements: Algorithmically  simulating the calls of Hyalessa maculaticollis for real-time instrumental  control

标题:平衡物理建模和音乐要求:数学地模拟Hyalessa maculaticollis的叫声以进行实时乐器控制
链接:https://arxiv.org/abs/2502.09459
作者:Staas de Jong
备注:The 2024 International Sound and Music Computing Conference
摘要:本文提出了一种算法,模拟的声音Hyalessa maculaticollis的音乐用途。用SuperCollider编写,它的输入参数可以实时控制昆虫的叫声相位、响度和感知的音高。为此,物理建模的鼓膜肌肉,鼓膜apodeme,鼓膜肋骨,鼓膜板,腹部气囊,鼓膜,和鳃盖的解剖力学。这也包括退相干,假设它,在H。maculaticollis,可能解释了在呼叫序列的最后阶段明显的音色变化。  总的来说,该算法似乎说明了关于在为音调用途:在设计和实现的许多阶段,可能有必要将音乐要求优先于现实物理建模;所得到的调整可能围绕使物理建模在感知上产生音调良好、单次打击、单源和音色表现力的声音事件;共振体的音高调整模拟在物理上成功时,可能在音乐上精确地失败,这是由于对不同音高引起不同声源的感知。音频的例子包括在内,源代码的结构和文件,以支持进一步发展的生物声学的音乐用途。
摘要:This paper presents an algorithm that simulates the calls of the Hyalessamaculaticollis cicada for musical use. Written in SuperCollider, its inputparameters enable real-time control of the insect call phase, loudness, andperceived musical pitch. To this end, the anatomical mechanics of the tymbalmuscles, tymbal apodeme, tymbal ribs, tymbal plate, abdominal air sac, tympana,and opercula are physically modeled. This also includes decoherence, followingthe hypothesis that it, in H. maculaticollis, might explain the change intimbre apparent during the final phase of a call sequence. Overall, the algorithm seems to illustrate three main points regarding thetrade-offs encountered when modeling bioacoustics for tonal use: that it may benecessary to prioritize musical requirements over realistic physical modelingat many stages of design and implementation; that the resulting adjustments mayrevolve around having physical modeling perceptually yield sonic events thatare well-pitched, single-attack, single-source, and timbrally expressive; thatthe pitch-adjusted simulation of resonating bodies may fail musically preciselywhen it succeeds physically, by inducing the perception of different soundsources for different pitches. Audio examples are included, and the source codeis structured and documented so as to support the further development of cicadabioacoustics for musical use.

【2】 Quantum Approaches for Dysphonia Assessment in Small Speech Datasets
标题:小语音数据集中语音障碍评估的量子方法
链接:https://arxiv.org/abs/2502.08968
作者:Ha Tran,  Bipasha Kashyap,  Pubudu N. Pathirana
摘要:发音困难是一种常见的疾病,会导致声音丧失、声音嘶哑或言语中断。为了评估它,研究人员一直在研究各种机器学习技术以及传统的医学评估。卷积神经网络(CNN)因其在音频分类和语音识别方面的成功而受到欢迎。然而,语音数据的有限可用性对CNN提出了挑战。这项研究评估了CNN对一种新的混合量子经典方法的性能,量子卷积神经网络(QNN),它非常适合于小数据集。对音频数据进行预处理,得到包含243个训练样本和61个测试样本的Mel谱图,并用于10个实验。开发了四个模型(两个QNN和两个CNN),第二个模型包含额外的层以提高性能。结果显示,在大多数实验中,QNN模型在准确性和稳定性方面始终优于CNN模型。
摘要:Dysphonia, a prevalent medical condition, leads to voice loss, hoarseness, orspeech interruptions. To assess it, researchers have been investigating variousmachine learning techniques alongside traditional medical assessments.Convolutional Neural Networks (CNNs) have gained popularity for their successin audio classification and speech recognition. However, the limitedavailability of speech data, poses a challenge for CNNs. This study evaluatesthe performance of CNNs against a novel hybrid quantum-classical approach,Quanvolutional Neural Networks (QNNs), which are well-suited for smalldatasets. The audio data was preprocessed into Mel spectrograms, comprising 243training samples and 61 testing samples in total, and used in ten experiments.Four models were developed (two QNNs and two CNNs) with the second modelsincorporating additional layers to boost performance. The results revealed thatQNN models consistently outperformed CNN models in accuracy and stabilityacross most experiments.

【3】 TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and  Text-to-Instrument
标题:TokenSynth:一种用于仪器克隆和文本到仪器的基于令牌的神经合成器
链接:https://arxiv.org/abs/2502.08939
作者:Kyungsu Kim,  Junghyun Koo,  Sungho Lee,  Haesun Joung,  Kyogu Lee
备注:5 pages, 1 figure, to be published in ICASSP 2025
摘要:神经音频编解码器的最新进展使得能够在各种音频生成任务中使用标记化音频表示,诸如文本到语音、文本到音频和文本到音乐生成。利用这种方法,我们提出了TokenSynth,这是一种新型的神经合成器,它利用仅解码器的Transformer从音频令牌和CLAP(对比音频预训练)嵌入生成所需的音频令牌,其中包含音色相关信息。我们的模型能够执行仪器克隆,文本到仪器的合成,和文本引导的音色操纵没有任何微调。这种灵活性实现了多样化的声音设计和直观的音色控制。我们评估了合成音频的质量、合成音频/文本与目标音频/文本之间的音色相似性以及合成准确性(即,它使用客观测量来跟踪输入的准确性。TokenSynth展示了利用先进的神经音频编解码器和Transformers来创建功能强大且多功能的神经合成器的潜力。源代码、模型权重和音频演示可在以下网址获得:https://github.com/KyungsuKim42/tokensynth
摘要:Recent advancements in neural audio codecs have enabled the use of tokenizedaudio representations in various audio generation tasks, such astext-to-speech, text-to-audio, and text-to-music generation. Leveraging thisapproach, we propose TokenSynth, a novel neural synthesizer that utilizes adecoder-only transformer to generate desired audio tokens from MIDI tokens andCLAP (Contrastive Language-Audio Pretraining) embedding, which hastimbre-related information. Our model is capable of performing instrumentcloning, text-to-instrument synthesis, and text-guided timbre manipulationwithout any fine-tuning. This flexibility enables diverse sound design andintuitive timbre control. We evaluated the quality of the synthesized audio,the timbral similarity between synthesized and target audio/text, and synthesisaccuracy (i.e., how accurately it follows the input MIDI) using objectivemeasures. TokenSynth demonstrates the potential of leveraging advanced neuralaudio codecs and transformers to create powerful and versatile neuralsynthesizers. The source code, model weights, and audio demos are available at:https://github.com/KyungsuKim42/tokensynth

【4】 SpeechCompass: Enhancing Mobile Captioning with Diarization and  Directional Guidance via Multi-Microphone Localization
标题:SpeechCompass:通过多麦克风定位,通过拨号和方向引导增强移动字幕
链接:https://arxiv.org/abs/2502.08848
作者:Artem Dementyev,  Dimitri Kavensky,  Samuel J. Yang,  Mathieu Parvaix,  Chiong Lai,  Alex Olwal
备注:Accepted to CHI 2025
摘要:移动设备上的语音转文本功能已被证明有助于听力和语音无障碍性、语言翻译、笔记和会议记录。然而,我们的基础性大规模调查(n=263)表明,无法区分和指示扬声器的方向,使他们在小组对话的挑战。SpeechCompass通过实时多麦克风语音定位解决了这一限制,其中语音的方向允许视觉分离和引导(例如,箭头)。我们介绍了高效的实时音频定位算法和自定义声音感知硬件上运行的低功耗微控制器和四个集成麦克风,我们在技术评估的特点。通过一项大规模的调查(n=494),我们进行了一项面对面的研究,与8个频繁使用移动语音转文本的用户进行了小组对话,他们提供了五种可视化风格的反馈。日记化和可视化本地化的价值在参与者中是一致的,每个人都同意小组对话方向性指导的价值和潜力。
摘要:Speech-to-text capabilities on mobile devices have proven helpful for hearingand speech accessibility, language translation, note-taking, and meetingtranscripts. However, our foundational large-scale survey (n=263) shows thatthe inability to distinguish and indicate speaker direction makes themchallenging in group conversations. SpeechCompass addresses this limitationthrough real-time, multi-microphone speech localization, where the direction ofspeech allows visual separation and guidance (e.g., arrows) in the userinterface. We introduce efficient real-time audio localization algorithms andcustom sound perception hardware running on a low-power microcontroller andfour integrated microphones, which we characterize in technical evaluations.Informed by a large-scale survey (n=494), we conducted an in-person study ofgroup conversations with eight frequent users of mobile speech-to-text, whoprovided feedback on five visualization styles. The value of diarization andvisualizing localization was consistent across participants, with everyoneagreeing on the value and potential of directional guidance for groupconversations.

【5】 Decision Tree Based Wrappers for Hearing Loss
标题:基于决策树的听力损失包装器
链接:https://arxiv.org/abs/2502.08785
作者:Miguel Rabuge,  Nuno Lourenço
摘要:听力学实体正在使用机器学习(ML)模型来指导他们对风险人群的筛查。特征工程(FE)专注于优化ML模型的数据,进化方法在特征选择和构建任务中是有效的。这项工作的目的是基准的进化FE包装,使用模型的基础上作为代理的决策树。FEDORA框架应用于听力损失(HL)数据集,能够降低数据维度并在统计上保持基线性能。与传统方法相比,FEDORA表现出卓越的性能,使用57个特征的最大平衡准确率为76.2%。该框架还生成了一个使用单一特征实现72.8%平衡准确度的个体。
摘要:Audiology entities are using Machine Learning (ML) models to guide theirscreening towards people at risk. Feature Engineering (FE) focuses onoptimizing data for ML models, with evolutionary methods being effective infeature selection and construction tasks. This work aims to benchmark anevolutionary FE wrapper, using models based on decision trees as proxies. TheFEDORA framework is applied to a Hearing Loss (HL) dataset, being able toreduce data dimensionality and statistically maintain baseline performance.Compared to traditional methods, FEDORA demonstrates superior performance, witha maximum balanced accuracy of 76.2%, using 57 features. The framework alsogenerated an individual that achieved 72.8% balanced accuracy using a singlefeature.

【6】 Are Expressions for Music Emotions the Same Across Cultures?
标题:不同文化中音乐情感的表达方式是否相同?
链接:https://arxiv.org/abs/2502.08744
作者:Elif Celen,  Pol van Rijn,  Harin Lee,  Nori Jacoby
备注:Submitted to CogSci
摘要:音乐唤起深刻的情感,但情感描述词在语言中的普遍性仍然存在争议。音乐情感跨文化研究的一个关键挑战是有偏见的刺激选择和分类的人工管理,主要依赖于西方音乐和语言。为了解决这个问题,我们提出了一个平衡的实验设计,在巴西,美国和韩国的9个在线实验,涉及N=672名参与者。首先,我们从这些国家的流行音乐中选取了一组平衡的样本。使用一个开放式的标记管道,我们收集情感术语来创建特定于文化的分类。最后,使用这些自下而上的分类法,参与者对每首歌的情感进行评级。这使我们能够在文化内部和文化之间绘制情感相似性。结果表明,一致性高唤醒,高效价的情绪,但在其他更大的变化。值得注意的是,机器翻译往往不足以捕捉音乐的特定含义。这些发现共同强调了需要一个领域敏感的,开放式的,自下而上的情绪诱导方法,以减少情绪研究中的文化偏见。
摘要:Music evokes profound emotions, yet the universality of emotional descriptorsacross languages remains debated. A key challenge in cross-cultural research onmusic emotion is biased stimulus selection and manual curation of taxonomies,predominantly relying on Western music and languages. To address this, wepropose a balanced experimental design with nine online experiments in Brazil,the US, and South Korea, involving N=672 participants. First, we sample abalanced set of popular music from these countries. Using an open-ended taggingpipeline, we then gather emotion terms to create culture-specific taxonomies.Finally, using these bottom-up taxonomies, participants rate emotions of eachsong. This allows us to map emotional similarities within and across cultures.Results show consistency in high arousal, high valence emotions but greatervariability in others. Notably, machine translations were often inadequate tocapture music-specific meanings. These findings together highlight the need fora domain-sensitive, open-ended, bottom-up emotion elicitation approach toreduce cultural biases in emotion research.

【7】 Enhanced LSTM by Attention Mechanism for Early Detection of Parkinson's  Disease through Voice Signals
标题:注意力机制增强LSTM,通过语音信号早期检测帕金森病
链接:https://arxiv.org/abs/2502.08672
作者:Arman Mohammadigilani,  Hani Attar,  Hamidreza Ehsani Chimeh,  Mostafa Karami
摘要:帕金森病(PD)是一种神经退行性疾病,其特征在于显著的运动和非运动表现。被称为统一帕金森病评定量表(Unified Parkinson's Disease Rating Scale,简称PDRS)的评估工具在评估与帕金森病(PD)相关的帕金森病程度方面发挥着至关重要的作用。这项研究提出了一种使用复杂的长短期记忆(LSTM)网络来预测BPRS分数的完整方法,该网络使用注意力机制,数据增强技术和鲁棒的特征选择来改进。这项工作中使用的数据来自UC Irvine机器学习存储库。它包括从帕金森病早期患者收集的一系列语音指标。递归特征消除(RFE)被用来实现有效的特征选择,而抖动的应用程序增强了数据集。长短期记忆(LSTM)网络经过精心设计,可以有效地捕捉数据集中的时间波动。此外,它还通过集成注意力机制来增强,这增强了网络识别序列重要性的能力。已经描述的方法提出了一种潜在的实用方法,用于对与帕金森病相关的医学数据进行更精确和个性化的分析。
摘要:Parkinson's disease (PD) is a neurodegenerative condition characterized bynotable motor and non-motor manifestations. The assessment tool known as theUnified Parkinson's Disease Rating Scale (UPDRS) plays a crucial role inevaluating the extent of symptomatology associated with Parkinson's Disease(PD). This research presents a complete approach for predicting UPDRS scoresusing sophisticated Long Short-Term Memory (LSTM) networks that are improvedusing attention mechanisms, data augmentation techniques, and robust featureselection. The data utilized in this work was obtained from the UC IrvineMachine Learning repository. It encompasses a range of speech metrics collectedfrom patients in the early stages of Parkinson's disease. Recursive FeatureElimination (RFE) was utilized to achieve efficient feature selection, whilethe application of jittering enhanced the dataset. The Long Short-Term Memory(LSTM) network was carefully crafted to capture temporal fluctuations withinthe dataset effectively. Additionally, it was enhanced by integrating anattention mechanism, which enhances the network's ability to recognize sequenceimportance. The methodology that has been described presents a potentiallypractical approach for conducting a more precise and individualized analysis ofmedical data related to Parkinson's disease.

eess.AS音频处理

【1】 Advances in Microphone Array Processing and Multichannel Speech  Enhancement
标题:麦克风阵列处理和多通道语音增强的进展
链接:https://arxiv.org/abs/2502.09037
作者:Gongping Huang,  Jesper R. Jensen,  Jingdong Chen,  Jacob Benesty,  Mads G. Christensen,  Akihiko Sugiyama,  Gary Elko,  Tomas Gaensler
备注:accepted by ICASSP 2025
摘要:本文回顾了麦克风阵列处理和多通道语音增强方面的开创性工作,重点介绍了历史成就、技术发展、商业化方面以及关键挑战。它为这些领域的进展和未来方向提供了宝贵的见解。本文探讨了麦克风阵列设计和优化的基础发展,展示了在嘈杂和混响环境中改善声音采集和增强语音清晰度的创新。然后介绍了该领域的最新进展和前沿研究,特别是深度学习技术的集成,如全神经波束形成器。本文还探讨了关键应用程序,讨论了它们的演变和当前最先进的技术,这些技术对用户体验有着重大影响。最后,本文概述了未来的研究方向,确定了可能推动这些领域进一步创新的挑战和潜在解决方案。本文旨在通过提供全面的概述和前瞻性的观点,激励正在进行的研究,并有助于麦克风阵列和多通道语音增强的持续增长和发展。
摘要:This paper reviews pioneering works in microphone array processing andmultichannel speech enhancement, highlighting historical achievements,technological evolution, commercialization aspects, and key challenges. Itprovides valuable insights into the progression and future direction of theseareas. The paper examines foundational developments in microphone array designand optimization, showcasing innovations that improved sound acquisition andenhanced speech intelligibility in noisy and reverberant environments. It thenintroduces recent advancements and cutting-edge research in the field,particularly the integration of deep learning techniques such as all-neuralbeamformers. The paper also explores critical applications, discussing theirevolution and current state-of-the-art technologies that significantly impactuser experience. Finally, the paper outlines future research directions,identifying challenges and potential solutions that could drive furtherinnovation in these fields. By providing a comprehensive overview andforward-looking perspective, this paper aims to inspire ongoing research andcontribute to the sustained growth and development of microphone arrays andmultichannel speech enhancement.

【2】 Predicting Cognitive Decline: A Multimodal AI Approach to Dementia  Screening from Speech
标题:预测认知衰退:通过言语筛查痴呆症的多模式人工智能方法
链接:https://arxiv.org/abs/2502.08862
作者:Lei Chi,  Arav Sharma,  Ari Gebhardt,  Joseph T. Colonel
备注:Submitted to IEEE ICAD 2025
摘要:最近的进展是完全通过记录患者的语音来检测早期痴呆症。多模态语音分析方法应用于PROCESS挑战,该挑战要求参与者使用临床访谈的录音来预测患者为健康对照、轻度认知障碍(MCI)或痴呆,并回归患者的简易精神状态检查(MMSE)评分。这项工作中实现的方法将声学特征(eGeMAPS和Prosody)与Whisper和RoBERTa模型的嵌入相结合,在回归(RMSE:2.7666)和分类(Macro-F1得分:0.5774)任务中都取得了有竞争力的结果。此外,一种新的两层分类设置被用来更好地区分MCI和痴呆。我们的方法在测试集上取得了很好的结果,在37个团队中,回归排名第七,分类排名第十一,超过了基线结果。
摘要:Recent progress has been made in detecting early stage dementia entirelythrough recordings of patient speech. Multimodal speech analysis methods wereapplied to the PROCESS challenge, which requires participants to use audiorecordings of clinical interviews to predict patients as healthy control, mildcognitive impairment (MCI), or dementia and regress the patient's Mini-MentalState Exam (MMSE) scores. The approach implemented in this work combinesacoustic features (eGeMAPS and Prosody) with embeddings from Whisper andRoBERTa models, achieving competitive results in both regression (RMSE: 2.7666)and classification (Macro-F1 score: 0.5774) tasks. Additionally, a noveltwo-tiered classification setup is utilized to better differentiate between MCIand dementia. Our approach achieved strong results on the test set, rankingseventh on regression and eleventh on classification out of thirty-seven teams,exceeding the baseline results.

【3】 ASVspoof 5: Design, Collection and Validation of Resources for Spoofing,  Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech
标题:ASVspoof 5:使用众包语音设计、收集和验证用于欺骗、Deepfake和对抗性攻击检测的资源
链接:https://arxiv.org/abs/2502.08857
作者:Xin Wang,  Héctor Delgado,  Hemlata Tak,  Jee-weon Jung,  Hye-jin Shim,  Massimiliano Todisco,  Ivan Kukanov,  Xuechen Liu,  Md Sahidullah,  Tomi Kinnunen,  Nicholas Evans,  Kong Aik Lee,  Junichi Yamagishi,  Myeonghun Jeong,  Ge Zhu,  Yongyi Zang,  You Zhang,  Soumi Maiti,  Florian Lux,  Nicolas Müller,  Wangyou Zhang,  Chengzhe Sun,  Shuwei Hou,  Siwei Lyu,  Sébastien Le Maguer,  Cheng Gong,  Hanjie Guo,  Liping Chen,  Vishwanath Singh
备注:Database link: this https URL, Database mirror link: this https URL, ASVspoof 5 Challenge Workshop Proceeding: this https URL
摘要:ASVspoof 5是一系列挑战中的第五版,这些挑战促进了对语音欺骗和深度伪造攻击的研究以及检测解决方案的设计。我们介绍了ASVspoof 5数据库,该数据库是以众包的方式从在不同声学条件下收集的数据中生成的(参见早期ASVspoof数据库的录音室质量数据)和约2,000名扬声器(参见~100更早)。该数据库包含由32种不同算法生成的攻击,也是众包的,并使用新的代理检测模型进行了不同程度的优化。其中包括使用传统和现代文本到语音合成和语音转换模型的混合生成的攻击,以及首次合并的对抗性攻击。ASVspoof 5协议包括七个说话者不相交的分区。它们包括两个不同的分区,用于训练不同的攻击模型集,另外两个用于开发和评估代理检测模型,然后是三个额外的分区,包括ASVspoof 5训练,开发和评估集。从另外的30 k扬声器收集的辅助数据集也可以用于训练扬声器编码器以实现攻击算法。本文还描述了使用一组自动说话者验证和欺骗/深度伪造基线检测器对新ASVspoof 5数据库的实验验证。除了用于生成欺骗/deepfake语音的协议和工具之外,本文中描述的资源已经被2024年ASVspoof 5挑战赛的参与者使用,现在都可以免费提供给社区。
摘要:ASVspoof 5 is the fifth edition in a series of challenges which promote thestudy of speech spoofing and deepfake attacks as well as the design ofdetection solutions. We introduce the ASVspoof 5 database which is generated incrowdsourced fashion from data collected in diverse acoustic conditions (cf.studio-quality data for earlier ASVspoof databases) and from ~2,000 speakers(cf. ~100 earlier). The database contains attacks generated with 32 differentalgorithms, also crowdsourced, and optimised to varying degrees using newsurrogate detection models. Among them are attacks generated with a mix oflegacy and contemporary text-to-speech synthesis and voice conversion models,in addition to adversarial attacks which are incorporated for the first time.ASVspoof 5 protocols comprise seven speaker-disjoint partitions. They includetwo distinct partitions for the training of different sets of attack models,two more for the development and evaluation of surrogate detection models, andthen three additional partitions which comprise the ASVspoof 5 training,development and evaluation sets. An auxiliary set of data collected from anadditional 30k speakers can also be used to train speaker encoders for theimplementation of attack algorithms. Also described herein is an experimentalvalidation of the new ASVspoof 5 database using a set of automatic speakerverification and spoof/deepfake baseline detectors. With the exception ofprotocols and tools for the generation of spoofed/deepfake speech, theresources described in this paper, already used by participants of the ASVspoof5 challenge in 2024, are now all freely available to the community.

【4】 Are Expressions for Music Emotions the Same Across Cultures?
标题:不同文化中音乐情感的表达方式是否相同?
链接:https://arxiv.org/abs/2502.08744
作者:Elif Celen,  Pol van Rijn,  Harin Lee,  Nori Jacoby
备注:Submitted to CogSci
摘要:音乐唤起深刻的情感,但情感描述词在语言中的普遍性仍然存在争议。音乐情感跨文化研究的一个关键挑战是有偏见的刺激选择和分类法的手动整理,主要依赖于西方音乐和语言。为了解决这个问题,我们提出了一个平衡的实验设计,在巴西,美国和韩国的9个在线实验,涉及N=672名参与者。首先,我们从这些国家的流行音乐中选取了一组平衡的样本。使用一个开放式的标记管道,我们收集情感术语来创建特定于文化的分类。最后,使用这些自下而上的分类法,参与者对每首歌的情感进行评级。这使我们能够在文化内部和文化之间绘制情感相似性。结果表明,一致性高唤醒,高效价的情绪,但在其他更大的变化。值得注意的是,机器翻译往往不足以捕捉音乐的特定含义。这些发现共同强调了需要一个领域敏感的,开放式的,自下而上的情绪诱导方法,以减少情绪研究中的文化偏见。
摘要:Music evokes profound emotions, yet the universality of emotional descriptorsacross languages remains debated. A key challenge in cross-cultural research onmusic emotion is biased stimulus selection and manual curation of taxonomies,predominantly relying on Western music and languages. To address this, wepropose a balanced experimental design with nine online experiments in Brazil,the US, and South Korea, involving N=672 participants. First, we sample abalanced set of popular music from these countries. Using an open-ended taggingpipeline, we then gather emotion terms to create culture-specific taxonomies.Finally, using these bottom-up taxonomies, participants rate emotions of eachsong. This allows us to map emotional similarities within and across cultures.Results show consistency in high arousal, high valence emotions but greatervariability in others. Notably, machine translations were often inadequate tocapture music-specific meanings. These findings together highlight the need fora domain-sensitive, open-ended, bottom-up emotion elicitation approach toreduce cultural biases in emotion research.

【5】 Visual-based spatial audio generation system for multi-speaker  environments
标题:用于多扬声器环境的基于视觉的空间音频生成系统
链接:https://arxiv.org/abs/2502.07538
作者:Xiaojing Liu,  Ogulcan Gurelli,  Yan Wang,  Joshua Reiss
摘要:在电影和视频游戏等多媒体应用中,空间音频技术被广泛用于通过模拟3D声音来增强用户体验:将单声道音频转换为双耳格式。然而,这个过程对于声音设计师来说通常是复杂和劳动密集型的,需要音频与视觉组件的空间位置精确同步。为了解决这些挑战,我们提出了一个基于视觉的空间音频生成系统-一个自动化系统,集成了人脸检测YOLOv 8的对象检测,单目深度估计和空间音频技术。值得注意的是,该系统在不需要额外的双耳数据集训练的情况下操作。所提出的系统进行评估,对现有的空间音频生成系统,使用客观的指标。实验结果表明,我们的方法显着提高音频和视频之间的空间一致性,提高语音质量,并在多说话人的情况下表现出鲁棒性。通过简化视听对齐过程,该系统使音响工程师能够有效地实现高质量的结果,使其成为多媒体制作专业人员的宝贵工具。
摘要:In multimedia applications such as films and video games, spatial audiotechniques are widely employed to enhance user experiences by simulating 3Dsound: transforming mono audio into binaural formats. However, this process isoften complex and labor-intensive for sound designers, requiring precisesynchronization of audio with the spatial positions of visual components. Toaddress these challenges, we propose a visual-based spatial audio generationsystem - an automated system that integrates face detection YOLOv8 for objectdetection, monocular depth estimation, and spatial audio techniques. Notably,the system operates without requiring additional binaural dataset training. Theproposed system is evaluated against existing Spatial Audio generation systemusing objective metrics. Experimental results demonstrate that our methodsignificantly improves spatial consistency between audio and video, enhancesspeech quality, and performs robustly in multi-speaker scenarios. Bystreamlining the audio-visual alignment process, the proposed system enablessound engineers to achieve high-quality results efficiently, making it avaluable tool for professionals in multimedia production.

机器翻译由腾讯交互翻译提供,仅供参考