今日论文合集:CS.SD语音与音频 | 共 15 篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily

[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准


快速导航

1. 语音识别与关键词检测 3 篇

2. 语音合成与声音生成 2 篇

3. 说话人识别、验证与分离 1 篇

4. 语音翻译与语音语言模型 3 篇

5. 数据集、基准与评测 1 篇

6. 其他/综合语音音频 5 篇


1. 语音识别与关键词检测 | 3 篇

1. CLEAR: Online Speech Content Leakage Estimation through Cross-ASR Disagreement

CLEAR:通过跨ASR不一致性进行在线语音内容泄露估计

AI 总结:CLEAR提出一种无参考的在线语音内容泄露估计方法,利用异构ASR系统的不一致性在运行时评估隐私保护效果,实现动态调整隐私强度并保留声学效用。

链接:https://arxiv.org/abs/2609.30415

机构:University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

作者:Bhawana Chhaglani, Tanvi Kandepuneni, Jeremy Gummeson, Prashant Shenoy

英文摘要:Signal-level speech privacy mechanisms suppress linguistic content while preserving acoustic information needed by downstream sensing applications. However, their privacy settings are typically evaluated/selected offline and remain fixed during deployment, even though speech-content leakage can vary substantially across utterances and speakers. Adapting privacy protection at runtime requires estimating how much speech remains recoverable, but conventional measures such as WER or PER require ground-truth transcripts and therefore cannot be computed online. We present CLEAR, a reference-free approach for estimating speech-content leakage at runtime using disagreement among heterogeneous ASR systems. Our key insight is that independently trained ASRs exhibit consistent behavior when linguistic content remains recoverable and increasingly disagree as privacy transformations obscure speech. Using configurable speech-suppression mechanism, we show that cross-ASR disagreement closely tracks transcript-grounded leakage across privacy operating points, achieving a correlation of 0.8. We further characterize the latency-accuracy trade-off of heterogeneous ASR subsets and use their hypotheses to identify potentially exposed words. These capabilities enable privacy to be treated as a runtime property rather than a fixed configuration: CLEAR can communicate residual speech exposure to users and provide feedback for dynamically adjusting privacy aggressiveness while retaining acoustic utility for downstream sensing tasks.


2. Training-Free Contextual ASR via SpeechLLM-Based Error-Aware Selective Retrieval

基于SpeechLLM的错误感知选择性检索的无训练上下文语音识别

AI 总结:提出无训练上下文ASR框架,利用SpeechLLM定位错误跨度并选择性检索术语,减少词典查询并提升多领域识别性能。

链接:https://arxiv.org/abs/2609.30694

机构:Hitachi, Ltd., Research & Development Group(株式会社日立制作所研发集团); Hitachi Advanced Systems Corporation(日立先进系统株式会社)

作者:Natsuo Yamashita, Ai Nemoto, Ryosuke Koichi, Masaaki Yamamoto

英文摘要:Recognition of domain-specific and low-frequency terms remains challenging for automatic speech recognition (ASR). Although contextual biasing can improve their recognition, directly providing a large terminology dictionary introduces many irrelevant biasing terms. Retrieval-based contextual biasing addresses this issue by selecting candidate terms from an external dictionary, but querying many recognized words requires numerous dictionary lookups and may yield poorly targeted candidates. We propose a training-free contextual ASR framework in which a pretrained speech large language model (SpeechLLM) jointly generates an ASR hypothesis and localizes error spans likely to involve domain-specific terms. Only the localized spans are used to retrieve phonologically similar terms from an external terminology dictionary. The same SpeechLLM then re-recognizes the audio conditioned on the first-pass hypothesis and the retrieved terms, without task-specific model training. To assess applicability across domains, we evaluate the framework on medical, air traffic control, and financial speech. The proposed method substantially reduces dictionary queries while improving the recall and ranking of relevant terminology candidates and second-pass ASR performance across all three domains.


3. BreathGRU: A Novel Semi-Supervised Bidirectional Gated Recurrent Unit Framework for Speech and Breath Segmentation for Respiratory Audio

BreathGRU:一种用于呼吸音频中语音与呼吸分割的新型半监督双向门控循环单元框架

AI 总结:针对呼吸音频分析中语音-呼吸分割的不足,提出半监督双向门控循环单元框架BreathGRU,结合特征提取、双向建模、伪标签精炼和维特比解码,在多项指标上优于现有方法,实现高效分割。

链接:https://arxiv.org/abs/2609.31165

作者:Sania Fatima Sayed, John W. Holloway, Reyer Zwiggelaar, Faisal I. Rezwan

英文摘要:Speech-breath segmentation is a fundamental preprocessing step in respiratory audio analysis, enabling applications such as respiratory acoustic biomarker extraction, lung function prediction and disease monitoring. Existing approaches, including threshold methods, Fourier Transform-based techniques, and unsupervised and pretrained voice activity detection (VAD) models, primarily focus on speech detection and often classify breathing events as non-speech or silence, limiting their applicability for precise breath detection. To address this limitation, we propose BreathGRU, a semi-supervised Bidirectional Gated Recurrent Unit (BiGRU) framework specifically designed for speech-breath segmentation. The proposed framework combines frame-level acoustic feature extraction with bidirectional recurrent modelling, pseudo-label refinement and duration-constrained Segmental Viterbi decoding to produce speech and breath segmentation. BreathGRU was evaluated against the existing approaches, using manually annotated recordings. Performance was assessed using event-based, time-based, overlap-based, duration-based and boundary-based segmentation metrics. Experiment results demonstrated that BreathGRU achieved the highest breath event recall (0.83), the lowest onset-localisation error (0.14s) and the highest Mean Match Intersection over Union (0.81), with competitive overall segmentation performance compared to large pretrained VAD models like Silero. Qualitative evaluation on manually annotated recordings further showed close agreement between BreathGRU and manual annotation, with better breath detection compared to Silero. These findings demonstrate that explicit breath event modelling provides advantages over general-purpose VAD models and establish BreathGRU as an effective speech-breath segmentation framework which can be applied for respiratory audio analysis and pulmonary healthcare applications.


2. 语音合成与声音生成 | 2 篇

4. Tracing and Relearning Detection Evidence in Text-to-Speech Systems

文本到语音系统中的检测证据追踪与再学习

AI 总结:本研究通过F5-TTS-BigVGAN流水线追踪检测证据来源,发现声学模型更新可减少固定检测器的检测证据,而检测器适应能恢复检测能力,显著降低EER。

链接:https://arxiv.org/abs/2609.30983

机构:Ewha W. University(梨花女子大学); KAIST AI(韩国科学技术院人工智能学院)

作者:Eunji Shin, Kyudan Jung, Jihwan Kim, Minwoo Lee, Jaegul Choo

英文摘要:Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-BigVGAN pipeline. Since vocoder reconstruction of a real mel can itself be separable from the source utterance, we fix the vocoder and trace the larger change in detector separation to acoustic generation. Adversarially fine-tuning the acoustic model, with no detector in its objective, raises EER against fixed detectors at comparable quality. However, adapting a detector only on the tuned model's VCTK outputs lowers its LibriSpeech EER from 19.42% to 7.46% and improves detection of unseen base F5-TTS outputs. These results suggest that acoustic-model updates can reduce the detection evidence available to fixed detectors, while detector adaptation keeps the updated outputs detectable in this pipeline.


5. TinyAudio: Compact and Efficient Text-to-Audio Generation for Low-Resource Deployment

TinyAudio:面向低资源部署的紧凑高效文本到音频生成

AI 总结:本文提出TinyAudio,一种仅87M参数的紧凑流匹配文本到音频模型,通过TA-DiT、TA-CLAP和TA-VAE实现低资源部署,在AudioCaps和TTA-Bench上达到竞争性质量,并支持CPU实时生成。

链接:https://arxiv.org/abs/2609.31525

机构:X-LANCE Lab, Shanghai Jiao Tong University(上海交通大学X-LANCE实验室); Shanghai Innovation Institute(上海创新研究院); SJTU Paris Elite Institute of Technology, Shanghai Jiao Tong University(上海交通大学巴黎卓越工程师学院); Nanyang Technological University(南洋理工大学)

作者:Junxi Liu, Xiquan Li, Wenhao Guan, Yifan Duan, Zhikang Niu, Yanru Huo, Ziyang Ma, Xie Chen

英文摘要:Text-to-audio (TTA) generation has advanced rapidly in generation quality and instruction following. However, representative systems often require around a billion parameters, limiting deployment on resource-constrained devices. This paper introduces TinyAudio, a compact flow-matching-based TTA model for low-resource deployment. At its core, TinyAudio uses TA-DiT, a 35M single-stream flow-matching Transformer. TinyAudio also includes TA-CLAP, a 32M audio-aligned text encoder, and TA-VAE, whose 20M decoder reconstructs 44.1 kHz audio from compressed latents. TinyAudio has only 87M parameters in total, over 90% fewer than representative billion-parameter pipelines, and uses 0.48 GB peak GPU memory. TinyAudio achieves competitive generation quality on AudioCaps and TTA-Bench. We further introduce TinyAudio-MF, a MeanFlow-accelerated model that enables real-time generation with a four-core CPU quota. Our results demonstrate a practical quality-footprint trade-off for low-resource TTA deployment.


3. 说话人识别、验证与分离 | 1 篇

6. Dialogue-Based Streaming Audio-Visual Target Speaker Extraction with Predictive Dialogue Information

基于预测性对话信息的流式音视频目标说话人提取

AI 总结:针对真实对话中轮流发言机制被忽视的问题,首次提出在线音视频目标说话人提取基准,并设计基于大语言模型的目标说话人语音活动投影模块,从重叠混合语音预测未来活动以引导低延迟分离器,结合历史、同步与预测性上下文,在多种骨干网络上持续提升流式提取性能,真实对话中获近1分贝增益。

链接:https://arxiv.org/abs/2609.30774

机构:Shenzhen Loop Area Institute(深圳河套学院); School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)人工智能学院); The Chinese University of Hong Kong(香港中文大学)

作者:Shuhan Zhang, Wenxuan Wu, Haizhou Li

英文摘要:In face-to-face, real-time communication, a talking agent must track the target speaker through natural pauses, turn-taking, and backchannels, often amid background cross-talk. Most target speaker extraction (TSE) studies, however, rely on simulated mixtures with full or sparse overlap and ignore the turn-taking of real conversations. We therefore introduce, to our knowledge, the first benchmark for online audio-visual TSE (AV-TSE), built from intact dyadic interactions with independent third-party interference. Observing that anticipating upcoming activity from semantic, acoustic, and facial cues benefits online AV-TSE, we propose an LLM-based target-speaker voice activity projection (TS-VAP) module. Unlike conventional VAP with separated speaker channels, it forecasts the future activity of the target and conversational partner directly from the overlapping mixture, drawing on the linguistic and conversational knowledge of a speech-LLM, and uses this prediction to guide a low-latency separator. We further combine this predictive context with historical and synchronous speaker context. Experiments show that TS-VAP consistently improves streaming extraction across multiple AV-TSE backbones, and that further combining historical, synchronous, and predictive context yields nearly 1 dB gain on real AV conversations. Project page: this https URL.


4. 语音翻译与语音语言模型 | 3 篇

7. Don't CLAP: Are Music-Text Models Bag-of-Words?

不要 CLAP:音乐-文本模型是词袋模型吗?

AI 总结:本研究通过属性交换扰动测试,发现CLAP等音乐-文本模型无法区分改变含义的标题扰动,其表示接近词袋,缺乏细粒度音乐语义绑定能力。

链接:https://arxiv.org/abs/2609.30540

机构:Georgia Institute of Technology(佐治亚理工学院)

作者:Yuan-Chiao Cheng, Alexander Lerch

英文摘要:Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective metric of faithfulness. We ask how accurately that score reflects the text: when an attribute is linked to an instrument (e.g., distorted guitar), does the text embedding capture that binding? To find out, we introduce an attribute swap perturbation: the caption of a real recording is edited by exchanging exactly one property, timbre, lead versus accompaniment, or order of first appearance, between two instruments. We then test four contrastive music-text models and one large audio-language model on whether the audio scores higher against the original caption than against the perturbed one. No contrastive model distinguishes the two captions reliably. The audio-language model does better, but further experiments show that its advantage rests largely on audio-agnostic language priors. Our results thus provide compelling evidence that the CLAP score and related metrics do not capture fine-grained musical meaning or attribute bindings; their representation is closer to a bag-of-words that leaves them insensitive to meaning-changing perturbations of the caption.


8. Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models

共生架构:冻结语言模型的后期音频扩展

AI 总结:提出一种共生架构,通过注入器向冻结LLM的KV缓存写入音频向量,实现无需微调的音频理解,兼具可扩展性与性能,接近微调ALM且保留文本能力。

链接:https://arxiv.org/abs/2609.30784

机构:Sakana AI

作者:Yotaro Kubo, Qi Sun, Yujin Tang

英文摘要:This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value (KV) cache, enabling the LLM to behave as an audio language model (ALM). The architectural advantages are twofold. First, it improves the scalability of ALMs: because the proposed method bypasses the LLM during audio injection, the injection cost is governed by the injector width rather than the backbone width, and can therefore scale more slowly than the cost of full-backbone prefilling. Second, since the training scheme does not update the LLM weights, the original capabilities of the LLM are preserved without the risk of degradation from fine-tuning. The effectiveness of the proposed method is evaluated on both audio-understanding tasks (automatic speech recognition, audio question answering, and acoustic scene classification) and text-only tasks. We confirm that, while activating fewer parameters during audio prefilling, our architecture outperforms the conventional method with a frozen LLM and approaches the performance of a fine-tuned ALM, all while preserving the backbone LLM's original text-only task performance by construction.


9. Acoustic-to-Text KV Compression for Full-Duplex Speech Models

全双工语音模型的声学到文本KV压缩

AI 总结:针对全双工语音模型内存密集问题,提出声学到文本KV压缩,利用聆听空隙将语音转文本并驱逐旧声学状态,减少64.6%峰值缓存,提升转录与问答性能。

链接:https://arxiv.org/abs/2609.31224

机构:Sungkyunkwan University(成均馆大学)

作者:Yejin Lee, Seungbeom Kim, Yongha Lee, Kyuhong Shim

英文摘要:Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interval listening-time slack. We propose acoustic-to-text KV compression, which introduces a transcription side channel to convert incoming speech into compact textual memory within this interval. When the cache exceeds a target budget during inference, older acoustic states are evicted while transcripts and recent acoustic context remain. We train the side channel with LoRA using cross-entropy on transcription segments. To preserve listening and speaking behavior, we apply knowledge distillation to the original model's token-level output distributions at native prediction positions. On ten-minute LongSpeech sessions, our MiniCPM-o 4.5 implementation reduces peak streaming KV-cache size by 64.6% compared with the same model without eviction. The proposed method also improves transcription, temporal question answering, and summarization over the baseline. Full-Duplex-Bench evaluations further show comparable pause-handling, turn-taking, and interruption performance.


5. 数据集、基准与评测 | 1 篇

10. AcoustiClaim: A Numeric Claim Benchmark with Instrument Ground Truth

AcoustiClaim:具有仪器真值的数值声明基准

AI 总结:AcoustiClaim提出数值声明基准,用仪器真值评分音频模型输出,弃权策略降低错误率,但部分量仍难准确预测。

链接:https://arxiv.org/abs/2609.30483

机构:Carnegie Mellon University(卡内基梅隆大学)

作者:Sheng-Tse Lin, Siyuan Zhai, Chien-Liang Kuo, Massa Baali, Bhiksha Raj

英文摘要:Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines the quantity, and classes each quantity by where its reference can be read. Four open-weight systems and one closed model, asked for ten quantities five ways on two corpora, fill 207 cells. Of these, 49 emit fewer than five distinct values, and eight of the 158 cells that can be ranked exceed a rank correlation of 0.3, the bar we set, three with an interval clear of it, five of them one closed model reading pitch. Error sits at or above a constant-predictor floor in every ranked cell but three. The reference decoder we train declines the five voice quantities in prose on 95% of mixtures, with nothing withheld, and states them on the clean twins, reproducing its targets' rule from audio alone. With a calibrated threshold, withholding lowers error on all ten quantities on the mixtures in the mean and on eight at every split, against at most 0.6% from a random selector. A linear baseline orders errors at least as well as ours. F0 s.d. and shimmer stay above the constant floor.


6. 其他/综合语音音频 | 5 篇

11. MuseTimbre: Zero-Shot Timbre Transfer by Controlling a Frozen Music Generator

MuseTimbre:通过控制冻结音乐生成器实现零样本音色迁移

AI 总结:MuseTimbre是首个通过调节预训练音乐生成器,从音频参考向复音声源迁移音色的零样本系统,利用多音高估计和微调CLAP编码器,在四个数据集上实现与基线相当的音高对齐和更紧密的音色匹配。

链接:https://arxiv.org/abs/2609.30548

机构:Georgia Institute of Technology(佐治亚理工学院); University of Rochester(罗切斯特大学)

作者:Yuan-Chiao Cheng, Zhiyao Duan

英文摘要:Instrument timbre transfer re-voices a performance using the timbre of another instrument. Extracting the target timbre from an audio reference capture more nuances than inferring it from a text prompt. Systems that read timbre from such a clip train a dedicated model for the task, which captures the timbre cleanly but stays a narrow, single-purpose system. More versatile approaches add control to a pretrained music generator, yet a reference clip entangles timbre with genre and melody, so these systems fall back on text to name the timbre. We present MuseTimbre, the first system, to our best knowledge, that transfers timbre from an audio reference to a polyphonic source through conditioning a pretrained music generator. This system employs a multi-pitch estimator to extract pitch information from the source and finetune a CLAP encoder to extract timbre information from the reference audio. Experiments show that across four datasets of real polyphonic recordings, MuseTimbre achieves pitch alignment on par with the baselines while matching the reference timbre far more closely. Results also show that the finetuned CLAP-based timbre extractor is robust to pitch variations, making it useful in timbre similarity measures.


12. Subject-Invariant Cross-Modal Decoding of Perceived Speech from Brain Recordings

基于脑记录感知语音的跨模态主题不变解码

AI 总结:本研究提出SICMD方法,融合fMRI与MEG信号实现跨受试者感知语音解码,显著提升Top-1、Top-10和Rankacc指标,并大幅降低训练成本。

链接:https://arxiv.org/abs/2609.30832

机构:School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院); Academy for Advanced Interdisciplinary Studies, Peking University(北京大学前沿交叉学科研究院); National Key Laboratory of General Artificial Intelligence(通用人工智能全国重点实验室)

作者:Aoke Zhang, Jing Chen

英文摘要:Perceived speech decoding based on non-invasive brain-computer interface (BCI) signals has been extensively studied in recent years. Research in this field primarily faces two challenges: extracting neural representations with rich spatiotemporal information and achieving cross-subject generalization. Although separate studies have proposed methods to cope with these issues, a unified approach that simultaneously tackles both challenges remains lacking. To fill this gap, we propose the Subject-Invariant Cross-Modal Perceived Speech Decoding (SICMD) method, which integrates functional magnetic resonance imaging (fMRI) and magnetoencephalography (MEG). We conduct comprehensive analyses of the fusion method, fusion position, encoder architecture, and model inputs. Our results demonstrate that the proposed method improves Top-1, Top-10, and Rankacc by more than 10.6%, 10.1%, and 1.7%, respectively, compared to baseline methods in cross-subject perceived speech decoding tasks, while reducing training costs by 88.8% and 60.5% compared to multi-subject and intra-subject decoding settings. Further visualization experiments also confirm the effectiveness of our approach.


13. Synth-JEPA: Joint Embedding Prediction for Renderer-Free Synthesizer Parameter Search

Synth-JEPA:用于无渲染合成器参数搜索的联合嵌入预测

AI 总结:Synth-JEPA通过联合嵌入预测学习参数与音频的互预测表示,实现无需渲染的合成器参数搜索,在域内优于基线,域外有竞争力,且可增加搜索计算提升匹配质量。

链接:https://arxiv.org/abs/2609.31024

机构:Sony Computer Science Laboratories(索尼计算机科学实验室); Queen Mary University of London(伦敦玛丽女王大学)

作者:Ben Hayes, Haokun Tian, Stefan Lattner

英文摘要:Sound matching can be formulated as optimizing synthesizer parameters against an audio-domain objective. However, objectives derived from generic audio representations are often difficult to optimize, while direct search requires rendering every candidate. We introduce Synth-JEPA, which learns mutually predictive audio and parameter representations from paired synthesizer data. At inference, candidate parameters are scored directly in this learned space, yielding a renderer-free objective whose audio geometry is shaped by parameter correspondences rather than generic audio similarity. We evaluate Synth-JEPA on Surge XT using held-out synthesizer sounds and out-of-domain NSynth and FSD50K targets, against inverse models, direct search, and learned proxy objectives. Synth-JEPA outperforms all baselines in-domain and remains competitive out-of-domain. Its matching quality continues to improve with additional test-time search, allowing compute to be traded for match quality. In pairwise listening tests, listeners preferred Synth-JEPA in 85% of trials overall. Together, these results show that an audio representation with a parameter-induced geometry allows synthesizer sound matching to be approached as an effective renderer-free search problem.


14. BAT-CLIP: Trimodal Alignment of Brain, Audio and Text

BAT-CLIP:脑、音频与文本的三模态对齐

AI 总结:提出BAT-CLIP,首个针对iEEG的CLIP式三模态对齐框架,联合对齐神经嵌入与音频和文本锚点,在Podcast基准上比双模态基线更稳健。

链接:https://arxiv.org/abs/2609.31180

机构:Yonsei University(延世大学); Seoul National University(首尔国立大学); Dartmouth College(达特茅斯学院)

作者:Suhyun Kim, Jinmo Han, Danny Dongyeop Han, Ahhyun Lucy Lee, Jewoon Lee, Yonghyeon Gwon, Zach Paris, Chun Kee Chung, Saewoong Bahk, Nam Soo Kim, Seong Jae Hwang, Jiook Cha

英文摘要:Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio or text-despite the brain's inherently multimodal speech processing. This induces a trade-off: audio anchoring preserves temporal structure but weakens linguistic separability, while text anchoring captures semantics yet discards acoustic detail. We propose BAT-CLIP, the first CLIP-style trimodal alignment framework for iEEG that jointly aligns neural embeddings to both pretrained audio and text anchors in a shared, frozen audio-text manifold. On the naturalistic Podcast benchmark, BAT-CLIP yields more robust representations than bimodal CLIP baselines. We also highlight the importance of using self-supervised foundation models for CLIP training.


15. AFA-Net: A Differential Attention Approach for Auditory Attention Detection

AFA-Net:一种用于听觉注意力检测的差分注意力方法

AI 总结:针对现有深度学习架构缺乏处理噪声EEG数据机制的问题,提出AFA-Net,用差分注意力机制聚焦任务相关神经活动,在2秒决策窗口达96.8%准确率且参数更少。

链接:https://arxiv.org/abs/2609.31402

机构:Johns Hopkins University(约翰斯·霍普金斯大学); University of Texas at Dallas(德克萨斯大学达拉斯分校)

作者:Philip H. Lee, Shreeram Suresh Chandra, Karan Thakkar, John H.L. Hansen

英文摘要:Auditory Attention Detection (AAD) utilizes electroencephalographic (EEG) signals to identify a target speaker in a multi-speaker environment. Despite considerable progress, existing deep learning architectures often lack explicit mechanisms for handling noisy EEG data. To address this limitation, we propose Auditory Focus Attention Networks (AFA-Net), a machine learning framework that replaces vanilla attention with a simple yet flexible differential attention mechanism to help focus on task-relevant neural activity. AFA-Net achieves an upward accuracy of 96.8% at the 2s decision window, while using substantially fewer parameters than most existing methods. To the best of our knowledge, AFA-Net is among the first frameworks to explicitly try to combat EEG noise to improve AAD.