The performance of acoustic machine learning systems is commonly evaluated using fully annotated test sets. In real-world deployments, however, exhaustively labeling large volumes of continuously collected audio data is often infeasible. Consequently, performance assessment typically relies on a small labeled subset of the available data, introducing a sampling bias that can severely distort evaluation metrics. This paper studies methods for compensating the bias in evaluation-labeled subsets under strict annotation-budget constraints. We study whether importance weighting techniques can mitigate this discrepancy by compensating for the selection bias. Specifically, we implement and compare three density-ratio estimation methods: kernel density estimation (KDE), logistic regression, and k-nearest neighbors (kNN), utilizing feature-space representations of the deployed audio. To emulate realistic deployment scenarios, the labeled subsets are generated using five distinct sampling strategies based on active learning techniques. Experiments conducted on an audio scene classification (ASC) benchmark demonstrate that importance weighting consistently yields more realistic accuracy estimates, significantly reducing the gap between subset-based metrics and the true evaluation performance.
https://arxiv.org/abs/2607.04463
Recent text-to-image generation models have demonstrated remarkable capabilities in synthesizing highly realistic images from text inputs alone. Although existing benchmarks can evaluate the generation capabilities of various models to some extent, they struggle to comprehensively and accurately measure performance across multiple dimensions, often failing to reveal the inherent deficiencies of models in specific categories. To address these limitations, we propose WeGenBench, a novel benchmark designed for the comprehensive, multi-perspective evaluation of text-to-image generation capabilities. Our benchmark comprises a total of 4,000 test prompts across two primary categories, meticulously balanced between Chinese and English to evaluate bilingual and cross-cultural generation capabilities. Beyond macroscopic scene classification, we annotate each prompt with multi-dimensional tags tailored to the distinct content and challenges of each language, thereby refining the generation tasks into more specific sub-categories. Through a cross-dimensional evaluation mechanism leveraging both scene classifications and multi-dimensional tags, WeGenBench can precisely pinpoint model shortcomings in specific generation categories. Furthermore, to measure generation quality more accurately, we design and validate several novel evaluation metrics by integrating Vision-Language Models (VLMs), which assess model performance on domain-specific tasks from three core aspects. Crucially, our approach yields both the assessment outcomes and the detailed reasoning trajectories, facilitating a rigorous verification of the accuracy and soundness of the evaluation results. Finally, we conduct systematic benchmarking on current state-of-the-art methods and provide an in-depth analysis of the limitations present in existing models.
https://arxiv.org/abs/2606.20100
This paper introduces the Event-Shifted Acoustic Scene (ESAS) dataset, a novel benchmark for evaluating the robustness of Acoustic Scene Classification (ASC) systems against unknown sound events. Existing ASC datasets typically contain recordings of clean and consistent audio, while real-world environments often include diverse and unexpected sound events. To bridge this gap, ESAS simulates real-world acoustic variability by injecting foreground sound events into background scenes with the assistance of large language models. In this work, we present the construction methodology, dataset statistics, and evaluation protocols. Furthermore, a comprehensive evaluation of state-of-the-art ASC systems is conducted using the ESAS benchmark. Experimental results reveal that existing ASC models suffer significant performance degradation when facing the event-shift challenge. The introduction of the ESAS dataset aims to drive future research toward event-robust ASC.
https://arxiv.org/abs/2606.06921
Foundation models offer a promising route to transferable remote sensing representations, but many current approaches depend on very large pretraining datasets and fixed sensor configurations, limiting their suitability for ecological and environmental applications, where observations often vary across platforms, spatial and spectral resolutions, and available modalities. We introduce FLORO, a multimodal geospatial foundation model designed to learn transferable representations from a small but highly diverse remote sensing corpus. FLORO is pretrained using masked autoencoding on a heterogeneous combination of Sentinel-1, Sentinel-2, SkySAT imagery, elevation, and UAV-derived data. To accommodate sensor variability, FLORO incorporates availability-aware inputs that indicate which spectral bands and auxiliary modalities are present in each sample, enabling a unified input space across heterogeneous sensor configurations. We evaluated FLORO on the PANGAEA benchmark under a frozen-encoder protocol across scene classification, segmentation, and regression tasks. Despite being pretrained on a smaller corpus than competing foundation models, FLORO achieved strong and stable transfer across optical, optical-SAR, and optical-elevation benchmarks spanning medium-resolution satellite, airborne, and ultra-high-resolution UAV imagery. FLORO obtained the second-best average segmentation performance across six PANGAEA benchmarks, trailing only a recently introduced foundation model pretrained on over two orders of magnitude more images, remained competitive on scene classification, and was robust in regression tasks, while qualitative results showed improved preservation of spatial structure in flood, urban, biomass, and canopy-height prediction settings. In a separate controlled experiment on EuroSAT-MS, geo-positional encoding further improved classification relative to absolute positional encoding.
https://arxiv.org/abs/2605.28174
Land Use Scene Classification (LUSC) from remote sensing imagery plays a critical role in environmental monitoring, urban planning, and sustainable resource management. In recent years, deep learning methods have significantly advanced the state of the art, with Convolutional Neural Networks (CNNs) dominating the field because of their strong ability to capture local spatial features. However, the emergence of Vision Transformers (ViTs) has introduced a new paradigm that models long-range dependencies through self-attention mechanisms, potentially enabling improved global context understanding. This paper presents a comparative assessment of Vision Transformers and CNN-based architecture for remote sensing land use scene classification. Representative CNN models, such as AlexNet, is evaluated alongside the Vision Transformer (ViT) using benchmark remote sensing datasets, including the UC Merced Land Use and EuroSAT Land Use datasets. The study examines classification accuracy, precision, recall, F1-score, and computational complexity to provide a comprehensive performance comparison. Experimental results demonstrate that CNNs perform robustly on datasets with limited training samples and strong local texture characteristics, whereas Vision Transformers exhibit superior performance in capturing global spatial relationships in complex scenes when sufficient training data are available. However, ViTs typically require greater computational resources and larger training datasets to achieve optimal performance. The findings of this study provide insights into the strengths and limitations of both architectures and offer guidance for selecting appropriate models for remote sensing land use scene classification applications.
https://arxiv.org/abs/2605.21268
We present Urban-ImageNet, a large-scale multi-modal dataset and evaluation benchmark for urban space perception from user-generated social media imagery. The corpus contains over 2 Million public social media images and paired textual posts collected from Weibo across 61 urban sites in 24 Chinese cities across 2019-2025, with controlled benchmark subsets at 1K, 10K, and 100K scale and a full 2M corpus for large-scale training and evaluation. Urban-ImageNet is organized by HUSIC, a Hierarchical Urban Space Image Classification framework that defines a 10-class taxonomy grounded in urban theory. The taxonomy is designed to distinguish activated and non-activated public spaces, exterior and interior urban environments, accommodation spaces, consumption content, portraits, and non-spatial social-media content. Rather than treating urban imagery as generic scene data, Urban-ImageNet evaluates whether machine perception models can capture spatial, social, and functional distinctions that are central to urban studies. The benchmark supports three tasks within one standardized library: (T1) urban scene semantic classification, (T2) cross-modal image-text retrieval, and (T3) instance segmentation. Our experiments evaluate representative vision, vision-language, and segmentation models, revealing strong performance on supervised scene classification but more challenging behavior in cross-modal retrieval and instance-level urban object segmentation. A multi-scale study further examines how model performance changes as balanced training data increases from 1K, 10K to 100K images. Urban-ImageNet provides a unified, theory-grounded, multi-city benchmark for evaluating how AI systems perceive and interpret contemporary urban spaces across modalities, scales, and task formulations. Dataset and benchmark are available at: this http URL and this http URL.
https://arxiv.org/abs/2605.09936
The exponential surge in high-resolution remote sensing data faces a severe bottleneck in satellite-to-ground transmission. Limited downlink bandwidth forces the use of extreme high-ratio compression, which irreversibly destroys high-frequency structural details essential for downstream machine perception tasks like object detection. While current super-resolution techniques attempt to recover these details, regression-based methods often yield over-smoothed textures, and generative diffusion models frequently introduce structural hallucinations that mislead detection systems. To address this trade-off, we propose the Structure-Aware Latent Diffusion (SALD) framework, an asymmetric edge-cloud collaborative SR system. At the resource-constrained edge, the system decouples imagery into a highly compressed low-frequency payload and a lightweight soft structural prior. Transmitting this decoupled representation minimizes bandwidth consumption. On the powerful cloud side, we introduce a Structure-Gated Large Kernel (SGLK) module and a Semantic-Guidance Engine (SGE) within the diffusion backbone. These modules leverage the transmitted structural priors to gate large-kernel convolutions, effectively capturing long-range dependencies inherent in aerial scenes while actively suppressing generative hallucinations. Extensive experiments on both the MSCM and UCMerced datasets demonstrate that, even under extreme bandwidth constraints, SALD achieves superior perceptual quality (LPIPS) and significantly enhances downstream performance in both scene classification and small-target detection.
https://arxiv.org/abs/2604.25319
Modern audio systems universally employ mel-scale representations derived from 1940s Western psychoacoustic studies, potentially encoding cultural biases that create systematic performance disparities. We present a comprehensive evaluation of cross-cultural bias in audio front-ends, comparing mel-scale features with learnable alternatives (LEAF, SincNet) and psychoacoustic variants (ERB, Bark, CQT) across speech recognition (11 languages), music analysis (6 collections), and European acoustic scene classification (10 European cities). Our controlled experiments isolate front-end contributions while holding architecture and training protocols minimal and constant. Results demonstrate that mel-scale features yield 31.2% WER for tonal languages compared to 18.7% for non-tonal languages (12.5% gap), and show 15.7% F1 degradation between Western and non-Western music. Alternative representations significantly reduce these disparities: LEAF reduces the speech gap by 34% through adaptive frequency allocation, CQT achieves 52% reduction in music performance gaps, and ERB-scale filtering cuts disparities by 31% with only 1% computational overhead. We also release FairAudioBench, enabling cross-cultural evaluation, and demonstrate that adaptive frequency decomposition offers practical paths toward equitable audio processing. These findings reveal how foundational signal processing choices propagate bias, providing crucial guidance for developing inclusive audio systems.
https://arxiv.org/abs/2604.10503
The adoption of vision-language models (VLMs) for wireless network management is accelerating, yet no systematic understanding exists of where these large foundation models outperform lightweight convolutional neural networks (CNNs) for spectrum-related tasks. This paper presents the first diagnostic comparison of VLMs and CNNs for spectrum heatmap understanding in non-terrestrial network and terrestrial network (NTN-TN) cooperative systems. We introduce SpectrumQA, a benchmark comprising 108K visual question-answer pairs across four granularity levels: scene classification (L1), regional reasoning (L2), spatial localization (L3), and semantic reasoning (L4). Our experiments on three NTN-TN scenarios with a frozen Qwen2-VL-7B and a trained ResNet-18 reveal a clear taskdependent complementarity: CNN achieves 72.9% accuracy at severity classification (L1) and 0.552 IoU at spatial localization (L3), while VLM uniquely enables semantic reasoning (L4) with F1=0.576 using only three in-context examples-a capability fundamentally absent in CNN architectures. Chain-of-thought (CoT) prompting further improves VLM reasoning by 12.6% (F1: 0.209->0.233) while having zero effect on spatial tasks, confirming that the complementarity is rooted in architectural differences rather than prompting limitations. A deterministic task-type router that delegates supervised tasks to CNN and reasoning tasks to VLM achieves a composite score of 0.616, a 39.1% improvement over CNN alone. We further show that VLM representations exhibit stronger cross-scenario robustness, with smaller performance degradation in 5 out of 6 transfer directions. These findings provide actionable guidelines: deploy CNNs for spatial localization and VLMs for semantic spectrum reasoning, rather than treating them as substitutes.
视觉语言模型(VLMs)在无线网络管理中的应用正加速推进,然而目前尚缺乏系统性研究阐明这些大型基础模型在与轻量级卷积神经网络(CNNs)对比时,于频谱相关任务中究竟在哪些方面表现更优。本文首次对非地面网络与地面网络(NTN-TN)协作系统中用于频谱热图理解的VLMs与CNNs进行了诊断性对比。我们提出了SpectrumQA基准,包含108K个跨四个粒度级别的视觉问答对:场景分类(L1)、区域推理(L2)、空间定位(L3)和语义推理(L4)。在三个NTN-TN场景中,使用冻结的Qwen2-VL-7B模型与训练好的ResNet-18进行实验,结果揭示了清晰的任务依赖互补性:CNN在严重度分类(L1)上达到72.9%准确率,在空间定位(L3)上达到0.552的IoU;而VLM仅通过三个上下文示例,在语义推理(L4)上实现了F1=0.576的独特能力——这是CNN架构 fundamentally 所不具备的。思维链(CoT)提示将VLM推理能力提升了12.6%(F1: 0.209→0.233),但对空间任务毫无影响,证实了这种互补性根植于架构差异而非提示限制。一个确定性任务类型路由器将监督任务分派给CNN、推理任务分派给VLM,实现了0.616的复合分数,较单独使用CNN提升了39.1%。我们还发现VLM表征展现出更强的跨场景鲁棒性,在6个迁移方向中的5个里性能下降更小。这些发现提供了可操作的指导:应为空间定位部署CNN,为语义频谱推理部署VLM,而非将它们视为替代关系。
https://arxiv.org/abs/2604.03774
YouTube Shorts have become central to news consumption on the platform, yet research on how geopolitical events are represented in this format remains limited. To address this gap, we present a multimodal pipeline that combines automatic transcription, aspect-based sentiment analysis (ABSA), and semantic scene classification. The pipeline is first assessed for feasibility and then applied to analyze short-form coverage of the Israel-Hamas war by state-funded outlets. Using over 2,300 conflict-related Shorts and more than 94,000 visual frames, we systematically examine war reporting across major international broadcasters. Our findings reveal that the sentiment expressed in transcripts regarding specific aspects differs across outlets and over time, whereas scene-type classifications reflect visual cues consistent with real-world events. Notably, smaller domain-adapted models outperform large transformers and even LLMs for sentiment analysis, underscoring the value of resource-efficient approaches for humanities research. The pipeline serves as a template for other short-form platforms, such as TikTok and Instagram, and demonstrates how multimodal methods, combined with qualitative interpretation, can characterize sentiment patterns and visual cues in algorithmically driven video environments.
YouTube Shorts已成为该平台上新闻消费的核心形式,然而关于地缘政治事件在此类短视频中的呈现方式研究仍较为匮乏。为填补这一空白,我们提出了一种融合自动转录、基于方面的情感分析(ABSA)及语义场景分类的多模态分析流程。该流程首先经过可行性验证,随后被应用于分析国有媒体对以色列-哈马斯战争的短视频报道。通过2300余条冲突相关Shorts及超过9.4万帧视觉画面,我们系统考察了主要国际广播机构的战争报道模式。研究发现:各媒体在转录文本中对特定方面的情感表达存在差异且随时间变化,而场景类型分类则反映了与现实事件一致的视觉线索。值得注意的是,针对特定领域优化的较小模型在情感分析任务中表现优于大型Transformer模型乃至大语言模型,这凸显了资源高效型方法在人文研究中的价值。该流程可为TikTok、Instagram等其他短视频平台提供参考模板,并展示了如何通过多模态方法结合质性解读,来刻画算法驱动视频环境中的情感模式与视觉线索。
https://arxiv.org/abs/2604.00994
This dataset provides a large collection of 10,915 synthetic hyperspectral image cubes paired with pixel-level vegetation trait maps, designed to support research in radiative transfer emulation, vegetation trait retrieval, and uncertainty quantification. Each hyperspectral cube contains 211 bands spanning 400--2500 nm at 10 nm resolution and a fixed spatial layout of 64 \times 64 pixels, offering continuous simulated surface reflectance spectra suitable for emulator development and machine-learning tasks requiring high spectral detail. Vegetation traits were derived by inverting Sentinel-2 Level-2A surface reflectance using a PROSAIL-based lookup-table approach, followed by forward PROSAIL simulations to generate hyperspectral reflectance under physically consistent canopy and illumination conditions. The dataset covers four ecologically diverse regions -- East Africa, Northern France, Eastern India, and Southern Spain -- and includes 5th and 95th percentile uncertainty maps as well as Sentinel-2 scene classification layers. This resource enables benchmarking of inversion methods, development of fast radiative transfer emulators, and studies of spectral--biophysical relationships under controlled yet realistic environmental variability.
https://arxiv.org/abs/2603.28390
Few-shot remote sensing image scene classification (FS-RSISC) aims at classifying remote sensing images with only a few labeled samples. The main challenges lie in small inter-class variances and large intra-class variances, which are the inherent property of remote sensing images. To address these challenges, we propose a transfer-based Dual Contrastive Network (DCN), which incorporates two auxiliary supervised contrastive learning branches during the training process. Specifically, one is a Context-guided Contrastive Learning (CCL) branch and the other is a Detail-guided Contrastive Learning (DCL) branch, which focus on inter-class discriminability and intra-class invariance, respectively. In the CCL branch, we first devise a Condenser Network to capture context features, and then leverage a supervised contrastive learning on top of the obtained context features to facilitate the model to learn more discriminative features. In the DCL branch, a Smelter Network is designed to highlight the significant local detail information. And then we construct a supervised contrastive learning based on the detail feature maps to fully exploit the spatial information in each map, enabling the model to concentrate on invariant detail features. Extensive experiments on four public benchmark remote sensing datasets demonstrate the competitive performance of our proposed DCN.
少样本遥感图像场景分类(FS-RSISC)旨在仅使用少量标注样本对遥感图像进行分类。其主要挑战在于类间差异小、类内差异大,这是遥感图像的固有特性。为应对这些挑战,我们提出了一种基于迁移的双对比网络(DCN),该网络在训练过程中集成了两个辅助监督对比学习分支。具体而言,其中一个分支为上下文引导对比学习(CCL)分支,另一个为细节引导对比学习(DCL)分支,二者分别聚焦于类间判别性与类内不变性。在CCL分支中,我们首先设计了冷凝器网络以捕获上下文特征,随后在获得的上下文特征上应用监督对比学习,从而促进模型学习更具判别性的特征。在DCL分支中,我们设计了熔炉网络以凸显显著的局部细节信息,并基于细节特征图构建监督对比学习,以充分挖掘各特征图的空间信息,使模型能够集中于不变的细节特征。在四个公开基准遥感数据集上的大量实验表明,我们提出的DCN具有竞争力的性能。
https://arxiv.org/abs/2603.23161
Remote sensing scene classification has experienced a paradigmatic transformation from traditional handcrafted feature methods to sophisticated artificial intelligence systems that now form the backbone of modern Earth observation applications. This comprehensive survey examines the complete methodological evolution, systematically tracing development from classical texture descriptors and machine learning classifiers through the deep learning revolution to current state-of-the-art foundation models and generative AI approaches. We chronicle the pivotal shift from manual feature engineering to automated hierarchical representation learning via convolutional neural networks, followed by advanced architectures including Vision Transformers, graph neural networks, and hybrid frameworks. The survey provides in-depth coverage of breakthrough developments in self-supervised foundation models and vision-language systems, highlighting exceptional performance in zero-shot and few-shot learning scenarios. Special emphasis is placed on generative AI innovations that tackle persistent challenges through synthetic data generation and advanced feature learning strategies. We analyze contemporary obstacles including annotation costs, multimodal data fusion complexities, interpretability demands, and ethical considerations, alongside current trends in edge computing deployment, federated learning frameworks, and sustainable AI practices. Based on comprehensive analysis of recent advances and gaps, we identify key future research priorities: advancing hyperspectral and multi-temporal analysis capabilities, developing robust cross-domain generalization methods, and establishing standardized evaluation protocols to accelerate scientific progress in remote sensing scene classification systems.
https://arxiv.org/abs/2603.26751
Scene understanding plays a critical role in enabling intelligence and autonomy in robotic systems. Traditional approaches often face challenges, including occlusions, ambiguous boundaries, and the inability to adapt attention based on task-specific requirements and sample variations. To address these limitations, this paper presents an efficient RGB-D scene understanding model that performs a range of tasks, including semantic segmentation, instance segmentation, orientation estimation, panoptic segmentation, and scene classification. The proposed model incorporates an enhanced fusion encoder, which effectively leverages redundant information from both RGB and depth inputs. For semantic segmentation, we introduce normalized focus channel layers and a context feature interaction layer, designed to mitigate issues such as shallow feature misguidance and insufficient local-global feature representation. The instance segmentation task benefits from a non-bottleneck 1D structure, which achieves superior contour representation with fewer parameters. Additionally, we propose a multi-task adaptive loss function that dynamically adjusts the learning strategy for different tasks based on scene variations. Extensive experiments on the NYUv2, SUN RGB-D, and Cityscapes datasets demonstrate that our approach outperforms existing methods in both segmentation accuracy and processing speed.
场景理解在增强机器人系统的智能和自主性方面扮演着关键角色。传统方法经常面临遮挡、边界模糊以及无法根据任务特定需求和样本变化调整注意力等问题的挑战。为了克服这些限制,本文提出了一种高效的RGB-D场景理解模型,该模型能够执行一系列任务,包括语义分割、实例分割、姿态估计、全景分割和场景分类。所提出的模型包含了一个增强融合编码器,可以有效利用来自RGB和深度输入数据的冗余信息。 对于语义分割任务,我们引入了标准化关注通道层和上下文特征交互层,旨在解决浅层特征误导以及局部-全局特征表示不足等问题。在实例分割任务中,我们的模型采用了一种非瓶颈1D结构,在减少参数的同时实现了更优的轮廓表示。此外,我们还提出了一种多任务自适应损失函数,可以根据场景变化动态调整不同的学习策略。 通过在NYUv2、SUN RGB-D和Cityscapes数据集上的大量实验验证,我们的方法不仅在分割准确性上超越了现有技术,而且在处理速度方面也表现出色。
https://arxiv.org/abs/2603.07570
Pretraining and fine-tuning have emerged as a new paradigm in remote sensing image interpretation. Among them, Masked Autoencoder (MAE)-based pretraining stands out for its strong capability to learn general feature representations via reconstructing masked image regions. However, applying MAE to multispectral remote sensing images remains challenging due to complex backgrounds, indistinct targets, and the lack of semantic guidance during masking, which hinders the learning of underlying structures and meaningful spatial-spectral features. To address this, we propose a simple yet effective approach, Spectral Index-Guided MAE (SIGMAE), for multispectral image pretraining. The core idea is to incorporate domain-specific spectral indices as prior knowledge to guide dynamic token masking toward informative regions. SIGMAE introduces Semantic Saliency-Guided Dynamic Token Masking (SSDTM), a curriculum-style strategy that quantifies each patch's semantic richness and internal heterogeneity to adaptively select the most informative tokens during training. By prioritizing semantically salient regions and progressively increasing sample difficulty, SSDTM enhances spectrally rich and structurally aware representation learning, mitigates overfitting, and reduces redundant computation compared with random masking. Extensive experiments on five widely used datasets covering various downstream tasks, including scene classification, semantic segmentation, object extraction and change detection, demonstrate that SIGMAE outperforms other pretrained geospatial foundation models. Moreover, it exhibits strong spatial-spectral reconstruction capability, even with a 90% mask ratio, and improves complex target recognition under limited labeled data. The source codes and model weights will be released at this https URL.
预训练和微调已成为遥感图像解释的新范式。其中,基于遮罩自动编码器(MAE)的预训练因其通过重构被遮罩的图像区域来学习通用特征表示的能力而脱颖而出。然而,由于复杂的背景、不清晰的目标以及在遮罩过程中缺乏语义指导,将MAE应用于多光谱遥感图像仍然具有挑战性,这阻碍了底层结构和有意义的空间-光谱特征的学习。为了解决这个问题,我们提出了一种简单而有效的方法——基于光谱指数的MAE(SIGMAE),用于多光谱图像预训练。核心思想是将领域特定的光谱指数作为先验知识,以指导动态标记遮罩向信息丰富的区域进行导向。 SIGMAE引入了语义显著性引导的动态令牌遮罩(SSDTM)策略,这是一种类似课程的学习方式,它量化每个补丁的语义丰富性和内部异质性,并在训练过程中自适应地选择最富有信息性的令牌。通过优先考虑语义显著区域并逐步增加样本难度,SSDTM增强了光谱丰富的和结构感知表示学习,减轻了过拟合问题,并与随机遮罩相比减少了冗余计算。 在五个广泛使用的数据集上进行了大量实验,这些数据集涵盖了包括场景分类、语义分割、目标提取和变化检测在内的各种下游任务。结果表明,SIGMAE优于其他预训练的地理空间基础模型。此外,即使在90%的遮罩比例下,它也展示了强大的光谱-空间重建能力,并且在标签有限的情况下提高了复杂目标识别的能力。 源代码和模型权重将在以下链接发布:[提供URL]
https://arxiv.org/abs/2603.07463
This paper introduces DashengTokenizer, a continuous audio tokenizer engineered for joint use in both understanding and generation tasks. Unlike conventional approaches, which train acoustic tokenizers and subsequently integrate frozen semantic knowledge, our method inverts this paradigm: we leverage frozen semantic features and inject acoustic information. In linear evaluation across 22 diverse tasks, our method outperforms previous audio codec and audio encoder baselines by a significant margin while maintaining competitive audio reconstruction quality. Notably, we demonstrate that this acoustic injection improves performance for tasks such as speech emotion recognition, music understanding, and acoustic scene classification. We further evaluate the tokenizer's generative performance on text-to-audio (TTA), text-to-music (TTM), and speech enhancement (SE). Our approach surpasses standard variational autoencoder (VAE)-based methods on TTA and TTM tasks, while its effectiveness on SE underscores its capabilities as a general-purpose audio encoder. Finally, our results challenge the prevailing assumption that VAE-based architectures are a prerequisite for audio synthesis. Checkpoints are available at this https URL.
这篇论文介绍了DashengTokenizer,这是一种连续音频标记器,专门设计用于同时处理理解和生成任务。与传统方法不同,后者训练声学标记器并随后整合固定的语义知识,我们的方法则反其道而行之:我们利用固定化的语义特征,并注入声学信息。在包含22项多样化任务的线性评估中,我们的方法显著超越了以往的音频编解码器和音频编码器基准,在保持竞争性的音频重构质量的同时实现了这一点。值得注意的是,我们展示了这种声学注射可以提高诸如语音情感识别、音乐理解和声音场景分类等任务的表现。此外,我们在文本到音频(TTA)、文本到音乐(TTM)以及语音增强(SE)的生成性能上评估了该标记器的能力。我们的方法在TTA和TTM任务中超越了基于变分自动编码器(VAE)的标准方法,并且其在SE上的有效性进一步证实了它作为通用音频编码器的能力。最后,我们的结果挑战了认为基于VAE架构是音频合成的先决条件这一普遍假设。相关模型检查点可以在提供的链接处获取。
https://arxiv.org/abs/2602.23765
Speech Enhancement (SE) in audio devices is often supported by auxiliary modules for Voice Activity Detection (VAD), SNR estimation, or Acoustic Scene Classification to ensure robust context-aware behavior and seamless user experience. Just like SE, these tasks often employ deep learning; however, deploying additional models on-device is computationally impractical, whereas cloud-based inference would introduce additional latency and compromise privacy. Prior work on SE employed Dynamic Channel Pruning (DynCP) to reduce computation by adaptively disabling specific channels based on the current input. In this work, we investigate whether useful signal properties can be estimated from these internal pruning masks, thus removing the need for separate models. We show that simple, interpretable predictors achieve up to 93% accuracy on VAD, 84% on noise classification, and an R2 of 0.86 on F0 estimation. With binary masks, predictions reduce to weighted sums, inducing negligible overhead. Our contribution is twofold: on one hand, we examine the emergent behavior of DynCP models through the lens of downstream prediction tasks, to reveal what they are learning; on the other, we repurpose and re-propose DynCP as a holistic solution for efficient SE and simultaneous estimation of signal properties.
语音增强(SE)在音频设备中通常依赖于辅助模块,如声源活动检测(VAD)、信噪比估计或声音场景分类来确保稳健的上下文感知行为和无缝用户体验。与SE一样,这些任务也常采用深度学习技术;然而,在设备上部署额外模型从计算角度来看是不可行的,而基于云的推理则会增加延迟并损害隐私。先前的研究中,SE采用了动态通道剪枝(DynCP)来通过自适应地禁用特定通道减少计算量,具体依据当前输入进行调整。 在本项研究工作中,我们探讨了是否可以从这些内部剪枝掩码中估计出有用的信号属性,从而消除单独模型的需求。研究表明,简单的、可解释的预测器可以达到高达93%的VAD准确率,84%的噪声分类准确率以及F0估计上R2值为0.86的表现。使用二进制掩码时,预测简化为加权求和操作,几乎不会增加额外开销。 我们的贡献是双重的:一方面,我们通过下游预测任务的角度来考察DynCP模型产生的新兴行为,揭示它们所学的内容;另一方面,我们将重新定义并提议使用DynCP作为高效语音增强和同时估计信号属性的整体解决方案。
https://arxiv.org/abs/2602.10666
Recent advancements in Multimodal Large Language Models (MLLMs) have enabled complex reasoning. However, existing remote sensing (RS) benchmarks remain heavily biased toward perception tasks, such as object recognition and scene classification. This limitation hinders the development of MLLMs for cognitively demanding RS applications. To address this, , we propose a Vision Language ReaSoning Benchmark (VLRS-Bench), which is the first benchmark exclusively dedicated to complex RS reasoning. Structured across the three core dimensions of Cognition, Decision, and Prediction, VLRS-Bench comprises 2,000 question-answer pairs with an average length of 71 words, spanning 14 tasks and up to eight temporal phases. VLRS-Bench is constructed via a specialized pipeline that integrates RS-specific priors and expert knowledge to ensure geospatial realism and reasoning complexity. Experimental results reveal significant bottlenecks in existing state-of-the-art MLLMs, providing critical insights for advancing multimodal reasoning within the remote sensing community.
最近的多模态大型语言模型(MLLMs)进展使得复杂推理成为可能。然而,现有的遥感(RS)基准测试仍然偏向于感知任务,例如物体识别和场景分类。这种限制阻碍了MLLMs在认知要求高的RS应用中的发展。为了解决这个问题,我们提出了一个视觉语言推理基准(VLRS-Bench),这是首个专门针对复杂RS推理的基准。VLRS-Bench涵盖了认知、决策和预测三个核心维度,包含2,000个问题-答案对,平均长度为71个单词,并涵盖14项任务及多达八个时间阶段。通过一个整合了遥感特定先验知识和专家知识的专业化管道构建而成的VLRS-Bench确保了地理空间的真实性以及推理复杂性。实验结果显示现有最先进的MLLMs存在显著瓶颈,这为推进多模态推理在遥感领域的进步提供了关键见解。
https://arxiv.org/abs/2602.07045
Aerial images play a vital role in urban planning and environmental preservation, as they consist of various structures, representing different types of buildings, forests, mountains, and unoccupied lands. Due to its heterogeneous nature, developing robust models for scene classification remains a challenge. In this study, we conduct a literature review of various machine learning methods for aerial image classification. Our survey covers a range of approaches from handcrafted features (e.g., SIFT, LBP) to traditional CNNs (e.g., VGG, GoogLeNet), and advanced deep hybrid networks. In this connection, we have also designed Aerial-Y-Net, a spatial attention-enhanced CNN with multi-scale feature fusion mechanism, which acts as an attention-based model and helps us to better understand the complexities of aerial images. Evaluated on the AID dataset, our model achieves 91.72% accuracy, outperforming several baseline architectures.
https://arxiv.org/abs/2601.18263
Recent years have witnessed the remarkable success of deep learning in remote sensing image interpretation, driven by the availability of large-scale benchmark datasets. However, this reliance on massive training data also brings two major challenges: (1) high storage and computational costs, and (2) the risk of data leakage, especially when sensitive categories are involved. To address these challenges, this study introduces the concept of dataset distillation into the field of remote sensing image interpretation for the first time. Specifically, we train a text-to-image diffusion model to condense a large-scale remote sensing dataset into a compact and representative distilled dataset. To improve the discriminative quality of the synthesized samples, we propose a classifier-driven guidance by injecting a classification consistency loss from a pre-trained model into the diffusion training process. Besides, considering the rich semantic complexity of remote sensing imagery, we further perform latent space clustering on training samples to select representative and diverse prototypes as visual style guidance, while using a visual language model to provide aggregated text descriptions. Experiments on three high-resolution remote sensing scene classification benchmarks show that the proposed method can distill realistic and diverse samples for downstream model training. Code and pre-trained models are available online (this https URL).
近年来,深度学习在遥感图像解读领域取得了显著的成功,这主要得益于大规模基准数据集的可用性。然而,对大量训练数据的高度依赖也带来了两大挑战:一是高昂的存储和计算成本,二是当涉及敏感类别时存在数据泄露的风险。为了解决这些问题,本研究首次将数据集蒸馏的概念引入遥感图像解读领域。具体而言,我们通过训练文本到图像扩散模型来将大规模遥感数据集浓缩成一个紧凑且具有代表性的精炼数据集。 为了提高合成样本的判别质量,我们提出了一种由分类器驱动的指导方法,在扩散训练过程中注入预训练模型中的分类一致性损失。此外,考虑到遥感影像丰富的语义复杂性,我们在训练样本上执行潜在空间聚类以选择代表性且多样化的原型作为视觉风格引导,并使用视觉语言模型提供聚合的文字描述。 在三个高分辨率遥感场景分类基准上的实验表明,所提出的方法能够为下游模型训练蒸馏出真实且多样的样本。代码和预训练模型可在在线平台获取(此链接)。
https://arxiv.org/abs/2601.15829