Zhengyu Wu,Boyang Pang,Xunkai Li et al.
Zhengyu Wu et al.
As a privacy-preserving collaborative paradigm, federated graph learning (FGL) enables distributed training of graph neural networks (GNNs) without exposing raw graph data. Subgraph-FL has become the dominant FGL paradigm, yet most studies ...
ConvShareViT: A Vision Transformer-Like Architecture for Free-Space Optical Accelerators [0.03%]
ConvShareViT:用于自由空间光学加速器的VisionTransformer-like架构
Riad Ibadulla,Thomas M Chen,Constantino Carlos Reyes-Aldasoro
Riad Ibadulla
This article introduces convolutional shared vision transformers (ConvShareViT), a novel deep learning architecture that adapts the vision transformer (ViT) architecture to the 4f free-space optical system. ConvShareViT replaces linear laye...
Physically Guided Diffusion Framework With Neural Information Bottleneck Regulation for Robust Small Object Detection [0.03%]
基于神经信息瓶颈约束的物理引导扩散小目标检测框架
Yongcheng Zhou,Shilei Tan,Wei Li et al.
Yongcheng Zhou et al.
Robust small object detection in adverse environments remains challenging due to physical degradations and unstable feature representations. Existing detectors often struggle to maintain semantic consistency in haze, rain, and low-light con...
RL-FM: Reinforcement Learning-Driven Flow Matching for Multimodal Remote Sensing Image Classification [0.03%]
基于强化学习的流匹配遥感图像分类方法
Wei Zhang,Hanqing Tao,Zichen Wang et al.
Wei Zhang et al.
Multimodal remote sensing image classification improves models' capacity to recognize complex land-cover patterns by integrating data from heterogeneous sensors such as hyperspectral image (HSI) and light detection and ranging (LiDAR). Howe...
MGA-CLIP: A Multigranularity Attribution Framework for Cross-Modal Explainability in CLIP [0.03%]
MGA-CLIP:用于CLIP跨模态可解释性的多粒度归属框架
Xiaotian Cheng,Tianyi Zhou,Siwen Yin et al.
Xiaotian Cheng et al.
The contrastive language-image pretraining (CLIP) model has demonstrated remarkable performance in multimodal tasks, but the interpretability of its similarity-based cross-modal alignment mechanism has attracted considerable attention. Howe...
Without Paired Labeled Data: End-to-End Self-Supervised Learning for Drone-View Geo-Localization [0.03%]
无配对标注数据下的无人机视角地理定位的端到端自监督学习方法研究
Zhongwei Chen,Zhao-Xu Yang,Hai-Jun Rong et al.
Zhongwei Chen et al.
Drone-view geo-localization (DVGL) aims to achieve accurate localization of drones by retrieving the most relevant GNSS-tagged satellite images. However, most existing methods heavily rely on strictly prepaired drone-satellite images for su...
Contaminative Data-Driven Koopman Resilient Distributed Filtering for Unknown Stochastic Nonlinear Systems [0.03%]
基于污染数据的Koopman抗损分布式 resilient distributed filtering 抗损分布式滤波鲁棒估计 未知随机非线性系统估计算法
Xiaoyuan Zheng,Zhenrui Sun,Xindi Yang et al.
Xiaoyuan Zheng et al.
This article proposes a Koopman-enhanced distributed filtering for unknown stochastic nonlinear systems using contaminated datasets. To overcome the limitation that conventional Koopman operators fail to handle unknown stochastic dynamics, ...
ToFe: Lagged Token Freezing and Reusing for Efficient Vision Transformer Inference [0.03%]
具有滞后令牌冻结和重用的高效视觉变压器推理(ToFe)
Haoyue Zhang,Jie Zhang,Chenyu Hu et al.
Haoyue Zhang et al.
Although vision transformers (ViTs) have shown remarkable success in various vision tasks, their computationally expensive self-attention mechanisms hinder their deployment on resource-constrained edge devices. Token reduction, which discar...
Adaptive Graph-Guided Feature Decomposition for Unsupervised Multiview Feature Selection [0.03%]
自适应图引导特征分解的无监督多视图特征选择方法
Aihong Yuan,Hong Lv,Jin Hu et al.
Aihong Yuan et al.
The exponential growth and heterogeneity of multiview data have made dimensionality reduction indispensable. Unsupervised multiview feature selection (UMFS) has emerged as a pivotal solution. However, existing embedded UMFS methods still fa...
Harnessing Knowledge From Pretrained VLMs for Unsupervised Person Search [0.03%]
基于无监督学习的预训练视觉语言模型在行人搜索中的应用研究
Yanling Tian,Shanshan Zhang,Di Chen et al.
Yanling Tian et al.
Person search is a unified task that includes the subtasks of pedestrian detection and re-identification (re-ID). It is expensive to label pedestrian bounding boxes and person identities for training. Purely unsupervised (US) person search ...