← 返回资源分享
VMamba: Visual State Space Model
论文
论文
发布时间2024-01-18
发表arXiv:2401.10166
作者:Lingxi Xie,Yue Liu,YaoWei Wang,Yunjie Tian,Qixiang Ye,Yuzhong Zhao,Hongtian Yu,Yunfan Liu
详细介绍
Designing computationally efficient network architectures persists as an ongoing necessity in computer vision. In this paper, we transplant Mamba, a state-space language model, into VMamba, a vision backbone that works in linear time complexity. At the core of VMamba lies a stack of Visual State-Space (VSS) blocks with the 2D Selective Scan (SS2D) module. By traversing along four scanning routes, SS2D helps bridge the gap between the ordered nature of 1D selective scan and the non-sequential structure of 2D vision data, which facilitates the gathering of contextual information from various sources and perspectives. Based on the VSS blocks, we develop a family of VMamba architectures and accelerate them through a succession of architectural and implementation enhancements. Extensive experiments showcase VMamba's promising performance across diverse visual perception tasks, highlighting its advantages in input scaling efficiency compared to existing benchmark models. Source code is available at https://github.com/MzeroMiko/VMamba.
代码仓库 (10)
mzeromiko/vmamba官方PyTorch
weitunglin/pixmambaPyTorch
zs1314/microscopic-mambaPyTorch
hunto/localmambaPyTorch
chenhongruixuan/mambacd
zs1314/skinmambaPyTorch
zs1314/octamambaPyTorch
yuhengsss/msvmambaPyTorch
longshaocong/dgmambaPyTorch
raytrun/mamba-clipPyTorch
