Vision Transformers & Attention

Attention mechanisms, vision transformers, and the architectures replacing convolutions across vision tasks. We unpack how attention is used, misused, and reinvented in current research, from efficient attention branches to Mamba-style state space models.

Weak-Mamba-UNet: How CNN, ViT, and Visual Mamba Collaborate to Segment Medical Images from Scribbles

Weak-Mamba-UNet: How CNN, ViT, and Visual Mamba Collaborate to Segment Medical Images from Scribbles

Weak-Mamba-UNet: How CNN, ViT, and Visual Mamba Collaborate to Segment Medical Images from Scribbles | AI Trend Blend Medical AI & Weakly-Supervised Learning · arXiv:2402.10887 · University of Oxford / Mianyang Visual Engineering Center · 25 min read Teaching Three Different Brains to Agree — How Weak-Mamba-UNet Segments Hearts from Scribbles Ziyang Wang at Oxford […]

Weak-Mamba-UNet: How CNN, ViT, and Visual Mamba Collaborate to Segment Medical Images from Scribbles Read More »

Mamba-3: Three Simple Ideas That Finally Fix What Transformers Get Wrong at Inference.

Mamba-3: Three Simple Ideas That Finally Fix What Transformers Get Wrong at Inference

Mamba-3: Three Simple Ideas That Finally Fix What Transformers Get Wrong at Inference | AI Trend Blend AITrendBlend Machine Learning NLP & LLMs About Efficient AI · arXiv:2603.15569 · CMU & Princeton · March 2026 · 22 min read Mamba-3: Three Simple Ideas That Finally Fix What Transformers Get Wrong at Inference Time Researchers at

Mamba-3: Three Simple Ideas That Finally Fix What Transformers Get Wrong at Inference Read More »

GateMamba: Feature Gated Mixer in State Space Model for Point Cloud 3D Object Detection.

GateMamba: Feature Gated Mixer in State Space Model for Point Cloud 3D Object Detection

GateMamba: Feature Gated Mixer in State Space Model for Point Cloud 3D Object Detection | AI Trend Blend AITrendBlend Machine Learning Computer Vision About Autonomous Driving AI · ISPRS Journal of Photogrammetry and Remote Sensing 236 (2026) 640–653 · 22 min read GateMamba: How Three Gated Mixers Taught a Mamba Network to Stop Ignoring Cyclists

GateMamba: Feature Gated Mixer in State Space Model for Point Cloud 3D Object Detection Read More »

The Moon's Many Faces: A Single Unified Transformer for Multimodal Lunar Reconstruction

The Moon’s Many Faces: A Single Unified Transformer for Multimodal Lunar Reconstruction

The Moon’s Many Faces: A Single Unified Transformer for Multimodal Lunar Reconstruction | AI Trend Blend Planetary AI & 3D Reconstruction · ISPRS J. Photogramm. Remote Sens. 236 (2026) 363–379 · TU Dortmund University · 26 min read The Moon’s Many Faces: How One Transformer Learned to Speak All Four Languages of Lunar Science Simultaneously

The Moon’s Many Faces: A Single Unified Transformer for Multimodal Lunar Reconstruction Read More »

Fusion-Mamba: Hidden State Space Fusion for Cross-Modality Object Detection

Fusion-Mamba: Hidden State Space Fusion for Cross-Modality Object Detection

Fusion-Mamba: Hidden State Space Fusion for Cross-Modality Object Detection | AI Trend Blend AITrendBlend Machine Learning Computer Vision About Computer Vision · arXiv:2404.09146 · Beihang University · 21 min read Mamba Goes Multimodal: How Fusion-Mamba Built a Hidden State Space to End Modality Disparity Researchers at Beihang University asked what happens when you stop treating

Fusion-Mamba: Hidden State Space Fusion for Cross-Modality Object Detection Read More »

BGPANet: How Bi-Granular Progressive Attention Cracked the Skin Cancer Diagnosis Problem

BGPANet: How Bi-Granular Progressive Attention Cracked the Skin Cancer Diagnosis Problem

BGPANet: How Bi-Granular Progressive Attention Cracked the Skin Cancer Diagnosis Problem | AI Medical Research AIMedical Research Machine Learning Medical AI About Medical Image AI · Expert Systems With Applications 321 (2026) 132169 · 16 min read BGPANet: The Bi-Granular Attention Breakthrough That Finally Taught AI to Diagnose Skin Cancer Like a Dermatologist How a

BGPANet: How Bi-Granular Progressive Attention Cracked the Skin Cancer Diagnosis Problem Read More »

CFFormer: Cross CNN-Transformer Attention Model

CFFormer: How Cross CNN-Transformer Attention Finally Solves the Blurry Ultrasound Problem

CFFormer: How Cross CNN-Transformer Attention Finally Solves the Blurry Ultrasound Problem | AI Trend Blend AITrendBlend Machine Learning Computer Vision Medical AI About Medical Image Segmentation · Expert Systems with Applications · 2025 · 24 min read CFFormer: How Cross CNN-Transformer Attention Finally Solves the Blurry Ultrasound Problem Researchers at University of Nottingham Ningbo built

CFFormer: How Cross CNN-Transformer Attention Finally Solves the Blurry Ultrasound Problem Read More »

PraNet-V2: Dual-Supervised Reverse Attention for Medical Image Segmentation.

PraNet-V2 Fixes Medical Segmentation By Modeling Background

Analysis by the aitrendblend editorial team · Medical image segmentation · Computational Visual Media, 2026 PraNet-V2 Dual Supervised Reverse Attention Polyp Segmentation Multi Organ CT Cardiac MRI Overview of the PraNet-V2 decoder. Three cascaded DSRA stages refine a coarse prediction using both a foreground head and an independently supervised background head. A flat polyp sitting

PraNet-V2 Fixes Medical Segmentation By Modeling Background Read More »

The WEMoE framework transforms critical MLP modules into dynamic mixture-of-experts structures while statically merging non-critical components. Input-dependent routing weights allow the model to adaptively blend task-specific knowledge, achieving superior multi-task performance over static merging methods.

WEMoE: How a Mixture-of-Experts Approach Is Solving the Multi-Task Model Merging Problem

WEMoE: How a Mixture-of-Experts Approach Is Solving the Multi-Task Model Merging Problem | MedAI Research Deep Learning · TPAMI, 2026 · 18 min read The Static Model Merging Problem — and How WEMoE Learned to Adapt WEMoE introduces a dynamic mixture-of-experts approach to multi-task model merging, transforming how we combine fine-tuned neural networks by routing

WEMoE: How a Mixture-of-Experts Approach Is Solving the Multi-Task Model Merging Problem Read More »

MedDINOv3: Revolutionizing Medical Image Segmentation with Adaptable Vision Foundation Models

MedDINOv3 Adapts A Vision Foundation Model For CT And MRI Segmentation

AI for medical imaging and healthcare Vision foundation models CT and MRI segmentation Self supervised pretraining Analysis by the aitrendblend editorial team A radiation oncologist planning a course of treatment needs the kidneys, liver, spinal cord, and every nearby organ outlined precisely enough that the radiation beam avoids them by design rather than by luck.

MedDINOv3 Adapts A Vision Foundation Model For CT And MRI Segmentation Read More »