Multimodal AI

Models that combine images, text, audio, and clinical signals. We cover fusion architectures, missing-modality robustness, and cross-modal alignment, with an emphasis on what actually improves when modalities are combined.

H2CL: Dual-Geometry Hyperbolic-Euclidean Image-Text Learning for Medical Hierarchical Classification.

H2CL: Dual-Geometry Hyperbolic-Euclidean Image-Text Learning for Medical Hierarchical Classification

A UNSW Sydney team built H²CL — a dual-geometry image-text framework that simultaneously operates in Euclidean and hyperbolic spaces, combining group contrastive learning with a hyperbolic entailment loss, to classify medical images across clinical taxonomies. The result: 7% accuracy gains…

H2CL: Dual-Geometry Hyperbolic-Euclidean Image-Text Learning for Medical Hierarchical Classification Read More »

The Moon's Many Faces: A Single Unified Transformer for Multimodal Lunar Reconstruction

The Moon’s Many Faces: A Single Unified Transformer for Multimodal Lunar Reconstruction

A team at TU Dortmund built a single 29.7-million-parameter foundation model that can translate any combination of lunar data — grayscale images, elevation maps, surface normals, and albedo — to any other, all in a single forward pass. The key…

The Moon’s Many Faces: A Single Unified Transformer for Multimodal Lunar Reconstruction Read More »

Fusion-Mamba: Hidden State Space Fusion for Cross-Modality Object Detection

Fusion-Mamba: Hidden State Space Fusion for Cross-Modality Object Detection

Researchers at Beihang University asked what happens when you stop treating cross-modal fusion as a spatial alignment problem and start treating it as a state-space modeling problem. The answer outperforms every transformer-based fusion method while running faster — by mapping…

Fusion-Mamba: Hidden State Space Fusion for Cross-Modality Object Detection Read More »

IRDFusion: Iterative Differential Feedback for Multispectral Object Detection.

IRDFusion: Iterative Differential Feedback for Multispectral Object Detection

Researchers at Jiangsu University asked a simple question: what if we treated cross-modal feature fusion the same way electrical engineers treat differential amplifiers? The answer — IRDFusion — hits 88.3 mAP50 on FLIR and works as a plug-and-play module inside…

IRDFusion: Iterative Differential Feedback for Multispectral Object Detection Read More »

SLGNet: Structural Priors and Language-Guided Modulation for Multimodal Object Detection.

SLGNet: Structural Priors and Language-Guided Modulation for Multimodal Object Detection

Researchers at the Chinese Academy of Sciences built a multimodal object detector that freezes 88% of its parameters, asks a language model to describe the scene, and still beats every fully fine-tuned competitor — including on aerial drone footage at…

SLGNet: Structural Priors and Language-Guided Modulation for Multimodal Object Detection Read More »

RideJudge: How an 8B Model Outperforms 32B Baselines at Ride-Hailing Dispute Resolution

RideJudge: How an 8B Model Outperforms 32B Baselines at Ride-Hailing Dispute Resolution

Researchers from Nanjing University and Didi Chuxing built a multimodal LLM framework that reads GPS maps like a detective, consults platform regulations like a lawyer, and delivers transparent verdicts on driver-passenger disputes — hitting 88.41% accuracy while a 32B-scale baseline…

RideJudge: How an 8B Model Outperforms 32B Baselines at Ride-Hailing Dispute Resolution Read More »

Think Before You Segment: How TGS-Agent Teaches AI to Reason About Sound Before Picking Up a Brush.

Think Before You Segment: How TGS-Agent Teaches AI to Reason About Sound Before Picking Up a Brush

A research team from MBZUAI and NUS built an agentic segmentation system that actually stops to figure out what it is looking for before drawing a single mask — and the results expose a deep flaw in how every prior…

Think Before You Segment: How TGS-Agent Teaches AI to Reason About Sound Before Picking Up a Brush Read More »