AFME Teaches Knowledge Graphs Which Modality To Trust

Analysis by the aitrendblend editorial team · Graph neural networks · 13 min read
Knowledge Graphs Link Prediction Multimodal Fusion Generative Adversarial Networks Self Attention
A knowledge graph with text and image nodes connected by dotted lines representing a missing relationship being predicted by AFME
A missing link between two entities, and four different kinds of evidence arguing about how to fill it in.
Picture a knowledge graph entry for the Elizabeth Tower. It has a grainy nineteenth century photograph, a paragraph of text that never quite says everything you need, and a web of relations to Sir Charles Barry, to the Great Fire of London, to the city itself. Now ask a model to decide whether that tower caused Big Ben to exist or the other way around, using only that mismatched pile of evidence. Most multimodal knowledge graph systems handle this by mashing every modality together and hoping the good signal outweighs the bad. A team from Minzu University of China built AFME because that hope is not a strategy.

Key points

  • AFME splits multimodal knowledge graph completion into two modules, one that cleans and weighs each modality before fusion, and one that uses a GAN to regenerate weak or missing modality features.
  • A relationship conditioned gate suppresses features irrelevant to the current triple before fusion even starts, rather than fusing first and hoping noise averages out.
  • Modality weights are set by a confidence score based on the L2 norm of each feature, softened by a learned relationship specific temperature, so the mix changes per triple rather than staying fixed.
  • The generator uses the structural embedding to steer a self attention refinement of the weaker modalities, trained adversarially against a discriminator with a WGAN gradient penalty for stability.
  • Across four benchmarks, MKG-W, MKG-Y, TIVA, and KVC16K, AFME reports the best MRR of every method tested, beating the next strongest baseline, NATIVE, by margins from under one percent up to nearly three percent depending on the dataset.
  • The enhancement module is portable. Bolting it onto three unrelated baseline models improved every one of them on the TIVA dataset, which is a stronger generalization claim than a single leaderboard number.

The unglamorous problem hiding inside every multimodal knowledge graph

Knowledge graph completion sounds abstract until you picture the actual data sitting in one of these graphs. An entity might have a blurry thumbnail, a two sentence description that was scraped and never edited, and a handful of relations that were extracted automatically and are only sometimes right. Add video or audio, as the TIVA and KVC16K datasets used in this paper both do, and the quality gap between modalities gets even wider. A model that treats a low resolution image the same way it treats a clean structural embedding is going to get pulled around by noise it should be ignoring.

That is the specific failure this paper by Zenglong Wang, Xuan Liu, Zheng Liu, Yu Weng, and Chaomurilige is trying to fix. Their framework, AFME, stands for Adaptive Fusion and Modality Information Enhancement, and the name is a fair description of what it actually does rather than marketing. It does not propose a new scoring function from scratch. It proposes a way to clean, weigh, and if necessary regenerate modality features before they ever reach a fairly conventional rotation based scoring function, on the argument that fusion quality, not scoring function novelty, is what has been holding this task back.

Why concatenation was never going to be enough

Most prior multimodal completion methods fall into a predictable pattern. IKRL added visual features with attention over which modality to trust. TBKGC and TransAE combined modal representations with structural embeddings using autoencoders or fixed weighting. RSME and OTKGE went further with sparse coding and optimal transport respectively, both trying to find a smarter static way to combine modalities. What none of them do, according to this paper, is let the fusion weight of a modality change based on what relationship is actually being predicted right now, or aggressively suppress the parts of a modality that have nothing to do with that relationship before fusion happens.

There is a second pattern the paper points at, this time in the GAN literature. KBGAN used adversarial training to generate harder negative samples, and MMKRL extended a similar idea to multimodal embeddings. Both treat the generator as a way to make training data more difficult, not as a way to actively repair a specific missing or noisy modality feature using the graph’s own structure as a guide. AFME’s generator does something closer to the second thing, using the structural embedding of an entity as a conditioning signal for regenerating its weaker modalities.

The core bet. Fuse less blindly and regenerate what is missing, instead of fusing everything you have and hoping the model learns to discount the bad parts on its own.

Two modules, one doing triage and one doing repair

AFME is built from two pieces that run in sequence. The Modal Information Fusion module, MoIFu, cleans and weighs the modalities that already exist. The Modal Information Enhancement module, MoIEn, uses a GAN to produce higher quality versions of the weaker ones. Both sit on top of a fairly standard setup, pretrained encoders per modality (BERT for text, a convolutional network for images, established encoders for video and audio), each projected into a shared embedding space, alongside trainable structural embeddings for the graph’s own entities and relations.

MoIFu, denoise first, then decide how much each modality is worth

Before any fusion happens, each modality feature passes through a gate conditioned on the relationship of the current triple. The relationship embedding interacts with the modality feature through an elementwise product, gets pushed through a sigmoid, and the result multiplies back into the original feature.

\( \text{gate}_m = \sigma\big(W_g \cdot (h_m \odot r) + b_g\big) \)
\( \tilde{h}_m = \text{gate}_m \odot h_m \)

Read plainly, this means the model learns, for a given relationship, which parts of a given modality actually matter and which parts are noise or irrelevant detail, and it turns down the volume on the latter before that modality gets anywhere near the fusion step. That is a meaningfully different design choice from denoising a modality in isolation, since the same image feature could be treated as mostly signal for one relationship and mostly noise for another, depending on what the relationship actually needs from it.

Once each modality has been cleaned, AFME still has to decide how much weight it deserves relative to the others. It computes a confidence score for each modality from the L2 norm of its cleaned feature, on the reasoning that a feature vector with more energy in it is carrying more information, then folds that confidence into a softmax that is also conditioned on a relationship specific temperature term.

$$ \omega_m(h, r) = \dfrac{\exp\!\big(\alpha_m \cdot U \odot \tanh(\tilde{h}_m) / \sigma(\tau_r)\big)}{\sum_{n \in M \cup \{S\}} \exp\!\big(\alpha_n \cdot U \odot \tanh(\tilde{h}_n) / \sigma(\tau_r)\big)} $$

The relationship specific temperature is the detail worth sitting with. A fixed softmax temperature applies the same sharpness to every relationship in the graph, whether that relationship depends heavily on a photograph or barely needs one. Letting the temperature vary by relationship means the model can be decisive about modality weighting for relations where one modality clearly dominates, and more balanced for relations where several modalities genuinely contribute.

MoIEn, let the graph structure repair what is weak or missing

Cleaning and weighting only goes so far when a modality is not noisy but simply absent, which happens constantly in real multimodal graphs. This is where the GAN comes in, and the specific design choice worth noting is what conditions the generator. Instead of starting from noise alone, the way a typical GAN generator would, AFME concatenates the entity’s structural embedding with a random noise vector, on the argument that the graph’s own relational structure carries information a purely visual or textual generator has no access to.

\( e_{mod} = e_s \oplus z \)
\( h_m^{(0)} = \sigma\big(W_d(e_{mod} \odot \tilde{h}_m) + b_d\big) \)

That initial guided feature then goes through a multi layer self attention block to sharpen internal consistency within the modality, followed by an inter modal interaction step that lets each modality’s representation absorb something from every other modality, weighted by a learned attention score between them.

$$ h_m^{(1)} = \text{softmax}\!\left(\frac{(h_m^{(0)}W_Q)(h_m^{(0)}W_K)^{\top}}{\sqrt{d}}\right)(h_m^{(0)}W_V) $$

A discriminator then checks the generated feature against a real one, concatenating both and scoring their alignment, while training uses the Wasserstein GAN objective with a gradient penalty term borrowed from WGAN-GP rather than the original minimax loss, specifically to avoid the vanishing gradients and mode collapse that plain adversarial training is known to produce.

$$ L_{gp} = \lambda \, \mathbb{E}_{\tilde{e}_m \sim P_{\tilde{e}_m}}\left[\big(\lVert \nabla_{\tilde{e}_m} D(\tilde{e}_m) \rVert_2 – 1\big)^2\right] $$

The full training objective sums the knowledge graph completion loss, a margin based negative sampling loss over the fused embeddings, together with the adversarial loss and the gradient penalty term, weighted by two coefficients the authors tune from a small grid.

$$ L = L_{kgc} + \lambda_1 L_{adv} + \lambda_2 L_{gp} $$

What the benchmarks actually show

The evaluation runs across four datasets that differ a lot in shape. MKG-W and MKG-Y, drawn from Wikidata and YAGO, cover text and images only, at roughly fifteen thousand entities each. TIVA and KVC16K go further, adding video and audio, with KVC16K in particular covering four full modalities across sixteen thousand entities pulled from a short video encyclopedia. That range matters because a fusion method that only works when two clean modalities are available is not proving much about the harder, messier case.

MethodMKG-W MRRMKG-Y MRRTIVA MRRKVC16K MRR
TransE29.1930.7383.858.54
RotatE33.6734.9584.5914.33
IKRL32.3633.2267.7111.11
MMKRL30.1036.8185.038.78
KBGAN29.4729.7185.4413.72
MMRNS35.0335.9383.1213.31
NATIVE (strongest baseline)36.4838.8191.6115.41
AFME (proposed)37.0939.4192.4415.83

The gain over NATIVE, the previous best performing method and itself an adversarial adaptive fusion approach, is the number that matters most here, since it is the comparison against the toughest opponent rather than against older translation based models. It comes out to about one point seven percent MRR on MKG-W, one point six percent on MKG-Y, under one percent on TIVA, and two point seven percent on KVC16K. The KVC16K result stands out because that is the dataset with all four modalities and the lowest baseline scores across the board, which is roughly where you would expect a better denoising and enhancement strategy to matter most, since there is more noisy or missing modality information for it to clean up.

A second comparison the authors highlight is more specific and, to their credit, more honest about what it does and does not prove. AFME still beats IKRL and VBKGC, two methods that use only structural, text, and image information, even when AFME is run on the same limited modality set. That rules out a mundane explanation for the headline numbers, namely that AFME is simply winning because it has access to more modalities than its older comparison points.

Taking the two modules apart

The ablation study removes each module in turn on the MKG-Y and TIVA datasets. Removing MoIFu and replacing relationship gated, confidence weighted fusion with plain concatenation costs the model measurable MRR on both datasets. Removing MoIEn and letting the generator produce features with no self attention refinement and no structural guidance costs even more, which is a useful signal about where the bulk of the performance is actually coming from.

DatasetVariantMRRHits@1
MKG-YFull AFME39.4135.45
MKG-YWithout MoIFu38.8734.70
MKG-YWithout MoIEn37.3333.35
TIVAFull AFME92.4291.58
TIVAWithout MoIFu91.9990.90
TIVAWithout MoIEn90.7589.55

On both datasets, losing MoIEn hurts more than losing MoIFu, and the gap between the two removed variants is not tiny. That fits a fairly intuitive story. Fusing modalities better still leaves you stuck if a modality is simply missing or badly degraded, while the enhancement module is the piece actually responsible for producing a usable substitute in that case. Neither module carries the whole result on its own, and the paper’s own framing of a collaborative interaction between the two is a reasonable read of the numbers rather than an overstatement.

significantly improves multi modal feature utilization and knowledge reasoning capabilities Wang, Liu, Liu, Weng, and Chaomurilige, Neural Networks, 2025

A generalization test that is more convincing than another leaderboard row

The most persuasive experiment in the paper is arguably not the main results table at all. The authors took the MoIEn module out of AFME entirely and bolted it onto three unrelated baseline architectures, TBKGC, MMKRL, and QEB, then measured whether those baselines improved on TIVA. All three did. That is a different and frankly more useful kind of evidence than beating a leaderboard, because it suggests the enhancement mechanism is not just tuned to work well specifically inside AFME’s own architecture, it is doing something generally useful for modality completion that a different model’s fusion strategy can also benefit from.

Why this result matters more than it looks. A module that only helps its own parent model could just be compensating for a weakness elsewhere in that same model. A module that helps three unrelated models is evidence the mechanism itself, not the surrounding architecture, is doing the work.

The hyperparameters that actually moved the needle

Three sensitivity experiments are worth knowing if you are trying to reproduce or adapt this work. The denoising intensity parameter, which controls how strongly the relationship gate suppresses raw modality features versus purified ones, degrades performance at both extremes, too little denoising leaves noise in, too much strips out genuinely useful detail along with it. The margin value in the negative sampling loss peaked at four within a tested range of two to six, balancing discriminative power against training stability. The number of self attention layers in the enhancement module peaked at four as well, with both fewer layers failing to capture cross modal dependencies and more layers adding computational cost without a matching gain. None of these are surprising shapes for a sensitivity curve to take, but having the actual peak values reported saves anyone reproducing this architecture a real amount of tuning time.

Where the paper is candid about its limits

The authors close by naming two specific gaps rather than presenting the result as finished.

  • There is no dedicated analysis of how AFME performs on rare relation types or long tail entities, and the four benchmark datasets mostly cover common entities and well represented relations, so it is genuinely unclear how the method holds up on the sparse, low frequency parts of a real world graph.
  • The evaluation stays within general purpose multimodal knowledge graphs. The authors explicitly propose testing on harder domains such as medical visual question answering and social network graphs as future work, which means claims about AFME’s behavior in those settings are not yet supported by anything in this paper.

It is also worth noting plainly what the paper does not claim. It does not claim to beat every baseline on every single metric on every dataset, the Hits@10 columns on MKG-W and MKG-Y are left blank in the source table rather than reported, and the improvement margins range from under one percent to nearly three percent depending on the benchmark, which is a real but not uniform advantage.

A reference implementation of the core mechanism

No code repository is linked from the paper, so the block below is an independent PyTorch implementation of the mechanism it describes, written to make the moving pieces concrete. It covers the relationship gated denoising step, the confidence weighted dynamic fusion softmax, a structure guided generator with self attention and inter modal interaction, a discriminator, the knowledge graph completion loss, the WGAN style adversarial loss with a gradient penalty, a training loop, an evaluation function, and a smoke test on random data.

# afme_demo.py # Independent reference implementation of the AFME mechanism. # Not the authors’ code, this reproduces the equations in the paper. import torch import torch.nn as nn import torch.nn.functional as F class RelationshipDenoiser(nn.Module): “””Gates each modality feature using the relationship embedding, matching gate_m = sigmoid(Wg . (h_m elementwise r) + bg).””” def __init__(self, dim): super().__init__() self.gate = nn.Linear(dim, dim) def forward(self, h_m, r): interaction = h_m * r gate = torch.sigmoid(self.gate(interaction)) return gate * h_m class DynamicWeightAllocation(nn.Module): “””Computes per modality fusion weights from an L2 confidence score and a relationship specific temperature, then a softmax over modalities plus the structural embedding.””” def __init__(self, dim): super().__init__() self.U = nn.Parameter(torch.randn(dim) * 0.02) self.temperature_head = nn.Linear(dim, 1) def forward(self, clean_features, relation_embedding): # clean_features is a dict, modality name to tensor of shape (dim,) tau = torch.sigmoid(self.temperature_head(relation_embedding)).clamp(min=1e-3) scores = {} for name, feat in clean_features.items(): confidence = torch.norm(feat, p=2) scores[name] = confidence * (self.U * torch.tanh(feat)).sum() / tau stacked = torch.stack(list(scores.values())) weights = F.softmax(stacked, dim=0) return {name: w for name, w in zip(scores.keys(), weights)} class Generator(nn.Module): “””Structure guided generator, initial gating, self attention refinement, then inter modal interaction across modalities.””” def __init__(self, dim, noise_dim=64, heads=4): super().__init__() self.init_proj = nn.Linear(dim + noise_dim, dim) self.attn = nn.MultiheadAttention(dim, heads, batch_first=True) def forward(self, structural_embedding, noise, clean_features): e_mod = torch.cat([structural_embedding, noise], dim=-1) e_mod = torch.sigmoid(self.init_proj(e_mod)) h0 = {name: e_mod * feat for name, feat in clean_features.items()} stacked = torch.stack(list(h0.values())).unsqueeze(0) # (1, num_modalities, dim) h1, _ = self.attn(stacked, stacked, stacked) h1 = h1.squeeze(0) # inter modal interaction, each modality attends to every other one sim = h1 @ h1.transpose(0, 1) / (h1.shape[-1] ** 0.5) sim.fill_diagonal_(float(“-inf”)) cross_weights = F.softmax(sim, dim=-1) h_final = cross_weights @ h1 return {name: h_final[i] for i, name in enumerate(h0.keys())} class Discriminator(nn.Module): “””Scores how well a generated modality feature matches a real one.””” def __init__(self, dim): super().__init__() self.net = nn.Sequential( nn.Linear(dim * 2, dim), nn.LeakyReLU(0.2), nn.Linear(dim, 1), ) def forward(self, generated, real): return self.net(torch.cat([generated, real], dim=-1)) def kgc_score(h_joint, r, t_joint): return -torch.norm(h_joint + r – t_joint, p=2) def kgc_loss(pos_scores, neg_scores, margin=4.0): pos_term = F.logsigmoid(margin + pos_scores).mean() neg_term = F.logsigmoid(-margin – neg_scores).mean() return -(pos_term + neg_term) def gradient_penalty(discriminator, real, fake, other): alpha = torch.rand(1) interpolated = (alpha * real + (1 – alpha) * fake).requires_grad_(True) score = discriminator(interpolated, other) grads = torch.autograd.grad( outputs=score, inputs=interpolated, grad_outputs=torch.ones_like(score), create_graph=True, retain_graph=True, )[0] return ((grads.norm(2) – 1) ** 2) def train_step(gen, disc, denoiser, weigher, opt_g, opt_d, structural_embedding, relation_embedding, real_features, noise_dim=64, lambda_adv=1e-3, lambda_gp=1e-4): clean = {name: denoiser(feat, relation_embedding) for name, feat in real_features.items()} weights = weigher(clean, relation_embedding) noise = torch.randn(noise_dim) generated = gen(structural_embedding, noise, clean) # discriminator step opt_d.zero_grad() d_loss = 0.0 for name in real_features: real_score = disc(real_features[name].detach(), real_features[name].detach()) fake_score = disc(generated[name].detach(), real_features[name].detach()) gp = gradient_penalty(disc, real_features[name], generated[name].detach(), real_features[name]) d_loss = d_loss + (fake_score.mean() – real_score.mean()) + lambda_gp * gp d_loss.backward() opt_d.step() # generator plus fusion step opt_g.zero_grad() joint = sum(weights[name] * generated[name] for name in generated) pos_score = kgc_score(joint, relation_embedding, joint.roll(1)) neg_score = kgc_score(joint.roll(2), relation_embedding, joint) loss_kgc = kgc_loss(pos_score, neg_score) adv_terms = [disc(generated[name], real_features[name]).mean() for name in generated] loss_adv = -torch.stack(adv_terms).mean() total_loss = loss_kgc + lambda_adv * loss_adv total_loss.backward() opt_g.step() return total_loss.item(), d_loss.item() if __name__ == “__main__”: # Smoke test on random data, confirming shapes and a full step run. torch.manual_seed(0) dim = 32 modalities = [“text”, “image”, “video”, “audio”] denoiser = RelationshipDenoiser(dim) weigher = DynamicWeightAllocation(dim) gen = Generator(dim) disc = Discriminator(dim) opt_g = torch.optim.Adam(list(gen.parameters()) + list(denoiser.parameters()) + list(weigher.parameters()), lr=1e-4) opt_d = torch.optim.Adam(disc.parameters(), lr=1e-4) structural_embedding = torch.randn(dim) relation_embedding = torch.randn(dim) real_features = {name: torch.randn(dim) for name in modalities} for step in range(5): g_loss, d_loss = train_step( gen, disc, denoiser, weigher, opt_g, opt_d, structural_embedding, relation_embedding, real_features, ) print(f”step {step}, generator loss {g_loss:.4f}, discriminator loss {d_loss:.4f}”)

Two caveats on the code above, in the interest of accuracy. The full paper batches this process across thousands of entities and triples at once, while the smoke test above runs on a single entity’s worth of features for clarity. And the inter modal interaction step here assumes every modality is present, whereas a real deployment needs to handle the case where a modality is absent entirely for a given entity, which the paper’s enhancement module is specifically designed to address.

What this adds up to for anyone building on multimodal graphs

The most transferable idea in this paper is not really the GAN, GANs have been tried in knowledge graph completion before. It is the decision to condition both the denoising step and the fusion weighting on the relationship being predicted, rather than treating modality quality as a fixed, entity level property. A photograph that is highly informative for a physical location relation may be nearly useless for a causal relation between two historical events, and a fusion mechanism that cannot tell the difference is leaving performance on the table regardless of how sophisticated its generator is.

The generalization experiment, where the enhancement module improved three unrelated baseline architectures, is the strongest piece of evidence in the paper that this idea is not narrowly tied to AFME’s specific scoring function. That makes MoIEn a reasonable candidate for anyone maintaining an existing multimodal completion pipeline who wants an incremental upgrade rather than a full architecture rewrite.

Conclusion

AFME’s contribution is a genuinely two part answer to a problem that most prior work treated as one part. Cleaning and weighting modality features by relationship context, through MoIFu, and regenerating weak or missing modality features using the graph’s own structural signal, through MoIEn, turn out to be complementary rather than redundant, based on both the ablation study and the fact that losing either one costs real performance on MKG-Y and TIVA.

The conceptual shift worth remembering is treating modality quality as something that depends on the relationship being asked about, not as a fixed property of the modality itself. Once fusion weights are allowed to move per relationship rather than staying static across the whole graph, a framework gets room to be decisive where one modality clearly dominates and balanced where several genuinely contribute, which a single global weighting scheme cannot do.

Where this could go next follows fairly directly from the paper’s own stated gaps. Testing on graphs with heavier long tail entity and rare relation coverage, and on harder domains such as medical visual question answering or social network reasoning, would tell us whether the relationship conditioned gating and the structure guided generator hold up outside the relatively well populated benchmarks used here.

The honest limitations are worth repeating rather than glossing over. Blank Hits@10 entries for two of the four benchmark datasets in the source table, a generalization experiment run on a single dataset with three baseline models, and an explicit admission that rare entities and relation types were not specifically evaluated, all mark where the next version of this framework needs to go rather than undercutting what has already been shown.

What stays with you after reading this paper is not the percentage improvement over NATIVE. It is the gating equation itself, the quiet idea that a model can be taught to ask, for this specific relationship, how much do I actually trust this specific piece of evidence, before it ever tries to fuse anything together.

Frequently asked questions

What problem is AFME actually solving in multimodal knowledge graphs

It addresses the fact that different modalities attached to the same entity, text, images, video, and audio, vary a great deal in quality and completeness, and that simply concatenating or statically weighting them fails to capture which modality actually matters for a given relationship, and performs poorly when a modality is noisy or missing outright.

How is the relationship driven denoising step different from ordinary feature cleaning

Instead of cleaning a modality feature in isolation, the gate is conditioned on the embedding of the specific relationship being predicted, so the same modality feature can be treated as mostly useful for one relationship and mostly irrelevant for another, rather than receiving one fixed cleaning treatment regardless of context.

Why does the generator use the structural embedding instead of starting from pure noise

A generator that starts from noise alone has no connection to the entity’s actual position in the knowledge graph, which tends to produce features disconnected from the graph’s real semantics. Conditioning on the structural embedding gives the generator a grounded signal to work from when filling in a weak or missing modality feature.

Does AFME actually beat the strongest prior method or just older baselines

It beats NATIVE, described in the paper as the strongest prior baseline, on mean reciprocal rank across all four tested datasets, with the improvement ranging from under one percent to nearly three percent depending on the dataset, alongside gains on Hits@1 and, where reported, Hits@10.

Can the enhancement module be used with a different base model

The paper tests exactly this by attaching the MoIEn module to three other baseline architectures, TBKGC, MMKRL, and QEB, and reports performance improvements on the TIVA dataset for all three, which is evidence the module generalizes beyond AFME’s own scoring function.

Is there published code for AFME

[CODE REPOSITORY NEEDED, add a link once the authors release code or confirm none is planned]. The PyTorch example in this article is an independent illustration of the mechanism the paper describes, not the authors’ own implementation.

Read the full paper for the complete derivations, the sensitivity analysis, and the dataset statistics referenced above.

Read the paper (DOI)

Related reading on this site

Wang, Z., Liu, X., Liu, Z., Weng, Y., and Chaomurilige. A link prediction method for multi modal knowledge graphs based on Adaptive Fusion and Modality Information Enhancement. Neural Networks, Vol. 191, Article 107771, 2025. https://doi.org/10.1016/j.neunet.2025.107771

This analysis is based on the published paper and an independent evaluation of its claims.

Leave a Comment

Your email address will not be published. Required fields are marked *