Skip to content

Architecture Atlas

Five ways to build a model from several modalities, drawn the same way: inputs on the left, what the model produces on the right. Each card says when the design fits, what it costs, an example from materials science and whether MEIDNet implements it. At the end, the advisor recommends one for your data and your goal.

New to the topic? Start with Learn multimodality.

Modality A Modality B concatenateab one model prediction

Early fusion

Not in MEIDNet

Idea. Join the modalities at the input: turn each into features, concatenate them into one vector, train one model on it.

When it fits. Every material has every modality; the features have a fixed length and similar scales; you want a quick baseline.

Strengths. Simple. The model can use interactions between features from the first layer. Weaknesses. A missing modality breaks the input. Very different shapes (a graph and a curve) are hard to concatenate. A large modality can drown a small one.

Materials example. Composition descriptors, the processing temperature and a few peak positions from an XRD pattern, concatenated and given to a gradient-boosted regressor.

References: Snoek et al. 2005; Baltrušaitis et al. 2019.

Modality A Modality B model A model B guess A guess B combine average or vote

Late fusion

Not in MEIDNet

Idea. One model per modality, each makes its own prediction; the predictions are combined by averaging, voting or a small model on top.

When it fits. Modalities are often missing, models for each already exist, or robustness matters more than squeezing out the last bit of accuracy.

Strengths. Modular. Tolerates a missing modality. Each model can be trained on its own, larger dataset. Weaknesses. Cannot learn how the modalities interact. There is no shared representation, so no translation from one modality to another and no inverse design.

Materials example. A structure-based graph network and a spectrum-based convolutional network each predict a phase label; the final label averages their probabilities.

References: Snoek et al. 2005; Baltrušaitis et al. 2019.

Structure Properties EGNN MLP shared space (128-d) align average decoders structure andproperties

Shared latent space (joint and coordinated)

Core idea used by MEIDNet Supported

Idea. One encoder per modality maps into a common space. In a coordinated representation the latents stay separate but are trained to agree; in a joint representation they are merged into one vector that the decoders read. MEIDNet does both: the structure latent and the property latent are aligned by contrastive learning, then averaged into the joint latent; the decoders read the joint latent and the property latent alone.

When it fits. You want to translate between modalities (properties → structure for inverse design), retrieve one modality from another, or search a space that both understand. You have paired examples.

Strengths. Translation in both directions, retrieval, and one space to optimise in. Each modality keeps the encoder that suits its shape. Weaknesses. Needs paired data. The space is only as good as the alignment. Several losses must be balanced.

Materials example. MEIDNet: crystal structures and DFT properties of 18,928 cubic perovskites, searched for compositions with a target band gap and formation enthalpy (the paper's experiment).

References: Ngiam et al. 2011; Baltrušaitis et al. 2019; Babu et al. 2026; Moro, Loh et al. 2025.

A tokens (atoms) B tokens (spectrum) querieskeys, values cross-attention fused tokens→ prediction

Cross-attention and transformer fusion

Not in MEIDNet

Idea. Each modality becomes a sequence of tokens (atoms, segments of a spectrum, words). Attention lets every token of one modality look up the tokens of the other that matter to it, layer after layer.

When it fits. The modalities have parts that interact in detail and are not aligned in advance; the dataset is large.

Strengths. Learns which parts relate to which, and handles sequences of different lengths without manual alignment. Weaknesses. Needs much data and compute; the cost grows with the product of the two sequence lengths; harder to interpret than a single shared vector.

Materials example. Spectrum segments attend to the atoms of the structure, so the model learns which sites shape which spectral features.

References: Tsai et al. 2019 (Multimodal Transformer); Lu et al. 2019 (ViLBERT).

properties of the batch structures true pairs:pulled together other pairs:pushed apart

Contrastive learning

Used by MEIDNet Supported

Idea. A training objective rather than a wiring. In a batch of paired examples, each example's two embeddings must be more similar to each other than to any other example's. It is how the latents of a shared space are coordinated.

When it fits. You have pairs (a structure and its properties, a pattern and its structure) and want a space in which nearest neighbours are meaningful across modalities.

Strengths. Learns from the pairing alone; gives retrieval for free; scales to large datasets. Weaknesses. Needs batches with enough other examples; sensitive to the temperature; two genuinely similar materials in one batch are still pushed apart.

In MEIDNet. Symmetric InfoNCE with temperature 0.01 and weight 5, switched on gradually over the warm-up (the paper's curriculum). The formula and a playground.

References: van den Oord et al. 2018 (InfoNCE); Radford et al. 2021 (CLIP).

Side by side

early fusion late fusion shared latent cross-attention contrastive (objective)
where the modalities meet input features predictions a common latent space inside the network, token by token the latent space, through the loss
every modality needed for every sample yes no paired examples for training usually paired examples
learns interactions between modalities yes no yes yes, in detail for whole samples
translates one modality into another no no yes with a decoder retrieval
data needed small to medium small per model medium large medium to large
in MEIDNet no no yes, the core no yes

Beyond these five. Conditional generative models such as diffusion (MatterGen) or variational autoencoders (CDVAE) also translate from properties to structures, and they generate free atomic arrangements, which MEIDNet does not. MEIDNet instead decodes compositions onto a prototype family and checks them against chemical rules.

Which architecture should I use?

Tick the data you have and choose what you want to do. The recommendation says why, what it costs and whether MEIDNet can do it today.

A note on words

"Early fusion" in the MEIDNet paper names the model in which the structure latent and the property latent are averaged into one joint latent: z_joint = (z_c + z_p) / 2 in meidnet/model.py. In the survey literature, early fusion usually means joining raw features at the input, as in the first card; averaging two learned latents, each from its own encoder, is intermediate or model-level fusion, the shared-latent card. Both names describe the same code. The checkpoints shipped here are the paper's early-fusion models, and the code implements only this fusion: there is no late-fusion option in MEIDNet.

"Curriculum" in the paper is the contrastive warm-up: the alignment weight grows from zero to its full value over training.contrastive_warmup_epochs, so that the decoders learn first.

References

  • C. G. M. Snoek, M. Worring, A. W. M. Smeulders, "Early versus late fusion in semantic video analysis", ACM Multimedia (2005). doi:10.1145/1101149.1101236
  • T. Baltrušaitis, C. Ahuja, L.-P. Morency, "Multimodal machine learning: a survey and taxonomy", IEEE TPAMI 41, 423–443 (2019). doi:10.1109/TPAMI.2018.2798607
  • J. Ngiam et al., "Multimodal deep learning", ICML (2011).
  • Y.-H. H. Tsai et al., "Multimodal Transformer for unaligned multimodal language sequences", ACL (2019). doi:10.18653/v1/P19-1656
  • J. Lu, D. Batra, D. Parikh, S. Lee, "ViLBERT: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks", NeurIPS (2019). arXiv:1908.02265
  • A. van den Oord, Y. Li, O. Vinyals, "Representation learning with contrastive predictive coding" (2018). arXiv:1807.03748
  • A. Radford et al., "Learning transferable visual models from natural language supervision", ICML (2021). arXiv:2103.00020
  • R. E. A. Goodall, A. A. Lee, "Predicting materials properties without crystal structure: deep representation learning from stoichiometry" (Roost), Nat. Commun. 11, 6280 (2020). doi:10.1038/s41467-020-19964-7
  • W. B. Park et al., "Classification of crystal structure using a convolutional neural network", IUCrJ 4, 486–494 (2017). doi:10.1107/S205225251700714X
  • C. Zeni et al., "A generative model for inorganic materials design" (MatterGen), Nature 639, 624–632 (2025). doi:10.1038/s41586-025-08628-5
  • T. Xie et al., "Crystal diffusion variational autoencoder for periodic material generation" (CDVAE), ICLR (2022). arXiv:2110.06197
  • V. Moro, C. Loh et al., "Multimodal foundation models for material property prediction and discovery", Newton (2025). doi:10.1016/j.newton.2025.100016
  • A. Babu, R. Almeida Gouvêa, P. Vandergheynst, G.-M. Rignanese, "MEIDNet: Multimodal generative AI framework for inverse materials design", npj Comput. Mater. (2026). doi:10.1038/s41524-026-02153-3