The authors describe Latent-Attention Masked Autoencoders, or LAMAE, a model pretrained on more than 1.2 million MIMIC-IV hospital stays. Instead of combining heart tests only after training, it shares information between them during pretraining through a latent-attention module, which the authors say also lets it cope when a modality is missing.

The reported comparisons are against modality-specific pretraining and against contrastive and vision-language baselines, on hospital-stay tasks including in-hospital mortality, ICD-10 and DRG coding, and length of stay. The authors say the advantage holds even when only one modality is available at test time. All of this is the authors' own evaluation of their own model, on one dataset, in a paper posted as a preprint.