ERC Starting Grant · 2023
Graphs without Labels: Multimodal Structure Learning without Human Supervision
Multimodal learning focuses on training models with data in more than one modality, such as videos capturing visual and audio information or documents containing image and text. Current approaches use such data to train large-scale deep learning models without human supervision by sampling pair-wise data e.g., an image-text pair from a website and train the network e.g. to identify matching vs. not matching pairs to learn better representations. We argue that multimodal learning can do more: by combining information from different sources, multimodal models capture cross-modal semantic entities, and as most multimodal documents are a collection of connected modalities and topics, multimodal…
From the public funding record at EU CORDIS. Describes the funded project, not the reviews below.