ERC Starting Grant · 2025
Establishing a Spatio-Temporal Language for Scene Representation
We have recently experienced a boost in AI through the performance of ChatGPT that has moved from a purely scientific endeavor to deployment in various businesses and real-world applications. Also in computer vision, we have seen very large progress that was enabled by scaling to foundation models that can be trained in an unsupervised fashion on enormous amounts of data. However, looking at these models in detail reveals that they are in fact truly terrible in understanding the structure of scenes and the properties of objects. Inferring such properties without explicit labeling requires visual observations over time, as well as motion and interaction. The current state-of-the-art models…
From the public funding record at EU CORDIS. Describes the funded project, not the reviews below.