InternVid
A large-scale video-text dataset for multimodal understanding and generation, part of the InternVideo foundation model project.
Visit datasetA curated, hand-verified collection of open-source datasets for training and benchmarking multimodal generative AI models — spanning video-text, image-caption, visual question answering, emotion, and RGB-D data.
A large-scale video-text dataset for multimodal understanding and generation, part of the InternVideo foundation model project.
Visit datasetA standard benchmark dataset for sentence-based image description, widely used for training and evaluating image-captioning models.
Visit datasetStudies the multimodal interplay between the presence of stress and expressions of affect, with annotated recordings and baseline models.
Visit datasetA Visual Question Answering dataset pairing images with natural-language questions and answers, for training multimodal reasoning models.
Visit datasetA question-answering benchmark for artificial social intelligence, testing a model's ability to reason about human social behavior.
Visit datasetA large-scale, multi-view RGB-D object dataset from the University of Washington, widely cited for 3D object recognition research.
Visit datasetA curated collection of image and video datasets for generative AI and multimodal models, requiring at least images and corresponding captions.
Visit dataset