Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

Datasets:
beatsprom
/
multimodal-vision-language-video-models-2026

Tasks:
Feature Extraction
Question Answering
Text Generation
Modalities:
Tabular
Text
Time-series
Formats:
parquet
Languages:
English
Size:
< 1K
Tags:
multimodal
vision-language
vlm
video-generation
diffusion-transformer
world-models
Libraries:
Datasets
pandas
Polars
License:
Dataset card Data Studio Files Files and versions
xet
Community
1
multimodal-vision-language-video-models-2026
770 kB
Ctrl+K
Ctrl+K
  • 1 contributor
History: 3 commits
beatsprom's picture
beatsprom
Update README.md
7fbcf9d verified about 2 months ago
  • .gitattributes
    2.5 kB
    initial commit about 2 months ago
  • MULTIMODAL_VISION_LANGUAGE_VIDEO_FOUNDATION_MODELS_2026_30_SAMPLE.json
    557 kB
    Upload 3 files about 2 months ago
  • MULTIMODAL_VISION_LANGUAGE_VIDEO_FOUNDATION_MODELS_2026_30_SAMPLE.parquet
    207 kB
    xet
    Upload 3 files about 2 months ago
  • README.md
    2.46 kB
    Update README.md about 2 months ago
  • quickstart_multimodal_vector_search.py
    1.37 kB
    Upload 3 files about 2 months ago