HyperAIHyperAI

Command Palette

Search for a command to run...

Hebrew-CLIP Hebrew Dataset

Date

Organization

NVIDIA

License

nvidia-license

Hebrew-CLIP is a dataset released by NVIDIA in 2024 for training Hebrew visual-language models, aimed at promoting the application of visual-language models such as CLIP in Hebrew scenarios.

The dataset contains approximately 7.78 million Hebrew image descriptions, primarily composed of machine-translated data from DataComp-1B and multilingual data from LAION-5B. The data is stored in Parquet format and does not include actual images, only providing text descriptions and corresponding image embedding references, making it suitable for training and evaluating multimodal models.

Dataset Composition

The dataset consists of two Parquet files, each mainly containing the following fields:

  • key: A unique identifier for the description.
  • heb_caption: The Hebrew description text.
  • file_name: The corresponding image embedding file name.
  • file_index: The index position of the embedding within the file.

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp