HyperAIHyperAI

Command Palette

Search for a command to run...

Updesh Indic Synthetic Text Dataset

Date

a year ago

Size

16.09 GB

Organization

Microsoft

Updesh is an Indian language synthetic text dataset released by Microsoft in 2025 to facilitate post-training of Large Language Models (LLMs) for Indian languages. The dataset contains 6,800,000 inference data and 2,100,000 generated data in the following languages: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Nepali, Odia, Punjabi, Tamil, Telugu, and Urdu.

Citation

@misc{chitale2026updeshsynthesizinggroundedinstruction, title={UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic}, author={Pranjal A. Chitale and Varun Gumma and Sanchit Ahuja and Prashant Kodali and Manan Uppadhyay and Deepthi Sudharsan and Sunayana Sitaram}, year={2026}, eprint={2509.21294}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2509.21294}, }

Updesh_beta.torrent
Seeding 2Downloading 0Completed 97Total Downloads 153
  • Updesh_beta/
    • README.md
      1.2 KB
    • README.txt
      2.4 KB
      • data/
        • Updesh_beta.zip
          16.09 GB

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp