HyperAIHyperAI

Command Palette

Search for a command to run...

CantTalkAboutThis Topic Control Dataset

Date

Organization

NVIDIA

Paper URL

2404.03820

License

CC BY 4.0

The "Can't Talk About This" topic control dataset, released by NVIDIA in 2024 for training language models to maintain topical focus in task-oriented dialogues, is detailed in the paper titled «CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues”, aimed at enhancing the model’s robustness against distractions and improving alignment in instruction-following and safety tasks.

This dataset comprises 1,080 synthetic conversations spanning nine domains: health, banking, travel, education, finance, insurance, law, real estate, and computer troubleshooting.
Each conversation includes distracting elements designed to test and improve the model’s performance when handling irrelevant topics.
The dataset contains no personally identifiable or sensitive information and is entirely generated using synthetic data.

Dataset Composition

The dataset consists of the following fields:

  • domain: The category to which the conversation belongs.
  • scenario: The specific context or task being discussed.
  • system_instruction: Dialogue strategies provided to the model, typically comprising complex instructions regarding permissible and impermissible topics.
  • conversation: The full dialogue, including both the primary topic and distraction elements.
  • distractors: A list of distraction items, encompassing bot responses within the dialogue as well as user inputs that should be included by users responding to those bot replies.
  • conversation_with_distractors: The complete dialogue incorporating all distraction elements.

The dataset is divided into two subsets: training (Mixtral) and testing (Human Test Set). The training set was generated using the Mixtral-8x7B-Instruct model, while the test set features more complex and realistic human-labeled distractions for evaluating model performance.

Citation```bibtex

@inproceedings{sreedhar2024canttalkaboutthis, title={CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues}, author={Sreedhar, Makesh and Rebedea, Traian and Ghosh, Shaona and Zeng, Jiaqi and Parisien, Christopher}, booktitle={Findings of the Association for Computational Linguistics: EMNLP 2024}, pages={12232--12252}, year={2024}, organization={Association for Computational Linguistics} }

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp