Command Palette
Search for a command to run...
Pinocchio Pinocchio Factual Knowledge Evaluation Dataset
Date
Size
Publish URL

The Pinocchio dataset was jointly created by researchers from Tsinghua University, University of Illinois at Chicago, and University of Cambridge. Its purpose is to comprehensively evaluate the performance of large language models (LLMs) in factual knowledge storage and reasoning capabilities. **This dataset covers 20,000 diverse factual questions covering different sources, timelines, domains, regions, and languages.**The dataset contains 7 different tasks to test LLMs’ ability to reason over multiple facts, handle structured and unstructured knowledge, identify subtle factual differences, and resist adversarial examples. Pinocchio provides researchers with a powerful tool to understand the capabilities of models at multiple levels while pushing the boundaries of LLMs’ ability to advance factual knowledge.
Citation
Please use the below BibTeX entry to cite this dataset:
@article{HuPino2023,
author = {Xuming Hu and
Junzhe Chen and
Xiaochuan Li and
Yufei Guo and
Lijie Wen and
Philip S. Yu and
Zhijiang Guo},
title = {Towards Understanding Factual Knowledge of Large Language Models},
journal = {12th International Conference on Learning Representations, {ICLR} 2024, Messe Wien Exhibition and Congress Center, Vienna Austria
May 7th, 2024 to May 11th, 2024},
volume = {2024},
year = {2024},
url = {https://doi.org/10.48550/arXiv.2310.05177},
doi = {10.48550/ARXIV.2310.05177},
eprinttype = {arXiv},
eprint = {2310.05177},
timestamp = {Fri, 20 Oct 2023 12:04:38 +0200},
biburl = {https://dblp.org/rec/journals/corr/abs-2310-05177.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.