Search for a command to run...
DECEPEVAL: A BENCHMARK FOR EVALUATING DE-CEPTION IN LLM AGENTS
What can I help you find?
Datasets, papers, notebooks and GPUs