HyperAIHyperAI

Command Palette

Search for a command to run...

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Jason Wei Xuezhi Wang Dale Schuurmans Maarten Bosma Brian Ichter Fei Xia Ed Chi Quoc Le Denny Zhou

Abstract

We explore how generating a chain of thought -- a series of intermediate reasoning steps -- significantly improves the ability of large language models to perform complex reasoning. In particular, we show how such reasoning abilities emerge naturally in sufficiently large language models via a simple method called chain of thought prompting, where a few chain of thought demonstrations are provided as exemplars in prompting. Experiments on three large language models show that chain of thought prompting improves performance on a range of arithmetic, commonsense, and symbolic reasoning tasks. The empirical gains can be striking. For instance, prompting a 540B-parameter language model with just eight chain of thought exemplars achieves state of the art accuracy on the GSM8K benchmark of math word problems, surpassing even finetuned GPT-3 with a verifier.

One-sentence Summary

Google Research's Brain Team proposes chain-of-thought prompting, a method that supplies a few exemplars of intermediate reasoning steps in prompts, and shows that it significantly improves large language models' performance on complex arithmetic, commonsense, and symbolic reasoning tasks, achieving state-of-the-art accuracy on the GSM8K benchmark with PaLM 540B using only eight exemplars and surpassing even finetuned GPT-3 with a verifier.

Key Contributions

  • Chain-of-thought prompting elicits step-by-step reasoning by providing a few exemplars with reasoning chains in the prompt, without any model fine-tuning.
  • On arithmetic, commonsense, and symbolic reasoning tasks, the method yields substantial improvements. Prompting PaLM 540B with eight chain-of-thought exemplars achieves state-of-the-art accuracy on GSM8K, surpassing a fine-tuned GPT-3 with a verifier.
  • Chain-of-thought reasoning emerges with model scale, transforming flat scaling curves into steeply increasing ones, and enables out-of-distribution generalization to longer sequence lengths in symbolic reasoning.

Introduction

Language models have improved dramatically with scale, but scaling alone is insufficient for complex reasoning tasks like arithmetic, commonsense, and symbolic reasoning. Prior work attempted to inject reasoning either by training models on natural language rationales, which is costly to annotate, or by using few-shot prompting, but standard prompting performs poorly on reasoning problems and does not reliably improve with model size. The authors propose chain-of-thought prompting, a method that provides a few exemplars containing intermediate reasoning steps alongside inputs and outputs, enabling large language models to perform multi-step reasoning without fine-tuning. This approach combines the strengths of rationale-based reasoning and in-context learning, substantially boosting performance across a range of reasoning benchmarks.

Method

The authors propose chain of thought prompting to enhance the reasoning capabilities of large language models. Rather than relying exclusively on scaling model parameters, this approach leverages few-shot prompting augmented with intermediate reasoning steps. When humans solve complex multi-step problems, they naturally decompose the task into intermediate steps before arriving at a final conclusion. The authors aim to replicate this cognitive process by providing the model with exemplars that include a coherent series of intermediate reasoning steps.

As shown in the figure below, standard prompting typically requires the model to output the final answer directly based on the input question. In contrast, chain of thought prompting augments each few-shot exemplar with a detailed reasoning sequence that mimics a step-by-step thought process.

This methodological framework offers several key advantages for facilitating reasoning. First, it allows the model to decompose multi-step problems into intermediate stages, effectively allocating additional computation to tasks that require more extensive reasoning. Second, the generated chain of thought provides an interpretable window into the model behavior, enabling developers to trace the reasoning path and identify where logical errors occur. Finally, this technique is highly versatile and can be readily elicited in sufficiently large off-the-shelf language models without the need for task-specific fine-tuning. It is broadly applicable across various domains, including arithmetic, commonsense, and symbolic reasoning tasks.

Experiment

The experiments assess chain-of-thought prompting, where few-shot exemplars include step-by-step reasoning, on arithmetic reasoning, commonsense reasoning, and symbolic manipulation tasks using language models ranging from 350M to 540B parameters. The technique produces large performance gains only for models of approximately 100B parameters or more, particularly on complex problems, and ablation studies demonstrate that the natural-language chain of reasoning is essential, not just extra computation or equation output. The approach is robust across different annotators and exemplar sets, and it also enables length generalization to symbolic reasoning inputs longer than those seen during prompting.


Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp