Command Palette
Search for a command to run...
Neural Architectures for Named Entity Recognition
Neural Architectures for Named Entity Recognition
Guillaume Lample; Miguel Ballesteros; Sandeep Subramanian; Kazuya Kawakami; Chris Dyer
Abstract
State-of-the-art named entity recognition systems rely heavily on hand-crafted features and domain-specific knowledge in order to learn effectively from the small, supervised training corpora that are available. In this paper, we introduce two new neural architectures---one based on bidirectional LSTMs and conditional random fields, and the other that constructs and labels segments using a transition-based approach inspired by shift-reduce parsers. Our models rely on two sources of information about words: character-based word representations learned from the supervised corpus and unsupervised word representations learned from unannotated corpora. Our models obtain state-of-the-art performance in NER in four languages without resorting to any language-specific knowledge or resources such as gazetteers.
One-sentence Summary
Researchers from Carnegie Mellon University and Pompeu Fabra University propose two neural architectures for named entity recognition: a bidirectional LSTM-CRF model and a transition-based segment-labeling approach, both using character-based word representations and unsupervised word embeddings to achieve state-of-the-art performance in four languages without hand-crafted features or gazetteers.
Key Contributions
- The paper introduces two neural architectures for sequence labeling: a bidirectional LSTM with a CRF layer and a transition-based model that constructs and labels segments using a shift-reduce style approach.
- The architectures combine character-based word representations (capturing orthographic information) with unsupervised pretrained word embeddings, and they apply dropout training to ensure the model learns to rely on both evidence sources.
- Experiments on English, Dutch, German, and Spanish show the LSTM-CRF model obtains state-of-the-art NER results without any hand-engineered features, gazetteers, or language-specific resources, surpassing the best published scores in three of the four languages.
Introduction
Named entity recognition (NER) is essential for extracting structured information from text, but it suffers from a scarcity of labeled data and the unbounded creativity of names, which makes generalization hard. Existing approaches rely heavily on hand-engineered orthographic features and language-specific gazetteers, and even systems that use unsupervised data merely augment rather than replace these costly resources. The authors introduce neural architectures that eliminate the need for such external resources, using only a small amount of supervised data and unlabeled corpora. They combine a character-based word representation to capture orthographic patterns with distributional word embeddings, and they model output label dependencies via a bidirectional LSTM-CRF or a novel transition-based chunking model with stack LSTMs. This design achieves state-of-the-art or near-state-of-the-art NER performance across English, Dutch, German, and Spanish without any handcrafted features or gazetteers.
Dataset
The authors rely on two distinct data resources: large external corpora for pretraining word embeddings, and standard multilingual named entity recognition datasets for supervised fine-tuning and evaluation.
Word embedding pretraining data
- Sources:
- Spanish: Spanish Gigaword version 3
- Dutch & other languages: Leipzig Corpora Collection
- German: German monolingual training data from the 2010 Machine Translation Workshop
- English: English Gigaword version 4, with the LA Times and NY Times portions removed
- Processing: Embeddings are trained with skip‑n‑gram, a word2vec variant that respects word order. A minimum word frequency cutoff of 4 is applied, and the context window is 8 tokens.
- Dimensions: English vectors use 100 dimensions; all other languages use 64 dimensions.
- Usage: The pretrained vectors are used to initialize the model’s lookup table. During NER training, these embeddings are fine‑tuned.
Named entity recognition datasets
- Sources: CoNLL‑2002 (Spanish, Dutch) and CoNLL‑2003 (English, German) shared tasks.
- Composition: Each dataset contains four entity types – location, person, organization, and miscellaneous.
- Preprocessing: No POS tags are used. The only transformation is replacing every digit with a zero in the English dataset; all other languages are used as released.
- Usage: The standard train, development, and test splits of each dataset are employed for model training and evaluation. No additional cropping or filtering is applied.
Method
The authors leverage a hybrid tagging architecture that combines Long Short-Term Memory (LSTM) networks with Conditional Random Fields (CRF) for sequence labeling tasks such as Named Entity Recognition.
The core of the model relies on a bidirectional LSTM to capture context. While standard recurrent neural networks often fail to learn long-range dependencies, LSTMs utilize memory cells and gating mechanisms to control information flow. For a given sentence, a forward LSTM processes the sequence from left to right, while a backward LSTM processes it in reverse. The final representation for each word is obtained by concatenating its left and right context vectors, effectively embedding the word within its full sentence context.
To model the dependencies between output labels, the authors employ a CRF layer rather than making independent classification decisions. The CRF computes a score for a sequence of predictions by summing emission scores from the LSTM and transition scores from a learned matrix. During training, the model maximizes the log-probability of the correct tag sequence, encouraging valid label transitions.
As shown in the figure below, the overall architecture integrates these components. Word embeddings are fed into the bidirectional LSTM encoder. The resulting context representations are concatenated and linearly projected before being passed to the CRF layer, which outputs the final tag sequence.
A critical aspect of this design is the construction of the input word embeddings. The authors argue that learning independent representations for word types from limited data is difficult, so they construct word representations from their constituent characters.
Refer to the framework diagram for the character-based embedding generation.
In this module, character embeddings are processed by a bidirectional LSTM to capture morphological features like prefixes and suffixes. The final character-based representation is concatenated with a word-level embedding retrieved from a lookup table. These lookup tables are initialized with pretrained skip-n-gram embeddings that are fine-tuned during training. To ensure the model utilizes both character and word-level features effectively, dropout training is applied to the final embedding layer just before the input to the bidirectional LSTM.
As an alternative to the LSTM-CRF approach, the authors also explore a transition-based chunking model. This architecture utilizes a Stack-LSTM, which augments a standard LSTM with a stack pointer to maintain a summary embedding of a stack data structure. The model incrementally constructs chunks of the input using a transition inventory that includes SHIFT, OUT, and REDUCE operations, allowing it to directly predict labeled multi-token entities.
Experiment
The evaluation uses CoNLL-2002 and CoNLL-2003 datasets in multiple languages to test LSTM-CRF and Stack-LSTM models trained with SGD and dropout. The LSTM-CRF model attains leading scores on all languages without external resources, while the Stack-LSTM performs competitively but relies more on character-based representations. Component analysis shows that pretrained word embeddings contribute the largest improvement, followed by the CRF layer, dropout, and character embeddings.
The transition table demonstrates the Stack-LSTM's shift-reduce parsing process for named entity recognition. It shows how the model sequentially moves words from a buffer to a stack, combines them into labelled entity segments using the REDUCE action, and finally emits completed entities with the OUT action. SHIFT action moves a word embedding from the buffer onto the top of the stack, advancing the reading position. REDUCE(y) takes the top items on the stack, combines them via a learned function, and replaces them with a single segment labelled with entity type y. OUT action pops the top entity from the stack and appends it to the output sequence of recognized named entities.
On the English CoNLL-2003 NER test set, prior models that leverage external labeled data reach F1 scores up to 90.90, while the best performance without such data is 90.05. The proposed LSTM-CRF model surpasses all listed systems, including those augmented with gazetteers, without using any external resources. The highest F1 among listed models using external labeled data is 90.90 (Lin and Wu 2009; Passos et al. 2014), while the best without external resources reaches 90.05 (Passos et al. 2014). The authors' LSTM-CRF model outperforms every system in the comparison, beating even those that rely on gazetteers and knowledge bases.
On the German CoNLL-2003 test set, the LSTM-CRF model achieved an F1 score of 78.76, outperforming all prior methods including those that used external labeled data. The character-level representation component contributed a substantial gain, raising the F1 from 75.06 to 78.76. LSTM-CRF reaches 78.76 F1, exceeding the best external-data system (Gillick et al., 76.22) by 2.54 points. Without character embeddings, LSTM-CRF scores 75.06 F1, placing it near earlier non-external models like Qi et al. (75.72). All systems using external labeled data, such as Florian et al. and Gillick et al., score between 72.41 and 76.22, below the LSTM-CRF result.
On the Dutch CoNLL-2002 test set, the LSTM-CRF model achieves an F1 of 81.74, outperforming all prior systems that do not use external labeled data. Dutch is the only language where a model leveraging external resources (Gillick et al. at 82.84) surpasses the LSTM-CRF. Removing character-level representations reduces performance substantially for both architectures, with the Stack-LSTM showing an especially steep decline. LSTM-CRF (81.74) outperforms all earlier methods that rely solely on the dataset, including Carreras et al. (77.05) and Gillick et al. without external data (78.08). Gillick et al. with external labeled data reaches 82.84, making Dutch the sole exception where the LSTM-CRF is not the top system. Without character-level features, LSTM-CRF falls to 73.14 and Stack-LSTM to 69.90, confirming the heavier dependence of the Stack-LSTM on orthographic information.
On the Spanish CoNLL-2002 test set, the LSTM-CRF model achieves an F1 of 85.75, surpassing all prior systems, including those that leveraged external labeled data. Even without character-level features, the LSTM-CRF reaches 83.44, which is already higher than the best external-data baseline. By contrast, the Stack-LSTM without character information scores only 79.46, highlighting its strong dependence on orthographic information. The full LSTM-CRF model attains 85.75 F1, exceeding all previous methods, including those that used external labeled resources. LSTM-CRF without character features still scores 83.44, outperforming the best external-data model (Gillick et al.* at 82.95). Stack-LSTM without character representations drops to 79.46, far below other systems, confirming its reliance on character-level information.
The evaluation uses the CoNLL named entity recognition test sets across English, German, Dutch, and Spanish to compare the proposed LSTM-CRF and Stack-LSTM models against prior systems. On all languages except Dutch, the LSTM-CRF model without external resources sets new state-of-the-art performance, surpassing even models that rely on gazetteers and external labeled data. Ablation studies show that removing character-level representations causes a pronounced drop in F1 scores, particularly for the Stack-LSTM, confirming the critical role of orthographic features in both architectures.