HyperAIHyperAI

Command Palette

Search for a command to run...

Neural Machine Translation by Jointly Learning to Align and Translate

Dzmitry Bahdanau KyungHyun Cho Yoshua Bengio

Abstract

Neural machine translation is a recently proposed approach to machine translation. Unlike the traditional statistical machine translation, the neural machine translation aims at building a single neural network that can be jointly tuned to maximize the translation performance. The models proposed recently for neural machine translation often belong to a family of encoder-decoders and consists of an encoder that encodes a source sentence into a fixed-length vector from which a decoder generates a translation. In this paper, we conjecture that the use of a fixed-length vector is a bottleneck in improving the performance of this basic encoder-decoder architecture, and propose to extend this by allowing a model to automatically (soft-)search for parts of a source sentence that are relevant to predicting a target word, without having to form these parts as a hard segment explicitly. With this new approach, we achieve a translation performance comparable to the existing state-of-the-art phrase-based system on the task of English-to-French translation. Furthermore, qualitative analysis reveals that the (soft-)alignments found by the model agree well with our intuition.

One-sentence Summary

Researchers from Jacobs University Bremen, Germany and Université de Montréal propose a neural machine translation model that jointly learns to align and translate, allowing the decoder to automatically soft-search for relevant source parts to overcome the fixed-length vector bottleneck in standard encoder–decoder architectures, and achieve English-to-French translation performance comparable to state-of-the-art phrase-based systems.

Key Contributions

  • The paper conjectures that the fixed-length context vector in encoder–decoder neural machine translation is a bottleneck and introduces a soft-search mechanism that lets the decoder attend to relevant source annotations dynamically at each step, removing the need to compress the entire input sentence into a single vector.
  • Experiments on English-to-French translation show that the proposed RNNsearch model significantly outperforms the conventional encoder–decoder (RNNencdec) for all sentence lengths and is much more robust to long source sentences.
  • The approach achieves translation performance comparable to a state-of-the-art phrase-based statistical system, and qualitative analysis reveals that the learned soft alignments agree with intuitive word-level translation correspondences.

Introduction

Neural machine translation offers an end-to-end alternative to phrase-based systems by using a single network that directly maps a source sentence to a target translation, but the dominant encoder–decoder architecture forces the entire source meaning into a single fixed-length vector. This compression becomes a severe bottleneck for long sentences, causing translation quality to degrade quickly as input length grows. The authors introduce a model that jointly learns to align and translate, replacing the fixed-length context vector with a dynamic soft-search over source positions at each decoding step. Instead of encoding the full sentence into one vector, the approach computes a sequence of annotations and adaptively selects the most relevant ones when generating each target word. This structure dramatically improves robustness to sentence length, yields substantial gains over the basic encoder–decoder, and reaches translation performance comparable to conventional phrase-based systems on English-to-French tasks.

Dataset

The authors use the WMT’14 English-French parallel corpus, which combines several sources:

  • Europarl (61 million words)
  • News commentary (5.5 million words)
  • United Nations (421 million words)
  • Two crawled corpora (90 million and 272.5 million words) The raw total is 850 million words.

To create the training data, they apply the data selection method of Axelrod et al. (2011), reducing the combined corpus to 348 million words. No monolingual data is added.

For validation, they concatenate the news-test-2012 and news-test-2013 sets. Evaluation uses the WMT’14 English-French test set, which contains 3,003 unseen sentences.

Preprocessing steps:

  • Standard tokenization
  • Vocabulary limited to the 30,000 most frequent words in each language; any other word is replaced with a special [UNK] token
  • No lowercasing, stemming, or other special preprocessing is applied

Method

The authors build upon the standard RNN Encoder-Decoder framework, where an encoder reads an input sequence of vectors x=(x1,…,xTx)\mathbf{x} = (x_1, \dots, x_{T_x})x=(x1​,…,xTx​​) into a fixed-length vector ccc. Typically, an RNN computes hidden states ht=f(xt,ht−1)h_t = f(x_t, h_{t-1})ht​=f(xt​,ht−1​), and the context vector is derived from these states. The decoder then predicts the next word yty_tyt​ by decomposing the joint probability into ordered conditionals: p(y)=∏t=1Tp(yt∣{y1,…,yt−1},c)p(\mathbf{y}) = \prod_{t=1}^{T} p(y_t \mid \{y_1, \dots, y_{t-1}\}, c)p(y)=∏t=1T​p(yt​∣{y1​,…,yt−1​},c) where each conditional probability is modeled as p(yt∣{y1,…,yt−1},c)=g(yt−1,st,c)p(y_t \mid \{y_1, \dots, y_{t-1}\}, c) = g(y_{t-1}, s_t, c)p(yt​∣{y1​,…,yt−1​},c)=g(yt−1​,st​,c), with sts_tst​ being the decoder hidden state.

To overcome the bottleneck of encoding all source information into a fixed-length vector, the authors propose a novel architecture that learns to align and translate simultaneously. In this new model, the probability of each target word is conditioned on a distinct context vector cic_ici​: p(yi∣y1,…,yi−1,x)=g(yi−1,si,ci)p(y_i \mid y_1, \dots, y_{i-1}, \mathbf{x}) = g(y_{i-1}, s_i, c_i)p(yi​∣y1​,…,yi−1​,x)=g(yi−1​,si​,ci​) where the decoder hidden state is computed as si=f(si−1,yi−1,ci)s_i = f(s_{i-1}, y_{i-1}, c_i)si​=f(si−1​,yi−1​,ci​). The context vector cic_ici​ is calculated as a weighted sum of the input annotations (h1,…,hTx)(h_1, \dots, h_{T_x})(h1​,…,hTx​​): ci=∑j=1Txαijhjc_i = \sum_{j=1}^{T_x} \alpha_{ij} h_jci​=∑j=1Tx​​αij​hj​ The weights αij\alpha_{ij}αij​ are computed via a softmax over alignment scores eije_{ij}eij​: αij=exp⁡(eij)∑k=1Txexp⁡(eik)\alpha_{ij} = \frac{\exp(e_{ij})}{\sum_{k=1}^{T_x} \exp(e_{ik})}αij​=∑k=1Tx​​exp(eik​)exp(eij​)​ Here, eij=a(si−1,hj)e_{ij} = a(s_{i-1}, h_j)eij​=a(si−1​,hj​) represents an alignment model, parametrized as a feedforward neural network, which scores how well the input around position jjj matches the output at position iii. This mechanism allows the decoder to selectively focus on relevant parts of the source sentence, effectively implementing attention.

As shown in the figure below:

To generate rich annotations that summarize both preceding and following words for each source word, the authors employ a bidirectional RNN as the encoder. A forward RNN f⃗\vec{f}f​ processes the input sequence from x1x_1x1​ to xTxx_{T_x}xTx​​ to produce forward hidden states (h⃗1,…,h⃗Tx)(\vec{h}_1, \dots, \vec{h}_{T_x})(h1​,…,hTx​​). Simultaneously, a backward RNN f←\overleftarrow{f}f​ reads the sequence in reverse, from xTxx_{T_x}xTx​​ to x1x_1x1​, yielding backward hidden states (h←1,…,h←Tx)(\overleftarrow{h}_1, \dots, \overleftarrow{h}_{T_x})(h1​,…,hTx​​). The final annotation for each word xjx_jxj​ is obtained by concatenating the corresponding forward and backward hidden states: hj=[h⃗j⊤;h←j⊤]⊤h_j = \left[ \vec{h}_j^\top; \overleftarrow{h}_j^\top \right]^\tophj​=[hj⊤​;hj⊤​]⊤ This sequence of annotations is then utilized by the decoder and the alignment model to compute the dynamic context vectors, relieving the encoder from the burden of compressing the entire source sentence into a single static representation.

Experiment

The experiments evaluate English-to-French translation on WMT'14 data, comparing a basic RNN Encoder-Decoder with the proposed RNNsearch model that uses a soft alignment mechanism. RNNsearch achieves higher BLEU scores, matches a phrase-based system on sentences with known words, and remains robust to increasing sentence lengths, whereas the basic model's performance degrades sharply on long inputs. Qualitative analysis shows the soft alignment handles non-monotonic word orders and article ambiguity intuitively, and RNNsearch translates long sentences accurately while the basic model often deviates from the source meaning.

The RNNsearch models consistently outperformed their RNNencdec counterparts, with the attention mechanism yielding higher BLEU scores on both full test sets and subsets without unknown words. On sentences containing no unknown words, the best RNNsearch variant exceeded the conventional phrase-based Moses system despite using no additional monolingual data, highlighting its ability to rival strong statistical baselines. The gains are particularly pronounced for long sentences, where the basic encoder–decoder tends to degrade rapidly. RNNsearch-30 surpassed RNNencdec-50 (21.50 vs. 17.82 BLEU on all sentences), showing the advantage of attention even with a smaller hidden state. When evaluating only sentences free of unknown words, RNNsearch-50* achieved 36.15 BLEU, outperforming Moses (35.63) without extra monolingual corpora.

The evaluation compares RNNsearch models with attention against basic RNNencdec on machine translation, finding that attention consistently yields higher BLEU scores, especially on long sentences where the baseline degrades. The best RNNsearch variant outperforms the conventional phrase-based Moses system on sentences without unknown words, even without using extra monolingual data, demonstrating competitive performance against strong statistical baselines. Notably, RNNsearch-30 surpasses RNNencdec-50, showing the attention mechanism’s advantage independent of model size.


Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp