Command Palette
Search for a command to run...
Rich feature hierarchies for accurate object detection and semantic segmentation
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick Jeff Donahue Trevor Darrell Jitendra Malik
Abstract
Object detection performance, as measured on the canonical PASCAL VOC dataset, has plateaued in the last few years. The best-performing methods are complex ensemble systems that typically combine multiple low-level image features with high-level context. In this paper, we propose a simple and scalable detection algorithm that improves mean average precision (mAP) by more than 30% relative to the previous best result on VOC 2012---achieving a mAP of 53.3%. Our approach combines two key insights: (1) one can apply high-capacity convolutional neural networks (CNNs) to bottom-up region proposals in order to localize and segment objects and (2) when labeled training data is scarce, supervised pre-training for an auxiliary task, followed by domain-specific fine-tuning, yields a significant performance boost. Since we combine region proposals with CNNs, we call our method R-CNN: Regions with CNN features. We also compare R-CNN to OverFeat, a recently proposed sliding-window detector based on a similar CNN architecture. We find that R-CNN outperforms OverFeat by a large margin on the 200-class ILSVRC2013 detection dataset. Source code for the complete system is available at http://www.cs.berkeley.edu/~rbg/rcnn.
One-sentence Summary
UC Berkeley researchers propose R-CNN (Regions with CNN features), a scalable detection algorithm that combines high-capacity convolutional neural networks with bottom-up region proposals and supervised pre-training plus fine-tuning, achieving a 53.3% mAP on PASCAL VOC 2012, a more than 30% relative improvement over the previous best result, and significantly outperforming the OverFeat sliding-window detector on the 200-class ILSVRC2013 detection dataset.
Key Contributions
- R-CNN applies high-capacity convolutional neural networks to bottom-up region proposals for object localization and segmentation, achieving 53.3% mAP on PASCAL VOC 2012, a relative improvement of more than 30% over prior state-of-the-art ensemble systems.
- Supervised pre-training on a large auxiliary classification dataset (ILSVRC) followed by domain-specific fine-tuning on scarce detection data yields an 8-point mAP boost on PASCAL VOC 2010, achieving 54% mAP compared to 33% for the deformable part model.
- R-CNN surpasses the contemporary OverFeat sliding-window detector by a large margin on the ILSVRC2013 200-class detection benchmark (31.4% mAP versus 24.3%), validating the advantage of region-based CNN features.
Introduction
Prior to this work, object detection performance on benchmarks like PASCAL VOC had plateaued, with only minor gains from ensembles and variants of shallow handcrafted features such as SIFT and HOG. These blockwise orientation histograms lack the hierarchical, multi-stage processing that characterizes the primate visual system and that might yield more informative representations. Although convolutional neural networks (CNNs) had recently shown dramatic improvements in image classification on ImageNet, it was unclear whether those gains would transfer to detection, a task that requires localizing objects and training large models with limited annotated data. Earlier CNN-based detectors relied on sliding windows, but deep networks with large receptive fields and strides made precise localization difficult, and the scarcity of detection labels hindered training high-capacity models from scratch.
The authors bridge the gap between classification and detection with R-CNN, which combines bottom-up region proposals with a CNN to extract fixed-length features per proposal and then classifies them with linear SVMs. To overcome data scarcity, they introduce a supervised pre-training and domain-specific fine-tuning paradigm: the CNN is first trained on the large ImageNet classification dataset and then fine-tuned on the smaller PASCAL detection data. This approach yields substantial gains, improving over strongly engineered HOG-based deformable part models by a wide margin and setting a new state of the art while remaining efficient due to shared features across categories.
Dataset
The authors use the ILSVRC2013 detection dataset in this work. Here is a breakdown of its composition, processing, and how the different subsets are applied.
-
Dataset composition and sources
-
Three official splits:
train(395,918 images),val(20,121 images), andtest(40,152 images). -
valandtestcome from the same image distribution, are exhaustively annotated (all instances of all 200 classes have bounding boxes), and have scene complexity similar to PASCAL VOC. -
trainis drawn from the ILSVRC2013 classification image distribution. It is not exhaustively annotated and contains more variable imagery, often showing a single centered object. -
Each class also has a verified negative image set (no instance of that class appears), but these sets are not used in this work.
-
Subset handling and filtering
-
Because
valmust serve both training and evaluation, the authors split it into two approximately class-balanced halves,val1andval2, using a randomised clustering procedure that minimises maximum class imbalance (maximum ~11%, median ~4%). -
The
trainset is used only as an auxiliary source of positive examples; no hard negatives are mined fromtrainbecause its annotations are incomplete. -
Region proposal details
-
Selective search is run in fast mode only on
val1,val2, andtest(not ontrain). -
Every image is first resized to a fixed width of 500 pixels to make proposal generation scale-invariant.
-
On
val, this yields an average of 2403 region proposals per image with a recall of 91.6% for ground-truth boxes at an IoU threshold of 0.5. -
Training data construction and usage
-
val1+trainNdataset: all selective search boxes and all ground-truth boxes fromval1, plus up toNground-truth boxes per class sampled from thetrainset (N= 0, 500, or 1000). -
CNN fine-tuning uses the entire
val1+trainNset (50k SGD iterations, same settings as for PASCAL VOC). -
SVM training takes all ground-truth boxes from
val1+trainNas positives. Hard negatives are mined from a randomly selected 5000-image subset ofval1(using the fullval1gave only a 0.5 mAP point gain while doubling training time). -
Bounding-box regressors are trained solely on
val1. -
The built-in negative image sets are not used in any training stage.
Method
The authors propose an object detection system comprising three distinct modules. The first module generates category-independent region proposals, defining the candidate detections. The second module is a large convolutional neural network that extracts a fixed-length feature vector from each region. The third module consists of class-specific linear support vector machines for classification.
For region proposals, the authors employ selective search to generate approximately 2000 candidate bounding boxes per image at test time. This approach allows for a controlled comparison with prior detection methods while remaining agnostic to the specific proposal algorithm.
To extract features, the system computes a 4096-dimensional feature vector from each proposal using a convolutional neural network. Since the network requires a fixed input size of 227×227 pixels, the authors warp the pixels within a tight bounding box around each arbitrary-shaped region to the required dimensions. Prior to warping, the bounding box is dilated to include exactly p=16 pixels of context around the original box. As shown in the figure below, this warping process produces a diverse set of normalized training samples across different object categories.
At test time, each warped proposal is forward-propagated through the network to compute its feature vector. For each object class, the extracted features are scored using the corresponding trained linear support vector machine. To eliminate redundant detections, the system applies greedy non-maximum suppression independently for each class, rejecting regions that have an intersection-over-union overlap with a higher-scoring selected region exceeding a learned threshold. The efficiency of this pipeline stems from sharing all network parameters across categories and utilizing low-dimensional feature vectors, allowing the system to scale to thousands of classes without approximate techniques.
The training process involves three main stages: supervised pre-training, domain-specific fine-tuning, and classifier training. The network is first discriminatively pre-trained on a large auxiliary dataset using image-level annotations. To adapt the network to the detection task and the warped proposal domain, the authors continue stochastic gradient descent training using only warped region proposals. The original classification layer is replaced with a randomly initialized (N+1)-way classification layer, where N is the number of object classes plus one for background. Region proposals with an intersection-over-union overlap of ≥0.5 with a ground-truth box are treated as positives, while the rest are negatives. The learning rate is set to 0.001, and mini-batches of size 128 are constructed by uniformly sampling 32 positive windows and 96 background windows.
Finally, for the object category classifiers, the authors optimize one linear support vector machine per class. To handle the large training data, they employ hard negative mining. Positive examples are defined as the ground-truth bounding boxes, while negative examples are regions with an intersection-over-union overlap below a carefully selected threshold of 0.3 with any ground-truth box. This threshold is crucial, as setting it to 0.5 or 0 significantly decreases mean average precision.
Experiment
The experiments evaluate R-CNN on PASCAL VOC and ILSVRC2013 object detection, showing that pre-trained CNN features with SVM classifiers substantially outperform hand-crafted features and competing methods, especially after fine-tuning. Ablation studies reveal that convolutional layers alone capture general visual patterns, while fine-tuning adapts the higher layers for domain-specific detection, and deeper network architectures yield further gains. The approach is also successfully extended to semantic segmentation, where combining features from whole regions and foreground masks achieves top accuracy on VOC 2011.
On PASCAL VOC 2010, R-CNN with bounding-box regression achieves a new state-of-the-art mAP of 53.7%, surpassing the previous top method SegDPM (40.4%) by a wide margin and dramatically improving over the UVA system (35.1%) that uses the same selective search proposals. The bounding-box regression step alone adds over 3 points of mAP by correcting localization errors, and R-CNN attains these gains without the context rescoring used by DPM and SegDPM. R-CNN without extra context rescoring outperforms SegDPM, which relied on additional context and segmentation cues. Compared to the UVA system using identical region proposals, R-CNN with bounding-box regression raises mAP from 35.1% to 53.7% while being much faster. Bounding-box regression lifts mAP by 3 to 4 points, fixing a large number of mislocalized detections. Large per-class gains are seen on hard categories like bird (25.6% SegDPM to 53.0% R-CNN BB) and cat (50.8% to 69.9%). Even without bounding-box regression, R-CNN (50.2% mAP) outpaces Regionlets (39.7%), the strongest competitor that also uses selective search proposals.
Without fine-tuning, the convolutional features (pool_5) alone deliver strong performance, while adding fully connected layers brings minimal benefit and the top layer (fc_7) actually degrades mAP. Fine-tuning on VOC 2007 substantially boosts accuracy at every layer, and a simple bounding-box regression stage adds a further 3–4 mAP points by correcting localization errors. Using only pool_5 features (6% of the network's parameters) nears the mAP of the full frozen CNN, and removing the fc_7 layer raises mAP, showing that most representational power lies in the convolutional layers. Fine-tuning lifts mAP across all layers, and appending a bounding-box regression step onto fine-tuned fc_7 features gives an approximate 4-point gain, largely by fixing mislocalized detections.
Using the deeper OxfordNet (O-Net) in R-CNN raises mean average precision on VOC 2007 test from 58.5% to 66.0% compared to the TorontoNet (T-Net) when both use bounding-box regression. This improvement is consistent across almost all object categories but comes with a roughly sevenfold increase in forward-pass time. R-CNN O-Net with bounding-box regression reaches 66.0% mAP, substantially outperforming R-CNN T-Net with bounding-box regression at 58.5%. The forward pass of O-Net is approximately seven times slower than that of T-Net, making the accuracy gain costly in compute time.
Starting from a baseline mAP of 20.9% with SVM training on val1 alone, adding more training data (val1+train5k) raises mAP to 24.1%, but using 1000 extra examples instead of 500 yields no further gain. Fine-tuning the CNN provides a larger improvement: fine-tuning on val1+train1k reaches 29.7% mAP, and incorporating bounding-box regression boosts performance to 31.0% on val2, with val2 mAP closely mirroring test set mAP throughout. Expanding the SVM training set from val1 to val1+train5k raises mAP from 20.9 to 24.1, but using train1k gives identical mAP. Fine-tuning the CNN on val1+train1k (with fc7 features) lifts mAP to 29.7, and adding bounding-box regression yields a final mAP of 31.0 on val2.
On the VOC 2011 validation set, using CNN features inside a second-order pooling framework, the best strategy concatenates full-window and foreground-mask features (full+fg) from the fc6 layer, reaching 47.9% mean accuracy and modestly exceeding the O2P baseline (46.4%). Foreground-only features slightly outperform full-window features (43.7% vs 43.0% at fc6), but combining them provides a 4.2-point gain, and fc6 consistently outperforms fc7 across all feature strategies. Combining full-window and foreground-mask features (full+fg, fc6) achieves 47.9% mean accuracy, surpassing the O2P result of 46.4%. The fc6 layer always scores higher than fc7: full+fg drops from 47.9% (fc6) to 45.8% (fc7), and similar gaps hold for full and fg alone.
R-CNN combined with bounding-box regression dramatically advances object detection on PASCAL VOC 2010, surpassing prior state-of-the-art methods without requiring external context. Fine-tuning the convolutional network on target data provides substantial accuracy gains, and adding bounding-box regression corrects localization errors, contributing consistent improvement across all feature layers. Using a deeper architecture further boosts performance at a significant computational cost, while experiments confirm that the convolutional layers capture most of the representational power, and these features also transfer effectively to semantic segmentation tasks when combined with foreground masks.