Command Palette
Search for a command to run...
Optical Character Recognition
Optical Character Recognition (OCR) is a technology that uses computer vision and algorithms to automatically recognize and convert text in images such as paper documents, scanned copies, or photographs into editable and searchable electronic text. The development of this concept has been a gradual evolution: as early as 1914 and 1929, physicist Emanuel Goldberg and engineer Gustav Tauschek respectively invented early mechanical template-based character recognition devices, laying the mechanical foundation for the technology; in the 1950s, IMR, founded by David H. Shepard, launched the first commercially available OCR device, "Gismo," and subsequently… IBM in 1959When promoting commercial scanning and character recognition systems, the technology was officially named... Optical Character Recognition (OCR)This has made it a globally accepted industry standard term.
In the course of technological evolution, several landmark papers have propelled OCR from rule-based comparison to the era of deep learning and large-scale models. In 1998, Yann LeCun et al. published a paper at AT&T Bell Labs... Gradient-based learning applied to document recognition Ray Smith proposed the LeNet-5 network and the MNIST dataset, laying the foundation for the application of convolutional neural networks in character recognition; in 2007, Ray Smith published... An Overview of the Tesseract OCR Engine This paper details the architecture and mechanisms of the open-source engine Tesseract, setting a benchmark in the open-source field. In 2017, Baoguang Shi et al. from Huazhong University of Science and Technology published... An End-to-End Trainable Neural Network for Image-Based Sequence Recognition A CRNN + CTC architecture was proposed, solving the end-to-end recognition problem of variable-length text and cursive characters; in 2023, Minghao Li et al. from Microsoft Research Asia published... TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models The introduction of Vision Transformer into OCR marks a new stage in the comprehensive advancement of recognition paradigms towards self-attention mechanisms and intelligent document parsing.
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.