arxivcs.CV2026-07-24
LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR
Xudong Liu, Bicheng Wan, Yulin Jin
End-to-end OCR systems based on vision-language models have achieved strong performance in complex document OCR, but their efficiency is limited by the large number of visual tokens produced from document images. Many of these tokens correspond to blank margins or visually redund…