Multi-page document visual question answering requires locating sparse evidence at both the page and region levels. Existing approaches typically emphasize one level over the other: page-centric methods focus on page acquisition, with region operations serving mainly as navigatio…
The development of accurate yet efficient deep learning models for Chest X-ray interpretation is critical for scalable screening and point-of-care diagnostics. While self-supervised learning reduces dependency on costly annotations and knowledge distillation compresses models, ef…
Retrieval-Augmented Generation (RAG) streamlines long-document understanding by leveraging retrieval mechanisms to restrict input images to a highly curated subset. However, existing multimodal RAG pipelines primarily face two critical challenges: first, standard semantic similar…