URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding
arxiv.org·18h
Flag this post

View PDF HTML (experimental)

Abstract:Recent multimodal large language models (MLLMs) still struggle with long document understanding due to two fundamental challenges: information interference from abundant irrelevant content, and the quadratic computational cost of Transformer-based architectures. Existing approaches primarily fall into two categories: token compression, which sacrifices fine-grained details; and introducing external retrievers, which increase system complexity and prevent end-to-end optimization. To address these issues, we conduct an in-depth analysis and observe that MLLMs exhibit a human-like coarse-to-fine reasoning pattern: early Transformer layers attend broadly across the document…

Similar Posts

Loading similar posts...