arxivcs.CV2026-07-23
ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning
Debjyoti Das Adhikary, Aritra Hazra, Partha Pratim Chakrabarti
Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing training-free alternatives instead ground captions in detected objects to curb hallucination, but apply only a single, fixed corr…