Detail View

ReFineVQA: Iterative Refinement of Video Description via Feedback Generation for Video Question Answering

Citations

WEB OF SCIENCE

Citations

SCOPUS

Metadata Downloads

DC Field Value Language
dc.contributor.author Shin, Jeongwan -
dc.contributor.author Hur, Chan -
dc.contributor.author Cho, Seongmin -
dc.contributor.author Choi, Jaeho -
dc.contributor.author Park, Hyeyoung -
dc.date.accessioned 2026-07-22T18:10:13Z -
dc.date.available 2026-07-22T18:10:13Z -
dc.date.created 2026-06-19 -
dc.date.issued 2026-03-10 -
dc.identifier.isbn 979-833155511-5 -
dc.identifier.issn 2642-9381 -
dc.identifier.uri https://scholar.dgist.ac.kr/handle/20.500.11750/60477 -
dc.description.abstract Video question answering is a non-trivial task that demands joint understanding of visual contents and linguistic questions as well as temporal reasoning across video frames. Recent agent-based approaches address this by conducting multi-step reasoning with large language models (LLMs) across frame-level captions generated by vision-language models, but encounter limited temporal coherence across frames. A possible direction based on video language models (VideoLMs) directly captures temporal dynamics via video-level descriptions, but often lacks fine-grained visual cues due to a restricted number of input frames and a large dependency on input prompts. To tackle these challenges, we propose RefineVQA, a training-free framework that can easily be plugged into existing VideoLMs with iterative, LLM-guided description refinements. Specifically, the VideoLM produces an initial description, followed by LLM feedback determining whether the description suffices for the question and guiding further visual extraction, which in turn enhances the description quality while preserving temporal context. Plugged into state-of-the-art VideoLMs, ReFineVQA yields consistent gains across diverse benchmarks-NExT-QA, EgoSchema, VideoMME, ActivityNet, and StreamingBench-even with a small external LLM of 3.8B parameters. © 2026 IEEE. -
dc.language English -
dc.publisher Institute of Electrical and Electronics Engineers Inc. -
dc.relation.ispartof Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026 -
dc.title ReFineVQA: Iterative Refinement of Video Description via Feedback Generation for Video Question Answering -
dc.type Conference Paper -
dc.identifier.doi 10.1109/WACV61042.2026.00738 -
dc.identifier.scopusid 2-s2.0-105041301405 -
dc.identifier.bibliographicCitation 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026, pp.7647 - 7657 -
dc.identifier.url https://wacv.thecvf.com/Conferences/2026 -
dc.citation.conferenceDate 2026-03-06 -
dc.citation.conferencePlace US -
dc.citation.conferencePlace Tucson -
dc.citation.endPage 7657 -
dc.citation.startPage 7647 -
dc.citation.title 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026 -
Show Simple Item Record

File Downloads

  • There are no files associated with this item.

공유

qrcode
공유하기

Related Researcher

최재호
Choi, Jae-Ho최재호

Department of Electrical Engineering and Computer Science

read more

Total Views & Downloads

???jsp.display-item.statistics.view???: , ???jsp.display-item.statistics.download???: