<?xml version="1.0" encoding="UTF-8"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns="http://purl.org/rss/1.0/" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel rdf:about="https://scholar.dgist.ac.kr/handle/20.500.11750/58108">
    <title>Repository Collection: null</title>
    <link>https://scholar.dgist.ac.kr/handle/20.500.11750/58108</link>
    <description />
    <items>
      <rdf:Seq>
        <rdf:li rdf:resource="https://scholar.dgist.ac.kr/handle/20.500.11750/60477" />
        <rdf:li rdf:resource="https://scholar.dgist.ac.kr/handle/20.500.11750/60061" />
      </rdf:Seq>
    </items>
    <dc:date>2026-08-01T04:19:43Z</dc:date>
  </channel>
  <item rdf:about="https://scholar.dgist.ac.kr/handle/20.500.11750/60477">
    <title>ReFineVQA: Iterative Refinement of Video Description via Feedback Generation for Video Question Answering</title>
    <link>https://scholar.dgist.ac.kr/handle/20.500.11750/60477</link>
    <description>Title: ReFineVQA: Iterative Refinement of Video Description via Feedback Generation for Video Question Answering
Author(s): Shin, Jeongwan; Hur, Chan; Cho, Seongmin; Choi, Jaeho; Park, Hyeyoung
Abstract: Video question answering is a non-trivial task that demands joint understanding of visual contents and linguistic questions as well as temporal reasoning across video frames. Recent agent-based approaches address this by conducting multi-step reasoning with large language models (LLMs) across frame-level captions generated by vision-language models, but encounter limited temporal coherence across frames. A possible direction based on video language models (VideoLMs) directly captures temporal dynamics via video-level descriptions, but often lacks fine-grained visual cues due to a restricted number of input frames and a large dependency on input prompts. To tackle these challenges, we propose RefineVQA, a training-free framework that can easily be plugged into existing VideoLMs with iterative, LLM-guided description refinements. Specifically, the VideoLM produces an initial description, followed by LLM feedback determining whether the description suffices for the question and guiding further visual extraction, which in turn enhances the description quality while preserving temporal context. Plugged into state-of-the-art VideoLMs, ReFineVQA yields consistent gains across diverse benchmarks-NExT-QA, EgoSchema, VideoMME, ActivityNet, and StreamingBench-even with a small external LLM of 3.8B parameters. © 2026 IEEE.</description>
    <dc:date>2026-03-09T15:00:00Z</dc:date>
  </item>
  <item rdf:about="https://scholar.dgist.ac.kr/handle/20.500.11750/60061">
    <title>MVDoppler-Pose: Multi-Modal Multi-View mmWave Sensing for Long-Distance Self-Occluded Human Walking Pose Estimation</title>
    <link>https://scholar.dgist.ac.kr/handle/20.500.11750/60061</link>
    <description>Title: MVDoppler-Pose: Multi-Modal Multi-View mmWave Sensing for Long-Distance Self-Occluded Human Walking Pose Estimation
Author(s): Choi, Jae-Ho; Hor, Soheil; Yang, Shubo; Arbabian, Amin
Abstract: One of the main challenges in reliable camera-based 3D pose estimation for walking subjects is to deal with self-occlusions, especially in the case of using low-resolution cameras or at longer distance scenarios. In recent years, millimeter-wave (mmWave) radar has emerged as a promising alternative, offering inherent resilience to the effect of occlusions and distance variations. However, mmWave-based human walking pose estimation (HWPE) is still in the nascent development stages, primarily due to its unique set of practical challenges including the quality of the observed radar signal dependent on the subject&amp;apos;s motion direction. This paper introduces the first comprehensive study comparing mmWave radar to camera systems for HWPE, highlighting its utility for distance-agnostic and occlusion-resilient pose estimation. Building upon mmWave&amp;apos;s unique advantages, we address its intrinsic directionality issue through a new approach - the synergetic integration of multi-modal, multi-view mmWave signals, achieving robust HWPE against variations both in distance and walking direction. Extensive experiments on a newly curated dataset not only demonstrate the superior potential of mmWave technology over traditional camera-based HWPE systems, but also validate the effectiveness of our approach in over-coming the core limitations of mmWave HWPE.</description>
    <dc:date>2025-06-14T15:00:00Z</dc:date>
  </item>
</rdf:RDF>

