Detail View

Toward Universal V2X Cooperation for VRU Safety: An Infrastructure-Centric Explainable Framework via Vision-Language Models

Citations

WEB OF SCIENCE

Citations

SCOPUS

Metadata Downloads

DC Field Value Language
dc.contributor.advisor 최지웅 -
dc.contributor.author Yunho Jeong -
dc.date.accessioned 2026-09-01T19:33:56Z -
dc.date.available 2026-09-01T19:33:56Z -
dc.date.issued 2026 -
dc.identifier.uri https://scholar.dgist.ac.kr/handle/20.500.11750/60829 -
dc.identifier.uri http://dgist.dcollection.net/common/orgView/200001016429 -
dc.description Vehicle-to-Everything (V2X) cooperation, Vulnerable Road User (VRU) anticipation, Vision-Language Model, Chain-of-Thought reasoning, Knowledge Distillation -
dc.description.abstract Vulnerable Road Users (VRUs) account for more than half of global traffic fatalities, and their safety at urban intersections is structurally constrained by the line-of-sight occlusions that single-vehicle perception cannot resolve. Conventional V2X cooperative perception schemes mitigate this limitation by exchanging raw sensor data or intermediate features, but they impose substantial communication bandwidth requirements and architectural compatibility constraints across heterogeneous vehicle platforms. This thesis proposes the VLM-based Explainable VRU Anticipation (VEVA) framework, an infrastructure-centric pipeline that anticipates the safety state of observed VRU and explains the reasoning behind the prediction in natural language, so that the resulting text can be used as a structured V2X cooperation message. The framework is structured in two phases. In Phase 1, a large-scale teacher VLM is prompted with a four-step chain-of-thought template that decomposes the safety estimation into sub-tasks, producing a reasoning-augmented annotation corpus on top of an existing binary-label dataset. In Phase 2, a student VLM is fine-tuned on this corpus through knowledge distillation, to learning teacher model’s reasoning pattern. On the WatchoutPed benchmark, the resulting student shows the leading performance in the quantitative metrics relative to five pedestrian-anticipation baseline models. These results demonstrate that the framework reliably interprets VRU behavioral intent to anticipate their safety states, while the resulting structured rationale serves as a universal V2X cooperation message that can be interpreted by any language-model-capable receiving vehicle regardless of its underlying architecture. Keywords: Vehicle-to-Everything (V2X) cooperation, Vulnerable Road User (VRU) anticipa- tion, Vision-Language Model, Chain-of-Thought reasoning, Knowledge Distillation|교통 약자는 전 세계 교통 사망 사고의 절반 이상을 차지하며, 도시 교차로에서의 사각지대 문제는 단일 차량 인지만으로 해결하기 어렵다. 기존 V2X 협력 인지 방식은 센서 데이터나 중간 특징을공유하여 이를 보완하지만, 통신 대역폭 요구가 크고 이종 차량 간 호환성 문제를 야기한다. 본 논문은 노변 카메라 영상으로부터 구조화된 자연어 안전 메시지를 생성하는 인프라 중심 프레임워크인 VEVA를 제안한다. 1단계에서는 대규모 교사 비전-언어 모델에 4단계 연쇄적 사고 프롬프트를적용하여 위치, 움직임, 신체 언어, 맥락 근거를 생성하고, 기존 이진 분류 데이터셋을 추론 증강데이터셋으로 확장한다. 2단계에서는 경량 학생 모델에 해당 데이터셋을 지도 학습으로 증류하여노변 배포가 가능한 단일 모델을 확보한다. WatchoutPed 벤치마크에서 본 모델은 다섯 가지 비교모델 대비 네 가지 종합 지표 모두에서 최고 성능을 달성하였으며, 정밀도와 재현율을 동시에 끌어올림으로써 분별 능력 자체가 향상되었음을 확인하였다. 제안 기법을 통해 생성된 자연어 근거자체가 차량 구조에 무관하게 해석 가능한 V2X 메세지로 기능하여 보편적인 V2X 협력의 가능성을제시한다.핵심어: 차량 사물 통신 협력, 교통 약자 위험 예측, 비전-언어 모델, 연쇄적 사고 추론, 지식 증류 -
dc.description.tableofcontents 1. Introduction 1

2. Background 5
2.1 V2X Cooperative Perception 5
2.2 Vision-Language Models 6
2.3 Chain-of-Thought Reasoning 7
2.4 Knowledge Distillation 9
2.5 Vulnerable Road User Anticipation 9

3. Method 12
3.1 Overview and Problem Formulation 12
3.1.1 Problem Formulation 12
3.1.2 Dataset 13
3.2 Phase 1: Reasoning-Augmented Annotation Generation 14
3.2.1 Input Construction 14
3.2.2 Teacher Model and Four-Step Chain-of-Thought Prompting 14
3.2.3 Structured Reasoning-Augmented Annotation 15
3.3 Phase 2: Knowledge Distillation via Supervised Fine-Tuning 16
3.3.1 Student Model and Parameter-Efficient Adaptation 17
3.3.2 Learning Objective 17

4. Experimental Results 18
4.1 Experimental Setup 18
4.1.1 Baselines 18
4.1.2 Evaluation Metrics 19
4.1.3 Parameter Setting 19
4.2 Quantitative Evaluation 20
4.2.1 VRU Safety Anticipation 20
4.2.2 Rationale Sufficiency 22
4.3 Qualitative Example 24

5. Conclusion 26
-
dc.format.extent 34 -
dc.language eng -
dc.publisher DGIST -
dc.title Toward Universal V2X Cooperation for VRU Safety: An Infrastructure-Centric Explainable Framework via Vision-Language Models -
dc.title.alternative 교통 약자 안전을 위한 보편적 V2X 협력: 비전-언어 모델 기반 인프라 중심 설명가능 프레임워크 -
dc.type Thesis -
dc.identifier.doi 10.22677/THESIS.200001016429 -
dc.description.degree Master -
dc.contributor.department Artificial Intelligence Major -
dc.date.awarded 2026-08-01 -
dc.publisher.location Daegu -
dc.description.database dCollection -
dc.citation XT.AM 정66 202608 -
dc.date.accepted 2026-07-21 -
dc.contributor.alternativeDepartment 학제학과인공지능전공 -
dc.subject.keyword Vehicle-to-Everything (V2X) cooperation, Vulnerable Road User (VRU) anticipation, Vision-Language Model, Chain-of-Thought reasoning, Knowledge Distillation -
dc.contributor.affiliatedAuthor Yunho Jeong -
dc.contributor.affiliatedAuthor Ji-Woong Choi -
dc.contributor.alternativeName 정윤호 -
dc.contributor.alternativeName Ji-Woong Choi -
dc.rights.embargoReleaseDate 2028-08-31 -
Show Simple Item Record

File Downloads

  • There are no files associated with this item.

공유

qrcode
공유하기

Total Views & Downloads

???jsp.display-item.statistics.view???: , ???jsp.display-item.statistics.download???: