Detail View

Social Reasoning-Aware Trajectory Prediction via Multimodal Language Model

Citations

WEB OF SCIENCE

Citations

SCOPUS

Metadata Downloads

Title
Social Reasoning-Aware Trajectory Prediction via Multimodal Language Model
Issued Date
2026-06
Citation
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, v.48, no.6, pp.6035 - 6052
Type
Article
Author Keywords
Trajectory ; Numerical models ; Predictive models ; Cognition ; Pedestrians ; Analytical models ; Context modeling ; Visualization ; Training ; Mathematical models ; Multimodal language model ; multi-task training ; social reasoning ; multi-agent trajectory forecasting
Keywords
BEHAVIORS
ISSN
0162-8828
Abstract

Recent advancements in language models have demonstrated its capacity of context understanding and generative representations. Leveraged by these developments, we propose a novel multimodal trajectory predictor based on a vision-language model, named VLMTraj, which fully takes advantage of the prior knowledge of multimodal large language models and the human-like reasoning across diverse modality information. The key idea of our model is to reframe the trajectory prediction task into a visual question answering format, using historical information as context and instructing the language model to make predictions in a conversational manner. Specifically, we transform all the inputs into a natural language style: historical trajectories are converted into text prompts, and scene images are described through image captioning. Additionally, visual features from input images are also transformed into tokens via a modality encoder and connector. The transformed data is then formatted to be used in a language model. Next, in order to guide the language model in understanding and reasoning high-level knowledge, such as scene context and social relationships between pedestrians, we introduce an auxiliary multi-task question and answers. For training, we first optimize a numerical tokenizer with the prompt data to effectively separate integer and decimal parts, allowing us to capture correlations between consecutive numbers in the language model. We then train our language model using all the visual question answering prompts. During model inference, we implement both deterministic and stochastic prediction methods through beam-search-based most-likely prediction and temperature-based multimodal generation. Our VLMTrajvalidates that the language-based model can be a powerful pedestrian trajectory predictor, and outperforms existing numerical-based predictor methods. Extensive experiments show that VLMTrajcan successfully understand social relationships and accurately extrapolate the multimodal futures on public pedestrian trajectory prediction benchmarks.

더보기
URI
https://scholar.dgist.ac.kr/handle/20.500.11750/60874
DOI
10.1109/TPAMI.2025.3582000
Publisher
IEEE COMPUTER SOC
Show Full Item Record

File Downloads

  • There are no files associated with this item.

공유

qrcode
공유하기

Related Researcher

배인환
Bae, Inhwan배인환

Department of Electrical Engineering and Computer Science

read more

Total Views & Downloads

???jsp.display-item.statistics.view???: , ???jsp.display-item.statistics.download???: