Detail View

Attention Dilution and Overlap Resolver for Complex Prompts in Text-to-Image Diffusion Models

Citations

WEB OF SCIENCE

Citations

SCOPUS

Metadata Downloads

Title
Attention Dilution and Overlap Resolver for Complex Prompts in Text-to-Image Diffusion Models
Alternative Title
동적 환경에서의 의미론적으로 일관된 시각적 세분화
DGIST Authors
Seungjun HyunSunghoon Im
Advisor
임성훈
Issued Date
2026
Awarded Date
2026-08-01
Type
Thesis
Description
Image Generation, Diffusion Models, Semantic Misalignment, Attention Overlap, Attention Dilution
Abstract

Text-to-image diffusion models have recently achieved impressive advances, generating high-quality and realistic images. Nevertheless, they continue to struggle with semantic misalignment, particularly when parsing complex prompts that involve multiple objects and diverse attributes. In this work, we analyze the cross-attention behavior of text-to-image diffusion models and identify two primary contributors to semantic misalignment: cross- attention overlap and cross-attention dilution. Building on these insights, we introduce ADOR, a training-free framework that mitigates semantic misalignment in a single forward pass without relying on external information. ADOR comprises two complementary modules: the Attention Overlap Disentangler (AO-Disentangler) and the Attention Dilution Reviver (AD-Reviver). The AO-Disentangler employs distance-based masking to reduce cross- attention overlap between noun phrases, thereby improving the disentanglement of object–attribute pairs. The AD-Reviver addresses the decline in average cross-attention intensity that arises with longer prompts by applying L2 normalization or selective amplification, ensuring that semantic concepts remain represented during generation. Our evaluation demonstrates that ADOR achieves state-of-the-art performance on standard benchmarks while remaining computationally efficient.
Keywords: Image Generation, Diffusion Models, Semantic Misalignment, Attention Overlap, Attention Dilution|시각적 세분화는 이미지 이해부터 복잡한 비디오 분석에 이르기까지 다양한 사용자 요구에 의해 주도되는 컴퓨터 비전의 근본적인 작업입니다. 이러한 응용 분야 전반의 핵심 과제는 의미론적 일관성을 유지하는 것이며, 특히 모델이 도메인 변화, 시간적 단절, 복잡한 언어를 특정 비디오 콘텐츠에 정렬해야 하는 어려움에 직면하는 동적이고 실제적인 환경에서 이는 더욱 중요합니다. 우리는 의미론적 불일치의 세 가지 핵심 영역을 점진적으로 다루며, 지도 신호를 지능적으로 정제함으로써 견고한 세분화가 달성될 수 있음을 입증합니다. 첫째, 모델이 합성 데이터로부터 여러 실제 도메인으로 일반화하는 데 실패하는 이미지 수준의 의미론적 세분화를 위한 일반화 문제를 다룹니다. 우리는 단일 네트워크에서 다양한 타겟 도메인 분포를 시뮬레이션하여 광범위한 시각적 스타일을 포괄하는 직접 적응 프레임워크를 제안합니다. 이 프레임워크는 의미론적으로 모호한 영역을 식별하고 필터링하여 모델이 도메인 불변의 일관된 특징을 학습하도록 강제하는 양방향 적응형 영역 선택 전략을 구현함으로써 학습 신호를 동시에 정제합니다. 둘째, 시간적 일관성 유지가 가장 중요한 비디오 인스턴스 세분화로 이 원칙을 확장합니다. 가려짐 현상 중 오래된 특징으로 인한 메모리 오염으로 발생하는 추적 실패를 해결하기 위해, 우리는 새로운 메모리 관리 시스템을 제안합니다. 우리는 특징을 평균화하는 대신 가려짐 중에 최신의 유효한 객체 상태만을 저장하고 이 상태를 유지하는 메모리 메커니즘을 활용하여 일관성을 보장합니다. 또한, 기존 객체와 새로 나타나는 객체를 분리하여 관리함으로써 모호성을 해결하는 연관 전략으로 이를 보완합니다. 셋째, 지시형 세분화에서 비디오와 텍스트 간의 의미론적 정렬이라는 복잡한 과제를 다룹니다. 우리는 모델이 행동 기반 텍스트를 관련 없는 정적 프레임과 연관시키도록 강제되는, '의미론적 모순'이라 명명한 기존 학습의 근본적인 결함을 식별합니다. 우리는 명시적인 시간 주석을 사용하여 교차 양식 신호를 정제하는 시간 기반 학습 프레임워크를 제안합니다. 우리는 언어적 설명이 능동적으로 참(true)인 시공간적 세그먼트에서만 손실을 적용하는 선택적 지도 메커니즘을 통해 이를 달성합니다. 요약하면, 이러한 기여들은 이미지 수준의 일반화에서 시간적 일관성, 그리고 시공간-언어 정렬로 이어지는 일관된 진행 과정을 보여주며, 의미론적 불일치를 방지하기 위한 지도 신호의 세심한 정제가 동적 환경에서 의미론적으로 일관된 시각적 세분화를 달성하기 위한 핵심 원칙임을 확립합니다.

더보기
Table Of Contents
I. Introduction 1
II. RELATED WORK 3
2.1 Text-to-Image Diffusion Models 3
2.2 Semantic Misalignment 3
III. METHOD 5
3.1 Observation 5
3.2 Overview 6
3.3 Attention Overlap Disentangler 8
3.4 Attention Dilution Reviver 10
IV. EXPERIMENTS 12
4.1 Experiment settings 12
4.2 Comparison with Other Models on T2I-CompBench 13
4.3 Comparison under a Shared Base Model 16
4.4 Ablation on Overlap Separator and Dilution Restorer 17
V. Conclusion 19
VI. References 20
VII.요약문 24
URI
https://scholar.dgist.ac.kr/handle/20.500.11750/60819
http://dgist.dcollection.net/common/orgView/200001006675
DOI
10.22677/THESIS.200001006675
Degree
Master
Department
Department of Electrical Engineering and Computer Science
Publisher
DGIST
Show Full Item Record

File Downloads

  • There are no files associated with this item.

공유

qrcode
공유하기

Total Views & Downloads

???jsp.display-item.statistics.view???: , ???jsp.display-item.statistics.download???: