Detail View
Controllable Generation and Identity-Preserving Editing on Vision Foundation Models for Medical Imaging
Citations
WEB OF SCIENCE
Citations
SCOPUS
| DC Field | Value | Language |
|---|---|---|
| dc.contributor.advisor | 문인규 | - |
| dc.contributor.author | Dongkyu Won | - |
| dc.date.accessioned | 2026-09-01T19:29:43Z | - |
| dc.date.available | 2026-09-01T19:29:43Z | - |
| dc.date.issued | 2026 | - |
| dc.identifier.uri | https://scholar.dgist.ac.kr/handle/20.500.11750/60731 | - |
| dc.identifier.uri | http://dgist.dcollection.net/common/orgView/200001012168 | - |
| dc.description | Medical Image Synthesis, Vision Foundation Models, Self-Supervised Denoising, Latent Diffusion, Direct Latent Editing | - |
| dc.description.abstract | 최근 심층 학습의 발전은 의료 영상 분석에 큰 변화를 가져왔다. 그러나 의료 영상의 세 가지 구조적 특성, 즉 짝지어진 학습 데이터의 부족, 임상적으로 중요한 소견의 희소성, 그리고 전문의 주석의 높은 비용으로 인해 생성 모델의 직접적인 적용은 쉽지 않다. 본 학위논문에서는 잡음 생성, 영상 생성, 파운데이션 모델 기반 생성, 그리고 정체성 보존 잠재 편집의 네 단계에 걸쳐 생성 모델 기반 접근법을 제안한다. 먼저, 저선량 CT 잡음 제거를 위해 환자별 잡음 모델 앙상블을 기반으로 추론 시점에 Pseudo-LDCT/NDCT 영상쌍을 합성하고, 소수의 짝지어진 환자 데이터만으로도 사전학습된 디노이저를 자기지도 방식으로 추가 학습하는 방법을 제안하였다. 다음으로, 경추 척추협착증 등급 분류를 위해 소규모 임상 코호트에 학습한 클래스 조건부 잠재 확산 모델로 등급별 디스크 패치를 생성하고 이를 이용해 Focal Modulation Network 분류기를 증강하여 일반 등급과 희소 등급 모두에서 성능을 향상시켰다. 이어서, 의료 영상 생성기의 VAE/VQ-GAN 인코더-디코더 병목을 제거하기 위해 학습된 1단계를 동결된 비전 파운데이션 모델 특징 공간으로 대체하고 의료 도메인 텍스트 인코더로 조건화된 흐름 정합을 그 공간에서 직접 수행함으로써, 인코더-디코더 재학습 없이 자유 텍스트 보고서로부터 흉부 X선 영상을 합성하는 생성기를 제안하였다. 마지막으로, 기존 추론 시점 편집 기법들이 무너지는 비전 파운데이션 모델 기반 생성기 위에서 정체성을 유지하는 반사실적 편집을 가능하게 하기 위해, 흐름 모델 호출이나 역변환, 가중치 갱신 없이 인코딩된 잠재 벡터의 끝점에 사전 추출한 개념 방향을 더하는 폐쇄형 사전 궤도 편집 방법인 Direct Latent Editing을 제안하였다. 동일한 연산이 얼굴 속성, 흉부 X선 소견, 뇌 자기공명영상 진단 방향 모두에 그대로 일반화된다. 본 학위논문이 제안하는 방법들은 비전 파운데이션 모델 기반 의료 영상 분석에서 통제 가능 생성과 정체성 보존 편집의 활용 범위를 확장한다.|Recent advances in deep learning have changed medical image analysis. However, three structural properties of medical images, i.e., scarce paired training data, rare clinically important findings, and costly expert annotation, make the direct application of generative models difficult. In this thesis, we developed generative-model-based approaches across noise generation, image generation, foundation-model-based generation, and identity-preserving latent editing. For low-dose CT, we proposed a self-supervised denoising method built around an ensemble of subject-specific noise models that synthesizes Pseudo-LDCT/NDCT pairs at inference time, so that a small number of paired subjects can refine a pretrained denoiser. For cervical-stenosis grading, we trained a class-conditional Latent Diffusion Model on a small institutional cohort to generate grade-conditioned disc patches that augment a Focal Modulation Network classifier and improve both common and rare-grade performance. To remove the VAE/VQ-GAN encoder–decoder bottleneck of medical-image generators, we proposed a chest-radiograph generator that replaces the learned first stage with a frozen Vision Foundation Model feature space and runs flow matching directly in those features, with a medical-domain text encoder for conditioning. This generator synthesizes full chest radiographs from free-text reports, with no VAE encoder–decoder stage trained at any point. To allow identity-preserving counterfactual edits on Vision-Foundation-Model-based generators where existing inference-time editors collapse, we proposed Direct Latent Editing, a closed-form pre-trajectory edit that adds an offline-extracted concept direction to the encoded latent at the sampling endpoint, with no flow-model forwards, inversion, or weight updates. The same editing operation applies across face attributes, chest-radiograph findings, and brain-MRI diagnostic directions. Together, these methods extend controllable generation and identity-preserving editing for medical image analysis on Vision Foundation Models. | - |
| dc.description.tableofcontents | Ⅰ. Introduction 1 1.1 Background and Motivation 1 1.2 Contributions and Outline 2 1.3 Publications 3 1.3.1 Excluded Research 4 Ⅱ. Low-Dose CT Denoising via Pseudo-CT Image Pairs 5 2.1 Introduction 5 2.2 Related Work 6 2.2.1 LDCT Denoising 6 2.2.2 Self-Supervised Denoising 6 2.3 Method 7 2.3.1 Overview 7 2.3.2 Pretraining the Noise Models 7 2.3.3 Pseudo-CT Image Generation 8 2.3.4 Pseudo-CT Self-Supervision 9 2.3.5 Implementation Details 9 2.4 Experiments 10 2.4.1 Dataset 10 2.4.2 Baselines 10 2.4.3 Quantitative Results 10 2.4.4 Effect of the Noise Model 11 2.5 Discussion 11 2.6 Conclusions 13 Ⅲ. Cervical Stenosis Grading on Sagittal MRI via Conditional Latent Diffusion 15 3.1 Introduction 15 3.2 Related Work 16 3.2.1 Deep-Learning Models for Spinal-Stenosis Grading 16 3.2.2 Latent Diffusion Models 16 3.2.3 Focal Modulation Networks 17 3.3 Method 17 3.3.1 Dataset and Preprocessing 18 3.3.2 Conditional LDM Augmentation 18 3.3.3 Classifier: Focal Modulation Network 19 3.4 Experiments 20 3.4.1 Setup 20 3.4.2 Baselines 20 3.4.3 Quantitative Results 20 3.4.4 Qualitative Samples from the Conditional LDM 21 3.5 Discussion 23 3.6 Conclusions 23 Ⅳ. VAE-Free Chest Radiograph Generation via Flow Matching in Frozen DINOv3 Feature Space 24 4.1 Introduction 24 4.2 Related Work 25 4.2.1 Text-Conditioned Chest X-Ray Generation 25 4.2.2 VAE-Free Generation with Vision Foundation Models 25 4.2.3 Self-Supervised Features in Medical Imaging 25 4.3 Method 26 4.4 Experiments 28 4.4.1 Dataset 28 4.4.2 Baselines 28 4.4.3 Evaluation Metrics 28 4.4.4 Clinical Fidelity 29 4.4.5 Image Quality 29 4.4.6 Per-Class Clinical Fidelity 29 4.4.7 Qualitative Results 30 4.5 Conclusions 30 Ⅴ. Direct Latent Editing for VFM-Based Image Generators 32 5.1 Introduction 32 5.2 Related Work 34 5.2.1 VFM-Based Generators 34 5.2.2 Inference-Time Editing in Diffusion / Flow Generators 34 5.3 Method: Direct Latent Editing (DLE) 34 5.3.1 Preliminaries 34 5.3.2 DLE: Two-Phase Closed-Form Procedure 37 5.3.3 Why DLE Succeeds: Acting at the Sampling Endpoint 39 5.4 Experiments and Results 39 5.4.1 Cross-Domain Setup 39 5.4.2 CelebA Face Attribute Editing 42 5.4.3 MIMIC-CXR Chest Radiograph Editing 43 5.4.4 ADNI-1 Brain MRI Editing 44 5.4.5 Ablations 46 5.4.6 Extended Qualitative Results 50 5.5 Discussion 55 5.5.1 Positioning: What Is and Is Not New 55 5.5.2 Why a VFM Latent Space Makes This Work 56 5.5.3 Why Single-Direction Addition Suffices on the VFM Latent 57 5.5.4 Edit-Time Cost 57 5.5.5 Limitations 58 5.6 Conclusions 59 Ⅵ. Concluding Remarks 61 6.1 Conclusion 61 6.2 Future Work 62 References 64 |
- |
| dc.format.extent | 72 | - |
| dc.language | eng | - |
| dc.publisher | DGIST | - |
| dc.title | Controllable Generation and Identity-Preserving Editing on Vision Foundation Models for Medical Imaging | - |
| dc.type | Thesis | - |
| dc.identifier.doi | 10.22677/THESIS.200001012168 | - |
| dc.description.degree | Doctor | - |
| dc.contributor.department | Department of Robotics and Mechatronics Engineering | - |
| dc.contributor.coadvisor | Sang Hyun Park | - |
| dc.date.awarded | 2026-08-01 | - |
| dc.publisher.location | Daegu | - |
| dc.description.database | dCollection | - |
| dc.citation | XT.RD 원25 202608 | - |
| dc.date.accepted | 2026-07-21 | - |
| dc.contributor.alternativeDepartment | 로봇및기계전자공학과 | - |
| dc.subject.keyword | Medical Image Synthesis, Vision Foundation Models, Self-Supervised Denoising, Latent Diffusion, Direct Latent Editing | - |
| dc.contributor.affiliatedAuthor | Dongkyu Won | - |
| dc.contributor.affiliatedAuthor | Inkyu Moon | - |
| dc.contributor.affiliatedAuthor | Sang Hyun Park | - |
| dc.contributor.alternativeName | 원동규 | - |
| dc.contributor.alternativeName | Inkyu Moon | - |
| dc.contributor.alternativeName | 박상현 | - |
| dc.rights.embargoReleaseDate | 2029-08-31 | - |
File Downloads
- There are no files associated with this item.
공유
Total Views & Downloads
???jsp.display-item.statistics.view???: , ???jsp.display-item.statistics.download???:
