Detail View

Optimization of MuMax3 by Using Claude Code: A CUDA-Graph-Based Case Study in AI-Assisted Performance Engineering

Citations

WEB OF SCIENCE

Citations

SCOPUS

Metadata Downloads

Title
Optimization of MuMax3 by Using Claude Code: A CUDA-Graph-Based Case Study in AI-Assisted Performance Engineering
Issued Date
2026-06
Citation
JOURNAL OF MAGNETICS, v.31, no.2, pp.204 - 213
Type
Article
Author Keywords
micromagnetics ; spintronics ; mumax3 ; AI-assisted
Keywords
SKYRMIONS
ISSN
1226-1750
Abstract

MuMax3 is a widely used open-source GPU-accelerated micromagnetic simulator whose computational core - CUDA kernels and cuFFT-based demagnetization convolutions - has changed little since its original release. We report a case study in which an agentic large-language-model coding assistant (Claude Code, Anthropic) was used, under continuous human supervision, to profile and optimize this mature CUDA/Go codebase. Profiling with NVIDIA Nsight Systems revealed that for small and medium grids (64 & times;64 & times;1 to 256 & times;256 & times;1 cells), 75-79 % of step time is spent in CPU-side cuLaunchKernel driver calls rather than in GPU computation, because every solver step launches roughly 27 kernels sequentially on a single CUDA stream. Building on this finding, we implemented a "split-graph" execution strategy: the time-step-independent torque-evaluation kernel sequences of each solver stage are captured once as CUDA Graphs and replayed via cudaGraphLaunch, while the time-step-dependent update, error estimate, and adaptive-step-size logic remain ordinary stream-ordered calls. The optimization is exposed transparently through Run()/Steps(), guarded by a compatibility check (constant excitation, zero thermal field, no custom field terms, time-independent material parameters, mesh size below a tunable threshold) that falls back silently to the original code path when violated. Across five solvers (Heun, RK23, RK45DP, RK56, Backward Euler), the optimization yields up to 5.4 & times; throughput for a 64 & times;64 & times;1 grid with a fixed time step, 2.4-3.6 & times; for adaptive time-stepping, decreasing smoothly to approximate to 1.0 & times; near 10(6) cells, essentially independent of which physical field terms (exchange, anisotropy, DMI) are active. All results were verified bit-for-bit identical (ndiff=0) against the unmodified code, and the full 176-script mumax3 regression suite passes with zero failures. We discuss the workflow itself - including three episodes in which the assistant autonomously diagnosed and corrected its own defects - as a template for AI-assisted optimization of legacy scientific HPC codes.

더보기
URI
https://scholar.dgist.ac.kr/handle/20.500.11750/60887
DOI
10.4283/JMAG.2026.31.2.204
Publisher
KOREAN MAGNETICS SOC
Show Full Item Record

File Downloads

  • There are no files associated with this item.

공유

qrcode
공유하기

Related Researcher

유천열
You, Chun-Yeol유천열

Department of Physics and Chemistry

read more

Total Views & Downloads

???jsp.display-item.statistics.view???: , ???jsp.display-item.statistics.download???: