Publication

Hyperautonomy Artificial Intelligence Lab

Transformer-based Latent Diffusion Model for Inverse Design of Phononic Crystals with Structural Variability

본문

Conference
Asian Congress of Structural and Multidisciplinary Optimization (ACSMO) 2026
Author
Donghyu Lee, Taehun Kim, Juhwan Han, Sayhee Kim, Soo-Ho Jo, Byeng Dong Youn
Date
2026-05-18
Presentation Type
Oral

Abstract


This study introduces a robust inverse design framework for phononic crystals (PnCs) with diverse structural configurations,  utilizing  an  advanced  approach  based  on  Latent  Diffusion  Transformers  (LDT). Traditional  deep learning-based inverse design methods, such as Conditional Generative Adversarial Networks (CGANs) and Conditional Variational Autoencoders (CVAEs), often struggle with training instability, mode collapse, and the production of low-fidelity designs. To overcome these critical bottlenecks, the proposed framework leverages the latent diffusion process, which not only enhances generative stability and quality but also effectively addresses the non-uniqueness problem—a fundamental challenge in inverse design where multiple structural geometries can correspond to the same target property.

A key architectural strength of this research lies in the adoption of a Transformer-based backbone. Unlike fixed-output architectures that limit design expressiveness to a specific scale or resolution, the Transformer’s attention mechanism and padding capabilities allow for the handling of variable lattice counts. This flexibility enables the generation of supercells with diverse scales and configurations, providing a more versatile tool for material scientists. Furthermore, this study pioneers an expanded conditioning scheme by introducing a three-category band mask that includes defect bands in addition  to  conventional  pass  bands  and  band  gaps.  By  incorporating  defect-induced  spectral  features  into  the conditioning  process,  the  model  can  solve  significantly  more  complex  design  problems  involving  sophisticated topological features that were previously difficult to automate.

To ensure precise and high-fidelity structural generation, a hybrid conditioning strategy is implemented. This strategy integrates Feature-wise Linear Modulation (FiLM) for injecting material properties and lattice specifications, ensuring these parameters are consistently applied throughout the generative process. Simultaneously, cross-attention mechanisms combined with Classifier-Free Guidance (CFG) are employed to condition the model on complex band masks, allowing for a fine-tuned balance between design innovation and adherence to target specifications.

Validation results demonstrate that the Transformer-based latent diffusion framework exhibits marked superiority over prevalent models, generating more accurate, feasible, and innovative PnC designs even under complex structural constraints. By successfully navigating intricate design spaces and handling high-dimensional variability, this research sets a new benchmark in the field of deep learning-based metamaterial design. In conclusion, the proposed framework not only provides a robust and effective tool for the automated design of phononic crystals but also showcases the immense potential of generative AI to revolutionize traditional optimization processes in structural engineering and material science.