Transformer-based Latent Diffusion Model for Inverse Design of Phononic Crystals with Structural Variability
본문
- Conference
- Asian Congress of Structural and Multidisciplinary Optimization (ACSMO) 2026
- Date
- 2026-05-18
- Presentation Type
- Oral
Abstract
This study introduces a robust inverse design framework for phononic crystals (PnCs) with diverse structural configurations, utilizing an advanced approach based on Latent Diffusion Transformers (LDT). Traditional deep learning-based inverse design methods, such as Conditional Generative Adversarial Networks (CGANs) and Conditional Variational Autoencoders (CVAEs), often struggle with training instability, mode collapse, and the production of low-fidelity designs. To overcome these critical bottlenecks, the proposed framework leverages the latent diffusion process, which not only enhances generative stability and quality but also effectively addresses the non-uniqueness problem—a fundamental challenge in inverse design where multiple structural geometries can correspond to the same target property.
A key architectural strength of this research lies in the adoption of a Transformer-based backbone. Unlike fixed-output architectures that limit design expressiveness to a specific scale or resolution, the Transformer’s attention mechanism and padding capabilities allow for the handling of variable lattice counts. This flexibility enables the generation of supercells with diverse scales and configurations, providing a more versatile tool for material scientists. Furthermore, this study pioneers an expanded conditioning scheme by introducing a three-category band mask that includes defect bands in addition to conventional pass bands and band gaps. By incorporating defect-induced spectral features into the conditioning process, the model can solve significantly more complex design problems involving sophisticated topological features that were previously difficult to automate.
To ensure precise and high-fidelity structural generation, a hybrid conditioning strategy is implemented. This strategy integrates Feature-wise Linear Modulation (FiLM) for injecting material properties and lattice specifications, ensuring these parameters are consistently applied throughout the generative process. Simultaneously, cross-attention mechanisms combined with Classifier-Free Guidance (CFG) are employed to condition the model on complex band masks, allowing for a fine-tuned balance between design innovation and adherence to target specifications.
Validation results demonstrate that the Transformer-based latent diffusion framework exhibits marked superiority over prevalent models, generating more accurate, feasible, and innovative PnC designs even under complex structural constraints. By successfully navigating intricate design spaces and handling high-dimensional variability, this research sets a new benchmark in the field of deep learning-based metamaterial design. In conclusion, the proposed framework not only provides a robust and effective tool for the automated design of phononic crystals but also showcases the immense potential of generative AI to revolutionize traditional optimization processes in structural engineering and material science.
