Xuwang Yin

Independent AI researcher working at the intersection of energy-based models and adversarial robustness.

Unified discriminative-generative modeling

Generative AI and discriminative AI have traditionally been two separate worlds—different models, different training, different applications. But a model that truly understands should be able to both recognize and imagine. Building on the energy-based learning framework, my research unifies them in a single model, where classification is grounded in the model's ability to generate, and decisions can be explained through counterfactual examples.

In our ICLR 2026 paper, this approach scales energy-based models to ImageNet 256×256 for the first time, with a single model that generates, classifies, detects out-of-distribution inputs, and explains its decisions.

AI safety and interpretability

Previously at the Center for AI Safety, I worked on making LLMs transparent, controllable, and honest: reading and steering their internal representations (Representation Engineering), evaluating their robustness against automated red teaming (HarmBench, ICML 2024), analyzing the value systems that emerge in them (Utility Engineering, NeurIPS 2025), and measuring whether they lie under pressure (MASK, NeurIPS 2026).