MMAD: The Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
A comprehensive benchmark covering anomaly discrimination, localization, description, object analysis, and defect reasoning.
I am a Ph.D. candidate in Computer Science at the Southern University of Science and Technology (SUSTech), advised by Prof. Feng Zheng.
My research focuses on industrial anomaly detection, multimodal foundation models, and embodied AI. I aim to build inspection systems that can compare visual evidence, explain anomalies, and actively acquire useful observations.
I received my M.Sc. in Computer Science from SUSTech in 2023 and my B.Sc. in Computer Science from Xi’an Jiaotong University in 2020, where I was advised by Prof. Danfeng Shan. I was awarded the National Scholarship for Doctoral Students in 2025.
I was a visiting Ph.D. student at A*STAR IHPC–CFAR in Singapore, supervised by Joey Tianyi Zhou, and previously worked as a research intern at Tencent YouTu Lab and Huawei.
* Equal contribution. Unaccepted work is listed only when I am first or co-first author.
A comprehensive benchmark covering anomaly discrimination, localization, description, object analysis, and defect reasoning.
Extends patch-level denoising with multiple discriminators for robust fully unsupervised inspection under high noise ratios.
A tuning-free diffusion framework for precise local customization jointly guided by reference images and text.
Approximates client and server posteriors online to reduce aggregation error and local forgetting under heterogeneous data.
A patch-level noise discriminator and soft-weighted coreset that make normal-only anomaly detection robust to contaminated training data.
A comparison-aware multimodal assistant that learns fine-grained anomaly perception and explanation from paired visual evidence.
Frames inspection as active robotic perception, using normal 3D Gaussian templates for pose-aligned reference rendering and view selection.
A training-free anomaly agent that challenges its first judgment through multi-turn, tool-grounded refutation.
Organizes recent progress into five foundation-model-driven generalization mechanisms and outlines open challenges.
A unified PCB inspection benchmark and progressive curriculum for domain-specialized multimodal reasoning.
Introduces positive–negative semantic groups and opposition-based objectives for negation-aware visual grounding.
MINT-AD learns class-aware query embeddings to reduce inter-class interference in a unified anomaly detector.
For research collaboration or open-source work, email jiangx2020@mail.sustech.edu.cn.