MMAD: The Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
A comprehensive benchmark covering anomaly discrimination, localization, description, object analysis, and defect reasoning.
I am a Ph.D. candidate in Computer Science at the Southern University of Science and Technology (SUSTech), advised by Prof. Feng Zheng.
My research focuses on industrial anomaly detection, multimodal large language models, and agentic systems. I aim to build inspection systems that can compare visual evidence, explain anomalies, and actively acquire useful observations. My work has received 550+ citations, and the Awesome Industrial Anomaly Detection repository I maintain has 4k+ stars.
I received my M.Sc. in Computer Science from SUSTech in 2023 and my B.Sc. in Computer Science from Xi’an Jiaotong University in 2020, where I was advised by Prof. Danfeng Shan. I was awarded the National Scholarship for Doctoral Students in 2025.
From Oct. 2025 to Apr. 2026, I was a joint-training Ph.D. student at A*STAR IHPC–CFAR in Singapore, supervised by Joey Tianyi Zhou. Before that, I spent nearly four years as a research intern at Tencent YouTu Lab (2022–2025), and earlier interned at Huawei on continual and federated learning.
* Equal contribution. Unaccepted work is listed only when I am first or co-first author.
A comprehensive benchmark covering anomaly discrimination, localization, description, object analysis, and defect reasoning.
Extends patch-level denoising with multiple discriminators for robust fully unsupervised inspection under high noise ratios.
A tuning-free diffusion framework for precise local customization jointly guided by reference images and text.
Approximates client and server posteriors online to reduce aggregation error and local forgetting under heterogeneous data.
A patch-level noise discriminator and soft-weighted coreset that make normal-only anomaly detection robust to contaminated training data.
A comparison-aware multimodal assistant that learns fine-grained anomaly perception and explanation from paired visual evidence.
Frames inspection as active robotic perception, using normal 3D Gaussian templates for pose-aligned reference rendering and view selection.
A training-free anomaly agent that challenges its first judgment through multi-turn, tool-grounded refutation.
Organizes recent progress into five foundation-model-driven generalization mechanisms and outlines open challenges.
A unified PCB inspection benchmark and progressive curriculum for domain-specialized multimodal reasoning.
Introduces positive–negative semantic groups and opposition-based objectives for negation-aware visual grounding.
MINT-AD learns class-aware query embeddings to reduce inter-class interference in a unified anomaly detector.
For research collaboration or open-source work, email jiangx2020@mail.sustech.edu.cn.