Arash Lagzian

I am a visiting scholar at the National University of Singapore, under the supervision of Prof. Dianbo Liu, specializing in reasoning and knowledge representation in large language models (LLMs) and visual language models (VLMs). My current research focuses on advancing understanding in these cutting-edge fields and their applications in AI.

Previously, I worked on both image anomaly detection and natural language processing areas at Sharif University of Technology under the supervision of Prof. Hamid Beigy, where I gained a solid foundation in AI-driven visual analysis and problem-solving techniques.

I am passionate about AI research and its potential to solve complex real-world problems, and I look forward to continuing to grow in this dynamic field.

Email / Google Scholar / Linkedin / Github / CV

Publications


Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders

Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu

Accepted in npj Artificial Intelligence (2026)

The rapid development of generative AI has transformed content creation, communication, and human development. However, this technology raises profound concerns in high-stakes domains, demanding rigorous methods to analyze and evaluate AI-generated content. While existing analytic methods often treat images as indivisible wholes, real-world AI failures generally manifest as specific visual patterns that can evade holistic detection and suit more granular and decomposed analysis. Here we introduce a content analysis tool, Language-Grounded Sparse Encoders (LanSE), which decompose images into interpretable visual patterns with natural language descriptions. Utilizing interpretability modules and large multimodal models, LanSE can automatically identify visual patterns within data modalities. Our method discovers more than 5,000 visual patterns with 93% human agreement, provides decomposed evaluation outperforming existing methods, establishes the first systematic evaluation of physical plausibility, and extends to medical imaging settings. Our method's capability to extract language-grounded patterns can be naturally adapted to numerous fields, including biology and geography, as well as other data modalities such as protein structures and time series, thereby advancing content analysis for generative AI.

IRuDD: A Large-Scale Industrial Rubber Defect Detection Dataset

Arash Lagzian, Saeed Mollaee, Leila Shahi, Hamid Beigy

Accepted in IEEE ICECCME 2026 (6th International Conference on Electrical, Computer, Communications and Mechatronics Engineering)

Object detection, driven by advances in deep learning, has become a fundamental component of modern computer vision systems and serves as the core technology behind many industrial inspection tasks, including surface defect detection. In the rubber manufacturing industry, reliable quality control is critical not only for ensuring product safety and customer satisfaction, but also for improving production efficiency and reducing material waste. Despite its practical importance, large-scale real-world datasets for rubber surface inspection remain scarce. In this paper, we introduce IRuDD, a large-scale Industrial Rubber Defect Detection dataset comprising 20,828 high-resolution RGB images collected directly from production lines in real manufacturing environments. All images contain genuine surface defects arising during the production process and are annotated with bounding boxes for supervised defect localization, provided in both YOLO and COCO formats. Notably, the dataset does not rely on synthetic defects or artificial pre-processing, offering a realistic benchmark that reflects real-world conditions. To establish strong baselines and demonstrate the practical value of IRuDD, we conduct extensive experiments using a range of state-of-the-art object detection models, including CNN-based, YOLO-based, and transformer-based detectors, and report comprehensive quantitative results. By focusing on supervised defect detection under realistic constraints, IRuDD fills an important gap in industrial vision research and provides a solid foundation for future work on rubber surface inspection. The complete IRuDD dataset and benchmark code will be made publicly available upon acceptance at: https://github.com/arashlagzian/IRuDD.

MIRAGE: Multi-perspective Inference-time Reasoning via Agent-Guided Exploration

Arash Lagzian, Srinivas Anumasa, Dianbo Liu

ICML Workshop on Multi-Agent Systems in the Era of Foundation Models: Opportunities, Challenges and Futures (2025)

Recent advances in Large Language Models (LLMs) have revolutionized artificial intelligence and how human interact with AIs. Despite impressive advancements, LLMs struggle with complex mathematical, scientific, and logical tasks. Inspired by human cognitive flexibility—our ability to dynamically switch mental perspectives—we propose MIRAGE (Multi-perspective Inference-time Reasoning via Agent-Guided Exploration), a novel inference-time creative thinking framework. MIRAGE includes a Selector that prioritizes effective conceptual perspectives (e.g., algebraic, probabilistic) and a Reasoner that sequentially solves tasks until a confident solution emerges, otherwise aggregating multiple perspectives. Tested on GSM8K, MATH500, MMLU-Pro, and Game-of-24 benchmarks, MIRAGE consistently outperforms methods like Chain-of-Thought and diverse prompting ensembles, significantly boosting accuracy with minimal inference overhead, providing a scalable solution for practical applications.

Multi-Novelty: Improve the Diversity and Novelty of Contents Generated by Large Language Models via Inference-time Multi-Views Brainstorming

Arash Lagzian, Srinivas Anumasa, Dianbo Liu

Erlier version accepted in ICLR Workshop Towards Agentic AI for Science: Hypothesis Generation, Comprehension, Quantification, and Validation (2025)

Large Language Models (LLMs) demonstrate remarkable proficiency in generating accurate and fluent text. However, they often struggle with diversity and novelty, leading to repetitive or overly deterministic responses. These limitations stem from constraints in training data, including gaps in specific knowledge domains, outdated information, and an overreliance on textual sources. Such shortcomings reduce their effectiveness in tasks requiring creativity, multi-perspective reasoning, and exploratory thinking, such as LLM based AI scientist agents and creative artist agents . To address this challenge, we introduce inference-time multi-view brainstorming method, a novel approach that enriches input prompts with diverse perspectives derived from both textual and visual sources, which we refere to as "Multi-Novelty". By incorporating additional contextual information as diverse starting point for chain of thoughts, this method enhances the variety and creativity of generated outputs. Importantly, our approach is model-agnostic, requiring no architectural modifications and being compatible with both open-source and proprietary LLMs. We evaluate our method and framework on over 909,500 generated outputs from various well-known LLMs, demonstrating significant improvements in output diversity and novelty while maintaining quality and relevance

KhabarChin: Automatic Detection of Important News in the Persian Language

Hamed Hematian Hemati, Arash Lagzian, Moein Salimi Sartakhti, Hamid Beigy, Ehsaneddin Asgari

Arxiv

Being aware of important news is crucial for staying informed and making well-informed decisions efficiently. Natural Language Processing (NLP) approaches can significantly automate this process. This paper introduces the detection of important news, in a previously unexplored area, and presents a new benchmarking dataset (Khabarchin) for detecting important news in the Persian language. We define important news articles as those deemed significant for a considerable portion of society, capable of influencing their mindset or decision-making. The news articles are obtained from seven different prominent Persian news agencies, resulting in the annotation of 7,869 samples and the creation of the dataset. Two challenges of high disagreement and imbalance between classes were faced, and solutions were provided for them. We also propose several learning-based models, ranging from conventional machine learning to state-of-the-art transformer models, to tackle this task. Furthermore, we introduce the second task of important sentence detection in news articles, as they often come with a significant contextual length that makes it challenging for readers to identify important information. We identify these sentences in a weakly supervised manner.

Teaching Experience

Sharif University of Technology

Head & Teaching Assistant