CORTEXA
← Browse
crossrefMachine Learning and Knowledge Extraction2026-06-18Cited by 0

XAI2Brain: A Perspective on Mechanistic Interpretability for Brain–AI Alignment

Richard Jiang, Yongchen Zhou, Boyuan Wang, Plamen Angelov, Qiang Ni

The convergence of artificial intelligence (AI), explainable AI (XAI), and neuroscience is fostering new opportunities for understanding both machine and biological intelligence through interpretable and human-centered learning paradigms. In this Perspective, we introduce XAI2Brain as a conceptual framework for brain–AI alignment, positioning mechanistic interpretability as an intermediate layer connecting neural network representations, human understanding, and neuroscience-inspired AI design. Rather than viewing XAI solely as a post hoc transparency tool, we emphasize its emerging role in enabling mechanistic analysis of internal model representations, concept-level reasoning, and interactive human–AI alignment. We define XAI2Brain as a multi-level conceptual framework rather than a deployable system, explicitly aimed at structuring brain–AI alignment across representation-level, mechanism-level, and interaction-level perspectives. We survey the evolution of XAI methodologies—from feature attribution and concept-based explanations to mechanistic and human-centric interpretability approaches—and discuss how these methods may support bidirectional knowledge transfer between AI systems and cognitive neuroscience. Importantly, we adopt a cautious stance on brain–AI analogy, explicitly recognizing that artificial neural representations are not equivalent to biological neural representations, and instead focusing on functional and informational correspondences rather than structural equivalence. Unlike conventional human-in-the-loop or reinforcement learning from human feedback paradigms that primarily optimize behavioral outputs, XAI2Brain focuses on cognitively interpretable and mechanistically grounded alignment between AI systems and human reasoning processes. This alignment promotes interactive human-in-the-loop intelligence, empowering humans to comprehend, guide, and refine AI systems, while enabling AI systems to better interpret human instructions, intentions, and contextual reasoning. We further discuss the challenges of scaling explainability to large generative and multimodal models, including issues of interpretability robustness, cognitive compatibility, evaluation, and ethical accountability. We also highlight key limitations of current mechanistic interpretability methods, including explanation instability, representation superposition, and lack of causal guarantees, underscoring that these challenges remain open research problems. Rather than proposing a complete artificial brain architecture, this Perspective outlines a research roadmap toward more interpretable, adaptive, and neuroscience-inspired AI systems capable of supporting future brain–AI integration and collaborative intelligence. We additionally clarify that this work follows a narrative perspective review methodology with structured thematic synthesis of the literature. By framing explainability as a bridge between mechanistic AI understanding, cognitive science, and human-centered interaction, XAI2Brain highlights the importance of interpretable alignment for the next generation of brain-inspired AI systems.

View free PDFSource page

Related papers

crossrefMachine Learning and Knowledge Extraction2024-11-06Cited by 1

Adaptive AI Alignment: Established Resources for Aligning Machine Learning with Human Intentions and Values in Changing Environments

Stephen Fox

AI Alignment is a term used to summarize the aim of making artificial intelligence (AI) systems behave in line with human intentions and values. There has been little consideration in previous AI Alignment studies of the need for AI Alignment to be adaptive in order to contribute…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2025-02-06Cited by 92

Advancing AI Interpretability in Medical Imaging: A Comparative Analysis of Pixel-Level Interpretability and Grad-CAM Models

Mohammad Ennab, Hamid Mcheick

This study introduces the Pixel-Level Interpretability (PLI) model, a novel framework designed to address critical limitations in medical imaging diagnostics by enhancing model transparency and diagnostic accuracy. The primary objective is to evaluate PLI’s performance against Gr…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2024-03-10Cited by 40

Augmenting Deep Neural Networks with Symbolic Educational Knowledge: Towards Trustworthy and Interpretable AI for Education

Danial Hooshyar, Roger Azevedo, Yeongwook Yang

Artificial neural networks (ANNs) have proven to be among the most important artificial intelligence (AI) techniques in educational applications, providing adaptive educational services. However, their educational potential is limited in practice due to challenges such as the fol…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-05-01

Morphology-Aware Multi-Scale Deep Representation Learning for Interpretable Knowledge Extraction in Brain Tumor MRI

Helala AlShehri, Mariam Busaleh

Robust brain tumor classification from magnetic resonance imaging (MRI) remains challenging due to complex structural heterogeneity and subtle inter-class variability. Beyond predictive accuracy, conventional convolutional neural networks predominantly rely on texture-dominant fe…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-07-02

An Interpretable Stacked Deep Learning Model for Diagnosis of Brain Tumor with Transparent Learning Dynamics

K. Kaivalya, N. Thirupathi Rao, Aditya Pal, Hari Mohan Rai, B. Omkar Lakshmi Jagan

The diagnosis and treatment planning for brain tumors remain a complex task in medical imaging, largely due to the intricate structure of such abnormalities. This study introduces an interpretable stacked deep learning framework consisting of three sequential stages: (i) tumor se…

View free PDFSource page
crossrefMachine Learning and Knowledge Extraction2026-06-18

XTrail-ID: An Explainable AI Human Footprint Trail Identification on Soil Substrate Using Unsupervised Machine Learning from UAV Imagery

Wazha Mmereki, Rodrigo S. Jamisola, Zoe C. Jewell, Tinao Petso, Oduetse Matsebe, Sky K. Alibhai

This paper investigates human–AI collaboration through explainable AI where we interpret the results of barefoot print clustering using unsupervised machine learning. This can be used to identify the number of individuals from barefoot prints on the ground as a tool in forensics…

View free PDFSource page