CORTEXA
← Browse
crossrefJournal of Imaging2025-07-18Cited by 2

A Novel 3D Convolutional Neural Network-Based Deep Learning Model for Spatiotemporal Feature Mapping for Video Analysis: Feasibility Study for Gastrointestinal Endoscopic Video Classification

Mrinal Kanti Dhar, Mou Deb, Poonguzhali Elangovan, Keerthy Gopalakrishnan, Divyanshi Sood, Avneet Kaur, Charmy Parikh, Swetha Rapolu, Gianeshwaree Alias Rachna Panjwani, Rabiah Aslam Ansari, Naghmeh Asadimanesh, Shiva Sankari Karuppiah, Scott A. Helgeson, Venkata S. Akshintala, Shivaram P. Arunachalam

Accurate analysis of medical videos remains a major challenge in deep learning (DL) due to the need for effective spatiotemporal feature mapping that captures both spatial detail and temporal dynamics. Despite advances in DL, most existing models in medical AI focus on static images, overlooking critical temporal cues present in video data. To bridge this gap, a novel DL-based framework is proposed for spatiotemporal feature extraction from medical video sequences. As a feasibility use case, this study focuses on gastrointestinal (GI) endoscopic video classification. A 3D convolutional neural network (CNN) is developed to classify upper and lower GI endoscopic videos using the hyperKvasir dataset, which contains 314 lower and 60 upper GI videos. To address data imbalance, 60 matched pairs of videos are randomly selected across 20 experimental runs. Videos are resized to 224 × 224, and the 3D CNN captures spatiotemporal information. A 3D version of the parallel spatial and channel squeeze-and-excitation (P-scSE) is implemented, and a new block called the residual with parallel attention (RPA) block is proposed by combining P-scSE3D with a residual block. To reduce computational complexity, a (2 + 1)D convolution is used in place of full 3D convolution. The model achieves an average accuracy of 0.933, precision of 0.932, recall of 0.944, F1-score of 0.935, and AUC of 0.933. It is also observed that the integration of P-scSE3D increased the F1-score by 7%. This preliminary work opens avenues for exploring various GI endoscopic video-based prospective studies.

View free PDFSource page

Related papers

crossrefJournal of Imaging2022-07-22Cited by 148

Brain Tumor Diagnosis Using Machine Learning, Convolutional Neural Networks, Capsule Neural Networks and Vision Transformers, Applied to MRI: A Survey

Andronicus A. Akinyelu, Fulvio Zaccagna, James T. Grist, Mauro Castelli, Leonardo Rundo

Management of brain tumors is based on clinical and radiological information with presumed grade dictating treatment. Hence, a non-invasive assessment of tumor grade is of paramount importance to choose the best treatment plan. Convolutional Neural Networks (CNNs) represent one o…

View free PDFSource page
crossrefJournal of Imaging2024-09-20Cited by 18

Convolutional Neural Network–Machine Learning Model: Hybrid Model for Meningioma Tumour and Healthy Brain Classification

Simona Moldovanu, Gigi Tăbăcaru, Marian Barbu

This paper presents a hybrid study of convolutional neural networks (CNNs), machine learning (ML), and transfer learning (TL) in the context of brain magnetic resonance imaging (MRI). The anatomy of the brain is very complex; inside the skull, a brain tumour can form in any part.…

View free PDFSource page
crossrefJournal of Imaging2026-07-08

Deep Learning-Based Multi-Class Pediatric Wrist Fracture Subtype Classification: A Pilot Study Comparing Convolutional Neural Network Architectures

Rohan A. Phadke, Samer G. Salman, Zane G. Salman, Sai M. Yedupati, Joshua Ong, Alireza Tavakkoli, et al.

Pediatric wrist fractures are among the most prevalent musculoskeletal injuries in children. Fracture subtype, including buckle/torus, greenstick, and Salter–Harris physeal injuries, directly influences management and prognosis. Subspecialty radiographic expertise required for su…

View free PDFSource page
crossrefJournal of Imaging2024-05-29Cited by 10

Hybridizing Deep Neural Networks and Machine Learning Models for Aerial Satellite Forest Image Segmentation

Clopas Kwenda, Mandlenkosi Gwetu, Jean Vincent Fonou-Dombeu

Forests play a pivotal role in mitigating climate change as well as contributing to the socio-economic activities of many countries. Therefore, it is of paramount importance to monitor forest cover. Traditional machine learning classifiers for segmenting images lack the ability t…

View free PDFSource page
crossrefJournal of Imaging2024-11-02Cited by 10

Convolutional Neural Network-Based Deep Learning Methods for Skeletal Growth Prediction in Dental Patients

Miran Hikmat Mohammed, Zana Qadir Omer, Barham Bahroz Aziz, Jwan Fateh Abdulkareem, Trefa Mohammed Ali Mahmood, Fadil Abdullah Kareem, et al.

This study aimed to predict the skeletal growth maturation using convolutional neural network-based deep learning methods using cervical vertebral maturation and the lower 2nd molar calcification level so that skeletal maturation can be detected from orthopantomography using mult…

View free PDFSource page
crossrefJournal of Imaging2024-12-12Cited by 3

Improved Generalizability in Medical Computer Vision: Hyperbolic Deep Learning in Multi-Modality Neuroimaging

Cyrus Ayubcha, Sulaiman Sajed, Chady Omara, Anna B. Veldman, Shashi B. Singh, Yashas Ullas Lokesha, et al.

Deep learning has shown significant value in automating radiological diagnostics but can be limited by a lack of generalizability to external datasets. Leveraging the geometric principles of non-Euclidean space, certain geometric deep learning approaches may offer an alternative…

View free PDFSource page