CORTEXA
← Browse
arxivcs.LGcs.CE2026-07-01

Interpretable vs Learned Encoders for High-Cardinality Fraud Detection

Xiao Han, Jingjing Liu, Moxuan Zheng, Zhen Zhang, Chenyu Wu

A total of seven categorical encoding methods were tested on the IEEE-CIS fraud benchmark dataset (590,540 records, 3.5% positives, 8 high-cardinality columns). The encoders were evaluated using a stratified 5-fold cross-validation (CV) with three repetitions. Five of the encoders had identical frozen LightGBM learners in the downstream phase, allowing for controlled comparisons of their performance to each other. CatBoost and TabNet were included as comparisons across paradigms using different learners. The entity embeddings produced the highest AUC-ROC (0.9612), with a statistically significant tie with that of CatBoost (0.9602) and statistically superior to tier group encoding (0.9548), whereas target encoding was only 0.0023 worse than tier group encoding and the auditor-friendly tier boundaries were maintained. Off-the-shelf TabNet did not outperform tree-based pipelines and collapsed under data scarcity. On AUC-PR, CatBoost leads (0.822 vs. 0.793); no encoder dominated both metrics. Per-column analysis confirmed the embedding advantage arises from joint multi-column representation.

View free PDFSource page

Related papers

arxivcs.CVcs.CEcs.LG2026-07-21

Benchmarking Deep Learning Approaches for AEC Engineering Drawing Layout Detection and Information Extraction

Tianyang Huang, Alessio Lombardi, Ahmed Elnagar, Ahmed Zalouk, George Paul, Sepehr Najjarpour, et al.

Information Extraction (IE) from Architecture, Engineering, and Construction (AEC) drawings remains hindered by manual inefficiency, while Layout Detection, a vital 'middleware' organizing graphical and textual hierarchies, is underexplored. General document layout models, optimi…

View free PDFSource page
arxivq-bio.BMcs.CEcs.LG2026-06-29

Structure-Regularized Interpretable TCR-Epitope Prediction

Jiarui Li, Zixiang Yin, Yunbei Zhang, Janet Wang, Samuel J. Landry, Zhengming Ding, et al.

T cell receptor (TCR)-epitope binding prediction is essential for understanding adaptive immunity and developing immunotherapies. Existing sequence- and structure-based models often generalize poorly to unseen epitopes and provide limited interpretability. Furthermore, the impact…

View free PDFSource page
arxivcs.LGcs.CEmath.NA2026-07-03

In-span learning: adapting reduced-order models using their own predictions

Amirpasha Hedayat, Laura Balzano, Karthik Duraisamy

Reduced-order models compress high-dimensional dynamics into low-dimensional representations that can be evaluated rapidly, but they lose accuracy when online dynamics drift beyond the training data. Adaptive methods address this by updating the subspace online with external, out…

View free PDFSource page
arxivcs.LGcs.CE2026-07-19

Apeliotes: A Diffusion-Based Modeling Framework for km-scale Multi-Level Atmospheric Fields

Evangelia Rafaela Frastali, Achyut Paudel, Maryam Golbazi, Frank Liu

High-resolution atmospheric data are required to resolve mesoscale and localized meteorological structures, however such datasets remain limited in many regions of the world. Existing high-resolution weather products are typically produced through dynamical downscaling, which is…

View free PDFSource page
arxivcs.LGcs.AIcs.CEcs.CL2026-07-06

URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment

Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Anton Morgunov, Arkadii Lin, Maksim Kuznetsov, et al.

Synthesis planning aiming to find pathways of reactions for a target molecule is one of the most important and challenging tasks in drug discovery. Recent progress has produced both specialized deep-learning retrosynthesis systems and general-purpose large language models, but ob…

View free PDFSource page
arxivcs.LGcs.AIcs.CEq-bio.QM2026-07-07

AbICL: In-Context Learning for Antigen-Specific Antibody Affinity Ranking

Zhiyuan Chen, Jing Hu, Junzhe Wang, Yueyang Huang, Xinyi Yang, Zhaoyang Wang, et al.

Accurate ranking of antibody candidates according to their binding affinity is essential for therapeutic antibody discovery. However, existing methods treat affinity comparisons independently and ignore the contextual information encoded in other labeled comparisons, limiting the…

View free PDFSource page