CORTEXA
← Browse
arxivcs.AI2026-07-09

The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs

Baha Rababah, Cuneyt Gurcan Akcora, Carson K. Leung

Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity. We show that these metrics fail to capture behavioral changes induced by quantization. We introduce correctness agreement, a decision-level metric that measures overlap in correct predictions between a base model and its quantized variants, independent of absolute accuracy. Across multiple models and quantization schemes from 8-bit to 2-bit, we find that behavioral divergence emerges under moderate quantization even when task performance appears preserved. To explain this effect, we analyze quantization as a structural operator on attention weights and quantify layer-wise distortions using statistical and distributional measures. Our results reveal non-linear breakpoints at low bit-widths and show that query and key projections are consistently more sensitive than value and output projections. These findings expose an illusion of equivalence between base and quantized models and motivate behavioral evaluation beyond conventional performance metrics.

View free PDFSource page

Related papers

arxivcs.CLcs.AI2026-07-23

Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders

Pablo Santiago Potes Velasco, María del Mar García Matabanchoy, Óscar Julián Pérez Ladino, Jhoan Stevan Mosquera Ortiz, Nicolás Lozano Mazuera, Gilber Alexis Corrales Gallego

Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated. This pilot study examines whether Qwen2.5-7B-Instruct internally represents Colombian identity, socioeconomic status, or stereotype-related inf…

View free PDFSource page
arxivcs.AI2026-07-23

Towards Faithful Graph Explanations with Synergistic Edge Effects via Granular Balls

Jiancu Chen, Shuyin Xia, Guan Wang, Degang Chen, Fan Chen

Instance-level explanations aim to reveal the rationale behind a model's decisions for a specific graph. Previous methods explain graph neural networks (GNNs) by selecting important edges to induce subgraphs, where edge importance is assessed by perturbing each edge and observing…

View free PDFSource page
arxivcs.HCcs.AI2026-07-31

The persuasive power of large language models does not depend on their perceived national origin

Ningzhi Liu, Yannic Hinrichs, Jonas R. Kunst

Conversational AI developed by geopolitical rivals reaches citizens worldwide, raising concerns that it could sway public opinion or be rejected as foreign propaganda, with consequences for democratic discourse and information sovereignty. Yet, whether an AI's perceived national…

View free PDFSource page
arxivcs.AI2026-07-23

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers

Vasudha Bhatnagar, Purnima Bindal, Vikas Kumar, Raj Kumari Bahl

Zero-shot summarization using Large Language Models (LLMs) has significantly advanced the abstractive summarization task by producing coherent and fluent summaries. However, underlying stochasticity of the large language models raises concerns about the stability and trustworthin…

View free PDFSource page
arxivcs.AI2026-07-23

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

Zhetong Zhang, Honghao Fu, Miao Xu, Yiwei Wang, Yujun Cai

As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuitively, we find that harmful videos paired with benign queries achieve higher attack success rates than the same videos paired wit…

View free PDFSource page