CORTEXA
← Browse
arxivcs.ITcs.LG2026-06-27

A Theoretical Interpretation of In-Context Learning via Probabilistic Modeling

Zhenyu Liu, Huaze Tang, Shao-Lun Huang

In-context learning (ICL) is an emerging paradigm that employs the semantic information inherent in large language models (LLMs) for generating answers to user queries. While the remarkable performance of ICL has been widely known, a general modeling and a rigorous theoretical analysis of this paradigm are still lacking. This work presents a probabilistic model for ICL and derives the performance of ICL for both general parametric distributions and exponential families. Based on the derived results, the work explains the impact of multiple factors such as the number of demonstrations, the sensitivity of the probabilistic model to the variation of its parameters, as well as the similarity between the demonstrations and the query on the performance of ICL.

View free PDFSource page

Related papers

arxivcs.LGcs.AIcs.CLcs.ITstat.ML2026-07-06

What Does a Discrete Diffusion Model Learn?

Rodrigo Casado Noguerales, Bernhard Schölkopf, Thomas Hofmann, Aran Raoufi

What does a discrete diffusion model learn: a denoiser, a score ratio, or a bridge plug-in predictor? At the level of jump rates, these are one object in different coordinates, and reading a neural network in the wrong coordinate changes the process being trained and sampled. Sta…

View free PDFSource page
arxivcs.ITcs.AIcs.LG2026-07-07

AirPASS: Over-the-Air Federated Learning via Pinching Antenna Systems

Seyed Mohammad Azimi-Abarghouyi, Christopher G. Brinton

This paper investigates over-the-air federated learning (AirFL) in wireless systems where the access point is equipped with a multi-waveguide pinching antenna system (PASS). We adopt the widely studied learning-oriented AirFL formulation, which seeks to maximize the number of sel…

View free PDFSource page
arxivcs.CRcs.ITcs.LG2026-07-14

Watermark Forensics for Generative Models: An Information-Theoretic Perspective

Xiaoyu Li, Zheng Gao, Xiaoyan Feng, Jiaojiao Jiang, Yulei Sui, Jiankun Hu

A watermark in a generative model's output is usually asked only whether a text is machine-made. The same mark can do more: attribute it to the user who produced it, extract a hidden payload, or localize the part that survives editing. These form a forensic ladder, and we ask wha…

View free PDFSource page
arxivcs.ITcs.DCcs.LGeess.SP2026-07-14

Mixed-Timescale Differential Coding for Downlink Model Broadcast in Wireless Federated Learning

Chung-Hsuan Hu, Zheng Chen, Erik G. Larsson

In standard federated learning systems, the parameter server broadcasts the global model to the participating devices in every iteration. Motivated by the temporal correlation between consecutive global models, differential coding can be applied to global model dissemination to r…

View free PDFSource page
arxivcs.LGcs.ITstat.ML2026-07-11

Conservation Laws for Diffusion Models

Ziv Aharoni, Henry D. Pfister

While autoregressive models optimize the exact data likelihood via the chain rule, diffusion models are typically trained with denoising objectives. We develop conservation laws based on generalized extrinsic information transfer (GEXIT) functions for a broad class of memoryless…

View free PDFSource page
arxivstat.MLcs.ITcs.LGmath.CAmath.CO2026-07-01

Function-Counting Theory for Low-Dimensional Data Structures

Konstantin Häberle, Helmut Bölcskei

The success of deep learning models in classification and regression is widely attributed to the low-dimensional structure that real-world data tend to exhibit, despite their high-dimensional representation. This work attempts to provide a mathematical framework for binary classi…

View free PDFSource page