CORTEXA
← Browse
arxivstat.MEcs.LG2026-07-12

Observation-Level Watermarking and Detection for Tabular Data

Dongyu Cui, Xuan Bi

With the development of generative AI, watermarking techniques have been widely used to detect the authenticity of AI-generated data and protect the rights of users and creators. While it is already well applied in data types including imaging and text data, watermarking tabular data is still under-explored. Existing methods primarily focus on numerical data, leaving discrete, categorical, and mixed data less studied. In this work, we propose STAMP (Single-observation Tabular Attribution and Marking Procedure), a novel framework for watermarking tabular data that can accommodate and preserve a wide range of distributions. We also develop a corresponding detection mechanism, which can reliably identify watermarks even when the sample size is as small as one. We establish theoretical guarantees for asymptotic consistency and detection accuracy. Finally, through extensive simulation studies and two real-data applications, we demonstrate that the proposed method is effective and robust to subsetting, while maintaining data fidelity and a high detection rate.

View free PDFSource page

Related papers

arxivstat.MEcs.LGstat.AP2026-06-30

Hierarchical Clustering As a Novel Solution to the Notorious Multicollinearity Problem in Observational Causal Inference

Yufei Wu, Zhiying Gu, Alex Deng, Jacob Zhu, Linsha Chen

Multicollinearity is a long lasting challenge in observational causal inference, especially in regressions -- highly correlated independent variables make it hard to isolate their individual impacts on outcomes of interest. While common solutions such as shrinkage estimators and…

View free PDFSource page
arxivstat.MLcs.LGstat.ME2026-07-22

Data-Poisoning Audits for Causal Effect Estimation

Kwangho Kim

Observational causal analyses increasingly pool records across sites, vendors, and collection systems, creating vulnerability to append-only attacks in which plausible records are strategically selected to alter a reported treatment effect. We develop a data-poisoning audit for a…

View free PDFSource page
arxivstat.MEcs.LG2026-07-03

CaSPECT: Discovering Causally Homogeneous Subgroups via Directed Spectral Clustering

Arghya Pratihar, Shinjon Chakraborty, Swagatam Das

We propose \textbf{CaSPECT}, a causal spectral clustering framework for discovering causally homogeneous subgroups from observational data. Rather than clustering in covariate space, CaSPECT defines similarity through the topology of a learned directed acyclic graph (DAG); a boot…

View free PDFSource page