CORTEXA
← Browse
arxivcs.SDcs.AI2026-07-08

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

Sirui Zhang, Tianle Wang, Xinyi Tong, Peiyang Yu, Jishang Chen, Liangke Zhao, Haoxin Zhang, Duo Xu, Xin Jin, Feng Yu, Songchun Zhu

Music aesthetic assessment is a challenging yet underexplored problem, requiring models to capture fine-grained, multi-dimensional human perceptual judgments. Progress in this area has been limited by the lack of large-scale datasets with structured aesthetic annotations. We introduce MADB, a large-scale dataset and benchmark comprising 9,999 tracks annotated by 30 trained annotators. Each track is rated by around 10 annotators across 10 perceptual dimensions and one overall score, with additional textual comments for multimodal analysis. We establish a unified evaluation framework over multiple pretrained models. Results reveal substantial gaps between model predictions and human judgments, exposing key limitations of current approaches. MADB provides a new benchmark for human-aligned music understanding. Project page: https://github.com/knownree/madb

View free PDFSource page

Related papers

arxivcs.SDcs.AI2026-07-07

From Textural Counterpoint to Feature Encoding: A Multi-Dimensional Machine Representation Study of Haydn's "The Lark" Integrating Electroacoustic Analysis

Yakun Liu, Zhiyu Jin, Hai Luan, Dong Liu, Xiaonan Li

Chamber music, as a highly precise multi-part interactive system, contains a logic of "role assignment and dynamic interaction" that provides an extremely valuable blueprint for exploring human-computer collaborative composition paradigms. Addressing the lack of role perception c…

View free PDFSource page
arxivcs.SDcs.AIcs.CRcs.LG2026-06-26

Room for Error: Large-Scale Simulation of Over-the-Air Acoustic Attacks

Andrew C. Cullen, Neil G. Marchant, Jiani Xie, Paul Montague, Sean Lamont, Maxwell Standen, et al.

While voice control is rapidly becoming a ubiquitous vector of human-AI communication, the risks facing these systems remain poorly understood. This is, in part, a product of the difficulties in scaling strictly digital adversarial workflows to the physical world. These scale bar…

View free PDFSource page
arxivcs.SDcs.AI2026-06-29

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

Shun Lei, Huaicheng Zhang, Dapeng Wu, Yaoxun Xu, Lishi Zuo, Wei Tan, et al.

Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics, and follow lyrics and prompts. Existing language model-based systems face a structural trade-off: mixed-token modeling preserves vocal-instrument coordination bu…

View free PDFSource page
arxivcs.SDcs.AI2026-07-06

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation

Linxi Li, Yuncong Yu, Qianwei Guo, Liwei Jin, Yechen Wang, Carsten Maple

While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects. Existing environmental audio datasets such as EnvSDD provide important initial resources, but remain limited in scale and generation provena…

View free PDFSource page
arxivcs.SDcs.AI2026-07-09

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

Wanyi Ning, Wei Zhou, Yingpeng Li, Yinshang Guo, Haitao Qian, Yiming Cheng

Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision are unavailable. We present PS4, a proxy-supervised training framework for TSE in real conversational…

View free PDFSource page