arxiveess.SPcs.SD2026-07-18
Efficient Audio-Visual Event Recognition via Knowledge Distillation and Dynamic INT8 Quantization of a Hybrid Cross-Attention Network
Parinaz Binandeh Dehaghani, Danilo Pena, A. Pedro Aguiar
Audio-visual event recognition (AVER) has achieved significant performance improvements through transformer-based multimodal architectures. However, the high computational complexity, large memory footprint, and inference cost of these models hinder their deployment on edge and r…