WAVENET-BASED QUEEN BEE PRESENCE DETECTION FROM BEEHIVE ACOUSTIC SIGNALS: A COMPARATIVE EVALUATION WITH CNN AND MULTIMODAL FUSION
Keywords:
WaveNet, Queen Bee Detection, Beehive Acoustics, Raw Audio, Deep Learning, Convolutional Neural Network, Acoustic Monitoring, Precision ApicultureAbstract
Continuous monitoring of honeybee colonies is important for non-invasive assessment of colony condition and early detection of queen loss. Acoustic monitoring provides a practical alternative to frequent manual hive inspection; however, reliable classification of queen presence requires models capable of capturing temporal and spectral characteristics of hive sounds. This study presents a comparative evaluation of raw-waveform WaveNet, a spectrogram-based convolutional neural network (CNN), an RBF-SVM baseline using MFCC features, and a decision-level WaveNet-CNN fusion model for queen bee presence detection from beehive acoustic recordings. The dataset comprises 6,000 recordings, including 4,000 queen-present and 2,000 queen-absent recordings. All recordings were resampled to 16 kHz and standardized to 5 seconds. The WaveNet architecture uses exponentially increasing dilated causal convolutions to model temporal dependencies directly from raw audio waveforms, while the CNN operates on 64 × 64 log-Mel spectrograms. Decision-level fusion was evaluated using a validation-set weight-selection procedure to determine whether complementary spectral information improved the raw-waveform model. On the current held-out test split, WaveNet achieved an accuracy of 99.33%, compared with 97.00% for the RBF-SVM, 85.45% for the spectrogram CNN, and 99.00% for the WaveNet-CNN fusion model. Thus, fusion did not improve upon the standalone WaveNet model under the evaluated conditions. These results indicate that raw-waveform temporal modeling is a promising approach for acoustic queen-presence detection, while also highlighting the need for source-independent evaluation, repeated statistical validation, and comparison with recent deep-learning audio models. The findings should therefore be interpreted as evidence of strong performance on the evaluated dataset rather than as definitive evidence of cross-hive generalization.












