Have a personal or library account? Click to login
Combination of Resnet and Spatial Pyramid Pooling for Musical Instrument Identification Cover

Combination of Resnet and Spatial Pyramid Pooling for Musical Instrument Identification

Open Access
|Apr 2022

Abstract

Identifying similar objects is one of the most challenging tasks in computer vision image recognition. The following musical instruments will be recognized in this study: French horn, harp, recorder, bassoon, cello, clarinet, erhu, guitar saxophone, trumpet, and violin. Numerous musical instruments are identical in size, form, and sound. Further, our works combine Resnet 50 with Spatial Pyramid Pooling (SPP) to identify musical instruments that are similar to one another. Next, the Resnet 50 and Resnet 50 SPP model evaluation performance includes the Floating-Point Operations (FLOPS), detection time, mAP, and IoU. Our work can increase the detection performance of musical instruments similar to one another. The method we propose, Resnet 50 SPP, shows the highest average accuracy of 84.64% compared to the results of previous studies.

DOI: https://doi.org/10.2478/cait-2022-0007 | Journal eISSN: 1314-4081 | Journal ISSN: 1311-9702
Language: English
Page range: 104 - 116
Submitted on: Nov 16, 2021
Accepted on: Feb 25, 2022
Published on: Apr 10, 2022
Published by: Bulgarian Academy of Sciences, Institute of Information and Communication Technologies
In partnership with: Paradigm Publishing Services
Publication frequency: 4 issues per year

© 2022 Christine Dewi, Rung-Ching Chen, published by Bulgarian Academy of Sciences, Institute of Information and Communication Technologies
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 License.