How to ensure universal sound source localization?
Hi,
I am very interested in your work. I noticed that OV-AVSS first relies on a universal sound source localization module to propose mask proposals, but how to ensure that it has universal capabilities?
I noticed that the training data of OV-AVSS is AVS, which is divided into 40 classes as seen and 30 classes as unseen. However, previous algorithms, such as AVSegFormer, are also trained on AVS, which do not claim to be universal. I also noticed that the qualitative results given in the paper can indeed segment some unseen classes.
Looking forward to receiving your reply, thank you in advance!
关闭于 2024-12-23 3 条评论