A repurposing Audio-Visual Speech Enhancement Using Multimodal Deep Convolutional Neural Networks, Hou Et Al. 2017, for musical audio source separation with an FFT-based approach of predicting a separated source's magnitude and phase FFT with video and audio input.
Trained on the MUSICES dataset, Zhou Et Al., 2020. Tools are available in this repo to download this dataset (available as YouTube videos.) The dataset was introduced along with this repo as an expansion of the MUSIC dataset, containing video and audio of musicians playing instruments from 11 classes.
