Revising research practices for singing data collection
Abstract As AI voice synthesis enables increasingly sophisticated vocal deepfakes and non-consensual voice cloning, the governance, licensing and access of singing datasets has become an urgent concern for data-contributors, who face significant harms from downstream and non-consensual usage of their singing data. Singing datasets are foundational to the development of high fidelity voice AI synthesis, yet current data collection practices pose challenges: data-contributors have an event-centric contribution to…

ace








