Revising research practices for singing data collection
Abstract As AI voice synthesis enables increasingly sophisticated vocal deepfakes and non-consensual voice cloning, the governance, licensing and access of singing datasets has become an urgent concern for data-contributors, who face significant harms from downstream and non-consensual usage of their singing data. Singing datasets are foundational to the development of high fidelity voice AI synthesis, yet current data collection practices pose challenges: data-contributors have an event-centric contribution to…
Singing datasets enable development of high-fidelity AI voice synthesis models.
Singers who contribute singing data face significant harms from downstream non-consensual use including vocal deepfakes and voice cloning.
Analysis is based on only three singing datasets, limiting generalizability of power-to-interest findings.
Evidence
- Peer-reviewedAI & SOCIETY2026-07-12
How should this claim be treated?
Truvace Impact Record TRV-2026-0125, v2: “Revising research practices for singing data collection.” Truvace, 2026-07-19. /record/TRV-2026-0125 (accessed at citation time). sha256 ebe1763c8aae53c4…
Calibration history
Every change to this record since certification, in the open.
Corrected automated sector classification after taxonomy audit
Certified into the record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0125 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace