Ask HN: Are there AI models for generating sounds based on a text and reference?
I've been searching for a commercialized AI model that takes a reference sound and a text instruction to generate a new sound, but I can't find one. While multimodal inputs to image or text outputs are solved, Audio + text -> Audio seems stuck where image generation was in 2016. I tried ElevenLabs SFX (text-only) and Stable Audio 3 (which supports reference but gave terrible results). Is this just not a common need, or am I missing something?