The architecture bet: 100+ small models instead of one big one
Modulate's technical pitch rests on a specific architectural choice. Instead of relying on one large foundation model, Velma's Ensemble Listening Model (ELM) combines more than 100 specialized audio models to analyze emotion, tone, intent, and synthetic speech directly from raw audio rather than from a transcript [1]. The company claims this ensemble approach is up to 1,000x more efficient, twice as accurate on true positives, and produces 7x fewer false positives than single large-model approaches - though it's worth noting these are Modulate's own benchmark claims rather than figures from an independent study cited in coverage [2]. The new funding is earmarked specifically to scale that architecture: AI/ML research, product and engineering headcount, developer relations, and new APIs and deployment options so outside developers can integrate Modulate's audio models directly rather than building their own [1].



