Appearance
Audio
Audio AI splits into three sub-markets: TTS / voice agents (ElevenLabs leads), music generation (Suno, Udio, Lyria), and transcription APIs (Whisper, AssemblyAI, Deepgram). All are usage-based and Flowstate has no coverage today.
What's tracked
| Tool | Vendor | Pricing model | Coverage today | Notes |
|---|---|---|---|---|
| ElevenLabs | ElevenLabs | Per-seat + heavily usage-based (TTS, voice cloning, voice agents) | Invisible | Easy to overspend on voice agents. Track seats as contract SaaS spend; model production voice agents as Agents. |
| Suno | Suno | Per-seat subscription with credit caps | Invisible | Music generation. Track as contract SaaS spend. |
| Udio | Udio | Per-seat subscription with credit caps | Invisible | Music generation. |
| Lyria 3 | Per-token via Vertex | Invisible | Could surface via Gemini connector if bundled into Vertex billing. | |
| Whisper | OpenAI | Per-token via OpenAI API | Invisible (rolls into OpenAI API spend) | Spend lands on the OpenAI bill — see Foundation APIs. |
| AssemblyAI | AssemblyAI | Per-minute API | Invisible | Per-minute pricing means usage spikes. Track as contract SaaS spend. |
| Deepgram | Deepgram | Per-minute API | Invisible | Same as AssemblyAI. Track as contract SaaS spend. |
What Flowstate misses today
All of it. The economic risk in audio is voice agents — ElevenLabs, in particular, has burned holes in budgets when teams ship 24/7 voice surfaces. A voice agent running in production is an autonomous workload, not a copilot: model it as an Agent linked to an AI service account, forecast it aggressively from observed minutes-per-day, and revisit the number monthly.
Whisper spend is already on your OpenAI bill, so the right place to look is your foundation API line item rather than a separate Whisper entry.