Director of Research, Text to Speech
Deepgram- Base salary
- $213k–$328k Published base salary range · Top quartile for Engineering (461 listings)
- Location
- Remote - US Remote eligibility
- Employment
- Full-time Director+
About the job
Company Overview
Deepgram is the leading platform for the Voice AI economy, providing real-time APIs for speech-to-text, text-to-speech, and production-grade voice agents. Over 200,000 developers and 1,300+ organizations build voice offerings powered by Deepgram, including Twilio, Cloudflare, Sierra, and others. Deepgram's voice-native foundation models are accessed via cloud APIs or self-hosted, with unmatched accuracy, low latency, and cost efficiency.
Role
We are looking for a Director of Research to own our Text-to-Speech program end to end — research strategy, technical bets, and models that ship. This is a hands-on leadership role: you set direction and stay in the details, shaping architectures, experiments, training strategy, and evaluation. You'll also build the team and operating model that let exceptional researchers move fast without lowering the bar.
Responsibilities
- Own the TTS research and model roadmap, deciding which technical directions can materially move speech-generation quality.
- Drive advances across neural audio modeling, prosody, expressiveness, controllability, multilingual speech, voice identity, data and training strategy, post-training, and inference performance.
- Stay deeply technical: review research, challenge assumptions, design experiments, diagnose model failures, and take on the highest-leverage problems.
- Build evaluation and benchmarking that explains why models improve, using automated metrics and human perceptual assessment.
- Lead a mix of individual contributors and tech lead managers, hire and develop both, and set direction across sub-teams.
- Partner with engineering and product leadership on ship-readiness, and represent Deepgram's TTS research internally and externally.
Qualifications
Must have
- Deep expertise in modern TTS, speech generation, or audio generative modeling, with a track record of personally training and improving large-scale neural models.
- Command of the modern speech-generation stack and open problems behind naturalness, expressiveness, controllability, robustness, voice consistency, and inference cost.
- History of setting research direction under uncertainty, prioritizing experiments, allocating compute and researcher time, and killing approaches that aren't working.
- Experience leading researchers and research engineers through other technical leaders, while staying technically influential.
- AI as your default mode of work, with a specific, earned view of what it still can't do in speech research.
- Ability to make complex technical tradeoffs legible to product, engineering, and executive audiences.
Nice to have
- TTS or generative-audio models deployed at meaningful production scale.
- Built or substantially scaled a high-performing AI research organization.
- Sophisticated evaluation systems for generative speech; expressive or multilingual generation, voice cloning and adaptation, or controllable generation.
- Recognized external contributions in speech synthesis, neural audio codecs, speech language models, or multimodal models.
- Experience in fast-moving startup or research environments that routinely take models from idea to production.
Compensation & Benefits
Estimated base salary range: $213,000 – $328,300 USD, depending on location. Offers equity and bonus.
Location
Remote within the United States.
Skills & tags
Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.