Skip to main content

Director of Research, Text to Speech

Deepgram
Remote - USAUpdated 46d ago
Base salary
$213k–$328k
Published base salary range
Location
Remote - USA
Remote eligibility
Employment
Full-time
Director+
Role family
Engineering
AI / ML
Apply on jobs.ashbyhq.com
Job actionsApply now
Job actionsApply now

About the job

About Deepgram

Deepgram is a leading platform for the Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and production-grade voice agents. Over 200,000 developers and 1,300+ organizations build voice offerings powered by Deepgram, including Twilio, Cloudflare, Sierra, and others. Deepgram's voice-native foundation models are accessed via cloud APIs or self-hosted/on-premises software, with a focus on accuracy, low latency, and cost efficiency. Backed by a recent Series C, Deepgram has processed over 50,000 years of audio and transcribed more than 1 trillion words.

Role Overview

We are looking for a Director of Research to own our Text-to-Speech (TTS) program end to end, including research strategy, technical bets, and models that ship. This is a hands-on leadership role: you set direction and stay in the details, shaping architectures, experiments, training strategy, and evaluation. You will also build the team and operating model that let exceptional researchers move fast without lowering the bar.

Responsibilities

  • Own the TTS research and model roadmap, deciding which technical directions can materially improve speech-generation quality, including expensive and non-obvious ones, and recognizing when an approach should change or die.
  • Drive advances across neural audio modeling, prosody and expressiveness, controllability, multilingual speech, voice identity and consistency, data and training strategy, post-training, and inference performance, ensuring they land as measurable production gains.
  • Stay deeply technical: review research, challenge assumptions, design experiments, diagnose model failures, and take on the highest-leverage problems yourself.
  • Build evaluation and benchmarking that explains why models improve, not just whether they did, using automated metrics alongside human perceptual assessment.
  • Lead a mix of individual contributors and tech lead managers, hire and develop both, hold an exceptionally high technical bar, grow senior researchers into technical leaders, and set direction across sub-teams while pushing decisions down to the people closest to the work.
  • Partner with engineering and product leadership on ship-readiness, and represent Deepgram's TTS research internally and externally.

Qualifications

Must Have

  • Deep expertise in modern TTS, speech generation, or audio generative modeling, with a track record of personally training and improving large-scale neural models.
  • Command of the modern speech-generation stack and the open problems behind naturalness, expressiveness, controllability, robustness, voice consistency, and inference cost.
  • A history of setting research direction under genuine uncertainty: prioritizing experiments, allocating compute and researcher time, and killing approaches that aren't working.
  • Experience leading researchers and research engineers through other technical leaders, developing tech lead managers or equivalent, setting direction across sub-teams, while staying technically influential yourself.
  • AI as your default mode of work, not an occasional tool, with a specific, earned view of what it still can't do in speech research.
  • The ability to make complex technical tradeoffs legible to product, engineering, and executive audiences.

Nice to Have

  • TTS or generative-audio models deployed at meaningful production scale.
  • Built or substantially scaled a high-performing AI research organization.
  • Sophisticated evaluation systems for generative speech; expressive or multilingual generation, voice cloning and adaptation, or controllable generation.
  • Recognized external contributions (publications, open source, patents, invited talks) in speech synthesis, neural audio codecs, speech language models, or multimodal models.
  • Experience in fast-moving startup or research environments that routinely take models from idea to production.

Compensation & Benefits

Compensation ranges by location: San Francisco, NYC, Seattle: Estimated Base Salary $262.6K – $328.3K; Everywhere else in the U.S.: Estimated Base Salary $213K – $266.3K. Offers Equity and Bonus.

Application Instructions

Apply via the provided application link. Note: All legitimate Deepgram recruiting communication comes from an @deepgram.com email address. If you receive a message claiming to be Deepgram, forward it to careers@deepgram.com.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.