Your search has found 2 jobs

Want to work on one of the hardest unsolved problems in voice AI — making it actually sound like a human conversation?

Most voice AI falls apart the moment a conversation gets messy. Someone interrupts, emotions shift, the flow breaks — and the model can't keep up.

A small, ambitious SF startup is tackling exactly these problems, building speech models that handle natural conversation the way humans actually experience it. They have a working prototype and early commercial traction across several high-profile industry verticals.

The role

As a Senior Research Scientist, your focus is post-training — curating data, fine-tuning pre-trained speech models, and building the evaluation infrastructure that validates it all. You'll work on large-scale models with access to significant data resources.

What you'll do

  • Shape the data that goes into post-training — sourcing, cleaning and structuring it for large speech models

  • Supervised fine-tuning of pre-trained speech models

  • Build evaluation workflows — automated and human-in-the-loop

  • Drive measurable improvements in hallucination rates, instruction-following and generalisation

What you'll bring

  • PhD in ML or related field with a strong publications record

  • Hands-on experience training large speech models — ASR, TTS, or speech-to-speech

  • Solid post-training and SFT experience

The founding team includes a founding engineer from a billion-dollar AI company where they co-created one of the first generative models in the field, alongside the co-creator of the first generative voice at one of the world's largest tech companies.

Compensation is between $400k-$500k base with generous equity.

Based in San Francisco, onsite. Relocation support for those in US and willing to make the move.

Location: San Francisco
Job type: Permanent
Emp type: Full-time
Salary type: Annual
Salary: negotiable
Job published: 30/04/2026
Job ID: 34047

Build speech AI at trillion-parameter scale, driving human to AI conversation

This team has built their entire speech stack in-house, including proprietary LLM-based ASR and TTS, both already outperforming SOTA benchmarks. It's already powering real-time, human-to-AI conversations at scale, every day.

Now they're working at trillion-parameter scale to push toward a genuine end-to-end speech-to-speech LLM, one that understands and responds with genuine emotional intelligence and natural, human-like conversation. That means solving problems like long-context reasoning, pronunciation accuracy and maintaining consistency in noisy, real-world environments, in a domain where no model has cracked this yet.

They're hiring Senior and Staff-level Speech Scientists (Principal-level also considered) to help drive that work.

What you'll do:

  • Build SOTA speech models from the ground up, at genuinely large scale
  • Own problems end-to-end, from research through to production
  • Solve hard, domain-specific speech challenges as part of the push toward speech-to-speech LLM research
  • Help shape technical direction as part of a team working at the frontier of this space

What you'll bring:

  • Deep, hands-on expertise in at least one of: Speech/Audio LLMs, TTS or audio generation, or large-scale speech understanding
  • Experience shipping speech systems at real scale, not just research-stage prototypes

Nice to have:

  • Experience pre-training speech foundation models (HuBERT, Wav2Vec or similar)
  • Multimodal experience

Package: Up to $350K–$400K base (DOE), plus substantial equity. On-site, South Bay or Seattle.

Location: Bay Area
Job type: Permanent
Emp type: Full-time
Salary type: Annual
Salary: negotiable
Job published: 16/04/2025
Job ID: 33086