Your search has found 7 jobs

The next frontier in LLMs isn't just training models. It's teaching them to reason, use tools, follow intent and improve through post-training.

That's exactly what this team is focused on.

You'll be joining a Series A company building the next generation of conversational AI, working alongside an ex-NVIDIA & Meta research leader and one of the co-creators behind some well known open-source models.

Already experiencing rapid growth with strong commercial traction, their technology powers hundreds of millions of conversations every month, giving you the opportunity to work on research that directly improves models deployed at real-world scale.

They're looking for researchers who've gone beyond implementing methods from papers and have owned meaningful capability improvements themselves.

What you'll do

  • Design and run post-training experiments using techniques such as DPO, GRPO, SFT, rejection sampling and related approaches
  • Build and scale reinforcement learning and post-training infrastructure
  • Develop reward models, evaluation frameworks and assessment rubrics that improve model quality and behaviour
  • Work with vendors and data partners to build high-quality post-training datasets and feedback pipelines
  • Own end-to-end projects spanning data collection, modelling, evaluation and infrastructure, identifying where models break down and driving improvements in reasoning, controllability and alignment

What you'll bring

We're looking for someone who's owned significant post-training projects that have delivered meaningful capability improvements.

You'll understand the practical challenges behind modern post-training techniques, not just the theory. You've worked across data, modelling, evaluation and RL infrastructure, know how to improve model quality through robust evaluation and preference optimisation, and understand what it takes to make capability improvements repeatable.

Location: Fully remote worldwide.

Compensation: Depends on experience/location. In SF, comp is up to $500,000 base salary, plus stock.

Location: Remote, worldwide
Job type: Permanent
Emp type: Full-time
Salary type: Annual
Salary: negotiable
Job published: 06/08/2026
Job ID: 36031

Machine Learning Engineer, Inference

Want to solve realtime inference problems where milliseconds genuinely matter?

This role is with a fast-growing voice AI company building the realtime speech infrastructure layer behind hundreds of millions of production conversations every month. Their systems power enterprise voice experiences used at massive scale across customer support, ordering, and conversational automation.

This is not another generic AI platform role focused on wrapping APIs or building dashboards.

The work here sits deep in the runtime stack, optimising realtime speech systems under production latency constraints. Think streaming inference, scheduler design, GPU utilisation, concurrency optimisation, dynamic batching, and making state-of-the-art speech models actually behave correctly in realtime environments.

You’ll join a lean engineering team working directly on the inference systems behind low-latency conversational speech models. The challenge is not simply generating outputs, it’s generating speech naturally, reliably, and fast enough for real human interaction.

Your work will include:

  • Building and optimising realtime TTS streaming infrastructure
  • Improving scheduler and batching systems for production workloads
  • Reducing TTFA/TTFB while maintaining speech quality and stability
  • GPU profiling and identifying kernel-level bottlenecks
  • Optimising TensorRT, Triton, ONNX Runtime, and custom serving systems
  • Managing KV cache systems, speculative decoding, and streaming inference
  • Supporting heterogeneous deployment environments across NVIDIA and AMD GPUs
  • Collaborating closely with model researchers to productionise cutting-edge speech systems

A large part of the role involves solving difficult runtime problems where latency consistency, concurrency, and throughput directly impact user experience. The team already operates beyond the performance of most publicly available realtime speech systems, but there’s still substantial room to push the infrastructure further.

You’ll likely have strong depth across inference systems, runtime optimisation, distributed serving, or GPU performance engineering. Experience with tools like TensorRT, Triton, vLLM, CUDA Graphs, ONNX Runtime, or custom schedulers would be highly valuable.

The environment suits engineers who naturally investigate bottlenecks, enjoy working close to hardware constraints, and care deeply about performance engineering. If reducing latency by 30ms feels meaningful, you’ll probably enjoy this team.

The stack includes Rust, C++, Python, CUDA, TensorRT, Triton, Kubernetes, AWS, and custom realtime inference infrastructure.

Compensation is highly competitive and flexible depending on experience, including strong salary, equity, and benefits.

Location: Remote across the US or Europe.

If you’re excited by realtime AI systems problems where optimisation work directly shapes production performance at scale, this would be worth exploring.

All applicants will receive a response.

Location: San Francisco, CA
Job type: Permanent
Emp type: Full-time
Salary type: Annual
Salary: negotiable
Job published: 21/05/2026
Job ID: 35998

Want to work on one of the hardest unsolved problems in voice AI — making it actually sound like a human conversation?

Most voice AI falls apart the moment a conversation gets messy. Someone interrupts, emotions shift, the flow breaks — and the model can't keep up.

A small, ambitious SF startup is tackling exactly these problems, building speech models that handle natural conversation the way humans actually experience it. They have a working prototype and early commercial traction across several high-profile industry verticals.

The role

As a Senior Research Scientist, your focus is post-training — curating data, fine-tuning pre-trained speech models, and building the evaluation infrastructure that validates it all. You'll work on large-scale models with access to significant data resources.

What you'll do

  • Shape the data that goes into post-training — sourcing, cleaning and structuring it for large speech models

  • Supervised fine-tuning of pre-trained speech models

  • Build evaluation workflows — automated and human-in-the-loop

  • Drive measurable improvements in hallucination rates, instruction-following and generalisation

What you'll bring

  • PhD in ML or related field with a strong publications record

  • Hands-on experience training large speech models — ASR, TTS, or speech-to-speech

  • Solid post-training and SFT experience

The founding team includes a founding engineer from a billion-dollar AI company where they co-created one of the first generative models in the field, alongside the co-creator of the first generative voice at one of the world's largest tech companies.

Compensation is between $400k-$500k base with generous equity.

Based in San Francisco, onsite. Relocation support for those in US and willing to make the move.

Location: San Francisco
Job type: Permanent
Emp type: Full-time
Salary type: Annual
Salary: negotiable
Job published: 30/04/2026
Job ID: 34047

Training builds capability. Post-training decides what it becomes.

This team are rethinking how large multimodal models learn after pre-training — developing post-training and reinforcement learning methods that help models reason, plan, and interact in real time.

Founded by the researchers behind several of the most influential modern AI architectures, this lab are pushing alignment and learning efficiency beyond standard RLHF. They’re scaling preference-based training (RLHF, DPO, hybrid feedback loops) to new model types and creating systems that learn from interaction rather than static data.

You’ll work at the intersection of post-training, RL, and model architecture — designing reward models, scalable evaluation frameworks, and training strategies that make large-scale learning measurable and reliable. It’s applied research with direct impact, supported by serious compute and a tight researcher-to-GPU ratio.

You’ll bring experience in large-scale post-training or reinforcement learning (RLHF, DPO, or SFT pipelines), a solid grasp of LLM or multimodal training systems, and the curiosity to explore new optimisation and alignment methods. A publication record at top venues (NeurIPS, ICLR, ICML, CVPR, ACL) is a plus, but impact matters more than titles.

The team are based in San Francisco, working mostly in person. $1 million+ total compensation. Base salary circa $300K – $600K (negotiable) plus stock and bonus — exact package depends on experience.

If you want to work where post-training meets architecture — shaping how foundation models learn, reason, and adapt — this is that opportunity.

All applicants will receive a response.

Location: San Francisco
Job type: Permanent
Emp type: Full-time
Salary type: Annual
Salary: negotiable
Job published: 11/02/2026
Job ID: 34012

GPU Optimisation Engineer — Real-Time Inference

Want to push GPU performance to its limits — not in theory, but in production systems handling real-time speech and multimodal workloads?

This team is building low-latency AI systems where milliseconds actually matter. The target isn’t “faster than baseline.” It’s sub-50ms time-to-first-token at 100+ concurrent requests on a single H100 — while maintaining model quality.

They’re hiring a GPU Optimisation Engineer who understands GPUs at an architectural level. Someone who knows where performance is really lost: memory hierarchy, kernel launch overhead, occupancy limits, scheduling inefficiencies, KV cache behaviour, attention paths. The work sits close to the metal, inside inference execution — not general infra, not model research.

You’ll operate across the kernel and runtime layers, profiling large-scale speech and multimodal models end-to-end and removing bottlenecks wherever they appear.

What you’ll work on

  • Profiling GPU bottlenecks across memory bandwidth, kernel fusion, quantisation, and scheduling

  • Writing and tuning custom CUDA / Triton kernels for performance-critical paths

  • Improving attention, decoding, and KV cache efficiency in inference runtimes

  • Modifying and extending vLLM-style systems to better suit real-time workloads

  • Optimising models to fit GPU memory constraints without degrading output quality

  • Benchmarking across NVIDIA GPUs (with exposure to AMD and other accelerators over time)

  • Partnering directly with research to turn new model ideas into fast, production-ready inference

This is hands-on optimisation work across the stack. No layers of bureaucracy. No “platform ownership” theatre. Just deep performance engineering applied to models that are actively evolving.

What tends to work well

  • Strong experience with CUDA and/or Triton

  • Deep understanding of GPU execution (memory hierarchy, scheduling, occupancy, concurrency)

  • Experience optimising inference latency and throughput for large generative models

  • Familiarity with attention kernels, decoding paths, or LLM-style runtimes

  • Comfort profiling with low-level GPU tooling

The company is revenue-generating, its models are used by global enterprises, and the SF R&D team is expanding following a recent raise. This is growth hiring, not backfill.

Package & location

  • Base salary: up to ~$300,000 (negotiable based on depth)

  • Equity: Meaningful stock

  • Location: San Francisco preferred (relocation and visa sponsorship can be provided)

If you care about real-time constraints, GPU architecture, and squeezing every last millisecond out of large models, this is worth a conversation.

All applicants will receive a response.

Location: San Francisco, CA
Job type: Permanent
Emp type: Full-time
Salary type: Annual
Salary: negotiable
Job published: 11/02/2026
Job ID: 34843

Want to build speech AI that actually sounds human?

You'll be joining a well-funded speech AI startup with strong customer traction. They're building ultra-realistic voice technology that handles natural laughter, breathing, seamless language switching, and accurate pronunciation across languages and accents.

As their Senior Research Scientist, you'll work hands-on to expand their foundation models and push the boundaries of what's possible in speech AI: exploring multilingual capabilities, long-context generation, full-duplex modeling for natural conversations with interruptions, and novel architectures that balance speed with control.

What you'll do

  • Conduct research to advance their core speech models and extend product capabilities
  • Develop and experiment with new model architectures and training approaches
  • Work on large-scale model training and data systems
  • Collaborate with the team to take research from concept to deployed systems

What you'll bring

  • 3+ years of experience in speech synthesis, audio generation, or generative modeling
  • Experience with audio generation using LLMs
  • Solid background in modern language model architectures
  • Proven ability to ship research into production systems
  • Experience training large-scale models

Nice to have

  • Published research in speech or generative modeling
  • Experience with real-time speech systems or multimodal models

Ideally in SF, but can also consider remote worldwide. Comp is up to $250K base DOE, plus equity.

Location: San Francisco, CA
Job type: Permanent
Emp type: Full-time
Salary type: Annual
Salary: negotiable
Job published: 23/12/2025
Job ID: 34579

Build speech AI at trillion-parameter scale, driving human to AI conversation

This team has built their entire speech stack in-house, including proprietary LLM-based ASR and TTS, both already outperforming SOTA benchmarks. It's already powering real-time, human-to-AI conversations at scale, every day.

Now they're working at trillion-parameter scale to push toward a genuine end-to-end speech-to-speech LLM, one that understands and responds with genuine emotional intelligence and natural, human-like conversation. That means solving problems like long-context reasoning, pronunciation accuracy and maintaining consistency in noisy, real-world environments, in a domain where no model has cracked this yet.

They're hiring Senior and Staff-level Speech Scientists (Principal-level also considered) to help drive that work.

What you'll do:

  • Build SOTA speech models from the ground up, at genuinely large scale
  • Own problems end-to-end, from research through to production
  • Solve hard, domain-specific speech challenges as part of the push toward speech-to-speech LLM research
  • Help shape technical direction as part of a team working at the frontier of this space

What you'll bring:

  • Deep, hands-on expertise in at least one of: Speech/Audio LLMs, TTS or audio generation, or large-scale speech understanding
  • Experience shipping speech systems at real scale, not just research-stage prototypes

Nice to have:

  • Experience pre-training speech foundation models (HuBERT, Wav2Vec or similar)
  • Multimodal experience

Package: Up to $350K–$400K base (DOE), plus substantial equity. On-site, South Bay or Seattle.

Location: Bay Area
Job type: Permanent
Emp type: Full-time
Salary type: Annual
Salary: negotiable
Job published: 16/04/2025
Job ID: 33086