Your search has found 1 job

The next frontier in LLMs isn't just training models. It's teaching them to reason, use tools, follow intent and improve through post-training.

That's exactly what this team is focused on.

You'll be joining a Series A company building the next generation of conversational AI, working alongside an ex-NVIDIA & Meta research leader and one of the co-creators behind some well known open-source models.

Already experiencing rapid growth with strong commercial traction, their technology powers hundreds of millions of conversations every month, giving you the opportunity to work on research that directly improves models deployed at real-world scale.

They're looking for researchers who've gone beyond implementing methods from papers and have owned meaningful capability improvements themselves.

What you'll do

  • Design and run post-training experiments using techniques such as DPO, GRPO, SFT, rejection sampling and related approaches
  • Build and scale reinforcement learning and post-training infrastructure
  • Develop reward models, evaluation frameworks and assessment rubrics that improve model quality and behaviour
  • Work with vendors and data partners to build high-quality post-training datasets and feedback pipelines
  • Own end-to-end projects spanning data collection, modelling, evaluation and infrastructure, identifying where models break down and driving improvements in reasoning, controllability and alignment

What you'll bring

We're looking for someone who's owned significant post-training projects that have delivered meaningful capability improvements.

You'll understand the practical challenges behind modern post-training techniques, not just the theory. You've worked across data, modelling, evaluation and RL infrastructure, know how to improve model quality through robust evaluation and preference optimisation, and understand what it takes to make capability improvements repeatable.

Location: Fully remote worldwide.

Compensation: Depends on experience/location. In SF, comp is up to $500,000 base salary, plus stock.

Location: Remote, worldwide
Job type: Permanent
Emp type: Full-time
Salary type: Annual
Salary: negotiable
Job published: 06/08/2026
Job ID: 36031