Job title: Senior Research Scientist LLM
Job type: Permanent
Emp type: Full-time
Industry: Generative AI
Functional Expertise: LLMs Reasoning/XAI Reinforcement Learning
Salary type: Annual
Salary: negotiable
Location: Remote, worldwide
Job published: 06/08/2026
Job ID: 36031

Job Description

The next frontier in LLMs isn't just training models. It's teaching them to reason, use tools, follow intent and improve through post-training.

That's exactly what this team is focused on.

You'll be joining a Series A company building the next generation of conversational AI, working alongside an ex-NVIDIA & Meta research leader and one of the co-creators behind some well known open-source models.

Already experiencing rapid growth with strong commercial traction, their technology powers hundreds of millions of conversations every month, giving you the opportunity to work on research that directly improves models deployed at real-world scale.

They're looking for researchers who've gone beyond implementing methods from papers and have owned meaningful capability improvements themselves.

What you'll do

  • Design and run post-training experiments using techniques such as DPO, GRPO, SFT, rejection sampling and related approaches
  • Build and scale reinforcement learning and post-training infrastructure
  • Develop reward models, evaluation frameworks and assessment rubrics that improve model quality and behaviour
  • Work with vendors and data partners to build high-quality post-training datasets and feedback pipelines
  • Own end-to-end projects spanning data collection, modelling, evaluation and infrastructure, identifying where models break down and driving improvements in reasoning, controllability and alignment

What you'll bring

We're looking for someone who's owned significant post-training projects that have delivered meaningful capability improvements.

You'll understand the practical challenges behind modern post-training techniques, not just the theory. You've worked across data, modelling, evaluation and RL infrastructure, know how to improve model quality through robust evaluation and preference optimisation, and understand what it takes to make capability improvements repeatable.

Location: Fully remote worldwide.

Compensation: Depends on experience/location. In SF, comp is up to $500,000 base salary, plus stock.