Member of Technical Staff, Research
ABOUT ABUNDANT
Abundant is an applied research lab focused on scaling reinforcement learning for safe and reliable agentic capabilities. We are an extremely talent-dense team of researchers, roboticists, founders, and operators whose work includes two-tower retrieval https://dl.acm.org/doi/10.1145/2959100.2959190, BERT https://arxiv.org/abs/1810.04805, web-scale graph neural networks https://arxiv.org/abs/1806.01973, and the Waymo Driver https://waymo.com/blog/2020/10/waymo-is-opening-its-fully-driverless.
THE ROLE
You will own a research agenda where research meets production, co-designing strategies with researchers from frontier AI labs. You are the “PM of the model”: you pick the problem, design the data and reward signal that moves it, and see it through to a live training run. The job rewards a founder mentality and extreme ownership, including going into unfamiliar domains, from chemical engineering to complex legal logic, and getting to ground truth fast on problems others have written off.
WHAT YOU’LL DO
- Foundational capability research and execution. Architect and execute a research agenda to find simple, generalizable ideas that advance model reasoning at scale, and own it from hypothesis through to running in live systems.
- Design tasks and verifiers that stay hard. Build long-horizon agentic tasks that are still difficult a model generation from now, with graders where partial credit means something and the obvious exploit does not score.
- Diagnose failure modes. Read trajectories until you can say why a model fails a class of work, then turn that into the data or reward signal that fixes it. This is the core loop of the job.
- Model alignment and data strategy. Partner with the world’s most advanced AI research teams on datasets and benchmarks that shape how frontier models behave, focused on alignment, safety, and reward signal design.
- Autonomous problem selection. Identify, scope, and run long projects, choosing the problems that matter most for scaling data toward AGI.
- System infrastructure. Work with engineering on data pipelines, internal tooling, and deep learning implementations. If a workflow produces bad data, you fix the workflow.
WHO YOU ARE
- You write a lot of code. Strong software engineering and deep Python, comfortable owning a system end to end and debugging through layers you did not write.
- You have built evaluations, RL environments, agent harnesses, verifiers, or execution sandboxes, and you have opinions about which of them can be trusted.
- You have shipped research into production, particularly post-training, distillation, and high-stakes evaluation.
- You have worked on large-scale agent systems: orchestration frameworks, tool APIs, distributed execution, observability, and logging.
- You get to ground truth in an unfamiliar domain by building something and reading outputs, not by reading about it.
- You notice the misspecified task, the ambiguous prompt, and the grader that can be gamed.
- You have a thoughtful perspective on the societal and safety impacts of deploying general-purpose AI systems.
NICE TO HAVE
- A PhD (CS, ML, NLP) and publications at NeurIPS, ICML, or ICLR. Neither is required.
- Work on public benchmarks or task suites, for example Terminal-Bench, SWE-bench-style harnesses, or HCAST.
- RLHF or RLAIF experience and a view on how training choices show up in model behavior.
- Experience running work through large expert or annotation operations.
COMPENSATION
Base Salary
$250,000 - $450,000
Cash Bonus
Sizable performance bonus tied to project and company milestones
Equity
Generous early-stage grant
Benefits
Health, dental, vision + flexible PTO
Sourced from a public career listing. Jobverse is an aggregator, not the employer.