Jobverse
Machine Learning Engineer · France

ML Research Engineer

whitecircle·Paris

View & apply on company site →

TLDR: We are looking for several ML Engineers to train, post-train, and evaluate the LLMs at the core of our platform. This is hands-on modern model training work: large-scale data pipelines, SFT/RLHF/DPO-style alignment, reward models, distributed multi-GPU training, and evaluation.

About us

White Circle https://whitecircle.ai/ is an AI Safety company building the safety, reliability, and optimization layer for AI systems. At the core of our platform are policies – simple natural-language rules that define what an AI model should and shouldn’t do. We automatically test, enforce, and continuously improve these policies at scale.

  • We’ve raised $11M from top funds, founders, and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, Datadog, Sentry, and others
  • We process over 100M+ API calls every month
  • We fine-tune and train our own LLMs so they run faster and cheaper than any open or proprietary model

We’re a small, highly focused team. If you want to work deeply on hard problems, see your work ship to production quickly, and influence how AI safety is actually built – you’re the one we need.

What you’ll do

  • Turn petabytes of unstructured text into a structured, explorable view (topics, clusters, segments, trends, anomalies): iterate from “unknown unknowns” to stable definitions we can track.
  • Build scalable representation pipelines: sampling strategies, preprocessing/normalization, embeddings at scale, indexing, and retrieval to make the corpus searchable and analyzable.
  • Use LLMs pragmatically: labeling/classification, weak supervision, data enrichment, summarization, and automated diagnostics of inbound volumes (with cost/quality controls).
  • Deliver insights that change decisions: translate findings into product and operational actions (what data we have, what’s missing, where quality breaks, what to prioritize next).
  • Ship self-serve analytics: datasets, data models, and lightweight tools/dashboards so the team can explore and answer questions without ad-hoc requests.
  • Partner closely with engineering/research: align pipelines with production constraints (latency/cost/privacy), and integrate outputs into workflows.

You'll fit right in if you

  • Have strong Python + SQL with an engineering mindset: you can build reliable pipelines, not just notebooks.
  • Have solid applied NLP/ML experience on real-world text: embeddings, clustering, topic modeling, semantic search, classification; you understand failure modes and how to debug them.
  • Are comfortable at scale: distributed processing, large-scale storage-querying, and performance-cost tradeoffs.
  • Know how to evaluate fuzzy problems: offline/online metrics, human-in-the-loop labelling, inter-annotator agreement, drift monitoring, and reproducibility.
  • Have prior work with safety/moderation datasets, policy/rule systems, or high-volume logging/observability

A big plus

  • A public builder footprint: open-source models, datasets, or training frameworks on HuggingFace/GitHub, benchmarks, papers (workshop or main conference), or technical posts with real usage
  • Experience training models at a frontier or near-frontier lab, or leading open-source model releases with documented adoption
  • Experience with RL methods for LLMs beyond standard RLHF: online RL, GRPO-style methods, or novel alignment approaches
  • Experience with moderation, safety, or classification models at scale
  • Multilingual model training experience

Compensation & benefits

  • Competitive compensation, including equity
  • Flexible time off
  • Office in central London/Paris with flexible hybrid setup
  • Relocation support if you’re moving to Paris, available after your probationary period
  • Premium private health insurance
  • Mental health support, including coverage for therapy when you need it
  • Lunch and dinner covered when you work from the office
  • Learning and development support for courses, conferences, and opportunities to grow your skills
  • All the hardware, subscriptions, tools, and services you need
  • Team off-sites twice a year: we’ve recently been to the Alps, Saint-Tropez, and Marbella

Process

1. Intro call with Talent Team

2. Test assignment

3. Technical interview with Head of Applied Research

4. Final conversation with our CEO

View & apply on company site →

Sourced from a public career listing. Jobverse is an aggregator, not the employer.

← All Machine Learning Engineer jobs