Evals & Continuous Learning Engineer

Pydantic
Pydantic

London, UK

Posted on Jun 25, 2026
<section id="the-opportunity-section"><h2 id="the-opportunity" role="presentation"> <a href="#"></a><span role="heading" aria-level="2">The Opportunity</span> </h2> <p>Shipping reliable AI applications means closing the loop: capture what happened in production, turn it into evaluation data, measure quality, and feed improvements back into the system.</p> <p>We already have a real foundation in production — evals and experiments built into Logfire, our observability platform, plus our open source <a href="https://ai.pydantic.dev/evals/"><code>pydantic-evals</code></a> library. We're looking for someone who has worked on evaluation or LLM-observability platforms to own this end to end — and to push it toward genuine continuous learning, where systems measurably improve from their own production data.</p> <p>This might be you if you've worked on a product like Braintrust, Langfuse, LangSmith, Arize/Phoenix, or Humanloop — or built serious internal eval tooling.</p> </section>
Pydantic is an equal opportunity employer.