Senior AI Evaluation Engineer
TalentCo are delighted to be partnering with an established SaaS business at the forefront of secure customer engagement and Conversational AI, looking to hire a high-calibre Senior AI Evaluation Engineer into their growing AI organisation.
The core role of the Senior AI Evaluation Engineer will be to define, build and operate the evaluation systems used to measure the quality, reliability and safety of production RAG, conversational and Agentic AI applications.
You'll work across retrieval, generation, multi-turn conversations, agent reasoning and tool use, combining exploratory investigation with robust Python automation. This isn't a conventional QA or test automation role; you'll define what good AI performance looks like, translate it into measurable criteria, and build the datasets, tooling and release controls needed to evaluate non-deterministic AI systems.
This hire can be based remote-first in the UK, with occasional travel to their HQ (just north of London) as required.
As Senior AI Evaluation Engineer you will:
* Define evaluation strategies, metrics, acceptance thresholds and release gates across conversational, RAG and Agentic AI applications.
* Build and maintain automated evaluation frameworks in Python using tools such as Ragas and DeepEval.
* Create representative evaluation datasets, golden test cases, adversarial cases and regression suites across complex AI workflows.
* Evaluate RAG and agent behaviour across retrieval quality, grounding, hallucination, tool use, reasoning, memory, safety and multi-turn conversations.
* Integrate evaluation into CI/CD workflows and work closely with AI, Software Engineering, Product and Security teams to identify failures and improve system performance.
We want to hear from you if:
* You have a strong Software Engineering, AI Evaluation, Quality Engineering or Test Automation background, with significant experience evaluating production RAG and Agentic AI systems.
* You're highly proficient in Python and comfortable building maintainable evaluation libraries, test harnesses and automated evaluation pipelines.
* You have hands-on experience with Ragas, DeepEval or comparable LLM evaluation frameworks.
* You understand modern LLM applications including embeddings, vector retrieval, reranking, prompting, tool calling and agent orchestration.
* You have experience with AI security, red-teaming or responsible AI, ideally within regulated environments.
What do they offer?
* Base salary of £75k–£95k, dependent on experience
* Comprehensive benefits package including bonus and Share Incentive Scheme
* Remote-first working with flexible arrangements
* 25 days holiday, private healthcare options and pension scheme
* Opportunity to take ownership of AI evaluation across a growing Conversational AI platform
More AI Engineer roles like this, weekly
Roles like this expire in about a week. Get new AI Engineer openings across the UK in your inbox, free, unsubscribe any time.