LLM EngineerseniorLondon Area, United KingdomhybridcontractSoftware Development and Staffing and RecruitingLLM evaluationRAGPrompt engineeringPythonA/B testingVector databasesAI safety testingResponsible AIposted
Senior AI Evaluation Engineer (Contract)
London, UK (Hybrid)
|
Long-Term Contract
|
Inside IR35
About the Role
We are seeking an experienced
Senior AI Evaluation Engineer
to join a major UK retail transformation programme, helping deliver next-generation Generative AI solutions that enhance customer experiences and business operations.
Working as part of a multidisciplinary AI engineering team, you will define and implement robust evaluation frameworks for enterprise-scale LLM applications, Retrieval-Augmented Generation (RAG) platforms, and agentic AI systems. You'll play a key role in ensuring AI solutions are accurate, reliable, safe, and production-ready through automated testing, benchmarking, and continuous quality improvement.
This is an excellent opportunity to work on cutting-edge AI technologies including microservices, event-driven architectures, cloud-native platforms, and large-scale conversational AI systems within a fast-paced retail environment.
Key Responsibilities
* Define and implement end-to-end evaluation strategies for Generative AI applications including LLM-powered conversational systems, RAG solutions, and multi-agent AI workflows.
* Design evaluation frameworks measuring response relevance, faithfulness, groundedness, hallucination rate, retrieval accuracy, latency, consistency, and customer experience quality.
* Create and maintain benchmark datasets representing real-world retail customer journeys, support interactions, operational queries, and business scenarios.
* Evaluate and benchmark GPT models, open-source LLMs, prompt strategies, embedding models, and retrieval configurations to optimise production performance.
* Develop automated evaluation pipelines using Python and modern evaluation frameworks, enabling continuous regression testing across prompts, models, and RAG pipelines.
* Conduct prompt engineering experiments, A/B testing, prompt versioning, and response optimisation to improve accuracy and reduce hallucinations.
* Perform human-in-the-loop evaluation alongside business stakeholders to validate AI behaviour and identify opportunities for continuous improvement.
* Design automated scoring methodologies using LLM-as-a-Judge, rule-based validation, and qualitative evaluation techniques.
* Validate RAG systems through retrieval quality assessment, knowledge grounding verification, citation validation, chunking optimisation, and semantic search evaluation.
* Deliver comprehensive safety, risk, and compliance testing including prompt injection protection, jailbreak testing, PII detection, bias assessment, toxicity evaluation, and Responsible AI governance.
* Build dashboards and reporting to monitor AI quality metrics, evaluation trends, production performance, and model drift.
* Collaborate with AI engineers, platform teams, product owners, and business stakeholders to translate evaluation findings into actionable improvements.
* Integrate evaluation frameworks into CI/CD pipelines supporting continuous deployment of production AI services.
* Produce technical documentation covering evaluation methodologies, testing standards, governance processes, and quality assurance frameworks.
Essential Skills \& Experience
* Strong commercial experience evaluating Large Language Models (LLMs) in production environments.
* Excellent understanding of Retrieval-Augmented Generation (RAG), vector search, embeddings, prompt engineering, and conversational AI.
* Experience designing enterprise AI evaluation frameworks and quality assessment methodologies.
* Strong Python development skills with experience building automated evaluation and experimentation pipelines.
* Experience with LangSmith, DeepEval, RAGAS, PromptTools, or similar LLM evaluation platforms.
* Experience benchmarking GPT models and open-source LLMs including Llama, Mistral, Gemini, Claude, or similar.
* Strong knowledge of LLM output evaluation, automated scoring, and NLP quality assessment techniques.
* Experience designing benchmark datasets and regression testing frameworks.
* Knowledge of Responsible AI principles, AI safety, governance, and compliance testing.
* Experience working with cloud platforms such as AWS or Azure.
* Familiarity with microservices, event-driven architectures, Docker, Kubernetes, and CI/CD pipelines.
* Strong analytical and problem-solving skills with the ability to translate model behaviour into measurable business improvements.
* Excellent communication skills and experience working with cross-functional engineering and business teams.
Technology Stack
Python • OpenAI • Azure OpenAI • LangChain • LangGraph • LangSmith • DeepEval • RAGAS • PromptTools • FastAPI • PostgreSQL • pgvector • Pinecone • FAISS • Docker • Kubernetes • AWS • Azure • GitHub Actions • MLflow • Pandas • NumPy • scikit-learn • CI/CD
What's on Offer
* Long-term contract opportunity
* Hybrid working in London
* Inside IR35 engagement
* Opportunity to work on enterprise-scale AI transformation within one of the UK's leading retail environments
* Exposure to the latest Generative AI, LLM evaluation, RAG, and cloud-native technologies
* Collaborative engineering culture focused on innovation, quality, and continuous improvement
More LLM Engineer roles like this, weekly
Roles like this expire in about a week. Get new LLM Engineer openings across the UK in your inbox, free, unsubscribe any time.