Artificial Intelligence Engineer

Queen Square Recruitment

LLM EngineerseniorLondon Area, United KingdomhybridcontractSoftware Development and Staffing and RecruitingLLM evaluationRAGPrompt engineeringPythonA/B testingVector databasesAI safety testingResponsible AIposted
Unlock apply linkApply links and the original listing are a Pro feature: £4.99/mo or £25 once.
Senior AI Evaluation Engineer (Contract) London, UK (Hybrid) | Long-Term Contract | Inside IR35 About the Role We are seeking an experienced Senior AI Evaluation Engineer to join a major UK retail transformation programme, helping deliver next-generation Generative AI solutions that enhance customer experiences and business operations. Working as part of a multidisciplinary AI engineering team, you will define and implement robust evaluation frameworks for enterprise-scale LLM applications, Retrieval-Augmented Generation (RAG) platforms, and agentic AI systems. You'll play a key role in ensuring AI solutions are accurate, reliable, safe, and production-ready through automated testing, benchmarking, and continuous quality improvement. This is an excellent opportunity to work on cutting-edge AI technologies including microservices, event-driven architectures, cloud-native platforms, and large-scale conversational AI systems within a fast-paced retail environment. Key Responsibilities * Define and implement end-to-end evaluation strategies for Generative AI applications including LLM-powered conversational systems, RAG solutions, and multi-agent AI workflows. * Design evaluation frameworks measuring response relevance, faithfulness, groundedness, hallucination rate, retrieval accuracy, latency, consistency, and customer experience quality. * Create and maintain benchmark datasets representing real-world retail customer journeys, support interactions, operational queries, and business scenarios. * Evaluate and benchmark GPT models, open-source LLMs, prompt strategies, embedding models, and retrieval configurations to optimise production performance. * Develop automated evaluation pipelines using Python and modern evaluation frameworks, enabling continuous regression testing across prompts, models, and RAG pipelines. * Conduct prompt engineering experiments, A/B testing, prompt versioning, and response optimisation to improve accuracy and reduce hallucinations. * Perform human-in-the-loop evaluation alongside business stakeholders to validate AI behaviour and identify opportunities for continuous improvement. * Design automated scoring methodologies using LLM-as-a-Judge, rule-based validation, and qualitative evaluation techniques. * Validate RAG systems through retrieval quality assessment, knowledge grounding verification, citation validation, chunking optimisation, and semantic search evaluation. * Deliver comprehensive safety, risk, and compliance testing including prompt injection protection, jailbreak testing, PII detection, bias assessment, toxicity evaluation, and Responsible AI governance. * Build dashboards and reporting to monitor AI quality metrics, evaluation trends, production performance, and model drift. * Collaborate with AI engineers, platform teams, product owners, and business stakeholders to translate evaluation findings into actionable improvements. * Integrate evaluation frameworks into CI/CD pipelines supporting continuous deployment of production AI services. * Produce technical documentation covering evaluation methodologies, testing standards, governance processes, and quality assurance frameworks. Essential Skills \& Experience * Strong commercial experience evaluating Large Language Models (LLMs) in production environments. * Excellent understanding of Retrieval-Augmented Generation (RAG), vector search, embeddings, prompt engineering, and conversational AI. * Experience designing enterprise AI evaluation frameworks and quality assessment methodologies. * Strong Python development skills with experience building automated evaluation and experimentation pipelines. * Experience with LangSmith, DeepEval, RAGAS, PromptTools, or similar LLM evaluation platforms. * Experience benchmarking GPT models and open-source LLMs including Llama, Mistral, Gemini, Claude, or similar. * Strong knowledge of LLM output evaluation, automated scoring, and NLP quality assessment techniques. * Experience designing benchmark datasets and regression testing frameworks. * Knowledge of Responsible AI principles, AI safety, governance, and compliance testing. * Experience working with cloud platforms such as AWS or Azure. * Familiarity with microservices, event-driven architectures, Docker, Kubernetes, and CI/CD pipelines. * Strong analytical and problem-solving skills with the ability to translate model behaviour into measurable business improvements. * Excellent communication skills and experience working with cross-functional engineering and business teams. Technology Stack Python • OpenAI • Azure OpenAI • LangChain • LangGraph • LangSmith • DeepEval • RAGAS • PromptTools • FastAPI • PostgreSQL • pgvector • Pinecone • FAISS • Docker • Kubernetes • AWS • Azure • GitHub Actions • MLflow • Pandas • NumPy • scikit-learn • CI/CD What's on Offer * Long-term contract opportunity * Hybrid working in London * Inside IR35 engagement * Opportunity to work on enterprise-scale AI transformation within one of the UK's leading retail environments * Exposure to the latest Generative AI, LLM evaluation, RAG, and cloud-native technologies * Collaborative engineering culture focused on innovation, quality, and continuous improvement