£•••kLLM EngineermidLondon Area, United KingdomhybridcontractIT Services and IT Consulting and Information ServicesPythonLarge Language Models (LLMs)RAG architecturesPrompt EngineeringDeepEvalPromptToolsRagasPromptfooposted
AI Evaluation Engineer
Location:
London, UK
Work Model:
Hybrid (1- 2 Days per Week in Office)
Contract Duration:
6 Months
Rate:
£600
-
£700 Per Day
(Inside IR35)
Key Responsibilities
* Define and implement end-to-end evaluation strategies for
Generative AI conversational systems
, including LLMs, RAG pipelines, and AI agents.
* Establish and maintain
evaluation metrics, benchmarks, and quality standards
for AI-powered applications.
* Create and manage
test datasets, benchmark suites, and evaluation scenarios
.
* Benchmark different
GPT variants, open-source LLMs, and prompt engineering strategies
.
* Develop and maintain
automated evaluation pipelines
to support continuous testing and quality monitoring.
* Conduct
human-in-the-loop evaluations
to validate qualitative response quality and business relevance.
* Drive
prompt optimization and response quality improvements
to improve accuracy and consistency.
* Validate
RAG systems and knowledge grounding
, ensuring responses are aligned with retrieved information.
* Perform
safety, risk, compliance, and Responsible AI testing
across AI solutions.
* Analyze model performance and translate findings into actionable recommendations for improvement.
Required Skills
* Strong understanding of
Large Language Models (LLMs)
,
RAG architectures
, and
Prompt Engineering
* Hands-on experience with
Python
for data analysis, experimentation, and evaluation automation
* Experience designing and implementing
AI evaluation frameworks
* Strong knowledge of
LLM evaluation metrics and NLP quality assessment
* Experience with
benchmarking, scoring, and validating AI model outputs
* Familiarity with evaluation tools such as
DeepEval, PromptTools, Ragas, Promptfoo, or similar frameworks
* Experience building automated testing and evaluation workflows
* Strong analytical and problem-solving skills
Experience Required
* End-to-End Evaluation Framework Design
* Evaluation Metrics Definition \& Quality Assessment
* Test Dataset Creation \& Benchmarking
* LLM Performance Evaluation and Scoring
* Prompt Testing, Optimization \& A/B Evaluation
* RAG Evaluation \& Knowledge Grounding Validation
* Hallucination Detection \& Response Validation
* Safety, Risk \& Compliance Testing
* Python-Based Experimentation Frameworks
* Automation \& Evaluation Tooling
Nice to Have
* Experience with
Responsible AI principles and governance frameworks
* Experience evaluating
Generative AI Agents and Autonomous Workflows
* Knowledge of
customer support and retail conversational AI use cases
* Familiarity with
LangSmith, MLflow, TruLens, Azure AI Evaluation SDK
, or similar platforms
* Experience working within enterprise-scale AI transformation programs
* Understanding of observability and monitoring for AI applications
More LLM Engineer roles like this, weekly
Roles like this expire in about a week. Get new LLM Engineer openings across the UK in your inbox, free, unsubscribe any time.