AI Quality Engineer

SoCode Recruitment

AI EngineermidLondon Area, United KingdomhybridfulltimeTechnology, Information and InternetPythonPlaywrightpytestCI/CDLLM evaluationRubric designGolden datasetsposted 10 Sep
Unlock apply linkApply links and the original listing are a Pro feature: £4.99/mo or £25 once.

Job Title: AI Quality Engineer (Senior options available too)

Location: Hybrid - 2 days per week in Central London

Salary: Flexible Depending on Experience

Are you an AI Quality Engineer with solid experience of running layered evaluation frameworks, looking to join a business at the forefront of AI-Powered investigative technology?

As the platform continues to scale, the volume of customers, product releases and regulatory challenges is increasing. With AI-generated output forming an important part of the product, maintaining confidence in quality is a complex engineering challenge.

You’ll have the opportunity to help shape how AI-driven functionality is evaluated, tested and improved, working closely with engineering and product teams to build confidence in releases.

You’ll work across the evaluation systems and test infrastructure that enable the team to confidently release non-deterministic AI functionality.

This is a hands-on role covering test automation, AI evaluation, CI infrastructure and investigation of complex test and quality issues.

Key Responsibilities

  • Enhance and run layered evaluation frameworks, including automated checks, LLM-as-judge scoring, defined rubrics and targeted human review.
  • Maintain versioned golden datasets to ensure evaluations remain reproducible and auditable.
  • Monitor key quality signals across investigative outputs, including groundedness, hallucination rate, entity resolution and source quality.
  • Run and improve a tiered CI testing model using Python, Playwright and pytest.
  • Maintain test coverage across six authenticated personas, three browsers and multi-tenant environments.
  • Investigate flaky tests, inconsistent results and changes in evaluation metrics, identifying and addressing root causes.
  • Work collaboratively with engineering and product teams to improve quality and confidence in releases.

The team uses Claude Code across the testing workflow, including selector synchronisation, failure triage and evaluation scoring. You'll need to apply your own judgement when interpreting results rather than relying solely on whether a test has passed or failed.

There are two potential routes into the role. You may already be a senior engineer with experience across much of the above, or you may have strong test automation foundations and the ability and motivation to develop your experience in AI evaluation.

Essential Experience

  • Strong test automation engineering experience.
  • Experience with Python or a comparable programming language.
  • Experience with a modern automation framework such as Playwright, pytest or similar.
  • Strong understanding of CI/CD.
  • An analytical and questioning approach to testing and quality.
  • The ability to investigate test failures and inconsistent metrics and identify underlying causes.
  • Strong communication skills, with the ability to explain complex technical issues to different audiences.
  • A collaborative approach to working with engineers and product teams.
  • Evidence that you can learn and become competent in technically challenging areas quickly.
  • Genuine interest in the challenges involved in testing non-deterministic AI output.

At a Senior Level

It would be beneficial to also have experience of

  • LLM evaluation.
  • Rubric design.
  • LLM-as-judge approaches.
  • AI regression testing.
  • Evaluation harnesses.
  • Improving quality practices across teams.

Benefits include

  • Employee shares and equity programme, giving you the opportunity to own a meaningful part of the business.
  • Private health insurance \& Life insurance.
  • Unlimited holidays.
  • Annual professional development fund to support learning and development.