Data Annotation / Prompt Engineer
Inviting applications for the role of Data Annotation / Prompt Engineer. In this role, you will drive the quality, consistency, and governance of AI and Large Language Model (LLM) outputs while supporting annotation, prompt engineering, AI evaluation, and quality assurance activities.
Responsibilities
* Identify and manage project deliverables.
* Lead the design, optimization, and standardization of annotation processes and prompts across multiple AI use cases.
* Collaborate with the AI team to define and govern annotation schemas, taxonomies, and evaluation rubrics.
* Review complex or high-risk AI outputs and provide authoritative quality judgments.
* Ensure consistency, accuracy, and auditability of annotation and evaluation results.
* Define reusable prompt frameworks and best practices aligned with business and compliance objectives.
* Evaluate prompt performance at scale and guide improvements based on data-driven insights.
* Establish quality benchmarks, calibration processes, and escalation protocols.
* Identify bias, hallucinations, and systemic risks in AI outputs and recommend mitigation strategies.
* Support AI governance, compliance, and responsible AI initiatives.
* Participate in calibration sessions to align on annotation standards and evaluation criteria.
* Work closely with AI team Data Scientists, content reviewers, and product teams.
* Mentor and coach junior and intermediate annotators / prompt engineers.
* Develop and maintain AI evaluation frameworks and quality standards.
* Evaluate AI-generated outputs using Completeness, Groundedness, Precision, Guardrails, and Bias metrics.
* Create and maintain benchmark datasets, golden datasets, and ground-truth references.
* Support RAG evaluation, regression testing, and conversational AI assessments.
* Analyze evaluation results and provide actionable recommendations.
* Maintain evaluation documentation, auditability, and traceability.
Minimum Qualifications
* Fluency in written and spoken English.
* Mastery of grammar and content evaluation.
* Knowledge of software QA methodologies, tools, and processes.
* Hands-on experience in SQL writing, debugging, data extraction, and validation.
* Experience writing test plans and test cases.
* Strong teamwork and communication skills.
* Ability to adapt quickly between tasks.
* Ability to handle sensitive data with discretion.
* Understanding of Generative AI, LLMs, and prompt engineering concepts.
* Experience in data annotation, content review, AI evaluation, or quality assurance.
* Ability to assess AI-generated content using structured evaluation criteria.
Preferred Qualifications
* Experience with Generative AI and LLM-powered applications.
* Knowledge of AI evaluation methodologies and frameworks.
* Experience using Python, Jupyter Notebooks, or similar analysis tools.
* Familiarity with Confluence, JIRA, SharePoint, or GitHub.
* Experience creating annotation guidelines, evaluation playbooks, and quality standards.
* Experience with reporting and dashboard tools such as Excel or Power BI.
More AI roles like this, weekly
Roles like this expire in about a week. Get new AI openings across the UK in your inbox, free, unsubscribe any time.