Platform Engineer (AI evaluation platform)

Intellias

fulltimeIT Services and IT Consultingposted 30 Jul
Unlock apply linkApply links and the original listing are a Pro feature: £4.99/mo or £25 once.
We are looking for a Platform Engineer — CI/CD Gate \& Online Evaluation to help establish automated quality controls for enterprise AI platforms and agent-based systems. In this role, you will integrate evaluation workflows into deployment pipelines, configure online quality monitoring, and build operational visibility for AI performance and reliability. You will work closely with platform, DevOps, AI engineering, and governance teams to ensure that AI agents meet quality standards before and after production deployment. Project overview: Our customer is a multinational corporation with more than a century of history and offices in over 180 countries. Their most ambitious goal at the time is to introduce a range of Reduced-Risk Products (RRPs). The target audience is more than 1 billion consumers around the globe. IT platform hosts 700+ applications. Intellia's mission is to help the client with the engineering of a comprehensive software ecosystem for a game-changing IoT product on the margin of innovative consumer experience and cutting-edge technology. Our teams are involved in the engineering of core platform components for best-in-class eCommerce, Digital Marketing and IoT solutions. As an Engineer, you will become a part of Core Architecture Team and be responsible for the architecture, implementation of best practices in our Digital Engineering Enterprise Platform. The Platform is a set of services and internet applications that accelerate the development and delivery of software applications by taking care of common SDLC challenges. The Platform provides access and consumption for engineering teams to a set of services, technologies, practices for their development and for operating their application, ensuring a set of compliance and best practices. Requirements: Skills: * AWS AgentCore Evaluation API (on-demand mode — CreateEvaluation in CI/CD pipelines) * GitHub Actions integration for evaluation deployment gate * Online evaluation configuration (AgentCore online mode, configurable sampling rate) * CloudWatch dashboards for per-agent per-evaluator evaluation scores over time * CloudWatch alarm configuration for quality degradation thresholds Experience: * 4+ years platform or DevOps engineering * CI/CD pipeline integration for quality gates * CloudWatch dashboard and alarm design Nice-to-have * AWS AgentCore Evaluation online mode sampling configuration * AWS AgentCore Runtime version comparison using evaluation (regression detection) * CloudWatch EMF for structured evaluation score metrics Responsibilities: * Design and implement CI/CD quality gates using AWS AgentCore Evaluation to validate AI agents and workflows before deployment. * Integrate evaluation execution into GitHub Actions and deployment pipelines to support automated release decisions. * Configure and maintain on-demand evaluation workflows for pre-release quality validation and regression detection. * Implement and manage online evaluation capabilities, including production sampling strategies and evaluation scheduling. * Build CloudWatch dashboards that provide visibility into evaluation scores, quality trends, and agent performance over time. * Configure CloudWatch alarms and notification mechanisms to detect quality degradation, performance regressions, and reliability issues. * Collaborate with AI engineers to define evaluation thresholds, release criteria, and quality acceptance standards. * Support platform-wide observability by integrating evaluation results into monitoring and operational reporting workflows. * Analyze evaluation outcomes and identify quality regressions across agent versions, prompts, tools, and workflows. * Develop and maintain automation that enables continuous validation of AI systems throughout the software delivery lifecycle. * Partner with platform and DevOps teams to improve deployment safety, operational readiness, and governance controls for AI workloads. * Contribute to the evolution of enterprise AI quality engineering practices, deployment standards, and monitoring frameworks.