Fully Remote (EU timezone overlap preferred)
Permanent
We are partnering with a tech scale-up to find a top-tier Applied AI Engineer . Our client is heavily investing in governed AI agents designed to work directly with customers and internal teams.
Data and Intelligence sit at the very center of their product ecosystem. Their platform is evolving to carry advanced intelligence: real-time decisioning, predictive modeling, and governed AI agents. They need an exceptional engineer to drive this shift.
If you want to build production agents that answer real business questions and act with real money in a high-stakes, highly governed environment, this is the role for you.
About the Role
Our client’s decision engine is governed by a formal decision register: over 100 decisions across 15 platform modules, spanning real-time risk gates, reward orchestration, payment routing, and responsible gambling interventions.
As an Applied AI Engineer on the Intelligence team, you are the product engineer of the data and agent platform. You will build production agents that are grounded, auditable, secured against adversarial input, and gated by human approval. Your first major deliverable will be a production SQL BI analyst agent: a Slack-native agent that answers executive business questions with governed SQL, validated queries, and cited evidence.
From there, the role balances four crafts in roughly equal measure:
Build and own Slack-native agents end-to-end that translate natural-language questions into governed SQL over the analytical warehouse.
Develop triage and routing systems, retrieval-grounded (RAG) responses over versioned knowledge bases, and multi-turn conversational state machines with strict escalation logic integrated into helpdesk/CRM platforms.
Build the layer between agents and back-office APIs. Harden agents against prompt injection, tool misuse, and data exfiltration. Design multi-turn confirmation workflows where low-risk actions graduate to full autonomy as evidence accumulates.
Build using LangGraph, Anthropic Agent SDK, Model Context Protocol (MCP), or equivalent. Engineer closed feedback loops, decision audit logging, and strict evaluation harnesses.
Build and productize models for churn, LTV, bonus sensitivity, and composite player risk scores (identity, payment, gameplay, bot-play, multi-accounting).
Ship models as versioned, SLA-backed contracts with the decision engine—served across real-time (Kafka/MSK) and batch (ClickHouse) tiers, complete with drift monitoring and automated retraining.
Own multi-vector withdrawal risk scoring, including cited rationale, evidence-aware aggregation, and automatic re-scoring.
Codify business rules with domain owners. Simulate and backtest every threshold change against historical data, designing holdouts and control groups to measure true uplift.
Build review queues, one-click action proposal cards for high-risk mutations, and searchable session replays exposing prompts, reasoning chains, and tool calls.
Design KPI views and decision audit dashboards that move toward AI-assisted anomaly detection, built in collaboration with the BI team.
4+ years in software, data science, or ML engineering, with 1+ years specifically building LLM-powered agents in production (tool use, structured outputs, memory, multi-step orchestration via LangGraph/Anthropic SDK).
You have shipped production RAG systems and thoroughly understand grounding, chunking, hallucination control, and refusal triggers.
You treat customer-facing agents as an attack surface and know how to defend against prompt injection and data leakage.
Proven experience with feature engineering, training, serving, and monitoring. Experience in fraud, risk, or abuse detection (imbalanced classes, adversarial users) is a massive plus.
Deep knowledge of experiment design, score calibration, and uplift measurement.
Strong Python for backend/services, TypeScript for frontend approval surfaces, and excellent SQL (ClickHouse preferred).
You can spin up lightweight internal apps (e.g., Streamlit, frontend frameworks) so you don't have to wait for another team to make your work visible.
Experience keeping the boundary between probabilistic models and deterministic execution clean, utilizing LLM observability tools (Langfuse, LangSmith, etc.).
How to Apply
We are managing the introduction process for this client exclusively through hiCalibre .
To get introduced with the client via the hiCalibre platform and ensure your application is fast-tracked, create your profile here now: hicalibre.io/join
Roles like this expire in about a week. Get new AI Engineer openings across the UK in your inbox, free, unsubscribe any time.