Valarian Technologies is a dual-use technology company building critical tools to safeguard the future in an era of evolving global security challenges. We're rethinking security beyond traditional military domains, addressing asymmetric threats that impact our technological advantage, economic strength, and democratic institutions.
We build Acra – the platform foundation for everything we do as a dual-use technology company. The platform’s name, rooted in the Greek word for citadel (or, fortress), reflects the design and purpose of our infrastructure-agnostic secure enclaves: protecting critical data. Some of the government and commercial workflows include: increased operational resiliency for mission-critical systems and functions; enabling organisations to more quickly and widely adopt emerging technologies while ensuring the integrity of their intellectual property; information flow during disaster response scenarios, and zero-trust / least-privilege environments for M\&A, attorney-client privileged communications, etc. And we’ve only scratched the surface.
At our core, we're driven by a shared mission and a belief in making a tangible impact on our world. Whether you join our London HQ or the wider global organisation, you’ll be a part of collaborative, high-performing teams, creating cutting-edge software, platforms, and infrastructure.
The Role
As an AI Harness Engineer at Valarian, you will own the experimental design, evaluation methodologies, and benchmark infrastructure for our AI models and autonomous agentic workloads. In an emerging domain with no standard playbook, you will serve as the bridge between rigorous data science, statistical validation, and production agent scaffolding.
You will design robust evaluation datasets, architect calibrated LLM-as-a-judge pipelines, and build metrics that accurately capture multi-step agent performance under non-deterministic conditions. Your insights and error analyses will directly drive iterative improvements to our harness scaffolding, tool orchestration, and system safety guardrails.
Design and maintain scalable Python evaluation harnesses that measure task completion, trajectory quality, tool-use correctness, cost, latency, and variance across multi-step agentic systems.
Apply rigorous statistical methods—including hypothesis testing, confidence intervals, bootstrapping, and power analysis—to determine whether performance changes are true system improvements or stochastic noise.
Develop automated judging systems validated against human ground truth. Measure alignment using agreement metrics (Cohen’s kappa, correlation, precision/recall) and systematically detect judge failure modes such as position bias, verbosity bias, self-preference, and prompt sensitivity.
Define sampling strategies, detailed annotation guidelines, and scoring rubrics. Measure inter-annotator agreement and manage label noise to ensure benchmark integrity over time.
Conduct hands-on failure analysis on agent runs, cluster root causes into actionable error categories, and clearly communicate findings to both engineering teams and stakeholders.
Translate evaluation results directly into architectural improvements across agent scaffolding, prompt design, tool execution loops, context window compaction, retries, and safety guardrails.
Demonstrated expertise in hypothesis testing, bootstrapping, power analysis, and statistical significance within probabilistic machine learning systems.
Proven experience building, auditing, and validating LLM judge setups against human annotations, including tracking agreement metrics and mitigating judge biases.
Strong track record designing evaluation datasets, crafting precise scoring rubrics, managing label noise, and measuring inter-annotator agreement.
Deep comfort evaluating complex agent trajectories, tool execution sequences, latency, cost, and variance across repeated runs rather than relying solely on single-shot benchmarks.
Exceptional judgment in defining metrics tied to real-world outcomes, with a keen eye for detecting benchmark gaming, target drift, or shortcut learning.
Ability to dissect complex execution logs, categorize failure modes, and synthesize clear, data-driven recommendations.
Practical proficiency with LLM APIs, prompt engineering, structured outputs, function/tool calling, and retrieval systems in Python.
Clear understanding of planning loops, tool orchestration, memory management, and multi-agent coordination, including their trade-offs and common failure modes.
Ability to write clean, maintainable Python code to iterate on harness components such as guardrails, retry logic, state handling, and context management.
Our benefits are designed to ensure our employees feel taken care of and are proud to be a part of the Valarian team. We are committed to consistently enhancing our benefit package, taking into account the overall well-being and needs of our teammates. Here are the key benefits accessible to all employees at Valarian Technologies:
Life at Valarian
Our culture is built on inclusivity, compassion and flexibility – we want everyone to be empowered to achieve their goals at Valarian.
The work we do is vital, but so are the connections that make it happen. We thrive on the shared energy, spontaneous conversations, and mutual trust built when we spend time together. We operate on a hybrid model, gathering in our London office 3 days a week to support one another and collaborate.
*We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.*
Valarian Technologies Limited is an equal opportunity employer and welcomes applications from individuals regardless of race, colour, religion, sex, sexual orientation, gender, identity or expression, national origin, age, disability, genetic information, marital status, veteran, amnesty, or any other legally protected characteristic.
We are committed to ensuring a fair and inclusive recruitment process and providing employment opportunities to all applicants. Decision recruitment, hiring, and employment are based solely on qualifications, skills, and experience relevant to the job requirements.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Roles like this expire in about a week. Get new AI Engineer openings across the UK in your inbox, free, unsubscribe any time.