Founding Engineer - AI Safety & Optimisation

Few&Far

London Area, United KingdomfulltimeTechnology, Information and Media and Software Developmentposted
Unlock apply linkApply links and the original listing are a Pro feature: £4.99/mo or £25 once.
Founding ML Engineer — AI Safety \& Optimisation London · In-office £120k–£170k · Meaningful equity *Retained search on behalf of a confidential client - details shared on intro call* About the Client We're working with an early-stage, well-funded AI startup (backed by top-tier VCs) building systems that need to understand and control how complex, large-scale AI behaviour plays out in the real world - before it goes wrong. The work sits right at the intersection of ML performance and safety: models need to be capable, but also predictable, aligned, and robust once they're live in front of real customers. Small team, high ownership, direct access to founders. This is a founding/early hire, not a cog-in-a-machine role. The Role The core of this role is RL, fine-tuning, and reward design -with safety as the design constraint, not an afterthought. You'll take training methods and turn them into systems that are fast and effective, but also well-behaved: models that stay within intended bounds, resist drift, and fail safely rather than silently. You'll own the loop end-to-end: reward/training design, post-training and distillation, and production optimisation - all with an eye on catching and correcting unwanted behaviour before it reaches a customer. What You'll Do * Design reward functions and training setups that optimise for capability *and* safe, predictable behaviour * Post-train, fine-tune, and distil models with alignment and robustness front of mind * Build evaluation and monitoring approaches that catch drift, edge cases, and failure modes early * Optimise inference for scale without compromising on safety guardrails — sub-second, high-volume, production-grade * Build the data pipelines that feed training, evaluation, and safety testing * Take a method from prototype to production, simplifying aggressively while preserving the safety properties that matter What We're Looking For * 2+ years shipping ML in a startup environment, ideally with end-to-end ownership * Strong hands-on experience with RL, fine-tuning, and/or distillation — bonus points if you've thought hard about reward hacking, alignment, or failure modes * Excellent Python engineering — clean, maintainable, production-grade * Comfortable with real-time/low-latency inference systems * Fluent with statistics, probability, and high-dimensional reasoning * A genuine interest in safety-conscious ML, not just raw performance chasing * Fast, high-bar, ownership mentality Nice to have: distributed training experience, distilling frontier models into small open-weights models for production, background in anomaly/behavioural detection, interpretability or evals work, familiarity with cloud infra (AWS/GCP). *And a quick note: if you're reading this and tick maybe 60% of the boxes above, please still get in touch. The best hires I've made rarely matched every bullet on paper - don't rule yourself out*