Machine Learning Engineer

Datatech Analytics

England, United KingdomfulltimeFinancial Services and Bankingposted 24 Aug
Unlock apply linkApply links and the original listing are a Pro feature: £4.99/mo or £25 once.
Global Financial Services Institution | Mid–Senior | Location London Make ML models faster. Building a model is one thing. Making it faster, leaner and reliable in production is where the engineering gets interesting. Our client is looking for a Machine Learning Engineer focused on inference optimisation and performance. You’ll work across risk, payments and client products, where low latency and efficiency genuinely matter. The work is predominantly CPU-based, low batch and latency sensitive, with models running in a regulated environment. GPU background? Absolutely still relevant. If you understand why your CUDA/GPU optimisation improved performance, the CPU inference toolchain can be learned. What you’ll do You’ll take production models and find ways to make them perform better: * Profile models and identify bottlenecks * Reduce inference latency and improve throughput * Apply quantisation while protecting model accuracy * Optimise computation graphs and threading * Benchmark changes against a clear baseline * Deploy and monitor optimised models * Identify the next performance improvement We want engineers who can explain the numbers: What was the problem? What did you change? What difference did it make? What you’ll need * Strong Python plus C++, Rust, Go or Java * Production experience optimising ML inference * Practical experience with quantisation * Experience with at least two of: ONNX Runtime, OpenVINO, oneDNN, IPEX, TVM, TensorRT, vLLM or llama.cpp * PyTorch or TensorFlow * XGBoost or LightGBM * Docker, Kubernetes and CI/CD * A strong performance-engineering mindset Useful experience * CPU optimisation: AVX-512, VNNI, AMX or NUMA * Real-time or streaming inference * Kernel-level optimisation * Financial services * Open-source ML tooling Senior level For Senior Engineers, we want a specific example of an optimisation you led. What was the baseline? What did you change? What was the measurable result? What this isn't This isn't a research role, distributed GPU training role or pure ML platform position. The focus is simple: make production ML models faster, more efficient and more reliable. You’ll also work within a regulated environment, so experience with model documentation, validation and monitoring is important. Confidential enquiries: justin.toomey@datatech.org.uk