Global Financial Services Institution | Mid–Senior | Location London
Make ML models faster.
Building a model is one thing.
Making it faster, leaner and reliable in production is where the engineering gets interesting.
Our client is looking for a Machine Learning Engineer focused on inference optimisation and performance. You’ll work across risk, payments and client products, where low latency and efficiency genuinely matter. The work is predominantly CPU-based, low batch and latency sensitive, with models running in a regulated environment.
GPU background?
Absolutely still relevant.
If you understand why your CUDA/GPU optimisation improved performance, the CPU inference toolchain can be learned.
What you’ll do
You’ll take production models and find ways to make them perform better:
* Profile models and identify bottlenecks
* Reduce inference latency and improve throughput
* Apply
quantisation
while protecting model accuracy
* Optimise computation graphs and threading
* Benchmark changes against a clear baseline
* Deploy and monitor optimised models
* Identify the next performance improvement
We want engineers who can explain the numbers:
What was the problem? What did you change? What difference did it make?
What you’ll need
* Strong
Python
plus C++, Rust, Go or Java
* Production experience optimising ML inference
* Practical experience with
quantisation
* Experience with at least two of: ONNX Runtime, OpenVINO, oneDNN, IPEX, TVM, TensorRT, vLLM or llama.cpp
* PyTorch or TensorFlow
* XGBoost or LightGBM
* Docker, Kubernetes and CI/CD
* A strong performance-engineering mindset
Useful experience
* CPU optimisation: AVX-512, VNNI, AMX or NUMA
* Real-time or streaming inference
* Kernel-level optimisation
* Financial services
* Open-source ML tooling
Senior level
For Senior Engineers, we want a specific example of an optimisation you led.
What was the baseline? What did you change? What was the measurable result?
What this isn't
This isn't a research role, distributed GPU training role or pure ML platform position.
The focus is simple: make production ML models faster, more efficient and more reliable.
You’ll also work within a regulated environment, so experience with model documentation, validation and monitoring is important.
Confidential enquiries:
justin.toomey@datatech.org.uk
More AI roles like this, weekly
Roles like this expire in about a week. Get new AI openings across the UK in your inbox, free, unsubscribe any time.