Senior Machine Learning Engineer, MLOps

MFK Recruitment

Brentford, England, UKremotefulltimeIT Services and IT Consultingposted
Unlock apply linkApply links and the original listing are a Pro feature: £4.99/mo or £25 once.
Senior Machine Learning Engineer, MLOps Salary: £70,000 to £100,000, depending on experience Location: West London Working arrangement: Predominantly office and customer-site based, with some remote working available Employment: Permanent, full-time MFK Recruitment is recruiting a Senior Machine Learning Engineer, MLOps for an innovative UK technology company developing advanced Artificial Intelligence and Machine Learning solutions for defence, security and other demanding real-world environments. The company specialises in computer vision, perception and autonomy, combining modern deep learning with neuroscience-inspired technology to create AI systems that are accurate, robust and reliable. MFK Recruitment has successfully recruited four Engineers to this company during the past five years, and all four are still with the business. This speaks volumes about its culture, technical challenges and long-term career opportunities. The Role This is a production-focused Machine Learning Engineering position specialising in MLOps infrastructure, model deployment and monitoring. It is not primarily a research or neural-network architecture role. You will develop and improve the infrastructure that takes advanced Machine Learning and computer vision models from research into reliable production environments. Working closely with Machine Learning Engineers, Computer Vision Engineers and Software Engineers, you will ensure that models can be trained efficiently, deployed consistently, monitored effectively and maintained throughout their production lifecycle. The role will involve Python, Docker, CI/CD, experiment tracking, model registries and production monitoring. Depending on your experience, you may also contribute to distributed training, GPU inference optimisation and deployment to edge devices. The company will consider candidates from different MLOps, Machine Learning platform and AI infrastructure backgrounds. You do not need experience with every technology listed below. Responsibilities * Build and maintain production infrastructure for Machine Learning systems. * Develop automated pipelines for model testing, packaging and deployment. * Implement experiment tracking, model versioning and model-registry processes. * Deploy containerised Machine Learning services across production environments. * Monitor model performance, infrastructure reliability and production behaviour. * Support reproducible model training, evaluation and deployment workflows. * Work with Machine Learning and Computer Vision Engineers to productionise deep-learning models. * Help optimise models for reliable and efficient production inference. * Contribute to platform architecture, technical planning and engineering standards. * Support customer deployments in secure and operational environments. Essential Experience * Strong commercial Python development experience. * Experience deploying and supporting Machine Learning or deep-learning models in production. * Strong Docker and containerisation skills. * Experience building and maintaining CI/CD pipelines. * Experience using MLflow, Weights \& Biases, Neptune, ClearML or another experiment-tracking or model-registry platform. * Experience monitoring production Machine Learning services or infrastructure. * Good understanding of the Machine Learning lifecycle, from training and evaluation through to deployment and ongoing maintenance. * Linux and shell-scripting experience. * Strong software engineering practices, including Git, testing and code reviews. * The ability to work effectively with Machine Learning Engineers, Software Engineers and technical customers. Desirable Experience Experience in any of the following areas would be advantageous, but candidates are not expected to have all of them: * Computer vision, CNN or other deep-learning workloads. * PyTorch or TensorFlow training and inference pipelines. * Distributed training using Ray, DeepSpeed, Horovod, PyTorch FSDP or similar. * GPU and CUDA environments. * Model optimisation using TensorRT, ONNX Runtime, OpenVINO, quantisation or pruning. * Kubernetes, Kubeflow, ECS or another container-orchestration platform. * Prometheus, Grafana, Datadog or another observability platform. * NVIDIA Jetson or another edge AI platform. * NVIDIA DeepStream, GStreamer or FFmpeg. * Triton Inference Server, Ray Serve, TorchServe, BentoML or Seldon Core. * Airflow, Prefect, Dagster, Metaflow or another workflow-orchestration platform. * DVC, LakeFS, Delta Lake, Feast or another data-versioning or feature-store technology. * Terraform, Pulumi, CloudFormation or Ansible. * GPU cluster management or high-performance computing. * Aerial, satellite or ISR imagery. * Defence, government or security-sector projects. Location and Working Arrangement The position is based in West London and will involve working from the company’s office and customer sites. Some remote working is available, and occasional travel to other customer locations may be required. Security Clearance Candidates must be willing and eligible to undergo BPSS and Security Check clearance. Existing clearance would be advantageous but is not essential. SC clearance commonly requires approximately five years of UK residency. However, individual circumstances can be assessed, and candidates should not assume that a shorter period of UK residency or time spent overseas will automatically make them ineligible.