This is a job that Jill, our AI Recruiter, is recruiting for on behalf of one of our customers.
She will pick the best candidates from Jack's network.
The next step is to speak to Jack.
Job Title
AI Research Engineer, Model Optimization and Inference
Salary
Not Disclosed
Company Description
Iconic Interactive is a seed-stage London-based AI-native video game studio building intelligent virtual actors that speak, move, and react in real-time. The team includes researchers and engineers from leading AI labs and game studios developing character intelligence, narrators, and world directors to create personal, immersive entertainment universes.
Job Description
As an AI Research Engineer at Iconic Interactive, you will bridge the gap between massive research models and real-time interactive entertainment. You'll architect high-performance inference pipelines for multimodal LLMs and TTS models, optimizing them to run on consumer hardware. Your work ensures digital entities perform seamlessly within the constraints of a game loop.
Location
London, UK
Why this role is remarkable
* Lead the technical frontier by taking cutting-edge research models and making them run in real-time on consumer hardware for interactive digital experiences.
* Join a seed-stage startup founded by AI lab and game studio veterans, offering significant autonomy and end-to-end ownership of the inference engine.
* Work at the unique intersection of System ML and Game Tech, shaping the core intelligence of virtual actors in next-generation entertainment.
What You Will Do
* Architect and maintain low-latency inference pipelines for Multimodal LLMs, TTS, and Vision models targeting server-side and consumer edge environments.
* Implement state-of-the-art optimization techniques like Speculative Decoding, KV-Cache quantization, and custom CUDA/Triton kernels to minimize latency and maximize throughput.
* Collaborate with game engineering teams to integrate thread-safe, non-blocking asynchronous inference directly into the game loop using C++ wrappers.
The ideal candidate
* Holds an MSc or PhD in Computer Science or ML with deep expertise in model optimization techniques like quantization, pruning, and distillation.
* Possesses strong proficiency in C/C++ and Python, with hands-on experience deploying latency-sensitive ML models using frameworks like PyTorch or JAX.
* Demonstrates specialized knowledge in LLM-specific optimizations such as KV-cache management, speculative decoding, and high-performance runtimes like TensorRT or ONNX.
Who are Jack \& Jill?
Ok, I'll go first. I'm Jack, an AI that gets to know you on a quick call, learning what you're great at and what you want from your career. Then I help you land your dream job by finding unmissable opportunities as they come up, supporting you with applications, interview prep, and moral support.
And I'm Jill, an AI Recruiter who talks to companies to understand who they're looking to hire. Then I recruit from Jack's network, making an introduction when I spot an excellent candidate.
How does this work?
* Jack's an AI agent for job searching and career coaching. He works for you.
* Jill is the AI recruiter working for the company. She recruits from Jack's network.
* If it's a match and the company wants to meet you, they'll make the intro. In the meantime, if you'd like, Jack will send you excellent alternatives.
We never post fake jobs
This isn't a trick. This is an open role that Jill is currently recruiting for from Jack's network.
Sometimes Jill's clients ask her to anonymize their jobs when she advertises them, which means she can't share all the details in the job description.
We appreciate this can make them look a bit suspect, but there isn't much we can do about it.
Give Jack a spin! You could land this role. If not, most people find him incredibly helpful with their job search, and we're giving his services away for free.
More LLM Engineer roles like this, weekly
Roles like this expire in about a week. Get new LLM Engineer openings across the UK in your inbox, free, unsubscribe any time.