Our Mission is to democratise the global research industry by providing unbiased business intelligence through analysis of mass public data, using Machine Learning and NLP techniques. We have recently been acquired by Beroe, but have kept the 'startup feeling' within our team, which is now looking to grow.
Curious, Helpful, Independent, Responsible, Remarkable
The role sits at the front of our data pipeline, designing and training the NLP components that our ETL team productionizes downstream.
While we use agentic LLMs throughout our stack, this role is rooted in traditional NLP craft : you should be as comfortable reasoning about vector comparison, clustering behavior, and tokenization edge cases as you are prompting a model. You'll work primarily in vector space — building, tuning, and evaluating embedding-based systems that transform raw text into structured, meaningful signal.
Output from this role — models, embedding pipelines, evaluation harnesses, labeled datasets and classifiers etc — becomes the spec that our ETL Engineers build into production services, so clear documentation and handoff discipline matter as much as the research itself.
We recognize the importance of proprietary data, and is in the process of creating an extensive network graph capturing decades of analyst insights as well as public information, allowing Graph RAG into the brains of an all-knowing procurement expert.
We work a 4-day week, with on-call every other Friday (business hours only).
*vibes* or leaderboard scores;
You have a deep, intuitive understanding of embedding spaces — how to construct them, index and query them efficiently (e.g. FAISS), and critically evaluate whether they're actually capturing what you think they are.
You find elegance in solving a problem with the smallest model that reliably works, and you treat GPU spend and spin-up time as a cost to justify, not a resource to reach for by default.
You reach for machine learning when it's the right tool, and you're equally willing to argue against it — or against a heavier model — when a simpler statistical or rule-based method will do the job more reliably or cheaply.
You don't trust a model until you've stress-tested it. You design evaluation frameworks before you design the model, and you're honest about failure modes.
You use agentic LLMs and coding assistants to move faster, but your judgment — not the model's confidence — is the final word on quality.
You value clean, reproducible research code, appreciate versioned datasets, and document your methodology so others can build on it without reverse-engineering your notebook.
You take ownership of open-ended research questions, and you're comfortable defining the problem, but stay clear headed to back out of a rabbit hole to a find second opinion.
You understand that your output is someone else's input — you write for the ETL Engineer who has to productionize your work, not just for yourself.
You stay current with NLP research, but you evaluate new techniques on evidence, not hype, as well as considering maintenance before recommending they enter the pipeline.
:
If you believe you are up for the challenge of being a Forestreeter, we have the following advice before you apply:
.
*we love to see your Github* . It does not need to be perfect – progression is what we want to see.
We are a small team, and we insist on Engineers hiring Engineers : the interviewers are the same people you will be working alongside. However this may mean that we may not have lightning response time in the middle of a deployment!
Roles like this expire in about a week. Get new AI Research openings across the UK in your inbox, free, unsubscribe any time.