Overview
Novogaia is an applied AI drug discovery company. We build machine learning systems that decode the chemistry of natural organisms, starting with fungi, to find the next generation of medicines.
We are a small team of AI engineers, computational biologists and chemists building foundation models for molecular structure prediction from mass spectrometry data. We are seeking a computational data scientist who can define what a trustworthy spectral dataset looks like, and build the schema, QC gates, and annotation process that gets us there. This is a data and cheminformatics role, not a wet-lab role: though you'll work closely with analytical chemistry collaborators who run the instruments.
The Role
* Developing a deep understanding of Novogaia's compound library, spectral data, and how both feed into our models
* Working closely with our machine learning team to assess model training/validation leakage, de-duplicate against public datasets, and help select representative subsets for benchmarking
* Working with analytical chemist collaborators to route ambiguous or high-value spectra for expert review, so results can be compared systematically against model predictions
* Building and evaluating predictive models to infer molecular properties from molecular structure
* Validating processing workflows for raw spectra arriving from analytical partners, including validating metadata, batch tracking, versioned releases, maintaining provenance, and licensing tags at the record level
In your first year, you'll build the data foundation everything else depends on: a documented schema, a QC process, and a dataset our modeling and evaluation teams can trust.
What We Require
* Background in analytical mass spectrometry or cheminformatics (PhD or equivalent industry experience), ideally with exposure to natural products or small-molecule drug discovery
* Deep familiarity with structural representation methods such as SMILES, SMARTS, SAFE, as well molecular fingerprinting and structural embedding
* Familiarity with statistics and ML concepts, for close collaboration with the rest of the team
* Hands-on experience working with LC-MS/MS data and standard formats and open-source tools (e.g. mzML, MSConvert) and spectral databases (e.g. GNPS, MassBank, MoNA)
* Scripting ability in Python for developing algorithms and Nextflow/Snakemake for building pipelines
* Familiarity with utilizing relational database schemas and ontologies to host and structure the variety of datatypes and datasets you will encounter
What We Value
* Ability to intuitively interpret and assess mass spectrometry data and corroborate automated QC checks
* Strong scientific judgment and a willingness to flag data that isn't ready, even under deadline pressure
* Ability to turn "make this dataset AI-ready" into a concrete schema, checklist, and pipeline
* Motivation to build data infrastructure other people will confidently rely on
* Experience with natural product dereplication and compound classification
* Curiosity, low ego, and a willingness to get close to the modeling and evaluation side of the work, even if it's outside your original training
More AI roles like this, weekly
Roles like this expire in about a week. Get new AI openings across the UK in your inbox, free, unsubscribe any time.