Data Scientist
AI for Science \& Knowledge Discovery
Technology – Data Science Organization
Do you want to build AI that helps the world’s scientists discover, understand, and advance human knowledge?
At Elsevier, data science isn’t about models, metrics, and pipelines for their own sake. It’s about doing science for science: building AI that helps researchers, clinicians, educators, authors, editors, institutions, and innovators make better decisions, uncover hidden connections, and accelerate scientific progress.
We are growing our Data Science organization within Technology, and we are looking for data scientists who want to work at the frontier of applied AI — machine learning, natural language processing, information retrieval, knowledge discovery, and state-of-the-art generative AI built specifically for science.
You will take on some of the hardest problems in science: building intelligent systems that can reason across scientific publications, research data, knowledge graphs, ontologies, metadata, taxonomies, citations, and content spanning every scientific discipline.
This is where advanced AI meets a mission that truly matters — helping researchers advance science and human progress, together.
About The Role
As a Data Scientist at Elsevier, you will design, build, evaluate, and productionize the AI and machine learning solutions that power the scientific and knowledge-discovery applications used by research and healthcare communities around the world.
You will work with some of the richest and most challenging data anywhere — enormous, heterogeneous, and intellectually deep: scientific literature, raw research datasets, structured and unstructured content, metadata, citations, ontologies, knowledge graphs, domain taxonomies, and large-scale behavioral and usage signals.
You will pair strong data science fundamentals with the most capable modern AI: classical machine learning, deep learning, large language models, retrieval-augmented generation, semantic search, ranking, entity extraction, classification, recommendation, and evidence-grounded generative AI.
You will own meaningful problems from day one — whether that means shaping a critical model component, designing an end-to-end AI capability, or driving the technical direction of complex data science solutions across teams.
We value practical problem-solving, scientific curiosity, engineering discipline, and the thoughtful use of technology. We don’t chase trends for their own sake. We choose the right method for the problem, measure impact rigorously, and build systems that are reliable, scalable, explainable, and genuinely useful in the real world.
What You’ll Do
You will work on high-impact AI and data science that helps people explore, understand, connect, and act on complex scientific information.
In This Role You Will
- Design and build machine learning, NLP, and generative AI systems for scientific discovery, knowledge extraction, decision support, and intelligent content understanding.
- Work with large-scale, complex, and heterogeneous data — scientific publications, research datasets, knowledge graphs, ontologies, taxonomies, citations, metadata, and content from every scientific discipline.
- Apply the right technique to each problem, from classical ML methods such as classification, regression, clustering, ranking, and feature engineering to modern approaches using deep learning, embeddings, LLMs, retrieval, and generative AI.
- Develop capabilities for semantic search, information retrieval, entity extraction, content classification, recommendation, ranking, summarization, question answering, and evidence-grounded generation.
- Build, evaluate, and continuously improve models that help researchers discover relevant knowledge faster, make better decisions, and advance science.
- Train, fine-tune, prompt, evaluate, and integrate models into robust production systems.
- Design evaluation frameworks that measure quality, relevance, reliability, model performance, user value, and real-world impact.
- Write clean, tested, production-quality Python and contribute to reusable data science components and packages.
- Build and maintain scalable data pipelines for preprocessing, model inference, experimentation, monitoring, and continuous improvement.
- Collaborate closely with engineering, product, UX, analytics, research, and domain experts to turn ambiguous scientific and business challenges into robust technical solutions.
- Support deployment, monitoring, model maintenance, drift detection, automated retraining, and ongoing optimization of data science systems.
- Communicate technical concepts, model behavior, insights, trade-offs, and recommendations clearly to both technical and non-technical audiences.
What Makes This Opportunity Unique
At Elsevier, you won’t be solving generic AI problems. You will be building AI for the global knowledge ecosystem — the systems that help science itself move forward.
You will help build AI that understands scientific language, connects ideas across disciplines, surfaces trustworthy evidence, reveals hidden relationships between concepts, and supports researchers as they push the boundaries of human understanding.
The data is vast. The problems are among the hardest in science. The mission could not matter more.
You May Work With
- Scientific publications and full-text content.
- Research metadata, citations, abstracts, author networks, and institutional data.
- Scientific raw-data repositories and structured research datasets.
- Knowledge graphs, ontologies, taxonomies, vocabularies, and entity networks.
- Multidisciplinary content across life sciences, physical sciences, engineering, medicine, social sciences, and beyond.
- User behavior and interaction signals that sharpen discovery, relevance, personalization, and decision support.
- Generative AI systems that must be accurate, scalable, explainable, source-grounded, and trusted.
Your work will help researchers save time, discover connections, generate new insights, and accelerate the pace of science.
What We’re Looking For
We are interested in people who bring a mix of curiosity, technical depth, collaboration, and a drive for real impact — and we will shape the role around the strengths you bring. You might be early in your career, with strong foundations in machine learning, analytics, Python, experimentation, and applied problem-solving, and eager to grow your expertise in modern AI, NLP, and scientific data. You might already design, build, evaluate, and deliver data science components and models independently across the full lifecycle. Or you might lead the design of complex AI and machine learning solutions — shaping technical approaches, weighing trade-offs, mentoring others, and influencing stakeholders. Wherever you sit on that spectrum, we want to hear from you.
Core Qualifications
- A background in data science, machine learning, artificial intelligence, NLP, statistics, applied mathematics, computer science, or a related quantitative field.
- Strong Python skills and a habit of writing clean, maintainable, well-tested code.
- A solid grasp of machine learning fundamentals, including supervised and unsupervised learning, feature engineering, model evaluation, model selection, and performance measurement.
- Experience working with structured, semi-structured, or unstructured data, especially large-scale text or content datasets.
- Familiarity with common data science and ML tools such as Pandas, NumPy, SciPy, Scikit-learn, PyTorch, TensorFlow, or Matplotlib.
- The ability to translate complex and ambiguous requirements into practical, measurable, data-driven solutions.
- Strong analytical thinking, problem-solving skills, and attention to quality.
- Clear communication, with the ability to explain technical concepts, insights, and trade-offs to different audiences.
- A collaborative approach to working with engineering, product, and business stakeholders.
- A genuine interest in building production-ready systems that deliver real user value.
AI, NLP and Modern Machine Learning Experience
Relevant experience might include any of the following — we don’t expect all of it:
- Large language models and applied LLM workflows.
- Retrieval-augmented generation, semantic search, embeddings, ranking, and information retrieval.
- NLP techniques such as classification, entity extraction, clustering, summarization, topic modeling, question answering, and text generation.
- Deep learning and neural network approaches for language, content understanding, or prediction.
- Evaluation of LLM outputs, AI-generated content, source-grounded answers, and evidence-based systems.
- Agentic or orchestrated AI workflows, including multi-step reasoning, tool use, and frameworks such as LangChain, LangGraph, or similar.
- Knowledge systems, knowledge graphs, ontologies, taxonomies, or citation-aware AI.
- Knowing when to reach for classical machine learning, deep learning, LLM-based approaches, rules, retrieval, or hybrid systems.
- Model monitoring, drift detection, retraining strategies, performance reporting, and continuous improvement.
Tools and Technologies
Experience With Some Of The Following Is Valuable
- Python, SQL, Git, CI/CD, unit testing, and collaborative software development practices.
- Scikit-learn, PyTorch, TensorFlow, Hugging Face, or similar ML/AI frameworks.
- Cloud platforms such as AWS or Azure.
- MLOps tools such as MLflow, Kubeflow, SageMaker, or similar.
- Big data and distributed processing frameworks such as Spark, Hadoop, Databricks, or equivalent.
- Data formats and interfaces such as JSON, XML, REST APIs, microservices, and relational or document databases.
- Experiment tracking, reproducible workflows, model registries, and structured development environments.
- System optimization for scale, latency, cost, maintainability, and reliability.
Nice to Have
- Experience in scientific, technical, medical, academic, publishing, research, healthcare, or other knowledge-intensive domains.
- Experience with search, recommendation, ranking, information retrieval, or knowledge-discovery systems.
- Experience working with large-scale unstructured text or document collections.
- Experience building source-grounded, citation-aware, or evidence-based AI systems.
- Familiarity with ontologies, taxonomies, entity resolution, metadata enrichment, or knowledge graphs.
- Experience productionizing ML or AI systems, including deployment, monitoring, retraining, and optimization.
- Experience with MLOps, cloud-native development, or distributed data processing.
- Interest in emerging approaches such as AI-assisted development, spec-driven development, agentic workflows, and human-in-the-loop evaluation.
- Experience communicating model performance, risk, limitations, and impact to non-technical stakeholders.
Why Join Us
Because your work will matter.
At Elsevier, you will apply AI and data science to problems that support the global research and healthcare communities. You will help build tools that make scientific knowledge more discoverable, connected, trustworthy, and actionable.
You Will Have The Opportunity To
- Tackle some of the hardest AI challenges in science and knowledge discovery.
- Build intelligent systems using vast, rich, and heterogeneous scientific data.
- Combine machine learning fundamentals with state-of-the-art AI — LLMs, retrieval, generative AI, and knowledge systems.
- Contribute to products and platforms used by researchers, clinicians, institutions, and decision-makers around the world.
- Help accelerate discovery, widen access to knowledge, and support human progress.
- Work alongside talented colleagues across data science, engineering, product, design, analytics, and deep domain expertise.
- Grow your career in an organization that values curiosity, quality, inclusion, flexibility, and meaningful impact.
Work in a Way That Works for You
We promote a healthy work/life balance and support flexible working. We know people do their best work when they have the flexibility, trust, and support they need to thrive.
We offer initiatives and benefits that may include wellbeing support, flexible working arrangements, shared parental leave, study assistance, sabbaticals, and ongoing opportunities for learning and development.
Working With Us
We are an equal-opportunity employer committed to helping you succeed.
You will find an inclusive, collaborative, agile, innovative, and supportive environment where everyone has a part to play. We value diverse perspectives, thoughtful debate, practical problem-solving, and colleagues who care deeply about what they do and how they do it.
About Elsevier
Elsevier is a global leader in information and analytics. We help researchers and healthcare professionals advance science and improve health outcomes for the benefit of society.
Building on our publishing heritage, we combine quality information, vast datasets, advanced analytics, and innovative technologies to support visionary science and research, health education, interactive learning, and exceptional healthcare and clinical practice.
At Elsevier, your work contributes to the world’s grand challenges and a more sustainable future. We harness technology to support science and healthcare in partnership with the communities we serve.
Together, we create possibilities. Join us.
If performed in NLD Amsterdam (Radarweg), the base pay range is €49,000 - €81,700. This job may be subject to a collective labor agreement in the Netherlands. Please consult with the hiring team for further details.
We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits. Click here to access benefits specific to your location.