At Apheris, we are building the future of how AI is applied in pharmaceutical R&D.
We enable leading pharmaceutical teams to discover and develop drugs faster. We host the industry’s largest federated data networks for drug discovery AI, spanning co-folding, ADMET, and antibody developability.
Across these networks, models are trained on proprietary industry datasets to achieve higher performance and broader applicability while keeping data control and IP protected. We deliver these superior models through drug discovery applications that enable teams to run them at scale, further customize them, and integrate them into existing R&D workflows.
About the role
We're looking for a senior ML engineer to own the training and evaluation pipelines behind our ADMET and toxicity networks - the machinery that turns partner data into trained, benchmarked, released models, run after run.
This is a hands-on engineering role at the point where molecular ML meets federation. Your code runs inside partner environments, on data you cannot see, alongside scientists at some of the largest pharma companies in the world.
You'll work closely with the scientific lead for toxicity: they own what we model and why, you own how it gets built, validated, evaluated and shipped - reliably enough that a partner will stake a drug program on the result.
The ADMET network is live and expanding into toxicity. You'll be building the pipelines that expansion runs on.
What you will do
Own the model pipelines end to end. Take partner data from landing to released, benchmarked model weights - data preparation, training, evaluation, release - for ADMET and toxicity endpoints.
Build the validation that partners run alongside us. Schema and data contracts, validators that return actionable errors, and QC/profiling reports that a partner can act on without us seeing their raw data.
Make federated runs reproducible and auditable. Versioned configs, pinned data snapshots, provenance for every released model - so we can say exactly what produced a given set of weights.
Build evaluation that survives scrutiny. Leakage-safe splitting, held-out benchmarks and honest performance reporting, so the numbers we put in front of partners hold up.
Gehalt für diese Stelle
Was denken andere über diese Stelle?
Stimme ab und sieh sofort, wie andere die Anforderungen, das Gehalt und mehr einschätzen.
Anonym: niemand sieht, wie Du abstimmst.
Was denken andere über diese Stelle?
Stimme ab und sieh sofort, wie andere die Anforderungen, das Gehalt und mehr einschätzen.
Anonym: niemand sieht, wie Du abstimmst.
+
Entdecke die Menschen hinter Apheris
Wirf einen Blick aufs Team: sieh, wer hier arbeitet, und entdecke bekannte Gesichter aus Deinem Netzwerk.
Harden research into product. Turn prototypes and research code into tested, modular systems, and hand them cleanly to engineering for scaling into Foundry.
Work across the boundary. Translate scientific requirements into pipeline behavior with the science team, and surface data or modeling risks early to partners and internally.
What we expect from you
5+ years building ML systems in Python, with genuine software engineering discipline: version control, tested modular code, code review, and interfaces other people can use.
Hands-on molecular ML or cheminformatics - RDKit, fingerprints and descriptors, or graph/transformer models - applied to property, activity or toxicity prediction.
You have built training and evaluation pipelines that other people run, not one-off notebooks.
You understand how molecular ML goes wrong: data leakage, split design, applicability domain, dataset shift, and over-optimistic benchmarks - and you design against them by default.
You can write validators and data contracts that hold up under partial visibility, where you cannot inspect the data yourself.
Comfortable with PyTorch or an equivalent modern ML stack.
You work well with scientists: you can take an ambiguous scientific requirement and turn it into defined, testable pipeline behavior.
Nice to have
Federated learning, privacy-preserving ML, or other multi-party training environments.
ML Ops or ML infrastructure experience, particularly Kubernetes-based training, evaluation or deployment workflows.
Production-grade model delivery in regulated, enterprise, pharmaceutical or biotech settings.
Familiarity with public ADMET, toxicity and bioactivity data resources (ChEMBL, Tox21, ToxCast) and the gotchas in each.
Open-source contributions or a publication record in cheminformatics, molecular ML, or applied machine learning.
What we offer you
Industry-competitive compensation, including early-stage virtual share options
Remote-first working - work where you work best
Wellbeing budget, mental health support, work-from-home budget, co-working stipend, and learning budget
Generous holiday allowance
Office days at our Berlin HQ or a different European location (3x per year)
A high-calibre, execution-focused team with experience from leading organizations
Job-Alert für Data Scientist erstellenWerde benachrichtigt, wenn neue passende Jobs erscheinen