Home / Insights / Hiring for AI in Drug Discovery
Pharma · AI

Hiring for AI in drug discovery: the four roles every R&D function got wrong.

R&D hired the visible role, computational chemists and ML scientists, then found the bottleneck was never the model. It was the data feeding it and the loop validating it. Four roles most R&D functions under-hired.

— Key Takeaways
4
roles every R&D function under-hired
0
come from the med-chemistry pipeline
DMTA
the loop a modeling team can’t close alone

The model was never the bottleneck.

When AI arrived in drug discovery as a board-level priority, most R&D functions did the intuitive thing: they hired modeling talent: computational chemists and ML scientists who could build generative models for molecules and predictors for binding, toxicity, and ADMET. If AI is going to discover drugs, hire the people who build the AI. Sound on its face.

Then the acceleration didn’t arrive on schedule. Models were built; pipelines didn’t visibly speed up. The reason was rarely the model. It was that the model ran on experimental data fragmented across assay systems, ELNs, and acquired datasets: inconsistently labeled, poorly governed, not remotely model-ready. A strong model on weak data doesn’t fail loudly. It produces confident predictions no one can validate, reproduce, or defend: the more expensive failure.

Pharma overinvested in computational chemists and underinvested in the data engineers who actually feed them. The computational chemists weren’t the wrong hire. They were the incomplete hire.

The four roles every R&D function got wrong.

The visible hire: the model. Computational chemists and ML scientists, filled first and fastest, because the role is legible, fits the PhD-science culture, and is what the field talks about at conferences. It was the incomplete hire.
built on four roles R&D under-hired
The scientific data engineer

What they doBuilds the pipelines that turn heterogeneous experimental data (assays, high-throughput screens, omics, ELN/LIMS exports, acquired datasets) into clean, harmonized, model-ready data with provenance intact. The rate-limiting role, and the one most often missing.

Where to recruit themTechnology and data-heavy industries, not the bench, with enough scientific literacy not to destroy meaning while cleaning. Rarely reached through the med-chem channel.

The ML platform engineer

What they doMLOps for science: productionizing models so they’re reproducible, versioned, and reusable: GPU orchestration, experiment tracking, model registries, the deployment path that makes a validated model a standing capability rather than a notebook one-off.

Where to recruit themML-platform and infrastructure teams in technology, a different labor market than R&D recruiting knows.

The translational lab-in-the-loop scientist

What they doCloses the loop between in-silico prediction and wet-lab reality, driving the design-make-test-analyze cycle so predictions get synthesized and tested, and results flow back to retrain the model. Where output becomes validated leads instead of a ranked list no one acts on.

Where to recruit themA thin band of dual-fluency talent trusted by the bench. Scarce, and rarely the same people who build the models.

The scientific data steward

What they doGovernance: provenance, FAIR principles, ontologies, and the integrity standards that make AI-derived results defensible, for IP, partnership diligence, and the GxP and data-integrity expectations regulators bring. When a model’s output may support a filing, “where did this data come from” isn’t clerical.

Where to recruit themA blend of data governance, informatics, and regulatory awareness. Unglamorous, under-hired, and what separates defensible science from a liability.

Why this is a hiring problem, not a tooling problem.

A platform with no one to feed it, operate it, close its loop, or govern it is shelfware with a good demo.

The reflex is to buy a platform (a data lake, an ML suite, an integrated discovery environment) and assume the roles come bundled. They don’t. The four roles live in different labor markets than pharma’s PhD-science pipeline: data engineering and MLOps from technology, informatics and governance from the data world, translational science from a thin dual-fluency band. Reaching them means competing with technology employers on comp and culture, through channels R&D talent teams don’t normally use.

The bottom line.

About the Desk

Thunderhawk Pharma Practice

Pharma & Life Sciences Desk · Thunderhawk Technology Partners

The Thunderhawk Pharma & Life Sciences Desk covers workforce strategy across drug discovery, clinical development, and regulatory affairs, from AI and data infrastructure to the GxP-critical roles that keep programs defensible. We write about the skills and hiring motions that turn scientific investment into validated, durable capability.

— Insights Newsletter

One sharp piece. Once a month.

The kind of writing on workforce, AI, and enterprise hiring you'd actually want to read on a Sunday morning. No vendor pitches, no ad copy, no fluff.