R&D hired the visible role, computational chemists and ML scientists, then found the bottleneck was never the model. It was the data feeding it and the loop validating it. Four roles most R&D functions under-hired.
When AI arrived in drug discovery as a board-level priority, most R&D functions did the intuitive thing: they hired modeling talent: computational chemists and ML scientists who could build generative models for molecules and predictors for binding, toxicity, and ADMET. If AI is going to discover drugs, hire the people who build the AI. Sound on its face.
Then the acceleration didn’t arrive on schedule. Models were built; pipelines didn’t visibly speed up. The reason was rarely the model. It was that the model ran on experimental data fragmented across assay systems, ELNs, and acquired datasets: inconsistently labeled, poorly governed, not remotely model-ready. A strong model on weak data doesn’t fail loudly. It produces confident predictions no one can validate, reproduce, or defend: the more expensive failure.
What they doBuilds the pipelines that turn heterogeneous experimental data (assays, high-throughput screens, omics, ELN/LIMS exports, acquired datasets) into clean, harmonized, model-ready data with provenance intact. The rate-limiting role, and the one most often missing.
Where to recruit themTechnology and data-heavy industries, not the bench, with enough scientific literacy not to destroy meaning while cleaning. Rarely reached through the med-chem channel.
What they doMLOps for science: productionizing models so they’re reproducible, versioned, and reusable: GPU orchestration, experiment tracking, model registries, the deployment path that makes a validated model a standing capability rather than a notebook one-off.
Where to recruit themML-platform and infrastructure teams in technology, a different labor market than R&D recruiting knows.
What they doCloses the loop between in-silico prediction and wet-lab reality, driving the design-make-test-analyze cycle so predictions get synthesized and tested, and results flow back to retrain the model. Where output becomes validated leads instead of a ranked list no one acts on.
Where to recruit themA thin band of dual-fluency talent trusted by the bench. Scarce, and rarely the same people who build the models.
What they doGovernance: provenance, FAIR principles, ontologies, and the integrity standards that make AI-derived results defensible, for IP, partnership diligence, and the GxP and data-integrity expectations regulators bring. When a model’s output may support a filing, “where did this data come from” isn’t clerical.
Where to recruit themA blend of data governance, informatics, and regulatory awareness. Unglamorous, under-hired, and what separates defensible science from a liability.
The reflex is to buy a platform (a data lake, an ML suite, an integrated discovery environment) and assume the roles come bundled. They don’t. The four roles live in different labor markets than pharma’s PhD-science pipeline: data engineering and MLOps from technology, informatics and governance from the data world, translational science from a thin dual-fluency band. Reaching them means competing with technology employers on comp and culture, through channels R&D talent teams don’t normally use.
The kind of writing on workforce, AI, and enterprise hiring you'd actually want to read on a Sunday morning. No vendor pitches, no ad copy, no fluff.