Research programme

Research

Lysosomal storage diseases are rare, devastating conditions in which a single enzyme deficiency allows toxic metabolites to accumulate. They have proved stubbornly difficult to cure through conventional drug discovery, and the therapies that do exist are expensive, short-lived in the blood, and often provoke an immune response.

We combine generative AI with engineering biology to design optimised biotherapeutics — from custom enzyme variants to cell therapies — and to manufacture them at a cost that makes treatment sustainable, moving from computational design to molecular repair.

Therapeutic target

Lysosomal storage diseases

Enzymes are the building blocks of cellular life, acting as natural catalysts able to accelerate almost any reaction. Enzymatic deficiencies are associated with devastating rare diseases, treatable only by supplying the missing enzyme intravenously — enzyme replacement therapy, or ERT.

ERT works, but poorly. Recombinant enzymes lose catalytic activity in blood and often provoke a severe immune response, and current manufacturing routes are inefficient enough to make treatment dramatically expensive. Our mission is a next generation of replacement therapies that are both effective and sustainable to use, focusing on Fabry disease.

A cell with its nucleus and organelles, with one lysosome enlarged in a callout. Inside it an enzyme is shown binding accumulated substrate and breaking it down into smaller products.
Lysosomal dysfunction and the therapeutic opportunity for engineered replacement enzymes.

Design

AI for protein engineering

Designing a biologic means finding the amino acid or nucleotide changes that maximise its therapeutic properties without sacrificing manufacturing yield — a search space far too large to explore by hand.

We build generative models that propose enzyme variants directly, learning the sequence determinants of stability, activity and expressibility from data. Our backbone frameworks are variational autoencoders, optimised for training and design efficiency using experience built across twenty years of statistical learning and optimisation.

Our methods are built on robust engineering practice — Python, Git, GitHub and GPU computing — and deployed through the Nextflow workflow management system.

A variational autoencoder. An amino acid sequence passes through an encoder into a Dirichlet latent space, drawn as a triangle scattered with points, and a decoder reconstructs a sequence from it.
A variational autoencoder learns a latent space of protein sequences; sampling that space proposes new enzyme variants.

Production

Engineering biology for bioproduction

Advances in DNA synthesis, sequencing and computer-aided design let us engineer the genomes of living cells directly and tackle problems intractable with standard technologies. We pioneered CAD software for synthetic genome engineering, work that was instrumental in designing Saccharomyces cerevisiae 2.0 — the first synthetic eukaryotic genome ever built.

We now apply that expertise to production hosts, engineering Komagataella phaffii and Chinese hamster ovary cell lines to produce human lysosomal enzymes efficiently and at a cost that makes treatment viable.

Two engineered production hosts, each carrying a designed plasmid: Komagataella phaffii feeding a stirred bioreactor, and a Chinese hamster ovary cell feeding a culture flask. Both routes converge on human lysosomal enzymes.
Host engineering for high-yield expression of human lysosomal enzymes.

Validation

Lab automation

AI models can generate thousands of plausible biologics, but experimental testing decides which are actually therapeutic. Classical low-throughput protocols cannot keep pace with a generative design loop.

So we miniaturise and automate: high-throughput electroporation, liquid handling robotics and rapid assays for protein expression in both K. phaffii and CHO. We work closely with the Edinburgh Genome Foundry to scale this work.

An automated workflow: Komagataella phaffii and Chinese hamster ovary cells expressing lysosomal enzymes, a liquid-handling robot filling microplates, a plate reader, and a screen showing the resulting assay data.
Automated expression and assay pipeline closing the design–build–test loop.

Where the methods came from

The lab's statistical genetics work on inherited cancer risk — dissecting how high-frequency p53 variants shape susceptibility and treatment response — built the modelling toolkit we now apply to protein design. That lineage runs through our publication record, and the questions it raised about learning from sparse biological data still shape how we work.

Funding

EPSRC BBSRC Fujifilm Diosynth Biotechnologies IBioIC Wellcome UKRI Centre for Doctoral Training in Biomedical AI University of Edinburgh