Enzyme Engineering for In Vitro Diagnostics
Resurrecting Extinct Plant Enzymes for Modern Diagnostics: Engineering Ancient Biosynthetic Pathways into Stable IVD Reagents
Ancestral sequence reconstruction (ASR) converts phylogenetic information from extant homologs into statistically inferred ancient protein.
Why Ancient Enzymes Matter
Diagnostic enzyme development has traditionally drawn on a relatively narrow set of well-characterized catalysts isolated from extant organisms. This pool is deep but not unlimited, and it is biased toward enzymes that happen to be abundant, easy to express, or already familiar to assay developers. Ancestral sequence reconstruction (ASR) offers a different entry point. Rather than screening nature's current inventory, ASR uses phylogenetic analysis of extant homologous protein sequences to produce statistical estimates of ancient sequences, which are then resurrected in the laboratory to test hypotheses about how catalytic activities and pathway assembly evolved. For the diagnostic developer, the practical consequence is access to protein scaffolds that no longer exist in any living organism but can still be synthesized, expressed, and characterized.
The scientific rationale for looking backward is that ancestral enzymes are not simply older versions of modern ones. Studies of resurrected plant enzymes indicate that exaptation is widespread, such that even low or undetectable levels of ancestral activity toward a substrate can later become the apparent primary activity of descendant enzymes. Ancient pathway flux also often differs from modern-day metabolic networks. These observations matter for diagnostics because they imply that ancestral scaffolds may occupy different regions of sequence and stability space than the enzymes currently used in reagent development, potentially offering combinations of activity and robustness that are difficult to find by screening extant homologs alone.
A second rationale is stability. Work on ancestral sequence reconstruction has highlighted the use of ASR to mine for thermostable enzymes that can serve as superior starting points for protein engineering, particularly where the innate stability of a modern enzyme limits the success of an engineering campaign. In one reported case, an ancestral non-heme iron-dependent enzyme displayed an improved evolvability profile relative to its modern counterpart, and engineering the ancestral protein accessed variants with enhanced thermostability and expression, increased rates, and a broader substrate scope. Such findings are background evidence rather than a guarantee of diagnostic performance, but they explain why ancient scaffolds are increasingly considered as candidate starting materials for reagent enzymes.
Statistical inference, not speculation
ASR does not guess at ancient sequences. It derives them from a multiple sequence alignment of extant homologs and a phylogenetic model, producing a statistically estimated ancestral state at each position.
- Phylogenetic tree inference from homologous sequences
- Statistical ancestral sequence inference
- Gene synthesis of the inferred sequence
Plant specialized metabolism
Resurrected plant enzymes have been used to test how catalytic activities and pathway assembly evolved, revealing exaptation and ancient flux patterns distinct from modern networks.
- Exaptation is widespread in plant enzyme evolution
- Intramolecular epistasis can constrain or permit preference switches
- Ancient pathway flux often differs from modern metabolism
Scaffolds for reagent development
Ancestral enzymes are evaluated as starting points, not finished reagents. Biochemical characterization determines whether a scaffold is worth engineering toward an assay.
- Activity and kinetic characterization
- Thermostability and pH tolerance screening
- Comparison against extant homologs
From Phylogeny to Protein
The ASR workflow begins with sequence data. Homologs of the enzyme family of interest are collected, aligned, and used to infer a phylogenetic tree. The quality of this tree and the alignment directly shape the ancestral estimates, so curation of the input set is a substantive scientific step rather than a formality. In published work on family I.3 bacterial lipases, for example, eighty-three protein sequences sharing a minimum sequence identity with a reference lipase were used to infer the phylogenetic tree before the last universal common ancestor sequence of the family was reconstructed. The same logic applies to plant enzyme families, where the depth and diversity of available homologs determine how confidently an ancestral state can be estimated. Teams facing similar bottlenecks often pair this approach with engineering diagnostic enzyme when moving from discovery into validation.
Once the tree is established, statistical ancestral sequence reconstruction infers the most probable ancestral state at each position, producing a complete ancestral protein sequence. That sequence is then synthesized as a gene and expressed recombinantly in a heterologous host. Expression outcomes vary: in the lipase study, the reconstructed ancestral gene was expressed as inclusion bodies in E. coli, and the insoluble protein was refolded by urea dilution and purified by affinity chromatography. The purified ancestral lipase exhibited an optimum temperature and pH that differed from typical modern counterparts, and it tolerated a range of metal ions and organic solvents. These are paper-specific results, but they illustrate the kind of characterization data that resurrection generates.
For diagnostic purposes, the transition from inferred sequence to usable protein is where the method contract matters most. ASR is not directed evolution, and it is not rational design. It does not begin with iterative rounds of random mutagenesis and high-throughput screening of a modern enzyme, nor does it begin with structure-guided site-directed mutagenesis of a current catalyst. It begins with phylogenetic analysis of extant homologs and statistical inference of an ancient sequence. Downstream engineering may follow, but the resurrection step itself is a reconstruction exercise, and keeping that boundary clear helps developers interpret what an ancestral scaffold does and does not represent.
| Stage | Core activity | Typical output | Diagnostic relevance |
|---|
Engineering Resurrected Scaffolds
Resurrection produces a candidate, not a product. The next phase is engineering the ancestral scaffold toward diagnostic performance. This is where directed evolution and related approaches become relevant, but they are applied to the resurrected backbone rather than used as the resurrection method itself. The rationale for starting from an ancestral scaffold is that innate stability can determine whether an engineering campaign succeeds at all; harnessing innately stable enzymes has been described as a way to overcome stability-related bottlenecks and accelerate biocatalyst engineering. In diagnostic terms, a more stable starting point can widen the range of formulation and storage conditions that remain accessible after engineering.
Engineering for diagnostic performance typically involves several parallel objectives. Catalytic activity and substrate specificity must be tuned to the analyte and detection chemistry of the intended assay. Background signal must be controlled, since low background and high specificity are central to assay performance. Thermostability and pH tolerance must be characterized across the range of conditions the reagent will encounter. Where the assay format demands it, matrix and inhibitor tolerance also become engineering targets. These objectives are not independent; improving one property can perturb another, which is why iterative characterization accompanies each engineering round.
The comparison between an ancestral enzyme and its modern counterpart is a useful experimental control. In the reported non-heme iron-dependent enzyme work, the ancestral protein and the modern enzyme were compared in the laboratory, and the ancestor showed an improved evolvability profile. Engineering the ancestral protein accessed variants with enhanced thermostability and expression, increased rates, and a broader substrate scope than the modern counterparts. This does not mean every ancestral scaffold will outperform every modern enzyme, and it does not authorize transferring that result to diagnostic assays as a project-dependent outcome. It does mean that evolvability is a property worth measuring when selecting a starting scaffold.
Define the assay objective
Specify the analyte, detection chemistry, sample matrix, and storage format before selecting or engineering a scaffold, so that activity, specificity, and stability targets are set by the assay rather than by convenience.
Characterize the resurrected scaffold
Measure activity, kinetics, thermostability, pH tolerance, and expression behavior of the ancestral protein and its modern homologs to establish a baseline comparison.
Engineer toward diagnostic performance
Apply directed evolution and related modification strategies to improve activity, specificity, low background behavior, and robustness, with characterization after each round.
Validate against assay conditions
Test engineered variants under the ionic strength, pH, temperature, and matrix conditions of the intended assay, including stability and shelf-life assessment.
Stability and Matrix Tolerance
Thermostability is often the first property examined in a resurrected enzyme because it is both mechanistically informative and commercially consequential. The lipase reconstruction described above produced an ancestral enzyme with an optimum temperature and pH that differed from modern family members, and the authors concluded that reconstructed ancestral enzymes can have improved physicochemical properties suitable for industrial applications. That conclusion is stated for that enzyme family and should be read as background rather than as a universal property of ancestral proteins. For diagnostics, the relevant question is narrower: does the ancestral scaffold retain sufficient activity under the thermal and pH conditions of the assay and its storage format? In adjacent workflows, enzyme engineering modification can support sample preparation and assay readouts without disrupting the core protocol.
Matrix tolerance is the second practical hurdle. Clinical samples are chemically complex, and enzymes that perform well in buffer may behave differently in serum, plasma, or other matrices. Inhibitor tolerance and matrix effects are therefore characterized early in development rather than deferred to late-stage validation. Because ancestral enzymes may differ from modern homologs in surface properties and stability, matrix behavior cannot be assumed from sequence alone; it must be measured. This is one reason why biochemical validation of catalytic activity is a required element of any resurrection project intended for diagnostic use.
Stability and matrix tolerance interact with formulation decisions. An enzyme that tolerates elevated temperature during processing may still be sensitive to freeze-thaw cycling or to the excipients used in a liquid format. Conversely, a scaffold with moderate solution stability may perform well when lyophilized. The practical approach is to treat stability as a set of measurable properties across a defined stress panel, then use those measurements to guide both engineering and formulation. This keeps the engineering objective tied to the actual reagent format rather than to an abstract stability target.
Thermostability screening
Resurrected enzymes are screened across a temperature range to identify scaffolds with useful thermal behavior and to compare them against extant homologs.
- Optimum temperature determination
- Activity retention across a thermal gradient
- Comparison with modern counterparts
pH and solvent tolerance
Ancestral enzymes have been reported to tolerate a range of metal ions and organic solvents, but tolerance profiles are family-specific and must be measured.
- pH optimum and pH stability range
- Metal ion effects on activity
- Solvent and additive tolerance
Clinical sample behavior
Assay-relevant matrices can suppress or alter enzyme activity, so matrix and inhibitor tolerance are characterized as part of development.
- Matrix effect screening
- Inhibitor tolerance assessment
- Assay-condition validation
Expression and Purification
Recombinant expression is the bridge between an inferred sequence and a physical reagent. Heterologous hosts are used to produce the ancestral protein, and the choice of host, vector, and cultivation conditions shapes yield, solubility, and post-translational characteristics. The lipase example illustrates that expression does not always yield soluble protein directly: the ancestral gene was expressed as inclusion bodies in E. coli, and recovery required refolding by urea dilution followed by affinity chromatography. Refolding is therefore a legitimate and sometimes necessary part of the workflow, not a sign that the project has failed.
Purification strategy follows from the expression outcome. Affinity chromatography is a common capture step, and subsequent polishing steps are selected to meet purity targets appropriate to the intended use. For diagnostic reagents, purity is not only a matter of removing host proteins; it also concerns consistency of the active species and the absence of contaminating activities that could generate background signal. Purity analysis and activity characterization therefore proceed together, because a preparation that is pure by one measure may still be heterogeneous in catalytic behavior. Practically, many labs complement this strategy with therapeutic enzymes enzyme to keep upstream reagents and downstream analytics aligned.
Scale-up introduces its own considerations. A process that produces sufficient material for characterization may not translate directly to larger batches, and parameters that were adequate at small scale may need re-examination. Technology transfer between development and production stages requires documentation of the process, the analytical methods, and the acceptance criteria. For ancestral enzymes, an additional consideration is that the reference sequence is itself an inferred construct, so sequence verification and batch-to-batch consistency take on particular importance: each batch should be traceable to the intended ancestral sequence and should meet defined activity and stability specifications.
Heterologous production
Ancestral genes are synthesized and expressed in heterologous hosts, with E. coli among the commonly used systems for initial evaluation.
- Codon optimization for the chosen host
- Soluble expression or inclusion body formation
- Refolding where required
Recovery and polishing
Capture and polishing steps are selected to deliver a preparation that meets purity and activity specifications for diagnostic use.
- Affinity capture
- Polishing chromatography
- Purity analysis
Process transfer
Process parameters, analytical methods, and acceptance criteria are documented so that production batches remain consistent with development material.
- Process documentation
- Analytical method transfer
- Batch consistency validation
Formulation and Shelf Life
Formulation determines whether a well-characterized ancestral enzyme becomes a practical reagent. Liquid formats require stabilizers and buffers that preserve activity over the intended shelf life, while lyophilized formats require a formulation that survives freezing, drying, and reconstitution. Glycerol-free and lyo-ready formats are often preferred for diagnostic reagents because they simplify handling and support ambient or reduced cold-chain distribution. The choice between formats is driven by the assay platform, the expected storage conditions, and the stability profile of the enzyme itself.
Shelf-life assessment is an empirical exercise. Accelerated stability studies, real-time stability studies, and stress testing are used to estimate how activity changes over time and under adverse conditions. Freeze-thaw and shipping stress testing addresses the handling that a reagent experiences between manufacture and use. Excipient, buffer, and stabilizer screening identifies formulation components that improve retention of activity. For ancestral enzymes, these studies are particularly informative because the scaffold may have a different stability profile from the modern enzymes with which developers are familiar, and assumptions carried over from those enzymes may not hold.
Point-of-care and near-patient testing impose additional constraints. Reagents may need to tolerate ambient temperatures, limited refrigeration, or extended transport. Cold-chain reduction strategies aim to make reagents more robust to these conditions, and lyophilization is a common route to ambient-stable formats. Instrument platform adaptation is a related consideration: the reagent must perform within the volume, timing, and detection constraints of the target analyzer. Formulation work therefore sits at the intersection of enzyme biochemistry and assay engineering, and it benefits from being planned alongside, rather than after, the engineering campaign. Where throughput or format constraints appear, enzyme engineering cdx is a natural capability to evaluate alongside the methods above.
Quality Control and Documentation
Quality control for a diagnostic enzyme begins with the definition of the molecule. For a resurrected enzyme, that definition includes the inferred ancestral sequence, the synthesized gene, and the expression construct. Sequence verification confirms that the produced protein matches the intended design. Activity assays confirm that the protein functions as expected, and stability assays confirm that it retains function under defined conditions. Together, these measurements form the analytical basis for release specifications and for comparability across batches.
Batch-to-batch consistency is a central requirement for diagnostic reagents because assay performance depends on reproducible enzyme behavior. Consistency is assessed through a combination of activity, purity, and stability measurements performed on each batch against predefined criteria. Where a process is transferred or scaled, comparability studies support the conclusion that material from the new process performs equivalently to material from the original process. For ancestral enzymes, the inferred nature of the reference sequence makes documentation of the design rationale and the reconstruction methodology especially important, since these records support the traceability of the reagent.
Regulatory and technical documentation supports the submission and lifecycle management of the reagent. Documentation typically covers the enzyme's identity, production process, characterization data, stability data, and quality control methods. Because ASR is a reconstruction methodology rather than a conventional isolation route, clear description of the phylogenetic analysis, the ancestral inference, and the validation of the expressed protein helps reviewers understand what the reagent is and how it was derived. This documentation is most effective when it is generated alongside development rather than assembled retrospectively.
FAQ
What is ancestral sequence reconstruction and how does it differ from directed evolution?
Ancestral sequence reconstruction uses phylogenetic analysis of extant homologous protein sequences to statistically infer ancient ancestral sequences, which are then resurrected by gene synthesis and recombinant expression. Directed evolution, by contrast, starts from an extant enzyme and introduces random or targeted mutations followed by screening. ASR infers an ancient sequence from phylogenetic data; it does not require iterative rounds of random mutagenesis as its core method. The two approaches can be combined, with ASR providing the starting scaffold and directed evolution used subsequently to tune performance.
How are resurrected enzymes expressed and purified?
The inferred ancestral sequence is synthesized as a gene and expressed in a heterologous host, with E. coli among the commonly used systems for initial evaluation. Expression may yield soluble protein or inclusion bodies; in the latter case, refolding followed by purification is required. Affinity chromatography is a common capture step, and additional polishing steps are selected to meet purity specifications. The purified ancestral enzyme is then characterized biochemically for activity, kinetics, and stability.
Why are ancestral enzymes considered useful starting points for diagnostic reagents?
Published work indicates that reconstructed ancestral enzymes can display improved physicochemical properties, and ASR has been described as a way to mine for thermostable enzymes that serve as superior starting points for protein engineering. In one reported case, an ancestral enzyme showed an improved evolvability profile relative to its modern counterpart, and engineering the ancestor accessed variants with enhanced thermostability and expression. These findings are background evidence; whether a specific ancestral scaffold suits a diagnostic application depends on biochemical characterization under assay-relevant conditions.
What quality control and documentation are needed for a resurrected diagnostic enzyme?
Quality control covers sequence verification, activity assays, stability assays, and purity analysis, with predefined acceptance criteria applied to each batch. Batch-to-batch consistency is assessed against those criteria, and comparability studies support process transfers or scale changes. Documentation describes the enzyme's identity, the reconstruction methodology, the production process, characterization data, and stability data. Because the reference sequence is inferred rather than isolated, records of the phylogenetic analysis and ancestral inference are particularly important for traceability.
References
- Albani Rocchetti G, Carta A, Mondoni A, et al. Selecting the best candidates for resurrecting extinct-in-the-wild plants from herbaria. Nature plants. 2022;8(12):1385-1393. View on PubMed
- Barkman TJ. Applications of ancestral sequence reconstruction for understanding the evolution of plant specialized metabolism. Philosophical transactions of the Royal Society of London. Series B, Biological sciences. 2024;379(1914):20230348. View on PubMed
From ancestral scaffold to diagnostic reagent
Ancestral sequence reconstruction opens a route to enzyme scaffolds that are not available from extant screening. Turning an inferred sequence into a reproducible diagnostic reagent requires phylogenetic analysis, gene synthesis, recombinant expression, purification, biochemical characterization, engineering, formulation, and quality control. Discuss your target enzyme family and assay format with our team to scope a development path from resurrection through reagent readiness.