When the known polymerase, reverse transcriptase, nuclease, ligase, helicase, protease, phosphatase, or reporter enzyme families cannot meet the required diagnostic profile, should the project search closer relatives, mine unexplored natural diversity, or create a new sequence?
Creative Enzymes provides sequence-to-experiment discovery support for new diagnostic enzyme scaffolds. A project can combine family and profile searches, metagenomic mining, sequence-similarity and structure-similarity analysis, protein language models, domain and motif checks, environmental and genomic context, diversity-aware candidate selection, gene and construct design, recombinant expression, biochemical screening, application-functional confirmation, and a traceable candidate dossier. Most outputs are intended for research use. Selected programs may support industrial diagnostic reagent raw-material development under an agreed scope; a computational candidate or discovery hit is not a finished test and does not establish regulatory authorization.
Enzyme discovery is not a database keyword search. The useful search space depends on what is already known, how far the required property lies from known enzymes, and which experiments can distinguish the desired molecular activity from a misleading surrogate. A narrow family search can be the most efficient route when the catalytic mechanism is trusted and the main gap is temperature, inhibitor tolerance, substrate use, expression, or licensing flexibility. Remote homolog and metagenomic mining becomes valuable when known families are too uniform or when an environmental prior suggests a different property regime. True de novo design is a different, higher-risk problem: a model must propose a new sequence or scaffold that supports a specified catalytic arrangement, then expression, folding, and activity must be established experimentally.
Begin from experimentally supported enzymes, trusted motifs, catalytic residues, domain architecture, and known positive and negative controls. Search homologs and orthologs, cluster the family, identify under-sampled clades, and choose candidates that cover useful evolutionary and property diversity.
Primary advantage: the mechanism and assay are often easier to interpret. Main risk: apparent diversity may still occupy a narrow functional range, and annotation can be transferred too broadly.
Best fit: a known enzyme class works, but available members do not satisfy the complete diagnostic target profile.Use profile methods, domain signatures, structure retrieval, protein embeddings, genome context, and environmental metadata to find candidates outside close sequence neighborhoods. Explore public or client-provided metagenomes, metatranscriptomes, genome catalogs, uncharacterized proteins, and predicted structures.
Primary advantage: access to biochemical diversity that is absent from common reference panels. Main risk: incomplete genes, uncertain annotation, difficult expression, missing context, and provenance gaps.
Best fit: the required property may be associated with a distinctive biome, lineage, architecture, or remote structural solution.Specify catalytic or binding geometry, required substrates and cofactors, forbidden activities, operating conditions, and a tractable experimental assay. Generative backbone, sequence, inverse-folding, motif-scaffolding, or reaction-conditioned methods can propose candidates after physical and sequence filters.
Primary advantage: exploration beyond observed natural sequences. Main risk: a plausible fold or model score does not prove catalytic turnover, specificity, expression, or diagnostic compatibility.
Best fit: no natural scaffold plausibly meets the requirement and the project accepts a research-stage, experiment-intensive route.
(Creative Enzymes Diagnostic)
The route can change after evidence is collected. A failed family search may reveal that a motif used as a query was too restrictive. A metagenomic panel may identify a weak but unusual hit that becomes an engineering seed. A de novo design attempt may clarify the active-site geometry but show that a natural scaffold search is more credible. We set decision gates before large synthesis or screening commitments so the project can expand, reroute, or stop without calling every untested sequence a candidate.
A discovery campaign begins with the molecular event, not the name of a product. “Find a better polymerase” is too broad. The brief must state whether the enzyme should extend DNA or RNA templates, displace strands, process a probe, ligate a nick, remove a modification, cleave a reporter after target activation, digest a matrix component, generate a chromogenic product, or protect another reagent. It should define the intended substrate and product, cofactors, temperature and pH window, reaction time, companion enzymes, formulation, forbidden side activities, and the application readout that will ultimately judge relevance.
Positive anchors are treated according to evidence quality. A sequence with a direct biochemical paper is not equivalent to a computational annotation copied through multiple database records. A commercial enzyme with a supplier activity claim can be a practical benchmark, but proprietary composition or sequence may limit its use as a search seed. A motif may define a catalytic superfamily without defining the exact substrate. We record why each anchor is trusted and which part of the target profile it supports.
| Discovery objective | Useful search anchors | Early computational filters | Primary experimental proof | Application confirmation |
|---|---|---|---|---|
| New polymerase or reverse transcriptase scaffold | Characterized family members, catalytic motifs, structural domains, template and primer requirements, known RT or strand-displacement activity | Complete catalytic architecture, exonuclease domains, insertions, predicted stability, expression liabilities, environmental temperature prior | Defined primer extension, template use, processivity or displacement, fidelity or side-activity studies as scoped | PCR, RT-PCR, LAMP, RCA, sequencing, or other locked reaction with appropriate controls |
| New nuclease, nickase, or CRISPR effector | Active-site architecture, guide or recognition elements, target grammar, known cis/trans behavior, genomic context | Domain arrangement, catalytic residues, guide-associated loci, compactness, nuclease contamination risk, structure similarity | Target binding or cleavage, product pattern, reporter cleavage, guide dependence, no-target and non-target controls | Defined amplification product, reporter, matrix, temperature, and reaction sequence |
| New ligase or end-processing enzyme | Substrate termini, cofactor use, family motifs, known nick or end preferences, structural class | Catalytic domain completeness, cofactor-binding sites, accessory domains, cellular context, likely oligomeric state | Direct ligation or end-conversion products with substrate and cofactor series | Library preparation, probe ligation, circularization, repair, or adapter workflow |
| New reporter or signal enzyme | Reaction chemistry, chromogenic or luminescent substrate, cofactor, known active-site geometry, negative cross-reactivities | Substrate-pocket features, secretion or disulfide requirements, oligomerization, cofactor dependence, background reaction risk | Product formation, substrate profile, background, kinetics, thermal and buffer behavior | Signal generation with the intended conjugate, sensor, calibrator, matrix, or detection instrument |
| New sample-processing enzyme | Target matrix component, cleavage or conversion chemistry, known enzyme classes, compatible sample conditions | Signal peptides, membrane association, proteolysis, pH and temperature prior, off-target substrate risk | Direct matrix-substrate conversion plus effects on representative target molecules | Extraction yield, inhibitor reduction, target recovery, and downstream assay compatibility |
The candidate universe can include curated proteins, broad public protein records, whole genomes, metagenomic assemblies, predicted proteins, predicted structures, client-owned sequences, and newly generated designs. These sources do not have the same error profile. Curated references provide stronger functional anchors but limited diversity. Unreviewed records provide breadth but may propagate incorrect annotations. Metagenomic contigs can expose uncultured diversity yet contain partial genes, assembly errors, uncertain start sites, or sparse metadata. Generated sequences may be structurally plausible but have no natural provenance or functional history.
| Sequence space | What it contributes | Key uncertainty | Required trace |
|---|---|---|---|
| Curated reference space | Trusted anchorsReviewed functions, experimental literature, structures, catalytic residues, assay methods, and known biochemical boundaries. | Coverage biasWell-studied families and organisms may dominate; absence from the curated set is not evidence of absence. | Accession and evidenceDatabase, release, record version, publication, experimental annotation, and retrieval date. |
| Broad genome/protein space | Family breadthOrthologs, paralogs, predicted proteins, taxonomic range, domain combinations, and uncharacterized branches. | Annotation propagationNames can be assigned by similarity without confirmation of exact substrate, activity, or domain context. | Record lineageAccession, source database, organism, genome or proteome, annotation type, and transformations. |
| Metagenomic space | Uncultured diversityEnvironmental sequence variation, novel loci, distinct temperature or chemistry priors, and remote architectures. | Assembly and gene qualityPartial ORFs, frameshifts, chimeric contigs, uncertain taxonomy, and inconsistent sample metadata. | Study and sampleContig, coordinates, biome, study and sample accessions, assembly, gene-calling method, and available collection metadata. |
| Client-owned space | Project relevanceHistorical hits, failed constructs, proprietary screens, internal libraries, and exact application labels. | Dataset compatibilityLabels, assay methods, sequence versions, construct boundaries, and failure codes may be inconsistent. | Ownership and methodsSource file, permissions, client identifier, assay provenance, confidentiality, and allowed project use. |
| Generated sequence space | Beyond-natural proposalsBackbones, sequences, motif scaffolds, active-site arrangements, or recombined hypotheses proposed by a model. | Function uncertaintyFold confidence, sequence naturalness, or model likelihood does not establish catalysis or useful expression. | Generation manifestModel or method, version, conditioning inputs, seed where applicable, filters, parent or motif relationship, and selection reason. |

(Creative Enzymes Diagnostic)
Resources such as UniProtKB, NCBI RefSeq and GenBank, InterPro, MGnify, BRENDA, structural databases, and family-specific repositories can contribute to a project. Their use is documented by database name, release or access date, query, accession, and filtering decisions. If a client supplies a private database or metagenome, the agreed data-governance and confidentiality requirements are applied to the discovery manifest. We can also record data-source licenses, notices, collection information, and access-and-benefit-sharing metadata that are available to the technical team.
No single search method captures catalytic function, distant homology, structure, environmental adaptation, and manufacturability. Close sequence similarity can retrieve reliable family members but miss remote solutions. A conserved catalytic motif can be present in proteins with different substrates. A structure match can identify a related fold without proving the same chemistry. Protein language-model embeddings can organize remote sequence relationships but may not resolve subtle activity differences within a family. Environmental temperature is a useful prior, not a thermostability measurement. We therefore combine passes and retain the evidence contributed by each.
Large-scale tools such as MMseqs2 can accelerate protein sequence search and clustering, while Foldseek supports fast structural comparison. Profile methods and InterPro signatures can detect family, domain, and site relationships. Sequence similarity networks can reveal clusters and boundary regions that a simple ranked hit list hides. Protein language models and structure-aware representations can add remote candidates, especially when sequence identity is weak. The analysis retains method-specific scores rather than collapsing every signal into a single unexplained rank.
Studies of enzyme superfamilies have documented serious functional misannotation outside highly curated records. A name transferred from the nearest database hit can therefore be wrong at the exact substrate or reaction level. Candidates are checked against domain architecture, catalytic residues, coverage, known negative functions, context, structure, and experimental controls. The final functional label comes from the measured assay, not from the FASTA header.
Biomes can focus a search. Hot springs, compost, hydrothermal systems, hypersaline environments, cold ecosystems, acidic sites, or host-associated microbiomes may enrich for proteins adapted to distinctive conditions. This logic has yielded diagnostic-relevant examples: a thermostable viral metagenome-derived polymerase was experimentally used in RT-PCR, and Cas12a orthologs mined from warm-environment metagenomes showed elevated-temperature target and trans-nuclease activity. These examples justify environmental stratification, but they do not support assuming that every sequence from a hot sample is thermostable or every viral polymerase performs reverse transcription.
Generative design can be conditioned on a known fold, active-site motif, substrate, transition-state geometry, metal coordination, binding interface, or a combination of sequence, structure, and function. Research has produced de novo luciferases and, more recently, designed hydrolases in selected systems. Yet catalysis involves more than a static pocket: protonation, solvent, dynamics, multiple reaction states, substrate entry and product exit, oligomerization, cofactors, and competing reactions can matter. A de novo branch should therefore have a defined catalytic hypothesis, a feasible direct assay, an adequate negative-control system, and acceptance of a lower-evidence starting point.
Generated candidates pass physical and sequence filters, predicted-fold checks, active-site geometry checks, similarity and novelty review, and synthesis constraints. Known natural enzymes and deliberately damaged active-site controls remain in the physical screen. When the catalytic problem is not sufficiently specified or the assay cannot distinguish weak true activity from contamination, the responsible decision may be to improve the assay or return to natural mining rather than generate more sequences.
A ranked spreadsheet can hide why a candidate exists. We instead assemble a candidate evidence passport that connects the sequence to its discovery route, source, evidence, risk, construct, and planned experiment. The passport makes portfolio review possible across bioinformatics, protein science, assay development, sourcing, and legal or compliance teams. It also prevents the candidate identity from changing silently when a database record, gene model, tag, or construct boundary is updated.
The passport travels with the candidate from in silico selection to gene order, expression, assay plate, data file, and transfer report.

(Creative Enzymes Diagnostic)
The passport also distinguishes “unknown” from “failed.” An untested candidate has no functional label. A sequence that could not be synthesized, a construct that did not express, an insoluble material, a purified protein below the assay limit, a contaminated sample, and a true inactive enzyme are different outcomes. These labels matter if the data will later support active learning or the AI-ready experimental dataset and screening data analysis service.
A common mining failure is to sort by one score and synthesize the top candidates. Protein databases contain many near duplicates, and model scores can favor the most familiar family region. Ten highly ranked sequences may therefore be ten versions of the same hypothesis. If that hypothesis fails because the family lacks the required property, the entire panel fails together. A discovery portfolio should cover evidence strength, sequence and structure diversity, environmental priors, construct risk, and uncertainty while retaining controls that make the screen interpretable.

(Creative Enzymes Diagnostic)
Panel size is project-specific. It depends on the number of credible clusters, screen throughput, gene and construct cost, expected expression difficulty, number of property axes, assay precision, and whether the first panel is intended to find a direct lead or learn where to search next. We do not promise a fixed hit rate. Instead, each portfolio has coverage metrics and a reason for inclusion. Redundant candidates can be retained when they test meaningful changes such as domain boundaries, environmental source, cofactor motif, or predicted active-site geometry; accidental redundancy is removed.
| Portfolio decision | Question answered | Useful inclusion rule | Failure prevented |
|---|---|---|---|
| Retain a characterized positive control | Can the assay detect the expected molecular function under the screen conditions? | Use a material with direct evidence and a compatible substrate, even if it lacks the desired final property | Calling the entire panel inactive when the screen or substrate is faulty |
| Retain negative-family or catalytic controls | Does the readout distinguish the intended chemistry from contamination or a related side reaction? | Choose a close relative with known wrong function or a justified catalytic-site disruption | Promoting nonspecific signal as a discovery hit |
| Cap near-duplicate clusters | Is synthesis capacity covering distinct hypotheses? | Select cluster representatives by evidence, completeness, construct risk, and property prior | Spending most of the panel on one overrepresented lineage |
| Reserve remote candidates | Can structure, context, or embedding evidence identify function beyond close sequence neighbors? | Require at least two complementary evidence channels and an intact catalytic hypothesis | Allowing one opaque model score to dominate a high-risk selection |
| Reserve information-gain candidates | Which uncertain region should the next search expand or abandon? | Choose candidates that discriminate between competing family, motif, context, or environmental hypotheses | Producing only confirmation data from familiar sequences |
| Apply synthesis and provenance gates | Can the exact sequence be built, traced, and evaluated under the intended project terms? | Require a stable sequence, source record, construct plan, unresolved-risk label, and client-approved exclusions | Discovering late that a hit is partial, untraceable, or outside the agreed source scope |
Computational discovery ends where functional evidence begins. Gene synthesis and expression are not administrative steps: a candidate can fail because the ORF is partial, the start site is wrong, a native signal peptide or membrane segment was retained, a domain boundary is missing, the selected host cannot support folding or cofactors, or the tag interferes with function. Construct alternatives may be more informative than ordering many additional sequences. Expression, solubility, purification, aggregation, and active fraction therefore become part of discovery evidence.
A primary screen is designed around the catalytic event rather than convenience alone. Fluorescence can provide throughput, but a fluorogenic signal may be influenced by contaminants, optical interference, substrate instability, or unintended cleavage. Orthogonal confirmation may use electrophoresis, chromatography, mass spectrometry, product sequencing, direct absorbance, binding analysis, or a second substrate format. The exact method is chosen from the reaction and decision, not from a universal discovery panel.
Material normalization is essential when comparing diverse natural or generated sequences. Equal culture volume, total protein, purified mass, and active enzyme concentration are not equivalent. A candidate with weak soluble expression can appear inactive even when its intrinsic enzyme is useful; a contaminated preparation can appear highly active. The project defines whether the first screen ranks crude lysate, soluble fraction, normalized purified material, or active fraction, and which conclusion is permitted from that stage.
A polymerase hit may extend a model primer but fail in a complex amplicon. A reverse transcriptase may accept a short RNA template yet stall on structured targets. A nuclease may cleave a purified substrate but show high reporter background or poor mismatch discrimination. A ligase may join a nick but reject the adapter chemistry. A matrix-processing enzyme may remove an inhibitor while damaging the analyte. Application-functional testing is therefore separate from the primary molecular assay.
The confirmation system uses representative targets, substrates, inhibitors, matrices, partner enzymes, temperature history, reporter chemistry, and storage state within scope. For complete reaction-system development, discoveries can connect to molecular diagnostic enzyme and master mix development, CRISPR diagnostic assay development, NGS library preparation enzyme system development, or nucleic acid extraction enzyme system optimization.
Discovery should end with a decision, not a pile of sequences. A candidate can be a direct biochemical lead, a promising but weak engineering scaffold, a family-expansion seed, a production-risk candidate, a de novo learning result, or a stopped route. The promotion criteria are defined in the discovery brief and applied to confirmed material. Rankings can change after purification, orthogonal confirmation, application testing, independent expression, or formulation challenge.

(Creative Enzymes Diagnostic)
A discovery hit that needs production improvement can enter our AI-guided expression, solubility, and manufacturability optimization service. A hit with a confirmed molecular function but insufficient catalytic performance can enter activity and kinetic performance optimization. Projects that require deeper structure and substrate interpretation can use structural modeling and enzyme-substrate interaction analysis. This routing prevents discovery from absorbing every later development problem.
Project size depends on the depth and quality of starting evidence, number of plausible enzyme classes, database and metagenome scope, sequence redundancy, search novelty, physical screen throughput, expression difficulty, required property axes, and whether the goal is a direct lead or an informative first map. A small, carefully diversified panel can be more valuable than a large redundant panel. Conversely, a broad family with weak annotations may require more sampling before a negative conclusion is justified.
Mining searches natural, metagenomic, client-owned, or predicted sequence space for a new scaffold. Engineering changes an existing scaffold through mutations, recombination, domain changes, or iterative evolution. Mining is appropriate when the current starting enzyme or candidate pool is inadequate. A mined hit often becomes the parent for a later engineering program.
It can refer to discovering an enzyme that is new to the application from uncharacterized natural sequence space, or to generating a sequence or scaffold not observed in nature. We distinguish these routes explicitly. True generative de novo design is feasibility-gated and higher risk because predicted fold or active-site geometry does not prove catalytic activity.
Projects can be considered for polymerases, reverse transcriptases, strand-displacement enzymes, nucleases, nickases, CRISPR effectors, helicases, ligases, end-processing enzymes, proteases, phosphatases, glycosidases, reporter enzymes, matrix-processing enzymes, and other defined functions. Feasibility depends on a credible molecular hypothesis, accessible sequence space, buildable candidates, and an assay that can prove the intended activity.
Yes, subject to data access, metadata, provenance, quality, and project scope. Biome, temperature, pH, salinity, host association, viral origin, or other environmental information can focus a search. These fields are treated as priors; the required property must still be measured experimentally.
They are useful evidence but should not be treated uniformly. Reviewed records with direct experiments carry more weight than automatically transferred labels. Exact substrate or reaction annotations can be wrong even when family assignment is correct. We retain annotation provenance and cross-check sequence coverage, catalytic residues, domain architecture, structure, context, negative families, and experimental data.
A project can combine profile searches, conserved-domain and site signatures, sequence similarity networks, predicted or experimental structure search, protein language-model embeddings, genome context, and environmental metadata. Remote candidates normally require more than one supporting evidence channel and stronger experimental controls because the functional inference is less direct.
No. A predicted fold can support domain completeness, architecture, active-site geometry, or structural relationship. It does not establish catalytic turnover, substrate specificity, cofactor use, expression, active fraction, side activities, stability, or diagnostic compatibility. These properties require experiments.
There is no universal number. It depends on credible cluster count, diversity, screen throughput, expression risk, candidate cost, number of property objectives, strength of prior evidence, and whether the goal is a direct hit or information for a second search. We propose a coverage-balanced panel and identify the controls and hypotheses represented by each candidate.
Yes, under agreed confidentiality, access, data-governance, and permitted-use terms. We can preserve client identifiers, source files, sequence versions, transformations, and access restrictions in the project manifest. The relationship between private data, public evidence, generated candidates, and experimental results remains traceable.
Ordinary technical discovery work does not constitute a legal opinion. We can record sequence sources, accessions, database notices, sample and biome metadata, collection information when available, client exclusions, and unresolved provenance. Patentability, freedom to operate, and access-and-benefit-sharing obligations should be reviewed by qualified legal or compliance specialists for the intended jurisdiction and use.
Non-expression is not automatically evidence of no function. We review gene and construct integrity, start site, boundaries, tags, signal peptides, transmembrane regions, codon design, host, temperature, solubility, cofactors, and purification behavior. The candidate can be redesigned, transferred to production optimization, retained as an unresolved sequence, or stopped according to the evidence and project value.
The primary assay includes positive, negative, blank, substrate, cofactor, and material controls. Orthogonal confirmation can establish product identity, cleavage pattern, substrate dependence, catalytic-residue dependence, or another independent signal. Contamination, optical interference, spontaneous substrate change, and nonspecific activity are investigated before promotion.
Yes, representative application-functional confirmation can be included. The upstream reaction, target, substrate, reporter, partner enzymes, matrix, instrument, and analysis rule are locked sufficiently to protect interpretation. Full formulation and assay development are scoped separately when those system variables also need optimization.
A well-designed negative campaign can still deliver the discovery brief, source manifest, searched space, rejected and tested candidates, construct and material results, assay methods, raw data, failure labels, family or environment boundaries, and recommended re-mining or stop decision. We do not relabel a negative result as success, but the evidence can prevent repetition and improve the next search.
No. A discovery hit has evidence only for the tested sequence, construct, material, assay, and conditions. Manufacturing readiness requires process development, quality attributes, analytical methods, lot comparability, stability, formulation, supply controls, and application validation. Finished diagnostic products require additional design control, analytical and clinical validation, quality systems, and regulatory work by the responsible manufacturer or sponsor.
Send us the required molecular activity, known positive and negative enzymes, diagnostic application, current performance gap, sequence or metagenome resources, experimental throughput, source constraints, and the evidence needed for your next decision. Creative Enzymes can propose a staged route from a traceable sequence universe to a diversity-balanced candidate portfolio, direct biochemical proof, diagnostic-context confirmation, and a defined hit-promotion plan.
Contact Creative Enzymes