Creative Enzymes helps teams define, structure, quality-control, analyze, and hand off diagnostic-enzyme screening datasets so that a result can be traced from sequence and sample to plate, raw file, processed endpoint, engineering decision, and future modeling use.
A spreadsheet can contain thousands of measurements and still fail as an engineering dataset. If a sequence cannot be connected to the construct actually built, a purified sample cannot be connected to its expression batch, a well cannot be connected to the correct plate map, or a normalized score cannot be recreated from raw measurements, the apparent volume of data overstates its usable evidence.
This service addresses the experimental and analytical data layer of diagnostic-enzyme development. We begin with the decision the dataset must support: selecting candidates for confirmation, comparing enzyme variants across days or lots, learning which sequence changes influence an assay endpoint, balancing multiple POCT-relevant properties, or preparing a dataset for later machine-learning work. The decision determines which entities need stable identifiers, which controls and replicates are necessary, which nuisance variables must be exposed, which transformations are legitimate, and which type of holdout test is credible.
Define the entity model, factor levels, sample and plate identifiers, controls, replicate logic, randomization or blocking, endpoints, QC rules, and analysis plan.
Capture execution metadata, raw exports, instrument and reagent context, deviations, failed wells, reruns, and the exact relationship between sample, plate, well, and file.
Validate identity, preserve raw evidence, version transformations, diagnose assay quality, analyze outcomes, confirm hits, and package a leakage-aware dataset with limitations.
Sequence names such as “mutant 12” or plate-local labels such as “A7” are convenient during an experiment but unsafe as global identifiers. The same well coordinate recurs on every plate, a sample can be diluted into several assay runs, a construct can produce more than one expression batch, and a candidate sequence can differ from the expression construct through tags, signal peptides, cloning scars, or processing. We create an identity and lineage plan that keeps these levels separate while connecting them with explicit keys.
A lineage record should answer practical questions without guesswork. Which amino-acid sequence was expressed? Was the tested material soluble fraction, purified enzyme, or lysate? Which lot of substrate and cofactor was used? Was the well part of the planned map or a manual substitution? Which raw channel and time point produced the reported endpoint? Which software or formula version transformed the value? Was the sample excluded, and if so, was the cause biological, analytical, or logistical?

(Creative Enzymes)
An assay data contract is a project-specific specification that connects biological intent, experimental execution, data capture, quality decisions, and downstream analysis. It prevents important meanings from being reconstructed after the screen, when the team may no longer remember why a blank cell meant “not tested,” “below detection,” “failed dispense,” or “file unavailable.” The contract is scaled to the project: it can be a concise table for a pilot screen or a more formal collection of schemas, dictionaries, and validation rules for a multi-round program.

(Creative Enzymes)
A label is not simply the column selected as a model target. “Activity,” “stability,” and “specificity” can represent very different measurements depending on substrate, temperature, incubation time, matrix, reaction window, instrument response, normalization reference, and calculation. We define labels at the level needed for comparison and avoid merging endpoints that only share a familiar name. When categorical labels such as active/inactive or pass/fail are required, the underlying continuous values, thresholds, and quality state should remain available whenever feasible.
Technical failures and biological negatives are also separated. A variant that expresses but has no detectable activity under a valid assay is evidence about sequence–function behavior. A well affected by failed dispensing, saturation, contamination, plate-map mismatch, or missing raw data does not carry the same label. Likewise, “not tested” must not be encoded as zero. Reason-coded missingness protects downstream analysis from learning logistics as biology.
Screening throughput does not automatically create independent evidence. When all high-priority variants occupy one plate, all controls are located at one edge, parent enzymes are tested on a different day from mutants, or a library branch is handled by one operator, variant identity becomes entangled with nuisance variables. We review the intended comparison and design a plate and run structure that makes relevant sources of variation estimable within the available capacity.
Positive, negative, blank, matrix, calibration, process, and reference-enzyme controls are selected for the measurement and decision. Control placement should support spatial and temporal diagnostics, not only calculation of a single summary statistic.
Technical replicates, independent preparations, biological batches, repeat plates, and confirmation experiments answer different questions. The plan identifies the experimental unit and prevents a large number of wells from being mistaken for independent samples.
Variants may be randomized within meaningful blocks while parent, reference, and bridge samples span days, lots, instruments, or plates. Practical constraints are recorded so the analysis can model rather than ignore them.

(Creative Enzymes)
For buffer, cofactor, enzyme concentration, temperature, incubation, substrate, formulation, or stress-factor studies, we can structure a design-of-experiments table that distinguishes target factors from nuisance factors and preserves the actual executed settings. Coded factor levels alone are insufficient for reuse; the package should include physical values, units, preparation details, interaction terms under consideration, constraints, and deviations. When the experimental space is sequential, the design records why each condition was added and which previous evidence informed it.
For multiparameter enzyme optimization, the data plan retains separate endpoints before applying a composite score. A single weighted score can be useful for a defined decision, but it hides tradeoffs and becomes misleading if weights change. We therefore preserve component endpoints, uncertainty, hard constraints, and Pareto relationships so that a later multiparameter POCT enzyme optimization decision can be recalculated without rerunning the experiment.
Instrument exports are preserved as received and associated with a file-level identity, acquisition context, and, where practical, a checksum. Corrections do not overwrite the source. Instead, each transformation creates a versioned layer with named inputs, parameters, formulas or code, output columns, and a record of the analyst or process that produced it. This approach supports both scientific review and efficient reanalysis when an endpoint definition changes.
| Data layer | Typical content | Control principle | Decision use |
|---|---|---|---|
| Raw | Instrument signals, timestamps, kinetic traces, spectra, images, or exported well values | Immutable; retain acquisition metadata and original units | Reconstruction, artifact review, and alternative processing |
| Context | Plate map, sample identity, protocol, reagent lots, operator, instrument, deviations | Versioned keys; preserve executed rather than only planned state | Join biology to execution and expose nuisance variables |
| Processed | Blank correction, calibration, kinetic-window estimates, normalized well values | Named transformation, parameters, reference controls, and version | Comparable well-level measurement with traceable assumptions |
| Aggregate | Replicate mean or median, variability, confidence interval, sample-level endpoint | Explicit experimental unit, replicate type, and aggregation rule | Ranking, comparison, and uncertainty assessment |
| Decision | QC state, threshold result, candidate rank, confirm/retest/hold/reject action | Separate evidence from recommendation; timestamp and version the rule | Program action and audit trail |
Normalization defines what the raw signal means relative to controls, standards, baselines, or reference enzymes. It should be selected for the assay’s response mechanism and applied consistently within a defined scope. Plate-local normalization may help account for plate-to-plate signal range, but it can also conceal a genuine systematic change if controls themselves drift. Global scaling can preserve between-plate differences but amplify run effects. Spatial correction can reduce position-related artifacts, yet an aggressive method can remove true biology when variant placement is confounded with position.
We therefore retain the raw value, reference values, processed value, transformation version, and quality state together. Alternative normalization methods can be compared against control behavior and expected biology. Any method change is evaluated on representative data; it is not silently applied to the historical dataset. The NCATS Assay Guidance Manual’s distinction among raw, normalized well-level, aggregate, and derived results provides a useful conceptual framework, while the exact processing rules remain assay-specific.
Signals below a detection or quantification limit, saturated readings, incomplete kinetic windows, failed curve fits, and observations outside a calibration range should not all become the same number. We can define qualifiers and reason codes such as below range, above range, non-estimable, technically failed, excluded by predeclared QC, not collected, or not applicable. The exported numeric field can then be interpreted together with the qualifier rather than forcing a potentially false value into analysis.
No single statistic can establish that an enzyme screen is decision-ready. A plate may show acceptable separation between control groups while containing an edge pattern, a dispensing gradient, a swapped quadrant, a deteriorating reagent, or a cluster of failed samples. Conversely, a strict generic threshold may reject useful data from an assay whose intended comparison is supported by paired references and confirmation testing. We select diagnostics and acceptance logic around the assay, stage, and consequences of error.
Range, saturation, time trace, dispense status, image or curve quality, contamination, duplicate identity, and reason-coded flags.
Control locations, dynamic range, variability, heat maps, row/column effects, edge patterns, drift, reference consistency, and layout integrity.
Instrument, operator, reagent lot, timing, bridge samples, batch effects, replicate concordance, and protocol deviations.
Round and library composition, missing groups, distribution shift, rerun policy, confirmation rate, and comparability across datasets.
Metrics may include signal window, coefficient of variation, robust dispersion, Z′ or related separation statistics, replicate correlation or concordance, calibration diagnostics, curve-fit uncertainty, and control-chart behavior. The Z′ factor introduced by Zhang and colleagues is widely used for assay evaluation, but a value should be interpreted with its control design, plate layout, and intended application. We do not impose a universal numerical cutoff without understanding the assay and decision.
The appropriate analysis depends on whether the screen is intended to detect a difference, rank candidates, estimate a kinetic parameter, identify a threshold-crossing variant, map a sequence–function landscape, or balance several endpoints. We document the estimand—the quantity the analysis is intended to learn—and keep it aligned with the experimental unit. A result derived from repeated wells on one enzyme preparation is not automatically evidence of batch reproducibility, and a high primary-screen value is not yet a confirmed lead.
Apply predeclared QC, review plate and run diagnostics, resolve identity problems, document deviations, and determine which observations can support the planned comparison.
Calculate assay-appropriate endpoints, replicate summaries, uncertainty, effects relative to references, rank stability, and sensitivity to defensible processing choices.
Advance, confirm, retest, hold, or reject using transparent rules that include capacity, risk, diversity, multiple objectives, and the cost of false positive or false negative decisions.
Hit thresholds can be fixed in advance, estimated from reference behavior, defined by robust distributions, or implemented as ranked capacity limits. Each strategy makes different assumptions. Where many hypotheses are formally tested, multiplicity and false-discovery considerations may be relevant. Where the purpose is candidate triage rather than population inference, effect size, uncertainty, confirmation capacity, and diversity may be more useful than a p-value alone. We choose and explain the analysis rather than applying a generic “top 5%” rule.
Confirmation experiments should be designed around plausible failure modes. These can include fresh sample preparation, independent expression or purification batches, repeated concentration series, alternative substrate levels, an orthogonal readout, interference controls, matrix testing, or an application-proximal assay. Creative Enzymes can connect data analysis with enzyme activity and stability analysis, assay interference and matrix-effect evaluation, and batch-to-batch consistency assessment when those studies are within the agreed project scope.

(Creative Enzymes)
A diagnostic enzyme may need adequate catalytic activity, low background, tolerance to inhibitors, storage stability, expression yield, compatibility with a dry format, and performance in a target matrix. We can analyze each endpoint, identify hard constraints, display correlations and tradeoffs, and construct decision views such as Pareto fronts or scenario-specific rankings. A composite desirability score is versioned with its scaling and weights. The raw component endpoints remain available so a different product format or customer priority can be evaluated later.
Randomly dividing rows into training and test sets is often inappropriate for protein-engineering data. Near-identical sequences can occur on both sides of a split. Variants derived from the same parent, plate, expression batch, or experimental round can share signals that will not be available for a new protein family or future campaign. Replicate wells can be separated while still representing the same test article. Preprocessing performed on the entire dataset can pass information from the held-out set into the model. These routes create optimistic validation without improving real-world decisions.

(Creative Enzymes)
| Intended use | Possible evaluation design | Risk a naive row split may hide |
|---|---|---|
| Predict untested combinations near one parent | Hold out variants or combinatorial regions while controlling sequence similarity and replicate identity | Near-duplicate variants make interpolation appear easier than it will be in the intended region |
| Select the next experimental round | Round- or time-based holdout that emulates training on past rounds and predicting a later round | Future-round labels or assay adjustments leak backward |
| Transfer to a new parent or enzyme family | Parent-, family-, or sequence-cluster holdout | Shared backgrounds dominate performance while distant generalization remains unknown |
| Transfer across production or assay conditions | Batch-, lot-, instrument-, site-, or campaign-level holdout | The model learns operational fingerprints instead of biological performance |
| Rank confirmed candidates | Locked confirmation set with an analysis plan established before unblinding | Repeated threshold and model tuning converts the test set into training feedback |
We can assess label distribution, sequence representation, missingness, group sizes, endpoint reliability, covariate balance, duplicate risk, and possible shift between train, validation, and test partitions. If the dataset cannot support a credible split, the correct output may be a gap assessment and a proposed collection plan rather than a misleading performance estimate. Informed training-set design has been shown to matter in machine-learning-assisted directed evolution, but the best strategy depends on the protein, landscape, assay, and decision.
Reusable does not mean public. Confidential data can follow FAIR-inspired principles—findable within the agreed environment, accessible to authorized users under defined conditions, interoperable through explicit formats and terminology, and reusable because provenance and context are sufficient. We adapt the level of formalization to client needs and can draw on concepts from BioAssay Ontology, the Ontology for Biomedical Investigations, or ISA-Tab where they improve clarity and exchange. We do not force every project into a standard that adds complexity without decision value.
The handoff can be designed for a human review workflow, a database import, or a computational pipeline. Common tables may be delivered in CSV or spreadsheet-compatible form, with machine-readable formats such as JSON where useful. Naming, column order, decimal precision, time and unit conventions, controlled terms, and null representation are documented. If the client has an existing laboratory information management system or analysis environment, we map the package to its import constraints during project scoping.
We define the scientific and operational decision, future deployment case, available evidence, current data flow, file formats, identifiers, assay maturity, and known pain points. Existing raw exports and plate maps are reviewed for reconstructability.
We specify entities, factors, units, controls, replicates, plate/run structure, metadata capture, endpoints, QC states, analysis rules, and candidate actions. For existing datasets, this becomes a remediation and mapping plan.
Files are inventoried; schemas, types, identifiers, joins, duplicates, missing fields, plate maps, and sequence/sample links are checked. Ambiguous mappings are raised rather than silently guessed.
Raw evidence is preserved, transformations are versioned, and well/plate/run/sample/campaign diagnostics are applied. Exclusions and reruns receive explicit reason codes and dispositions.
Endpoints, uncertainty, effects, ranking, thresholds, multivariate tradeoffs, and sensitivity are examined. Candidates are assigned transparent advance, confirm, retest, hold, or reject actions.
When relevant, we define a deployment-aligned split, assess leakage and representation, produce manifests and dictionaries, document limitations, and identify the next data that would reduce uncertainty.
| Deliverable | What it contains | When it is most useful |
|---|---|---|
| Experimental data design memo | Decision, experimental unit, factor structure, controls, replicates, plate/run allocation, endpoints, and planned analysis | Before a pilot, screen, DOE, or new engineering round |
| Assay data contract and dictionary | Entities, fields, units, allowed values, identifiers, QC states, missingness, transformations, and version rules | For consistent collection across scientists, sites, or rounds |
| Plate map and metadata templates | Planned and executed layouts, sample keys, controls, bridge samples, run context, and deviation capture | For plate-based enzyme variant screens |
| Cleaned and traceable dataset | Linked raw/context/processed/aggregate tables with preserved source references and reason-coded changes | When legacy files or campaign exports require reconstruction |
| QC and screening-analysis report | Diagnostics, quality dispositions, endpoint results, candidate ranking, uncertainty, sensitivity, and confirmation recommendations | For primary, confirmation, or multiparameter screens |
| AI-readiness and split package | Representativeness assessment, leakage audit, group definitions, split manifest, locked-test policy, and reuse limitations | Before training or benchmarking a predictive model |
| Data-gap and next-experiment plan | Missing regions, confounded factors, label weaknesses, high-value controls, and prioritized collection options | When the existing dataset cannot yet answer the intended question |
An initial audit may find that sequence-to-sample mapping is incomplete, raw files are unavailable, plate maps were overwritten, assay versions were pooled, controls do not support the desired normalization, all candidates are confounded with a batch, or the label distribution is too narrow for the proposed model. We report these conditions directly. Depending on the decision, the most defensible next step may be to reconstruct lineage, analyze only a qualified subset, run a bridging experiment, repeat selected controls, collect a deliberately informative pilot set, or narrow the claim.
We do not guarantee that a dataset will yield a predictive model, that a particular algorithm will outperform conventional selection, that a primary hit will confirm, or that an improved enzyme will meet application requirements. Our role is to make the evidence, assumptions, transformations, uncertainty, and decision logic visible enough for the next action to be scientifically testable.
This service can stand alone for an existing dataset or provide the data foundation for a broader AI-driven diagnostic enzyme engineering program. It is deliberately distinct from adjacent services: it defines and analyzes the evidence asset, while other pages focus on candidate design, molecular interpretation, or iterative program execution.
Use AI-assisted mutation library design when the primary need is to choose positions, substitutions, combinations, and a buildable sequence set. The present service defines how those candidates, test articles, and outcomes will be represented and analyzed.
Use structural modeling and enzyme–substrate interaction analysis when the question concerns binding geometry, catalytic contacts, conformational hypotheses, or testable structural explanations.
Use the closed-loop DBTL enzyme evolution service when multiple design–build–test–learn rounds must be governed as one program. This page supplies the data contract, QC, analysis, and reuse layer within or outside that loop.
Where physical test articles are needed, projects may connect to enzyme expression and purification and other agreed experimental services. Scope, materials, assays, data volume, and handoff format are defined before work begins.
Candidate and property-focused options include variant design and screening, thermostability and lyophilization-stability engineering, activity and kinetic optimization, specificity and cross-reactivity reduction, and expression, solubility, and manufacturability optimization. Modality-focused options include polymerase and reverse-transcriptase engineering, LAMP, RPA, and isothermal-enzyme optimization, and CRISPR/Cas diagnostic enzyme engineering. Discovery and comparability options include de novo enzyme discovery and mining and second-source and sequence-equivalency engineering.
Often we can improve traceability, standardize fields, reconstruct plate/sample relationships, separate raw and derived layers, define missingness and QC states, and assess leakage or representation. The achievable result depends on the availability of source files, sequence/construct mappings, plate maps, protocols, and assay-version information. Ambiguous lineage will be documented rather than inferred as fact.
No. The same data architecture improves ordinary screening decisions, reproducible normalization, candidate ranking, confirmation planning, cross-run comparison, and future reuse. A data audit may conclude that a transparent statistical analysis is more appropriate than machine learning for the current decision.
Scope can include plate-reader exports, kinetic traces, endpoint tables, curve data, sample and plate maps, sequence tables, process metadata, and client-generated summaries. Compatibility is reviewed from example files before the project is finalized. Complex proprietary formats may require an agreed export from the client’s system.
No. Z′ can be a useful separation statistic, but acceptable performance depends on assay design, control construction, stage, endpoint, spatial behavior, replication, and the decision being made. We combine appropriate metrics with plate maps, control trends, replicate behavior, and assay-specific acceptance logic.
A validly tested low- or no-activity variant can be informative and should not automatically be discarded. It is kept distinct from a technical failure, missing observation, censored value, or untested candidate. This distinction is important for both engineering interpretation and model training.
No. Utility depends on label reliability, sample size, diversity, landscape complexity, coverage of the intended deployment domain, assay noise, confounding, and the performance threshold required for the decision. We provide an evidence-based readiness assessment and can recommend the most valuable next data to collect.
Yes, when the relevant endpoints and context are available. We can preserve component measurements, evaluate correlations and tradeoffs, apply hard constraints, compare Pareto-efficient candidates, and create versioned scenario-specific scores. A single composite score is not treated as the only truth.
Send the scientific decision, assay protocol, sequence/construct and sample tables, plate maps, a representative raw export, the current processed result, control definitions, and a description of the desired handoff or future prediction problem. Sensitive materials can be discussed within the agreed confidentiality process.
The service is configured for the client’s assay and decision. The following resources inform our approach to screening data levels, assay quality, metadata, reuse, training-set design, and leakage:
We can scope an experimental data contract, legacy-data audit, screening analysis, or AI-readiness package around the evidence you have and the decision your next experiment must support.