An NGS library is not simply fragmented DNA in a tube. It is a population of molecules with the required insert size, end structure, adapter orientation, index configuration, amplifiability, and platform-facing sequences. Creative Enzymes provides NGS library preparation enzyme system development for reagent companies, sequencing-workflow developers, molecular research teams, and industrial partners. Projects can include enzymatic fragmentation or tagmentation, end repair and A-tailing, adapter ligation, high-fidelity library amplification, cleanup compatibility, formulation, robustness, automation, and transfer studies. The work is scoped around the client's input material, adapter design, downstream enrichment or sequencing process, and acceptance evidence.
Development begins with the structure that must leave library preparation, not with a preferred enzyme catalog. Every upstream choice is tested against that molecular output.
Library preparation converts an input nucleic-acid population into a platform-compatible population while inevitably discarding, failing to convert, or overrepresenting some molecules. The central development problem is therefore a conversion problem. Input mass is useful, but mass does not reveal the number of available fragment ends, damage state, sequence composition, or the fraction that can accept two functional adapters. A short-fragment sample can present many more ends than the same mass of long DNA. A degraded sample can contain blocked or chemically modified termini. A high post-PCR yield can arise from repeated copies of relatively few original molecules.

Fig. 1. Input-to-library molecule journey. Each checkpoint asks how many intended molecules remain and whether their molecular structure is correct.
We translate the desired library into a molecule specification: input type and amount; intended insert range; single- or paired-end read geometry; adapter and index architecture; unique molecular identifier requirements; target-enrichment compatibility; whether amplification is allowed; acceptable short-insert and adapter-dimer fractions; storage and automation constraints; and the downstream evidence used to call a batch acceptable. The specification prevents local optimization from creating a downstream failure. For example, increasing ligation-product fluorescence is not useful if the gain is mainly adapter dimer. Reducing cleanup loss is not useful if the retained distribution falls outside the capture or read-length requirement.
This service can support a new reagent system, replacement of one or more enzymes in an established workflow, second-source development, recovery of a low-conversion process, adaptation to a new sample class, or transfer from a manual protocol to a plate-based format. When input preparation is itself variable, the library project can be connected to our nucleic acid extraction enzyme system optimization work so that extraction yield, fragment integrity, inhibitors, and library conversion are not optimized in isolation.
Fragmentation determines more than the average insert size. It can alter end chemistry, sequence representation, damage sensitivity, input consumption, automation burden, and the number of subsequent enzymatic operations. We compare routes using matched input and downstream library evidence. A route is not selected merely because it is faster or uses fewer tubes.
Acoustic or other physical shearing can provide a fragmentation step that is largely separated from the library enzyme mix. It may be preferred when a mature instrument process exists or when direct enzymatic sequence-context effects must be minimized.
A nuclease-based route can be integrated with end repair and A-tailing. Fragment size is influenced by enzyme activity, time, temperature, input amount, DNA quality, salts, and mixing.
A loaded transposase fragments DNA and introduces adapter tags in the same reaction. Downstream steps complete platform-facing adapters and indexes according to the chosen architecture.

Fig. 2. Fragmentation route selector. The route is chosen from input state, desired insert distribution, end chemistry, workflow constraints, and bias evidence.
For enzymatic fragmentation, we establish a response surface rather than a single incubation. The study can vary intact-input mass, sample concentration, fragmenting activity, enzyme-to-substrate ratio, time, temperature, buffer, and mixing. Stop behavior matters: a reaction that continues during deck movement or warming can produce plate-position or operator-dependent insert distributions. We therefore test the specified stop or transition step under realistic hold times. If fragmentation, end repair, and A-tailing share a tube, the final fragmentation condition must also create a suitable chemical environment for the next enzymes.
Tagmentation demands a different model. Fragment-size behavior depends on access to DNA and active transposome, while the introduced tags define downstream primer and adapter logic. A nominally attractive fragment trace does not establish that both ends are correctly configured or that sequence representation is acceptable. Development can include transposase loading, adapter-tag composition, DNA normalization sensitivity, reaction quench, tag-completion or index PCR, cleanup, and pilot sequencing. The original high-density in-vitro transposition work showed the value of combining fragmentation and adapter incorporation, but a project-specific system still requires its own input range and evidence.
| Selection question | Mechanical shearing | Enzymatic fragmentation | Tagmentation | Evidence used in development |
|---|---|---|---|---|
| Is fragmentation separable from enzyme formulation? | Yes; the instrument process is external | No; nuclease activity is part of the reagent system | No; fragmentation and tag introduction are coupled | Matched input, fragment trace, conversion assay, and pilot data |
| How is insert size controlled? | Instrument settings, vessel, volume, input quality | Activity, time, temperature, substrate, buffer, mixing | Transposome loading, DNA amount/concentration, time, chemistry | Full distribution and tails, not only modal size |
| Which end-state work follows? | End repair, phosphorylation, A-tailing as required | May be separate or integrated with repair/A-tailing | Tag completion and adapter/index completion are architecture-specific | End-structure or functional-ligation evidence |
| Main development risk | Instrument and transfer variability | Overdigestion, condition sensitivity, representation bias | Input normalization sensitivity, tag geometry, insertion bias | Robustness panel and sequencing metrics tied to intended use |
In a ligation-based short-read workflow, fragmented DNA may require blunting or fill-in, removal of incompatible overhangs, 5-prime phosphorylation, 3-prime A addition, adapter ligation, cleanup, and optional amplification. Vendors often combine some of these operations, but combination does not eliminate chemical dependencies. Nucleotides, cofactors, salts, reducing agents, crowding components, inactivation products, residual nuclease, and temperature history pass from one operation to the next.

Fig. 3. Enzyme handoff map. Each stage must deliver the right substrate and chemical environment to the next stage.
End preparation is evaluated by function as well as bulk yield. Where appropriate, the development plan uses defined end substrates, challenging end structures, or ligation-readout assays to distinguish repair from ligation. A polymerase, kinase, exonuclease, or A-tailing activity may perform well on an ideal control but underperform after the selected fragmentation reaction. Conversely, a downstream ligase failure can make end preparation appear weak. Module-level tests prevent one enzyme from being adjusted to compensate for an unidentified defect elsewhere.
Is the input double-stranded, intact, nicked, chemically damaged, already fragmented, or an amplicon? Which termini are actually present?
Do carried salts, cofactors, nucleotides, PEG-like components, stop reagents, or bead eluates support the next activity?
Does the incubation sequence stop the preceding reaction and activate the next one without harming adapters or enzymes?
Can conversion of intended inserts be separated from free adapter, adapter dimer, short products, and amplification artifacts?
Creative Enzymes can screen native or engineered enzyme candidates, adjust activity ratios, and formulate multi-enzyme mixes for the required temperature program. If an enzyme needs altered activity, specificity, inhibitor tolerance, or storage behavior, the program can connect to enzyme engineering and modification. Candidate selection is followed by lot-aware assays and fit-for-purpose impurity controls through our enzyme QC and QA support. A combined mix is advanced only when the downstream library evidence is comparable to or better than the agreed reference under the project's conditions.
Adapter ligation is governed by molecules and ends, not DNA mass alone. Shorter inserts provide more ligatable ends per nanogram than long fragments. Low-input reactions can contain a large molar excess of adapter relative to available ends, which favors adapter-derived species if the adapter architecture and cleanup do not suppress them. Too little adapter can leave intended inserts incompletely converted. The useful operating window therefore depends on input amount, fragment distribution, adapter design, ligase system, crowding environment, reaction volume, and cleanup.
Estimate the effective concentration from input mass, size distribution, damage, and conversion-ready fraction. The estimate guides a titration; it is not treated as exact truth.
Define Y, stubby, full-length, indexed, UMI-containing, phosphorylated, blocked, or other structural features and their annealing quality.
Co-optimize ligase activity, enhancer, time, temperature, bead ratio, wash, elution, and the downstream PCR or enrichment requirement.
An apparent short peak should not automatically be labeled adapter dimer. Its size relative to the known adapter structure, response to insert-free controls, amplification behavior, and sequencing composition are considered. A short biological insert with correctly ligated adapters can migrate near adapter-derived products and may be valuable for cfDNA or degraded-DNA applications. Conversely, removing every short molecule can erase the intended sample signal. We define which short species are undesirable before selecting bead ratios or designing a size cutoff.
SPRI-type bead cleanup is a unit operation with its own loss and bias. Bead-to-sample ratio, sample chemistry, mixing, incubation, magnet time, wash dryness, residual ethanol, elution volume, and aspiration height affect recovery. Double-sided size selection can sharpen a distribution but adds two boundaries and more opportunities for loss. During automation development, edge wells, delayed columns, mixing geometry, dead volume, and tip behavior are included. Enzyme optimization cannot compensate reliably for a bead process that changes across the plate.
| Observed result | Possible causes | Discriminating checks | Development response |
|---|---|---|---|
| High fluorometric yield, low library qPCR | Incomplete adapters, non-amplifiable DNA, residual input, damaged ends, inhibition | Pre/post-ligation controls, dilution series, adapter-specific qPCR, fragment trace | Isolate end preparation, ligation, and inhibition before increasing PCR |
| Prominent short product | Adapter dimer, primer dimer, true short inserts, excessive fragmentation | Insert-free control, adapter structure calculation, no-PCR trace, sequencing composition | Adjust adapter input, cleanup boundary, fragmentation, or PCR according to identity |
| Broad or shifted insert distribution | Fragmentation drift, stop delay, bead-ratio variation, degraded input | Time course, plate-position study, pre-cleanup trace, input integrity | Stabilize the upstream operation rather than masking it with size selection |
| Uneven index representation after pooling | Quantification error, index PCR differences, pipetting, library-size effects | Functional concentration, replicate indexing, pool reconstruction, index balance | Define normalization and pooling rules with appropriate controls |
Indexing can occur through full-length indexed adapters or through PCR that completes adapter sequences and adds indexes. Unique dual indexes can help identify or control certain sample-assignment artifacts, but index design does not remove errors introduced before indexing. Where unique molecular identifiers are used, their location, diversity, read structure, ligation efficiency, consensus strategy, and downstream software must be specified. Oligonucleotide sequence ownership and platform licenses remain client responsibilities unless separately scoped.
Library bias is any systematic difference between the molecular population of interest and the sequenceable population. It can arise before library preparation, and it can accumulate at several enzymatic and physical steps. Aird and colleagues identified library PCR as a major source of base-composition bias in the studied Illumina libraries, but PCR is not the only source. Fragmentation preference, damaged-end repair, ligation efficiency, cleanup retention, hybrid capture, cluster generation, and analysis can each reshape representation.

Fig. 4. Library bias budget. Risk is traced across the workflow so that a downstream adjustment is not credited with correcting an upstream loss.
PCR-free development can reduce PCR-induced duplication and amplification bias, but it is not a universal specification. The input must provide enough correctly converted molecules for the downstream process, and the adapter architecture must be complete without amplification. Functional concentration, not input mass alone, determines whether the route is practical. If PCR is needed, we select a high-fidelity polymerase and define primer concentration, denaturation, extension, cycle number, and stopping rule. More cycles can increase tube yield while reducing library complexity and increasing duplicates. The minimum cycle count is therefore determined against functional-yield and sequencing requirements, not a cosmetic electropherogram target.
Sequence representation is assessed with controls suited to the application. A microbial whole-genome library may be evaluated for coverage uniformity and GC response. A targeted library adds on-target rate, fold-80-like uniformity measures, and duplicate behavior after enrichment. A cfDNA workflow may require preservation of short-fragment distributions and molecule families. An FFPE workflow may require damage-related artifact assessment. No single metric is declared universally sufficient.
The same enzyme mixture should not be assumed to serve intact genomic DNA, cfDNA, FFPE DNA, amplicons, and ultra-low-input DNA without adjustment. We choose representative samples, contrived controls, and reference materials that expose the relevant failure modes. Each lane below changes both the chemistry and the evidence plan.
Diagnostic-adjacent sequencing projects may use libraries for pathogen characterization, inherited-variant research, oncology research, or assay-development studies. The laboratory workflow, reference materials, bioinformatics, and acceptance criteria must reflect that purpose, but this service does not confer a diagnostic claim. For projects that combine sequencing with an orthogonal detection method, our CRISPR diagnostic assay development support and planned digital PCR and digital LAMP reagent development pages describe separate system-development questions.
A development candidate advances through evidence levels. Early assays are faster and more diagnostic; later assays are more integrated and expensive. Running pilot sequencing before enzyme modules are understood can reveal failure without locating it. Conversely, declaring success from synthetic-substrate activity alone ignores the complete workflow.
Defined substrates, activity ratios, impurity checks, and formulation stress establish what each enzyme can do.
Fragment size, end preparation, ligation controls, and cleanup recovery show how input becomes library.
Adapter-specific qPCR, fragment analysis, yield, and dilution behavior estimate usable molecules.
Coverage, duplicates, insert size, index balance, GC behavior, and application metrics test representation.
Lots, operators, days, holds, plates, instruments, storage, and instructions define reproducible operation.

Fig. 5. Evidence ladder. Candidate decisions move from mechanistic assays to sequencing and transfer without treating any single measurement as universal proof.
Routine QC methods answer different questions. Fluorometry estimates double-stranded DNA mass but does not identify adapter completeness. Electrophoresis or capillary analysis shows a size distribution but can combine structurally different molecules. Adapter-specific qPCR estimates amplifiable library molecules under the primer design and reaction conditions. Sequencing shows the behavior of molecules that survive platform entry and analysis but is affected by loading, instrument, run configuration, and bioinformatics. We use orthogonal measurements and define which method releases the reagent versus which method characterizes development.
| Measurement | What it can answer | What it cannot prove alone | Typical development use |
|---|---|---|---|
| Input integrity and amplifiability | Starting fragment state and ability to support a test amplicon | Library conversion or final sequence representation | Stratify samples and explain outliers |
| Fluorometric DNA concentration | Approximate double-stranded DNA mass | Correct adapters, index structure, amplifiability, complexity | Mass recovery and pooling support |
| Fragment analysis | Apparent size distribution and prominent short species | Exact molecular identity or platform functionality | Fragmentation, cleanup, and dimer investigation |
| Adapter-specific qPCR | Amplifiable molecules recognized by the selected primer pair | Unbiased genome representation or all platform steps | Functional yield, dilution response, pooling |
| Pilot sequencing | Integrated insert, coverage, duplicates, indexes, error, and application metrics | The unique biochemical cause of every failure | Candidate confirmation and intended-workflow comparison |
Record input classes, adapter/index structure, sequencing or enrichment interface, workflow constraints, reference process, acceptance metrics, exclusions, and ownership of oligos, software, licenses, and validation.
Decide whether fragmentation is mechanical, nuclease-based, or transposase-based; identify end-preparation, ligation, cleanup, amplification, and quantification dependencies; and create module controls.
Evaluate candidate enzymes, activity ratios, buffers, temperature programs, time sensitivity, adapter levels, beads, and input states using compact experiments that answer the highest-risk decisions.
Test full libraries across representative and boundary inputs, short holds, plate positions, operators, days, reagent lots, freeze-thaw or storage conditions, and the intended manual or automated process.
Compare candidates using predefined library and sequencing metrics. Investigate discrepancies with module controls rather than changing multiple reagents simultaneously.
Deliver agreed formulas, component specifications, preparation instructions, in-process controls, release methods, acceptance rules, stability evidence, deviation guidance, and a change-control baseline.
Deliverables are project-specific. They can include a development plan, risk register, enzyme and formulation screen, response-surface data, candidate ranking, bill of materials, preparation record, test methods, sample and control plan, robustness report, pilot-sequencing comparison, draft specifications, and transfer protocol. For clients planning larger reagent lots, the package can connect to enzyme production and scale-up. If a liquid workflow must become a dry or ambient-stable format, lyophilization of molecular diagnostic reagents is treated as a separate formulation program because drying can change enzyme activity, adapter integrity, rehydration, and reaction kinetics.
Projects can start with a complete development need or a specific symptom such as low ligation conversion, excessive adapter dimer, unstable fragment size, loss during cleanup, strong tube yield with weak qPCR, GC-dependent coverage, high duplication, or plate-position effects. Existing reagents and data are reviewed before new experiments are proposed. Creative Enzymes also supplies molecular diagnostic enzymes and kits that may support feasibility studies, subject to project fit and the stated research or industrial-use conditions.
Tell us the input type, current workflow, adapter and index design, target insert range, sequencing or enrichment interface, existing data, main failure, automation needs, and desired development stage. We will use that information to define the module boundaries, comparison strategy, evidence plan, and transfer deliverables.
Contact Creative Enzymes