Case Study – Version Controlling a CRISPR-Cas9 editing protocol

By Dr Neil Williamson from NW Bio-X Limited

Strain identity is one half of the data management problem. The other is the protocol itself: the method that must perform consistently regardless of who runs it, which reagents they use, or which instrument they work on. The composite illustrative case study below demonstrates that CellRepo has standalone value as a method version control system, independent of Genosignature®.

Human cell engineering has become close to routine over the past decade. CRISPR-Cas9 and its derivatives, base editors and prime editors, alongside improved delivery systems including lipid nanoparticles and AAV vectors, have transformed what was once a specialist capability requiring significant infrastructure into a workflow now offered as a catalogue service by CROs. The provenance problem scales directly with the ease of engineering: the faster and cheaper it becomes to generate a modified cell line, the more critical it becomes to retain a robust record of exactly how that modification was made. This applies with particular force to patient-derived induced pluripotent stem cell (iPSC) lines. A 2025 international stem cell banking workshop identified DNA barcoding, combined with genomic and phenotypic assays, as an emerging tool for tracking clonal dynamics and supporting manufacturing quality control in human pluripotent stem cell lines [6], a capability since demonstrated directly through a multi-kingdom genetic barcoding system validated in human pluripotent stem cells among other cell types [7].

Composite Case Study

A team optimising a CRISPR-Cas9 knock-in protocol begins with guide RNA design, selecting a design tool, a scoring algorithm, and parameter thresholds for on-target activity and predicted off-target effects. Handover 1, design parameters: without a record of which tool version and thresholds were used, the design step cannot be reproduced exactly. The construct is synthesised and cloned into a delivery vector, with a specific sgRNA scaffold and supplier. Delivery into the target cell line follows, by lipid nanoparticle, AAV, or electroporation. Handover 2, cell-state at transfection: cell confluency and passage number at the moment of transfection directly influence editing efficiency, yet are rarely logged against the specific clone they produced.

Figure 1. CRISPR-Cas9 editing protocol workflow. Numbered markers indicate the four handovers at which protocol metadata is most commonly lost as the work passes from design through to a validated clone.

A selection and screening strategy is applied to identify successfully edited clones, and editing efficiency is assessed by sequencing and indel quantification. Handover 3, screening method: CRISPR-Cas9 enables scarless genome edits that are very difficult to distinguish from naturally occurring mutations without a marker linking the clone back to its intended edit [8], which makes the screening method applied to a given clone essential metadata in its own right. Off-target effects are validated using a chosen method, commonly targeted deep sequencing of predicted off-target sites or an unbiased genome-wide assay. Handover 4, validation method: which assay cleared a given clone, and against which predicted sites, is rarely retained alongside the clone itself.

Each of these steps has parameters that materially affect the outcome but are rarely captured with sufficient precision to be reproduced exactly. When the protocol is revisited months later, perhaps because a downstream clone behaves unexpectedly, the team frequently cannot confidently establish which guide RNA design tool version was used, what confluency the cells were at during transfection, or which off-target validation method was applied to the clone now in question.

With CellRepo, each protocol run is committed as a version: the guide design parameters, the delivery method and reagent lot, the cell state at transfection, and the validation method are all linked directly to the resulting clone and its downstream data. A poor or unexpected result can be traced back to the exact protocol version that produced it, rather than being reconstructed from memory or abandoned in favour of repeating the experiment from scratch.

Read full paper on “The Missing Layer in Biology’s Data Revolution” here

Return To Blog

Share this post:

Contact Us: