By Dr Neil Williamson from NW Bio-X Limited
The operational impact of data management failures is most clearly traced through a real engineering biology workflow. The composite illustrative case study below is based on common patterns observed across multiple engineering biology organisations. It does not represent any single company or project.
Composite Illustrative Case Study
A pathway is first identified by the bioinformatics team, either a natural pathway or one constructed by retrosynthesis. Parts are ordered and domesticated, a process that removes internal Type IIS restriction sites (BsaI, BsmBI) to allow assembly methods such as Golden Gate or MoClo (Engler, Kandzia, and Marillonnet, 2008), and codon-optimised or harmonised for the intended host. The assembled constructs are screened and transformed into model and non-model chassis by the strain engineering team. Handover 1, construct origin: the specific domesticated part and assembly version behind a given construct is rarely carried forward once the strain engineering team takes over.
Strains are selected and screened, and the best-performing construct is taken forward, often assessed across multiple candidate media. At this point, progress is reviewed at a project meeting. Handover 2, selection logic: the decision criteria applied at that review are rarely captured in a form that can be retrieved later. The strain may by now have passed through several research scientists and at least three teams: bioinformatics, strain engineering, and analytics. The decision criteria applied are rarely captured in a form that can be retrieved later. The strain may by now have passed through several research scientists and at least three teams: bioinformatics, strain engineering, and analytics.
If further development is required, such as upregulation of a precursor pathway, the team iterates the engineering steps. Handover 3, variant lineage: establishing which iteration round produced the working strain becomes progressively harder with each cycle. Master and working cell banks are then created and validated, each requiring an unambiguous chain of custody. Handover 4, bank identity: passage history, QC records, storage conditions, and verified strain identity all need to travel with the bank, not just alongside it.
The strain progresses through upstream and downstream process development, moving from 500 mL shake flask to 3 L bench bioreactor to a 200 L pilot vessel, before technology transfer and scale-up to 3,000 L through an external CDMO. Handover 5, reagent metadata: media composition, supplier, and lot number behind a given yield are rarely retained at the resolution needed to reproduce it later. Before technology transfer and scale-up to 3,000 L via an external CDMO, the vessel size, aeration rate, feeding strategy (batch, fed-batch, or continuous), and verified cell bank identity must be captured and carried forward. Handover 6, CDMO interface data: without a verification step, there is no direct way to confirm that the strain the CDMO received is the strain that was actually sent. A Genosignature implementation specifically scoped for the receiving CDMO would allow that organisation to confirm, by sequencing, that the strain delivered to their facility is the intended strain before a single litre of fermentation capacity is committed.
Throughout this journey to commercialisation, a safety data package for regulatory submission and a documented IP strategy both depend on the same underlying record. Every handoff in this pipeline is a point at which the receiving team or external partner must reproduce the previous team’s results. A detailed, version-controlled record at each stage shortens the time and cost associated with IP strategy development and substantially de-risks the tech transfer itself.