As biology becomes easier to engineer, and more and more engineered biology is being created, a fundamental problem has quietly emerged: living cells do not carry reliable proof of how they were built, who built them or where they came from.
Labels get lost. Files drift out of sync. Documentation breaks as strains move between teams, companies, and countries. In software, this problem was solved decades ago with version control. In biology, it remains largely unsolved.
Genosignatures are GitLife’s answer. They are a new class of DNA barcodes that embed identity, provenance, and version history directly into the genome, permanently linking a living cell to its digital record. The digital record is maintained in CellRepo, our cloud-based biological version control platform which retains a full digital history of the development lineage of a genome. See image below.
What are Genosignatures?
A Genosignature is a short, engineered DNA sequence inserted into an organism’s genome that acts as a unique, machine‑readable identifier for that specific biological version.
Unlike external identifiers such as tube labels, database IDs, or file names, a Genosignature is:
-
- Physically embedded in the cell
-
- Inherited through cell division
-
- Readable at any time via DNA sequencing
In practical terms, this means the cell itself becomes the source of truth.
How Genosignatures work
1. Designed as Digital Identifiers First
Genosignatures are not arbitrary genetic tags. They originate as unique digital identifiers, similar in spirit to commit hashes or UUIDs in software systems. These identifiers can be systematically encoded into DNA using structured, bio‑orthogonal designs that:
- Avoid similarity to natural genetic elements
- Minimise biological interference
- Support robust decoding after sequencing
This ensures that each Genosignature is globally unique and unambiguous, even at scale.
2. Inserted into Neutral Genomic Regions with Error Correcting Technology
Once designed, the Genosignature is inserted into a genetically neutral site, a stable region of the genome shown not to affect growth, fitness, or function, using state-of-the-art predictive models. Gitlife have demonstrated this approach across multiple widely used organisms, including Escherichia coli, Bacillus subtilis,P.putidas, C. Glutamicum, K. phaffii, S. cereviciae, etc.
Multiple insertion techniques have been demonstrated, including:
-
- λ‑Red recombineering
-
- CRISPR/Cas‑based genome editing
-
- Cre‑Lox recombination
-
- Targetron‑based insertion systems
Once inserted, the Genosignature becomes a stable genomic feature, inherited like any other piece of DNA. Experiments in E. coli have shown that these barcodes remain stable (100% sequence identity) after 500 generations of growth in a bio-reactor, ensuring that the link between a biological sample and its digital record is preserved over time. Should any changes emerge over many rounds of subculture, our newest Genosignatures incorporate novel error correcting technology that can be used to retrieve all of the information contained within the barcode even if it were to become partially corrupted through multiple mutations.
Harry Jackson, Principal Synthetic Biologist of Gitlife explains why this is important: “A barcode is only useful if you can trust it decades and thousands of generations later. Stability and error correction ensure that the link between a living cell and its digital record remains intact over time.“
3. Sequencing Links the Cell to Its History
The power of Genosignatures lies in what happens after insertion.
By sequencing the Genosignature, researchers can retrieve the complete digital record associated with that specific biological version, including:
-
- its design rationale;
-
- engineering steps;
-
- contributors;
-
- and validation data, information that a whole genome sequence by itself cannot provide
This approach applies version‑control principles to living systems, where each major engineering step corresponds to a new, identifiable biological state.
In effect, the cell becomes a pointer to its own provenance.
The figure below illustrates GitLife’s Genosignature design and validation pipeline. If you are interested in learning more, please refer to the Nature Communications paper referenced below from our founder’s lab, which presents foundational work in biological barcoding and its experimental validation.
Why Genosignatures are different from traditional DNA barcodes?
Embedding DNA sequences into genomes as barcodes is not entirely new. DNA barcodes have been used for years in research to track cell lineages, measure fitness in pooled experiments, or distinguish strains within a single laboratory (our whitepaper is a good introduction). In some industrial settings, companies also insert proprietary DNA sequences into their strains to help with internal identification.
They are often random or experiment‑specific, meaningful only within a single organisation or study, and disconnected from any structured digital record. As a result, they cannot reliably answer questions about provenance, authorship, or which version of a strain is being used.
Genosignatures take a different approach. They are designed as persistent identifiers, embedded once and carried for the lifetime of the biological asset, explicitly linking a cell to its full digital history in CellRepo. This shifts DNA barcoding from an isolated laboratory technique into a foundation for biological version control and trust at the community level.
The figure above illustrates GitLife’s Genosignature design and validation pipeline. Readers interested in learning more can refer to the Nature Communications paper from our founder’s lab, which presents foundational work in biological barcoding and its experimental validation.
Nature communications paper: Tellechea-Luzardo, Jonathan, et al. “Versioning biological cells for trustworthy cell engineering.” Nature communications 13.1 (2022): 765. https://www.nature.com/articles/s41467-022-28350-4