Genomics, Transcriptomics & Systems Biology

reference genome

When you photograph a document, it helps enormously to have a clean master copy on file. To check whether your photo has a smudge, you do not re-read the whole page from scratch — you line it up against the master and look only at where they differ. In genomics, that agreed-upon master copy of a species' DNA is the reference genome, and almost everything in the field is read by comparison to it.

A reference genome is a high-quality, carefully assembled representative sequence for a species, maintained and updated by the community as a shared coordinate system. Its real power is that once it exists, you no longer need to assemble each new individual from scratch. You sequence a new person's DNA into short reads and simply map (align) each read to its matching spot on the reference, a far easier problem than building a genome from nothing. Differences between the reads and the reference — a changed letter here, a missing chunk there — are recorded as that individual's variants, and that compact list of differences is what most genomics actually studies.

The catch is that a reference is a representative, not a universal truth. The human reference is a composite of a handful of people and was historically biased toward a few populations, so DNA that is common in under-represented groups can be missing from it entirely, leading to reads that map poorly or variants that look 'abnormal' only because the reference is incomplete. The field is moving toward pangenomes — references built from many diverse individuals at once — precisely to fix this. Treating any single reference as 'the' genome of a species, or as a standard of normality, is a real and consequential error.

To find a person's variants, software aligns their reads to the human reference and reports differences as compact entries like 'chromosome 7, position 117559590, C changed to T' rather than re-storing all 3.2 billion bases.

A reference is a shared coordinate system; individual genomes are recorded as differences from it.

A reference genome is a representative, not a standard of 'normal'. It is biased toward the people who built it, which is why diverse pangenomes are now being adopted.

Also called
reference sequencereference assembly参考序列參考序列