Genomics, Transcriptomics & Systems Biology

bioinformatics

/ BY-oh-in-fer-MAT-iks /

A single human genome is about 3.2 billion letters, a sequencing run produces hundreds of millions of reads, and a single-cell experiment yields tables with millions of numbers. No one reads any of this by eye. Bioinformatics is the discipline of using computers and statistics to store, search, compare, and make sense of biological data at this scale — it is the indispensable companion to every genome-scale experiment, the part where raw data becomes biology.

In practice, bioinformatics is the toolkit and mindset behind everything else in this field. It includes the algorithms that align one sequence to another or assemble a genome, the databases that store sequences and let the world retrieve them, the file formats that carry the data, and crucially the statistics that decide whether a pattern is real or just noise. When you measure 20,000 genes at once and ask which changed in a disease, you are testing 20,000 hypotheses simultaneously, and by chance alone many will look 'significant'; bioinformatics supplies the multiple-testing corrections and controls that keep you from fooling yourself.

Bioinformatics matters because modern molecular biology is now as much a data science as a wet-lab science — many discoveries are made entirely at the keyboard by re-analyzing public data. But its results are only as trustworthy as their assumptions: a clever pipeline run on biased input, with the wrong statistical model or an unstated parameter choice, can produce confident nonsense. This is why the field places heavy weight on reproducibility — sharing data, code, and exact methods — so that a computational claim, like any scientific claim, can be checked by someone else. Garbage in, garbage out is not a joke here; it is the central risk.

An RNA-seq study finds 1,000 genes 'significantly changed' in a disease; after correcting for testing thousands of genes at once, only 80 survive — a sober reminder that without proper statistics, most of that first list was noise.

Bioinformatics is where data becomes biology — and where good statistics separate signal from noise.

A sophisticated pipeline does not rescue bad data or a wrong statistical model — garbage in, garbage out. Testing thousands of features at once demands multiple-testing correction, or noise masquerades as discovery.

Also called
computational biology生物信息生物資訊