Computational & In Silico Methods

cheminformatics

Cheminformatics is the craft of teaching computers to handle molecules as data, the way bioinformatics handles genes and sequences. Before any model can rank or predict anything, someone has to store millions of structures, search them, compare them, and clean the messy data, and that plumbing is cheminformatics.

Its building blocks include ways to write molecules as text or codes, such as line notations that capture a structure in a single string, plus methods to compute molecular descriptors and molecular fingerprints, to measure similarity between molecules, to detect duplicates and standardize charges and tautomers, and to organize libraries for fast searching. These foundations make virtual screening, QSAR, and machine learning even possible.

Though it sounds like mere bookkeeping, good cheminformatics is what separates a reliable analysis from a misleading one. Subtle issues, such as the same molecule written two different ways, inconsistent salt forms, or activity values measured by incomparable assays, can quietly corrupt a model. Much of the real work in computational drug discovery is the careful, unglamorous data curation that cheminformatics provides.

Garbage in, garbage out applies forcefully here: a machine-learning model is only as trustworthy as the curated chemical and assay data feeding it.

Also called
chemoinformatics化学信息处理化學資訊處理