In recent years, there has been much talk of "Big Data," or the science of massive datasets (see our feature in
Tangente 181, 2018). In biology, Big Data refers to
omics data.
Sequencing the genome
-------------------
Sequencing techniques aim to determine the nucleotide sequence of a fragment of deoxyribonucleic acid (DNA) from our genome. The data produced are called reads or sequences. They are represented as strings built from an alphabet of four letters, corresponding to the four nucleotides that make up DNA: A (adenine), C (cytosine), T (thymine), and G (guanine).
Since the first human genome was fully sequenced in 2003, tremendous progress has been made in this field. Advances in technology have increased the amount of data generated while reducing production costs. The first sequencing of the complete human genome—which received extensive media coverage—cost about three billion dollars and took more than ten years to sequence just over three billion nucleotides. Today, any human genome can be sequenced in a matter of days for under €1,000!
These technological advances have other benefits as well. The growing range of sequencing protocols now makes it possible to study an organism at different biological scales (see box).