Chapter 4. Sequences and Strings
In this chapter you will begin to write Perl programs that manipulate biological sequence data, that is, DNA and proteins. Once you have the sequences in the computer, you’ll start writing programs that do the following with the sequence data:
Transcribe DNA to RNA
Concatenate sequences
Make the reverse complement of sequences
Read sequence data from files
You’ll also write programs that give information about your sequences. How GC-rich is your DNA? How hydrophobic is your protein? You’ll see programming techniques you can use to answer these and similar questions.
The Perl skills you will learn in this chapter involve the basics of the language. Here are some of those basics:
Scalar variables
Array variables
String operations such as substitution and translation
Reading data from files
Representing Sequence Data
The majority of this book deals with manipulating symbols that represent the biological sequences of DNA and proteins. The symbols used in bioinformatics to represent these sequences are the same symbols biologists have been using in the literature for this same purpose.
As stated earlier, DNA is composed of four building blocks: the nucleic acids, also called nucleotides or bases. Proteins are composed of 20 building blocks, the amino acids, also called residues. Fragments of proteins are called peptides. Both DNA and proteins are essentially polymers, made from their building blocks attached end to end. So it’s possible to summarize the primary ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access