Nucleic Acids: DNA, RNA & Genes

Nucleotides — The Building Blocks

Nucleic acids are polymers composed of repeating units called nucleotides. Each Nucleotide consists of three subunits: a 5-carbon monosaccharide (pentose sugar), a nitrogen-containing Nitrogenous Base, and Phosphoric Acid. In RNA, the sugar is Ribose; in DNA, it is Deoxyribose — identical except for the absence of an oxygen on carbon 2 of deoxyribose. The nitrogenous base is attached to carbon 1 of the pentose sugar, while the phosphoric acid forms an ester linkage with the hydroxyl group at carbon 5. When a base combines with a pentose sugar alone, the resulting compound is called a Nucleoside. Adding phosphoric acid to a nucleoside produces a nucleotide. Nucleotides can carry one, two, or three phosphate groups (e.g., AMP, ADP, ATP in RNA; dAMP, dADP, dATP in DNA). ATP is a particularly important nucleotide that serves as the cell's energy currency, storing energy in the high-energy phosphate bonds between the second and third phosphate groups.
Pentose Sugars: Ribose in RNA nucleotides, deoxyribose in DNA nucleotides — the only difference is the missing OH group at carbon 2 in deoxyribose
Base Attachment: The nitrogenous base connects to carbon 1 of the pentose sugar via a glycosidic bond
Phosphate Attachment: Phosphoric acid () links to carbon 5 of the pentose sugar through a phosphoester bond
Nucleoside vs Nucleotide: A nucleoside is base + sugar only; a nucleotide is nucleoside + phosphate group
Polymerization: Nucleotides link together through Phosphodiester Linkage between the phosphate of one nucleotide and the sugar of the next, forming polynucleotide chains

Nitrogenous Bases Classification

•
Purines (double-ringed): Adenine (A) and Guanine (G) — found in both DNA and RNA
•
Pyrimidines (single-ringed): Cytosine (C) and Thymine (T) in DNA; Cytosine (C) and Uracil (U) in RNA
•
Key Difference: Thymine is exclusive to DNA; Uracil replaces thymine in RNA

Ribonucleotides and Deoxyribonucleotides

•
Adenine: Adenosine (nucleoside) → AMP, ADP, ATP (RNA) / d-Adenosine → dAMP, dADP, dATP (DNA)
•
Guanine: Guanosine → GMP, GDP, GTP (RNA) / d-Guanosine → dGMP, dGDP, dGTP (DNA)
•
Cytosine: Cytidine → CMP, CDP, CTP (RNA) / d-Cytidine → dCMP, dCDP, dCTP (DNA)
•
Thymine: d-Thymidine → dTMP, dTDP, dTTP (DNA only — no RNA equivalent)
•
Uracil: Uridine → UMP, UDP, UTP (RNA only — no DNA equivalent)

DNA — The Double Helix

DNA is the hereditary material that controls the properties and potential activities of a cell. It was first isolated in 1869 by Friedrich Miescher from the nuclei of pus cells, earning the name 'nucleic acid' from its nuclear origin and acidic nature. DNA is constructed from four kinds of deoxyribonucleotides — dAMP, dGMP, dCMP, and dTMP — linked by Phosphodiester Linkage in specific sequences to form long polynucleotide chains. DNA occurs primarily in chromosomes within the cell nucleus, with smaller amounts in mitochondria and chloroplasts.
Miescher's Discovery: First isolated nucleic acids from pus cell nuclei in 1869, naming them for their acidic nature and nuclear location
Chromosomal Location: DNA resides mainly in chromosomes in the nucleus, with trace amounts in mitochondria and chloroplasts
Four Nucleotides: DNA uses dAMP, dGMP, dCMP, and dTMP — each containing deoxyribose sugar, a phosphate group, and one of four bases (A, G, C, T)
The structure of DNA was determined through the combined work of several scientists. In 1951, Erwin Chargaff analysed base ratios and showed that adenine equals thymine and guanine equals cytosine in any DNA sample. Maurice Wilkins and Rosalind Franklin used X-Ray Diffraction to obtain critical structural data. Building on this evidence, James Watson and Francis Crick proposed the Double Helix model in 1953. DNA consists of two polynucleotide strands coiled around each other in antiparallel orientation — one strand runs 5' to 3' while the other runs 3' to 5'. The two strands are held together by Hydrogen Bonding between complementary bases: adenine pairs with thymine through two hydrogen bonds, and guanine pairs with cytosine through three hydrogen bonds. Each complete turn of the helix spans about 34 Å and contains approximately 10 base pairs.
Complementary base pairing rules: adenine always pairs with thymine (2 hydrogen bonds), and guanine always pairs with cytosine (3 hydrogen bonds)
=Adenine — a purine base(—)
=Thymine — a pyrimidine base (DNA only)(—)
=Guanine — a purine base(—)
=Cytosine — a pyrimidine base(—)
→
Uracil (U) replaces thymine, so A pairs with U via 2 H-bonds instead
Antiparallel Strands: The two strands run in opposite directions — one 5' to 3', the other 3' to 5' — giving the helix its stable structure
Chargaff's Rule: In any DNA sample, A ≈ T and G ≈ C — this complementary ratio was key evidence for the base-pairing model
Two H-Bonds (A=T): The adenine-thymine pair is held by two hydrogen bonds, making it relatively easier to separate
Three H-Bonds (G≡C): The guanine-cytosine pair has three hydrogen bonds, providing greater stability — regions with more G-C pairs are harder to denature
Helix Dimensions: Each turn of the double helix is ~34 Å long and contains ~10 base pairs; one Ångström equals one 100-millionth of a centimetre
DNA Content is Species-Specific: The total amount of DNA is fixed for a species (depends on chromosome number) — germ cells (sperm, ova) contain half the DNA of somatic cells

Key Figures in DNA Structure Discovery

1
F. Miescher (1869): First isolated nucleic acids from pus cell nuclei
2
E. Chargaff (1951): Discovered base ratios — A = T, G = C in all DNA
3
M. Wilkins & R. Franklin: Used X-ray diffraction to reveal DNA's helical structure
4
J. Watson & F. Crick (1953): Built the scale model confirming the double helix

Chargaff's Data — Base Composition (% of total bases)

•
Human: A = 30.9%, T = 29.4%, G = 19.9%, C = 19.8%
•
Sheep: A = 29.3%, T = 28.3%, G = 21.4%, C = 21.0%
•
Wheat: A = 27.3%, T = 27.1%, G = 22.7%, C = 22.8%
•
Yeast: A = 31.3%, T = 32.9%, G = 18.7%, C = 17.1%

DNA Content per Nucleus (species-specific comparison)

•
Chicken — somatic cells: ~2.4 pg/nucleus (red blood cells ~2.3 pg, liver & kidney ~2.4 pg)
•
Chicken — germ cells (sperm): ~1.3 pg/nucleus (half of somatic)
•
Carp — somatic cells: ~3.3 pg/nucleus (RBC, liver, kidney all ~3.3 pg)
•
Carp — germ cells (sperm): ~1.6 pg/nucleus (half of somatic)

RNA — Structure, Types & Function

RNA is a polymer of Ribonucleotides, each containing Ribose sugar, a phosphate group, and one of four bases: adenine (A), guanine (G), cytosine (C), and uracil (U). Unlike DNA, RNA exists as a single-stranded molecule. However, this single strand can fold back on itself to form regions of double-helical character through intramolecular base pairing between complementary sequences. The base pairing follows the same principle as DNA: cytosine pairs with guanine and uracil pairs with adenine. RNA is synthesized from DNA through a process called transcription. It is found in the nucleolus, ribosomes, cytosol, and in smaller amounts throughout the cell.
Single-Stranded Polymer: RNA is a single polynucleotide chain, unlike DNA's double helix — but it can fold to create local double-stranded regions
Uracil Replaces Thymine: RNA uses uracil (U) instead of thymine (T); uracil pairs with adenine via two hydrogen bonds
Ribose Sugar: Every RNA nucleotide contains ribose (with OH at carbon 2), distinguishing it from deoxyribose in DNA
Self-Complementary Folding: The single strand folds back on itself so that C pairs with G and U pairs with A within the same molecule, giving local double-helical regions
Synthesized from DNA: RNA is made by transcription — DNA serves as the template strand
There are three main types of RNA, each with a distinct role in the cell: Messenger RNA (mRNA), Transfer RNA (tRNA), and Ribosomal RNA (rRNA). All three are synthesized from DNA in the nucleus and then transported to the cytoplasm to perform their functions.
mRNA — The Messenger (3–4% of total RNA): Carries the genetic information from DNA in the nucleus to ribosomes in the cytoplasm. It is a single-stranded molecule of variable length, determined by the size of the gene and the protein it codes for. For example, a protein of 1,000 amino acids requires an mRNA of ~3,000 nucleotides (3 nucleotides per amino acid)
tRNA — The Transfer Agent (10–20% of total RNA): Small molecules of 75–90 nucleotides each. There is at least one specific tRNA for each of the 20 amino acids. tRNA picks up amino acids in the cytoplasm and transfers them to the ribosome, where they are linked together to form proteins
rRNA — The Molecular Machinery (up to 80% of total RNA): The most abundant type of RNA, strongly associated with ribosomal proteins (40–50% of ribosome mass). rRNA forms the structural and catalytic core of the ribosome, where mRNA and tRNA interact to translate genetic information into protein

Comparison of RNA Types

•
mRNA: 3–4% of cellular RNA, variable length, carries genetic code from DNA to ribosomes
•
tRNA: 10–20% of cellular RNA, 75–90 nucleotides long, transfers amino acids to ribosome
•
rRNA: Up to 80% of cellular RNA, structural + catalytic core of ribosomes, facilitates protein synthesis

Genes — Units of Biological Inheritance

All the information for the structure and functioning of a cell is stored in DNA. This information is organized into discrete units called Genes. A gene is a segment of DNA that contains the code for producing a specific Polypeptide chain (protein). In the bacterium Escherichia coli, each strand of DNA contains approximately 5 million bases arranged in a particular linear order, divided into units of several hundred bases each — each unit being a gene. The E. coli genome consists of 4,639,221 base pairs coding for at least 4,288 proteins. The concept of a gene as a DNA sequence coding for a polypeptide is central to understanding how genetic information flows from DNA to RNA to protein.
DNA as Information Storage: DNA stores all instructions for cell structure and function in the sequence of its bases
Gene Definition: A gene is a specific segment of DNA that codes for a particular polypeptide (protein)
Linear Organization: Genes are arranged in a linear sequence along the DNA strand, each comprising several hundred bases
E. coli Genome Example: 4,639,221 base pairs divided into at least 4,288 genes, each coding for a specific protein
Polypeptide Product: Each gene's base sequence determines the amino acid sequence of a polypeptide chain through the processes of transcription and translation

Gene Facts from E. coli

•
Each DNA strand contains ~5 million bases
•
Total genome: 4,639,221 base pairs
•
Codes for at least 4,288 different proteins
•
Information is divided into units of several hundred bases each (genes)

DNA vs RNA — Key Differences

While both DNA and RNA are nucleic acids composed of nucleotide polymers, they differ in several structural and functional aspects. Understanding these differences is fundamental to grasping how genetic information is stored and expressed in cells.

Structural Differences Between DNA and RNA

•
Sugar: DNA has deoxyribose (no OH at C-2); RNA has ribose (OH at C-2)
•
Strands: DNA is double-stranded (double helix); RNA is single-stranded (may fold back on itself)
•
Bases: DNA uses adenine, guanine, cytosine, thymine; RNA uses adenine, guanine, cytosine, uracil
•
Size: DNA is very long (millions of base pairs); RNA is shorter (75–3,000+ nucleotides depending on type)
•
Location: DNA is primarily in the nucleus; RNA is in the nucleolus, ribosomes, cytosol
•
Function: DNA stores genetic information long-term; RNA carries and expresses that information