Why this mattersEvery algae product is encoded here. Engineering starts here.
DNA β the instruction manual of all life
Why the instruction system matters
The blueprint behind every product
Every molecule microalgae produce commercially β astaxanthin, DHA, phycocyanin, pharmaceutical compounds β is made by an enzyme. Every enzyme is a protein. Every protein is built from instructions encoded in DNA. So the DNA is not just a biology curiosity. It is the master specification for every product in the industry. Edit the DNA and you change the product. Understand DNA and you understand why genetic engineering is the most powerful lever in the entire field.
These three weeks build the complete picture: what DNA is and how it stores information, how that information gets read and turned into proteins (the central dogma of molecular biology), and how modern tools like CRISPR allow us to rewrite those instructions β and what that means commercially for microalgae.
The central dogma of molecular biology β the most important sentence in biology
DNA β RNA β Protein β Trait / Product
DNA is copied into RNA (transcription) Β· RNA is read to build a protein (translation) Β· The protein does work that creates a visible trait or a commercial product
This one sentence β DNA makes RNA makes protein β describes how genetic information flows in every living thing on Earth, from bacteria to blue whales to Haematococcus algae. Every algae product you want to make, every productivity improvement you want to engineer, every strain you want to improve β all of it operates through this single pathway. Master this, and genetic engineering becomes logical rather than magical.
Week 10 Β· Part 1 of 4
What DNA is β structure and information storage
DNA stands for deoxyribonucleic acid. That name tells you exactly what it is: a nucleic acid (a long chain of nucleotide units) built around a deoxyribose sugar. But the chemistry is less important than the concept: DNA is an information storage molecule. It stores instructions in the sequence of its chemical letters.
The double helix β a ladder that holds information
DNA has a shape called a double helix β famously described by Watson and Crick in 1953. Imagine a ladder that has been twisted into a spiral. The two sides of the ladder are made of alternating sugar and phosphate molecules β the "backbone." The rungs of the ladder are made of pairs of molecules called bases.
There are exactly four bases in DNA. Their names, and the rule that governs how they pair, is the single most important fact in molecular biology:
A
Adenine
always pairs with T
A purine. Two hydrogen bonds to T. Found in ATP β the energy currency you already know.
T
Thymine
always pairs with A
A pyrimidine. Only in DNA β replaced by Uracil (U) in RNA. Two hydrogen bonds.
G
Guanine
always pairs with C
A purine. Three hydrogen bonds to C β a stronger pair than AβT. Higher GβC content = more heat-stable DNA.
C
Cytosine
always pairs with G
A pyrimidine. Three hydrogen bonds. High GβC content in thermophilic algae β enables survival at high temperatures.
The base-pairing rule is why DNA can copy itself
Because A always pairs with T, and G always pairs with C, knowing one strand's sequence instantly tells you the other strand's sequence. When a cell divides and needs to copy its DNA, it simply unzips the two strands and builds a matching strand on each β using base-pairing as the template. This is the elegance that makes heredity possible. One cell becomes two genetically identical cells every time.
How information is encoded β the sequence is the message
The four bases β A, T, G, C β are the letters of the genetic alphabet. Just as the 26 letters of the English alphabet can be arranged to write every book ever written, these four letters can be arranged in sequences that encode every instruction needed to build and run a living cell.
The human genome contains about 3 billion base pairs. The genome of Chlamydomonas reinhardtii (a model algae species) contains about 120 million base pairs. Nannochloropsis has around 30 million. These billions of letters, read in the right way, contain the instructions for making every protein in the organism.
Scale of DNA β from molecule to genome
What is a gene?
A gene is a specific stretch of DNA that encodes the instructions for building one protein (or sometimes a functional RNA molecule). The human genome has about 20,000β25,000 genes. Nannochloropsis has about 10,000β12,000. Each gene has a defined structure:
Anatomy of a gene β from control switch to stop signal
PromoterON/OFF SwitchWhere RNA polymerase binds to begin reading. Transcription factors bind here to turn the gene up or down. This is how cells regulate which proteins to make and when.
5' UTRLeaderUntranslated region. Helps position the ribosome.
Coding sequence (CDS)The recipeThe actual protein-coding sequence. Read in triplets (codons). Each codon = one amino acid. Starts with ATG (start codon). Ends with a stop codon.
3' UTRTrailerControls mRNA stability and lifespan.
TerminatorStop signalSignals RNA polymerase to stop and release the RNA copy.
Genetic engineers modify all five sections. Swap the promoter β change when/how much of the protein is made. Edit the coding sequence β change the protein's structure. These are the levers of metabolic engineering in algae.
Week 11 Β· Part 2 of 4
From DNA to protein β the two-step read
The instructions in DNA are not read directly. They are first copied into a temporary messenger molecule (RNA), which is then carried out of the nucleus to the ribosomes, where it is used to build a protein. This two-step process β transcription then translation β is universal to all life.
π
Step 1 β Transcription: DNA is copied into mRNA
Location: nucleus Β· Enzyme: RNA polymerase
An enzyme called RNA polymerase binds to the promoter region of a gene and "unzips" the double helix locally. It reads one strand of the DNA and builds a complementary copy β but using RNA nucleotides instead of DNA nucleotides. This copy is called messenger RNA (mRNA). The mRNA sequence mirrors the DNA sequence, except T (thymine) is replaced by U (uracil) in RNA. The mRNA then peels off and travels out of the nucleus through pores in the nuclear membrane into the cytoplasm.
Analogy: the DNA blueprint never leaves the head office (nucleus). Instead, a photocopy (mRNA) is made and sent to the factory floor (ribosome). The original is preserved; the copy is used and eventually degraded.
Algae context: measuring which mRNAs are present in an algae cell (called transcriptomics) tells researchers which genes are currently "switched on." Under nitrogen stress, mRNAs for fatty acid synthesis enzymes spike β a direct readout that the cell has shifted to lipid production mode. This is how researchers identify which genes to target for engineering.
βοΈ
Step 1b β RNA splicing: introns are removed (eukaryotes only)
Location: nucleus Β· Complex: spliceosome
In eukaryotic algae (Chlorella, Haematococcus, Nannochloropsis etc.), genes contain non-coding interruptions called introns, between the protein-coding sequences called exons. After transcription, the spliceosome (a molecular machine) cuts out all the introns and joins the exons together, producing a mature mRNA ready for translation. Prokaryotic algae (Spirulina/cyanobacteria) have almost no introns β one reason they are simpler to engineer.
Analogy: the initial mRNA transcript is like a film reel with scenes (exons) interspersed with blank leader tape (introns). Splicing cuts out the blank tape and joins the real scenes into a continuous film.
Engineering note: when inserting a foreign gene into an algae genome, you often use cDNA (a version of the gene with introns already removed) so the algae's splicing machinery does not interfere with an unfamiliar gene structure.
βοΈ
Step 2 β Translation: mRNA is read to build a protein
Location: ribosome (in cytoplasm, or on endoplasmic reticulum) Β· Molecule: tRNA
The mRNA travels to a ribosome β a molecular machine made of proteins and ribosomal RNA. The ribosome reads the mRNA three letters at a time. Each three-letter sequence is called a codon. Each codon specifies one of the 20 amino acids β or a stop signal. Transfer RNA molecules (tRNA) act as adaptors: each tRNA carries the matching amino acid for its codon and delivers it to the ribosome. The ribosome links amino acids together one by one, in the order specified by the mRNA, building a long chain. This chain is the protein.
Analogy: the mRNA is a ticker tape with three-letter codes. The ribosome is a machine that reads the tape. Each code is an address β the tRNA delivers the right amino acid parcel to that address. The resulting chain of amino acid parcels, folded into a 3D shape, is the protein.
Algae context: the codon usage of different algae species varies β some prefer certain codons for the same amino acid. When engineering algae, inserted genes must be "codon-optimised" for the host species, otherwise the ribosomes translate them inefficiently. This is a routine but critical step in algae genetic engineering.
π§
Step 3 β Protein folding and post-translational modification
The newly built amino acid chain does not work yet. It must fold into a precise 3D shape β determined by its amino acid sequence β before it becomes a functional protein. Molecular chaperones (helper proteins) assist folding. Many proteins are then further modified: sugars are added (glycosylation), phosphate groups are attached (phosphorylation), or portions are cleaved off. These modifications fine-tune the protein's function. The folded, modified protein is now ready for work.
Analogy: the amino acid chain is like a flat-pack piece of furniture delivered in pieces. Folding is the assembly step. Post-translational modifications are the adjustments β tightening screws, adding cushions β that make it fully functional.
Commercial note: phycocyanin is a protein that has a pigment molecule (phycocyanobilin) covalently attached to it post-translationally. You cannot just translate the phycocyanin gene in a bacterium and get the blue colour β you also need the pigment-attachment machinery, which is why producing authentic phycocyanin outside of cyanobacteria is difficult.
The genetic code β 64 codons, 20 amino acids
There are 4Β³ = 64 possible three-letter combinations of A, T, G, C. These 64 codons encode only 20 amino acids plus 3 stop signals. This means most amino acids are encoded by multiple codons β the code is redundant. This redundancy provides a buffer: many mutations in the third letter of a codon change nothing (a "silent" mutation), which is why life is robust to low levels of random DNA damage.
Sample codons β the genetic code in action (DNA sequences)
ATG
Methionine (Start)
TTT / TTC
Phenylalanine
GGT / GGC / GGA / GGG
Glycine (4 codons)
CAT / CAC
Histidine
CCT / CCC / CCA / CCG
Proline (4 codons)
GAA / GAG
Glutamic acid
TGG
Tryptophan (only 1 codon)
TAA / TAG / TGA
STOP β end of protein
Week 11 Β· Part 3 of 4
What proteins actually do
Proteins are the cell's workforce. While DNA stores information and RNA carries it, proteins are the molecules that actually do things. Every chemical reaction in the cell, every structural element, every signal received from the environment β all are executed by proteins. In microalgae, the proteins of interest commercially are almost always enzymes.
βοΈ
Enzymes
Proteins that catalyse (speed up) chemical reactions β often by millions of times. Without enzymes, most biological reactions would be too slow to sustain life. Every metabolic pathway is a chain of enzyme-catalysed steps.
RuBisCO (Calvin Cycle) Β· Astaxanthin synthase Β· Fatty acid synthase Β· CRISPR-Cas9 is itself an enzyme
ποΈ
Structural proteins
Proteins that provide physical shape and support. They make up cell walls, cytoskeletons, and the scaffolding that holds organelles in place. Some are the most abundant proteins on Earth.
Proteins embedded in membranes that move molecules across β either passively (channels) or actively using ATP (pumps). Control what enters and leaves every compartment.
Transcription factors that bind to gene promoters and turn genes on or off. They are the master control switches that determine which proteins the cell makes at any given time.
NRR1 (nitrogen response regulator β triggers lipid production under N starvation in Chlamydomonas)
π¨
Pigment-protein complexes
Proteins with bound pigment molecules β the light-harvesting complexes of photosynthesis and the commercial pigments of algae. The protein scaffold determines the pigment's optical properties.
The enzymeβproduct connection β the commercial chain
For every commercial algae product, there is an enzyme (or a chain of enzymes) that makes it. This chain from gene to enzyme to product is the direct commercial logic of genetic engineering. Here are five key examples:
Insert entire biosynthetic gene clusters from rare/unculturable algae into fast-growing hosts. Access chemistry from organisms too slow or difficult to grow commercially.
Week 12 Β· Part 4 of 4
CRISPR β editing the instruction manual
Now that you understand DNA, genes, and proteins, genetic engineering becomes straightforward in concept. The goal is always the same: change the DNA sequence of a specific gene to change which protein is made, how much of it is made, or what it does. CRISPR-Cas9 is the technology that made this practical, cheap, and precise.
CRISPR stands for Clustered Regularly Interspaced Short Palindromic Repeats β a mouthful that describes a natural immune system in bacteria. Scientists co-opted this bacterial defence mechanism and turned it into a molecular scalpel for editing DNA in any organism, including algae.
01
Design a guide RNA
A short RNA sequence (the "guide") is designed to match the DNA target β the exact stretch of the algae's genome you want to edit. This guide is the address label. It can be designed in software in minutes, then synthesised in a laboratory.
02
Cas9 finds and cuts the target
The guide RNA binds to the Cas9 protein (the molecular scissors). The guide-Cas9 complex scans the genome, finds the matching sequence, binds to it, and makes a precise double-strand cut β breaking both strands of the DNA helix at exactly that location.
03
The cell repairs the break
The algae cell detects the break and repairs it β either imprecisely (knocking out the gene β disabling it) or precisely (if a template is provided, inserting new DNA at that location β editing or inserting a new gene). The outcome is determined by which repair pathway is active and whether a template was supplied.
04
Screen for successful edits
Only a fraction of cells receive and retain the edit. Edited cells are selected using markers (fluorescence, antibiotic resistance) or by directly sequencing their DNA. Successful lines are expanded into stable engineered strains.
05
Verify the phenotype
Does the edit produce the desired change in the protein and ultimately the product? Measure the target compound (astaxanthin, lipid content, protein yield) in the edited strain vs the wild type. If it works as intended β you have a new commercial strain.
06
Navigate regulation
GMO algae face regulatory scrutiny in most markets. Whether CRISPR edits count as GMO varies by country and edit type. Knockouts (removing sequences) are treated differently from insertions of foreign DNA. This regulatory landscape is a major commercial constraint and timeline factor.
What CRISPR has already achieved in algae (as of 2025)
Chlamydomonas reinhardtii (model alga): CRISPR used to knock out competing pathways, increasing lipid accumulation 2β3Γ. Nannochloropsis: CRISPR knockouts of starch synthesis genes redirect carbon to oils, boosting lipid yield significantly without stress protocols. Spirulina: Gene insertion of higher-activity RuBisCO variants tested for improved COβ fixation. Phaeodactylum tricornutum (diatom): EPA content increased 40% by editing fatty acid desaturase genes. These are not theoretical futures β they are published results. The challenge is moving from lab-scale to commercially viable, regulatory-approved production strains.
The master insight of weeks 10β12
DNA is not just a biology topic. It is the engineering specification for every commercial product microalgae make. The pathway DNA β RNA β protein β product is the production line. Every gene is a recipe. Every enzyme is a machine in that line. Understanding this means you can read a biotech company's claims with precision: are they overexpressing an existing gene, knocking out a competing pathway, inserting foreign genes, or editing regulatory regions? Each strategy has different yields, risks, regulatory implications, and timelines. The future of the microalgae industry belongs to companies that can not only grow algae well, but write and edit their instruction manuals with skill.
Quick-reference summary
Concept
Definition
Algae / commercial relevance
DNA
Double-stranded helix of AβT and GβC base pairs. Information stored in sequence.
The master specification for every enzyme that makes every commercial product.
Gene
A stretch of DNA encoding one protein. Has promoter, coding sequence, terminator.
Each commercial product has one or more key genes. Knowing which gene β knowing the engineering target.
Transcription
DNA copied into mRNA by RNA polymerase. Happens in nucleus.
Transcriptomics reveals which genes are active under stress β maps the cell's commercial response.
Translation
mRNA read by ribosome in codons; tRNA delivers amino acids; protein chain built.
Codon optimisation required when inserting foreign genes into algae. Ribosome is the protein factory.
Enzyme
A protein that catalyses a specific chemical reaction.
Every step from glucose β astaxanthin / DHA / phycocyanin is catalysed by a specific enzyme encoded in a specific gene.
CRISPR-Cas9
Guide RNA directs Cas9 scissors to cut DNA at a precise location; cell repairs in desired way.
Used to overexpress product genes, knock out competing pathways, insert foreign biosynthetic clusters. Regulatory status varies by country and edit type.
Promoter
DNA region upstream of a gene β where transcription factors bind to control expression level.
Swap in a strong promoter β make more enzyme β make more product. A primary lever of metabolic engineering.
Self-check β end of week 12
All questions require reasoning from DNA to commercial outcome. Attempt before revealing.
1. A company claims they have engineered Nannochloropsis to produce 3Γ more EPA without any nitrogen starvation stress. Walk through the molecular steps of how this could be achieved β from DNA edit to final higher EPA yield.
The engineering strategy would work through the central dogma pathway. EPA (eicosapentaenoic acid) is synthesised by a chain of fatty acid desaturase and elongase enzymes. Each enzyme is encoded by a specific gene with its own promoter. Step 1 β Identify the rate-limiting enzyme: researchers use transcriptomics to find which enzyme in the EPA pathway is most limiting (lowest expression, most likely Ξ5-desaturase or PUFA synthase). Step 2 β Design CRISPR or overexpression construct: design a DNA construct where a stronger promoter (e.g. from a highly expressed housekeeping gene) drives the target gene at higher levels than the native promoter allows. Codon-optimise if needed. Step 3 β Transform the algae: deliver the construct into Nannochloropsis cells using biolistics (gene gun) or electroporation. Step 4 β Screen: identify cells that integrated the construct and show stable high expression using PCR and sequencing. Step 5 β Measure EPA: use gas chromatography to quantify fatty acid content. If EPA is 3Γ higher, the bottleneck enzyme has been relieved. A secondary strategy: knock out the gene for a competing desaturase that diverts carbon away from EPA toward other lipids. Both strategies work through the same pathway: change the DNA β change which mRNA is produced (and how much) β change how much enzyme the ribosome builds β change the amount of EPA catalysed per unit time.
2. Why does Spirulina (a prokaryote/cyanobacterium) have different genetic engineering challenges than Nannochloropsis (a eukaryote)? Name at least three specific molecular differences.
Three key molecular differences. First β introns: Nannochloropsis genes contain introns (non-coding sequences) that must be spliced out of mRNA before translation. Spirulina genes have almost none. When inserting a foreign gene into Nannochloropsis, you typically use cDNA (introns-removed version). In Spirulina, you can often use the gene as-is. Second β nuclear vs nucleoid: Nannochloropsis has a membrane-bound nucleus containing its chromosomal DNA. Getting DNA into the nucleus requires specific delivery mechanisms (electroporation, biolistics) and the construct must include nuclear localisation signals for some proteins. Spirulina has no nucleus β its DNA floats in the cytoplasm (nucleoid), making direct transformation somewhat simpler. Third β codon usage: eukaryotic algae like Nannochloropsis have specific codon preferences that differ from prokaryotes and from each other. Genes must be codon-optimised for the host species or they are translated inefficiently. Additionally, Nannochloropsis has three genomes to consider (nuclear, chloroplast, mitochondrial) β depending on which compartment you want to target the product to, the gene must be directed to the right genome with appropriate targeting sequences. Spirulina has only one main chromosome plus a small plasmid.
3. A mutation changes a single DNA base in the BKT gene (beta-carotene ketolase) in Haematococcus β from a G to an A at position 542. The codon changes from GGT to AGT. GGT codes for Glycine; AGT codes for Serine. Would you expect this change to affect astaxanthin production? How would you determine whether it does?
This is a missense mutation β a single base change that changes one amino acid (Glycine at position 181 of the protein to Serine). Whether it affects astaxanthin production depends entirely on whether that specific amino acid position matters for the enzyme's function. If position 181 is in or near the enzyme's active site (where it binds its substrate and performs catalysis), changing its character from Glycine (very small, no side chain) to Serine (slightly larger, has a hydroxyl group) could alter the active site geometry and reduce or eliminate catalytic activity β resulting in less astaxanthin. If position 181 is in a structurally unimportant region, the change may have no effect whatsoever (a "neutral" missense). To determine the outcome experimentally: (1) Grow the mutant Haematococcus under standard stress conditions and measure astaxanthin yield by HPLC compared to wild type. (2) Extract and purify BKT protein from both; test enzymatic activity in vitro by measuring how quickly it converts Ξ²-carotene to astaxanthin. (3) Computationally: model the 3D structure of BKT and assess whether position 181 is near the active site (using tools like AlphaFold protein structure prediction). This type of mutation analysis β screening for beneficial or neutral mutations β is standard practice in directed evolution of enzymes for commercial algae production.
4. An algae biotech company has inserted the gene for human insulin into Chlamydomonas. The gene is correctly sequenced and the promoter is active. But they are getting very little insulin protein. Give three possible molecular explanations and for each, describe how you would test it.
Possible explanation 1 β Codon bias: the human insulin gene uses human codon preferences. Chlamydomonas ribosomes may rarely encounter some of those codons and stall frequently, producing very little protein. Test: synthesise a codon-optimised version of the insulin gene using Chlamydomonas-preferred codons and re-measure protein output. If it increases substantially, codon bias was the problem. Possible explanation 2 β mRNA instability: the 3' UTR of the human insulin gene may not be recognised as stable by Chlamydomonas mRNA stability factors, leading to rapid mRNA degradation before translation occurs. Test: perform quantitative PCR (qRT-PCR) to measure mRNA levels. If mRNA is present at high levels but protein is low, the problem is translation or protein stability. If mRNA is low, the problem is transcription or mRNA instability. Replace the 3' UTR with a Chlamydomonas native 3' UTR and re-test. Possible explanation 3 β Protein instability/degradation: insulin is a secreted human protein that requires specific folding and disulfide bond formation, normally done in the human endoplasmic reticulum. Chlamydomonas may lack the specific chaperones or disulfide bond-forming machinery to fold human insulin correctly. Misfolded protein is typically identified and degraded rapidly by the cell's proteasome. Test: add a proteasome inhibitor and check whether protein accumulates (indicating degradation was the issue). Also check for inclusion bodies (aggregates of misfolded protein) by electron microscopy. Solution: target the insulin gene to the chloroplast, which has a more reducing environment with different (sometimes more suitable) protein folding machinery for some recombinant proteins β a strategy that has worked for other human proteins expressed in algae.
5. You are advising an investor evaluating two algae biotech companies. Company A has a wild-type Haematococcus strain producing 3% astaxanthin by dry weight, grown in a 2-phase stress protocol. Company B has a CRISPR-engineered Nannochloropsis strain producing 1.5% astaxanthin grown continuously without stress. Which is commercially more interesting, and what additional information would you need to make a confident recommendation?
Company B's 1.5% is lower in absolute content than A's 3%, but the commercial picture is not determined by astaxanthin percentage alone β it is determined by cost per kilogram of astaxanthin produced. The key additional questions: (1) Productivity rate: Nannochloropsis grows much faster than Haematococcus (doubling time ~12h vs 48β72h). If B's faster growth rate produces more total biomass per day, the lower astaxanthin percentage may still yield more astaxanthin per pond per year than A's slow but rich culture. (2) Production cost per kg: the 2-phase stress protocol in Company A requires running two separate production stages (growth + stress), which adds operational complexity, time, and cost. Company B's continuous production is simpler to operate and scale. Ask for cost-of-goods-sold per kg of astaxanthin, not just yield percentage. (3) Regulatory status: is Company B's CRISPR-edited Nannochloropsis approved or approvable in target markets? Natural astaxanthin regulatory approval is well-established (EU, US, Japan). A GMO Nannochloropsis astaxanthin faces a longer, uncertain regulatory path in most markets β this could add years and millions to the timeline. (4) Stability of the engineered strain: does the CRISPR edit remain stable over many generations of continuous culture, or does the cell silencing or mutate the insert over time? This is a known challenge for transgenic algae in extended production. (5) IP position: does Company B have protected intellectual property on the CRISPR construct and production process? A weaker IP position means competitors could replicate the approach quickly, eroding any first-mover advantage. On balance, Company B's model is more intellectually compelling if the numbers work β but regulatory risk is the pivotal question that determines timeline to revenue.
Coming up β Week 13β14
The tree of life β where do algae fit?
Evolution, classification, and the surprising fact that "algae" is not a single group β it is dozens of unrelated lineages that independently evolved photosynthesis. Understanding this explains the extraordinary chemical diversity of algae products and why no single species dominates the entire industry.