Showing posts with label genome sequencing. Show all posts
Showing posts with label genome sequencing. Show all posts

Monday, December 4, 2017

How to Tame a Fox (and Build a Dog): a long term experiment and a story of perseverance and hope

The story of Dmitry Belaraev's long-term experiment in fox domestication was beautifully covered by Scientific American and The Discover Blogs as well as on Radiolab. I found the story and the science behind it fascinating, but I wasn't sure a book was needed to delve further in the story. After I read How to Tame a Fox (and Build a Dog), I learned how wrong I was.

Lee Alan Dugatkin co-authored the book with one of the original researchers, Lyudmila Trut. The authors are able to tell the story of Belaraev's seemingly Quixotic plan to tame foxes in a way that is compelling, even when I knew the major plot points.

The project's hypothesis was simple: continuous breeding of the tamest/gentlest foxes would lead to foxes that behaved like dogs. I found it amazing that they were so committed to scientific rigor that they ran a parallel experiment where they chose the most aggressive foxes. At the time, Belaraev faced obstacles from the ruling Communist party. Today, scientists would also question funding such a long-term experiment. Most geneticists use animals with short lifespans (e.g. fruit flies, C. elegans, yeast) so they can get results quickly and at lower costs. The experiment started producing results far earlier than even the principal investigators would have guessed. Even after a few generations, they found that the foxes showed more dog-like traits (floppy ears, tails, and other classic signs of neotony.) After 10 generations, they found that females went into estrus earlier and some male pups showed changes in fur color.

These changes were all predicted by Belaraev, who had a pretty revolutionary idea at the time, which he called "destabilizing selection". He suggested that the changes that occur in the course of domestication aren't simply due to the accumulation of mutations, but rather changes in the expression of existing genes.

Graphical Abstract from Parker et al Cell Reports 2017
Once genome sequencing arrived on the scene, the domesticated foxes were soon the subject of several studies. The first approach was to compare 700 genetic markers that had been used in the initial dog genome sequence with both the wild and domesticated foxes.  The researchers found several cases of convergence between the two domestication events, which included genetics variations that would ultimately lead to changes in the appearance and behavior of the tame animals. In 2010, a paper in Nature reported the genomic changes that accompanied the domestication of dogs from wolves.  Of course, the research is still in progress. One recent paper in Science suggests that dogs were domesticated in Eurasia and Eastern Asia from 14000 to 6000 years ago. The most recent data (published in my journal Cell Reports) suggest that different breeds of dogs had origins in different geographical locations. Importantly, the work published so far suggests that Belaraev was right: changes in gene expression not just mutations were key to domestication.

Perhaps the most fascinating parts of the book were the details concerning the barriers to science in Soviet-era Russia. These included the domination of genetics by a non-scientist named Trofim Lysenko, who was staunchly opposed to the ideas of Darwin and Mendel; these ideas were gaining acceptance at the time in the West. Lysenko even convinced Lenin that putting seeds in the cold would make crops grow better at low temperatures (they don't). Lysenko fabricated data to support his ideas and those of the party, which devastated Soviet agriculture for decades. Lysenko was celebrated by the Communist party simply for being a peasant and for going against "bourgeois Western science".  Interestingly, Khrushchev's daughter Rada, a journalist who trained as a biologist, argued against Lysenkoism and fought to protect science and the work of Belaraev. I think this serves as another example of why it's best to keep politics out of science (but not science out of politics!) The stories of the suppression of science in the height of soviet Russia are evocative of current anti-science rhetoric in these United States.

One of the reasons that Belaraev was able to persevere in the face of such odds was his enthusiasm and charisma. Belaraev was a social chameleon, who could easily adapt to his audience. These traits made people gravitate towards him and helped him argue for the benefits of the fox experiment to the powers that be. Officially, the rationale of the experiment was to breed foxes that would have more variety in their colors, which would help fetch higher prices for their fur. The domestication experiment was funded despite the crackdown on work in genetics. This isolationism crippled Russian scientists, who were cut off from the science in the rest of the world.

In the late 1980s the fox experiment was 30 years old. The farm started to have some trouble with funding and the researchers had to scale back the experiment. The situation was made worse in 1998 when the Russian economy collapsed. The scientists started sacrificing some of the foxes and selling their pelts. Shortly after this, they contacted a few select media outlets to cover the story, which helped them secure some funding for the project. Today, the experiment has been going on for 60 years, which is a long time for a laboratory experiment, but a short time frame for evolution.

How to Tame a Fox (and Build a Dog) was a fascinating and ultimately hopeful story. Even in such a difficult political climate, this scientist and his visionary experiment were able to persevere with some patience and ingenuity, a message that I think will resonate with scientists that are struggling to stay afloat in 2017.

Wednesday, March 15, 2017

Your next generation data storage solution: DNA



I recently stumbled upon the image on the right, which distills the changes in data storage in the past 40-plus years. Perhaps even more amazing to consider is that the future of data storage could become even smaller. The genetic material that stores all the information required to build a person or a pear or a penguin may be the key to creating even smaller data storage that never reaches obsolescence.

Our genome is often compared to a computer, where DNA is the code. In fact, DNA is a proven data storage system with billions of years of reliable use. While your old floppy disks may now be unreadable, the tools required to read and copy DNA are present in every genome, making it unlikely that we would lose the ability to decode DNA. These advantages led scientists to ask: could DNA also be used to store other types of data? Perhaps the information that would normally be encoded by 0's and 1's in your hard drive could be stored in sequences based on ACGT's.

The first publication to propose that DNA could be used for purposes other than building an organism comes in 1999 from Bancroft and colleagues in the journal Science. They suggest that genomic steganography could be a method for storing coded messages in DNA for use in espionage. Using a simple substitution cipher where each codon equals an alphanumeric value, the researchers synthesized a DNA sequence to encode the message "June 6 invasion: Normandy". The message was flanked by sequences to allow the recipient to decode the message. The final sequence of just 109 nucleotides of DNA was hidden within denatured human DNA and, just like the predecessor microdots used in espionage, embedded on top of a period in a typewritten message. 
Subsequent work from Bancroft's group and others in the early 00's suggested that DNA could help to address the need for increasing data storage. Computer scientists estimate that by 2020, there will be 4.4 x 1019 bytes (44 zettabytes) of digital data; to give you a sense of scale, 1 ZB would be about 152 million years of high resolution videoEven with the advances in storage potential, storing just 1 ZB requires more than 1000 kilograms of the cobalt alloy used to make hard drives. In contrast, 1 gram of DNA could store 4.6 x 1018 bytes. 

Early publications were proof of principal experiments that aimed to generate increasingly bigger data files encoded in DNA. The general approach, outlined above, convers a digital file to binary and to DNA. The beginnings were admittedly small, just as scientists had to sequence the genome of E. coli before they could complete the human genome. One problem is that DNA sequencing technology is improving at much faster rates than DNA synthesis techniques. Essentially, you could read the data you stored faster and cheaper than you could write it. Creating long accurate strands of DNA had technical and financial limitations. To circumvent this problem, George Church's lab used multiple copies of short DNA sequences to encode an entire book (53,246 words), 11 JPG images, and a JavaScript program. The paper, published in Science in 2012, also describes the recovery and reassembly process. The following year, a Nature paper from Ewan Birney's lab at the European Bioinformatics Institute reported a similar approach that increased the file size and decreased decoding errors. The final DNA file consisted of 739 KB of information, including text, pictures, videos, and audio files; they also added a PDF of the classic Nature paper from Watson and Crick describing the structure of DNA.   

In July 2016, researchers from the University of Washington collaborated with Microsoft to push the limits of DNA storage again (coverage in The Verge). Their storage reached 200 MB and included copies of the Universal Declaration of Human Rights, the top 100 books from Project Gutenberg, and the Crop Trust seed database; for fun, they encoded a video from the band OK Go for the song "This Too Shall Pass".

Most recently, a paper in Science from Yaniv Erlich and Dina Zielinski, who are working at the intersection of molecular biology and computer science, details a new storage architecture for more efficient DNA storage. They adapted fountain coding, which is currently used by streaming services like Netflix and Spotify to eliminate gaps in playback. The method greatly improved the storage density, getting closer to the theoretical limit for DNA storage (1.83 bits per nucleotide). Their DNA storage sample included the movie The Arrival of a Train, an entire computer operating system, a computer virus, and a Amazon gift card (which was quickly decoded by one of the researchers' Twitter followers). While the size of the data was smaller than previous attempts (only 2.2 MB), the method greatly improved data density and readability. One problem with previous storage methods is that reading the DNA leads to loss of the original sample. While it is easy to amplify DNA, it can sometimes introduce mistakes. Erlich and Zielinski's fountain technique permitted error-free amplification even after 10 complete reads. Their work achieved a density of 2.15 x 1018 bytes, which would allow storage of all the world's data in the trunk of a car.
Another stumbling block was that DNA was writable, but not re-writable, which limit the applications to archival data storage. Two recent papers (in Nature Communications and PNAS) report on a method that allows re-writing of DNA (bringing us from 8 track to cassette tapes) as well as reading from any point in the sample, rather than from a set starting spot (bringing us from cassette to CD).
 1 gram of DNA can store 4.5 x  1018 byte

While there has been
tremendous progress in increasing the amount and density of data storage, the major roadblock continues to be the amo
unt of time it takes to encode and decode data in DNA. Another place where inorganic data storage beat carbon-based products is in the cost, especially of synthesizing DNA. In the most recent paper, the cost was $3,500/MB, while the 2012 paper $12,400/MB.  

Despite these limitations, biologists are teaming up with computer scientists to explore the future of DNA data storage. This is largely driven by the need to store increasing amounts of digital data with decreasing resources. Estimates indicate that by 2040 global memory demand (3 x 1024 bytes) will exceed the supply of silicon necessary to build traditional data storage devices.

Obsolescence is another shortcoming of current storage methods. Just as it has become difficult to play your cassette tape collection (much to my chagrin), your old floppies and ZIP disks are not readable either. Scientists conjecture that because DNA is the basis for life on Earth, we will always have methods for DNA sequencing. This gives DNA a huge advantage for long-term archival storage. Luckily, DNA also has great fidelity over the long term. Scientists are increasingly able to recover readable sequences from ancient samples of DNA with the best results coming from samples stored at low temperature. Thus, you could imagine a long-term storage system, like a secure server in a remote tundra, where the DNA back up disk to re-start civilization would be stable and safe. 

This isn't completely crazy. The Svalbard Global Seed Vault is a huge storage site in the frozen tundra of Norway where scientists and governments are making contributions of plant seeds. The idea is to keep a stock of the original seed in case of the collapse of civilization. I am sure we could rent a shoe box-sized space there for storing all the relevant files from humankind (that means there probably won't be room for cat videos). It is certain that resource limitations will continue to make digital DNA storage, borne of a thought experiment over beer, not just a reality but a necessity.
   
References

Scientific American, Tech Turns to Biology as Data Storage Needs Explode
George Church interview in Popular Science
Ed Yong covered the DNA fountain technique in The Atlantic