Wednesday, March 15, 2017

Your next generation data storage solution: DNA



I recently stumbled upon the image on the right, which distills the changes in data storage in the past 40-plus years. Perhaps even more amazing to consider is that the future of data storage could become even smaller. The genetic material that stores all the information required to build a person or a pear or a penguin may be the key to creating even smaller data storage that never reaches obsolescence.

Our genome is often compared to a computer, where DNA is the code. In fact, DNA is a proven data storage system with billions of years of reliable use. While your old floppy disks may now be unreadable, the tools required to read and copy DNA are present in every genome, making it unlikely that we would lose the ability to decode DNA. These advantages led scientists to ask: could DNA also be used to store other types of data? Perhaps the information that would normally be encoded by 0's and 1's in your hard drive could be stored in sequences based on ACGT's.

The first publication to propose that DNA could be used for purposes other than building an organism comes in 1999 from Bancroft and colleagues in the journal Science. They suggest that genomic steganography could be a method for storing coded messages in DNA for use in espionage. Using a simple substitution cipher where each codon equals an alphanumeric value, the researchers synthesized a DNA sequence to encode the message "June 6 invasion: Normandy". The message was flanked by sequences to allow the recipient to decode the message. The final sequence of just 109 nucleotides of DNA was hidden within denatured human DNA and, just like the predecessor microdots used in espionage, embedded on top of a period in a typewritten message. 
Subsequent work from Bancroft's group and others in the early 00's suggested that DNA could help to address the need for increasing data storage. Computer scientists estimate that by 2020, there will be 4.4 x 1019 bytes (44 zettabytes) of digital data; to give you a sense of scale, 1 ZB would be about 152 million years of high resolution videoEven with the advances in storage potential, storing just 1 ZB requires more than 1000 kilograms of the cobalt alloy used to make hard drives. In contrast, 1 gram of DNA could store 4.6 x 1018 bytes. 

Early publications were proof of principal experiments that aimed to generate increasingly bigger data files encoded in DNA. The general approach, outlined above, convers a digital file to binary and to DNA. The beginnings were admittedly small, just as scientists had to sequence the genome of E. coli before they could complete the human genome. One problem is that DNA sequencing technology is improving at much faster rates than DNA synthesis techniques. Essentially, you could read the data you stored faster and cheaper than you could write it. Creating long accurate strands of DNA had technical and financial limitations. To circumvent this problem, George Church's lab used multiple copies of short DNA sequences to encode an entire book (53,246 words), 11 JPG images, and a JavaScript program. The paper, published in Science in 2012, also describes the recovery and reassembly process. The following year, a Nature paper from Ewan Birney's lab at the European Bioinformatics Institute reported a similar approach that increased the file size and decreased decoding errors. The final DNA file consisted of 739 KB of information, including text, pictures, videos, and audio files; they also added a PDF of the classic Nature paper from Watson and Crick describing the structure of DNA.   

In July 2016, researchers from the University of Washington collaborated with Microsoft to push the limits of DNA storage again (coverage in The Verge). Their storage reached 200 MB and included copies of the Universal Declaration of Human Rights, the top 100 books from Project Gutenberg, and the Crop Trust seed database; for fun, they encoded a video from the band OK Go for the song "This Too Shall Pass".

Most recently, a paper in Science from Yaniv Erlich and Dina Zielinski, who are working at the intersection of molecular biology and computer science, details a new storage architecture for more efficient DNA storage. They adapted fountain coding, which is currently used by streaming services like Netflix and Spotify to eliminate gaps in playback. The method greatly improved the storage density, getting closer to the theoretical limit for DNA storage (1.83 bits per nucleotide). Their DNA storage sample included the movie The Arrival of a Train, an entire computer operating system, a computer virus, and a Amazon gift card (which was quickly decoded by one of the researchers' Twitter followers). While the size of the data was smaller than previous attempts (only 2.2 MB), the method greatly improved data density and readability. One problem with previous storage methods is that reading the DNA leads to loss of the original sample. While it is easy to amplify DNA, it can sometimes introduce mistakes. Erlich and Zielinski's fountain technique permitted error-free amplification even after 10 complete reads. Their work achieved a density of 2.15 x 1018 bytes, which would allow storage of all the world's data in the trunk of a car.
Another stumbling block was that DNA was writable, but not re-writable, which limit the applications to archival data storage. Two recent papers (in Nature Communications and PNAS) report on a method that allows re-writing of DNA (bringing us from 8 track to cassette tapes) as well as reading from any point in the sample, rather than from a set starting spot (bringing us from cassette to CD).
 1 gram of DNA can store 4.5 x  1018 byte

While there has been
tremendous progress in increasing the amount and density of data storage, the major roadblock continues to be the amo
unt of time it takes to encode and decode data in DNA. Another place where inorganic data storage beat carbon-based products is in the cost, especially of synthesizing DNA. In the most recent paper, the cost was $3,500/MB, while the 2012 paper $12,400/MB.  

Despite these limitations, biologists are teaming up with computer scientists to explore the future of DNA data storage. This is largely driven by the need to store increasing amounts of digital data with decreasing resources. Estimates indicate that by 2040 global memory demand (3 x 1024 bytes) will exceed the supply of silicon necessary to build traditional data storage devices.

Obsolescence is another shortcoming of current storage methods. Just as it has become difficult to play your cassette tape collection (much to my chagrin), your old floppies and ZIP disks are not readable either. Scientists conjecture that because DNA is the basis for life on Earth, we will always have methods for DNA sequencing. This gives DNA a huge advantage for long-term archival storage. Luckily, DNA also has great fidelity over the long term. Scientists are increasingly able to recover readable sequences from ancient samples of DNA with the best results coming from samples stored at low temperature. Thus, you could imagine a long-term storage system, like a secure server in a remote tundra, where the DNA back up disk to re-start civilization would be stable and safe. 

This isn't completely crazy. The Svalbard Global Seed Vault is a huge storage site in the frozen tundra of Norway where scientists and governments are making contributions of plant seeds. The idea is to keep a stock of the original seed in case of the collapse of civilization. I am sure we could rent a shoe box-sized space there for storing all the relevant files from humankind (that means there probably won't be room for cat videos). It is certain that resource limitations will continue to make digital DNA storage, borne of a thought experiment over beer, not just a reality but a necessity.
   
References

Scientific American, Tech Turns to Biology as Data Storage Needs Explode
George Church interview in Popular Science
Ed Yong covered the DNA fountain technique in The Atlantic


Sunday, February 19, 2017

The great asparagus experiment: investigating the genetic basis for asparagus anosmia

After the great success of the cilantro experiment, we decided to undertake a similar approach with asparagus. As you probably know, some people find that asparagus makes their urine smell funny, while others experience no effect. In fact, observations on the effect of asparagus have been made as early as the 1700s. Ben Franklin included some negative comments about asparagus ("A few Stems of Asparagus eaten, shall give our Urine a disagreable Odour") in his essay on flatulence (I think my son's budding interest in Franklin will increase greatly knowing that Ben was equally interested in lightning and farts!)

In contrast to the cilantro experiment, where I dislike the taste of the test food, I have no idea what the big deal is with asparagus, but my husband tells me of the foul smell that is the result of enjoying this verdant spring vegetable. Depending on the study, 33-60% of people cannot detect the smell. In our experiment, we all enjoyed a side of asparagus that was lightly pan seared with olive oil. Our son was not too excited by the taste of the vegetable, but he ate it for the sake of science. Frankly, I think he was curious what the outcome would be. After our next bladder evacuations, we compared notes on the smell. It seems my son has inherited his father's perception of asparagus smell. Unfortunately for us, this fun experiment did not encourage our boy to eat more asparagus. Rather, he can now use the smell as an excuse to avoid it.

Next came the moment in the experiment to explain what happened and why. The first question is what it is about asparagus that makes urine smell different. Luckily, that one is firmly based in chemistry, so it is relatively straightforward to answer. The odor is the result of the metabolism of a chemical unique to asparagus: asparagusic acid. The infographic below from Compound Interest on the Chemistry of Asparagus includes the specifics on the chemical structures if you are curious. Asparagusic acid breaks down into four other sulfur-containing compounds, which happen to be volatile, evaporating quite readily and giving up their pungent odors in the process. Methanethiol and dimethyl sulfide are thought to be the culprits for the smell; this was determined by giving people purified forms of either compound, which can induce the smell without eating asparagus. It is thought that asparagusic acid helps asparagus keep pests away in the wild; the compound can prevent the growth of fungi as well as parasitic nematodes; consistent with this, the concentration of asparagusic acid is highest in the emerging shoots and other parts of the plant that are likely to be infected.






















The second question is why only some people experience the smell while others don't. The answer to this question is based in genetics and sense perception, so it is a bit more complicated. As with so many phenotypes, there is not one simple genetic explanation. There are two main physiological issues here: the metabolism of the asparagus and the perception of the smell of those metabolic byproducts. For some time, scientists thought that everyone was an excreter (meaning they could produce the smell in their urine), but that some people could not perceive the smell (the scientific term is anosmia). As more scientists started to ask people about their experience with asparagus, the story became murkier. There is variation in how people perceive the smell (from strong to mild) and there is also a small population of people that that do not metabolize asparagus the same way. 

There have been a number of papers looking at the genetic connection between anosmia and asparagus. A 2010 PLOS Genetics paper published by the personal genomics service 23andme seems to be the first paper linking asparagus anosmia to a particular genetic variation. The paper showed that several single nucleotide polymorphisms (SNPs) occurring near genes encoding for odor receptors were correlated with asparagusic acid anosmia. The most significant association was seen in a gene encoding olfactory receptor 2 (OR2M7), where the variation acted in a dominant fashion to decrease the likelihood of anosmia.

genome-wide association study (GWAS) published in the journal BMJ in 2016 looked at a group of nearly 7000 Europeans. Their results suggest that a majority of this group (60%) had anosmia, with it being slightly more common in women than in men. Consistent with the results from 23andme, they found several SNPs in the olfactory receptor 2 gene. Thus, there is good evidence for the genetic basis for asparagus anosmia.What is still unknown is what SNPs (if any) are associated with the inability to excrete aspargusic acid. In addition, it is unclear if the ability to detect this smell has any evolutionary basis or if it is just a random mutation.

These experiments taught me a lot about the science of food and taste perception. There are lots of opportunities to experiments with food in the future, While most of my previous experiments with food involve baking, there are lots of opportunities to experiment with food in the future. For example, this post from Discover Blogs talks about the large variety in how humans perceive tastes and the opportunity for citizen science to unravel that. I will be sure to post an update about our next foray into human genetics and food.

Sunday, January 8, 2017

My data, myself: My fitness tracker helps motivate me to exercise, but clinical research suggests your mileage may vary

It is the beginning of a new year, which for me means calculating my yearly mileage and thinking about next year's goals. I have been obsessed with tracking my performance in fitness since I started running in 2000. It started in a very low-tech way, writing down my mileage each day in a calendar and then adding it up each month and then at the end of the year. With the advent of the app MapMyRun, I became even more enthusiastic about tracking my running, which helped me to run more consistently.

Last Christmas, I got a new Garmin VivoActive smart watch, which has really changed the way I look at my activity levels. Like most smart watches, this one can track my exercise, steps, and sleep cycles. It also buzzes me if I have been inactive for too long. This is particularly useful when you have an office job. In a completely unsurprising Pavlovian response, I have now started to anticipate when the watch will buzz, so I get up and take a few laps around the office until the "move bar cleared" buzz comes in. The watch also gives me a little fireworks display when I reach my daily step goal (or multiples thereof, which are particularly edifying). When I sync my watch, the Garmin website will analyze the data, so I can see how I am doing for the week, month, or year. To some people, this may sound exhausting, but as a runner, bike commuter, and science nerd, I love all the charts and graphs.
2016 distance totals
I always thought that my obsession with my fitness data was some vestigial interest in analyzing data due to my background in science, but it seems that many people find fitness tracking to be motivating and fun. My data analysis has been fairly simple; I have only looked at trends over time (e.g., my running increased during the stress of grad school and decreased around the time of my pregnancy). It's amazing to see how other people have quantified their lives so thoroughly. Quantified Self has some great examples of people data mining their own lives to understand and improve themselves, one woman has analyzed more than two decades worth of data about her fitness, diet, and weight.
Based on my experience, I would think that fitness tracking apps and wearables can help people to stay motivated. However, the results of some of the published clinical trials suggest that it isn't clear if fitness trackers can help everyone achieve better health (as measured by activity levels, weight, and metabolic variables). There are currently 21 clinical trials using fitness trackers registered on Clinical Trials.gov, so more information about their effectiveness should be available soon. In September 2016, a study published in The Journal of the American Medical Association received a lot of buzz for reporting the results of a trial that compared the weight loss with the standard regimen (i.e., calorie restriction and physical activity) versus the standard regimen plus fitness tracking. They found that the fitness tracking group lost less weight than the control group, suggesting that wearables may not help people lose more weight, especially in the long term (excellent coverage of this study in STAT). There are some variables that can make activity trackers improve fitness; one study showed that cash incentives led to some increase in activity levels. However, this increased activity did not continue once the cash rewards were gone. An earlier study, sponsored by Fitbit and using an opt in approach to recruit people for the wearable group, found that people who use Fitbits for more than a year have larger decreases in insurance costs than non-users. Another randomized control study in postmenopausal women confirmed that the use of the Fitbit increased activity levels, but they did not collect data on weight loss or fitness levels. A major problem with fitness trackers is abandonment: about a third of people stop wearing their fitness tracker after 6 months.
Distance totals 1999-2016

These results suggest that, for most people, wearable devices will not encourage them to exercise or lose weight. Indeed, I would expect that it would be a very particular personality type that would be motivated by a watch's fireworks display or an online competition. It remains to be seen if people who already have a strong exercise pattern are more motivated by fitness apps and trackers. The challenge for fitness tracker designers and doctors alike is how to motivate people to make changes for their health.

Perhaps more interesting is the potential for the usage of wearables in the future. According to Kat Arney in Herding Hemingway's Cats and here in shorter form for BBC science, the data from our smart watches could ultimately be combined with the data from our personal genome sequences: "Combining the power of modern genetic analysis with bio-monitoring" could improve care, save lives, and revolutionize genetics. Larry Smarr, profiled in The Guardian, has been tracking his life for 15 years following more than 150 variables. Such complex bio-monitoring generates a lot of data and turning that data into usable information is the challenge. This is one reason why tech companies are getting into the health data market. A recent editorial from Nature warns that there could be unforeseen problems with sharing our health data, particularly in terms of personal privacy. In addition, they speculate that it could widen existing inequities and biases. In addition, it is not yet clear what benefit (if any) these data-driven approaches will have.

I don't expect to abandon my GPS watch anytime soon. Instead, I imagine the future me using a multi-function tracker that combines data from my metabolic profile with my genomic data to make recommendations on exercise, fitness goals, and vitamins. I may be living in a science fiction fantasy, but 15 years ago when I started tracking my miles with gmaps pedometer and a calendar, I would never think I would have a watch that could do all that for me.