This week marks the 10-year anniversary of the publications of a (nearly) completed human genome sequence. Much has been made already of this passage of time, as well as what we can look forward to in the next ten years.
What I thought I would do in this space is share a little personal story on my connection with this achievement. In late spring of 2001, I happened to search the Internet and PubMed for my name because I wanted to check to see if any presentations at conference or publications from previous laboratories in which I had worked had been released. To my surprise, I found a website in Japan with the title of something like "list of authors" which contained a collection of names of former colleagues from my days in the Genome Sequencing Center at Cold Spring Harbor Laboratory. That seemed strange and so investigating a bit I learned that we were included on the Nature paper describing the human genome - along with some 5000+ other authors (hence the special listing on this website, and no hits in PubMed). Well, needless to say but that was quite a thrill. I quickly updated my CV to include this landmark publication.
Back in 1997 to 1999, as the publicly funded project to sequence the human genome was ramping up and dollars were dangled in front of genome centers around the USA and the globe, we at CSHL were trying to deposit as much finished sequence into GenBank as possible. Monthly and quarterly totals of base pairs deposited were key to securing grant money. An introduction to all this came within my first two weeks as the Computational Fellow (post-doc) with Dick McCombie when I was told I would be leading the analysis segment of his Genome Sequencing course. I learned the ins and outs of a new computer system and new software tools (I came from a cell biology lab) just in time to teach the students. We worked hard during that 2-week course to sequence a 143-kbp BAC clone containing some critical HIV/AIDS-relevant genes: CCR2, CCR5 and CCR6. You can view the sequence entry I deposited to GenBank here, accession U95626.
From this initial BAC, we worked on many more to try to show that we could put high-quality sequence data together and to get as much sequence finished as possible. Of course, our main funding was to contribute to the Arabidopsis thaliana genome and so the human projects (BACs and cosmid/fosmid clones) took second priority. But we did contribute enough sequence to warrant inclusion on the paper and Dick was kind enough to remember everyone who had passed through his lab during those years.
Showing posts with label human genome. Show all posts
Showing posts with label human genome. Show all posts
Friday, February 18, 2011
Wednesday, June 16, 2010
Genome compexity and the number of genes
This month marks the 10th anniversary of (one of) the announcement(s) of the completion of the human genome. Several have taken this occasion to comment on the successes of genome-based biomedical research, or lack thereof.
At "The Loom," Carl Zimmer has a neat graphic depicting estimates of the number of genes present in the human genome. This number has, more or less, steadily fallen as progress in sequencing and then in filling in remaining gaps in the reference genome sequence has moved forward. My comments to that blog entry are:
The reduction in the estimates of the number of protein-coding genes in the human genome parallels our increased understanding of the complexities involved in regulation of gene activity. For example, many types of non-coding RNAs have been described as well as their roles in modulating the information flow from DNA to protein.
At the same time, I believe that the human genome’s reduced “tool kit” (in terms of number of protein-coding genes) shows a certain level of our genome’s sophistication. Think of the many different ways one can use a screwdriver – say to open a can of paint, or its handle as a hammer. In other words, different proteins can join to different networks in a tissue- or developmental-specific manner. In conjunction with this are the alternatively spliced mRNAs, which often lead to different protein isoforms (proteins that are mostly the same, but with perhaps one different functional subdomain). Think of a Phillips vs. regular screwdriver.
Thus, fewer genes has not meant there are fewer protein isoforms nor less complex protein-protein or protein-small molecule interaction networks. To the contrary, there is an increased complexity and that is one reason it has been difficult to define all the players in a particular human affliction such as type 2 diabetes or cancer.
Here, I would like to add a couple of other points:
1. Genetic variation, whether common, moderately rare or even unique to an individual or family, no doubt has a role in adding to the complexity of interactions among the (relatively) small number of genes and small molecules. For example, our research is considering transcription factor binding sites and seed sites for mRNA-mRNA interactions that are created by minor alleles of SNPs.
2. Where much of the above leads is toward differential pathway dynamics. Because the number of protein-coding genes is low while the human organism, its response to a number of different situations (consider how long tobacco smoking or a poor diet must be endured, on average, before life-threatening phenotypes emerge), and its great capacity to develop, survive and even thrive with numerous genetic aberrations are all complex, many of the answers we research seek simply remain to be discovered. We just do not know all the players - proteins, RNAs, genome state (e.g., methylation) and small molecules - and so cannot fully describe a type 2 diabetes pathway in a series of affected tissues, or a given cancer for that matter. Progress is being made - sequencing of the genomes of a tumor and healthy tissue from the same individual have uncovered common pathways and perhaps drug targets. There, however, remains much more to describe before the full potential of human genomics research will be realized.
At "The Loom," Carl Zimmer has a neat graphic depicting estimates of the number of genes present in the human genome. This number has, more or less, steadily fallen as progress in sequencing and then in filling in remaining gaps in the reference genome sequence has moved forward. My comments to that blog entry are:
The reduction in the estimates of the number of protein-coding genes in the human genome parallels our increased understanding of the complexities involved in regulation of gene activity. For example, many types of non-coding RNAs have been described as well as their roles in modulating the information flow from DNA to protein.
At the same time, I believe that the human genome’s reduced “tool kit” (in terms of number of protein-coding genes) shows a certain level of our genome’s sophistication. Think of the many different ways one can use a screwdriver – say to open a can of paint, or its handle as a hammer. In other words, different proteins can join to different networks in a tissue- or developmental-specific manner. In conjunction with this are the alternatively spliced mRNAs, which often lead to different protein isoforms (proteins that are mostly the same, but with perhaps one different functional subdomain). Think of a Phillips vs. regular screwdriver.
Thus, fewer genes has not meant there are fewer protein isoforms nor less complex protein-protein or protein-small molecule interaction networks. To the contrary, there is an increased complexity and that is one reason it has been difficult to define all the players in a particular human affliction such as type 2 diabetes or cancer.
Here, I would like to add a couple of other points:
1. Genetic variation, whether common, moderately rare or even unique to an individual or family, no doubt has a role in adding to the complexity of interactions among the (relatively) small number of genes and small molecules. For example, our research is considering transcription factor binding sites and seed sites for mRNA-mRNA interactions that are created by minor alleles of SNPs.
2. Where much of the above leads is toward differential pathway dynamics. Because the number of protein-coding genes is low while the human organism, its response to a number of different situations (consider how long tobacco smoking or a poor diet must be endured, on average, before life-threatening phenotypes emerge), and its great capacity to develop, survive and even thrive with numerous genetic aberrations are all complex, many of the answers we research seek simply remain to be discovered. We just do not know all the players - proteins, RNAs, genome state (e.g., methylation) and small molecules - and so cannot fully describe a type 2 diabetes pathway in a series of affected tissues, or a given cancer for that matter. Progress is being made - sequencing of the genomes of a tumor and healthy tissue from the same individual have uncovered common pathways and perhaps drug targets. There, however, remains much more to describe before the full potential of human genomics research will be realized.
Subscribe to:
Posts (Atom)