Showing posts with label USS. Show all posts
Showing posts with label USS. Show all posts

Wednesday, November 18, 2009

Saturation and Chaser


EDITED: Used bad mol per cell calculation earlier.

How much DNA does an average cell in a competent culture take up? I did two experiments today, which were aimed at trying to decide how much degeneracy we can tolerate in our degenerate USS experiment. There's several issues, but one simple one is just to know how much of an "optimal" USS is taken up in the presence of competing "suboptimal" USS...

In the first experiment, I added different amounts of hot USS-1 DNA (the new-fangled ones designed for Illumina sequencing... ~200 bp) to a fixed amount of competent cells to find out how much DNA was enough to saturate the cells with DNA. Here's the results for several amounts of USS-1 incubated with 200 ul of competent cells for 20 minutes:
Looks like less than 8 ng (or ~200 molecule per cell) is sub-saturating and greater than 8 ng is near saturation.

In the second experiment, I looked at how well USS-1 is taken up in the presence of a competitor mutant USS (USS-V6, which is taken up 10X worse than USS-1). Based on the results of the first experiment, I used two different amounts of hot USS-1, either sub-saturating (4ng) or saturating (16ng). For each of these, I had four samples in which different amounts of cold USS-V6 were added. Here's the results for subsaturating USS-1:

And now for saturating USS-1:

Though adding the competitor did reduce uptake, uptake decreased only by ~1/4, even when there was several-fold more competitor. Also, it appears that using saturating amounts of USS-1 yields a different decline in uptake with excess inhibitor than when USS-1 is sub-saturating.

So what's this all mean? I think it means that even when "optimal" USS are only 1/10 as common as "suboptimal" USS, uptake proceeds quite nicely. What does this mean practically for making our fancy oligo purchase? Not too much, except that I think we can err on the side of more degeneracy without too many problems. (This will be good for several reasons, probably the most important of which is that our preliminary data collection plans will work best with more degeneracy.)
(continued...)

Tuesday, November 17, 2009

USS uptake with Illumina adaptors

(I'm somehow deeply amused by the above image. I think it has something to do with the arrow going from the DNA molecule to the sun.)

I've finally gotten around to doing some uptake experiments with the control sequences I got for our degenerate USS experiments. The important thing was to make sure that the constructs would behave as our older ones do with our new-fangled design that will allow for Illumina sequencing directly from the purified periplasmic DNA (i.e. no library construction steps)...

And it worked quite well! I also scaled these experiments down, so that I could more easily do some saturation curves tomorrow.

Results today for uptake of 8 ng of 200 bp DNA (~160 million molecules) by ~200 million competent cells (in 200 ul):

O.G.-USS1 -> 33.5%
New-USS1 -> 35.4%
New-USSV6 -> 3.7%
New-USSR -> 0.7%

Right on target!
(continued...)

Wednesday, November 4, 2009

Decompression


Whew! Rosie and I just made what I consider a heroic effort to produce a grant application to Genome BC to use DNA sequencing to measure recombination biases during H. influenzae natural transformation.

It was heroic not only because we finally decided to apply only late last week, but because our co-funding support is tenuous at best. Genome BC requires that we match their funds at least equally with funds from another source. Rosie has funding, but it was applied for too long ago and the proposal only indirectly relates to our planned sequencing. We also applied for a CIHR grant recently, but will not have reviews until after the Genome BC committee meets in early January. The best hope of adequate co-funding comes from my NIH postdoctoral fellowship grant application (a resubmission) late this summer, for which I should have scores (or lack thereof) within a couple of weeks. If the application gets a good score, we can tell Genome BC that the major threat to the success of our application is ameliorated.

Almost immediately after submitting the grant application with Rosie last night, I had to turn to editing my buddy's manuscript (which I am an author on), which takes on the weighty topic of detecting rearrangements in complex mammalian genomes from limited sequencing data. I just turned my edits over to the corresponding author, and now, after letting the excess nitrogen out of my bloodstream, I need to decide what to do next.

Since aforementioned buddy also has taken several DNA samples off my hands for sequencing, I think I'd best turn to purifying uptake DNA from the periplasm of competent cells. I've already gotten things fairly well under way (see here, here, and here), but there's a bunch of uptake experiments waiting to be done. So tonight, I'll inoculate some cultures, so I can try a large-scale periplasmic DNA prep tomorrow, and tomorrow I'll also order some more radiolabel for doing more sensitive uptake experiments.
(continued...)

Thursday, October 1, 2009

Corrected Logos

Yesterday, I attempted to make some logos out of the degenerate USS simulated data that Rosie sent me. Turns out, I was doing it wrong. After checking out the original logo paper, I was able to figure out how to make my own logos in Excel. I wasn't supposed to plot the "information content" of each base; I was supposed to take the total information content (in bits) of each position (as determined by the equation in the last post) and then, for each base, multiply that amount by the observed frequency of that base to get the height of that element in the logo. So, below I put the proper logos for the selected and unselected degenerate USS sets (background corrected):




Here's the selected set:
Here's the unselected set:
Woo!
(continued...)

Wednesday, September 30, 2009

Struggling with the Background

Previously, I had shown some preliminary analysis of Rosie’s simulated uptake data of chromosomal DNA fragments. Rosie also sent me simulated uptake data of degenerate USS sequences (using a 12% degeneracy per position in this USS consensus sequence :

5’-AAAGTGCGGTCAAATTTCAGTCAATTTTT-3’).

So what can I do with this dataset? Well, first, since Rosie also provided me with the “scores” for each sequence, I could plot a histogram of the scores for the 100 selected and 100 unselected sequences, showing that the uptake algorithm seems to work pretty well.

Here's the histogram:
Notably, even the unselected sequences have rather high scores, when compared to the same analysis of genomic DNA fragments. This is unsurprising, since the sequences under selection in this degenerate USS simulation are all rather close to the consensus USS.

Here’s the histogram from the genomic uptake simulation again (just to compare):
(I think the reason for the difference in the “selected” distributions is due to a different level of stringency when Rosie produced the two simulated datasets.)

That’s all fine and good, but now what? With the genomic dataset, I could use the UCSC genome browser to plot the location of all the fragments I was sent, but this consists of two alignment blocks of sequence that look markedly similar.

The obvious thing to do was to make Weblogos of the two different datasets… The unselected set should have very little information in it, while the selected set should contain information. In doing this, I discovered a rather important issue… The on-line version of Weblogo does NOT, I repeat, does NOT account for the background distribution.

This is a problem. It means that whenever you make a Weblogo (on the webserver) from your alignment block, it is assuming that each base is equally likely to occur at a random position. This is why the y-axis in all Weblogos plots always has a maximum of 2 bits when using DNA sequence. Why is this a problem? First of all, if one is using an AT-rich genome, as we are, then the information content of any G or C is underestimated and any A or T is overestimated.

So how does Weblogo calculate the information content of each position in an alignment block? From the Weblogo paper (link found at the Weblogo website):
Rseq = the information at a particular position in the alignment
Smax = the maximum possible entropy
Sobs = the entropy of the observed distribution of symbols (bases)
N = the number of distinct symbols (4 for DNA)
pn = the observed frequency of symbol n

The log2 is there to put everything in terms of bits.

So for DNA (4 bases), the maximum entropy at a position is 2 bits. Makes perfect sense: 2 bits a base. However, this only makes sense if each base is equally probable for a randomly drawn sequence. Now for purposes of gaining an intuition for different motifs, this isn’t really a big deal, although it does complicate comparing motifs between genomes.

When this isn’t the case (probably much of the time), then different measures have been used, namely the “relative entropy” of a position. This is an odds ratio of the observed probability and the background probability. Apparently, the off-line version of Weblogo can account for non-uniform base composition, but I haven’t tried installing it yet, nor any other software out there that handles variation in GC content.

Why? Because what we need for our degenerate sequences is a different background distribution at each position! So, the first position in the core is 88% A, but the third position is 88% G!

To illustrate the problem, here is a Weblogo of the selected set:
Here’s the unselected set:
Looking closely, it is clear that there are differences in the amount of “information” at each position. So in the strong consensus positions of the USS, the selected set has higher “information” than the unselected set, while at weak consensus positions, that’s less true.

But the scaling of each base here is completely wrong. There isn’t nearly a bit of information at the first position in the unselected set. We expected 88% A. The fact that there are mostly As in the alignment block at the first position is NOT informative. In fact, if all was well, we’d get zero bits at all the unselected positions!

What to do? I tried to make my own logo, using the known true background distribution at each position. I won’t belabor the details too much at the moment, except to say that I had to figure out what a “pseudocount” was and how to incorporate it into the weight matrix, so as to not ever take the logarithm of zero.

Here’s the selected set of 100 sequences:
Here’s the unselected set of 100 sequences:
(Note that I somehow lost the first position when I did this, so the motif starts at the first position of the core.)

This actually looks quite a bit better, or at least more sensible. A few things worth noting:
  • I think if we did thousands of unselected sequences, we’d pretty much get zero information from that alignment, which is what we would want, since that’s just the background distribution.
  • Some values are negative. This is expected. Since these are scaled to log-odds ratios, when the frequency of seeing a certain base is less than the expected background frequency, a negative number emerges.
  • The scale is extremely reduced. Every position is worth less than 0.3 bits. This is also expected. One description I’ve seen about how information content can be thought about is how “surprised” one should be when making an observation (there’s even a unit of measure called a “suprisal”!). Since we are drawing from an extremely non-uniform distribution that actually favors the base that’s expected to be taken up by cells better, we are basically squashing our surprise way down. That is, getting an A at the first position of the core is highly favored, but it’s the most likely base to get anyways, even in the absence of selection.
  • The unimportant bases in the USS have the most information content in the selected set. At first this bothered me, but then I realized it was utterly expected for the same reason as above. For example, at position #18 above (sorry, it’s position 19 in the Weblogos), the selection algorithm doesn’t really care what base is there. That means that the selected set will let mutations at that position (from A to something else) come through, which will be surprising, when compared to the background distribution!
(ADDED LATER: Actually, this last point is wrong. The reason for so much information at the weak positions is related to the matrix that was used to select the sequences, not from surprise. I'll try and get a proper dataset later and redo this analysis. To some extent, the positions will still have some information, as partially explained in my erroneous explanation above, but not nearly so much.)

Whew! I’ve gotta quit now. There’s a lot more to think about here.
(continued...)

Friday, September 25, 2009

More simulated uptake

Thanks to Rosie eliminating the int function from her Perl model, I got to take a look at some more simulated uptake data. Last time, there were several issues, which now seem solved. This time to model uptake, she used the real genomic USS position weight matrix to stochastically select 500 bp fragments from the first 50 kb of the Haemophilus genome. I got 200 from the forward strand and 200 from the reverse complement strand, along with a set of random fragments. This is 4X spanning coverage of the 50 kb…

Below is the way the data looked in the UCSC genome browser, added as custom tracks (click on the figures to enlarge).
From the top, the tracks are:
(1) Chromosome position
(2) 400 random fragments (shades of brown).
(3) 400 selected fragments (shades of blue).
(4) Positions of “perfect core” USS motifs (5’AAGTGCGGT-3’) on either strand
(5) RefSeq gene annotations.

Here’s a bit of the 50 kb zoomed in:
That looks pretty good for such low coverage! (In our real experiment, we expect to get several hundred times more data.) It’s starting to look like a real model of how uptake might look! The random fragments look roughly randomly distributed, and the selected fragments clearly show a punctate distribution around the “perfect core” USS sites.

Here’s a histogram of the scores of the best site on a given fragment for the random and selected datasets:
Indeed, the distributions are quite distinct, though notably the distribution of random fragments looks bimodal. This may simply be a feature of the genome, since there are so many USS sites… Worth thinking about though.

There are other details obscured in the browser figures: the shading indicates the relative score of the best site on the fragment (on a log scale), and each fragment also has orientation shown as a small arrowhead within the box. I’ve also associated each fragment with its score. So in the browser, I can easily check things out more carefully:
In this zoom, it’s clear that there is an excellent site to the left (just under 18,000), a weaker site to its right (~18,400; fragments are overlapping with the left-most site; no perfect core), and a pair of sites on different strands with different scores to the right (~19250). I can also retrieve the sequence associated with a given fragment to see if I can spot the USS site within it. And if I really zoom in, the DNA sequence is listed at the top.

It’s not really so easy to see what’s going on with all these overlapping fragments, so my next task will be to convert this data from fragment positions to spanning coverage per chromosomal position (though I’ll probably bin positions, perhaps every 100 bp to keep things reasonably small for now). I will take a stab at doing this properly (with a script) but may wimp out and do it in Excel. If I can then muscle this data into a WIG formatted file, then I’ll be able to plot the data in a way where “good” sites will look like peaks in coverage…
(continued...)

Thursday, September 17, 2009

The Last Straw

Yesterday, Rosie kindly ran her Perl script over the USS construct I designed. The final thing I was worried about was whether or not my design had any USS or USS-like sequences in it, other than the one it's supposed to have. I'd checked the construct for any core USS motifs (5'-AAGTGCGGT-3'), but since we think that the motif is more complex than this, it was important to make sure that there were no extra sequences that got high scores using the USS position-weight matrix.
Fortunately, the construct looks good, so I can go ahead and order the control oligos and have high expectations that they'll work...

Here's how every 32 base pair window over the 199mer looks when scored with the USS PWM:
There's a single prominent high-scoring site right where it should be, and all of the surrounding area scores near background. The USS in the construct has a score (~10^-8) more than 10 orders of magnitude better than the next best sites. There's a slight increase for windows immediately adjacent to the USS, presumably because the AT-tracts in the USS are still contained in those windows. The rest of the construct only has scores at background.

Just to show that these other sites really do represent background levels of USS score, Rosie also ran a randomized version of the sequence:
Nothing better than 10^-18. Excellent.
(continued...)

Scale UP!

How much periplasmic DNA can I hope to get using my current protocols, and how much DNA will I need? As promised, here are some rough calculations regarding the oligo purchases that we want to make.

I used a molecular weight calculator available on-line to determine the size of the dsDNA I described in the last post.
I alternatively could’ve used Rosie’s Universal Constants (660 g / mol of base pair and 10^-18 g / single 1 kb DNA molecule) to make this calculation, but since I’m dealing with a known sequence, I might as well get an exact molecular weight. (I also made a minor mistake in the last post, and the molecule I describe is actually only 199 bp).

So, for our USS molecule, MW = 122,828.6 g / mol. And the oligo synthesis service we’re planning on using will be at the 1 micromole scale. That means if we took all of the two oligos, annealed, extended, and purified, we’d end up with 0.123 grams of input DNA pool! That’s really a very large amount.

My previous concerns about needing to do PCR to maintain the pool are unfounded. This scale should be sufficient for hundreds (or even thousands) of experiments...

What follows are my preliminary assumptions about yields from the periplasmic DNA prep. They are based on several different experiments, though I am erring on the side of being conservative with my estimates and are guides for future experiments only. In this post, I will address the issues of scale-up at the end; before that I’ll just refer to the approximate total culture volume and amount of DNA that I’d need to get a target amount of DNA, assuming all else works perfectly.

So, I’ve now done several experiments using a PCR fragment bearing the consensus USS, called USS-1. If I add 20 ng DNA / 1 ml competent cells, ~50% is taken up. That is, in rec-2 cells, my theoretical yield of periplasmic DNA is 10 ng. My actual yield is considerably lower; as evaluated by my radiolabeling experiments, I estimate I get ~25% of my theoretical maximum.

This means,

1 ml cells + 20 ng DNA → 2.5 ng recovered.
20 ml cells + 400 ng DNA → 50 ng.
40 ml cells + 800 ng DNA → 100 ng

But this is only for the consensus sequence. Our real experiments will be a mix of molecules, some of which will be efficiently taken up and others that won’t. For a cursory estimate, we might assume that ~50% of fragments will be “good” USS and the other half will be “bad”. This would further reduce the yield.

That means, I am likely to need ~80 ml cultures and a starting input DNA amount of ~1600 ng, just to get back a mere 100 ng of DNA back!

Most Illumina sequencing centers seem to want ~1 ug of DNA to make libraries, but a lot of ChIP-seq experiments seem to call for only ~100 ng. In our case, there will be no downstream library construction, so we can likely get away with small amounts of DNA, as long as it is quite pure and accurately quantified.

Regardless, this is going to take fairly large cultures, fairly large amounts of DNA, and a good scaled-up periplasmic prep.

BUT, one important thing to note is that our degenerate oligo preparation will be more than sufficient for a large number of experiments, even at this large scale. For the controls, I can merely buy minimum-scale synthesis long oligos at ~$200 a pop. Since I can safely PCR amplify these, I will be able to make a replenishable stock for use in scale-up experiments.

More on this in the future, but while I’m doing this, I might as well estimate what it will take to get a microgram of chromosomal DNA fragments out of competent cell periplasms.

My previous experiments with sonicated DNA gave pretty consistent DNA uptake measurements:

~50% of 200 ng 1-10kb DNA / 1 ml cells → 100 ng max. yield.
~10% of 200 ng 0.2-0.4kb DNA / 1 ml cells → 20 ng max. yield.

Given a 25% recovery rate from the periplasm, this means that for a microgram of DNA, I will need:

1-10kb DNA: 8 micrograms in a 40 ml culture
0.2-0.4kb DNA: 40 micrograms in a 200 ml culture (!)

This last is really asking a lot. That size of scale-up will require special thought…

Appendix on Scale-up Issues:
  1. Purity: I have not been adding RNase. I need to get all the RNA away, in order to accurately quantify the DNA. I am also concerned about salt. The CsCl in my DNA precipitates may not be getting washed out adequately by a single 80% ethanol wash.
  2. Cell concentration: It would help for technical reasons, if I could concentrate the cells quite a bit before doing the organic extractions. I have used a ratio of 1:1, cells : organic solvents. So a 1 ml competent cell prep (~a billion cells) gets mixed with 1 ml solvent. But I might be able to resuspend 10 ml of cells in 1 ml and then use 1 ml solvent. I just don’t know.
  3. DNA concentration: I want to make sure that I am saturating with DNA for my initial experiments, but I haven’t yet done a proper saturation curve to know what I should be using. This will decrease the total efficiency of DNA uptake, but my total yields will be higher, and I will be biasing things towards the best uptake sequences (which is a good place to start).
  4. DNase: I have not been treating cells with DNase prior to isolation. From what I can tell, this is not a problem, and the free DNA is washed away. But if I use very high DNA concentrations, I will probably want to use DNase, just to be sure I’m eliminating free DNA completely.
  5. Details, details: Scale-up is never quite as simple as just increasing the volume of everything. I will need to make sure that there are appropriate centrifuges, shakers, tubes, and everything else. Growth rates of cells and competence induction may be poor when going to larger volume cultures. I am also concerned about scaling up the organic extractions. It turns out that not all conicals are created equal; I’ve had disasters where the phenol has torn through the bottom of 50 ml conicals when doing large-scale organic extractions, depending on the brand of conical and rotor used. I’ll need to make sure that things like this don’t happen in advance before I mess up somebody else’s equipment!

(continued...)

Reverse engineering

Okay, so now that we’ve exposited all the brilliant experiments we’re planning to do while writing proposals, the actual reality of doing the experiments is starting to sink in. We’ve also managed to put down some fairly concrete goals for the next several months.

One of our experiments involves measuring the specificity of DNA uptake by naturally competent H. influenzae for fragments containing “the genomic USS motif”. The H. influenzae genome contains an abundant sequence motif, and fragments bearing it are taken up better than fragments that don’t. This “uptake signal sequence” was originally defined by its functional role in DNA uptake, but has since been characterized mostly by bioinformatics, with no direct uptake specificity data. The limited data from previous lab members suggests only an imperfect correspondence between the properties of the genomic motif and the specificity of DNA uptake.

The idea, then, is to feed competent cells small DNA fragments bearing a degenerate (highly mutated) version of the USS consensus sequence, recover those that are preferentially taken up, and sequence the resulting pool. USSs are ~32 bases, well within the reach of single-end Illumina reads, if they are positioned properly next to a sequencing primer.

I’ve previously discussed the expected properties of a degenerate USS pool. And though I think we need to consider this more, I will focus this post on the design of other parts of the construct that will allow us to circumvent subsequent sequencing library construction steps. Illumina sequencing uses specific sequences added to the ends of molecules to capture and sequence DNA of interest...

Properties needed for a USS-containing construct, where Illumina sequencing can be directly performed to sequence the USS:

(1) SIZE: ≥200 base pairs, dsDNA. 200 base fragments with USS are efficiently taken up by cells, and the size is sufficient for efficient cluster synthesis and sequencing using Illumina’s Genetic Analyzer.
(2) CAPTURE SEQUENCES: One end of a strand of each fragment needs to be able to anneal to one of the two “Flow Cell Primers” (FP) in the Illumina flow cell, while the other end of the same molecule needs to contain the reverse complement of the other FP.
(3) SEQUENCING PRIMER BINDING SITE: The reverse complement of Illumina’s sequencing primer needs to be immediately downstream of the reverse complement of the USS. (This could work the other way, but getting the “sense” USS directly from the sequencing reads seems optimal).
(4) TAG SEQUENCE: The first few (four) bases of each read should be in non-degenerate fixed sequence to facilitate the alignment of the degenerate USS reads.
(5) CONSTRUCTION: After consulting several oligo makers, we learned that we wouldn’t be able to get our degenerate constructs built into an oligo longer than 130 nt. This means that I will need to anneal two oligos together and extend with polymerase to generate a full-length construct.

The first trick was to actually find out what the normal Illumina adapter and primer sequences were. They were available on-line, and I think I’ve mostly reverse-engineered what the different bits do. And think I have a reasonable design:

I’ll order two oligos, one 130 nt and the other 106 nt. (At the end of this post, I will list the exact sequences of each part and some notes.) They’ll have 36 bp of reverse complementarity at their 3’-ends, so that I can anneal them and extend to produce full-length construct.
To illustrate what all the different parts of the construct are for, here’s a color-coded version, for which I’ll schematically diagram the Illumina cluster synthesis and sequence priming.
The key features are that the flow cell primers (FP) are on opposite strands on opposite ends and the sequencing primer (SP) sits adjacent to the USS (with the 4-bp tag at the beginning). I am using plasmid sequence present in the lab’s other USS constructs for the Gaps (1 and 2).

To sequence the 200mer (either before or after recovery from competent cell periplasms), the DNA would be melted and annealed to an Illumina flow cell. Below are shown two different parts of a flow cell surface, where the two different strands of a single molecule might anneal.
DNA synthesis from FP1 or FP2 generates a covalently attached version of each strand.
The original molecule is melted off and washed out of the flow cell, and a special in situ PCR method generates clusters of single-strands covalently bound to the flow cell surface. In each cluster, the strands are oriented in both directions.
Sequencing then proceeds from SP binding sites. In this design, the SP binding site will then read the complement of the USS (with the first four fixed bases), so the actual sequence generated would be the USS contained in the construct.
There are several small details to go over to make sure that this design will work. Because the oligos are so expensive, and the degenerate oligo will be precious, I also plan to buy several non-degenerate oligos corresponding to perfect consensus, randomized, and mutant USSs. These will act as controls for the annealing/extension step that generates the uptake substrates and as controls for measuring saturation curves to optimize the appropriate DNA uptake conditions. I will also be able to do PCR to regenerate the control constructs, while I should probably avoid amplifying the degenerate USS construct for fears of strongly biasing the representation of different sequences.

NEXT UP: Uh oh… What about yields? Dimensional analysis…

APPENDIX:

The different parts of the two oligos:

Notes on my reverse engineering:
  1. FP1 (25 nt): Composed of putative 20mer FP1 + first 5 bases of one adaptor (calling it A)
  2. SP1 (33 nt): Sequencing primer for single-end Illumina runs. Includes the 13 bases of the normal adaptor that normally results in a 13 bp inverted repeat palindrome on either side of adapted DNA fragments.
  3. USS (36 nt): Includes 4-base tag (ATGC) upstream of a 32-base genomic Gibbs consensus sequence with a set level of degeneracy at each position.
  4. G1 (36 nt): Additional sequence from pGEM7f ,corresponding to the portion of the spacer region where the two oligos are intended to anneal.
  5. G2 (46 nt): More sequence from pGEM7f, corresponding to the spacer region only on one of the two oligo.
  6. FP2’ (23 nt): Composed of the complement to the 20mer FP2 + first 3 bases of the other adaptor (calling it B).
  7. Total length after annealing and extension is 200 bases, where the USS is located from position 63 (after the spacer) to position 94. In the flow cell, the use of SP1 as a sequencing primer should read the complement of the USS sequence, so the actual sequence obtained will correspond to USS (with the first four bases always ATGC).

(continued...)

Tuesday, June 16, 2009

They're eating our DNA!!!

I am in the midst of proposal-writing, which makes blogging tougher, but when I hit a writing-block at the end of the day, I decided to dally with a random thing I'd been meaning to figure out.

Several friends of mine, when I tell them about my plans with the naturally competent bacteria, have said, "Dude, you should feed them human DNA!"

Sound silly? Maybe it is, but I went ahead and used TAGSCAN to search the human genome for the 10-bp core USS motif: AAAGTGCGGT and found 762 instances of the USS core. BLAST and BLAT didn't want to deal with me due to how short the query USS motif is.

That means Haemophilus influenzae might find nearly 800 bits of our genomes tasty! I remember from my reading that there's nearly 200 micrograms of DNA per milliliter of our lung mucus, which seems like a heck of a lot. Much of this DNA is probably human...

(For context, on average Haemophilus have a USS less than every 2 kilobases, but humans have an average USS density lower than one in 20 megabases. So while humans have half as many USSs, there's a 10,000 times lower density.)

It isn't surprising the human genome contains at least some "USS". A random 10mer string would have a 1/(4^10), or a little less than 1 / 1,000,000, chance of being the USS motif. But the human genome is more than 3 billion base pairs. So randomly, we might expect to find more than 6000 USS (3 billion / 1 million X 2). (The 2 is to also count the reverse complement.)

If I did that right, then the observed number of USS motifs in the human genome is rather less than expected. That's interesting...

However, that was a crude estimate of the expected number of USS motifs. The USS core sequence has 5 GC bases and 5 AT bases, giving a GC base composition of 50%, whereas the human genome has only 41% GC content. There's also the issue of dinucleotide frequencies. For example, the CpG dinucleotide is underrepresented in the human genome, but happens to appear in the USS. I've been trying to figure out a rational way to incorporate this type of information into my estimate of expected, but so far have failed to do so properly. Regardless, my estimate of expected is certainly too high. The question is: how much so?

To produce control numbers, maybe tomorrow I'll run a few other TAGSCANs for arbitrary 10mers with the same GC content and a CpG and see how many come up. I haven't really looked at the distribution of USS, except that the number of USS per chromosome is highly correlated to chromosome size (R^2 = 0.91). I might also predict that USS will fall into more GC-rich regions of the genome.

But for fun, assuming that the result held, what might it mean if the USS motif is significantly underrepresented in the human genome? I can hardly imagine that Haemophilus could be responsible ("it ate them!!!"), but maybe it could work in the other direction? Perhaps the USS motif is only coincidentally somewhat rare in humans, and as such makes a good sequence to use for preferring conspecific DNA uptake? If the USS motif was extremely abundant in the human genome, then naturally competent Haemophilus might not take up conspecific DNA as easily? Hmmm...

Regardless, think about that when you fall asleep, folks... There's bacteria inside you, and THEY'RE EATING YOUR DNA...

UPDATE:

Putting my proposal together continues, but in a bout of procrastination I did go ahead and run a few random 10mers with the same base composition (and a CG dinucleotide) through TAGSCAN, and it looks like the observed number of USS motifs is about expected for 10mers of similar composition.

The USS motif:
AAAGTGCGGT 762

Randomized USS motifs:
GTACGTAAGG 414
CGAGAGAGTT 761
GAATACGTGG 1112
AGATAGCGTG 714
GTAACGAGTG 458
AGACGTTAGG 713

Nothing to see here... Move along...

(continued...)