Thursday, September 17, 2009

Reverse engineering

Okay, so now that we’ve exposited all the brilliant experiments we’re planning to do while writing proposals, the actual reality of doing the experiments is starting to sink in. We’ve also managed to put down some fairly concrete goals for the next several months.

One of our experiments involves measuring the specificity of DNA uptake by naturally competent H. influenzae for fragments containing “the genomic USS motif”. The H. influenzae genome contains an abundant sequence motif, and fragments bearing it are taken up better than fragments that don’t. This “uptake signal sequence” was originally defined by its functional role in DNA uptake, but has since been characterized mostly by bioinformatics, with no direct uptake specificity data. The limited data from previous lab members suggests only an imperfect correspondence between the properties of the genomic motif and the specificity of DNA uptake.

The idea, then, is to feed competent cells small DNA fragments bearing a degenerate (highly mutated) version of the USS consensus sequence, recover those that are preferentially taken up, and sequence the resulting pool. USSs are ~32 bases, well within the reach of single-end Illumina reads, if they are positioned properly next to a sequencing primer.

I’ve previously discussed the expected properties of a degenerate USS pool. And though I think we need to consider this more, I will focus this post on the design of other parts of the construct that will allow us to circumvent subsequent sequencing library construction steps. Illumina sequencing uses specific sequences added to the ends of molecules to capture and sequence DNA of interest...

Properties needed for a USS-containing construct, where Illumina sequencing can be directly performed to sequence the USS:

(1) SIZE: ≥200 base pairs, dsDNA. 200 base fragments with USS are efficiently taken up by cells, and the size is sufficient for efficient cluster synthesis and sequencing using Illumina’s Genetic Analyzer.
(2) CAPTURE SEQUENCES: One end of a strand of each fragment needs to be able to anneal to one of the two “Flow Cell Primers” (FP) in the Illumina flow cell, while the other end of the same molecule needs to contain the reverse complement of the other FP.
(3) SEQUENCING PRIMER BINDING SITE: The reverse complement of Illumina’s sequencing primer needs to be immediately downstream of the reverse complement of the USS. (This could work the other way, but getting the “sense” USS directly from the sequencing reads seems optimal).
(4) TAG SEQUENCE: The first few (four) bases of each read should be in non-degenerate fixed sequence to facilitate the alignment of the degenerate USS reads.
(5) CONSTRUCTION: After consulting several oligo makers, we learned that we wouldn’t be able to get our degenerate constructs built into an oligo longer than 130 nt. This means that I will need to anneal two oligos together and extend with polymerase to generate a full-length construct.

The first trick was to actually find out what the normal Illumina adapter and primer sequences were. They were available on-line, and I think I’ve mostly reverse-engineered what the different bits do. And think I have a reasonable design:

I’ll order two oligos, one 130 nt and the other 106 nt. (At the end of this post, I will list the exact sequences of each part and some notes.) They’ll have 36 bp of reverse complementarity at their 3’-ends, so that I can anneal them and extend to produce full-length construct.
To illustrate what all the different parts of the construct are for, here’s a color-coded version, for which I’ll schematically diagram the Illumina cluster synthesis and sequence priming.
The key features are that the flow cell primers (FP) are on opposite strands on opposite ends and the sequencing primer (SP) sits adjacent to the USS (with the 4-bp tag at the beginning). I am using plasmid sequence present in the lab’s other USS constructs for the Gaps (1 and 2).

To sequence the 200mer (either before or after recovery from competent cell periplasms), the DNA would be melted and annealed to an Illumina flow cell. Below are shown two different parts of a flow cell surface, where the two different strands of a single molecule might anneal.
DNA synthesis from FP1 or FP2 generates a covalently attached version of each strand.
The original molecule is melted off and washed out of the flow cell, and a special in situ PCR method generates clusters of single-strands covalently bound to the flow cell surface. In each cluster, the strands are oriented in both directions.
Sequencing then proceeds from SP binding sites. In this design, the SP binding site will then read the complement of the USS (with the first four fixed bases), so the actual sequence generated would be the USS contained in the construct.
There are several small details to go over to make sure that this design will work. Because the oligos are so expensive, and the degenerate oligo will be precious, I also plan to buy several non-degenerate oligos corresponding to perfect consensus, randomized, and mutant USSs. These will act as controls for the annealing/extension step that generates the uptake substrates and as controls for measuring saturation curves to optimize the appropriate DNA uptake conditions. I will also be able to do PCR to regenerate the control constructs, while I should probably avoid amplifying the degenerate USS construct for fears of strongly biasing the representation of different sequences.

NEXT UP: Uh oh… What about yields? Dimensional analysis…

APPENDIX:

The different parts of the two oligos:

Notes on my reverse engineering:
  1. FP1 (25 nt): Composed of putative 20mer FP1 + first 5 bases of one adaptor (calling it A)
  2. SP1 (33 nt): Sequencing primer for single-end Illumina runs. Includes the 13 bases of the normal adaptor that normally results in a 13 bp inverted repeat palindrome on either side of adapted DNA fragments.
  3. USS (36 nt): Includes 4-base tag (ATGC) upstream of a 32-base genomic Gibbs consensus sequence with a set level of degeneracy at each position.
  4. G1 (36 nt): Additional sequence from pGEM7f ,corresponding to the portion of the spacer region where the two oligos are intended to anneal.
  5. G2 (46 nt): More sequence from pGEM7f, corresponding to the spacer region only on one of the two oligo.
  6. FP2’ (23 nt): Composed of the complement to the 20mer FP2 + first 3 bases of the other adaptor (calling it B).
  7. Total length after annealing and extension is 200 bases, where the USS is located from position 63 (after the spacer) to position 94. In the flow cell, the use of SP1 as a sequencing primer should read the complement of the USS sequence, so the actual sequence obtained will correspond to USS (with the first four bases always ATGC).

(continued...)

Wednesday, September 9, 2009

Eating chromosomal DNA fragments

Haemophilus influenzae cells will take up closely related DNA from the environment quite efficiently, when they are made naturally competent by resource limitation.

Previously, I had done some experiments using sonicated chromosomal DNA of two different size distributions. The take-away lesson was that, for a fixed DNA concentration, larger fragments were taken up better than smaller fragments. This could be due to two non-exclusive reasons:
  1. Larger fragments are more likely to contain an uptake signal sequence.
  2. The uptake machinery is saturated when I used the smaller fragments, since there are more fragments per unit mass.
I am not certain of the best way to measure the relative contributions of these two factors to the observed disparity in uptake, though I’m pretty sure a saturation curve would be the way to start things off, that is measuring the amount of uptake over a wide range of DNA concentrations.

But first, and more to a practical concern for our sequencing plans...

I repeated this experiment, but also prepared total DNA and periplasmic DNA (by the slick method of Kahn et al) to make sure that I could cleanly recover chromosomal DNA fragments trapped in the periplasm from bulk chromosomes, as I previously showed for a small USS-containing PCR fragment.

Here are the results of that experiment (in which I provided ~0.5 billion competent cells with 200 ng of end-labeled DNA fragments of two different size distributions, either 1-10kb or 200-400 bp, for 30 minutes):
In (a), the % uptake is clearly better for the larger size distribution than the smaller size distribution. In (b) and (c), I show that I can purify periplasmic chromosomal fragments away from the cell’s chromosome. In (b), the results using the larger fragments is shown, while in (c) the results using the smaller fragments are shown. (I ran gels with two different agarose concentrations to optimize the separation for the two different input pools).

One thing to note is that the size-distribution of DNA between the input and periplasmic preparation were effectively indistinguishable. I looked at traces of these lanes in the Molecular Dynamics ImageQuant software, and they looked pretty much exactly the same. This is a little bit confusing, given the two models discussed above and the fact that fewer small fragments were taken up compared with larger fragments. I might have expected that there would be a bias towards the larger fragments in the periplasm compared to the input, but this was not the case.

Another thing to note is that, unlike when I previously did this experiment with USS-containing PCR fragments, there is still evidence of periplasmic DNA in wild type after 30 minutes. I don’t think this is due to poor washing of free DNA away from the cells, but rather reflects that there had been insufficient time to translocate all of the DNA in the periplasm into the cytosol. There is also the possibility that some of the non-chromosomal DNA in the wild-type samples are indeed cytosolic, which I can’t tell without some way to distinguish ssDNA and dsDNA.
(continued...)

Friday, August 28, 2009

Homoplasy versus Recombination

Making up for lost blogging time!

Just to get started doing something with this alignment, I just grabbed a single 68,580 bp alignment block out of the XMFA file I generated (as described in the last post), which was itself a simple FastA formatted file. So I could use whatever programs could handle a simple FastA alignment and not worry about converting around to different formats yet, or coping with contig and rearrangement boundaries.

Here’s a picture from MAUVE of the aligned block I used (KW20 Rd reference coordinates 401,890 to 467,839). It’s the orange block shown below.
Based on some recent reading, I tried a couple different programs designed to detect signals of recombination from multiple alignment files. There is a serious caveat here, in that I do not understand exactly how these programs work, and since they were designed for use on sexual eukaryotes, it’s quite possible that the assumptions of the programs will not be accurate when applied to Haemophilus. Nevertheless, I sallied forth, just to see if they’d at least work with my sequences. I can come back and try to understand the meaning of the output more slowly...

LDhat

First, I tried LDhat, which is composed of a suite of programs aimed at detecting recombination in a population sample. I didn’t have any difficulty downloading and compiling the program, which was nice. I wish I could tell you more about it, but I never actually made it work, so there’s not much point yet. I got the “Convert” program to make the requisite files from an only slightly modified FastA, and I eventually got the “Pairwise” program to work. This program generates a “look-up table” which is required for the other programs to operate.

At first, the program kept choking on my data and was unable to produce the requisite look-up table for my data. So I fed it a pre-made file that can be found at LDhat’s website. The look-up table I used was first converted to a 15 taxa file using their “LKgen” utility. Then “Pairwise” worked and generated me a new look-up table that should theoretically have worked with the other programs. I next used the “Interval” program, but I could never get past here. It wouldn’t do the analysis, because the look-up table was somehow lacking. It was not “exhaustive”, whatever that means, and there was apparently a disparity between “npt” (which I assume stands for Nonparametric test) and the data.

No clue. Moving on. (I’ll try to return to this some other time, since the program “Rhomap” looks like what I’m looking for).

PhiTest

Next I turned to PhiTest, which was also easily downloaded and compiled. This program uses a principle that seems pretty sensible. It examines “incompatibilities” in phylogenetic signals. If two lineages of bacteria diverge and never recombine, then adjacent polymorphisms will most likely be “compatible”, that is both polymorphic sites will have the same phylogenetic signal (i.e. they will support the same tree topology). On the other hand, “incompatible” sites could have two possible histories, one in which there was recurrent mutation in different lineages and one in which there was recombination between lineages. This is illustrated nicely in a figure from the paper reporting PhiTest:

"Figure 1:
The dual nature of incompatibility. Two possible histories for a pair of incompatible sites are shown: (a) two incompatible sites explained by a recombination event and (b) two incompatible sites explained by a convergent mutation. Mutations in the first site are indicated by open circles and mutations in the second site are indicated by solid circles. To explain the incompatibility between the pair of sites either a recombination event must be invoked or a homoplasy must have occurred in the history of one of the sites."

Figure 1A:
Figure 1b:
So in (a) the black mutation only occurred once and was subsequently shared between the central lineages, whereas in (b) the black mutation occurred independently in the two lineages.

They more rigorously define “compatibility” thusly: “Two sites i and j are compatible if and only if there is a genealogical history that can be inferred parsimoniously that does not involve any recurrent or convergent mutations (known as homoplasies as in Figure 1b). If the two sites are not compatible, they are termed incompatible. Under an infinite-sites model (KIMURA 1969) of sequence evolution, the possibility of a homoplasy does not exist, and so incompatibility for a pair of sites implies that at least one recombination event must have occurred, as in Figure 1a.”

So the idea of recombination detection algorithms is to exploit the notion of incompatibility, but it is extremely important to note that there are really two possibilities for any incompatibility: (a) It is a true “homoplasy”, i.e. a recurrent mutation; (b) it indicates recombination. The basic idea of the NSS method (which is apparently what LDhat is using) is to look at neighboring pairs of sites and find regions where incompatibilities are clustered together, suggesting perhaps a hotspot.

The other, possibly more eukaryotic, aspect is that sites that are more distant will tend to be less compatible, since they’d more likely have had crossovers between them. Our provisional model for natural transformation would not necessarily have this feature, however, since we do not expect our recombination events to be crossover-like.

The authors of PhiTest produced a different measure than the NSS model that looked at the “Pairwise Homoplasy Index” for each pair of SNPs in an alignment. This gives a test statistic of the minimum number of homoplasies in the alignment, essentially based on a parsimony criterion. They could then apply permutation tests to measure the significance of a given PHI for an interval. Happily, the program will also output the significance under the NSS and MaxChi models.

For my alignment, PhiTest gave the probability of no recombination as 0 for all three models. So it looks like there’s quite a bit of homoplasy in my alignment, evidence of recombination. (It’s worth noting that PhiTest and LDhat both explicitly ignore gaps in the alignment.)

But that’s not good enough, what I really want is a plot of recombination rates across the alignment block. The figure I made from this post was the output of that analysis. It is only showing the ~9500 polymorphic sites and ignoring all the gaps and identical sequences. It’s showing, for each pair of SNPs, the probability of recombination (assuming that all homoplasies are due to recombination and not recurrent mutation).

It pretty much looks like there’s evidence of recombination almost everywhere! There’s just a couple areas that look like they have extensive compatibility, so it may actually be a lot easier for us to detect places where there’s been little recombination than places where there’s been lots.

I’ll need to go back to the raw data and figure out which set of SNPs are showing evidence of low recombination and see where those spots are in the alignment. They could be horizontally transferred segments with few USS?

I’m not sure what to make of this. It complicates things. To what extent can I trust these algorithms to find recombination and not recurrent mutation? The divergence between my sequences isn’t that high (only ~3%), so recurrent mutation doesn’t seem likely to be a huge problem, but still...

Anyways, I went ahead and used the other utility that comes with PhiTest, called Profile. This does PhiTest on sliding windows, providing a P-value for each segment. Here’s what it looked like, when I did a “Profile” of non-overlapping 500 bp chunks of my alignment:
The x-axis is the position in the alignment, and the y-axis is the p-value of the PhiTest for that particular 500 bp chunk. The line indicates a P-value threshold of 0.05, so any segment below that line has evidence of recombination, based on PHI. Segments that have higher P-values may either have compatible SNPs, or for the very high P-values, be sites with very little SNPs.

This seems promising, if we can trust that we’re measuring something meaningful. I have a strange feeling that using homoplasy to measure recombination is going to be fraught with difficulties, but it’s pretty much the best principle that we can work with to detect recombination in these population-level samples. Luckily, both PhiTest and the other methods are sophisticated enough that they are evaluating the homoplasy index for a given pair of SNPs relative to other surrounding SNPs, so when incompatibilities are detected, they are standing out from other possible pairings. This should help with the problem of true homoplasy, but then again, maybe it doesn't. I'll need to talk to my friend Corbin about this stuff.

Whew! Okay, enough on this for the moment...
(continued...)

More alignments!

I realized something disturbing a couple days ago, when I crashed my computer by running out of hard-drive space... The Haemophilus influenzae genome is actually replicating inside my computer! Not just on my lab bench! Several of the programs I've been using (namely the alignment programs) have been spitting out files containing the complete genome, and I think I must've had several hundred before I got rid of a bunch of intermediary files. In silico infection! A scourge!

Anyways, I guess that last entry deserves more explanation, but it'll take me more than a post to do it. First off, the motivation for doing these kinds of analyses:

There’s 4 complete Haemophilus influenzae genomes, 11 with draft assemblies (in 4-50 contigs), and 17 for which sequencing is in progress. On top of that, hundreds of strains have been sequenced at seven house-keeping genes (MLST studies).

Within this “population genomic” data, we should be able to do some estimates of recombination rates between isolates, as well as some sophisticated analyses of genetic variation at USS motifs. But how? Here’s the basic idea:

(1) Align the complete and draft genomes.
(2) Use programs written to detect recombination signals in population genetic data.
(3) Examine genetic variation in and around USS motifs (scored for information content across each genome) and see if they correlate with the determined recombination signals.

There are several issues with this at each step, but I’ll take the first one first. I’ll start with a post describing how I produced the multiple alignment file that I fed to the PhiTest program...

The first trick was downloading the files...
I already had the four complete sequences, but hadn’t had the easiest time getting the contigs of the draft assemblies. Do they really want me to download each contig by clicking one at a time? My problem was solved when I found the whole-genome shotgun (WGS) FTP site at NCBI. Of course, these files were named by their WGS code, so I simply hand-search for the correct four letter code, based on this list and downloaded the corresponding file. If I’d been more clever, I would’ve done this from the command line with a script containing the list of WGS data I wanted.

I used the search term "txid727[orgn]" in the Genome Database or the Taxonomy Browser to call up Haemphilus influenzae genomes. Using the Genome database yielded some spare entries that were plasmid sequences, and using Taxonomy browser yielded many isolates where the sequencing is still in progress. Anyways, I could get to a Genome page for each sequenced isolate, but when I clocked to get the contigs, it provided separate accessions for each contig. Boo! There had to be a better way, which took some digging, but was here. This is all the WGS data at NCBI. Whew! Luckily, I found the requisite four-letter codes from my other searches and downloaded the appropriate *.gbff.gz files and expanded them with gunzip.

Anyway, this gave me a bunch of GenBank files with all the contigs in each one. The contigs themselves were simply ordered by size, with the largest first and the smallest last, and the coordinate system corresponded to this ordering. While this is fine, it would probably help me looking at alignments, if I ordered the contigs based on the reference KW20 Rd genome. It would be a meaningless gesture in some ways, but would certainly facilitate looking at pictures.

I did this by using MAUVE’s MoveContig function. I reordered each set of WGS contigs pairwise against the reference KW20 Rd. (Again, I should've been using the command-line, but found it easier to just click my way through... shameful.) Here’s an example of what MoveContigs did comparing PittII to Kw20 Rd:

Alignment #1: Keeps the original order of the contigs (from largest to smallest).
KW20 Rd is on top; PittII is on bottom. The colored blocks indicate contiguous aligned blocks. The red lines indicate contig boundaries (KW20 has only two, for the start and stop of the chromosome). The diagonal line between the two genomes show which blocks belong together. It’s quite obvious from all the crosses that the order of the contigs is off.

The program then reordered the contigs and did another alignment. Then it did this a third time (it keeps iterating until it’s done). It decides this is as good as it can do and stops.

Alignment #3: Reordered contigs maximizing “synteny” between the reference and PittII.
Clearly reordering the contigs has made the picture clearer to see. There are far fewer alignment blocks and fewer diagonally crossing central connector lines. Of course, some of the contigs did span some breakpoints, so there’s obviously rearrangements between the genomes. At each contig boundary, there is always the possibility of some rearrangement between the strains, so this contig order is at best provisional. Another thing I had to keep in mind while tooling around with these alignment programs is that the genome is circular, but the programs don’t treat the strings that way, so there’s often things at the right or left ends that obviously are still syntenic, but just don’t look like it with these linear visualizations.

Here’s the alignments again, without the colored blocks, just so that the reordering of the contigs is quite obvious:

Alignment #1:
Alignment #3:
Notably, the little tiny contigs that failed to find something in the reference to align with remained at the right end.

Alright, so going through 11 of these MoveContig steps provided me with 11 new files that had the contigs in a different order (based on KW20 Rd) than what I got from GenBank. The only unfortunate thing about this was that MoveContigs only outputted FastA, instead of GenBank, so I lost all the annotations. Oh well. I would still be able to use the KW20 Rd annotations. I think that the annotations can still be added back, but I haven’t figured that out yet.

To show that this exercise wasn’t a waste, below I show the stripped-down alignments when I did the big 15-way multiple alignments (I hid the 3 other completely sequenced strains from view, and only the top several contig reorderings are visible.

Before Contig Reordering:
After Contig Reordering:
Much easier to look at... But still sort of meaningless.

Alright! So that’s a 15-way multiple genome alignment. And indeed, it’s associated with a gigantic alignment file in the XMFA format: an extension of the FastA format that allows for a description of alignments that have rearrangements / contig breakponts / etc between them.

The most difficult challenges:
(1) Handling the breakpoints, contig boundaries and real rearrangement breakpoints alike.
(2) Handling indels. Those within an aligned block should be scorable, but how?

So really, the first thing I can work with without much difficulty are contiguous aligned blocks. I grabbed one of these from the XMFA file to produce a single FastA file, which I called test1.fasta, which I planned to use for some preliminary analyses...

TO BE CONTINUED
(continued...)

Wednesday, August 26, 2009

Measuring recombination?

Whew! I've been slacking on the blog posts! Above: Probability of recombination between pairwise SNPs in a 15 strain alignment block of ~65 kb. Yellow indicates low probability of recombination...

The figure above was generated by PhiTest, a program that tries to identify regions of recombination between different taxa in a multiple alignment. Last night, I used MAUVE to generate a 15-way multiple alignment between the genome sequences of 15 different Haemophilus infleunzae isolates. (This took a while, but still was impressively quick.)

I grabbed a reasonably large alignment block (65,480 bp, including gaps) and stuck it into Phitest. It did everything reasonably quickly and gave me the unshocking result that there was a p = 0E+00 chance of no recombination occurring on the interval. So I decided to use the -g switch to provide a graphical output giving the chance of recombination between every pair-wise combination of SNPs (of which there were 9422; the program ignored indels).

This was fun, because it wrote such a gigantic file that my computer died due to lack of hard-drive space! That doesn't happen often! But it was mainly because I have several dozen gigabytes of stuff on my laptop that doesn't belong. I nudged a fraction of this aside and repeated the analysis to give the plot above.
(continued...)

Thursday, August 13, 2009

Uptake and Transformation with "Biorupted" samples

To get an idea of how uptake and transformation would work with different sized donor chromosome fragments, I took MAP7 DNA and sheared it in a “bioruptor”. I ended up with several samples with different size distributions, three of which I used for a pair of experiments:

LARGE: >40 kb (unsonicated)
MEDIUM: 1-10kb (1 X 10 min sonication)
SMALL: 100-400 bp (5 X 10 min sonication)

My naïve assumption was that % uptake would go down as the fragment size decreased, since fewer fragments would contain “uptake signal sequences” (USS), which have an average density in the genome of ~1kb.

I also thought that transformation rates would also go down for smaller fragments, but would not necessarily correlate that well with uptake, since additional steps of translocation and recombination could also potentially influence the efficiency of transformation. So for example, transformation might drop off more quickly than uptake, if degradation would affect smaller fragments more than larger fragments. (This seemed to be the case in Pifer and Smith, 1985.)

Keeping in mind that these are just one-off experiments and need to be repeated (like pretty much every experiment I’ve reported in this blog), the above predictions look like they’re true, but I’m not certain if my reasons are necessarily correct...

First, I’ll show the uptake data. I end-labeled the MEDIUM and SMALL donor fragments using Klenow and did a simple uptake experiment (comparing total radiolabel to that in cell pellets after 30 min of uptake) using wild type cells and as donors, either MEDIUM or SMALL chromosome fragments, and either saturating (500 ng) or sub-saturating (100 ng) amounts of input DNA per 0.5 ml of wild-type competent cells. Here’s the data:
So clearly, several-fold less chromosomal DNA is taken up when smaller fragments (100-400 bp) are used than when larger fragments (1kb-10kb) are used. As explained above, one reasonable explanation for this is that fewer fragments in the SMALL sample contain USS, so maybe only 1/10 to 1/2 will have a USS, whereas in the MEDIUM sample most fragments will contain at least 1 USS.

But there is an alternative explanation that I’d like to be able to distinguish (but am pretty sure I can’t with this one experiment). It could be that a given competent cell only takes up a fixed number of DNA fragments, independent of fragment size. So since the SMALL sample is composed of ~10-100X more fragments per unit mass, it could be that I’ve simply saturated the system with fragments in the case of SMALL, but not in the case of MEDIUM. This issue was addressed by Deich and Smith, 1980, and they concluded that indeed this was the case (that the number of molecules taken up was independent of fragment size), but while they do mention USS, they do not bring up USS density as a potential reason for their data.

I’d hoped that by doing a second DNA concentration (100 ng) that I might get hints as to which of the two above models is correct (or if they are both correct and both contribute to the observation), but I don’t really think I can say too much without repeating this several times and getting some error bars on that graph. Furthermore, I’m not really entirely sure what the expectations are for the two models. I’ll have to think on this some more....
-----
Okay, what about transformation rates using my “biorupted” fragments? Below are two graphs reporting the transformation rates of the KanR and NovR alleles from MAP7 to KW20 for the three different DNA pools (LARGE, MEDIUM, and SMALL):
(Note: In the case of the NovR/CFU SMALL sample, the number reported is actually the limit of detection, so NovR/CFU(small) is less than 4.4e-6. I didn't observe any NovR transformants for the SMALL fragments, despite having a decent limit of detection.)

First, I'll look at the difference between the LARGE and MEDIUM fragments. Both markers showed ~6-fold decrease in transformation rate in the briefly sonicated sample, compared to the large intact fragments. Possible explanations:
  1. Less USS per fragment: I doubt this is a major reason for the difference. While a larger percentage of fragments are expected to contain no USS in the MEDIUM sample, it shouldn’t be that large of a difference, since the mean density of USS motifs is ~1kb.
  2. Degradation by translocation or cytosolic nucleases: This seems to be a reasonable explanation. From an old set of experiments using a defined plasmid donor, Pifer and Smith, 1985 estimated that an average of ~1.5 kb of a leading 3’ end is degraded during translocation. Maybe the medium-sized fragments simply don’t survive translocation as well as large taken up fragments.
  3. Recombination efficiency: Maybe both the LARGE and MEDIUM fragments make it into the cytosol, but homology search and recombination are much better for larger fragments.
Things look a little more interesting when looking at the change between MEDIUM and SMALL fragments: While the KanR rate only changed modestly (less than 2-fold), the NovR rate went down below my limit of detection. Possible explanations:
  1. Distance to USS: The nearest USS to the NovR allele is more than 4oo bp away, but less than 400 bp for the KanR allele: I like this explanation. I really really need to figure out the identity of these antibiotic resistance alleles. We’re pretty sure it’s a mutation in the gyrB gene, but I don’t know the actual change. I looked at gyrB and it does contain a USS core motif and two other core motifs a few hundred bases before the start codon, but the gene is ~2.5 kb, so the actual gyrB mutation could easily be too far away from these USS. When I looked at the putative gene responsible for KanR (the ribosomal S7 gene), there was a single USS near the start, but none within. Again, without knowing the causative lesion, I can’t tell whether this is within the size distribution of the SMALL fragments.
  2. Differences in degradation rates at the two loci: This is possible. The other thing these data suggest is that, despite the ~1.5 kb average degradation reported by Pifer and Smith, 1985, there’s still plenty of small fragments that can recombine, since the KanR rates between MEDIUM and SMALL are not really dramatically different.
  3. Recombination signals: Also possible. And probably the hardest to tell, since the other effects need to be canceled out.
The main take-away (assuming that the results replicate) is that not only do different markers transform at different rates, the change in transformation rate for different sized-fragments also varies for different markers. The underlying reasons for this are probably interesting. So what next? I need to repeat this, but next time I’d like to:
  1. Extend this to additional markers to see if this variability also applies to other loci.
  2. Measure linkage between Kan and Nov. I’ve previously seen the known linkage between Kan and Nov using large fragments, but would expect linkage to vanish for small fragments when the KanR and NovR alleles never share the same fragment.
And again, I really need to know what the lesions are that are responsible for the MAP7 antibiotic resistances. I’ve looked around a fair amount, but it seems that many antibiotic resistances can be produced by mutations in more than one different gene, so narrowing it down isn’t that straightforward.
(continued...)

Monday, August 10, 2009

Back to the grind

Whew! Getting that albatross (= grant application) off my neck feels good, but now I need to formulate a plan for the next several weeks of lab work. What are my priorities? There’s quite a lot of things I could be doing, so I might as well make a list...

(A) Generate purified DNA from 86-028NP, PittEE, and PittGG donor DNA: Should be pretty easy. I tried this a couple times in the past couple weeks, but had a few problems:
  1. Got a goopy mess. I don’t know if it was my lysis or something about my extractions, but when I would try to pull the aqueous layers off of my phenol/chloroform extractions, this goopy junk at the interface kept interfering with my ability to complete the preps. I will try try again, but this next time I will do a few things differently: (i) start with fewer cells per volume, (ii) do a straight phenol extraction first, (iii) let lysis proceed longer.
  2. PittEE didn’t grow. I’d gone to the lab stocks, streaked for single colonies, grown up overnight cultures from single colonies, and stored them in glycerol in the -80. But when I returned to PittEE frozen stocks, they don’t grow up! Weird. Maybe PittEE is particularly bad at surviving outside log-phase, so I had merely frozen down a pile of dead cells. I’ll have to return to the lab stocks and restreak for single colonies.
  3. I never checked the putative nalidixic acid resistant phenotype of 86-028NP. I really hope that marker will work to select for transformants from 86-028NP to KW20, as this would really help several steps of the project. In our strain database, 86-028NP is listed as NalR, but this might only be a clinically relevant phenotype. I need to check it on plates.

(B) Check the uptake and transformation phenotypes of WT, rec-2, rec-1, and pilB: I’ve already gone through this exercise with wt and rec-2, but as a negative control, pilB mutants should neither take up DNA nor be transformed by it, and for my translocation experiments, rec-1 should take up DNA normally but fail to transform. I have all these strains frozen down as M-IV competent cells, so I just need to do the experiments to verify that my mutants have the expected phenotypes. The rec-1 mutant in particular is important, since I plan to use it for cytosolic DNA purifications, but it could be somewhat leaky, something I need to know. (The only rec-1 mutant available is the original isolated mutant from back in 1972. It’s older than me!

(C) Quantify uptake and transformation for different-sized donor DNA: A coupe weeks ago, I used a “Bioruptor” to sonicate MAP7 DNA to different size distributions, and I tested the transformation rate of antibiotic resistance into competent cells of some of these different sized DNA. I also want to check how well they’re taken up. Once I’ve done this, I’ll report back the full set of results. The expectation is that while even small fragments will be taken up well, only large fragments will transform well. There’s also an interesting expected relationship between the amount of uptake expected and the size distribution of the DNA. If the DNA fragments are quite small, then only a fraction are expected to contain uptake signals, but many fragments containing such signals will be taken up. On the other side, large DNA fragments will mostly all contain uptake signals, but only a few molecules will be taken up, but these will be larger... Hmmm....

(D) Design the degenerate USS construct. This will require its own post, but suffice it to say, we have a pretty good plan on how to design this construct so that we will be able to directly sequence the degenerate USS motif by short single-end Illumina sequencing without any processing steps upfront.

(E) Transformation experiment: If indeed 86-028NP can grow on nalidixic acid plates, then I will just go ahead and do half of the experimental-side of one of my specific aims. All I need to do is transform KW20 competent cells (which I have) with DNA taken from 86-028NP (which I’ll make), select for colonies that are NalR, pick several of them, grow them up, and extract their DNA. This would provide up with plenty of material to produce sequencing libraries and then send to sequencing. Then we can really get this project off the ground.

(F) Find somewhere that will do sequencing with a short turn-around: The local Illumina sequencers (at UBC’s Genome Center) offer what looks like excellent sequencing services for a reasonable cost, but their turn-around time is a bit too slow for us to hope to have any preliminary data for Rosie’s grant applications. If we want to get anything done more quickly, we will have to find someone else to get started.

(G) Design and produce a large DNA fragment from the 86-028NP genome to use for developing a cytosolic ssDNA prep: Since the most difficult experimental part of our plans is likely to be purifying (and cloning) ssDNA from the cytosol, I’d like to start with a defined construct I can use for purification experiments. I think that a good choice would be a large clone from 86-028NP, because this might be useful for other experiments as well. I will want something at least several kb on a plasmid. I will need to think about whether or not it matters what this fragment contains. As a first guess, I would want it to contain the putative NalR allele and several kb of flank.

What else should I be doing?
(continued...)

Thursday, August 6, 2009

Almost...


...there... However it turns out, I'll finally have this grant I've been working on out of my hands tomorrow. I'm actually thrilled to get back into the lab when it's all over...

(continued...)

Monday, July 27, 2009

Food-to-sex ratio?

In my recent experiments using radiolabeled USS-1 donor DNAs, I was impressed by just how well naturally competent H. influenzae will slurp up USS-containing DNA. I’d read about it, but observing it myself was really something. (It reminds me a bit of the first time I “saw” Mendelian segregation when I dissected my first yeast tetrads.)

The other thing that struck me, reading the antecedents of my uptake experiments, was that a large majority of taken up DNA is simply degraded and the nucleotide subunits used for DNA replication. In my experiments, while donor DNA remained intact in rec-2 cells, there was no intact donor DNA after an hour in wild-type cells. All the radiolabel was found in the chromosomal DNA.

Using linearized USS-containing plasmid donor DNA, Barouki and Smith (1985) nicely show that this chromosomal labeling is NOT dependent on recombination, as rec-1 mutant competent cells obtain similar levels of chromosomal labeling (rec-1 is the recA homolog in H. influenzae). Using restriction digestion, they also nicely show that in wild-type cells, only some of the donor DNA manages to recombine into the recipient chromosome.

The most important lanes in the above autoradiogram from Figure 3 of their paper are Lane E and Lane F, which show restriction-digested DNA after uptake in wild-type and rec-1. In Lane E (wild-type), the two arrowheads indicate the chromosomal restriction fragments indicative of transformation by the linear plasmid donor. In Lane F (rec-1), there is no appreciable transformation of the donor DNA into the chromosome (those restriction fragments are gone).

These results raise an interesting prospect: Could we interpret amount of recombination-independent radiolabeling relative to the recombination-dependent radiolabeling as a “food-to-sex” ratio? Undeniably, DNA is taken up and used by competent cells, but it’s clearly used in two different ways: Subunit recycling (food) and recombination (sex). By my eye, it seems like a lot more of the donor DNA, even in wild-type cells, is used for food than for sex.

Of course the amount of taken up DNA a cell could use for “sex” would be highly dependent on what DNA was taken up. In the case of the experiment above if homologous portions of the donor plasmid are removed, then the “sex” fragments disappear (Lanes I and J for wt and rec-1, respectively). Furthermore, the length of homologous DNA fragments taken up by competent cells is likely to matter, due to degradation during translocation and in the cytosol.

In Maughan and Redfield (2009), they show extensive natural variation among H. influenzae strains in the amount of uptake and transformation that competent cells will undertake. Do strains that can take up DNA well but fail to transform have a high food to sex ratio?

The assay that Barouki and Smith use doesn’t seem like the best way to measure a food-to-sex ratio, since the “sex” signal is somewhat buried behind the “food” signal. I wonder if there’s an experimental scheme that would allow one to measure such a ratio more accurately. Is there a way I could feed one strain’s chromosomes to another recipient strain and figure out how much DNA incorporated into chromosomes is recombination-dependent and how much is recombination-independent?

Anyway, thinking about this has led me to having a slightly clearer idea about the food versus sex hypotheses for the maintenance of natural competence by natural selection. Things are rarely black-and-white, so perhaps both models have their merits, but it seems like it might be possible to experimentally measure how much naturally competent cells use DNA for food or sex. Would this help in understanding these arguments?
(continued...)

Friday, July 24, 2009

Pile-ups

As I try and address the critiques of my original NIH postdoc fellowship application in my resubmission application, I started fleshing out ways I which our planned periplasmic donor DNA sequencing experiments will investigate the mechanism of DNA uptake.

I’m playing around with different figures that illustrate what I’ll do and what it might reveal. Here’s the basic idea:

(1) Incubate sheared chromosomal donor DNA with recipient competent cell preparations.
(2) Recover the DNA that is taken up in to the periplasm.
(3) Obtain paired-end sequence data for periplasmic and input DNA libraries.
(4) Compare the abundance of different sequences between the periplasmic and input DNA libraries to calculate the periplasmic uptake efficiency for sequences across the genome.

What would this data look like?

One mapped back to the genome, each paired-end read will define the span of an individual fragment from the sequenced pool.

Below, I illustrate what a few dozen of these spans would look like mapped to a short stretch of chromosome containing an uptake signal sequence in the center.

The first diagram shows what the input DNA pool would look like. The blue spans do not contain USS, while the red spans do.
The second diagram shows what a similar amount of sequencing of the periplasmic uptake DNA pool might look like. I assume that the presence of a USS motif is sufficient and necessary to strongly stimulate uptake (“all-or-nothing” model of uptake). Thus, spans containing USS would be much more abundant than spans that do not.
These sequence data could be plotted in several ways:

(1) Spanning coverage at each genomic position: The input DNA is expected to have roughly equal spanning coverage of eaach genomic position, but spanning coverage in the periplasmic uptake library is expected to be higher closer to USS motifs. As the distance between a genomic position and USS increases, fewer spans will contain both. Peaks will indicate USS, and peak height will indicate the effectiveness of individual USS loci. I estimate that one lane of sequencing the input will provide ~2500X spanning coverage per nucleotide for 500 bp donor DNA fragments.

(2) End coverage at each genomic position: If the position of USS in a fragment is irrelevant to the uptake mechanism, then plotting end coverage would have a different shape than the spanning coverage around. Since any fragment containing USS will be effectively taken up, I expect a more sawtooth-shaped distribution of end reads at USS determined by the spanning fragment length. I estimate that one lane of sequencing the input will provide 100X end coverage per nucleotide for 500 bp donor DNA fragments.

Below, I illustrate an idealized case of spanning and end coverage.
Different USS loci may behave in different fashions. Here are what two other scenarios might look like:

(1) If uptake is polarized by USS, such that the position of USS on a fragment is important, or if there was an uptake blocking sequence nearby a USS, the distribution might be skewed:
(2) If a fragment’s uptake is equally efficient with one or two USS motifs (a USS interference model), then coverage around two nearby USS might look like this:
I wonder what other mechanistic details might be found in the data...
(continued...)

Friday, July 17, 2009

Dose Response

How hungry are competent cells for DNA? I know that about a billion cells will consume ~65% of 20 nanograms tasty USS-1 fragment, but what if I offer the cells different amounts of USS-1?

To get a better hands-on feel for the DNA uptake process in wild-type and rec-2 mutant competent cells, I did a dose response experiment, where I incubated competent cells with different amounts of USS-1 DNA.

For this first experiment, I used 0.5 ml of competent cell cultures for each sample and did 6 different amounts of USS-1 DNA (12 samples total for wt and rec-2). I didn’t have enough radiolabeled fragment for all of my desired concentration, so I mixed in some cold USS-1 DNA to make up the difference. I let the DNA and cells incubate for 30 mins, then I washed the cells several times and determined the total radioactive counts in the cell pellet and washes to determine the % uptake and total uptake.

Here’s the results:

Total Uptake:

Percent uptake:

Interestingly, rec-2 does better at low concentrations of DNA than wild-type, but worse with high concentrations. The latter could be due to the periplasm getting too clogged with DNA, such that the outer membrane uptake machinery has to work too hard to get more DNA through, while in wild-type translocation of DNA frees up space in the periplasm. But the former (higher uptake in rec-2 at low DNA concentrations) doesn't really make much sense to me. Maybe not all free nucleotides created during degradation at the inner membrane remain in the cell, so that at low concentrations, rec-2 simply holds more label?
(continued...)

Sunday, July 12, 2009

Standing upon the shoulders of giants

It really is gratifying to have things work the way they're supposed to. Some kind of bug bit me on Saturday and I came in to see if the periplasmic DNA preparation reported by Kahn et al 1983 would work in my hands. And sure enough it did!

The experiment was much the same as before. I added radiolabeled USS-1 fragments to either wild-type or rec-2 competent cell preps, incubated for 5 minutes, and then either prepared total DNA or did the periplasmic extraction (TE/1.5M CsCl + phenol/acetone, 1:1).

Since wild-type cells will take up the fragment, but also incorporate labeled subunits from degradation of taken up DNA, I can tell if the periplasmic DNA prep managed to exclude chromosomal DNA. But first, I counted the radiolabel present in the different cellular fractions...

This time, about a quarter of the USS-1 fragment added was taken up within five minutes (wild-type: 26%; rec-2: 28%). I suspect these numbers are lower than the last time I did it, because my five minutes was really five minutes (whereas the first time, I think I was 2-3 minutes late).

The extraction: When I collected the aqueous phase, I also collected the organic phase, and the interface between the phases (which should contain the cells minus their outer membranes). I counted the radiolabel in these different fractions as before:
Wild-type cells had label in both the aqueous extract, as well as in the interface containing the cells, while rec-2 had nearly all the label in the aqueous extract. The organic phase had less than 1% rec-2.

But here's the important bit:
Lane 1: Input (1/3, or 4 ng)
Lane 2: Total DNA, wild-type + USS-1 for 5 min.
Lane 3: Total DNA, rec-2 +USS-1 for 5 min.
Lane 4: Peri DNA, wild-type + USS-1 for 5 min.
Lane 5: Peri DNA, rec-2 +USS-1 for 5 min.

The important point here is that in the total DNA extract of wild-type, both intact donor USS-1 and chromosomal labeling are evident, while in the peri-extract of wild-type, there is no chromosomal label.

This means that the extraction I did successfully purified periplasmic DNA over chromosomal DNA. Fabulous!

Now I need to scale this protocol up, and get cleaner DNA (i.e. use RNase), so hopfully I can see this without using radiolabel. If I can really get clean periplasmic DNA with little or no chromosomal contamination, I will move onto doing the "real" experiment with donor DNA made up from sheared genomic DNA of another isolate.

Yay!
(continued...)

Friday, July 10, 2009

Building a periplasm prep...

After my failed attempts at doing a large-scale periplasm prep right off the bat, I decided to spend this week going a bit more slowly. I repeated what others have already done successfully using radio-labeled DNA fragments as donors. This means that I can do smaller scale experiments and don't need particularly pure DNA.

And this time the experiments all worked. Here's what I did:

(1) I made competent cells of KW20 (RR722) and KW20 rec-2 (RR622). I confirmed that the wild-type strain transformed normally and the rec-2 strain not at all (or at least below my limit of detection). This confirmed that my competent cell preps were okay, and that the rec-2 strain seems to be correct.

(2) Following the DNA uptake assay protocol of Maughan and Redfield, 2009, I showed that USS-1 is taken up very well, but USS-R is only poorly taken up. To do this, I simply incubated ~12 ng of either radio-labeled USS-1 or USS-R with 0.5 ml of competent cells for 20 minutes, and then compared the radioactive counts in a washed cell pellet compared to the total counts:
Wild type and rec-2 both take up USS-1 well, but USS-R poorly, as expected. But rec-2 seems to take up USS-1 slightly better than wild type. This is also true in the next experiment. This may be significant but could also reflect slight differences in the competent cell prep of the two strains.

Possibly the coolest part of this for me was that I got numbers that were spot-on the former post-doc's numbers (found in her notebook) and older papers describing % uptake. That is: ~65% uptake for ~20 ng / ml of cells. This was very encouraging to me.

(3) I repeated the uptake assay described above using ~12 ng USS-1 donor DNA and incubated wild-type and rec-2 cells for either 5 or 60 minutes. This gave results similar to those shown above:
Most uptake was finished after only 5 mins, though additional incubation increased the level of uptake. The results were nearly identical for 60 min incubation as for 20 min incubation, so I don't need to do it for so long.. The rec-2 strain again showed slightly more uptake at all time points.

After this, I took it a step further: I also extracted the DNA from the cell pellets and ran them out on a gel. I also included the input donor DNA as a control. I dried down the gel and exposed it to a phosphor screen. This is what the gel looked like:


Lane 1: Donor DNA (50% of input; ~6 ng)
Lane 2: Wild-type + USS-1 for 5 min. Total DNA.
Lane 3: Wild-type + USS-1 for 60 min. Total DNA.
Lane 4: rec-2 + USS-1 for 5 min. Total DNA.
Lane 5: rec-2 + USS-1 for 60 min. Total DNA.

Alright! That's exactly what I hoped for! (Well, not quantitatively between lanes: this was a sloppy first experiment.) The gel shows that the natural competence phenotypes of the two strains: wild-type and rec-2.

Intact uptake DNA is the smaller (lower) band, while chromosomal DNA is the high molecular weight species. In wild type, donor DNA gets degraded and nucleotides can be incorporated into the chromosome over time. (Importantly, the labeling of the chromosome is NOT from transformation, but from incorporation of degraded nucleotides into the genome by DNA replication.) In rec-2, the radio-labeled donor DNA is trapped in the periplasm and isn't degraded. So there is no chromosomal labeling in this case.

This is effectively a repeat of an experiment from Barouki and Smith, 1985.

Next, I'll try exactly the same thing, but I'll also try the extraction from Kahn et al., 1983. If this successfully yields pure periplasmic DNA, then I expect that the extraction will not yield radiolabeled chromosome, even for the wild type sample. If that works, I can work on scaling the protocol up to do a real purification of uptake DNA.

Onward!
(continued...)

Tuesday, July 7, 2009

Imagine it exists, and maybe it does!

Our proposed experiments involve capturing DNA molecules at the different stages of natural transformation. One of the technical challenges we face will be producing a library of DNA molecules that have been translocated into the cytosol. We have some schemes for how we’ll do the purification of donor DNA from the cytosol, but even assuming that this works wonderfully, we still need to turn these into double-stranded DNA. We can’t use a specific primer to the 3’ends of translocated ssDNAs, because (a) we don’t know the exact 3’ ends and (b) it will be a complex mixture.

What to do?

Until now the only thing that had occurred to me is to use random priming of our ssDNA to convert cytosolic ssDNA into dsDNA (shown schematically above), but this approach has several limitations. The biggest problem is that we would only be able to accurately identify the 5’-end of translocated DNA. The 3’-end of the final dsDNA we produce would not represent the 3’-ends of the original ssDNA molecules using random primers. Furthermore, we would not know which end of our dsDNA was the original ssDNA’s 5’ or 3’ end. And finally, we would end up with a highly heterogeneous size distribution, which might complicate sequencing.

How can I circumvent this, get both ends, and know which is which? I need a strategy like RACE. I decided to imagine that a certain enzyme existed that might help me in this endeavor and then see if it actually existed and was already was commercially available. This strategy has worked for me in the past: Once, I’d wanted to know if there were restriction enzymes that only nicked at their recognition sites, so I typed “nickase neb” into Google, and sure enough NEB carries nickases! Go biotechnologists!

This time, I want to tack some type of single-stranded adaptor sequence onto the 3’ ends of my putative cytosolic ssDNAs, so I typed “ssDNA ligase” into Google, and Presto!... Epicentre produces a single-stranded ligase that they call CircLigase. Sweet!

This doesn’t fully solve the problem, since the ligase will normally take an ssDNA and circularize it (since the intra-molecular ligation will usually be favored). This is useful to plenty of folks who are interested in doing rolling circle amplification and rolling circle transcription, but I would rather not circularize my ssDNAs, but would like to favor ligation of an ssDNA adaptor specifically to the 3’-end. This will require a couple of bells and whistles.

If we take our ssDNA and then treat it with a phosphatase, we can rid the 5’-end of its terminal phosphate and both block circular ligation, as well as ligation of our adaptor to the 5’end. If our adaptor oligonucleotide also has a protected 3’end (not sure how to do this... an oligo with a terminal dideoxy nucleotide?), then we’d block the ligation of the adaptors to each other and force ligation only in the orientation we want (5’ of the adaptor to 3’ of the target).

Then, using a primer complementary to the adaptor, we can convert full-length ssDNA into dsDNA. Furthermore, the adaptor marks the original 3’ end of the fragment, so we can give a polarity to our cytosolic fragments. Here’s the scheme:

Afterwards, of course, we’d need to either amplify this product or de-protect both ends, so that we could ligate sequencing adaptors to the mixture.

This plan just might work, and I could make sure it works using defined substrates, rather than precious (as well as non-existent) cytosolic DNA fractions. The main thing I can’t think of off the top of my head is getting a hold of an oligo with a protected 3’-end (preferably reversibly so).

UPDATE: Looks like at least some oligo companies can include dideoxy bases in oligos. Awesome. This is not quite as ideal as a reversible protection of the 3' end...

UPDATE 2: Uh oh. How to amplify the product? There's no primer sequence at one end... This could involve a phosphorylation step and a ligation of a normal adaptor to that end? That's an extra unfortunate step. Also, the adaptor sequences will eat into the sequence read length, but that should be acceptable, since we only need to get tag sequences.
(continued...)

Thursday, July 2, 2009

Mutation versus Transformation: Structural Variation

Here's an idea: Compare structural changes (indels and rearrangements) in transformed and untransformed cultures.

One major goal of our planned experiments is to measure the transformation rate across the genome. This is a fairly ambitious prospect, mainly because of the amount of sequencing it will require.

If we assume that the average donor allele transforms recipient chromosomes 1% of the time, then we would need to sequence the average locus 100 times to see the donor allele just one time. But to get a reliable measurement of transformation rate would require considerably more... perhaps 10,000 times. This would give us an average of 100 donor alleles / 10,000 alleles sequenced. Even using the Illumina platform would get fairly expensive to measure the transformation rate for every SNP.

One way we could do a preliminary sequencing experiment would be to ignore single-nucleotide differences between donor and recipient and focus solely on indel and rearrangement differences (a.k.a. structural variation). This data would speak to interesting hypotheses regarding the role of natural competence in maintaining the core genome and diversifying the accessory genome, and it would require considerably less sequencing.

So the goal of the preliminary experiment would be to measure structural variation due to mutation versus structural variation due to transformation.

How would this work, and why would it be cheap and easy?

First of all, the experiment-side is extremely easy. Naturally competent recipient cultures would be split in two. One part would be used to prepare untransformed recipient chromosomes; the other part would be incubated with donor DNA for a while, allowed to recover, and then transformed chromosomes would be purified. (We could also include a selection for a donor marker at this point to increase the relative transformation rate.) That’s it.

The sequencing side would also be comparatively easy. Untransformed and transformed chromosome preparations would be sheared, end-repaired, and size-fractionated by gel to 500 bp (as precisely as possible). This DNA would be ready for sequencing library construction and paired-end sequencing.

For a 500 bp library, I previously estimated ~2500X “spanning coverage” in one lane of Illumina sequencing using conservative estimates of the sequencing parameters. “Spanning coverage” is defined as how many times a particular genomic position is found between two mapped paired-end reads. So we’d get to 10,000X spanning coverage in ~1/2 a full run.

How would this help us measure the transformation of donor structural variants into the recipient?

Let’s use an example to illustrate. In the alignment shown above (using GenomeMatcher), the donor genome (86-028NP) is shown on top, while the recipient genome is shown on bottom (Kw20). As is pretty clear, genes Hi_0512 and Hi_0513 are absent from the donor genome, indicating a deletion of those genes from the donor (or possibly an insertion into the recipient). These genes happen to be the HindII restriction enzyme and methylase. The flanking genes are syntenic (so Hi_0511 = tchA, etc.).

Since we know the size-distribution of the library (500 bp), it would be quite simple to spot the deletion allele. Paired-end reads with one end in tchA and the other in rpoC would define the deletion. By contrast the insertion allele would always have tchA and rpoC mappings on different fragments (with the other ends in the HindII methylase or restriction enzyme).

Here’s a way of illustrating what paired-end reads of different kinds of alleles relative to the recipient would look like:
So for our deletion, we’d see paired-ends that mapped to positions further apart than they should be. By having extremely high spanning coverage, we could count how often these kinds of mappings occurred versus the recipient mappings. This would give us the rate of deletion.

In our untransformed library, if we saw the deletion allele, we’d be seeing mutation, while in the transformed library we’d be seeing mutation and/or transformation. Since we know the donor genome sequence, we can distinguish paired-end reads that look like the donor sequence versus de novo mutations that occurred when we grew out the cells. Better still, we could spot transformation-induced de novo mutations by comparing to the untransformed library.

Why else do the untransformed chromosomes at all? Well, structural mutations are likely to occur at a much higher rate than single-nucleotide mutations in many instances. Independent of our interest in natural transformation, doing the control experiment may reveal regions of the genome that are unstable, along with the mutation rate of different types of structural variation. This last part is non-trivial. If, for example, we see a particular deletion that occurred on 50% of the untransformed chromosomes we looked at, it could be that this is an extremely frequent mutation, but it could also have simply occurred early in the grow-out of the culture.

As a control for transformation rates, we don’t have to worry about that, but doing the untransformed control would make us confident that changes we saw were induced by transformation and not simply due to such kinds of mutation.
(continued...)

Happy (Belated) Canada Day!



(Image: The Canadian-built robot arm attaching the space shuttle docked to the Hubble Space Telescope with the Earth in the background.)

Oh yeah, and my second attempt at preparing periplasm DNA was... inconclusive. But it was pretty interesting to try out. In particular, the TE/CsCl/phenol/acetone extraction was quite compelling visually, involving small bubbles breaking up and reforming. I need to start these experiments out at a smaller scale.

And luckily, we've received radiolabeled dATP, so I can do some more sensitive and controlled experiments next week. The use of radiolabeled uptake fragments will be significantly more sensitive and allow me to use small cultures and follow small amounts of uptake DNA.

Once I've got a functioning uptake assay, I can work out the best purification method and scale up from there.
(continued...)