Showing posts with label plans. Show all posts
Showing posts with label plans. Show all posts

Monday, September 28, 2009

Mismatch repair versus Segregation











Things have gone swimmingly with my strain construction plans, and indeed today I am extracting DNA that will presumably be sequenced. To recap, I made a couple of clinical isolates (86-028NP and PittGG) resistant to novobiocin (NovR) by transforming them with a bit of left-over NovR allele of the former postdoc. I then isolated the new strains’ DNA, and used these to transform the standard KW20 Rd strain. By selecting for NovR, we can be certain that the clones I pick took up DNA and recombined it into their genomes.

One technical issue arose, however, which required a little bit of thought: Should I have streaked for single colonies? I.e. once I had my transformants, it might be a good idea to streak out individual colonies to make sure I purified them away from any background or broke apart any doublet colonies. No big deal, but after talking it out with Rosie, we decided to skip it. Why? So that we might get lucky and distinguish recombination followed by mismatch repair versus recombination followed by segregation. In the following figures, I illustrate what I mean by this…

In this first one, the donor DNA is shown in red, and the recipient chromsome is shown in two colors, blue and green, to distinguish the strands. The lowercase letters indicate polymorphic sites in the donor genome. Little a is meant to be the selectable marker, in this case an allele of gyrB:
Donor DNA is incubated with competent recipient cells, and recombination of single-stranded DNA leaves patches of heteroduplex in the genome, shown as small red patches on either the blue or green strands.

After this, the cells have a chance to perform mismatch correction to fix any heteroduplex. I select for cells that have little a by plating to novobiocin plates, so only cells that end up a/a will survive an make colonies. (I am not going to show any examples of restoration repair, in which donor alleles are repaired back into recipient alleles… this will be invisible in our analysis.)

In the below example, I show the A/a and B/b heteroduplexes getting mismatch repaired into a/a and b/b, whereas C/c and D/d heteroduplexes remain unrepaired (they escape correction). What will happen in such as case is the generation of a sectored colony, in which (in principle) half the cells would have one genotype and the other half a different genotype:
In the above example, the original transformant segregates the c and d alleles into different cells, while a and b end up in all cells. If the whole resulting colony is grown up and sequenced, the a and b alleles will be the only ones observed, while at the other two loci, there will be a mix of C and c, along with a mix of D and d. We wouldn’t be able to tell “phase”, i.e. whether c and d were on the same or different chromosomes, unless we did streak for singles and the sequenced several clones. But as a first pass, this could be a really interesting analysis. It will also serve as excellent proof-of-principle for our more intense sequencing plans.

There is a caveat, however, which means we need to get a little bit lucky to be able to distinguish these phenomena (mismatch repair versus segregation). We won’t see two different genotypes, if the A/a heteroduplex isn’t mismatch corrected:
The issue isn’t that segregation didn’t happen; the problem is that one of the segregants dies under selection for little a.

Thus, if we see a pure genotype, then either all mismatches were corrected, or our selectable marker didn’t mismatch correct.

When I pre-screen my transformants to make sure they’re not spontaneous mutants, I might be able to pick a colony where I think segregation is occurring. If I get the standard sequencing traces back and see mixed bases in the chromatograms that corresponde to donor and recipient alleles, I’ll pick that kind of clone for sequencing…

One sort of sad note here, in terms of the more distant future, is that mismatch repair mutants, which should be quite useful for understanding transformation, will need to be transformed without selection if we hope to recover isolated segregants from individual transformants.
(continued...)

Thursday, September 17, 2009

The Last Straw

Yesterday, Rosie kindly ran her Perl script over the USS construct I designed. The final thing I was worried about was whether or not my design had any USS or USS-like sequences in it, other than the one it's supposed to have. I'd checked the construct for any core USS motifs (5'-AAGTGCGGT-3'), but since we think that the motif is more complex than this, it was important to make sure that there were no extra sequences that got high scores using the USS position-weight matrix.
Fortunately, the construct looks good, so I can go ahead and order the control oligos and have high expectations that they'll work...

Here's how every 32 base pair window over the 199mer looks when scored with the USS PWM:
There's a single prominent high-scoring site right where it should be, and all of the surrounding area scores near background. The USS in the construct has a score (~10^-8) more than 10 orders of magnitude better than the next best sites. There's a slight increase for windows immediately adjacent to the USS, presumably because the AT-tracts in the USS are still contained in those windows. The rest of the construct only has scores at background.

Just to show that these other sites really do represent background levels of USS score, Rosie also ran a randomized version of the sequence:
Nothing better than 10^-18. Excellent.
(continued...)

Scale UP!

How much periplasmic DNA can I hope to get using my current protocols, and how much DNA will I need? As promised, here are some rough calculations regarding the oligo purchases that we want to make.

I used a molecular weight calculator available on-line to determine the size of the dsDNA I described in the last post.
I alternatively could’ve used Rosie’s Universal Constants (660 g / mol of base pair and 10^-18 g / single 1 kb DNA molecule) to make this calculation, but since I’m dealing with a known sequence, I might as well get an exact molecular weight. (I also made a minor mistake in the last post, and the molecule I describe is actually only 199 bp).

So, for our USS molecule, MW = 122,828.6 g / mol. And the oligo synthesis service we’re planning on using will be at the 1 micromole scale. That means if we took all of the two oligos, annealed, extended, and purified, we’d end up with 0.123 grams of input DNA pool! That’s really a very large amount.

My previous concerns about needing to do PCR to maintain the pool are unfounded. This scale should be sufficient for hundreds (or even thousands) of experiments...

What follows are my preliminary assumptions about yields from the periplasmic DNA prep. They are based on several different experiments, though I am erring on the side of being conservative with my estimates and are guides for future experiments only. In this post, I will address the issues of scale-up at the end; before that I’ll just refer to the approximate total culture volume and amount of DNA that I’d need to get a target amount of DNA, assuming all else works perfectly.

So, I’ve now done several experiments using a PCR fragment bearing the consensus USS, called USS-1. If I add 20 ng DNA / 1 ml competent cells, ~50% is taken up. That is, in rec-2 cells, my theoretical yield of periplasmic DNA is 10 ng. My actual yield is considerably lower; as evaluated by my radiolabeling experiments, I estimate I get ~25% of my theoretical maximum.

This means,

1 ml cells + 20 ng DNA → 2.5 ng recovered.
20 ml cells + 400 ng DNA → 50 ng.
40 ml cells + 800 ng DNA → 100 ng

But this is only for the consensus sequence. Our real experiments will be a mix of molecules, some of which will be efficiently taken up and others that won’t. For a cursory estimate, we might assume that ~50% of fragments will be “good” USS and the other half will be “bad”. This would further reduce the yield.

That means, I am likely to need ~80 ml cultures and a starting input DNA amount of ~1600 ng, just to get back a mere 100 ng of DNA back!

Most Illumina sequencing centers seem to want ~1 ug of DNA to make libraries, but a lot of ChIP-seq experiments seem to call for only ~100 ng. In our case, there will be no downstream library construction, so we can likely get away with small amounts of DNA, as long as it is quite pure and accurately quantified.

Regardless, this is going to take fairly large cultures, fairly large amounts of DNA, and a good scaled-up periplasmic prep.

BUT, one important thing to note is that our degenerate oligo preparation will be more than sufficient for a large number of experiments, even at this large scale. For the controls, I can merely buy minimum-scale synthesis long oligos at ~$200 a pop. Since I can safely PCR amplify these, I will be able to make a replenishable stock for use in scale-up experiments.

More on this in the future, but while I’m doing this, I might as well estimate what it will take to get a microgram of chromosomal DNA fragments out of competent cell periplasms.

My previous experiments with sonicated DNA gave pretty consistent DNA uptake measurements:

~50% of 200 ng 1-10kb DNA / 1 ml cells → 100 ng max. yield.
~10% of 200 ng 0.2-0.4kb DNA / 1 ml cells → 20 ng max. yield.

Given a 25% recovery rate from the periplasm, this means that for a microgram of DNA, I will need:

1-10kb DNA: 8 micrograms in a 40 ml culture
0.2-0.4kb DNA: 40 micrograms in a 200 ml culture (!)

This last is really asking a lot. That size of scale-up will require special thought…

Appendix on Scale-up Issues:
  1. Purity: I have not been adding RNase. I need to get all the RNA away, in order to accurately quantify the DNA. I am also concerned about salt. The CsCl in my DNA precipitates may not be getting washed out adequately by a single 80% ethanol wash.
  2. Cell concentration: It would help for technical reasons, if I could concentrate the cells quite a bit before doing the organic extractions. I have used a ratio of 1:1, cells : organic solvents. So a 1 ml competent cell prep (~a billion cells) gets mixed with 1 ml solvent. But I might be able to resuspend 10 ml of cells in 1 ml and then use 1 ml solvent. I just don’t know.
  3. DNA concentration: I want to make sure that I am saturating with DNA for my initial experiments, but I haven’t yet done a proper saturation curve to know what I should be using. This will decrease the total efficiency of DNA uptake, but my total yields will be higher, and I will be biasing things towards the best uptake sequences (which is a good place to start).
  4. DNase: I have not been treating cells with DNase prior to isolation. From what I can tell, this is not a problem, and the free DNA is washed away. But if I use very high DNA concentrations, I will probably want to use DNase, just to be sure I’m eliminating free DNA completely.
  5. Details, details: Scale-up is never quite as simple as just increasing the volume of everything. I will need to make sure that there are appropriate centrifuges, shakers, tubes, and everything else. Growth rates of cells and competence induction may be poor when going to larger volume cultures. I am also concerned about scaling up the organic extractions. It turns out that not all conicals are created equal; I’ve had disasters where the phenol has torn through the bottom of 50 ml conicals when doing large-scale organic extractions, depending on the brand of conical and rotor used. I’ll need to make sure that things like this don’t happen in advance before I mess up somebody else’s equipment!

(continued...)

Monday, August 10, 2009

Back to the grind

Whew! Getting that albatross (= grant application) off my neck feels good, but now I need to formulate a plan for the next several weeks of lab work. What are my priorities? There’s quite a lot of things I could be doing, so I might as well make a list...

(A) Generate purified DNA from 86-028NP, PittEE, and PittGG donor DNA: Should be pretty easy. I tried this a couple times in the past couple weeks, but had a few problems:
  1. Got a goopy mess. I don’t know if it was my lysis or something about my extractions, but when I would try to pull the aqueous layers off of my phenol/chloroform extractions, this goopy junk at the interface kept interfering with my ability to complete the preps. I will try try again, but this next time I will do a few things differently: (i) start with fewer cells per volume, (ii) do a straight phenol extraction first, (iii) let lysis proceed longer.
  2. PittEE didn’t grow. I’d gone to the lab stocks, streaked for single colonies, grown up overnight cultures from single colonies, and stored them in glycerol in the -80. But when I returned to PittEE frozen stocks, they don’t grow up! Weird. Maybe PittEE is particularly bad at surviving outside log-phase, so I had merely frozen down a pile of dead cells. I’ll have to return to the lab stocks and restreak for single colonies.
  3. I never checked the putative nalidixic acid resistant phenotype of 86-028NP. I really hope that marker will work to select for transformants from 86-028NP to KW20, as this would really help several steps of the project. In our strain database, 86-028NP is listed as NalR, but this might only be a clinically relevant phenotype. I need to check it on plates.

(B) Check the uptake and transformation phenotypes of WT, rec-2, rec-1, and pilB: I’ve already gone through this exercise with wt and rec-2, but as a negative control, pilB mutants should neither take up DNA nor be transformed by it, and for my translocation experiments, rec-1 should take up DNA normally but fail to transform. I have all these strains frozen down as M-IV competent cells, so I just need to do the experiments to verify that my mutants have the expected phenotypes. The rec-1 mutant in particular is important, since I plan to use it for cytosolic DNA purifications, but it could be somewhat leaky, something I need to know. (The only rec-1 mutant available is the original isolated mutant from back in 1972. It’s older than me!

(C) Quantify uptake and transformation for different-sized donor DNA: A coupe weeks ago, I used a “Bioruptor” to sonicate MAP7 DNA to different size distributions, and I tested the transformation rate of antibiotic resistance into competent cells of some of these different sized DNA. I also want to check how well they’re taken up. Once I’ve done this, I’ll report back the full set of results. The expectation is that while even small fragments will be taken up well, only large fragments will transform well. There’s also an interesting expected relationship between the amount of uptake expected and the size distribution of the DNA. If the DNA fragments are quite small, then only a fraction are expected to contain uptake signals, but many fragments containing such signals will be taken up. On the other side, large DNA fragments will mostly all contain uptake signals, but only a few molecules will be taken up, but these will be larger... Hmmm....

(D) Design the degenerate USS construct. This will require its own post, but suffice it to say, we have a pretty good plan on how to design this construct so that we will be able to directly sequence the degenerate USS motif by short single-end Illumina sequencing without any processing steps upfront.

(E) Transformation experiment: If indeed 86-028NP can grow on nalidixic acid plates, then I will just go ahead and do half of the experimental-side of one of my specific aims. All I need to do is transform KW20 competent cells (which I have) with DNA taken from 86-028NP (which I’ll make), select for colonies that are NalR, pick several of them, grow them up, and extract their DNA. This would provide up with plenty of material to produce sequencing libraries and then send to sequencing. Then we can really get this project off the ground.

(F) Find somewhere that will do sequencing with a short turn-around: The local Illumina sequencers (at UBC’s Genome Center) offer what looks like excellent sequencing services for a reasonable cost, but their turn-around time is a bit too slow for us to hope to have any preliminary data for Rosie’s grant applications. If we want to get anything done more quickly, we will have to find someone else to get started.

(G) Design and produce a large DNA fragment from the 86-028NP genome to use for developing a cytosolic ssDNA prep: Since the most difficult experimental part of our plans is likely to be purifying (and cloning) ssDNA from the cytosol, I’d like to start with a defined construct I can use for purification experiments. I think that a good choice would be a large clone from 86-028NP, because this might be useful for other experiments as well. I will want something at least several kb on a plasmid. I will need to think about whether or not it matters what this fragment contains. As a first guess, I would want it to contain the putative NalR allele and several kb of flank.

What else should I be doing?
(continued...)

Thursday, August 6, 2009

Almost...


...there... However it turns out, I'll finally have this grant I've been working on out of my hands tomorrow. I'm actually thrilled to get back into the lab when it's all over...

(continued...)

Friday, July 24, 2009

Pile-ups

As I try and address the critiques of my original NIH postdoc fellowship application in my resubmission application, I started fleshing out ways I which our planned periplasmic donor DNA sequencing experiments will investigate the mechanism of DNA uptake.

I’m playing around with different figures that illustrate what I’ll do and what it might reveal. Here’s the basic idea:

(1) Incubate sheared chromosomal donor DNA with recipient competent cell preparations.
(2) Recover the DNA that is taken up in to the periplasm.
(3) Obtain paired-end sequence data for periplasmic and input DNA libraries.
(4) Compare the abundance of different sequences between the periplasmic and input DNA libraries to calculate the periplasmic uptake efficiency for sequences across the genome.

What would this data look like?

One mapped back to the genome, each paired-end read will define the span of an individual fragment from the sequenced pool.

Below, I illustrate what a few dozen of these spans would look like mapped to a short stretch of chromosome containing an uptake signal sequence in the center.

The first diagram shows what the input DNA pool would look like. The blue spans do not contain USS, while the red spans do.
The second diagram shows what a similar amount of sequencing of the periplasmic uptake DNA pool might look like. I assume that the presence of a USS motif is sufficient and necessary to strongly stimulate uptake (“all-or-nothing” model of uptake). Thus, spans containing USS would be much more abundant than spans that do not.
These sequence data could be plotted in several ways:

(1) Spanning coverage at each genomic position: The input DNA is expected to have roughly equal spanning coverage of eaach genomic position, but spanning coverage in the periplasmic uptake library is expected to be higher closer to USS motifs. As the distance between a genomic position and USS increases, fewer spans will contain both. Peaks will indicate USS, and peak height will indicate the effectiveness of individual USS loci. I estimate that one lane of sequencing the input will provide ~2500X spanning coverage per nucleotide for 500 bp donor DNA fragments.

(2) End coverage at each genomic position: If the position of USS in a fragment is irrelevant to the uptake mechanism, then plotting end coverage would have a different shape than the spanning coverage around. Since any fragment containing USS will be effectively taken up, I expect a more sawtooth-shaped distribution of end reads at USS determined by the spanning fragment length. I estimate that one lane of sequencing the input will provide 100X end coverage per nucleotide for 500 bp donor DNA fragments.

Below, I illustrate an idealized case of spanning and end coverage.
Different USS loci may behave in different fashions. Here are what two other scenarios might look like:

(1) If uptake is polarized by USS, such that the position of USS on a fragment is important, or if there was an uptake blocking sequence nearby a USS, the distribution might be skewed:
(2) If a fragment’s uptake is equally efficient with one or two USS motifs (a USS interference model), then coverage around two nearby USS might look like this:
I wonder what other mechanistic details might be found in the data...
(continued...)

Tuesday, July 7, 2009

Imagine it exists, and maybe it does!

Our proposed experiments involve capturing DNA molecules at the different stages of natural transformation. One of the technical challenges we face will be producing a library of DNA molecules that have been translocated into the cytosol. We have some schemes for how we’ll do the purification of donor DNA from the cytosol, but even assuming that this works wonderfully, we still need to turn these into double-stranded DNA. We can’t use a specific primer to the 3’ends of translocated ssDNAs, because (a) we don’t know the exact 3’ ends and (b) it will be a complex mixture.

What to do?

Until now the only thing that had occurred to me is to use random priming of our ssDNA to convert cytosolic ssDNA into dsDNA (shown schematically above), but this approach has several limitations. The biggest problem is that we would only be able to accurately identify the 5’-end of translocated DNA. The 3’-end of the final dsDNA we produce would not represent the 3’-ends of the original ssDNA molecules using random primers. Furthermore, we would not know which end of our dsDNA was the original ssDNA’s 5’ or 3’ end. And finally, we would end up with a highly heterogeneous size distribution, which might complicate sequencing.

How can I circumvent this, get both ends, and know which is which? I need a strategy like RACE. I decided to imagine that a certain enzyme existed that might help me in this endeavor and then see if it actually existed and was already was commercially available. This strategy has worked for me in the past: Once, I’d wanted to know if there were restriction enzymes that only nicked at their recognition sites, so I typed “nickase neb” into Google, and sure enough NEB carries nickases! Go biotechnologists!

This time, I want to tack some type of single-stranded adaptor sequence onto the 3’ ends of my putative cytosolic ssDNAs, so I typed “ssDNA ligase” into Google, and Presto!... Epicentre produces a single-stranded ligase that they call CircLigase. Sweet!

This doesn’t fully solve the problem, since the ligase will normally take an ssDNA and circularize it (since the intra-molecular ligation will usually be favored). This is useful to plenty of folks who are interested in doing rolling circle amplification and rolling circle transcription, but I would rather not circularize my ssDNAs, but would like to favor ligation of an ssDNA adaptor specifically to the 3’-end. This will require a couple of bells and whistles.

If we take our ssDNA and then treat it with a phosphatase, we can rid the 5’-end of its terminal phosphate and both block circular ligation, as well as ligation of our adaptor to the 5’end. If our adaptor oligonucleotide also has a protected 3’end (not sure how to do this... an oligo with a terminal dideoxy nucleotide?), then we’d block the ligation of the adaptors to each other and force ligation only in the orientation we want (5’ of the adaptor to 3’ of the target).

Then, using a primer complementary to the adaptor, we can convert full-length ssDNA into dsDNA. Furthermore, the adaptor marks the original 3’ end of the fragment, so we can give a polarity to our cytosolic fragments. Here’s the scheme:

Afterwards, of course, we’d need to either amplify this product or de-protect both ends, so that we could ligate sequencing adaptors to the mixture.

This plan just might work, and I could make sure it works using defined substrates, rather than precious (as well as non-existent) cytosolic DNA fractions. The main thing I can’t think of off the top of my head is getting a hold of an oligo with a protected 3’-end (preferably reversibly so).

UPDATE: Looks like at least some oligo companies can include dideoxy bases in oligos. Awesome. This is not quite as ideal as a reversible protection of the 3' end...

UPDATE 2: Uh oh. How to amplify the product? There's no primer sequence at one end... This could involve a phosphorylation step and a ligation of a normal adaptor to that end? That's an extra unfortunate step. Also, the adaptor sequences will eat into the sequence read length, but that should be acceptable, since we only need to get tag sequences.
(continued...)

Thursday, July 2, 2009

Mutation versus Transformation: Structural Variation

Here's an idea: Compare structural changes (indels and rearrangements) in transformed and untransformed cultures.

One major goal of our planned experiments is to measure the transformation rate across the genome. This is a fairly ambitious prospect, mainly because of the amount of sequencing it will require.

If we assume that the average donor allele transforms recipient chromosomes 1% of the time, then we would need to sequence the average locus 100 times to see the donor allele just one time. But to get a reliable measurement of transformation rate would require considerably more... perhaps 10,000 times. This would give us an average of 100 donor alleles / 10,000 alleles sequenced. Even using the Illumina platform would get fairly expensive to measure the transformation rate for every SNP.

One way we could do a preliminary sequencing experiment would be to ignore single-nucleotide differences between donor and recipient and focus solely on indel and rearrangement differences (a.k.a. structural variation). This data would speak to interesting hypotheses regarding the role of natural competence in maintaining the core genome and diversifying the accessory genome, and it would require considerably less sequencing.

So the goal of the preliminary experiment would be to measure structural variation due to mutation versus structural variation due to transformation.

How would this work, and why would it be cheap and easy?

First of all, the experiment-side is extremely easy. Naturally competent recipient cultures would be split in two. One part would be used to prepare untransformed recipient chromosomes; the other part would be incubated with donor DNA for a while, allowed to recover, and then transformed chromosomes would be purified. (We could also include a selection for a donor marker at this point to increase the relative transformation rate.) That’s it.

The sequencing side would also be comparatively easy. Untransformed and transformed chromosome preparations would be sheared, end-repaired, and size-fractionated by gel to 500 bp (as precisely as possible). This DNA would be ready for sequencing library construction and paired-end sequencing.

For a 500 bp library, I previously estimated ~2500X “spanning coverage” in one lane of Illumina sequencing using conservative estimates of the sequencing parameters. “Spanning coverage” is defined as how many times a particular genomic position is found between two mapped paired-end reads. So we’d get to 10,000X spanning coverage in ~1/2 a full run.

How would this help us measure the transformation of donor structural variants into the recipient?

Let’s use an example to illustrate. In the alignment shown above (using GenomeMatcher), the donor genome (86-028NP) is shown on top, while the recipient genome is shown on bottom (Kw20). As is pretty clear, genes Hi_0512 and Hi_0513 are absent from the donor genome, indicating a deletion of those genes from the donor (or possibly an insertion into the recipient). These genes happen to be the HindII restriction enzyme and methylase. The flanking genes are syntenic (so Hi_0511 = tchA, etc.).

Since we know the size-distribution of the library (500 bp), it would be quite simple to spot the deletion allele. Paired-end reads with one end in tchA and the other in rpoC would define the deletion. By contrast the insertion allele would always have tchA and rpoC mappings on different fragments (with the other ends in the HindII methylase or restriction enzyme).

Here’s a way of illustrating what paired-end reads of different kinds of alleles relative to the recipient would look like:
So for our deletion, we’d see paired-ends that mapped to positions further apart than they should be. By having extremely high spanning coverage, we could count how often these kinds of mappings occurred versus the recipient mappings. This would give us the rate of deletion.

In our untransformed library, if we saw the deletion allele, we’d be seeing mutation, while in the transformed library we’d be seeing mutation and/or transformation. Since we know the donor genome sequence, we can distinguish paired-end reads that look like the donor sequence versus de novo mutations that occurred when we grew out the cells. Better still, we could spot transformation-induced de novo mutations by comparing to the untransformed library.

Why else do the untransformed chromosomes at all? Well, structural mutations are likely to occur at a much higher rate than single-nucleotide mutations in many instances. Independent of our interest in natural transformation, doing the control experiment may reveal regions of the genome that are unstable, along with the mutation rate of different types of structural variation. This last part is non-trivial. If, for example, we see a particular deletion that occurred on 50% of the untransformed chromosomes we looked at, it could be that this is an extremely frequent mutation, but it could also have simply occurred early in the grow-out of the culture.

As a control for transformation rates, we don’t have to worry about that, but doing the untransformed control would make us confident that changes we saw were induced by transformation and not simply due to such kinds of mutation.
(continued...)

Happy (Belated) Canada Day!



(Image: The Canadian-built robot arm attaching the space shuttle docked to the Hubble Space Telescope with the Earth in the background.)

Oh yeah, and my second attempt at preparing periplasm DNA was... inconclusive. But it was pretty interesting to try out. In particular, the TE/CsCl/phenol/acetone extraction was quite compelling visually, involving small bubbles breaking up and reforming. I need to start these experiments out at a smaller scale.

And luckily, we've received radiolabeled dATP, so I can do some more sensitive and controlled experiments next week. The use of radiolabeled uptake fragments will be significantly more sensitive and allow me to use small cultures and follow small amounts of uptake DNA.

Once I've got a functioning uptake assay, I can work out the best purification method and scale up from there.
(continued...)

Tuesday, June 30, 2009

Periplasm Prep Planning II

I tried a modification of a periplasmic protein prep to try and purify uptake DNA, which didn't work. There are several possible reasons why the experiment might not have worked, but one simple reason could be that I failed to dissociate DNA from the membranes and cells when I did the chloroform extraction.

I know! Maybe I should try an extraction that has already been used for purifying uptake DNA...

Kahn, Barany, and Smith (1983) PNAS 80:6927. Rather than describe the paper at length here, I just want to show Table 1 and Figure 4b, which relate to my extraction plans:

The first two columns describe the extraction conditions (rows 1-5). Competent cell cultures were incubated with a radiolabeled plasmid and pelleted after DNA uptake (5 or 60 min). Cell pellets were then resuspended in the indicated aqueous and organic solutions (columns A and B) in a 1:1 mixture, gently mixed, centrifuged to separate the phases, and radioactive counts in each fraction were measured.

The remaining columns indicate the relative amount of uptake in the different fractions and the identity of the radiolabeled DNA, either transformed into the chromosome (C) or still a double-stranded donor DNA molecule (D).

In the fourth condition (row 4), the aqueous phase consists of mostly donor DNA! So chromosomal contamination is in the pellet, and the desired donor molecules are in the aqueous phase. Sounds like a scheme. That’s what I’ll proceed with tomorrow.

Why not condition 1, just TE and phenol? Looks good, right? Because of Figure 4B:

The bottom line is that the phenol condition degraded the donor molecules (lane D), whereas the phenol/acetone condition did not (lane I). Here’s the gory details:

Lanes A-D describe phenol extraction of intact donor DNA molecules (Table 1, row 1):
A: input plasmid donor molecule.
B: total DNA from cells 4 min into uptake.
C: DNA extracted and dialyzed out of the pellet.
D: DNA extracted into the aqueous layer (TE).

Lanes E-J describe the phenol/acetone extraction of intact donor DNA molecules:
E: input again.
F: input cut with HindIII.
G: total DNA after 8 min uptake.
H: same as G, but digested with HindIII.
I: phenol/acetone extracted DNA after 8 mins.
J: same as I, but digested with HindIII.

The point of the HindIII digestion was as an additional test for whether molecules were donor or chromosomal. Chromosomal DNA is resistant to HindIII (this being Haemophilus influenzae after all), while donor DNA is not. Also, the HindIII digests show that the recovered DNA is double-stranded, since ssDNA won’t get cut.

Condition 4 it is, then:
Aqueous: TE/1.5 M CsCl
Organic: phenol/acetone, 1:1

That's what I'll try next.

Hmm, I’ll have to remember how to clean the DNA of CsCl after the extraction...I seem to remember doing something in particular at one point...
(continued...)

Periplasm Prep

UPDATED BELOW

Experimental plan for the day: Medium-scale periplasm prep test-- to purify double-stranded DNA in the “protected state” (the periplasm).

I have two PCR products: (1) a “good” uptake sequence USS-1, and (2) a “bad” uptake sequence, USS-R . I want to compare their uptake into rec-2 cells, which can bring DNA through the outer membrane, but not the inner membrane. This means that if I can specifically enrich USS-1, but not USS-R--and can see the difference in a gel-- then I’ve got a functioning periplasmic DNA prep. Naturally, it’ll probably take several attempts to get working...

I will use a modification of this paper and see if it nets some DNA where it should be. In outline, I’ll: Add chloroform to washed cell pellets. Soak. Extract periplasm with TE. Clean and concentrate. Run on a gel.

Based on other studies with radiolabeled USS-1 uptake, I expect that for 20 ng added to a 1 ml culture, ~50% will be taken up. To see uptake DNA without radiolabel on a gel and for reasonable controls, I will need larger cultures of competent cells than I aliquoted and froze last week.

Here’s my protocol so far:
1) Defrost two tubes of rec-2 (0.3 OD/ml aliquot) into fresh sBHI@37 (2X25 ml); wait ~2-2.5 hrs.
2) At OD600=0.3 / ml, transfer cells to M-IV by filtration.
3) Incubate 100 min @37 to induce natural competence. (negative control: frozen tube of rec-2 (0.1 OD/ml) into fresh sBHI (25 ml).)
4) Split cultures 2X and incubate with 20ng 222bp PCR fragments (USS-1, USS-R, none) / 1 ml M-IV culture (~10^9 cells) for 15-30 min @37, DNase I, EDTA to kill DNase I and other nucleases. Also add USS-1 to non-competents.
5) Spin, wash pellet 3X PBS, chloroform (20-40ul), incubate 20 min @RT (chloroform pellet DNA extraction?, save washes).
6) Extract with 100-200ul cold TE, proteinase? RNase?, p/c extraction, PCR clean-up column (or ppt?) to concentrate.
7) 1.2% agarose gel. Lanes:

Size standard
USS-1 input (2X dilution)
USS-R input (2X dilution)
rec-2 + USS-1 -> chloroform extract
rec-2 + USS-R -> chloroform extract
rec-2 + no dna -> chloroform extract
non-competent rec-2 + USS-1 -> chloroform extract

UPDATE:

Didn't work. A few little mishaps aside (mainly that the chloroform and cells really didn't mix well), I got no USS out of the prep, but did have a fair amount of chromosomal contamination. So clearly, I didn't really get the periplasm specifically, but since I didn't see any USS come through, it may also be that my competent cells weren't really.

I'll try to go through this again tomorrow, but instead of going straight for the periplasm preparation, I'll just lyse the cells, extract the DNA, run it over a mini-prep column, and run it on a gel. I'm not going to worry about the periplasm specifically, but simply that the cells are taking up DNA. When the radiolabel shows up, I can repeat the uptake assay the lab has typically done.

(continued...)

Friday, June 26, 2009

Plan for the next few weeks

One proposal down, one to go... The next one isn’t due until August 8th, so I’ve got just over a month to get it done. This time, however, I am going to manage my time better, since I need to get some preliminary data and still keep learning how to use a computer.

So here’s my plan for the next several weeks:

(1) Work on the proposal for a limited time each day (~1-2 hrs). I’ll start by developing a detailed outline of what I want to say and the order I want to say it, rather than leaping straight into writing. Based on my experience with this last one and in the past, I find I am an extremely inefficient writer (both with my time and with my words), so hopefully I can improve by having more focused daily goals.

(2) Work on the computational stuff only 1-2 hrs per day. Still need to fix the browser to display the Hin genome. Still want to work out the best way to align the genomes and report differences... particularly enumerating structural variation (non-SNPs). I also need to keep a mind towards what file formats I expect to get from sequencing. It might be particularly useful to try and simulate the kinds of results I might expect from sequencing periplasmic uptake DNA, etc.

(3) The rest of the time will be dedicated to lab work. The priority is to use defined fragments (USS-1 and USS-R) to work out a periplasmic DNA purification protocol. I’ve got cleaned amplified USS-1 and USS-R fragments, and I’m making competent cells of wild-type and rec-2 today. I tried to grow up a pilA mutant to use as a no uptake control, but something was wonky with the strain. All I need now is label. And a calibrated Geiger counter. I’ll get these things done today.

(The image above was made using the Mac-specific application GenomeMatcher. It represents a BLAST alignment between KW20 and 086-28NP across an interval containing several inversions. It also has a bunch of useful-seeming utilities that I'd like to figure out. Now if I could just get it to use my MUMmer program, like it’s supposed to...)
(continued...)

Friday, June 5, 2009

Joint molecules...

There are several classes of transformation event that could be mediated by natural competence and recombination. The first thing that I had to wrap my head around (after having come from the double-stranded break world of recombination) was to realize that uptake DNA recombining into a host chromosome is single-stranded. I’m more used to drawing recombination models involving two broken ends of a double-stranded molecule.

So to kick things off, in the above figure, I’ve drawn what I think the joint molecule intermediates look like between polymorphic donor ssDNA invaders and recipient dsDNA substrates for the four major classes of transformation products I can envision.

The red (red/pink) lines indicate an ssDNA that’s found its way from outside the cell to the host chromosome (in blue/light blue). Aligned regions are base-paired. Unaligned regions are not...

The amount of transformation will depend not only on how many of a given joint molecule are formed, but also how they are resolved and the rate and directionality of “mismatch” correction.

One important aspect of the strand invasions depicted above will be the extent of heterology. The more polymorphisms in the donor sequence, the less stable the joint molecule will be.

Each class of polymorphism deserves special mention:

(1) Single nucleotide polymophisms: There are 12 possible heteroduplexes involving single-nucleotides, if we keep track of donor and recipient. So any of the 4 bases could go to any of the remaining 3... 3 X 4 = 12. In my preliminary analysis, I found nearly 39,029 SNPs between our reference Rd strain and the clinical isolate 86-028NP falling into each one of these different classes (with a paucity of G->C and C->G changes). Even without precisely measuring the frequency of transformation for each individual SNP, we may be able to estimate the relative transformation frequencies for these different classes of SNPs.

(2) Insertions: If the donor has an insertion relative to the recipient, recombination could yield an insertion. In the case of insertions, the uptake sequence can be anywhere on the fragment. base pairing of the substrate will require the formation of a loop. The size distribution of the input DNA will also dictate the size of possible insertional recombination. How the mismatch correction machinery handles insertional joint molecules like this is unclear to me. My preliminary analysis indicated 149 insertions in 86-028NP relative to Rd (mean size = 1162 bp, but the median only a few hundred).

(3) Deletions: I made this distinct from insertions for two reasons: (1) The uptake sequence (obviously) must flank the deleted segment. An absent piece of DNA can’t be used for uptake. This imposes directionality on deletions that’s distinct from insertions. (2) It’s unclear how the mismatch machinery would act on the two central joint molecules depicted... It restoration repair more likely in one case than in the other? My preliminary analysis indicates 137 deletions in 86-028NP relative to Rd (mean size 1904 bp, again with a much lower median).

(4) Rearrangements: I’ll give this very short shrift for now. If a particular piece of ssDNA spans a rearrangement breakpoint, then strand invasion of each end into two separate chromosomal positions could mediate an inversion or several other kinds of rearrangement. My preliminary analysis with MUMmer suggests ~9 inversions and 26 transpositions between the two strains.

There are many many ways in which rearrangements in the donor might produce fragments that could mediate a rearrangment of the recipient. It’s going to take some time to enumerate these. Indeed, even if the donor was perfectly colinear with the recipient, some fragments could still mediate rearrangements.

For example, looking at a close-up of the MAUVE output between Rd and 86-028NP shows that the rRNA gene cluster of 23S and 16S (in Red) often span the rearrangement breakpoints between the two strains:
This could arise in a traditional mutational way (i.e. recombination between inverted repeats would cause an inversion), but could also be mediated by transformation. For example, if a DNA fragment that contained a bit of rDNA and a bit of flank first invaded the wrong rDNA copy, the other end could then grab the original flank... causing a rearrangement.

That is, in the rearrangements-involving-repeats scenario, it would be quite possible to produce rearrangements distinct from either the donor or recipient DNA.

This is going to take a while to flesh out, as there are several possible outcomes to this sort of thing. It doesn't make things any easier that I also now have to consider that the chromosome is circular... sigh... I'd better try and remember what plectonemic and paranemic mean...

Nevertheless, the upside is that it will be possible to measure indel and rearrangement transformation rates with considerably more ease than the SNP class, due to the use of “spanning coverage”.
(continued...)

Searching for Cash (How I'll spend my summer vacation)

So it’s time to outline my plans for the next several weeks... Grant-writing and preliminary data gathering!

Two grant applications to do:
(1) Michael Smith Foundation for Health Research (due June 25 to Office of Research)
(2) National Institutes of Health (due August 8)

I’ll find out whether or not to submit a full application to (1) in the next several days (if my “letter of intent” was sufficiently cool-sounding, I guess). The proposal itself is short (only 3 pages), so I should be able to focus on having a well-written piece in the next couple of weeks. This will help me a lot with the other proposal too. (I’d also better solicit letters for that one soon, but I want to wait until they tell me to apply first!)

And for (2) much editing, writing, and analysis to do! Rosie and I had begun tackling my rejected NIH proposal after I first got here, but enough time has passed that we’re both out-of-the-loop on our own editing. There are several things to do, besides polish the writing:

First off, the basics of the proposal are these: (a) I want to feed genomic DNA from one Haemophilus influenzae strain to a competent cell culture of another strain. (b) I’ll purify the donor DNA from various cellular compartments (representing different stages in the natural transformation pathway). (c) Then, I’ll use sequencing to measure the relative abundance of different DNA sequences along the pathway. (d) This should give me a comprehensive view of transformation potential of one genome into another.

There are a huge number of possible analyses, each of which may or may not be interesting. One major issue is how to present the kinds of analysis I’d plan on doing (and showing the reviewers that I can do it). And how to show that it really is an important set of experiments that I am capable of doing.

Proposal writing: Alongside simply improving the writing, I need to specify more details on how I will conduct the data analysis and provide any form of preliminary data that I can. These will also help us with the larger DNA uptake proposals we plan to submit in the Fall.

Background: I need to make the importance of the research plan and the specific questions I’ll answer much more clearly written and accessible. I also need to better understand the history of uptake signal sequences and how they were discovered. I.e. what’s already known versus what I’m going to learn. Rosie and I have talked this section through pretty well. It just needs to get re-written now.

Preliminary analysis: In my first version, this was just more background, but since I'll have been here a few months by the time I resubmit, might as well show what I've been doing...
  1. I need a more direct comparison of the donor and recipient genomes I’ll be using. I’d previously culled data from this paper, but showing that I can do the comparison between the genomes myself will likely go a long way with the reviewers. I’ve got some pre-preliminary analyses here in this blog, but I'm still feretting around with that data and still learning how the alignment algorithms work.
  2. I want to demonstrate that I can purify donor DNA from the periplasm and/or cytosol, so that reviewers know that I can obtain the material I want to sequence. I'm going to start with the silly way, then move onto more sophiticated periplasmic space enrichments, as necessary.
  3. It could be good to have preliminary co-transformation data to get a more realistic estimate of sequence coverage needed to to our global transformation rate experiments. The biggest help here would be to know the identity of the antibiotic resistance alleles I've been using. And maybe to do whatever it takes to get really high rates...
Experimental methods: My first version seemed okay to me, but upon re-reading, it was extremely dry and didn't give too many specifics about the downstream data analysis.
  1. Should I keep the experiment using degenerate USS oligos in this grant proposal? It’s a nice experiment, but may require more explanation than I really have space for. I’ve got a better idea now about what the experiment might look like in real life, but it may draw away from the toher components too much to talk about here.
  2. I need to more clearly and succinctly discuss the sequencing, particularly distinguishing the difference between “spanning” and “sequence” coverage.
  3. The periplasm/cytosol experiments may be real overkill as written. We may want to bring up the possibility of doing competitive experiments between multiple donor genomes, if our single donor experiment goes well. Analysis-wise, it might be nice to show a mock figure of some possible expectations. For example, I discuss but do not illustrate what an uptake blocking sequence would look like.
  4. The co-transformation experiment can be improved. Based on the way things are going with the Illumina GA2, we could very well get a lot more out of this than I’d thought. If we could fully sequence 40 independent transformants, we could say a lot...
  5. For the bulk transformation experiment, it may make sense to break it into two rounds: In one round, we’d have enough coverage to do a strong analysis of large indel and rearrangement transformation. This could also use a figure or illustration. Since we’ll have high spanning coverage, measuring these rates will use considerably less sequencing cash. If this goes well and we’ve developed a good analytical pipeline, we can then sequence a lot more and pick up the little rearrangements and all the SNPs. At this point, we’d also have a much better idea of how much sequencing we’d need to do to get to whatever level of sensitivity we want to get.
Supplementary essays: Based on the reviews of the grant, the components of my application other than the proposal appear to have been lacking. Here’s some things I need to do to improve these things in the next round:
  1. Have a better career plan. I made this essay very short, but the reviewers had some qualms that my research experience did not point at a future coherent research direction. I’ve got plenty of notions in my head about how I might conduct an independent research program, but I need to figure out how to bridge my eukaryotic and prokaryotic research directions together more robustly.
  2. Submit 1 or 2 manuscripts from graduate school. I’d said I had two manuscripts in late stages of preparation, which was true. But it’d be nice to actually have these submitted. I’m meeting with my PhD adviser and a co-author next week in California to try and hammer out the last few details of one of the manuscripts. I may still be able to submit the other before the August deadline, since it’s written and only needs some figure-fixing and final edits. But it’s been a long time and will likely take me and my PhD adviser some time to get our brains wrapped around it again.
  3. Since I tout my desire to be a teacher, I need to formulate a precise training plan to learn to teach during my postdoc. Rosie has referred me to several possible leads on campus for this purpose...
  4. I need to re-orient my research experiences to seem less haphazard. It’s true that after graduate school, I spent a year exploring my options and helping some friends set up their labs, but there was indeed a method to my madness...
  5. Letters of support: I need to make sure these letters are perhaps a bit more gushing.

The Style Book. I’m about half way through. It’s mostly filled with things that seem common sense, but reading it has helped me a bit, I think...

(continued...)

Wednesday, May 27, 2009

As easy as that? Nahh...

For one of our planned experiments, I want to purify periplasmic uptake DNA away from chromosomal DNA in a clean efficient manner. There are likely several good ways to do this, some of which may be relatively complicated. But purity will be very important for our downstream sequencing plans, so complications are okay.

Nevertheless I did a silly little experiment today to see what kind of size bias our in-house GenElute columns from Sigma have...

The columns are based on DNA adsorption to silica under high salt conditions and elution under low salt. However, larger DNAs will have a difficult time eluting off of the column, even under low salt. That’s why the manufacturer states that the columns are only good up for up to 10 kb fragments.

The DNA I’ll feed to cells will be of a discrete size distribution much smaller than chromosomal DNA, so I simply mixed genomic DNA with DNA size standards and ran them over the column. Here’s the results on a 0.6% gel:
Lane 1: 1-kb ladder alone.
Lane 2: INPUT: MAP7 + 1-kb ladder.
Lane 3: OUTPUT: the input run over a silica column with high salt.

Lane 4: lambda ladder alone.
Lane 5: INPUT: MAP7 + lambda ladder.
Lane 6: OUTPUT: the input run over a silica column with high salt.

Lane 7: MAP7 DNA alone.

It looks like the large genomic DNA fragments were pretty efficiently cleaned away from the smaller ladder fragments. The largest Lambda fragment is ~23-kb, and it seems to have been depleted quite a bit as well. Some smaller sheared genomic is clearly coming through, though, as can be seen when comparing lanes 1 to 3 and lanes 4 to 6.

I don’t think this is really enough size bias for our purposes, but I’m really quite surprised at how well it worked, so maybe I’m just being pessimistic.

I wonder how well this will work in real life. Assuming all our ladder fragments were efficiently taken up by cells into the periplasm, the association of the uptake DNA with the membranes is a major concern. If the uptake DNA is only loosely associated with the membranes, then a standard plasmid mini-prep may very well work quite nicely. Since most of the chromosomal DNA will pellet with the other cellular debris and lysed membranes, the column would then take care of most of the contaminating large molecular weight DNA.

Hmmm... I need some real uptake fragments, so I can try this with cells...
(continued...)

Thursday, May 21, 2009

Sequencing, Ho!


In response to Rosie’s last post, I wanted to outline my take on two of our planned experiments that involve deep sequencing:
  1. Transform a sequenced isolate’s genomic DNA (sheared) into our standard Rd strain, collect the transformed chromosomes, and sequence the daylights out of them to measure the transformation potential across the whole genome.
  2. Feed a sequenced isolate’s genomic DNA to competent Rd, purify the DNA that gets taken up into the periplasmic space, and sequence the daylights out of this periplasm-enriched pool to measure the uptake potential across the whole genome.
As Rosie stated, Illumina’s GA2 sequencer is pretty incredible:

The GA2 can generate massive amounts of sequence data for a low cost. The technology is pretty involved, but in short, it involves a version of “sequencing-by-synthesis” from “polonies” (or clusters) amplified from single DNA molecules that have settled down on a 2D surface. (This reminds me tremendously of my first real job after college at Lynx, Inc., which developed MPSS). The original instrument could get ~32 bases from one end of an individual DNA fragment in a polony; the new GA2 instrument can get 30-75 bases from each end of an individual DNA fragment on a polony. The number of clusters in a given run is quite large (though apparently pretty variable), so in aggregate, several gigabases of sequence can be read in a single sequencing “run”, costing about $7000.

For reasonable estimates of coverage, we can use the following conservative parameters:
• 60 million paired-end reads per flow cell (7.5 million / lane)
• 50 bases per read (that’s a very conservative 25 bases per end)
That’s ~3 gigabases! For our ~2 megabase genome, we’d get conservatively get 1500X sequence coverage for a full run, costing ~$7000.

Apparently, if we’re doing things just right, this number would more than double:
• 80 million paired-ends per flow cell (10 million / lane)
• 2 X 50 bases per paired-end read
That’s 8 Gb, or 4000X coverage of a 2 Mb genome! (Or 500X coverage in a single lane)

(1) TRANSFORMATION: Unfortunately, a single run, as described above, is probably not quite what we’d need to get decent estimates of transformation potential for every single-nucleotide polymorphism of our donor sequences into recipient chromosomes. While some markers may transform at a rate of 1/100, we probably want to have the sensitivity to make decent measurements down to 1/1000 transformation rates. For this, I think we’ll still need several GA2 runs. However, we could get very nice measurements with even a few lanes for indel and rearrangement polymorphisms, using spanning coverage (see below).

But, in a single run, we could do some good co-transformation experiments by barcoding several independent transformants, pooling, and sequencing. So if we asked for 100X coverage, we could pool 5 independent transformants into each of 8 lanes (40 individuals total). This wouldn’t give us transformation rates per polymorphic site, but since we’d know which individual each DNA fragment came from, we’d be able to piece together quite a bit of information about co-transformation and mosaicism.

(2) PERIPLASMIC UPTAKE: On the other hand, even one run is extreme overkill for measuring periplasmic uptake efficiency across the genome, because in this instance, we don’t really need sequence coverage, but only “spanning coverage” (aka physical coverage). Since we can do paired-ends, any given molecule in our periplasmic DNA purification only needs to get a sequence tag from each end of the molecule. Then we’ll know the identity of the uptake DNA fragment, based on those ends (since we have a reference genome sequence). Sequence tags that map to multiple locations in the reference will create some difficulties, but the bulk should be uniquely mappable.

So, if we started with 500bp input DNA libraries, spanning coverage in only one single lane (rather than a full run of eight) will give us a staggering 2500X spanning coverage! (500 bp spans X 10 million reads / 2 Mb genome = 2500.) To restate: For <$1000, each nucleotide along the genome of our input would be spanned 2500 times. That is, we’d paired-end tag 2500 different DNA fragments that all contained any particular nucleotide in the genome. If we graphed our input with the x-axis chromosome position and the y-axis spanning coverage, in principle we’d simply get a flat line crossing the y-axis at 2500. In reality, we’ll get a noisy line around 2500. Some sequences may end up over- or under-represented by our library construction method or sequencing bias. In our periplasmic DNA purification, however, we expect to see all kinds of exciting stuff: Regions that are taken up very efficiently would have vastly more than 2500X coverage, while regions that are taken up poorly would have substantially less coverage. This resolution is most certainly higher than we need, but heck, I’ll take it. I would almost be tempted to use a barcoding strategy and compete several genomes against each other for the periplasmic uptake experiments. If we could accept, say, a mere 250X spanning coverage for our uptake experiments, we could pool 10 different genomes together and do it that way. We could even skip the barcodes with some loss of resolution if the genomes were sufficiently divergent, such that most paired-end reads would contain identifying polymorphisms... The periplasmic uptake experiment is our best one so far, assuming we can do a clean purification of periplasmic DNA. The cytosolic translocation experiment has some bugs, but should be similiarly cheap. The transformation experiment will, no matter how much we tweak our methods, require a lot of sequencing. But it's still no mouse.
(continued...)

Thursday, May 7, 2009

Degenerate oligos II

(EDITED to add another graph)

In the first post on our planned degenerate oligo experiment, I described how we might expect the input degenerate oligo pools to look. I used the binomial probability to calculate the chance that a given oligo would have k incorrect bases in an oligo of length n.

So Pn(k) = n! / [ k! * (n-k)! ] * p^(n-k) * (1-p)^k, where p was the probability of the correct base being added at a particular position.

As a thought experiment, I now want to assume that only some of the bases in the n-mer matter to the uptake machinery. I’ll call this number m. The first difficulty was to figure out what the probability of hitting h of the m important bases, given k substitutions in a particular oligo.

This took me a while to figure out. I knew it had something to do with "sampling without replacement" and was able to work out the probabilities by hand for k = 0, 1, or 2, but then I started to get flummoxed. Eventually, I stumbled across the right equation to use: The Hypergeometric Distribution. (That’s seriously cool-sounding.)

Since I don’t know how to write a clean looking equation in HTML (see the above morass given for the binomial probability), I’ll again use the notation ( x y ) to mean "x choose y", or x! / [ y! * (x-y)! ] and will carefully add a * where I mean to indicate multiplication. The hypergeometric is then given as:

f(h; n, m, k) = ( m h ) * ( n-m k-h) / ( n k )

h = hits, or the # substitutions in the important bases
n = total length of the oligo
m = # of important bases in the oligo
k
= # of substitutions in the oligo.

So if we define only 1/4 of the bases as important to uptake, then for a 32-mer, m = 8. Then, for a 12% degenerate oligo pool, we get a histogram that looks like this:
If we say that every substitution that hits an important base causes a 10-fold decrease in uptake efficiency, but every hit in am unimportant base causes no change in uptake efficiency, we can then tell what our histogram would look like for both our input oligo pool (IN) and our periplasm-enriched oligo pool (OUT):On the left side of the graph, the output is higher than the input, while on the right side of the graph, the opposite is true. This is what we expect, since the more substitutions in the oligo, the more likely they'll hit important bases.

ADDED:

And here's a third graph addressing Rosie's comment, which shows the ratio of OUT/IN for four sets of conditions:
Here it is again on a log-plot:


(continued...)

Tuesday, May 5, 2009

Degenerate Oligos


Rosie and I have been working our way through editing my unfunded NIH postdoc proposal. There's a lot of work to do on it, but I think we'll make it excellent by the time we're done.

This is our first aim (so far):
Determine the selectivity of the outer membrane uptake machinery for USS-like sequences using degenerate USS constructs.

This is a meaty Aim and will require some time to fully flesh out. I'm not even fully convinced it belongs in this grant. It may be better for the NIH R01 submission exclusively.

The idea is to feed Haemophilus synthetic DNA containing a USS with a few random mistakes in each independent DNA molecule. We'd then purify the DNA that was preferentially taken up by the bacteria into the periplasm and sequence it to high coverage using Illumina's GA2 platform. This experiment has a lot of potential, but also a number of issues we need to work out first before we can seriously discuss the possible analyses.

To start with, I want to try and outline the basic properties of the degenerate USS donor construct we’re planning on having made. Later, I’ll try and tackle issues such as sequencing error and competition between different USS-like sequences that we expect with the actual experiment, as well as the analysis and significance of the data we hope to obtain.

As donor DNA, we want to make a long oligo (~200mer), such that the central 32 bp correspond to a known and previously studied USS, except that there will be a chance at each base in the USS for a mistake to be introduced (i.e. one of the three alternative bases could be added, instead of the correct one).

This will entail doping each bottle of dNTPs in the oligo synthesizing machine (A, T, G, or C) with a few percent of each of the other bases. For example, for a 9% degenerate oligo, the bottle that would normally be filled with 100% dATP would instead be 91% dATP, 3% dTTP, 3% dGTP, and 3% dCTP.

What do I expect the input degenerate oligo pools to look like in a sequencer?

I’ll start somewhat simply. Assuming perfect degenerate oligo synthesis and no sequencing error, I can use binomial probabilities to calculate how many oligos per million would be expected to have a certain number of incorrect bases. And I can further calculate the coverage of a given oligo species per million sequence reads. (A million sequence reads is used as a convenient frame of reference for thinking about deep sequencing. Normalizing to parts per million is also something I've seen others do.)

Variables:

n = length of the oligo to be considered (n=32 hereafter)
p = probability that the correct base is added at a particular position
q = 1-p = probability that an incorrect base is added at a particular position
k = the number of incorrect bases in a particular oligo

Then,

( n k ) = total number of possible oligos with k incorrect bases
= “n choose k”
= n! / [ k! * (n-k)! ]

nPk = binomial probability of an oligo having k incorrect bases
= ( n k ) * p^(n-k) * q^k

ppm = expected number of sequence reads per million with k incorrect bases
= nPk * 10^6

seq = number of sequences with k mistakes
= ( n k ) * 3^k

coverage = expected reads per million of a given oligo species with k incorrect bases
= ppm / seq

Example,
n = 32
p = 0.88
q = 0.12



This means that if we produce a degenerate 32mer oligo, in which each position has a 12% chance of incorporating an incorrect base (any one of the alternate 3 bases with equal probability), we would expect 21% of the molecules to have three mistakes.

But there are 4,960 different ways in which a 32mer could have 3 mistakes ( n k ). Furthermore, there are 3^3 = 81 different ways that there could be a 3 base mistake in an oligo, yielding 133,920 distinct oligo species that have 3 mistakes (seq). So out of a million sequence reads, we would only expect to see a given species of oligo with three mistakes 1.57 times per million sequence reads (i.e. 1.57X coverage of each possible 3 mistake oligo).

Here are some graphs to illustrate what this looks like for different levels of degeneracy. These should probably be histograms, but for now, I’m keeping them as they are.

Graph #1: The expected number of sequence reads (per million) with k incorrect bases in the 32mer oligo are shown for different levels of degeneracy.


Graph #2: The expected fold-coverage (per million reads) of individual oligo species with k incorrect bases are shown for different levels of degeneracy.


The take-away message: With moderate levels of degeneracy, we can obtain large amounts of sequence data for oligos with a few incorrect bases, but we will have a substantially more difficult time in getting high coverage of individual oligo species for any more than 2 or 3 incorrect bases per oligo. This means we’ll have to carefully consider how we’ll do the analysis.

Much more on this in the future...
(continued...)