Nature © Macmillan Publishers Ltd 1997
articles
NATURE
|
VOL 390
|
11 DECEMBER 1997 581
coding sequence. Biological roles were assigned to 59% of the 853
ORFs using the classification scheme adapted from Riley29 (Fig. 1),
12% of ORFs matched hypothetical coding sequences of unknown
function from other organisms, and 29% were new genes. The
average relative molecular mass (M
r
) of the chromosome-encoded
proteins in B. burgdorferi is 37,529 ranging from 3,369 to 254,242,
values similar to those observed in other bacteria including
Haemophilus influenzae20 and Mycoplasma genitalium21. The
median isoelectric point (pI) for all predicted proteins is 9.7.
Analysis of codon usage in B. burgdorferi reveals that all 61 triplet
codons are used. When both AU- and GC-containing codons
specify a single amino acid, there is a marked bias (from 2-fold to
more than 20-fold, depending on the amino acid) in the use of AU-
rich codons. The most frequently used codons are AAA (Lys, 8.1%),
AAU (Asn, 5.9%), AUU (Ile, 5.9%), UUU (Phe, 5.7%), GAA (Glu,
5.0%), GAU (Asp, 4.2%) and UUA (Leu, 4.2%). The most common
amino acids are Ile (10.6%), Leu (10.3%), Lys (10.2%), Ser (7.8%)
and Asn (7.2%). The high value for Lys is in agreement with the
median calculated isoelectric point of 9.7.
Plasmid analysis
Analysis of the nucleotide sequence and Southern analyses on B.
burgdorferi DNA indicate that, in addition to the large linear
chromosome, isolate B31 contains linear plasmids of the following
approximate sizes: 56 kilobase pairs (kbp) (lp56), 54 kbp (lp54),
four plasmids of 28 kbp (lp28-1, lp28-2, lp28-3 and lp28-4), 38 kbp
(lp38), 36 kbp (lp36), 25 kbp (lp25) and 17 kbp (lp17); and circular
plasmids of the following sizes: 9 kbp (cp9), 26 kbp (cp26) and five
or six homologous plasmids of 32 kbp (cp32). These include all of
the plasmids previously identified in this strain, but comparisons
with other B31 cultures suggest that this isolate may have lost one 21
kbp linear and one or two 32 kbp circular plasmids during growth in
culture since its original isolation11–14,19,30. The sequences of all
plasmids were assembled as part of this project. However, the
assembled sequences of the cp32 and related lp56 plasmids could
not be determined with a high degree of confidence because of DNA
sequence similarity among them ($99% in several regions of
3,000–5,000 bp per plasmid)13,16 (Table 1). Improved assembly
strategies are being tested to achieve closure on these plasmids
(G. Sutton, unpublished). Plasmid lp17 is identical to that of lp16.9
from Barbour et al.15.
The 11 plasmids we have described contain a total of 430 putative
ORFs with an average size of 507 bp; plasmid G+C content ranges
from 23.1% to 32.3%. Only 71% of plasmid DNA represents
predicted coding sequences, a value significantly lower than that
on the chromosome. This indicates that average intergenic distances
are greater in the plasmids than in the chromosome, and that many
potential ORFs contain authentic frameshifts or stops (see E29, for
example), suggesting that they are decaying genes not encoding
functional proteins. Of the 430 plasmid ORFs, only 70 (16%) could
be identified and these include membrane proteins such as OspA-D,
decorin-binding proteins, the VlsE lipoprotein recombination cas-
sette, and the purine ribonucleotide biosynthetic enzymes GuaA
and GuaB. We found that 100 ORFs (23%) match other hypothe-
tical proteins from plasmids in this and related strains of B.
burgdorferi 15,16,31; 10 ORFs (2.3%) match hypothetical proteins
from species other than Borrelia; and 250 ORFs (58%) have no
database match.
We found that 47 paralogous gene families containing from 2 to
12 members account for 39% (169 ORFs) of the plasmid-encoded
genes with no known biological role (Fig. 1). Paralogue families 32
and 50, typified by previously identified B. burgdorferi plasmid
genes cp32 orfC and cp8.3 orf2, respectively, have some similarities
to proteins involved in replication, segregation and control of copy
number in other bacterial systems16,31. Previous studies have
reported examples of plasmid gene duplication, but the extent of
Table 2 Gene identification numbers are listed with the prefix BB as in Fig. 2. Each gene
identified is listed in its functional role category (adapted from Riley29). The percentage of
similarity and a two-letter abbreviation for genus and species for the best match are also
shown. An expanded version of this table with additional information is available on the
World-Wide Web at http://www.tigr.org/tdb/mdb/bbdb/bbdb.htm. Abbreviations of gene
names are: Ac, acetyl; BP, binding protein; biosyn, biosynthesis; cello, cellobiose; CPDase,
carboxypeptidase; Dcase, decarboxylase; DHase, dehydrogenase; flgr, flagellar/flagellum;
fru, fructose; GBP, glycine, betaine, L-proline; glu, glucose; Kase, kinase; mal, maltose; MC-
methyl-accepting chemotaxis; MTase, methyltransferase; NAG, N-acetylglucosamine; OH,
hydroxy; OP, oligopeptide; P, phosphate; PPTase, phosphotransferase; PPase, phospha-
tase; prt, protein; put, putative; RDase, reductase; RG, ribose/galactose; SAM, S-adenosyl-
methionine; Sase, synthetase/synthase; SP, spermidine/putrescine; ss, single-stranded;
sub, subunit; Tase, transferase.
Abbrevation of genus and species are: Ah, Aeromonas hydrophila; Ar, Agrobacterium
radiobacter; Al, Alteromonas sp.; Ab, Anabaena sp.; An, Anacystis nidulans; At, arabidopsis
thaliana; Av, Azotobacter vinelandii; Bf, Bacillus firmus; Bl, Cacillus licheniformis; Bm,
Bacillus megaterium; Bs, Bacillus stearothermophilus; Bs, Bacillus subtilis; Bb, Borrelia
burgdorferi; Bc, Borrelia coriaceae; Bh, Borrelia hermsii; Ba, Buchnera aphidicola; Ca,
Clostridium acetobutylicum; Cl, Clostridium longisporum; Cp, Clostridium perfringens; Cg,
Corynegacterium glutamicum; Cb, Coxiella burnetii; Cp, Cyanophora paradoxa; Dd,
Dictyostelium discoideum; Ec, Escherichia coli; Eh, Entamoeba histolytica; Ec,
Enterobacter cloacae; El, Enterococcus faecalis; Eh, Enterococcus hirae; Ha,
Haemophilus aegyptius; Hi, Haemophilus influenzae; Hp, Helicobacter pylori; Hs, Homo
sapiens; La, Lactobacillus acidophilus; Ll, Lactococcus lactis; Li, Leptospira interrogans
serovar lai; Mj, Methanococcus jannaschii; Mb, Methanosarcina barkeri; Ml,
Mycobacterium leprae; Mt, Mycobacterium tuberculosis; Mc, Mycoplasma capricolum;
Mg, Mycoplasma genitalium; Mh, Mycoplasma hominis; Mh, Mycoplasma hyorhinis; Mm,
Mycoplasma mycoides; Mp, Mycoplasma pneumoniae; Mx, Myxococcus xanthus; Ng,
Neisseria gonorrhoeae; Nm, Neisseria meningitidis; Os, Odontella sinensis; Pt,
Paramecium tetraurelia; Pa, Pediococcus acidilactici; Pf, Plasmodium falciparum; Pg,
Porphyromonas gingivalis; Pv, Proteus vulgaris; Pa, Pseudomonas aeruginosa; Pm,
Pseudomonas mevalonii; Pp, Pseudomonas putida; Rm, Rhizobium meliloti; Rc,
Rhodobacter capsulatus; Rs, Rhodobacter sphaeroides; Rp, Rickettsia prowazekii; Sc,
Saccharomyces cerevisiae; Sc, Salmonella choleraesius; St, Salmonella typhimurium; Sh,
Serpulina hyodysenteriae; Sd, Shigella dysenteriae; So, Spinacia oleracea; Sc,
Staphylococcus camosus; Se, Staphylococcus epidermidis; Sp, Streptococcus pyogen-
es; Sc, Streptomyces coelicolor; Ss, Sulfolobus solfataricus; Syn, Synechococcus sp.; Sp,
Synechocystis PCC6803; Tt, Thermoanaerobacterium thermosaccharolyticum; Tb, Ther-
mophilic bacterium RT8.B4.; Ttv, Thermoproteus tenax virus; Tm, Thermotoga maritima; Tat,
Thermus aquaticus thermophilus; Ta, Thermus aquaticus; Td, Treponema denticola; Tp,
Treponema pallidum; Ta, Triticum aestivum; Tb, Trypanosoma brucei mitochondrion; Vc,
Vibrio cholerae; Vp, Vibrio parahaemolyticus; Zm, Zymomonas mobilis.
Table 1 Genome features in Borrelia burgdorferi
Chromosome 910,725 bp (28.6% G+C)
Coding sequences (93%)
RNAs (0.7%)
Intergenic sequence (6.3%)
853 coding sequences
500 (59%) with identified database match
104 (12%) match hypothetical proteins
249 (29%) with no database match
………………………………………………………………………….…………………………………………………………………………….
Plasmids
cp9 9,386 bp (23.6% GC)
cp26 26,497 bp (26.3% GC)
lp17 16,828 bp (23.1% GC)
lp25 24,182 bp (23.3% GC)
lp28-1 26,926 bp (32.3% GC)
lp28-2 29,771 bp (31.5% GC)
lp28-3 28,605 bp (25.1% GC)
lp28-4 27,329 bp (24.4% GC)
lp36 36,834 bp (26.8% GC)
lp38 38,853 bp (26.1% GC)
lp54 53,590 bp (28.1% GC)
Coding sequences (71%)
Intergenic sequence (29%)
430 coding sequences
70 (16%) with identified database match
110 (26%) match hypothetical proteins
250 (58%) with no database match
………………………………………………………………………….…………………………………………………………………………….
Ribosomal RNA Chromosome coordinates
16S 444581–446118
23S 438590–441508
5S 438446–438557
23S 435334–438267
5S 435201–435312
………………………………………………………………………….…………………………………………………………………………….
Stable RNA
tmRNA 46973–47335
mpB 750816–751175
………………………………………………………………………….…………………………………………………………………………….
Transfer RNA
34 species (8 clusters,14 single genes)
………………………………………………………………………….…………………………………………………………………………….
*The telomeric sequences of the nine linear plasmids assembled as part of this study were
not determined; estimation of the number of missing terminal nucleotides by restriction
analysis suggests that less than 1,200 bp is missing in all cases. Comparisons with
previously determined sequences of lp 16.9 and one terminus of lp28-1 indicate that 25,
60 and 1,200 bp are missing, respectively.
R