244 lines
6.6 KiB
Plaintext
244 lines
6.6 KiB
Plaintext
J. Mol. Bid. (1981) 147, 195-197
|
||
|
||
Identification of Common Molecular Subsequences
|
||
|
||
The identification of maximally homologous subsequences among sets of long
|
||
|
||
sequences is an important problem in molecular sequence analysis. The problem is
|
||
|
||
straightforward
|
||
|
||
only if one restricts consideration to contiguous subsequences
|
||
|
||
(segments) containing no internal deletions or insertions. The more general problem
|
||
|
||
has its solution in an extension of sequence metrics (Sellers 1974; Waterman et al.,
|
||
|
||
1976) developed to measure the minimum number of “events” required to convert
|
||
|
||
one sequence into another.
|
||
|
||
These developments in the modern sequence analysis began with the heuristic
|
||
|
||
homology algorithm of Needleman & Wunsch (1970) which first introduced an
|
||
|
||
iterative matrix method of calculation. Numerous other heuristic algorithms have
|
||
|
||
been suggested including those of Fitch (1966) and Dayhoff (1969). More mathemat-
|
||
|
||
ically rigorous algorithms were suggested by Sankoff (1972), Reichert et al. (1973)
|
||
|
||
and Beyer et al. (1979) but these were generally not biologically satisfying or
|
||
|
||
interpretable. Success came with Sellers (1974) development of a true metric measure
|
||
|
||
of the distance between sequences. This metric was later generalized by Waterman
|
||
|
||
et al. (1976) to include deletions/insertions
|
||
|
||
of arbitrary length. This metric
|
||
|
||
represents the minimum number of “mutational events” required to convert one
|
||
|
||
sequence into another. It is of interest to note that Smith et al. (1980) have recently
|
||
|
||
shown that under some conditions the generalized Sellers metric is equivalent to the
|
||
|
||
original homology algorithm of Needleman & Wunsch (1970).
|
||
|
||
In this letter we extend the above ideas to find a pair of segments, one from each of
|
||
|
||
two long sequences, such that there is no other pair of segments with greater
|
||
|
||
similarity (homology). The similarity measure used here allows for arbitrary length
|
||
|
||
deletions and insertions.
|
||
|
||
Algorithm
|
||
|
||
The two molecular sequences will be h=alaz . . . an and IZj= blb,
|
||
|
||
b,. A
|
||
|
||
similarity a(a,b) is given between sequence elements a and b. Deletions of length k
|
||
|
||
are given weight Wt. To find pairs of segments with high degrees of similarity, we set up a matrix H. First set
|
||
|
||
Hto = Ho, = 0 for 0 I k I n and 0 I 1 I m.
|
||
|
||
Preliminary values of H have the interpretation of two segments ending in ai and bj, respectively. relationship
|
||
|
||
that H, is the maximum similarity These values are obtained from the
|
||
|
||
Hij=max{Hi-,,j-1+S(ai,bj),
|
||
|
||
~F,X {Hi-k,j- W,}, ~2" {Hi,j-,- W,}, 0}, (1)
|
||
|
||
1 li<n and 1 <j<m.
|
||
|
||
0922-2836/80/09019&03
|
||
|
||
$02.00/O
|
||
|
||
195 0 1980 Academic Press Inc. (London) Ltd.
|
||
|
||
196
|
||
|
||
T. P. SMITH
|
||
|
||
AND M. S. LVATER>lAS
|
||
|
||
The formula for H, follows by considering the possibilities
|
||
|
||
segments at any ai and b,. (1) If ai and bj are associated, the similarity is
|
||
|
||
for ending the
|
||
|
||
Hi-l,j-l +s(ai,bj).
|
||
|
||
(2) If ai is at the end of a deletion of length k, the similarity is
|
||
|
||
Hi-k,j-Wk
|
||
|
||
(3) If bj is at the end of a deletion of length I, the similarity is
|
||
|
||
Hi-k,j- cc',.
|
||
|
||
(4) Finally, a zero is included to prevent calculated negative similarity, indicating no similarity up to ai and bj.t
|
||
|
||
The pair of segments with maximum similarity is found by first locating the
|
||
maximum element of H. The other matrix elements leading to this maximum value
|
||
are than sequentially determined with a traceback procedure ending with an
|
||
element of H equal to zero. This procedure identifies the segments as well as
|
||
produces the corresponding alignment. The pair of segments with the next best
|
||
similarity is found by applying the traceback procedure to the second largest,
|
||
element of H not associated with the first traceback.
|
||
A simple example is given in Figure 1. In this example the parameters s(aibj) and
|
||
W, required were chosen on an a priori statistical basis. A match, ai = bj, produced
|
||
an s(aibj) value of unity while a mismatch produced a minus one-third. These values have an average for long, random sequences over an equally probable four letter set
|
||
of zero. The deletion weight must be chosen to be at least equal to the difference
|
||
between a match and a mismatch. The value used here was Wk= 1=0-t-1/3*k.
|
||
|
||
A
|
||
|
||
0.0 0.0 0.0 0.0 04 0.0 0.0 0.0 0.0 04 0.0 04 04 04
|
||
|
||
A
|
||
|
||
0.0 0.0 1,o 0.0 04 04 04 0.0 04 0.0 0.0 0.0 1,o 04
|
||
|
||
A
|
||
|
||
0.0 0.0 1.0 0.7 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 1.0 0.7
|
||
|
||
L7
|
||
|
||
0.0 0.0 0.0 0.7 0.3 0.0 I.0 04 04 04 I.0 1.o 0.0 0.i
|
||
|
||
G
|
||
|
||
0.0 0.0 04
|
||
|
||
1.0 0.3 0.0 04
|
||
|
||
0.i
|
||
|
||
1.0 0.0 0.0 0.7 0.7 I .o
|
||
|
||
c
|
||
|
||
0.0 I .o 0.0 o-o 2.0 1.3 0.3 1.0 0.3 2.0 0.7 0.3 0.3 0.3
|
||
|
||
(
|
||
|
||
0.0 1.0 0.1 0.0 1 .o 3.0 1.7 1.3 I.0 1.3 1.i 0.3 0.0 04
|
||
|
||
A
|
||
|
||
0.0 0.0 2.0 0.1 0.3 1.7 2.7 I.3
|
||
|
||
1.0 0.i
|
||
|
||
1 4
|
||
|
||
1.3 I .3 0.0
|
||
|
||
u
|
||
|
||
0.0 0.0 0.7 1.7 0.3 E-
|
||
|
||
2.1 2.3 1.0 0.5 1.7 2.0 I 4
|
||
|
||
1.0
|
||
|
||
u
|
||
|
||
0.0 0.0 0.3 0.3 1.3 I .o 1.3 2.3 2.0 0.7 1.7 2.7 1.7 I .o
|
||
|
||
G
|
||
|
||
0.0 0.0 0.0 l-3 0.0 1 .o 1.0 G
|
||
|
||
3.3 2.0 1.7 1.3 1.3 2.i
|
||
|
||
A
|
||
|
||
0.0 0.0 1 .o 0.0 I.0 0.3 0.7 0.7 56 3.0 I .i I.3 2.3 2.1)
|
||
|
||
c
|
||
|
||
~ 0.0 1.0 0.0 0.7 1,O I.0
|
||
|
||
0.7 1 ,7 1.7 3.0 1.;
|
||
|
||
1.3 I.0 2.0
|
||
|
||
G
|
||
|
||
0.0 0.0 0.7 1.0 0.3 0.7 1.7 0.3 2.i
|
||
|
||
I.7 23
|
||
|
||
2.3 1.0 2.0
|
||
|
||
G
|
||
|
||
0.0 0.0 0.0 1.7 0.7 0.3 0.3 1.3 1 ,3 1.3 1.3 2.3 24 2.0
|
||
|
||
FK:, 1. Hij matrix generated from the application ofeqn (1) to the sequences A-4-U-G-(‘-(!-$-~‘-~~(~-.~~
|
||
|
||
C-G-G and C-A-G-C-C-U-C-G-C-U-U-A-G.
|
||
|
||
The underlined elements indicate the trackback path fkom the
|
||
|
||
maximal element 3.30.
|
||
|
||
t Zero need not be included unless there are negative values ofs(a.b)
|
||
|
||
LETTERS TO THE EDITOR
|
||
|
||
197
|
||
|
||
Note. in this simple example, that the alignment obtained:
|
||
|
||
-G-C-C-A-U-U-G-G-C-C-UU-C.G-
|
||
|
||
contains both a mismatch and an internal deletion. It is the identification of the latter which has not been previously possible in any rigorous manner.
|
||
This algorithm not only puts the search for pairs of maximally similar segments on a mathematically rigorous basis but it can be efficiently and simply programmed on a computer.
|
||
|
||
Northern Michigan University
|
||
|
||
T. F. SMITH
|
||
|
||
Los Alamos Scientific Laboratory P.0. Box 1663, Los Alamos N. Mex. 87545. U.S.A.
|
||
Received 14 July 1980
|
||
|
||
M. S. WATERMAN
|
||
|
||
REFERENCES
|
||
Beyer, W. A., Smith, T. F., Stein. M. L. & Ulam, S. M. (1979). Math. Biosci. 19, 9-25. Dayhoff. M. 0. (1969). Atlas of Protein Sequence and Structure, National Biomedical Research
|
||
Foundation, Silver Springs, Maryland. Fitch, W. M. (1966). J. Mol. Biol. 16, 9-13. Needleman, S. B. & Wunsch, C. D. (1970). J. Mol. Biol. 48, 443-453. Reich&. T. A., Cohen, D. N. & Wong, A. K. C. (1973). J. Theoret. Biol. 42, 245-261. Sankoff, D. (1972). Proc. Nat. Acud. Sci., U.S.A. 61, 44. Sellers. P. H. (1974). J. Appl. Math. (Siam), 26, 787-793. Smith, T. F., Waterman, M. S. & Fitch, W. M. (1981). J. Mol. Evol. In the press. Waterman. M. S., Smith, T. F. & Beyer, W. A. (1976). Advan. Math. 20, 367-387.
|
||
|
||
,Votp added in proof: A weighting similar to that given above was independently developed by Walter Goad of Los Alamos Scientific Laboratory.
|
||
|