Files
DissLiteratur/storage/FI4CETHW/.zotero-ft-cache
T
2025-10-26 17:12:09 +01:00

244 lines
6.6 KiB
Plaintext
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
J. Mol. Bid. (1981) 147, 195-197
Identification of Common Molecular Subsequences
The identification of maximally homologous subsequences among sets of long
sequences is an important problem in molecular sequence analysis. The problem is
straightforward
only if one restricts consideration to contiguous subsequences
(segments) containing no internal deletions or insertions. The more general problem
has its solution in an extension of sequence metrics (Sellers 1974; Waterman et al.,
1976) developed to measure the minimum number of “events” required to convert
one sequence into another.
These developments in the modern sequence analysis began with the heuristic
homology algorithm of Needleman & Wunsch (1970) which first introduced an
iterative matrix method of calculation. Numerous other heuristic algorithms have
been suggested including those of Fitch (1966) and Dayhoff (1969). More mathemat-
ically rigorous algorithms were suggested by Sankoff (1972), Reichert et al. (1973)
and Beyer et al. (1979) but these were generally not biologically satisfying or
interpretable. Success came with Sellers (1974) development of a true metric measure
of the distance between sequences. This metric was later generalized by Waterman
et al. (1976) to include deletions/insertions
of arbitrary length. This metric
represents the minimum number of “mutational events” required to convert one
sequence into another. It is of interest to note that Smith et al. (1980) have recently
shown that under some conditions the generalized Sellers metric is equivalent to the
original homology algorithm of Needleman & Wunsch (1970).
In this letter we extend the above ideas to find a pair of segments, one from each of
two long sequences, such that there is no other pair of segments with greater
similarity (homology). The similarity measure used here allows for arbitrary length
deletions and insertions.
Algorithm
The two molecular sequences will be h=alaz . . . an and IZj= blb,
b,. A
similarity a(a,b) is given between sequence elements a and b. Deletions of length k
are given weight Wt. To find pairs of segments with high degrees of similarity, we set up a matrix H. First set
Hto = Ho, = 0 for 0 I k I n and 0 I 1 I m.
Preliminary values of H have the interpretation of two segments ending in ai and bj, respectively. relationship
that H, is the maximum similarity These values are obtained from the
Hij=max{Hi-,,j-1+S(ai,bj),
~F,X {Hi-k,j- W,}, ~2" {Hi,j-,- W,}, 0}, (1)
1 li<n and 1 <j<m.
0922-2836/80/09019&03
$02.00/O
195 0 1980 Academic Press Inc. (London) Ltd.
196
T. P. SMITH
AND M. S. LVATER>lAS
The formula for H, follows by considering the possibilities
segments at any ai and b,. (1) If ai and bj are associated, the similarity is
for ending the
Hi-l,j-l +s(ai,bj).
(2) If ai is at the end of a deletion of length k, the similarity is
Hi-k,j-Wk
(3) If bj is at the end of a deletion of length I, the similarity is
Hi-k,j- cc',.
(4) Finally, a zero is included to prevent calculated negative similarity, indicating no similarity up to ai and bj.t
The pair of segments with maximum similarity is found by first locating the
maximum element of H. The other matrix elements leading to this maximum value
are than sequentially determined with a traceback procedure ending with an
element of H equal to zero. This procedure identifies the segments as well as
produces the corresponding alignment. The pair of segments with the next best
similarity is found by applying the traceback procedure to the second largest,
element of H not associated with the first traceback.
A simple example is given in Figure 1. In this example the parameters s(aibj) and
W, required were chosen on an a priori statistical basis. A match, ai = bj, produced
an s(aibj) value of unity while a mismatch produced a minus one-third. These values have an average for long, random sequences over an equally probable four letter set
of zero. The deletion weight must be chosen to be at least equal to the difference
between a match and a mismatch. The value used here was Wk= 1=0-t-1/3*k.
A
0.0 0.0 0.0 0.0 04 0.0 0.0 0.0 0.0 04 0.0 04 04 04
A
0.0 0.0 1,o 0.0 04 04 04 0.0 04 0.0 0.0 0.0 1,o 04
A
0.0 0.0 1.0 0.7 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 1.0 0.7
L7
0.0 0.0 0.0 0.7 0.3 0.0 I.0 04 04 04 I.0 1.o 0.0 0.i
G
0.0 0.0 04
1.0 0.3 0.0 04
0.i
1.0 0.0 0.0 0.7 0.7 I .o
c
0.0 I .o 0.0 o-o 2.0 1.3 0.3 1.0 0.3 2.0 0.7 0.3 0.3 0.3
(
0.0 1.0 0.1 0.0 1 .o 3.0 1.7 1.3 I.0 1.3 1.i 0.3 0.0 04
A
0.0 0.0 2.0 0.1 0.3 1.7 2.7 I.3
1.0 0.i
1 4
1.3 I .3 0.0
u
0.0 0.0 0.7 1.7 0.3 E-
2.1 2.3 1.0 0.5 1.7 2.0 I 4
1.0
u
0.0 0.0 0.3 0.3 1.3 I .o 1.3 2.3 2.0 0.7 1.7 2.7 1.7 I .o
G
0.0 0.0 0.0 l-3 0.0 1 .o 1.0 G
3.3 2.0 1.7 1.3 1.3 2.i
A
0.0 0.0 1 .o 0.0 I.0 0.3 0.7 0.7 56 3.0 I .i I.3 2.3 2.1)
c
~ 0.0 1.0 0.0 0.7 1,O I.0
0.7 1 ,7 1.7 3.0 1.;
1.3 I.0 2.0
G
0.0 0.0 0.7 1.0 0.3 0.7 1.7 0.3 2.i
I.7 23
2.3 1.0 2.0
G
0.0 0.0 0.0 1.7 0.7 0.3 0.3 1.3 1 ,3 1.3 1.3 2.3 24 2.0
FK:, 1. Hij matrix generated from the application ofeqn (1) to the sequences A-4-U-G-(-(!-$-~-~~(~-.~~
C-G-G and C-A-G-C-C-U-C-G-C-U-U-A-G.
The underlined elements indicate the trackback path fkom the
maximal element 3.30.
t Zero need not be included unless there are negative values ofs(a.b)
LETTERS TO THE EDITOR
197
Note. in this simple example, that the alignment obtained:
-G-C-C-A-U-U-G-G-C-C-UU-C.G-
contains both a mismatch and an internal deletion. It is the identification of the latter which has not been previously possible in any rigorous manner.
This algorithm not only puts the search for pairs of maximally similar segments on a mathematically rigorous basis but it can be efficiently and simply programmed on a computer.
Northern Michigan University
T. F. SMITH
Los Alamos Scientific Laboratory P.0. Box 1663, Los Alamos N. Mex. 87545. U.S.A.
Received 14 July 1980
M. S. WATERMAN
REFERENCES
Beyer, W. A., Smith, T. F., Stein. M. L. & Ulam, S. M. (1979). Math. Biosci. 19, 9-25. Dayhoff. M. 0. (1969). Atlas of Protein Sequence and Structure, National Biomedical Research
Foundation, Silver Springs, Maryland. Fitch, W. M. (1966). J. Mol. Biol. 16, 9-13. Needleman, S. B. & Wunsch, C. D. (1970). J. Mol. Biol. 48, 443-453. Reich&. T. A., Cohen, D. N. & Wong, A. K. C. (1973). J. Theoret. Biol. 42, 245-261. Sankoff, D. (1972). Proc. Nat. Acud. Sci., U.S.A. 61, 44. Sellers. P. H. (1974). J. Appl. Math. (Siam), 26, 787-793. Smith, T. F., Waterman, M. S. & Fitch, W. M. (1981). J. Mol. Evol. In the press. Waterman. M. S., Smith, T. F. & Beyer, W. A. (1976). Advan. Math. 20, 367-387.
,Votp added in proof: A weighting similar to that given above was independently developed by Walter Goad of Los Alamos Scientific Laboratory.