Pattern Discovery in Time Series, Part II:
Implementation, Evaluation, and Comparison


Kristina Lisa Shalizi
Santa Fe Institute
1399 Hyde Park Rd.
Santa Fe, NM 87501, USA
and
Physics Department
University of San Francisco
2130 Fulton Street
San Francisco, CA 94117, USA

and

Cosma Rohilla Shalizi
Santa Fe Institute
1399 Hyde Park Rd.
Santa Fe, NM 87501, USA
and
Center for the Study of Complex Systems
University of Michigan
Ann Arbor, MI 48109, USA

James P. Crutchfield
Santa Fe Institute
1399 Hyde Park Rd.
Santa Fe, NM 87501, USA

Abstract

We present a new algorithm for discovering patterns in time series or other sequential data. In the prior companion work, Part I, we reviewed the underlying theory, detailed the algorithm, and established its asymptotic reliability and various estimates of its data-size asymptotic rate of convergence. Here, in Part II, we outline the algorithm's implementation, illustrate its behavior and ability to discover even ``difficult'' patterns, demonstrate its superiority over alternative algorithms, and discuss its possible applications in the natural sciences and to data mining.

Citation

Kristina Lisa Shalizi, Cosma Rohilla Shalizi, and James P. Crutchfield Pattern Discovery in Time Series, Part II:
Implementation, Evaluation, and Comparison
, Journal of Machine Learning Research (2002) to be submitted.
Santa Fe Insitute Working Paper 02-10-XXX.
arXiv.org/abs/cs.LG/02XXXXX.

To transfer a compressed PostScript version of the paper
click on its title or use one of the links below.

Compressed: size = 313 kb.
Uncompressed: size = 1,470 kb.
PDF: size = 360 kb.
File stored as PostScript, gzip compressed PostScript, and PDF.
Above, kb = kilobytes.
For FTP access to these files use ftp.santafe.edu:/pub/CompMech/papers.
Last modified: 28 October 2002, JPC