Pattern Discovery in Time Series, Part I:
Theory, Algorithm, Analysis, and Convergence


Cosma Rohilla Shalizi
Santa Fe Institute
1399 Hyde Park Rd.
Santa Fe, NM 87501, USA
and
Center for the Study of Complex Systems
University of Michigan
Ann Arbor, MI 48109, USA

Kristina Lisa Shalizi
Santa Fe Institute
1399 Hyde Park Rd.
Santa Fe, NM 87501, USA
and
Physics Department
University of San Francisco
2130 Fulton Street
San Francisco, CA 94117, USA

and

James P. Crutchfield
Santa Fe Institute
1399 Hyde Park Rd.
Santa Fe, NM 87501, USA

Abstract

We present a new algorithm for discovering patterns in time series and other sequential data. We exhibit a reliable procedure for building the minimal set of hidden, Markovian states that is statistically capable of producing the behavior exhibited in the data --- the underlying process's causal states. Unlike conventional methods for fitting hidden Markov models (HMMs) to data, our algorithm makes no assumptions about the process's causal architecture (the number of hidden states and their transition structure), but rather infers it from the data. It starts with assumptions of minimal structure and introduces complexity only when the data demand it. Moreover, the causal states it infers have important predictive optimality properties that conventional HMM states lack. Here, in Part I, we introduce the algorithm, review the theory behind it, prove its asymptotic reliability, and use large deviation theory to estimate its rate of convergence. In the sequel, Part II, we outline the algorithm's implementation, illustrate its ability to discover even ``difficult'' patterns, and compare it to various alternative schemes.

Citation

Cosma Rohilla Shalizi, Kristina Lisa Shalizi, and James P. Crutchfield Pattern Discovery in Time Series, Part I:
Theory, Algorithm, Analysis, and Convergence
, Journal of Machine Learning Research (2002) submitted.
Santa Fe Insitute Working Paper 02-10-060.
arXiv.org/abs/cs.LG/0210025.

To transfer a compressed PostScript version of the paper
click on its title or use one of the links below.

Compressed: size = 406 kb.
Uncompressed: size = 977 kb.
PDF: size = 257 kb.
File stored as PostScript, gzip compressed PostScript, and PDF.
Above, kb = kilobytes.
For FTP access to these files use ftp.santafe.edu:/pub/CompMech/papers.
Last modified: 28 October 2002, JPC