动态时间归整
亲密度
聚类分析
度量(数据仓库)
计算机科学
系列(地层学)
代表(政治)
相似性(几何)
相似性度量
模式识别(心理学)
人工智能
时间序列
距离矩阵
数据挖掘
序列(生物学)
机器学习
数学
算法
数学分析
古生物学
图像(数学)
政治
法学
生物
遗传学
政治学
标识
DOI:10.48550/arxiv.2309.03579
摘要
Measuring distance or similarity between time-series data is a fundamental aspect of many applications including classification, clustering, and ensembling/alignment. Existing measures may fail to capture similarities among local trends (shapes) and may even produce misleading results. Our goal is to develop a measure that looks for similar trends occurring around similar times and is easily interpretable for researchers in applied domains. This is particularly useful for applications where time-series have a sequence of meaningful local trends that are ordered, such as in epidemics (a surge to an increase to a peak to a decrease). We propose a novel measure, DTW+S, which creates an interpretable "closeness-preserving" matrix representation of the time-series, where each column represents local trends, and then it applies Dynamic Time Warping to compute distances between these matrices. We present a theoretical analysis that supports the choice of this representation. We demonstrate the utility of DTW+S in several tasks. For the clustering of epidemic curves, we show that DTW+S is the only measure able to produce good clustering compared to the baselines. For ensemble building, we propose a combination of DTW+S and barycenter averaging that results in the best preservation of characteristics of the underlying trajectories. We also demonstrate that our approach results in better classification compared to Dynamic Time Warping for a class of datasets, particularly when local trends rather than scale play a decisive role.
科研通智能强力驱动
Strongly Powered by AbleSci AI