计算机科学
搜索引擎索引
散列函数
数字音频
背景(考古学)
情报检索
可靠性(半导体)
音频挖掘
鉴定(生物学)
数据挖掘
音频信号
语音识别
语音编码
语音处理
计算机安全
语音活动检测
植物
量子力学
生物
物理
古生物学
功率(物理)
作者
Hendrik Schreiber,Meinard Müller
标识
DOI:10.1109/tmm.2014.2318517
摘要
In view of rapidly growing digital music collections and ubiquitous music consumption, the development of technologies for identifying, browsing, and managing audio content has become a major strand of research. In this context, audio identification (ID) systems for identifying audio recordings by means of short query audio clips have become of commercial relevance. In this paper, we take a closer look at a widely used audio ID system originally developed by Haitsma and Kalker and propose several modifications that yield significant improvements with regard to retrieval speed and storage requirements. As the main contribution, we introduce a measure that establishes a connection between the temporal correlation of hash values (used for indexing) and their ability to survive in the presence of noise and signal distortions. Based on this measure, we improve the overall performance of the audio ID system by means of four strategies. First, we change the way fingerprints (audio features) are generated to increase their reliability. Second, by prioritizing more reliable hash values when searching for reference entries, we achieve substantial gains in retrieval speed by a factor of almost seven. Third, by enlarging the query fingerprint, we increase our chances of identifying reliable hash values. Fourth, by indexing only the most reliable hashes, thus applying a sub-sampling strategy, we significantly lower the server side storage requirements by a factor of ten.
科研通智能强力驱动
Strongly Powered by AbleSci AI