散列函数
计算机科学
特征哈希
通用哈希
动态完美哈希
双重哈希
汉明空间
哈希表
可扩展性
局部敏感散列
线性哈希
判别式
理论计算机科学
人工智能
数据挖掘
汉明码
算法
区块代码
数据库
解码方法
计算机安全
作者
Jingkuan Song,Lianli Gao,Yan Yan,Dongxiang Zhang,Nicu Sebe
标识
DOI:10.1145/2733373.2806341
摘要
There is an increasing interest in using hash codes for efficient multimedia retrieval and data storage. The hash functions are learned in such a way that the hash codes can preserve essential properties of the original space or the label information. Then the Hamming distance of the hash codes can approximate the data similarity. Existing works have demonstrated the success of many supervised hashing models. However, labeling data is time and labor consuming, especially for scalable datasets. In order to utilize the supervised hashing models to improve the discriminative power of hash codes, we propose a Supervised Hashing with Pseudo Labels (SHPL) which uses the cluster centers of the training data to generate pseudo labels, based on which the hash codes can be generated using the criteria of supervised hashing. More specifically, we utilize linear discriminant analysis (LDA) with trace ratio criterion as a showcase for hash functions learning and during the optimization, we prove that the pseudo labels and the hash codes can be jointly learned and iteratively updated in an unified framework. The learned hash functions can harness the discriminant power of trace ratio criterion, and thus can achieve better performance. Experimental results on three large-scale unlabeled datasets (i.e., SIFT1M, GIST1M, and SIFT1B) demonstrate the superior performance of our SHPL over existing hashing methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI