Deep contrastive learning–based redundancy elimination for similar data in cloud storage systems
Ling Xiao, Qinbao Fang, Shihao Feng
Anhui Jianzhu University
内容与影响
Cloud storage systems contain a large amount of both duplicate and similar data, which imposes significant storage overhead. Redundancy elimination has become an important technique for improving storage efficiency. However, existing redundancy elimination methods still lack deep semantic understanding and expressive feature representations, especially for identifying similar blocks beyond exact duplicates. To address these limitations, this paper proposes DCLC, a redundancy elimination method based on deep contrastive learning. DCLC constructs a semantic representation model that integrates local features with global contextual information through self-supervised contrastive learning. Therefore, it generates more discriminative embedding vectors and improves the accuracy and efficiency of identifying similar blocks. Furthermore, a dynamic cache mechanism based on access locality processes most queries for similar blocks directly in the cache, which avoids unnecessary vectorization and approximate nearest-neighbor searches. Finally, DCLC arranges multiple candidate base blocks in sequence to form a cooperative reference unit, enabling the delta encoder to exploit long-range structural correlations and eliminate cross-block redundancy. Experimental results demonstrate that, compared with representative existing approaches, DCLC improves system throughput by 4.18 \(\times \) to 59.58 \(\times \) and achieves a maximum improvement of 248% in overall compression ratio under the evaluated workloads. The source code is available at: https://github.com/qcode-systems-lab/DCLC .
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
回答优先基于摘要、文献信息与可获取全文;依据不足时会明确说明。
学术脉络
学科主题
计算机 / AICloud Data Security Solutions
Advanced Data Storage Technologies · Big Data and Digital Economy