清晨好,您是今天最早来到科研通的研友!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您科研之路漫漫前行!

IDT: Dual-Task Adversarial Attacks for Privacy Protection

对抗制 计算机安全 隐私保护 对偶(语法数字) 计算机科学 任务(项目管理) 互联网隐私 工程类 人工智能 艺术 文学类 系统工程
作者
Pedro Faustini,Shakila Mahjabin Tonni,Annabelle McIver,Qiongkai Xu,Mark Dras
出处
期刊:Cornell University - arXiv [Cornell University]
标识
DOI:10.48550/arxiv.2406.19642
摘要

Natural language processing (NLP) models may leak private information in different ways, including membership inference, reconstruction or attribute inference attacks. Sensitive information may not be explicit in the text, but hidden in underlying writing characteristics. Methods to protect privacy can involve using representations inside models that are demonstrated not to detect sensitive attributes or -- for instance, in cases where users might not trust a model, the sort of scenario of interest here -- changing the raw text before models can have access to it. The goal is to rewrite text to prevent someone from inferring a sensitive attribute (e.g. the gender of the author, or their location by the writing style) whilst keeping the text useful for its original intention (e.g. the sentiment of a product review). The few works tackling this have focused on generative techniques. However, these often create extensively different texts from the original ones or face problems such as mode collapse. This paper explores a novel adaptation of adversarial attack techniques to manipulate a text to deceive a classifier w.r.t one task (privacy) whilst keeping the predictions of another classifier trained for another task (utility) unchanged. We propose IDT, a method that analyses predictions made by auxiliary and interpretable models to identify which tokens are important to change for the privacy task, and which ones should be kept for the utility task. We evaluate different datasets for NLP suitable for different tasks. Automatic and human evaluations show that IDT retains the utility of text, while also outperforming existing methods when deceiving a classifier w.r.t privacy task.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
科研通AI6.4应助悦耳小夏采纳,获得10
4秒前
葛甲率发布了新的文献求助10
19秒前
科研通AI6.4应助悦耳小夏采纳,获得10
22秒前
合适的飞绿完成签到,获得积分10
25秒前
科研通AI6.4应助悦耳小夏采纳,获得10
37秒前
科研通AI6.4应助悦耳小夏采纳,获得10
52秒前
58秒前
科研通AI6.2应助葛甲率采纳,获得10
58秒前
科研通AI6.4应助悦耳小夏采纳,获得10
1分钟前
shang完成签到,获得积分10
1分钟前
随心所欲完成签到 ,获得积分10
1分钟前
Doctor.TANG完成签到 ,获得积分10
1分钟前
乐乐完成签到 ,获得积分10
1分钟前
正直的晋鹏完成签到,获得积分10
1分钟前
科研通AI6.4应助悦耳小夏采纳,获得10
1分钟前
科研通AI6.4应助悦耳小夏采纳,获得10
1分钟前
科研通AI6.2应助悦耳小夏采纳,获得10
2分钟前
高贵飞丹完成签到,获得积分10
2分钟前
科研通AI6.4应助Pami采纳,获得10
2分钟前
XiaoLiu完成签到,获得积分0
2分钟前
loii应助文件撤销了驳回
2分钟前
深情安青应助悦耳小夏采纳,获得10
2分钟前
小李老博完成签到,获得积分10
2分钟前
科研通AI6.4应助悦耳小夏采纳,获得10
2分钟前
烟花应助悦耳小夏采纳,获得100
3分钟前
3分钟前
乔翼娇完成签到 ,获得积分10
3分钟前
淡然雅彤完成签到,获得积分10
3分钟前
3分钟前
maolao发布了新的文献求助10
3分钟前
科研通AI6.4应助悦耳小夏采纳,获得100
3分钟前
4分钟前
科研通AI6.4应助悦耳小夏采纳,获得10
4分钟前
Pami发布了新的文献求助10
4分钟前
JamesPei应助悦耳小夏采纳,获得10
4分钟前
幸福海之完成签到,获得积分10
4分钟前
踏实一德完成签到,获得积分10
4分钟前
五月完成签到,获得积分10
4分钟前
LL完成签到 ,获得积分10
4分钟前
科研通AI6.4应助悦耳小夏采纳,获得10
4分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Effects of Two Weeks of Red Light Therapy on Choroidal Thickness and Axial Length in Young Adults 700
内視鏡的に摘除しえた十二指腸乳頭部腫瘍の2例 660
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The Neuroscience of Language 400
Common Foundations of American and East Asian Modernisation: From Alexander Hamilton to Junichero Koizumi 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7676985
求助须知:如何正确求助?哪些是违规求助? 9242870
关于积分的说明 19919254
捐赠科研通 7247499
什么是DOI,文献DOI怎么找? 3286725
关于科研通互助平台的介绍 2444665
邀请新用户注册赠送积分活动 2289764