TUNE: a task for Turkish machine unlearning for data privacy

Benli, Doruk and Canoğlu, Ada and Gönençer, Nehir İlkim and Keküllüoğlu, Dilara (2026) TUNE: a task for Turkish machine unlearning for data privacy. In: 2nd Workshop on Natural Language Processing for Turkic Languages (SIGTURK 2026), Rabat, Morocco

Full text not available from this repository. (Request a copy)

Abstract

Most large language models (LLMs) are trained on massive datasets that include private information, which may be disclosed to third-party users in output generation. Developers put defences to prevent the generation of harmful and private information, but jailbreaking methods can be used to bypass them. Machine unlearning aims to remove information that may be private or harmful from the model's generation without retraining the model from scratch. While machine unlearning has gained some popularity to counter the removal of private information, especially in English, little to no attention has been given to Turkish unlearning paradigms or existing benchmarks. In this study, we introduce TUNE (Turkish Unlearning Evaluation), the first benchmark dataset for Turkish unlearning task for personal information. TUNE consists of 9842 input-target text pairs about 50 fictitious personalities with two training task types: (1) Q&A and (2) Information Request. We fine-tuned the mT5 base model to evaluate various unlearning methods, including our proposed approach. We find that while current methods can help unlearn unwanted private information in Turkish, they also unlearn other information we want to retain in the model.
Item Type: Papers in Conference Proceedings
Divisions: Faculty of Engineering and Natural Sciences
Depositing User: Doruk Benli
Date Deposited: 07 Aug 2026 12:16
Last Modified: 07 Aug 2026 12:16
URI: https://research.sabanciuniv.edu/id/eprint/54212

Actions (login required)

View Item
View Item