Metin, Niyazi Ahmet and Yılmaz, Sevde and Erdoğdu, Osman Enes and Meydan, Elif Sude and Sümer, Oğul and Keküllüoğlu, Dilara (2026) SarcasTürk: Turkish context-aware sarcasm detection dataset. In: 2nd Workshop on Natural Language Processing for Turkic Languages (SIGTURK 2026), Rabat, Morocco
Full text not available from this repository. (Request a copy)
Official URL: https://dx.doi.org/10.18653/v1/2026.sigturk-1.6
Abstract
Sarcasm is a colloquial form of language that is used to convey messages in a non-literal way, which affects the performance of many NLP tasks. Sarcasm detection is not trivial and existing work mainly focus on only English. We present SarcasTürk, a context-aware Turkish sarcasm detection dataset built from Ekşi Sözlük entries, a large-scale Turkish online discussion platform where people frequently use sarcasm. SarcasTürk contains 1,515 entries from 98 titles with binary sarcasm labels and a title-level context field created to support comparisons between entry-only and context-aware models. We generate these contexts by selecting representative sentences from all entries under a title using summarization techniques. We report baseline results for a fine-tuned BERTurk classifier and zero-shot LLMs under both no-context and context-aware conditions. We find that BERTurk model with title-level context has the best performance with 0.76 accuracy and balanced class-wise F1 scores (0.77 for sarcasm, 0.75 for no sarcasm). SarcasTürk can be shared upon contacting the authors since the dataset contains potentially sensitive and offensive language.
| Item Type: | Papers in Conference Proceedings |
|---|---|
| Divisions: | Faculty of Engineering and Natural Sciences |
| Depositing User: | Niyazi Ahmet Metin |
| Date Deposited: | 07 Aug 2026 12:21 |
| Last Modified: | 07 Aug 2026 12:21 |
| URI: | https://research.sabanciuniv.edu/id/eprint/54211 |

