A twitter corpus for named entity recognition in Turkish

Çarık, Buse and Yeniterzi, Reyyan (2022) A twitter corpus for named entity recognition in Turkish. In: 13th International Conference on Language Resources and Evaluation Conference, LREC 2022, Marseille, France

Full text not available from this repository. (Request a copy)

Abstract

This paper introduces a new Turkish Twitter Named Entity Recognition dataset. The dataset, which consists of 5000 tweets from a year-long period, was labeled by multiple annotators with a high agreement score. The dataset is also diverse in terms of the named entity types as it contains not only person, organization, and location but also time, money, product, and tv-show categories. Our initial experiments with pretrained language models (like BertTurk) over this dataset returned F1 scores of around 80%. We share this dataset publicly.
Item Type: Papers in Conference Proceedings
Uncontrolled Keywords: Named Entity Recognition; Turkish; Twitter
Divisions: Faculty of Engineering and Natural Sciences
Depositing User: Reyyan Yeniterzi
Date Deposited: 09 Apr 2023 22:06
Last Modified: 09 Apr 2023 22:06
URI: https://research.sabanciuniv.edu/id/eprint/45293

Actions (login required)

View Item
View Item