| Title |
A dataset for detecting emotional manipulation techniques in Lithuanian text |
| Authors |
Butkienė, Rita ; Šukys, Algirdas ; Dambrauskas, Edgaras ; Žitkus, Voldemaras ; Ablonskis, Linas ; Vaičiukynas, Evaldas ; Danėnas, Paulius ; Butleris, Rimantas |
| DOI |
10.1016/j.dib.2026.113000 |
| Full Text |
|
| Is Part of |
Data in brief.. Amsterdam : Elsevier. 2026, vol. 67, art. no. 113000, p. 1-18.. ISSN 2352-3409 |
| Keywords [eng] |
Dataset annotation ; Emotional manipulation ; GPT ; Manipulative span identification ; Manipulative technique detection ; Prompt engineering |
| Abstract [eng] |
This paper presents a dataset of Lithuanian language comments annotated with emotional manipulation techniques. The source material comprises comments from Lithuanian news portal texts. A total of 1000 comments were selected and manually annotated by four human annotators using Label Studio. The annotators identified text fragments corresponding to fourteen emotional manipulation techniques, producing span-level human annotations that form the primary component of the dataset. In addition to the manual annotations, the dataset includes machine-generated outputs created by the GPT-4.1 model. These outputs were produced by executing prompts designed for the detection and classification of emotional manipulation span. Several prompting strategies were applied, and each prompt was run five times to capture variability in model behaviour. Because GPT-4.1 does not always return verbatim text excerpts, all generated spans were post-processed to extract precise fragments from the original comments and compute their character-level offsets. The resulting dataset encompasses four components: unprocessed source comments, human-annotated spans, prompt templates, and GPT-4.1 generated predictions. The dataset provides the first publicly available Lithuanian resource annotated for emotional manipulation techniques and can support research on manipulation and persuasion phenomena in a morphologically rich, low-resource language. The human annotations may be used as a benchmark for evaluating computational models for span extraction and technique classification. The inclusion of prompts and corresponding GPT-4.1 outputs enables the study of large language model behaviour under different prompting strategies and facilitates prompt engineering research without repeating inference runs. The dataset may also be reused for training or evaluating multilingual and cross-lingual models and offers a foundational framework for the development of annotation schemes and guidelines in related corpus construction efforts. |
| Published |
Amsterdam : Elsevier |
| Type |
Journal article |
| Language |
English |
| Publication date |
2026 |
| CC license |
|