Research groups use INCEpTION to build annotated corpora across a wide range of fields and languages. Below are 191 papers from 2025 and 2026 that used it for their own annotation work.
- A Corpus of Persuasion Techniques in Slavic Languages
LREC 2026 Social science & politics Bulgarian/Polish parliamentary debates and Russian social media Stance & argumentationText & span classification
- A Dataset for Named Entity Recognition and Relation Extraction from Art-historical Image Descriptions
arXiv (cs.CL) Digital humanities & literature German and English art-historical image descriptions from 13 museum catalogues, auction listings and scholarly databases (FRAME dataset) Named entity recognitionRelation extractionCoreference & anaphoraEntity linking & normalisation
- A Learner-Oriented Annotated Resource of French Multiword Expressions for Text Adaptation in Foreign Language Reading
LREC 2026 Education & learner corpora French pedagogical corpora (multiword expressions) Terminology & lexicographyWord sense & lexical semantics
- A Novel Dataset and Three Ways to Approach Automatic Metaphor Detection in German Religious Online Forums
LREC 2026 NLP & computational linguistics German religious online forum posts Word sense & lexical semanticsText & span classification
- A Study of Coreference Resolution for Russian-Language Texts from a Limited Subject Domain
Dialogue 2026 (Computational Linguistics and Intellectual Technologies) NLP & computational linguistics Russian popular-science texts about computational linguistics: the HaRuCo corpus of 21 texts containing 9,905 mention annotations, with domain-entity class labels, singletons and abstract mentions. Coreference & anaphoraNamed entity recognitionEntity linking & normalisation
- A Study of a Supporting Application in Patient-Centered Breast Cancer Care
Biomedical & clinical Japanese hospital medical records (progress notes, nursing records, discharge summaries) for three breast cancer cases, tagged for patient-centered information Text & span classification
- A Synthetic Conversational Dataset for Type 2 Diabetes Management
LREC 2026 Biomedical & clinical English synthetic patient-caretaker diabetes dialogues Relation extraction
- A knowledge graph-based post-stroke gait assessment system: A pilot study
Medical Engineering & Physics Biomedical & clinical English clinical and biomechanical literature on post-stroke gait deviations, their causes and compensatory mechanisms, annotated to build a gait-analysis knowledge graph Event extractionRelation extractionNamed entity recognition
- A novel credibility dataset and annotation framework for COVID-19-related German tweets
Language Resources and Evaluation NLP & computational linguistics 643 German-language COVID-19-related tweets from Twitter/X, annotated on four layers (entities/named entities, informativeness, topic, credibility) Text & span classificationNamed entity recognition
- A software pipeline for medical information extraction with large language models, open source and suitable for oncology
npj Precision Oncology Biomedical & clinical Unstructured clinical free text: 100 TCGA colorectal-cancer pathology reports for TNM-stage extraction and eight fictitious clinical letters on pulmonary embolism used for anonymization/IE evaluation Named entity recognitionEntity linking & normalisation
- ABCD-LINK: Annotation Bootstrapping for Cross-Document Fine-Grained Links
EACL 2026 NLP & computational linguistics English peer reviews and news articles Relation extractionDiscourse & rhetorical structureText & span classification
- AI-Based Feedback in Counselling Competence Training of Prospective Teachers
arXiv (cs.HC) Education & learner corpora German simulated teacher-parent counselling conversations Discourse & rhetorical structureText & span classificationSpeech, multimodal & transcription
- Aborder l'énonciation sur corpus oral : approches et outils
Linguistics & typology The PluriDiRe corpus — 36 dyadic spoken French conversations drawn from C-Oral-Rom, annotated for direct reported speech (150 occurrences) on two layers, cues and segments Discourse & rhetorical structureSpeech, multimodal & transcription
- Adapting Clinical Event Annotation to Dutch Primary Care: An Event Annotation Framework for Post-Acute Infection Syndromes
medRxiv Biomedical & clinical Dutch general-practitioner SOEP consultation notes from a Dutch primary-care network (UMC Utrecht), covering post-acute infection syndromes; a 200-note pilot, one note per patient for 200 patients, doubly annotated. Event extractionNamed entity recognitionRelation extraction
- Analyzing students' conceptual understanding over the course of a teaching unit: Tracking changes in knowledge structures over time
Unterrichtswissenschaft Education & learner corpora Short written responses and artifacts (drawings) from N = 300 German secondary-school students (grades 11-13) in a chemistry unit on chemical kinetics, collected via Moodle Text & span classificationError annotation & correction
- Annotate Rhetorical Relations with INCEpTION: A Comparison with Automatic Approaches
NLP & computational linguistics English cricket news reports: 10 articles segmented into 57 elementary discourse units and labelled with RST relations Discourse & rhetorical structureText & span classification
- Annotating Compositionality Scores for Irish Noun Compounds is Hard Work
arXiv (cs.CL) Linguistics & typology Irish noun compounds in the Dúchas Irish Folklore Collection (1937-1939 school transcriptions) and the Universal Dependencies Irish Dependency Treebank, scored for compositionality, domain specificity, familiarity and confidence Word sense & lexical semanticsTerminology & lexicography
- Annotating Conversational Phases and Communication Techniques: A Corpus of German Teacher-Parent Counseling Conversations
LREC 2026 Education & learner corpora German simulated teacher-parent counseling dialogues (transcripts) Text & span classificationDiscourse & rhetorical structureSpeech, multimodal & transcription
- Annotating candy speech in German YouTube comments
LAW XIX (19th Linguistic Annotation Workshop) NLP & computational linguistics German YouTube comments Sentiment, emotion & appraisal
- Annoter les expressions linguistiques de l'émotion dans des transcriptions de l'oral: faisabilité et reproductibilité
Humanités numériques en pédagogie et en recherche à la faculté des langues, 2e édition (Université de Strasbourg, 2024) Psychology & mental health French transcriptions of spoken narratives by acquired brain injury patients (GREMO-LCA project): 10 transcriptions of 58,625 patient tokens annotated twice, plus 12 further transcriptions of 85,223 patient tokens. Sentiment, emotion & appraisalSpeech, multimodal & transcription
- AraEventCoref: An Arabic Event Coreference Dataset and LLM Benchmarks
ACM Transactions on Asian and Low-Resource Language Information Processing NLP & computational linguistics 50 Arabic news articles from the SANAD dataset (Arabiya and Khaleej sources) in the political and cultural domains Coreference & anaphoraEvent extraction
- AraMIP: Extending MIPVU Towards Metaphor Identification in Arabic
arXiv cs.CL 2026 NLP & computational linguistics Arabic sentences (Modern Standard Arabic), a 300-sentence pilot corpus of 5277 words Word sense & lexical semanticsText & span classification
- Are rubrics all you need? Towards rubric-based automatic short answer scoring via guided rubric-answer alignment
LAK 2026 (16th International Learning Analytics and Knowledge Conference), Bergen, Norway Education & learner corpora German short answers written by middle- and high-school students (Gemeinschaftsschule and Gymnasium) in Schleswig-Holstein, Germany, collected in Moodle formative assessments across chemistry, biology, mathematics and physics; the new ALICE-LP dataset covering 118 questions. Text & span classificationError annotation & correction
- Automated epilepsy and seizure type phenotyping with pre-trained language models
medRxiv Biomedical & clinical Free-text outpatient epilepsy clinic progress notes from the University of Pennsylvania, 2011-2024; 309 notes annotated Text & span classification
- Automated epilepsy and seizure type phenotyping with transformer-based language models
npj Digital Medicine Biomedical & clinical 309 free-text English clinical progress notes from epilepsy clinic visits at the University of Pennsylvania, labelled with epilepsy and seizure types Text & span classification
- Automatic Annotation of Legal References (Allegationes) in the Liber Extra's Ordinary Gloss
Umanistica Digitale Law The medieval Latin Ordinary Gloss to the Liber Extra (canon-law decretals); an expert annotated legal references (allegationes) in 12 of 185 titles, 4578 annotations Named entity recognitionEntity linking & normalisation
- Automatic detection of complex quotation patterns in Aggadic literature
Cogent Arts & Humanities Classics & ancient languages Hebrew Aggadic/Midrashic rabbinic literature of Late Antiquity, with biblical quotations linked to book, chapter and verse Entity linking & normalisationText & span classification
- Automating Performance Status Annotation in Oncology Using Llama-3
Studies in Health Technology and Informatics (IOS Press) Biomedical & clinical Dutch-language oncology medical notes: 161,869 notes from 1,390 patients with palliative esophagogastric cancer, of which 300 validation and 300 test notes were manually labelled. Text & span classificationNamed entity recognition
- Avvakum Im(personal): A Categorization Approach to Information Structure and Coreference in the Historical Russian Text from the 17th Century
Junge Slavistik im Dialog Linguistics & typology The Life of Archpriest Avvakum, a 17th-century Old Russian text, annotated for third-person plural impersonal (3PL-IMP) constructions and givenness SyntaxCoreference & anaphoraDiscourse & rhetorical structure
- Building the Infrastructure for the German Medical Text Corpus Project (GeMTeX)
Intelligent Health Systems - From Technology to Data and Knowledge (MIE/Studies in Health Technology and Informatics, 2025) Biomedical & clinical German clinical documents (e.g. discharge summaries and findings reports) collected at six German university hospital data integration centres for the GeMTeX reference corpus, with over 40,000 patient broad consents obtained. Named entity recognitionEntity linking & normalisation
- CATS: An Annotation Scheme of Causality and Temporal Structure
LREC 2026 NLP & computational linguistics Portuguese news texts (Lusa) Discourse & rhetorical structureEvent extractionRelation extraction
- CLARIAH-EUS: A Strategic Network Helping Basque Country Researchers to Participate in European Research Infrastructures
Linköping Electronic Conference Proceedings Digital humanities & literature Basque learner texts in the HABE-IXA corpus, labelled with language errors for the CORPErrore resource Error annotation & correction
- Characterizing students’ energy learning trajectories
Disciplinary and Interdisciplinary Science Education Research Education & learner corpora German open-text responses written by 165 secondary-school students (9th and 10th grade) in northern Germany during a ten-week physics instructional unit on the energy concept. Text & span classification
- Cheap Annotation of Complex Information: A Study on the Annotation of Information Status in German TEDx Talks
LAW XIX (19th Linguistic Annotation Workshop) NLP & computational linguistics German TEDx talk transcripts Coreference & anaphoraDiscourse & rhetorical structure
- CitiLink-Minutes: A Multilayer Annotated Dataset of Municipal Meeting Minutes
arXiv (cs.CL) Social science & politics 120 European Portuguese municipal council meeting minutes (2021-2024) from six Portuguese municipalities, over one million tokens Metadata & bibliographicNamed entity recognitionText & span classificationRelation extraction
- ClaimPT: A Portuguese Dataset of Annotated Claims in News Articles
arXiv (cs.CL) Social science & politics European Portuguese news articles, annotated for claim vs non-claim plus claim span, object, claimer, time, stance and topic, with attribute and identity links Stance & argumentationNamed entity recognitionCoreference & anaphoraMetadata & bibliographicText & span classification
- CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations
arXiv (cs.CL) NLP & computational linguistics Code-mixed Hindi-English (Hinglish) call-centre conversations with audio and transcripts, annotated with nine discourse relations over elementary discourse units Discourse & rhetorical structureSpeech, multimodal & transcription
- Cococorpus: a corpus of copredication
21st Joint ACL-ISO Workshop on Interoperable Semantic Annotation (ISA-21) NLP & computational linguistics English sentences extracted from BookCorpus: about 1,500 gold-standard manually annotated sentences, of which roughly 200 contain copredication, covering three kinds of dot-types. Word sense & lexical semanticsSemantic roles & framesRelation extraction
- Comparing and Modeling Argumentation in German Political Communication across Arenas
KONVENS 2026 Social science & politics German political language: plenary speeches, committee meetings and press conferences on COVID-19, a 17k-sentence corpus Stance & argumentationDiscourse & rhetorical structure
- Complex Nominals in Thai: A Framework for Universal Dependencies Annotation
Linguistics & typology 100 Thai sentences containing four types of complex nominal (complement clauses, relative clauses, reduced relatives and synthetic compounds), annotated in Universal Dependencies Syntax
- ConGA: Guidelines for Contextual Gender Annotation. a Framework for Annotating Gender in Machine Translation
LREC 2026 NLP & computational linguistics English-Italian parallel MT output (gENder-IT) Translation & parallel alignmentNamed entity recognitionCoreference & anaphora
- ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
arXiv (cs.CL) Social science & politics English-language online conspiracy-theory articles (LOCO corpus and GlobalResearch), annotated in 80-120 word excerpts Text & span classificationStance & argumentationLLM output evaluation
- Context Matters - Analysis and Integration of Contextual Factors for Stance Classification Models
Doctoral thesis, Technische Universität Darmstadt NLP & computational linguistics German tweets about Covid-19 governmental containment measures: 200 tweets labelled by four experts plus 140 unique tweets per student for 21 German-speaking social-science students, labelled Unrelated / Comment / Support / Refute. Stance & argumentationText & span classification
- Context-Aware Citation Networks: A Human–AI Dataset, Analysis, and Tool
Frontiers in Artificial Intelligence and Applications Law European Court of Human Rights case law (Grand Chamber, Chamber and Committee judgments and decisions), with each citation instance in the THE LAW section labelled for Complaint Article and Judicial Consideration Metadata & bibliographicText & span classificationRelation extraction
- Contextualized Prompting For Stance Detection On Social Media
arXiv (cs.CL) Social science & politics German-language Twitter posts about governmental COVID-19 containment measures (the new covid19-de dataset), 2020-2022 Stance & argumentationText & span classification
- Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations
arXiv (cs.CL) Digital humanities & literature Literary source texts in English and Russian translated into Dutch and Catalan (poems, short stories, thrillers), in human-translation, post-edited and machine-translation modalities Translation & parallel alignmentError annotation & correctionLLM output evaluation
- Cross-Lingual Emotion Recognition in Balinese Text using Multilingual-LLMs under Peer-Collaborations Settings
EACL 2026 NLP & computational linguistics Balinese text Sentiment, emotion & appraisal
- Cross-lingual and cross-country approaches to argument component detection: a comparative study.
EACL 2026 Social science & politics French and US-English political debate/speech text Stance & argumentationTranslation & parallel alignment
- DUO_DE A1: An Annotated Corpus of Online Learning Material for Beginning Learners of German as a Foreign Language
LREC 2026 Education & learner corpora German A1 online language-course learning material SyntaxMetadata & bibliographic
- DeLTA: A Description Logic-based Annotation Schema for Constructing Expressive OWL DL Axioms from Text
DL 2026: 39th International Workshop on Description Logics Software & security English: two illustrative use cases — UK building regulations (a provision from Clause 1.24 of UK Approved Document F, part of the CODE-ACCORD corpus of annotated building regulations) and scientific claims from the Bucur et al. claims dataset. Relation extractionSemantic roles & framesEntity linking & normalisation
- Decoding the conqueror's gaze: A computational approach to Ennio Flaiano's (post)colonialism
Computational Humanities Research Digital humanities & literature Ennio Flaiano's Italian colonial novel Tempo di uccidere (1947), 5,793 sentence-level segments labelled for narrative focalization Discourse & rhetorical structureText & span classification
- Designing and Customizing AI Adaptive Dialogs in Middle School Science Classrooms
PhD dissertation, UC Berkeley Education & learner corpora English written science explanations by middle-school students in US classrooms, collected as responses to adaptive dialog prompts. Text & span classification
- Detecting Legal Citations in United Kingdom Court Judgments
EMNLP 2025 Law English UK court judgments (Cambridge Law Corpus) Named entity recognitionMetadata & bibliographic
- Detecting Literary Evaluations: Can Large Language Models Compete with Human Annotators?
DHd 2026 Nicht nur Text, nicht nur Daten (DHd2026) Digital humanities & literature 35 German-language fictional narratives published between 1800 and 2015 (corpus "Evaluative Structures in Narrative Fiction"), 10 sampled for LLM comparison LLM output evaluationSentiment, emotion & appraisal
- Developing Annotation Guidelines for CSAM Prevention Interventions: Psychosocial Risk and Protective Factors Grounded in Research and Clinical Practice
LREC 2026 Psychology & mental health German therapist-client prevention chat transcripts Text & span classificationNamed entity recognition
- Developing the German Medical Text Corpus (GeMTeX): Legal Compliance and Semantic Enrichment
LREC 2026 Biomedical & clinical German clinical documents from six university hospitals Entity linking & normalisationNamed entity recognitionOther
- Development and Validation of an AWE System “Write On with Cambi!”
Education & learner corpora US student essays in grades 3-8 (opinion, argumentative, informative and explanatory writing), labelled for compositional elements and conventions errors Stance & argumentationError annotation & correctionText & span classification
- Development of Automated Gait Assessment and Intelligent Virtual Reality Based Gait Training Systems for Post-Stroke Rehabilitation
PhD thesis in Exercise Sciences, University of Auckland Biomedical & clinical English clinical gait-analysis literature, specifically a curated deviation-cause table of post-stroke gait deviations and their causes, annotated to populate a gait knowledge graph (the downstream evaluation used a twenty-patient post-stroke gait dataset). Relation extractionEvent extractionNamed entity recognition
- DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information
arXiv (cs.CL) NLP & computational linguistics Synthetic dialog transcripts and speech-derived (ASR) transcripts in 11 languages covering emergency calls, medical anamnesis, therapy sessions, insurance, customer support, clinical interviews, police reports and group therapy Named entity recognitionSpeech, multimodal & transcriptionTranslation & parallel alignment
- Die Digitalisierung des Darmstädter Tagblatts (1740–1986)
IDS-Open, Band 16 (2026) Digital humanities & literature German historical newspaper text: the Darmstädter Tagblatt (1740-1986), with a gold standard of authorship/byline attributions built from 18 issues of the scanned 1950-1970 volumes containing roughly 4,500 articles. Named entity recognitionMetadata & bibliographicOCR, layout & document structure
- Die Rolle der Diskursgrammatik bei der Detektion und Analyse sprachlicher Praktiken
Diskursgrammatik Social science & politics German moralising language across genres from DeReKo — newspaper commentaries, letters to the editor, interviews, court reports, Bundestag plenary protocols, non-fiction and Wikipedia discussion posts Stance & argumentationSemantic roles & framesSentiment, emotion & appraisal
- Digital Transformation in Legal History Through Automatic Machine Learning Annotation
ERCIM News Law Medieval Latin canon law: legal references (allegationes) in the Glossa Ordinaria to the Liber Extra, 12 of 185 titles annotated by a domain expert (4,578 annotations) Entity linking & normalisationNamed entity recognitionMetadata & bibliographic
- Directionality in translation: Throwing new light on an old question
SKASE Journal of Translation and Interpretation NLP & computational linguistics Keystroke-logged English/German translation processes (L1>L2 and L2>L1) by translation students, rendered as linear representations and annotated for cohesive items, editing operations and translation phases Translation & parallel alignmentCoreference & anaphoraError annotation & correction
- Division, not reconciliation: Mapping news media polarisation during Australia's indigenous voice to parliament referendum
Media International Australia Social science & politics 354 Australian news articles from 26 mainstream and fringe outlets covering the 2023 Indigenous Voice to Parliament referendum Stance & argumentationSentiment, emotion & appraisalText & span classificationNamed entity recognition
- EHRI Annotator: A Web-Based Tool for Named Entity Recognition and Linking in Holocaust-Related Texts
LREC 2026 History & archives Holocaust testimonies in English, German, Hungarian Entity linking & normalisationNamed entity recognition
- ELEXIS-Ro – Adding Romanian to the ELEXIS Corpus
Computational Linguistics in Bulgaria NLP & computational linguistics The Romanian component of the ELEXIS multilingual parallel corpus, machine-translated from English and validated for lemmas, POS and morphological features Error annotation & correctionSyntaxTerminology & lexicography
- Empathy Cause Identification: Towards Unveiling Empathic Triggers in Online Interactions
Psychology & mental health English acne-support forum threads from the AcnEmpathize dataset, with the sentence causing empathy in each empathic reply marked and linked to the post it responds to Sentiment, emotion & appraisalRelation extraction
- Empathy Speaks in Metaphors: The Empathy-Metaphor Corpus of Figurative Language in Empathetic Text
LREC 2026 Psychology & mental health English acne peer-support forum posts Word sense & lexical semanticsSentiment, emotion & appraisal
- Enhancing Information Extraction from German Regulatory Documents Using Deep Learning
International Conference on Computing in Civil and Building Engineering (ICCCBE) 2024 Software & security German road and traffic construction regulations (FGSV standards); sentences carrying quantitative requirements were manually extracted from five FGSV regulatory documents and converted from PDF to plain text. Named entity recognition
- Enhancing Pragmatic Processing: A Two-Dimension Approach to Detecting Intentions in Spanish
NLP & computational linguistics Spanish tweets from three topics, annotated for global and segment-level communicative intentions (the INTENT-ES corpus); 454 tweets in the segment-intention agreement round Text & span classification
- Entity Framing and Role Portrayal in the News
arXiv (cs.CL) Social science & politics 1,378 news articles in Bulgarian, English, Hindi, European Portuguese and Russian on the Ukraine-Russia War and Climate Change, with 5,800+ entity mentions labelled for role archetypes Text & span classificationNamed entity recognitionStance & argumentation
- Evaluating RAG for French immigration law: a benchmark and baseline study
arXiv (cs.IR) Law Synthetic French residence-permit applicant profiles annotated against French immigration law sources (CESEDA, Service-Public, bilateral agreements) Text & span classificationEntity linking & normalisationQA & reading comprehension
- Evaluating the Adaptability of Large Language Models to Linguistic Variation
LREC 2026 NLP & computational linguistics French multi-genre texts (NEM.fr) Named entity recognitionLLM output evaluation
- Expert statements: Where science communication discourse meets peer review discourse
Discourse and Interaction Social science & politics English science-media-centre expert statements Discourse & rhetorical structureStance & argumentationText & span classification
- Extending CARDIO:DE: Additional annotation guidelines and evaluation of NLP approaches for clinical applications
International Journal of Medical Informatics Biomedical & clinical German cardiovascular clinical routine doctor's letters: the CARDIO:DE corpus of 500 discharge letters from Heidelberg University Hospital collected in 2020-2021 (400 fully annotated, 100 held out), yielding 304,582 token-based annotations. Named entity recognition
- Extracción de Información utilizando Modelos Generativos en Documentos del Pasado Reciente
Universidad de la República, Facultad de Ingeniería (Montevideo), MSc thesis in Data Science and Machine Learning History & archives Spanish (Uruguayan) archival documents — the 'fichas' (index cards) of the OCOA from the Berrutti archive of the Uruguayan dictatorship, held as microfilm scans; 1,000 manually curated transcriptions were produced and 515 documents were labelled with entities, relations and events. Event extractionNamed entity recognitionRelation extraction
- Fables-DTR: A Corpus of Fables Annotated for Discourse and Temporal Relations
LREC 2026 Digital humanities & literature Aesop's fables in English, European Portuguese, Polish Discourse & rhetorical structureEvent extractionRelation extraction
- Fauna e Flora setecentista: das Entidades Nomeadas aos problemas de normalização
PROPOR 2026 History & archives Eighteenth-century (1758) European Portuguese historical sources from the Torre do Tombo, transcribed by palaeographers and held in the CIDEHUSDigital repository: 87 texts covering five Alentejo municipalities (Évora, Beja, Portalegre, Vila Viçosa, Elvas), manually normalised to modern orthography before annotation. Named entity recognitionEntity linking & normalisationTerminology & lexicography
- Fine-Tuning Large Language Models with Greek Learner Corpus Data: Towards Enhanced Grammatical Error Detection
Education & learner corpora Essays by learners of Greek as a second language from the Greek Learner Corpus II (GLCII), annotated with grammatical error types Error annotation & correction
- Fine-grained Fallacy Detection with Human Label Variation
NAACL 2025 Social science & politics Italian social media posts on migration, climate change, public health Stance & argumentationText & span classification
- Fine-grained Named-Entity Recognition for the East-India Company domain
Anthology of Computers and the Humanities History & archives Early modern Dutch archival texts of the Dutch East India Company (VOC), 15 fine-grained entity tags and 8,000 mentions Named entity recognition
- Formen und Funktionen von Moralisierungen in der Gesundheitskommunikation
Social science & politics German Bundestag plenary debates from 2020 on the blood-donation ban for homosexual and transgender people, annotated for moralisation practices (72 instances across eight speeches) Stance & argumentationText & span classification
- Frame Semantics for EU Reporting Requirements: An Annotated Benchmark, Guidelines, and a Path Towards RRMV Population
NXDG 2026 (NeXt-Generation Data Governance Workshop) Law English-language EU legislative provisions: a corpus of 91 reporting requests drawn from EU legislation (Directives, Regulations, implementing acts) from the SORTIS project, with a 19-request sample double-annotated. Semantic roles & frames
- From Controlled Annotation to Context: Dependency-Based Projection of Multiword Expression
UniDive (working-group submission; venue not printed in the text) Education & learner corpora French: a CEFR-graded French as a Foreign Language (FFL) pedagogical corpus of roughly 584,000 words drawn from 40 pedagogical resources (textbooks and assessment materials), annotated for multiword expressions projected from the CEFR-graded PolyLexFLE MWE lexicon. Word sense & lexical semanticsSyntaxTerminology & lexicographyError annotation & correction
- From Debates to Diplomacy: Argument Mining Across Political Registers
12th Workshop on Argument Mining (ArgMining 2025) Social science & politics English: ArgUNSC, a new corpus of 144 United Nations Security Council speeches, manually annotated with claims, premises and support/attack links; compared against the USElecDeb corpus of U.S. presidential debates. Stance & argumentationRelation extraction
- From Heart to Words: Generating Empathetic Responses via Integrated Figurative Language and Semantic Context Signals
Findings of ACL 2025 Psychology & mental health English peer-support forum posts from the AcnEmpathize acne mental-health support community; a sampled subset of 2,492 speaker-response pairs covering 1,110 unique speaker posts, drawn from a collection of over 12K posts. Sentiment, emotion & appraisalText & span classification
- From text to data: Open-source large language models in extracting cancer related medical attributes from German pathology reports
International Journal of Medical Informatics Biomedical & clinical 522 German-language cancer pathology reports from University Medical Center Hamburg-Eppendorf, annotated by oncology-trained professionals for tumour attributes (grading, TNM status, vascular/perineural invasion, UICC stage) Named entity recognition
- GOLEMcoref: A Multilingual Coreference Dataset of Fiction
ACL 2026 Digital humanities & literature Fiction in Bahasa Indonesia, Chinese, Dutch, English, Italian, Korean, Spanish Coreference & anaphora
- GSAP-ERE: Fine-Grained Scholarly Entity and Relation Extraction Focused on Machine Learning
arXiv (cs.CL) NLP & computational linguistics Full text of 100 English machine-learning publications, annotated with 10 scholarly entity types and 18 relation types (63K entity and 35K relation mentions) Relation extractionNamed entity recognition
- GeMTeX’s De-Identification in Action: Lessons Learned & Devil’s Details
German Medical Data Sciences 2025: GMDS Illuminates Health (Studies in Health Technology and Informatics) Biomedical & clinical German clinical routine documents (discharge letters and similar records converted from MS-Word/PDF to plain text) drawn from hospital information systems at six German university hospital sites, plus the synthetic GRASCCO corpus as a pilot; the curated GeMTeX corpus reached 9,009 documents and 19,475,024 tokens by June 2025. Named entity recognition
- GePaDeSE: A New Resource for Clause-Level Aspect in German Parliamentary Debates
LREC 2026 Social science & politics German parliamentary debates (Bundestag speeches) Text & span classificationSemantic roles & frames
- Genderly: a data-centric gender bias detection system
Complex & Intelligent Systems NLP & computational linguistics English sentences harvested from web/Google Search API queries, labelled for gender-bias subtypes (generic pronoun, benevolent sexism, women's dehumanization) Text & span classificationHate speech & toxicitySentiment, emotion & appraisal
- Generation of Training Data to Distinguish Adverse Events from Medical Conditions
Studies in Health Technology and Informatics (IOS Press), 'Opening the Personal Gate between Technology and Health Care' Biomedical & clinical French-language patient messages from health discussion forums: 200 messages totalling 2,219 sentences, annotated for medications, adverse events and medical conditions. Named entity recognitionLLM output evaluation
- Global synthesis of peer-reviewed articles reveals blind spots in climate impacts research
Social science & politics English-language open-access peer-reviewed articles on socioeconomic impacts of climate hazards; 7,924 sentences from 39 full-text articles annotated Text & span classification
- Grammatik in der Interaktion – eine Fallstudie zu den interaktionalen Funktionen des Indefinitpronomens man in Lessings Dramen
Diskursgrammatik Digital humanities & literature German Baroque-to-Classical drama texts (Gryphius, Lessing, Goethe, Schiller), here two Lessing comedies, annotated for person-referring pronouns and the interactional functions of indefinite man OtherCoreference & anaphora
- Graph Embeddings to Empower Entity Retrieval
Information Retrieval Research NLP & computational linguistics English DBpedia-Entity search queries, annotated by experts with both named entities and concepts linked to Wikidata (the new "Radboud annotations") Entity linking & normalisationNamed entity recognition
- HARE: an entity and relation centric evaluation framework for histopathology reports
EMNLP 2025 Biomedical & clinical English histopathology reports (hospital and TCGA) Named entity recognitionRelation extractionLLM output evaluation
- Heroes, Villains, and Victims: Character Narratives in the WPS Agenda of the UNSC
KONVENS 2025 Social science & politics 54 English-language speeches from UN Security Council open debates on the Women, Peace and Security agenda, 2000-2019 Text & span classificationNamed entity recognitionSemantic roles & frames
- Humor and Political Discourse: a Corpus-Based Study of Italian Politicians on X
PhD thesis, Università degli Studi di Bergamo / Università degli Studi di Pavia Social science & politics Italian tweets posted by 38 Italian politicians (23 men, 15 women) on X, 7,552 tweets in total, annotated for humorous value, communicative function (e.g. aggressive), type of humor (e.g. mock-impoliteness) and humor markers (e.g. repetition) by three annotators. Sentiment, emotion & appraisalText & span classificationDiscourse & rhetorical structure
- Implicit Messages of Narratives and Evaluative Text Structures: A Network-Based Approach
Scientific Study of Literature Digital humanities & literature 35 German-language literary narratives (Kleist, Thomas Mann, Hesse, Joseph Roth and others), 77,068 tokens annotated for literary evaluations, encodings and oppositions Sentiment, emotion & appraisalRelation extractionDiscourse & rhetorical structure
- Improving Online Job Advertisement Analysis via Compositional Entity Extraction
EMNLP 2025 Finance & business German online job advertisements Named entity recognitionRelation extraction
- In Vielfalt geeint? Europäische Identitätskonstruktionen im bundesdeutschen Diskurs seit 1990
Dissertation, Universität Trier, Fachbereich II — Germanistische Linguistik Social science & politics German: 186 German Bundestag plenary speeches (the Euro-PARL sub-corpus) from debates on ratifying European integration treaties between 1992 and 2021, drawn from a larger German press and plenary-protocol corpus (Bild, FAZ, Süddeutsche Zeitung, taz, Der Spiegel, Die Zeit). Stance & argumentationText & span classification
- Integrating TEI, NER/NEL, Textometry, and Linked Data for a Semantically Enriched Interview Corpus
LREC 2026 History & archives Serbian oral-history interview transcripts Entity linking & normalisationNamed entity recognitionSummarisation & simplification
- Is Literal Annotation Enough? Building an Annotation Framework for Metonymic Named Entities in Marathi
LREC 2026 NLP & computational linguistics Marathi news text (Devanagari) Named entity recognitionWord sense & lexical semantics
- It-Sr-NER: il corpus parallelo serbo-italiano per l'apprendimento del serbo come lingua straniera
Филолог – часопис за језик књижевност и културу Education & learner corpora The It-Sr-NER Serbian-Italian parallel corpus: 10,000 aligned sentences from classical and modern Serbian and Italian literature, with named entities linked to Wikidata Entity linking & normalisationNamed entity recognitionTranslation & parallel alignment
- Knowledge Graphs Generation from Cultural Heritage Texts: Combining LLMs and Ontological Engineering for Scholarly Debates
arXiv (cs.CL) History & archives English Wikipedia articles about cultural-heritage items of disputed authenticity (documents, artifacts), annotated with item metadata, scholarly agents and authenticity opinions as a gold standard for LLM knowledge-graph extraction Relation extractionEntity linking & normalisationStance & argumentationMetadata & bibliographic
- L'IA au service du FLE: classification et annotation des expressions polylexicales pour un apprentissage inclusif
Congrès Orbicom - GTnum IA2GE - ITI LIRIC: Communication, Intelligence Artificielle, Remédiation, Éthique et Inclusion (CIAREI), Strasbourg 2025 Education & learner corpora French as a Foreign Language (FLE) pedagogical corpus of roughly 500,000 tokens drawn from FLE textbooks, DELF/TCF exam collections and graded readers, each text carrying a CEFR level A1-C2 and covering hexagonal, Québécois and other francophone varieties; about 2,800 nominal multiword expressions plus ~1,000 verbal ones form the PolyLexFLE lexicon. Terminology & lexicographyWord sense & lexical semanticsError annotation & correction
- L'oralité du discours représenté: entre unité énonciative et unités prosodiques
HAL preprint (hal-05022564), from the IXe colloque interdisciplinaire Ci-Dit Linguistics & typology Spoken French: a sub-corpus of 36 conversation recordings totalling two hours, taken from C-Oral-Rom (Cresti & Moneglia 2005) and available online in the Orféo database, with prosodic phenomena transcribed following the ICOR convention. Discourse & rhetorical structureSpeech, multimodal & transcription
- LLM-Human Alignment in Evaluating Teacher Questioning Practices: Beyond Ratings to Explanation
AIMeCon 2025 Education & learner corpora German mathematics lesson transcripts from the Global Teaching Insights (GTI) classroom observation study, with rater highlights of teacher questioning practices LLM output evaluationText & span classificationSpeech, multimodal & transcription
- LLMs Out-of-the-Box Do Not Generate Context-Appropriate Word Order in Russian
SIGDIAL 2026 NLP & computational linguistics Russian movie scripts (released 1992-2012, genres crime, comedy, drama, thriller) taken from a Kaggle movie-scripts dataset; a subset of 10 scripts was manually annotated, from which 1,543 transitive declarative sentences were extracted. OCR, layout & document structureText & span classification
- LLMs as chainsaws: evaluating open-weights generative LLMs for extracting fauna and flora from multilingual travelogues
Digital humanities & literature 58 OCR'd literary-historical travelogues from the 18th to 20th centuries in English, French, Dutch and German (~5,000 tokens each), annotated with PERSON, LOCATION, ORGANISATION, FAUNA, FLORA, BIOME, landform, weather and myth entities Named entity recognition
- Labor Lex: A New Portuguese Corpus and Pipeline for Information Extraction in Brazilian Legal Texts
EMNLP 2025 Law Brazilian Portuguese labor-court documents Relation extractionNamed entity recognition
- Learning Trajectories Toward Energy Understanding: Linking Cognitive, Metacognitive, and Affective Perspectives to Enable Adaptive Feedback
Doctoral dissertation, Kiel University Education & learner corpora German student responses and interviews from physics education in Schleswig-Holstein, Germany Text & span classification
- Legal Experts Disagree With Rationale Extraction Techniques for Explaining ECtHR Case Outcome Classification
arXiv (cs.CL) Law English European Court of Human Rights case documents, where legal experts judged whether machine-extracted rationale highlights support the violation outcome, and wrote their own gold rationales LLM output evaluationStance & argumentation
- Linked Ancient Greek and Latin (LAGL) and Wikidata: Structuring and Reusing Data of Classical Literature
Journal of Open Humanities Data Classics & ancient languages Ancient Greek and Latin sources (Athenaeus' Deipnosophists, the Suda lexicon and others), token-level annotation of named entities and of bibliographic citations of 1,236 Classical authors and works Named entity recognitionMetadata & bibliographicEntity linking & normalisation
- LuxDiagRC: A Diagnostic Reading Comprehension Corpus for Luxembourgish with Linguistic and Cognitive Annotation Layers
EACL 2026 Education & learner corpora Luxembourgish reading-comprehension texts QA & reading comprehensionSyntaxWord sense & lexical semantics
- MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking
arXiv (cs.CL) NLP & computational linguistics BBC news article titles with their accompanying images, in Hindi, Japanese, Indonesian, Vietnamese and Tamil (from the M3LS dataset), linked to Wikidata Entity linking & normalisationNamed entity recognitionSpeech, multimodal & transcription
- Mapping Meaning in Latin with Large Language Models: A Multi-Task Evaluation of Preverbed Motion Verbs and Spatial Relation Detection in LLMs
CLiC-it 2025 (11th Italian Conference on Computational Linguistics) Classics & ancient languages Classical Latin (3rd c. BCE - 2nd c. CE literary texts) Relation extractionWord sense & lexical semanticsLLM output evaluation
- Med2Story Referential: A Domain-Specific Extension of ISO 24617-9 for Clinical Narratives Annotation
LREC 2026 Biomedical & clinical European Portuguese clinical reports (AML, fictitious patients) Named entity recognitionCoreference & anaphoraRelation extraction
- Medical-FLAVORS-AECC: Spanish Oncological Metaphors Dataset
LREC 2026 Biomedical & clinical Peninsular Spanish cancer forum posts (AECC) Word sense & lexical semanticsEntity linking & normalisation
- Metaphors in Literary Machine Translation: Close but no cigar?
MT Summit 2025 NLP & computational linguistics English source and Dutch machine translations of literary fiction: 100 sentences sampled from a novel test set (four chunks of about six sentences per novel) containing 333 source metaphors, with outputs of five MT/LLM systems plus the official published translations. Error annotation & correctionTranslation & parallel alignmentWord sense & lexical semanticsLLM output evaluation
- Mining Legal Arguments to Study Judicial Formalism
arXiv (cs.CL) Law 272 Czech Supreme Court and Supreme Administrative Court decisions (the MADON dataset), with 9,183 paragraphs labelled for eight legal argument types plus holistic formalism labels Stance & argumentationText & span classification
- Moral Framing in Politics (MFiP): A new resource and models for moral framing
EMNLP 2025 Social science & politics German parliamentary debates Semantic roles & framesText & span classificationStance & argumentation
- MultiGraSCCo: A Multilingual Anonymization Benchmark with Annotations of Personal Identifiers
LREC 2026 Biomedical & clinical German synthetic clinical notes (GraSCCo) translated into 10 languages Named entity recognitionTranslation & parallel alignment
- Multidisciplinary End-to-End Document-Level Relation Extraction from Scientific Literature
ICDAR 2025 NLP & computational linguistics English scientific abstracts on coastal areas: 600 abstracts sampled from a 64,000-paper Scopus collection (1980-2023) matching 'coastal areas' or 'littoral', of which 215 were manually annotated and 416 form the released CoastRED corpus. Relation extractionNamed entity recognitionEntity linking & normalisation
- NER4all or Context is All You Need: Using LLMs for low-effort, high-performance NER on historical texts. A humanities informed approach
arXiv (cs.CL) History & archives 55 OCR'd pages of the 1921 German Baedeker travel guide to Berlin and surroundings, annotated with PER, ORG and LOC Named entity recognition
- Named Entity Recognition in the Historical Meeting Protocols of the Tartu City Council
DHNB Publications History & archives Estonian historical meeting protocols of the Tartu City Council, 1918-1940, transcribed from the National Archives of Estonia Named entity recognition
- NarratEX Dataset: Explaining the Dominant Narratives in News Texts
EMNLP 2025 Social science & politics News articles in Bulgarian, English, Portuguese, Russian Text & span classificationStance & argumentation
- Nicht nur Namen und Orte: Warum historische Annotation mehr kann (und soll)
DHd 2026 Nicht nur Text, nicht nur Daten (DHd2026) History & archives Early New High German historical records: Bernese Turmbücher and property/tower-book documents from the Ökonomien des Raums and The Flow projects, plus English medieval court rolls Named entity recognitionRelation extractionEvent extraction
- Old Swedish Entities: Developing Named Entity Recognition for Medieval Charters Written in Old Swedish
Linnaeus University, MA thesis in Digital Humanities (VT26) History & archives Old Swedish medieval charters from the Swedish National Archives' Svenskt Diplomatariums huvudkartotek (SDHK): 230 charters written 1380-1382 were automatically annotated and manually verified into a corpus of 417 charters, with an expert-annotated test set of 75 charters from 1375-1382. Named entity recognitionError annotation & correction
- On Complex Argumentation Structures in Tweets
Argumentation et Analyse du Discours, 34 | 2025 Social science & politics German tweets by members of the German federal parliament on the topic of climate change, posted 2016-2022; filtered from ~30,000 tweets down to 102 reply tweet pairs (204 tweets). Stance & argumentationDiscourse & rhetorical structure
- On-Premise Medical Information Extraction from German Doctor’s Letters under Clinical Constraints
PhD dissertation, Department of Computational Linguistics, Heidelberg University Biomedical & clinical German clinical routine text: CARDIO:DE, 500 de-identified German doctor's letters (discharge letters) from the Cardiology Department of Heidelberg University Hospital, annotated with nine medication entity classes and seven medication relation classes (27,155 annotations in v1.1; 83,869 entities and 59,810 relations reported overall). Named entity recognitionRelation extractionText & span classification
- Ontology-conformal recognition of materials entities using language models
Nature Scientific Reports Life & materials sciences English materials-science literature (materials mechanics/fatigue) Named entity recognitionEntity linking & normalisation
- Operationalizing Empathic and Supportive Communication in Natural Language Processing
Biomedical & clinical 63 transcribed German breaking-bad-news role-play conversations between medical students and standardised patients, annotated for appraisal-theoretic empathic opportunities, responses and SPIKES protocol steps Sentiment, emotion & appraisalSpeech, multimodal & transcriptionDiscourse & rhetorical structure
- Outiller la description linguistique de la continuité référentielle dans un corpus d'écrits scolaires français et italien avec une perspective didactique. (Tooling the linguistic description of referential continuity in a corpus of French and Italian school writings through NLP-based annotation, from a didactic perspective)
PhD thesis, Université Grenoble Alpes / Università di Milano-Bicocca Education & learner corpora French and Italian primary-school pupils' written texts from the longitudinal Scolinter corpus (built on Scoledit): a study sub-corpus of 148 French and 150 Italian texts (CE1-CM2 / Years 1-5), plus 140 French and 20 Italian texts annotated by student pairs in a teaching lab and 15 full French texts annotated in expert working sessions. Coreference & anaphora
- Overview of the CLPsych 2025 Shared Task: Capturing Mental Health Dynamics from Social Media Timelines
NAACL 2025 Psychology & mental health English social media timelines (Reddit/TalkLife) Text & span classificationSummarisation & simplificationSentiment, emotion & appraisal
- PAPEA: A modular pipeline for the automation of protest event analysis
Political Science Research and Methods Social science & politics German local newspaper articles reporting protests in Bremen, Dresden, Leipzig and Stuttgart, 2000-2020 (Leipziger Volkszeitung, Sächsische Zeitung, Weser-Kurier, Stuttgarter Zeitung) Event extractionText & span classification
- PREMOVE: A multilayer manually annotated dataset of PREverbed MOtion VErbs in Ancient Greek and Latin
Research Square preprint (submitted to a journal) Classics & ancient languages Ancient Greek and Latin: the PREMOVE Base Corpus of 35 texts by 29 authors from the 8th century BCE to the 2nd century CE, from which 2,835 preverbed motion-verb occurrences were manually annotated across 42 layers (verb stems, preverbs, actionality, telicity, spatial roles, participant structure, textual and geographic metadata); texts retrieved via the Perseus Digital Library and PHI. Word sense & lexical semanticsSemantic roles & framesSyntax
- Patterns of multimodal evaluation: How written language and emojis interact in Instagram comments
Patterns in Language and Communication Linguistics & typology A manually tagged subcorpus of 29,373 German Instagram comments on body-positivity posts (#bodylove, #bodyacceptance, #bodypositivity) from 2020-2021, covering written language and emojis Sentiment, emotion & appraisalDiscourse & rhetorical structureSpeech, multimodal & transcription
- Predictive power of semantic information in the use context of multiword terms for their structural disambiguation
European Public & Social Innovation Review Life & materials sciences English environmental/river multiword terms Terminology & lexicographySemantic roles & framesRelation extraction
- Preliminary Evaluation of an Open-Source LLM for Lay Translation of German Clinical Documents
NAACL 2025 Biomedical & clinical German clinical documents / tumor board protocols LLM output evaluationSummarisation & simplificationError annotation & correction
- Prompting Is All You Need – Until It Isn’t: Exploring the Limits of LLMs for Negation Detection in German Clinical Text
German Medical Data Sciences 2025: GMDS Illuminates Health (Studies in Health Technology and Informatics) Biomedical & clinical German clinical discharge letters from the Chest Pain Unit at Heidelberg University Hospital; 214 randomly selected sentence samples containing 268 negated and 362 non-negated nouns. Error annotation & correctionNamed entity recognitionLLM output evaluation
- Protests in Germany 2000–2020. A dataset based on local protest event data in Bremen, Dresden, Leipzig, and Stuttgart
German Politics Social science & politics German local newspaper articles (2000-2020) from four German cities — Bremen (Weser Kurier), Dresden, Leipzig and Stuttgart (Stuttgarter Zeitung) — yielding the ProLoc dataset of 8,190 protest events; 607 articles were double-annotated during guideline development. Event extractionText & span classificationMetadata & bibliographic
- RAGE: Roman and Greek Emotions
LREC 2026 Classics & ancient languages Ancient Greek and Latin literature Sentiment, emotion & appraisalSemantic roles & framesEntity linking & normalisation
- Reflection Before Action: Designing a Framework for Quantifying Thought Patterns for Increased Self-awareness in Personal Decision Making
arXiv (cs.HC) Psychology & mental health English pre-decision reflections written by participants in conversation with a conversational agent about big life decisions (having children, buying versus renting a house) Text & span classification
- Relation Extraction for Dutch Maritime History
Thesis History & archives Early Modern Dutch VOC (Dutch East India Company) archival records from the 17th and 18th centuries Relation extractionNamed entity recognition
- Relation Extraction from Text of Various Domains: A Serbian Case Study
Computational Linguistics in Bulgaria 2026 NLP & computational linguistics Serbian text across several domains, including literary texts, from the TeslaNER+ dataset Relation extractionNamed entity recognition
- Representing Normative Regulations in OWL DL for Automated Compliance Checking Supported by Text Annotation
arXiv (cs.AI) Law Normative building-construction regulations (Russian Building Code examples), annotated against the ifcOWL ontology Semantic roles & framesRelation extractionTerminology & lexicography
- Rhetorical echoes: source-audience style alignment in youtube immigration debates
Social Network Analysis and Mining Social science & politics English-language YouTube user comments and video transcripts from 14 U.S. news and politics channels discussing immigration (2020-2024); 3,275 comments and 80 transcripts span-annotated for divisive rhetorical techniques Stance & argumentationDiscourse & rhetorical structureText & span classification
- Say Again? The Limits of Whisper with Conversation. A Case Study on the KIParla Corpus.
LREC 2026 Linguistics & typology Spoken Italian conversation (KIParla corpus) Speech, multimodal & transcriptionDiscourse & rhetorical structure
- SemEval 2025 Task 10: Multilingual Characterization and Extraction of Narratives from Online News
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), Vienna, Austria Social science & politics Online news articles in five languages — Bulgarian, English, Hindi, Portuguese and Russian — covering the Ukraine-Russia War and Climate Change domains. Text & span classificationNamed entity recognitionStance & argumentationSummarisation & simplification
- Semantic Roles of 'Heart' in the Prose Version of the Middle High German Die Lilie
Metaphor Papers Digital humanities & literature The Middle High German devotional prose text Die Lilie (late 13th century), from the Reference Corpus Middle High German 2.1 — all 29 occurrences of the lemma hërze 'heart' Semantic roles & framesWord sense & lexical semantics
- Semplificazione e inclusione nell'italiano istituzionale. Un primo studio sul corpus ALIAS
Italiano LinguaDue Linguistics & typology The ALIAS corpus of Italian institutional texts on university alias careers (regulations, web pages, forms), plus student rewritings, annotated for gender-fair language strategies OtherText & span classificationSummarisation & simplification
- Small is Beautiful : addressing resource scarcity, language variation, and transfer challenges for automatic detection of Harmful language
PhD thesis, Sorbonne Université NLP & computational linguistics NArabizi, the romanised North African Arabic (Algerian dialect) with French code-switching found in web forums: the NArabizi Treebank of 1,287 sentences, extended with 2,169 named entity mentions (PER, LOC, ORG, COMP, OTH, PERderiv, PERderivA). Named entity recognitionSyntax
- SocCor: A Multimodal-based Multilingual Soccer Corpus for Text Data Analytics
KONVENS 2025 NLP & computational linguistics Multilingual UEFA EURO 2024 media coverage — livetickers, match reports and ASR-transcribed live commentary in German, French, Spanish, English, Italian, Dutch and Turkish Named entity recognitionEntity linking & normalisation
- Span Labeling with Large Language Models: Shell vs. Meat
BEA 2025 Education & learner corpora English written responses by test-takers of the Duolingo English Test; 142 responses annotated by the authors for five categories of shell language Discourse & rhetorical structureStance & argumentationLLM output evaluation
- Speaking on Their Behalf: Detecting Indirect Speech in Historical Danish and Norwegian Texts
EACL 2026 Digital humanities & literature 19th-century Danish and Norwegian novels Text & span classificationDiscourse & rhetorical structure
- Standards-aligned annotations reveal organizational patterns in argumentative essays at scale
Frontiers in Education Education & learner corpora Nearly 20,000 English argumentative essays written by US grade 6-8 students across nine assessment prompts, annotated sentence by sentence for organisational elements (Introduction, Controlling Idea, Conclusion) Discourse & rhetorical structureStance & argumentationText & span classification
- Strategic Transparency or Deliberate Ambiguity? A Corpus-Assisted Multimodal Analysis of Airline CSR Communication on LinkedIn
Finance & business 220 English-language LinkedIn posts from Delta Air Lines, British Airways, ITA Airways and China Southern Airlines, screened for CSR-related content and transparency signal strength Text & span classification
- Structuring cultural heritage content and context: integrating llms in ontology-driven knowledge graph extraction
PhD thesis, Università di Bologna (Dottorato in Patrimonio Culturale nell'Ecosistema Digitale, ciclo 38) Digital humanities & literature English Wikipedia articles about historical forgeries, hoaxes and authenticity debates: a 581-article corpus, of which a pilot set of seven articles (Donation of Constantine, Eremin Letter, Getty Kouros, Historia Augusta, Life of Homer, Marriage Charter of Empress Theophanu, Protocols of the Elders of Sion) was annotated, yielding 45 CH items, 235 entities, 215 interpretation acts, 132 evidences, 115 features and 308 Wikidata alignments. Entity linking & normalisationRelation extractionStance & argumentationEvent extractionMetadata & bibliographic
- Superquestions and some ways to answer them
Journal of Argumentation in Context Finance & business English earnings conference call Q&A transcripts Discourse & rhetorical structureStance & argumentationText & span classification
- Supporting teaching and learning through an AI-enhanced interactive chat activity
Education & learner corpora Written science explanations from 1,460 US students in grades 6-8 about sound-wave propagation, scored for Knowledge Integration level with spans marked for 17 disciplinary idea types Text & span classification
- Taxonomic Trace Links in Requirements Engineering
Software & security Requirements from Swedish road and railway infrastructure projects (Trafikverket), labelled with classes from a construction-domain taxonomy to build a traceability ground truth Text & span classification
- Taxonomy-Driven Knowledge Graph Construction for Domain-Specific Scientific Applications
Findings of ACL 2025 NLP & computational linguistics English climate science publications retrieved from Semantic Scholar (25 papers annotated) Named entity recognitionEntity linking & normalisationRelation extraction
- Temporal Structure in Clinical Narratives in Portuguese: Insights from Cross-Document Annotation
LREC 2026 Biomedical & clinical European Portuguese medical reports (Acute Myeloid Leukemia, IPO-Porto) Event extractionRelation extractionCoreference & anaphora
- The Basel Land Records Ground Truth: An Annotated Dataset for Information Extraction on German-Language Administrative Records
Journal of Open Humanities Data History & archives 829 excerpts (50,000+ tokens) from the Historical Land Records of Basel, 1400-1700, in premodern German, with nested entity, event and relation annotations Named entity recognitionEvent extractionRelation extraction
- The Biblical Heritage in Ancient Latin Christian Literature: Advancing Intertextual Mapping Through Sentence Embeddings
Umanistica Digitale Classics & ancient languages Augustine of Hippo's Latin commentary De Genesi ad litteram, with intertextual references mapped to Jerome's Vulgate and pre-Vulgate (Vetus Latina) biblical versions Entity linking & normalisationRelation extractionText & span classification
- The ClimateCheck Dataset: Mapping Social Media Claims About Climate Change to Corresponding Scholarly Articles
5th Workshop on Scholarly Document Processing (SDP 2025) Social science & politics English: 1,325 unique climate-related claims in lay language taken from social media (435 of them used for the shared task) paired with scientific abstracts from climate and environmental science articles, yielding 3,048 manually annotated claim-abstract pairs. Text & span classificationStance & argumentation
- The ELEXIS-WSD Parallel Sense-Annotated Corpus and South Slavic Languages: Subcorpora for Croatian, Serbian, and Slovene
South Slavic Languages in the Digital Environment JuDig Linguistics & typology Croatian, Serbian and Slovene subcorpora of the ELEXIS-WSD parallel sense-annotated corpus (2,024 sentences per language), with named entities linked to Wikidata Word sense & lexical semanticsEntity linking & normalisationNamed entity recognitionSyntaxTranslation & parallel alignment
- The German Medical Text Corpus: Early 2026 Update
LREC 2026 Biomedical & clinical German clinical routine documents from university hospitals Named entity recognitionEntity linking & normalisation
- The Incremental Process of Building an Annotation Scheme for Clinical Narratives in Portuguese: the Contribution of Human Variation Analysis
LAW XIX (19th Linguistic Annotation Workshop) Biomedical & clinical Portuguese synthetic clinical reports (AML patient) Event extractionRelation extractionNamed entity recognition
- The Moralization Corpus: Frame-Based Annotation and Analysis of Moralizing Speech Acts across Diverse Text Genres
LREC 2026 Social science & politics German political debates, news articles and online discussions Stance & argumentationSemantic roles & framesText & span classification
- The Object Order in the German Middle Field through the Lens of Information Theory: A Diachronic Study
Journal of Germanic Linguistics Linguistics & typology Historical and modern German from the fifteenth to twentieth century: sentences with both a dative and an accusative object in the middle field, drawn from the Anselm corpus (about 50 texts, 14th-16th century, 405,000+ tokens), GerManC, and TIGER. SyntaxDiscourse & rhetorical structureCoreference & anaphora
- Theoretical implications of automated discourse parsing in student writing
IJCoL – Italian Journal of Computational Linguistics, 11-2 | 2025 Education & learner corpora Italian argumentative essays written by 12th-grade students at 13 upper secondary schools in the autonomous province of Bolzano during the 2021/2022 school year; the ITACA corpus of 635 essays (424,693 tokens), of which 388 texts were manually annotated and a further evaluation set drawn from the remaining 247 texts. Discourse & rhetorical structureWord sense & lexical semanticsLLM output evaluation
- To Eat and beyond: A FrameNet-Inspired Annotation of Food and Its Uses over Time
LREC 2026 Digital humanities & literature Historical English texts (recipes, medicine/science) Semantic roles & framesWord sense & lexical semanticsLLM output evaluation
- Towards Clinical Applications of NLP: Detecting Emotion Regulation via Emotional Categories and Expression Modes in French Transcriptions
LREC 2026 Psychology & mental health French patient interview transcriptions (acquired brain injury) Sentiment, emotion & appraisalText & span classification
- Towards Early Maternal Morbidity Risk Identification by Concept Extraction from Clinical Notes in Spanish Using Fine-Tuned Transformer-Based Models
Applied System Innovation Biomedical & clinical 200 de-identified Spanish-language maternal electronic health record notes from a high-obstetric-risk referral clinic in Colombia (births 2015-2019), annotated with UMLS semantic types Named entity recognitionEntity linking & normalisation
- Towards reliable Spanish clinical text de-identification through comparative evaluation of language model approaches
Journal of Biomedical Semantics Biomedical & clinical Spanish electronic health records: the ObstEHR dataset of 500 real obstetrics EHRs from a healthcare institution in Medellín, Colombia, covering women who gave birth between 2015 and 2019, annotated for personally identifiable information (NAME, PROFESSION, LOCATION and further PII types). Named entity recognition
- Towards streamlining reproducibility studies in academic research
MSc in Business Analytics thesis, Athens University of Economics and Business NLP & computational linguistics English-language scientific research papers (PDFs of academic articles) processed for research-artifact mentions; annotations from the Research Artifact Analysis tool are mapped onto UIMA CAS files produced in INCEpTION. The thesis itself is written in English with a Greek abstract. OtherMetadata & bibliographicOCR, layout & document structure
- Tracing Sensory Events: A Frame-Based Approach to Automatically Capture and Diachronically Analyse Olfactory and Gustatory Information from Texts
Digital humanities & literature English and Italian historical texts (17th-20th century) describing smell and taste events, annotated with olfactory and gustatory frame elements (Odeuropa) Semantic roles & framesEvent extractionCoreference & anaphoraNamed entity recognition
- UD-CHILDES-BG: a dependency treebank of Bulgarian child and child-directed speech
ACL 2026 Linguistics & typology Bulgarian child and child-directed speech (CHILDES) Syntax
- UlyssesLegalNER-Br: from Legislative to Legal, a comprehensive corpus of Brazilian legal documents for Named Entity Recognition
PROPOR 2026 Law Brazilian Portuguese legal documents: 560 public documents spanning bills, case law (jurisprudência) and laws, annotated with 9 categories and 23 fine-grained entity types, extending the legislative-only UlyssesNER-Br corpus. Named entity recognitionError annotation & correction
- Uncovering Temporal Framing in the News
ACL 2026 Social science & politics English and German news articles Text & span classificationStance & argumentation
- UniCite: A Dataset and Unified Hierarchical Taxonomy for Multi-Dimensional Citation Analysis
LREC 2026 NLP & computational linguistics English scientific citation contexts (2018-2024 publications) Text & span classificationSentiment, emotion & appraisalMetadata & bibliographic
- Unveiling Empathic Triggers in Online Interactions via Empathy Cause Identification
IJCNLP-AACL 2025 Psychology & mental health English forum posts and replies from the acne.org online support community (AcnEmpathize corpus) Sentiment, emotion & appraisalText & span classification
- Utilizing Geoparsing for Mapping Natural Hazards in Europe
Water (MDPI), vol. 17, 3520 History & archives Multilingual historical hazard literature in English, German and French: a scholarly compilation on European natural hazards 1301-1500, OCRed into 21 plain-text files of about 858,825 characters, from which 4,413 toponyms were annotated. Named entity recognitionEntity linking & normalisationEvent extraction
- UzUDT: Uzbek Universal Dependencies Treebank
LREC 2026 Linguistics & typology Uzbek literary texts Syntax
- VEIL: A Benchmark for Value-Preserving Entity Identification Limitation
LREC 2026 NLP & computational linguistics English web text (DCLM/Common Crawl paragraphs) Named entity recognitionText & span classificationCoreference & anaphora
- ValuesML: A new multilingual dataset for values detection in news and political manifestos
Behavior Research Methods Social science & politics News articles and political party manifestos in nine languages (Bulgarian, Dutch, English, French, German, Greek, Hebrew, Italian, Turkish), 2648 texts and 74,231 sentences Stance & argumentationSentiment, emotion & appraisalText & span classification
- Wandel im Pronomengebrauch vom Frühneuhochdeutschen zum frühen Neuhochdeutschen
IDS-Open, Band 16 (2026) Digital humanities & literature Historical German: two corpora built by the DFG research group 'Praktiken der Personenreferenz' — 31 German dramas (Gryphius, Lessing, Schiller, Goethe) from the 17th-19th centuries, and German medical texts (plague tracts and Bäderkunden) from the 15th-17th centuries; individual texts range from 1,867 to 54,126 tokens. SyntaxWord sense & lexical semanticsError annotation & correctionOCR, layout & document structureRelation extraction
- Whose Pragmatics? Cultural Grounding as a Bottleneck for Stereotype Detection in Egyptian Arabic Social Media
ACL 2026 Social science & politics Egyptian Arabic social media comments (Facebook/TikTok/Instagram) Sentiment, emotion & appraisalHate speech & toxicityLLM output evaluation
No papers match that combination.