Natural Language Processing Seminar 2026–2027
The NLP Seminar is organised by the Linguistic Engineering Group at the Institute of Computer Science, Polish Academy of Sciences (ICS PAS). It takes place on (some) Mondays, usually at 10:15 am, often online – please use the link next to the presentation title. All recorded talks are available on YouTube. |
7 September 2026 |
Varvara Magomedova (University of Nova Gorica) |
Slovenian dialect research has produced a lot of literature, but not as many data sources. The lack of available syntactically annotated corpora limits research for those, who do not have access to native speakers. This talk presents the first phase of Veje – to become a Slovenian dialect treebank. The present version of the pipeline is designed to reduce the manual work. The pipeline has two modules – normalization and annotation. In the first stage, dialectal surface forms are normalized towards standard Slovenian through a sequence of corrected lexical replacements, deterministic phonological rules, Sloleks-based morphological validation, and a constrained GaMS residual corrector. The process preserves token order and alignment wherever possible and records substitutions that cannot be morphologically verified for later review. In the second stage, normalized sentences are parsed with the non-standard Slovenian CLASSLA-Stanza model, while SloBERTa identifies whatever could not be normalized and sends it to GaMS for a second check of the annotation. The resulting extended CoNLL-U representation retains the original dialect form alongside its normalized equivalent, Universal Dependencies annotation, dialect metadata, and audit information. The current output is intentionally silver-standard: native speakers and dialectologists will correct selected data into a gold subset, which will subsequently support evaluation and model fine-tuning. I will also show the web-platform developed based on user needs study held with corpora users. |
21 September 2026 |
Mirosław Koziarski (independent researcher) |
The Dictionary of the Polish Language, edited by Jan Karłowicz, Adam Kryński, and Władysław Niedźwiedzki and commonly known as the Varsovian Dictionary, is one of the most important and extensive works of Polish lexicography. Its eight volumes, published between 1900 and 1927 and comprising almost 7,800 pages, document Polish from a wide range of periods, regions, registers, and domains. Although scans of the dictionary are available in digital libraries, its content has largely remained locked within images of printed pages. |
In this presentation, I will introduce the project of a digital edition of the Varsovian Dictionary, which develops the methodology and tools designed as part of my doctoral dissertation. I will discuss the successive stages of processing: source selection and analysis, image processing and segmentation, OCR—now also supported by AI-based tools—text correction and normalisation, hierarchical parsing of dictionary entries, validation, indexing, and the generation of derived data. The core component is a parser that transforms a linear transcription into a structured XML representation. It identifies several dozen types of segments, including headwords, variants, grammatical information, senses and subsenses, definitions, qualifiers, usage examples, etymologies, cross-references, and phraseological units, as well as the relations between them. Given the complexity and inconsistency of the source material, the process is iterative and combines automatic methods with manual verification. |
The online edition extends an earlier research prototype into a gradually expanding platform. Users can compare a formatted entry with its source transcription, raw XML, and the corresponding facsimile. Structuring the content also makes it possible to create new paths of access to the data: a corpus of examples, collections of derivative pairs, a consolidated index of abbreviations, quantitative statistics, and a matrix comparing the dictionary’s headword inventory with those of other Polish dictionaries. |
Finally, I will present the current state of the project, its quality-control mechanisms, and the limitations of automatic processing of historical lexicographic material. I will also outline planned developments. |
22 October 2026 |
Nina Smirnova (GESIS – Leibniz Institute for the Social Sciences) |
Talk summary will be made available shortly. |
9 November 2026 |
Piotr Pęzik, Filip Żarnecki, Wojciech Janowski, Jakub Kwiatkowski, Paweł Wilk, Łukasz Stolarski (University of Łódź) |
Talk summary will be made available shortly. |
19 November 2026 |
Sabina Tomkins (University of Michigan) |
Talk summary will be made available shortly. |
23 November 2026 |
|
Persuasion in Parliament and Beyond: Computational Analysis of Political Discourse |
|
Mini-conference supported by the Polish Academy od Sciences (MOST PAN) |
|
|
|
The talk presents the ParlaMint corpora of parliamentary debates of 29 European countries and autonomous regions, covering at least the period from 2015 to 2022, and containing over 1 billion words. The corpora are uniformly encoded, contain rich metadata about their 24 thousand speakers, and are linguistically annotated, as well as machine translated to English. We present the compilation of the corpora, including the encoding infrastructure, use of GitHub, the production of individual corpora, the common pipeline for producing their distribution, and use of CLARIN services for dissemination. We also present associated efforts, such corpora of additional countries, and the ParlaSpeech corpora. Finally, the use of the corpora and further work are discussed. |
|
|
|
The Polish Parliamentary Corpus, which is also represented in ParlaMint, served as a model for a series of Polish corpora containing similar types of material, which are ideally suited to research into persuasion techniques. The presentation will introduce these datasets: the Round Table Corpus, documenting the opposition’s negotiations with the communist authorities of the Polish People’s Republic in 1989, and the Local Government Debates Corpus – a collection of transcripts from the proceedings of provincial assemblies, city councils, county councils and local councils from 2018 to 2026. |
|
|
|
This talk introduces novel human-annotated disinformation datasets that enable the analysis of manipulation techniques and malicious intent in Polish and English. It presents reasoning approaches that explicitly incorporate persuasion and intent to improve LLM-based disinformation detection across domains, genres, and languages. Finally, it introduces a multilingual benchmark of human- and AI-generated persuasive content, showing that AI-generated persuasive text exhibits linguistic differences and poses new challenges for detection. Overall, the talk presents new datasets, reasoning frameworks, and empirical insights toward more transparent and generalizable disinformation detection in the era of generative AI. |
|
|
|
Computational work on persuasion has largely focused on detecting rhetorical techniques in text. This talk would ask a complementary question: do those techniques tell us anything about whether an argument actually changes someone's mind? Multi-Strategy Persuasion Scoring (from Labruna, EACL 2026) is a zero-shot framework in which a large language model reasons about each of six persuasion strategies independently and produces a per-strategy score, which can then be aggregated directly or used as input to a lightweight classifier. The strategy has been tested across three datasets, proving that strategy-guided reasoning consistently outperforms both direct pairwise comparison and generic chain-of-thought baselines. |
|
|
Please see also the talks given in 2000–2015 and 2015–2026. |





that tries to make this question precise enough to answer. We use syllogistic logic — a small, fully controlled fragment of natural language — as a testbed, and ask models not whether a conclusion holds but which premises are needed to derive it. This lets us separate two things usually conflated under "compositional generalization": extending a learned pattern to longer chains of reasoning, and recovering the underlying rules from complex cases to apply them to simpler ones. Models turn out to be markedly better at the first than the second, and their success depends systematically on the shape of the inference rather than its difficulty. I will end with two more encouraging results — that training on examples of the task itself substantially helps smaller models, and that even imperfect neural reasoners make effective assistants to a symbolic prover — and with what all this suggests about what "learning logic" could mean for a system that learns from data.
17 November 2025 (NOTE: the seminar will start at 16:00)
Marzena Karpińska (Microsoft)
In this presentation, I will look at how well language models perform when extracting information from texts of up to 128,000 tokens (approximately 100,000 words) in 26 languages, including Polish. The results of the experiments show that as the length of the context increases, the differences between languages with large and small data resources also increase. Surprisingly, even minimal changes in the command (adding the possibility that the information does not exist) cause a significant decrease in effectiveness, especially with longer texts.
11 March 2024
Mateusz Krubiński (Charles University in Prague)
Talk summary will be made available soon.