|
Size: 871
Comment:
|
← Revision 834 as of 2026-10-06 20:31:09 ⇥
Size: 22432
Comment:
|
| Deletions are marked like this. | Additions are marked like this. |
| Line 3: | Line 3: |
| = Natural Language Processing Seminar 2016–2017 = | = Natural Language Processing Seminar 2026–2027 = |
| Line 5: | Line 5: |
| ||<style="border:0;padding:0">The NLP Seminar is organised by the [[http://nlp.ipipan.waw.pl/|Linguistic Engineering Group]] at the [[http://www.ipipan.waw.pl/en/|Institute of Computer Science]], [[http://www.pan.pl/index.php?newlang=english|Polish Academy of Sciences]] (ICS PAS). It takes place on (some) Mondays, normally at 10:15 am, in the seminar room of the ICS PAS (ul. Jana Kazimierza 5, Warszawa). ||<style="border:0;padding-left:30px">[[seminarium|{{attachment:seminar-archive/pl.png}}]]|| | ||<style="border:0;padding-bottom:10px">The NLP Seminar is organised by the [[http://nlp.ipipan.waw.pjl/|Linguistic Engineering Group]] at the [[http://www.ipipan.waw.pl/en/|Institute of Computer Science]], [[http://www.pan.pl/index.php?newlang=english|Polish Academy of Sciences]] (ICS PAS). It takes place on (some) Mondays, usually at 10:15 am, often online – please use the link next to the presentation title. All recorded talks are available on [[https://www.youtube.com/ipipan|YouTube]]. ||<style="border:0;padding-left:30px">[[seminarium|{{attachment:seminar-archive/pl.png}}]]|| |
| Line 7: | Line 7: |
| ||<style="border:0;padding-top:10px">It's summer holiday season, please come back in October! And now see [[http://nlp.ipipan.waw.pl/NLP-SEMINAR/previous-e.html|the talks given between 2000 and 2015]] and [[http://zil.ipipan.waw.pl/seminar|2015-16]].|| | ||<style="border:0;padding-top:5px;padding-bottom:5px">'''7 September 2026'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Varvara Magomedova''' (University of Nova Gorica)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[https://www.youtube.com/watch?v=hveN6krWBuQ|{{attachment:seminarium-archiwum/youtube.png}}]] '''[[attachment:seminarium-archiwum/2026-09-07.pdf|Veje – a future treebank of Slovenian dialectal texts]]'''  {{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Slovenian dialect research has produced a lot of literature, but not as many data sources. The lack of available syntactically annotated corpora limits research for those, who do not have access to native speakers. This talk presents the first phase of Veje – to become a Slovenian dialect treebank. The present version of the pipeline is designed to reduce the manual work. The pipeline has two modules – normalization and annotation. In the first stage, dialectal surface forms are normalized towards standard Slovenian through a sequence of corrected lexical replacements, deterministic phonological rules, Sloleks-based morphological validation, and a constrained GaMS residual corrector. The process preserves token order and alignment wherever possible and records substitutions that cannot be morphologically verified for later review. In the second stage, normalized sentences are parsed with the non-standard Slovenian CLASSLA-Stanza model, while SloBERTa identifies whatever could not be normalized and sends it to GaMS for a second check of the annotation. The resulting extended CoNLL-U representation retains the original dialect form alongside its normalized equivalent, Universal Dependencies annotation, dialect metadata, and audit information. The current output is intentionally silver-standard: native speakers and dialectologists will correct selected data into a gold subset, which will subsequently support evaluation and model fine-tuning. I will also show the web-platform developed based on user needs study held with corpora users.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''21 September 2026'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Mirosław Koziarski''' (independent researcher)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[https://www.youtube.com/watch?v=OlNVqLLv_SE|{{attachment:seminarium-archiwum/youtube.png}}]] '''[[attachment:seminarium-archiwum/2026-09-21.pdf|The Digital Edition of the “Varsovian Dictionary”]]'''  {{attachment:seminarium-archiwum/icon-pl.gif|Talk in Polish.}}|| ||<style="border:0;padding-left:30px;padding-bottom:5px">''The Dictionary of the Polish Language'', edited by Jan Karłowicz, Adam Kryński, and Władysław Niedźwiedzki and commonly known as the ''Varsovian Dictionary'', is one of the most important and extensive works of Polish lexicography. Its eight volumes, published between 1900 and 1927 and comprising almost 7,800 pages, document Polish from a wide range of periods, regions, registers, and domains. Although scans of the dictionary are available in digital libraries, its content has largely remained locked within images of printed pages.|| ||<style="border:0;padding-left:30px;padding-bottom:5px">In this presentation, I will introduce the project of a digital edition of the Varsovian Dictionary, which develops the methodology and tools designed as part of my doctoral dissertation. I will discuss the successive stages of processing: source selection and analysis, image processing and segmentation, OCR—now also supported by AI-based tools—text correction and normalisation, hierarchical parsing of dictionary entries, validation, indexing, and the generation of derived data. The core component is a parser that transforms a linear transcription into a structured XML representation. It identifies several dozen types of segments, including headwords, variants, grammatical information, senses and subsenses, definitions, qualifiers, usage examples, etymologies, cross-references, and phraseological units, as well as the relations between them. Given the complexity and inconsistency of the source material, the process is iterative and combines automatic methods with manual verification.|| ||<style="border:0;padding-left:30px;padding-bottom:5px">The online edition extends an earlier research prototype into a gradually expanding platform. Users can compare a formatted entry with its source transcription, raw XML, and the corresponding facsimile. Structuring the content also makes it possible to create new paths of access to the data: a corpus of examples, collections of derivative pairs, a consolidated index of abbreviations, quantitative statistics, and a matrix comparing the dictionary’s headword inventory with those of other Polish dictionaries.|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Finally, I will present the current state of the project, its quality-control mechanisms, and the limitations of automatic processing of historical lexicographic material. I will also outline planned developments.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''19 October 2026'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Justyna Gromada''', '''Natalia Krawczyk''' (Orange Research)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Evaluation of Conversational Agents and Interpretable Satisfaction Modeling in Sales Dialogues'''  {{attachment:seminarium-archiwum/icon-pl.gif|Talk in Polish.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available shortly.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''22 October 2026'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Nina Smirnova''' (GESIS – Leibniz Institute for the Social Sciences)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''The title of the talk will be made available soon'''  {{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available shortly.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''9 November 2026'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Piotr Pęzik''', '''Filip Żarnecki''', '''Wojciech Janowski''', '''Jakub Kwiatkowski''', '''Paweł Wilk''', '''Łukasz Stolarski''' (University of Łódź)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''LLMs in corpus search. From orchestration to exploration'''  {{attachment:seminarium-archiwum/icon-pl.gif|Talk in Polish.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available shortly.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''19 November 2026'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Sabina Tomkins''' (University of Michigan)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''The title of the talk will be made available soon'''  {{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available shortly.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''23 November 2026'''||<rowspan=4 style="border:0;padding-left:30px;">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/pas-logo.png|PAN|width=120}}]]|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Persuasion in Parliament and Beyond: Computational Analysis of Political Discourse'''|| ||<style="border:0;padding-left:30px;padding-bottom:5px">Mini-conference supported by the Polish Academy od Sciences (MOST PAN)  {{attachment:seminarium-archiwum/icon-en.gif|Talks in English.}}|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''!ParlaMint – Comparable and Interoperable Parliamentary Corpora''' (Tomaž Erjavec, Jožef Stefan Institute)|| ||<style="border:0;padding-left:30px;padding-bottom:10px">The talk presents the [[https://www.clarin.eu/parlamint|ParlaMint corpora of parliamentary debates]] of 29 European countries and autonomous regions, covering at least the period from 2015 to 2022, and containing over 1 billion words. The corpora are uniformly encoded, contain rich metadata about their 24 thousand speakers, and are linguistically annotated, as well as machine translated to English. We present the compilation of the corpora, including the encoding infrastructure, use of !GitHub, the production of individual corpora, the common pipeline for producing their distribution, and use of CLARIN services for dissemination. We also present associated efforts, such corpora of additional countries, and the [[https://clarinsi.github.io/parlaspeech/|ParlaSpeech]] corpora. Finally, the use of the corpora and further work are discussed.|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Beyond the Polish Parliamentary Corpus''' (Maciej Ogrodniczuk, Institute of Computer Science, Polish Academy of Sciences)|| ||<style="border:0;padding-left:30px;padding-bottom:10px">The [[https://clip.ipipan.waw.pl/PPC|Polish Parliamentary Corpus]], which is also represented in ParlaMint, served as a model for a series of Polish corpora containing similar types of material, which are ideally suited to research into persuasion techniques. The presentation will introduce these datasets: the [[https://clip.ipipan.waw.pl/PRTC|Round Table Corpus]], documenting the opposition’s negotiations with the communist authorities of the Polish People’s Republic in 1989, and the Local Government Debates Corpus – a collection of transcripts from the proceedings of provincial assemblies, city councils, county councils and local councils from 2018 to 2026.|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Leveraging Persuasion and Intent for Analysis and Reasoning-based Detection of Disinformation with Large Language Models''' (Arkadiusz Modzelewski, Uniwersytet Padewski / NASK)|| ||<style="border:0;padding-left:30px;padding-bottom:10px">This talk introduces novel human-annotated disinformation datasets that enable the analysis of manipulation techniques and malicious intent in Polish and English. It presents reasoning approaches that explicitly incorporate persuasion and intent to improve LLM-based disinformation detection across domains, genres, and languages. Finally, it introduces a multilingual benchmark of human- and AI-generated persuasive content, showing that AI-generated persuasive text exhibits linguistic differences and poses new challenges for detection. Overall, the talk presents new datasets, reasoning frameworks, and empirical insights toward more transparent and generalizable disinformation detection in the era of generative AI.|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''From Detecting Persuasion to Predicting Its Success: Persuasion Strategies as Predictive Signals''' (Tiziano Labruna, Fondazione Bruno Kessler)|| ||<style="border:0;padding-left:30px;padding-bottom:10px">Computational work on persuasion has largely focused on detecting rhetorical techniques in text. This talk would ask a complementary question: do those techniques tell us anything about whether an argument actually changes someone's mind? Multi-Strategy Persuasion Scoring (from Labruna, EACL 2026) is a zero-shot framework in which a large language model reasons about each of six persuasion strategies independently and produces a per-strategy score, which can then be aggregated directly or used as input to a lightweight classifier. The strategy has been tested across three datasets, proving that strategy-guided reasoning consistently outperforms both direct pairwise comparison and generic chain-of-thought baselines.|| ||<style="border:0;padding-left:30px;padding-bottom:15px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''A Corpus of Persuasion Techniques in Slavic Languages''' (Jakub Piskorski, Joint Research Centre of the European Commission)|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''30 November 2026'''||<rowspan=4 style="border:0;padding-left:30px;">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/pas-logo.png|PAN|width=120}}]]|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''AI without Borders'''|| ||<style="border:0;padding-left:30px;padding-bottom:5px">Mini-conference supported by the Polish Academy od Sciences (MOST PAN)  {{attachment:seminarium-archiwum/icon-en.gif|Talks in English.}}|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Universalist modelling and processing of idiomaticity in the PARSEME and !UniDive framework – recent developments and the state of Polish''' (Agata Savary, Paris-Saclay University)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">Multiword expressions (MWEs), like "biały kruk" (lit. a white crow, 'a rare person or thing') or "krew kogoś zalewa" (lit. blood floods someone, 'someone get furious'), a "być komuś na rękę (lit. to be on hand to someone, 'to be convenient'), or "wchodzić w grę" (lit. to enter the play, 'to come into play'), are combinations of words which exhibit idiosyncratic behavior on the lexical, morphological, syntactic, and semantic levels. Their most outstanding feature is their semantic non-compositionality, i.e. the fact that their meaning cannot be straightforwardly deduced form the meanings of their components.|| ||<style="border:0;padding-left:30px;padding-bottom:5px">MWEs pose severe challenges in semantically-oriented NLP tasks. For instance paraphrasing sentences containing idioms is hindered by their lexical and morpho-syntactic inflexibility. Thus, when applying usual paraphrasing techniques to "krew kogoś zalewa", we often obtain incorrect paraphrases, in which the idiomatic meaning is lost, e.g. "ktoś jest zalany krwią" (someone is flooded with blood) or "krew kogoś oblewa" (lit. blood poured onto someone). One of the ways to tackle this challenge is to identify MWEs in running text and apply dedicated treatment to them.|| ||<style="border:0;padding-left:30px;padding-bottom:10px">MWE identification has been the object of many efforts, and is one of the tasks where supervised encoder-based methods still outperform more recent generative LLMs. Paraphrasing of MWEs is an understudied task, especially in multilingual contexts. These two tasks have been the subject of longstanding effort within two European COST networks: PARSEME (2013-2017) and !UniDive (2022–2026), which are notably the organizers of 5 multilingual evaluation campaigns on these topics. I will summarize major challenges and findings from these campaigns and highlight future work.|| ||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Detecting AI generated content on surface- and idea-level''' (Marzena Karpińska, Simon Fraser University)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Can Neural Network Models Learn Logic?''' (Jakub Szymanik, University of Trento)|| ||<style="border:0;padding-left:30px;padding-bottom:10px">Large language models now produce chains of reasoning that look like proofs, but it remains unclear whether they have acquired the rules of inference or merely the patterns that usually accompany them. I will describe a line of joint work with Manuel Vargas Guzmán and Maciej Malicki {{attachment:seminarium-archiwum/info.png|Vargas Guzmán, M., Szymanik, J., & Malicki, M. (2024). Testing the limits of logical reasoning in neural and hybrid models. Findings of the Association for Computational Linguistics: NAACL 2024, 2267–2279. • Bertolazzi, L., Vargas Guzmán, M., Bernardi, R., Malicki, M., & Szymanik, J. (2026). Teaching small language models to learn logic through meta-learning. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL), 8049–8080. • Vargas Guzmán, M., Szymanik, J., & Malicki, M. (2026). Hybrid models for natural language reasoning: The case of syllogistic logic. Proceedings of the 23rd International Conference on Principles of Knowledge Representation and Reasoning (KR), 1143–1152.}} that tries to make this question precise enough to answer. We use syllogistic logic — a small, fully controlled fragment of natural language — as a testbed, and ask models not whether a conclusion holds but which premises are needed to derive it. This lets us separate two things usually conflated under "compositional generalization": extending a learned pattern to longer chains of reasoning, and recovering the underlying rules from complex cases to apply them to simpler ones. Models turn out to be markedly better at the first than the second, and their success depends systematically on the shape of the inference rather than its difficulty. I will end with two more encouraging results — that training on examples of the task itself substantially helps smaller models, and that even imperfect neural reasoners make effective assistants to a symbolic prover — and with what all this suggests about what "learning logic" could mean for a system that learns from data.|| ||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Faithful Explanations in the Age of Large Language Models''' (Mateusz Lango, Charles University in Prague)|| ||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''AI Argues Differently: Distinctive Persuasive Patterns and Communicative Strategies of LLMs''' (Agnieszka Faleńska, University of Stuttgart)|| ||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''In Search for Universal Language Similarity Metric with Application in Multilingual AI''' (Michał Ptaszyński, Kitami Institute of Technology)|| ||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Answering Questions and Temporal Reasoning in Large Language Models''' (Adam Jatowt, University of Innsbruck)|| ||<style="border:0;padding-left:30px;padding-bottom:15px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Safety Vulnerabilities in Spoken Language Models and Web Agents''' (Karolina Stańczak, ETH Zurich)|| ||<style="border:0;padding-top:10px">Please see also [[http://nlp.ipipan.waw.pl/NLP-SEMINAR/previous-e.html|the talks given in 2000–2015]] and [[http://zil.ipipan.waw.pl/seminar-archive|2015–2026]].|| {{{#!wiki comment ||<style="border:0;padding-top:5px;padding-bottom:5px">'''17 November 2025''' '''(NOTE: the seminar will start at 16:00)'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Marzena Karpińska''' (Microsoft) || ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''!OneRuler: testing multilingual language models on long contexts'''  {{attachment:seminarium-archiwum/icon-pl.gif|Talk in Polish}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">In this presentation, I will look at how well language models perform when extracting information from texts of up to 128,000 tokens (approximately 100,000 words) in 26 languages, including Polish. The results of the experiments show that as the length of the context increases, the differences between languages with large and small data resources also increase. Surprisingly, even minimal changes in the command (adding the possibility that the information does not exist) cause a significant decrease in effectiveness, especially with longer texts.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''11 March 2024'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Mateusz Krubiński''' (Charles University in Prague)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Talk title will be given shortly'''  {{attachment:seminarium-archiwum/icon-en.gif|Talk in Polish.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available soon.|| }}} |
Natural Language Processing Seminar 2026–2027
The NLP Seminar is organised by the Linguistic Engineering Group at the Institute of Computer Science, Polish Academy of Sciences (ICS PAS). It takes place on (some) Mondays, usually at 10:15 am, often online – please use the link next to the presentation title. All recorded talks are available on YouTube. |
7 September 2026 |
Varvara Magomedova (University of Nova Gorica) |
Slovenian dialect research has produced a lot of literature, but not as many data sources. The lack of available syntactically annotated corpora limits research for those, who do not have access to native speakers. This talk presents the first phase of Veje – to become a Slovenian dialect treebank. The present version of the pipeline is designed to reduce the manual work. The pipeline has two modules – normalization and annotation. In the first stage, dialectal surface forms are normalized towards standard Slovenian through a sequence of corrected lexical replacements, deterministic phonological rules, Sloleks-based morphological validation, and a constrained GaMS residual corrector. The process preserves token order and alignment wherever possible and records substitutions that cannot be morphologically verified for later review. In the second stage, normalized sentences are parsed with the non-standard Slovenian CLASSLA-Stanza model, while SloBERTa identifies whatever could not be normalized and sends it to GaMS for a second check of the annotation. The resulting extended CoNLL-U representation retains the original dialect form alongside its normalized equivalent, Universal Dependencies annotation, dialect metadata, and audit information. The current output is intentionally silver-standard: native speakers and dialectologists will correct selected data into a gold subset, which will subsequently support evaluation and model fine-tuning. I will also show the web-platform developed based on user needs study held with corpora users. |
21 September 2026 |
Mirosław Koziarski (independent researcher) |
The Dictionary of the Polish Language, edited by Jan Karłowicz, Adam Kryński, and Władysław Niedźwiedzki and commonly known as the Varsovian Dictionary, is one of the most important and extensive works of Polish lexicography. Its eight volumes, published between 1900 and 1927 and comprising almost 7,800 pages, document Polish from a wide range of periods, regions, registers, and domains. Although scans of the dictionary are available in digital libraries, its content has largely remained locked within images of printed pages. |
In this presentation, I will introduce the project of a digital edition of the Varsovian Dictionary, which develops the methodology and tools designed as part of my doctoral dissertation. I will discuss the successive stages of processing: source selection and analysis, image processing and segmentation, OCR—now also supported by AI-based tools—text correction and normalisation, hierarchical parsing of dictionary entries, validation, indexing, and the generation of derived data. The core component is a parser that transforms a linear transcription into a structured XML representation. It identifies several dozen types of segments, including headwords, variants, grammatical information, senses and subsenses, definitions, qualifiers, usage examples, etymologies, cross-references, and phraseological units, as well as the relations between them. Given the complexity and inconsistency of the source material, the process is iterative and combines automatic methods with manual verification. |
The online edition extends an earlier research prototype into a gradually expanding platform. Users can compare a formatted entry with its source transcription, raw XML, and the corresponding facsimile. Structuring the content also makes it possible to create new paths of access to the data: a corpus of examples, collections of derivative pairs, a consolidated index of abbreviations, quantitative statistics, and a matrix comparing the dictionary’s headword inventory with those of other Polish dictionaries. |
Finally, I will present the current state of the project, its quality-control mechanisms, and the limitations of automatic processing of historical lexicographic material. I will also outline planned developments. |
22 October 2026 |
Nina Smirnova (GESIS – Leibniz Institute for the Social Sciences) |
Talk summary will be made available shortly. |
9 November 2026 |
Piotr Pęzik, Filip Żarnecki, Wojciech Janowski, Jakub Kwiatkowski, Paweł Wilk, Łukasz Stolarski (University of Łódź) |
Talk summary will be made available shortly. |
19 November 2026 |
Sabina Tomkins (University of Michigan) |
Talk summary will be made available shortly. |
23 November 2026 |
|
Persuasion in Parliament and Beyond: Computational Analysis of Political Discourse |
|
Mini-conference supported by the Polish Academy od Sciences (MOST PAN) |
|
|
|
The talk presents the ParlaMint corpora of parliamentary debates of 29 European countries and autonomous regions, covering at least the period from 2015 to 2022, and containing over 1 billion words. The corpora are uniformly encoded, contain rich metadata about their 24 thousand speakers, and are linguistically annotated, as well as machine translated to English. We present the compilation of the corpora, including the encoding infrastructure, use of GitHub, the production of individual corpora, the common pipeline for producing their distribution, and use of CLARIN services for dissemination. We also present associated efforts, such corpora of additional countries, and the ParlaSpeech corpora. Finally, the use of the corpora and further work are discussed. |
|
|
|
The Polish Parliamentary Corpus, which is also represented in ParlaMint, served as a model for a series of Polish corpora containing similar types of material, which are ideally suited to research into persuasion techniques. The presentation will introduce these datasets: the Round Table Corpus, documenting the opposition’s negotiations with the communist authorities of the Polish People’s Republic in 1989, and the Local Government Debates Corpus – a collection of transcripts from the proceedings of provincial assemblies, city councils, county councils and local councils from 2018 to 2026. |
|
|
|
This talk introduces novel human-annotated disinformation datasets that enable the analysis of manipulation techniques and malicious intent in Polish and English. It presents reasoning approaches that explicitly incorporate persuasion and intent to improve LLM-based disinformation detection across domains, genres, and languages. Finally, it introduces a multilingual benchmark of human- and AI-generated persuasive content, showing that AI-generated persuasive text exhibits linguistic differences and poses new challenges for detection. Overall, the talk presents new datasets, reasoning frameworks, and empirical insights toward more transparent and generalizable disinformation detection in the era of generative AI. |
|
|
|
Computational work on persuasion has largely focused on detecting rhetorical techniques in text. This talk would ask a complementary question: do those techniques tell us anything about whether an argument actually changes someone's mind? Multi-Strategy Persuasion Scoring (from Labruna, EACL 2026) is a zero-shot framework in which a large language model reasons about each of six persuasion strategies independently and produces a per-strategy score, which can then be aggregated directly or used as input to a lightweight classifier. The strategy has been tested across three datasets, proving that strategy-guided reasoning consistently outperforms both direct pairwise comparison and generic chain-of-thought baselines. |
|
|
Please see also the talks given in 2000–2015 and 2015–2026. |





that tries to make this question precise enough to answer. We use syllogistic logic — a small, fully controlled fragment of natural language — as a testbed, and ask models not whether a conclusion holds but which premises are needed to derive it. This lets us separate two things usually conflated under "compositional generalization": extending a learned pattern to longer chains of reasoning, and recovering the underlying rules from complex cases to apply them to simpler ones. Models turn out to be markedly better at the first than the second, and their success depends systematically on the shape of the inference rather than its difficulty. I will end with two more encouraging results — that training on examples of the task itself substantially helps smaller models, and that even imperfect neural reasoners make effective assistants to a symbolic prover — and with what all this suggests about what "learning logic" could mean for a system that learns from data.