Locked History Actions

Diff for "seminar"

Differences between revisions 576 and 834 (spanning 258 versions)
⇤ ← Revision 576 as of 2023-10-31 14:03:40 →
Size: 13459
Comment:
← Revision 834 as of 2026-10-06 20:31:09 → ⇥
Size: 22432
Comment:
Deletions are marked like this. Additions are marked like this.
Line 3: Line 3:
= Natural Language Processing Seminar 2023–2024 = = Natural Language Processing Seminar 2026–2027 =
Line 7: Line 7:
||<style="border:0;padding-top:5px;padding-bottom:5px">'''9 October 2023'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Agnieszka Mikołajczyk-Bareła''', '''Wojciech Janowski''' (!VoiceLab), '''Piotr Pęzik''' (University of Łódź / !VoiceLab), '''Filip Żarnecki''', '''Alicja Golisowicz''' (!VoiceLab)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''[[attachment:seminarium-archiwum/2023-10-09.pdf|TRURL.AI: Fine-tuning large language models on multilingual instruction datasets]]''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk delivered in Polish.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">This talk will summarize our recent work on fine-tuning a large generative language model on bilingual instruction datasets, which resulted in the release of an open version of Trurl (trurl.ai). The motivation behind creating this model was to improve the performance of the original Llama 2 7B- and 13B-parameter models (Touvron et al. 2023), from which it was derived in a number of areas such as information extraction from customer-agent interactions and data labeling with a special focus on processing texts and instructions written in Polish. We discuss the process of optimizing the instruction datasets and the effect of the fine-tuning process on a number of selected downstream tasks.||
||<style="border:0;padding-top:5px;padding-bottom:5px">'''7 September 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Varvara Magomedova''' (University of Nova Gorica)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[https://www.youtube.com/watch?v=hveN6krWBuQ|{{attachment:seminarium-archiwum/youtube.png}}]] '''[[attachment:seminarium-archiwum/2026-09-07.pdf|Veje – a future treebank of Slovenian dialectal texts]]''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Slovenian dialect research has produced a lot of literature, but not as many data sources. The lack of available syntactically annotated corpora limits research for those, who do not have access to native speakers. This talk presents the first phase of Veje – to become a Slovenian dialect treebank. The present version of the pipeline is designed to reduce the manual work. The pipeline has two modules – normalization and annotation. In the first stage, dialectal surface forms are normalized towards standard Slovenian through a sequence of corrected lexical replacements, deterministic phonological rules, Sloleks-based morphological validation, and a constrained GaMS residual corrector. The process preserves token order and alignment wherever possible and records substitutions that cannot be morphologically verified for later review. In the second stage, normalized sentences are parsed with the non-standard Slovenian CLASSLA-Stanza model, while SloBERTa identifies whatever could not be normalized and sends it to GaMS for a second check of the annotation. The resulting extended CoNLL-U representation retains the original dialect form alongside its normalized equivalent, Universal Dependencies annotation, dialect metadata, and audit information. The current output is intentionally silver-standard: native speakers and dialectologists will correct selected data into a gold subset, which will subsequently support evaluation and model fine-tuning. I will also show the web-platform developed based on user needs study held with corpora users.||
Line 12: Line 12:
||<style="border:0;padding-top:5px;padding-bottom:5px">'''16 October 2023'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Konrad Wojtasik''', '''Vadim Shishkin''', '''Kacper Wołowiec''', '''Arkadiusz Janz''', '''Maciej Piasecki''' (Wrocław University of Science and Technology)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Evaluation of information retrieval models in zero-shot settings on different documents domains''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk delivered in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Information Retrieval over large collections of documents is an extremely important research direction in the field of natural language processing. It is a key component in question-answering systems, where the answering model often relies on information contained in a database with up-to-date knowledge. This not only allows for updating the knowledge upon which the system responds to user queries but also limits its hallucinations. Currently, information retrieval models are neural networks and require significant training resources. For many years, lexical matching methods like BM25 outperformed trained neural models in Open Domain setting, but current architectures and extensive datasets allow surpassing lexical solutions. In the presentation, I will introduce available datasets for the evaluation and training of modern information retrieval architectures in document collections from various domains, as well as future development directions.||
||<style="border:0;padding-top:5px;padding-bottom:5px">'''21 September 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Mirosław Koziarski''' (independent researcher)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[https://www.youtube.com/watch?v=OlNVqLLv_SE|{{attachment:seminarium-archiwum/youtube.png}}]] '''[[attachment:seminarium-archiwum/2026-09-21.pdf|The Digital Edition of the “Varsovian Dictionary”]]''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk in Polish.}}||
||<style="border:0;padding-left:30px;padding-bottom:5px">''The Dictionary of the Polish Language'', edited by Jan Karłowicz, Adam Kryński, and Władysław Niedźwiedzki and commonly known as the ''Varsovian Dictionary'', is one of the most important and extensive works of Polish lexicography. Its eight volumes, published between 1900 and 1927 and comprising almost 7,800 pages, document Polish from a wide range of periods, regions, registers, and domains. Although scans of the dictionary are available in digital libraries, its content has largely remained locked within images of printed pages.||
||<style="border:0;padding-left:30px;padding-bottom:5px">In this presentation, I will introduce the project of a digital edition of the Varsovian Dictionary, which develops the methodology and tools designed as part of my doctoral dissertation. I will discuss the successive stages of processing: source selection and analysis, image processing and segmentation, OCR—now also supported by AI-based tools—text correction and normalisation, hierarchical parsing of dictionary entries, validation, indexing, and the generation of derived data. The core component is a parser that transforms a linear transcription into a structured XML representation. It identifies several dozen types of segments, including headwords, variants, grammatical information, senses and subsenses, definitions, qualifiers, usage examples, etymologies, cross-references, and phraseological units, as well as the relations between them. Given the complexity and inconsistency of the source material, the process is iterative and combines automatic methods with manual verification.||
||<style="border:0;padding-left:30px;padding-bottom:5px">The online edition extends an earlier research prototype into a gradually expanding platform. Users can compare a formatted entry with its source transcription, raw XML, and the corresponding facsimile. Structuring the content also makes it possible to create new paths of access to the data: a corpus of examples, collections of derivative pairs, a consolidated index of abbreviations, quantitative statistics, and a matrix comparing the dictionary’s headword inventory with those of other Polish dictionaries.||
||<style="border:0;padding-left:30px;padding-bottom:15px">Finally, I will present the current state of the project, its quality-control mechanisms, and the limitations of automatic processing of historical lexicographic material. I will also outline planned developments.||
Line 17: Line 20:
||<style="border:0;padding-top:5px;padding-bottom:5px">'''30 October 2023'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Agnieszka Faleńska''' (University of Stuttgart)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Steps towards Bias-Aware NLP Systems''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:5px">For many, Natural Language Processing (NLP) systems have become everyday necessities, with applications ranging from automatic document translation to voice-controlled personal assistants. Recently, the increasing influence of these AI tools on human lives has raised significant concerns about the possible harm these tools can cause.||
||<style="border:0;padding-left:30px;padding-bottom:15px">In this talk, I will start by showing a few examples of such harmful behaviors and discussing their potential origins. I will argue that biases in NLP models should be addressed by advancing our understanding of their linguistic sources. Then, the talk will zoom into three compelling case studies that shed light on inequalities in commonly used training data sources: Wikipedia, instructional texts, and discussion forums. Through these case studies, I will show that regardless of the perspective on the particular demographic group (speaking about, speaking to, and speaking as), subtle biases are present in all these datasets and can perpetuate harmful outcomes of NLP models.||

||<style="border:0;padding-top:5px;padding-bottom:5px">'''13 November 2023'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Piotr Rybak''' (Institute of Computer Science, Polish Academy of Sciences)||
||<style="border:0;padding-left:30px;padding-bottom:5px">'''Advancing Polish Question Answering: Datasets and Models''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk delivered in Polish.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Although question answering (QA) is one of the most popular topics in natural language processing, until recently it was virtually absent in the Polish scientific community. However, the last few years have seen a significant increase in work related to this topic. In this talk, I will discuss what question answering is, how current QA systems work, and what datasets and models are available for Polish QA. In particular, I will discuss the resources created at IPI PAN, namely the PolQA and MAUPQA datasets and the Silver Retriever model. Finally, I will point out further directions of work that are still open when it comes to Polish question answering.||

||<style="border:0;padding-top:5px;padding-bottom:5px">'''11 December 2023''' (a series of short invited talks by Coventry Univerity researchers)||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Xiaorui Jiang''' (Coventry University)||
||<style="border:0;padding-left:30px;padding-bottom:5px">'''NLP for automating systematic reviews for evidence-based healthcare''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:10px">Systematic literature review (SLR) is the standard tool for synthesising medical and clinical evidence from the ocean of publications. SLR is extremely expensive. SLR is extremely expensive. AI can play a significant role in automating the SLR process, such as for citation screening, i.e., the selection of primary studies-based title and abstract. [[http://systematicreviewtools.com/|Some tools exist]], but they suffer from tremendous obstacles, including lack of trust. In addition, a specific characteristic of systematic review, which is the fact that each systematic review is a unique dataset and starts with no annotation, makes the problem even more challenging. In this study, we present some seminal but initial efforts on utilising the transfer learning and zero-shot learning capabilities of pretrained language models and large language models to solve or alleviate this challenge. Preliminary results are to be reported.||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Xiaorui Jiang''' (Coventry University)||
||<style="border:0;padding-left:30px;padding-bottom:5px">'''Scientific text mining and summarisation''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:10px">It is a difficult task to understand and summarise the development of scientific research areas. This task is especially cognitively demanding for postgraduate students and early-career researchers, of the whose main jobs is to identify such developments by reading a large amount of literature. Will AI help? We believe so. This short talk summarises some recent initial work on extracting the semantic backbone of a scientific area through the synergy of natural language processing and network analysis, which is believed to serve a certain type of discourse models for summarisation (in future work). As a small step from it, the second part of the talk introduces how comparison citations are utilised to improve multi-document summarisation of scientific papers.||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Xiaorui Jiang''', '''Alireza Daneshkhah''' (Coventry University)||
||<style="border:0;padding-left:30px;padding-bottom:5px">'''NLP for reducing GP workload: An early progress report''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:15x">In face of a post-COVID global economic slowdown and aging society, the primary care units in the National Healthcare Services (NHS) are receiving increasingly higher pressure, resulting in delays and errors in healthcare and patient management. AI can play a significant role in alleviating this investment-requirement discrepancy, especially in the primary care settings. A large portion of clinical diagnosis and management can be assisted with AI tools for automation and reduce delays. This short presentation reports the initial studies worked with an NHS partner on developing NLP-based solutions for the automation of clinical intention classification (to save more time for better patient treatment and management) and an early alert application for Gout Flare prediction from chief complaints (to avoid delays in patient treatment and management).||

||<style="border:0;padding-top:15px;padding-bottom:5px">'''8 January 2024''' (a series of presentation of DARIAH.Lab project results) ||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''DARIAH.Lab project team''' (Institute of Computer Science, Polish Academy of Sciences)||
||<style="border:0;padding-left:30px;padding-bottom:5px">'''Talk title will be available soon''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk delivered in Polish.}}||
||<style="border:0;padding-top:5px;padding-bottom:5px">'''19 October 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Justyna Gromada''', '''Natalia Krawczyk''' (Orange Research)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Evaluation of Conversational Agents and Interpretable Satisfaction Modeling in Sales Dialogues''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk in Polish.}}||
Line 44: Line 25:
||<style="border:0;padding-top:15px;padding-bottom:5px">'''29 January 2024'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Adam Przepiórkowski''' (Institute of Computer Science, Polish Academy of Sciences)||
||<style="border:0;padding-left:30px;padding-bottom:5px">'''Talk title will be available soon''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk delivered in Polish.}}||
||<style="border:0;padding-top:5px;padding-bottom:5px">'''22 October 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Nina Smirnova''' (GESIS – Leibniz Institute for the Social Sciences)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''The title of the talk will be made available soon''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}||
Line 49: Line 30:
||<style="border:0;padding-top:5px;padding-bottom:5px">'''9 November 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Piotr Pęzik''', '''Filip Żarnecki''', '''Wojciech Janowski''', '''Jakub Kwiatkowski''', '''Paweł Wilk''', '''Łukasz Stolarski''' (University of Łódź)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''LLMs in corpus search. From orchestration to exploration''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk in Polish.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available shortly.||
Line 50: Line 35:
||<style="border:0;padding-top:5px;padding-bottom:5px">'''19 November 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Sabina Tomkins''' (University of Michigan)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''The title of the talk will be made available soon''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available shortly.||
Line 51: Line 40:
||<style="border:0;padding-top:5px;padding-bottom:5px">'''23 November 2026'''||<rowspan=4 style="border:0;padding-left:30px;">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/pas-logo.png|PAN|width=120}}]]||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Persuasion in Parliament and Beyond: Computational Analysis of Political Discourse'''||
||<style="border:0;padding-left:30px;padding-bottom:5px">Mini-conference supported by the Polish Academy od Sciences (MOST PAN) &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talks in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''!ParlaMint – Comparable and Interoperable Parliamentary Corpora''' (Tomaž Erjavec, Jožef Stefan Institute)||
||<style="border:0;padding-left:30px;padding-bottom:10px">The talk presents the [[https://www.clarin.eu/parlamint|ParlaMint corpora of parliamentary debates]] of 29 European countries and autonomous regions, covering at least the period from 2015 to 2022, and containing over 1 billion words. The corpora are uniformly encoded, contain rich metadata about their 24 thousand speakers, and are linguistically annotated, as well as machine translated to English. We present the compilation of the corpora, including the encoding infrastructure, use of !GitHub, the production of individual corpora, the common pipeline for producing their distribution, and use of CLARIN services for dissemination. We also present associated efforts, such corpora of additional countries, and the [[https://clarinsi.github.io/parlaspeech/|ParlaSpeech]] corpora. Finally, the use of the corpora and further work are discussed.||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Beyond the Polish Parliamentary Corpus''' (Maciej Ogrodniczuk, Institute of Computer Science, Polish Academy of Sciences)||
||<style="border:0;padding-left:30px;padding-bottom:10px">The [[https://clip.ipipan.waw.pl/PPC|Polish Parliamentary Corpus]], which is also represented in ParlaMint, served as a model for a series of Polish corpora containing similar types of material, which are ideally suited to research into persuasion techniques. The presentation will introduce these datasets: the [[https://clip.ipipan.waw.pl/PRTC|Round Table Corpus]], documenting the opposition’s negotiations with the communist authorities of the Polish People’s Republic in 1989, and the Local Government Debates Corpus – a collection of transcripts from the proceedings of provincial assemblies, city councils, county councils and local councils from 2018 to 2026.||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Leveraging Persuasion and Intent for Analysis and Reasoning-based Detection of Disinformation with Large Language Models''' (Arkadiusz Modzelewski, Uniwersytet Padewski / NASK)||
||<style="border:0;padding-left:30px;padding-bottom:10px">This talk introduces novel human-annotated disinformation datasets that enable the analysis of manipulation techniques and malicious intent in Polish and English. It presents reasoning approaches that explicitly incorporate persuasion and intent to improve LLM-based disinformation detection across domains, genres, and languages. Finally, it introduces a multilingual benchmark of human- and AI-generated persuasive content, showing that AI-generated persuasive text exhibits linguistic differences and poses new challenges for detection. Overall, the talk presents new datasets, reasoning frameworks, and empirical insights toward more transparent and generalizable disinformation detection in the era of generative AI.||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''From Detecting Persuasion to Predicting Its Success: Persuasion Strategies as Predictive Signals''' (Tiziano Labruna, Fondazione Bruno Kessler)||
||<style="border:0;padding-left:30px;padding-bottom:10px">Computational work on persuasion has largely focused on detecting rhetorical techniques in text. This talk would ask a complementary question: do those techniques tell us anything about whether an argument actually changes someone's mind? Multi-Strategy Persuasion Scoring (from Labruna, EACL 2026) is a zero-shot framework in which a large language model reasons about each of six persuasion strategies independently and produces a per-strategy score, which can then be aggregated directly or used as input to a lightweight classifier. The strategy has been tested across three datasets, proving that strategy-guided reasoning consistently outperforms both direct pairwise comparison and generic chain-of-thought baselines.||
||<style="border:0;padding-left:30px;padding-bottom:15px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''A Corpus of Persuasion Techniques in Slavic Languages''' (Jakub Piskorski, Joint Research Centre of the European Commission)||
Line 52: Line 53:
||<style="border:0;padding-top:10px">Please see also [[http://nlp.ipipan.waw.pl/NLP-SEMINAR/previous-e.html|the talks given in 2000–2015]] and [[http://zil.ipipan.waw.pl/seminar-archive|2015–2023]].|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''30 November 2026'''||<rowspan=4 style="border:0;padding-left:30px;">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/pas-logo.png|PAN|width=120}}]]||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''AI without Borders'''||
||<style="border:0;padding-left:30px;padding-bottom:5px">Mini-conference supported by the Polish Academy od Sciences (MOST PAN) &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talks in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Universalist modelling and processing of idiomaticity in the PARSEME and !UniDive framework – recent developments and the state of Polish''' (Agata Savary, Paris-Saclay University)||
||<style="border:0;padding-left:30px;padding-bottom:5px">Multiword expressions (MWEs), like "biały kruk" (lit. a white crow, 'a rare person or thing') or "krew kogoś zalewa" (lit. blood floods someone, 'someone get furious'), a "być komuś na rękę (lit. to be on hand to someone, 'to be convenient'), or "wchodzić w grę" (lit. to enter the play, 'to come into play'), are combinations of words which exhibit idiosyncratic behavior on the lexical, morphological, syntactic, and semantic levels. Their most outstanding feature is their semantic non-compositionality, i.e. the fact that their meaning cannot be straightforwardly deduced form the meanings of their components.||
||<style="border:0;padding-left:30px;padding-bottom:5px">MWEs pose severe challenges in semantically-oriented NLP tasks. For instance paraphrasing sentences containing idioms is hindered by their lexical and morpho-syntactic inflexibility. Thus, when applying usual paraphrasing techniques to "krew kogoś zalewa", we often obtain incorrect paraphrases, in which the idiomatic meaning is lost, e.g. "ktoś jest zalany krwią" (someone is flooded with blood) or "krew kogoś oblewa" (lit. blood poured onto someone). One of the ways to tackle this challenge is to identify MWEs in running text and apply dedicated treatment to them.||
||<style="border:0;padding-left:30px;padding-bottom:10px">MWE identification has been the object of many efforts, and is one of the tasks where supervised encoder-based methods still outperform more recent generative LLMs. Paraphrasing of MWEs is an understudied task, especially in multilingual contexts. These two tasks have been the subject of longstanding effort within two European COST networks: PARSEME (2013-2017) and !UniDive (2022–2026), which are notably the organizers of 5 multilingual evaluation campaigns on these topics. I will summarize major challenges and findings from these campaigns and highlight future work.||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Detecting AI generated content on surface- and idea-level''' (Marzena Karpińska, Simon Fraser University)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Can Neural Network Models Learn Logic?''' (Jakub Szymanik, University of Trento)||
||<style="border:0;padding-left:30px;padding-bottom:10px">Large language models now produce chains of reasoning that look like proofs, but it remains unclear whether they have acquired the rules of inference or merely the patterns that usually accompany them. I will describe a line of joint work with Manuel Vargas Guzmán and Maciej Malicki {{attachment:seminarium-archiwum/info.png|Vargas Guzmán, M., Szymanik, J., & Malicki, M. (2024). Testing the limits of logical reasoning in neural and hybrid models. Findings of the Association for Computational Linguistics: NAACL 2024, 2267–2279. • Bertolazzi, L., Vargas Guzmán, M., Bernardi, R., Malicki, M., & Szymanik, J. (2026). Teaching small language models to learn logic through meta-learning. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL), 8049–8080. • Vargas Guzmán, M., Szymanik, J., & Malicki, M. (2026). Hybrid models for natural language reasoning: The case of syllogistic logic. Proceedings of the 23rd International Conference on Principles of Knowledge Representation and Reasoning (KR), 1143–1152.}} that tries to make this question precise enough to answer. We use syllogistic logic — a small, fully controlled fragment of natural language — as a testbed, and ask models not whether a conclusion holds but which premises are needed to derive it. This lets us separate two things usually conflated under "compositional generalization": extending a learned pattern to longer chains of reasoning, and recovering the underlying rules from complex cases to apply them to simpler ones. Models turn out to be markedly better at the first than the second, and their success depends systematically on the shape of the inference rather than its difficulty. I will end with two more encouraging results — that training on examples of the task itself substantially helps smaller models, and that even imperfect neural reasoners make effective assistants to a symbolic prover — and with what all this suggests about what "learning logic" could mean for a system that learns from data.||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Faithful Explanations in the Age of Large Language Models''' (Mateusz Lango, Charles University in Prague)||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''AI Argues Differently: Distinctive Persuasive Patterns and Communicative Strategies of LLMs''' (Agnieszka Faleńska, University of Stuttgart)||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''In Search for Universal Language Similarity Metric with Application in Multilingual AI''' (Michał Ptaszyński, Kitami Institute of Technology)||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Answering Questions and Temporal Reasoning in Large Language Models''' (Adam Jatowt, University of Innsbruck)||
||<style="border:0;padding-left:30px;padding-bottom:15px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Safety Vulnerabilities in Spoken Language Models and Web Agents''' (Karolina Stańczak, ETH Zurich)||

||<style="border:0;padding-top:10px">Please see also [[http://nlp.ipipan.waw.pl/NLP-SEMINAR/previous-e.html|the talks given in 2000–2015]] and [[http://zil.ipipan.waw.pl/seminar-archive|2015–2026]].||
Line 56: Line 73:
||<style="border:0;padding-top:5px;padding-bottom:5px">'''2 April 2020'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Stan Matwin''' (Dalhousie University)||
||<style="border:0;padding-left:30px;padding-bottom:5px">'''Efficient training of word embeddings with a focus on negative examples''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk delivered in Polish.}} {{attachment:seminarium-archiwum/icon-en.gif|Slides in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">This presentation is based on our [[https://pdfs.semanticscholar.org/1f50/db5786913b43f9668f997fc4c97d9cd18730.pdf|AAAI 2018]] and [[https://aaai.org/ojs/index.php/AAAI/article/view/4683|AAAI 2019]] papers on English word embeddings. In particular, we examine the notion of “negative examples”, the unobserved or insignificant word-context co-occurrences, in spectral methods. we provide a new formulation for the word embedding problem by proposing a new intuitive objective function that perfectly justifies the use of negative examples. With the goal of efficient learning of embeddings, we propose a kernel similarity measure for the latent space that can effectively calculate the similarities in high dimensions. Moreover, we propose an approximate alternative to our algorithm using a modified Vantage Point tree and reduce the computational complexity of the algorithm with respect to the number of words in the vocabulary. We have trained various word embedding algorithms on articles of Wikipedia with 2.3 billion tokens and show that our method outperforms the state-of-the-art in most word similarity tasks by a good margin. We will round up our discussion with some general thought s about the use of embeddings in modern NLP.||
||<style="border:0;padding-top:5px;padding-bottom:5px">'''17 November 2025''' '''(NOTE: the seminar will start at 16:00)'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Marzena Karpińska''' (Microsoft) ||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''!OneRuler: testing multilingual language models on long contexts''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk in Polish}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">In this presentation, I will look at how well language models perform when extracting information from texts of up to 128,000 tokens (approximately 100,000 words) in 26 languages, including Polish. The results of the experiments show that as the length of the context increases, the differences between languages with large and small data resources also increase. Surprisingly, even minimal changes in the command (adding the possibility that the information does not exist) cause a significant decrease in effectiveness, especially with longer texts.||


||<style="border:0;padding-top:5px;padding-bottom:5px">'''11 March 2024'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Mateusz Krubiński''' (Charles University in Prague)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Talk title will be given shortly''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in Polish.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available soon.||

Natural Language Processing Seminar 2026–2027

The NLP Seminar is organised by the Linguistic Engineering Group at the Institute of Computer Science, Polish Academy of Sciences (ICS PAS). It takes place on (some) Mondays, usually at 10:15 am, often online – please use the link next to the presentation title. All recorded talks are available on YouTube.

seminarium

7 September 2026

Varvara Magomedova (University of Nova Gorica)

https://www.youtube.com/watch?v=hveN6krWBuQ Veje – a future treebank of Slovenian dialectal texts  Talk in English.

Slovenian dialect research has produced a lot of literature, but not as many data sources. The lack of available syntactically annotated corpora limits research for those, who do not have access to native speakers. This talk presents the first phase of Veje – to become a Slovenian dialect treebank. The present version of the pipeline is designed to reduce the manual work. The pipeline has two modules – normalization and annotation. In the first stage, dialectal surface forms are normalized towards standard Slovenian through a sequence of corrected lexical replacements, deterministic phonological rules, Sloleks-based morphological validation, and a constrained GaMS residual corrector. The process preserves token order and alignment wherever possible and records substitutions that cannot be morphologically verified for later review. In the second stage, normalized sentences are parsed with the non-standard Slovenian CLASSLA-Stanza model, while SloBERTa identifies whatever could not be normalized and sends it to GaMS for a second check of the annotation. The resulting extended CoNLL-U representation retains the original dialect form alongside its normalized equivalent, Universal Dependencies annotation, dialect metadata, and audit information. The current output is intentionally silver-standard: native speakers and dialectologists will correct selected data into a gold subset, which will subsequently support evaluation and model fine-tuning. I will also show the web-platform developed based on user needs study held with corpora users.

21 September 2026

Mirosław Koziarski (independent researcher)

https://www.youtube.com/watch?v=OlNVqLLv_SE The Digital Edition of the “Varsovian Dictionary”  Talk in Polish.

The Dictionary of the Polish Language, edited by Jan Karłowicz, Adam Kryński, and Władysław Niedźwiedzki and commonly known as the Varsovian Dictionary, is one of the most important and extensive works of Polish lexicography. Its eight volumes, published between 1900 and 1927 and comprising almost 7,800 pages, document Polish from a wide range of periods, regions, registers, and domains. Although scans of the dictionary are available in digital libraries, its content has largely remained locked within images of printed pages.

In this presentation, I will introduce the project of a digital edition of the Varsovian Dictionary, which develops the methodology and tools designed as part of my doctoral dissertation. I will discuss the successive stages of processing: source selection and analysis, image processing and segmentation, OCR—now also supported by AI-based tools—text correction and normalisation, hierarchical parsing of dictionary entries, validation, indexing, and the generation of derived data. The core component is a parser that transforms a linear transcription into a structured XML representation. It identifies several dozen types of segments, including headwords, variants, grammatical information, senses and subsenses, definitions, qualifiers, usage examples, etymologies, cross-references, and phraseological units, as well as the relations between them. Given the complexity and inconsistency of the source material, the process is iterative and combines automatic methods with manual verification.

The online edition extends an earlier research prototype into a gradually expanding platform. Users can compare a formatted entry with its source transcription, raw XML, and the corresponding facsimile. Structuring the content also makes it possible to create new paths of access to the data: a corpus of examples, collections of derivative pairs, a consolidated index of abbreviations, quantitative statistics, and a matrix comparing the dictionary’s headword inventory with those of other Polish dictionaries.

Finally, I will present the current state of the project, its quality-control mechanisms, and the limitations of automatic processing of historical lexicographic material. I will also outline planned developments.

19 October 2026

Justyna Gromada, Natalia Krawczyk (Orange Research)

http://zil.ipipan.waw.pl/seminarium-online Evaluation of Conversational Agents and Interpretable Satisfaction Modeling in Sales Dialogues  Talk in Polish.

Talk summary will be made available shortly.

22 October 2026

Nina Smirnova (GESIS – Leibniz Institute for the Social Sciences)

http://zil.ipipan.waw.pl/seminarium-online The title of the talk will be made available soon  Talk in English.

Talk summary will be made available shortly.

9 November 2026

Piotr Pęzik, Filip Żarnecki, Wojciech Janowski, Jakub Kwiatkowski, Paweł Wilk, Łukasz Stolarski (University of Łódź)

http://zil.ipipan.waw.pl/seminarium-online LLMs in corpus search. From orchestration to exploration  Talk in Polish.

Talk summary will be made available shortly.

19 November 2026

Sabina Tomkins (University of Michigan)

http://zil.ipipan.waw.pl/seminarium-online The title of the talk will be made available soon  Talk in English.

Talk summary will be made available shortly.

23 November 2026

PAN

Persuasion in Parliament and Beyond: Computational Analysis of Political Discourse

Mini-conference supported by the Polish Academy od Sciences (MOST PAN)  Talks in English.

http://zil.ipipan.waw.pl/seminarium-online ParlaMint – Comparable and Interoperable Parliamentary Corpora (Tomaž Erjavec, Jožef Stefan Institute)

The talk presents the ParlaMint corpora of parliamentary debates of 29 European countries and autonomous regions, covering at least the period from 2015 to 2022, and containing over 1 billion words. The corpora are uniformly encoded, contain rich metadata about their 24 thousand speakers, and are linguistically annotated, as well as machine translated to English. We present the compilation of the corpora, including the encoding infrastructure, use of GitHub, the production of individual corpora, the common pipeline for producing their distribution, and use of CLARIN services for dissemination. We also present associated efforts, such corpora of additional countries, and the ParlaSpeech corpora. Finally, the use of the corpora and further work are discussed.

http://zil.ipipan.waw.pl/seminarium-online Beyond the Polish Parliamentary Corpus (Maciej Ogrodniczuk, Institute of Computer Science, Polish Academy of Sciences)

The Polish Parliamentary Corpus, which is also represented in ParlaMint, served as a model for a series of Polish corpora containing similar types of material, which are ideally suited to research into persuasion techniques. The presentation will introduce these datasets: the Round Table Corpus, documenting the opposition’s negotiations with the communist authorities of the Polish People’s Republic in 1989, and the Local Government Debates Corpus – a collection of transcripts from the proceedings of provincial assemblies, city councils, county councils and local councils from 2018 to 2026.

http://zil.ipipan.waw.pl/seminarium-online Leveraging Persuasion and Intent for Analysis and Reasoning-based Detection of Disinformation with Large Language Models (Arkadiusz Modzelewski, Uniwersytet Padewski / NASK)

This talk introduces novel human-annotated disinformation datasets that enable the analysis of manipulation techniques and malicious intent in Polish and English. It presents reasoning approaches that explicitly incorporate persuasion and intent to improve LLM-based disinformation detection across domains, genres, and languages. Finally, it introduces a multilingual benchmark of human- and AI-generated persuasive content, showing that AI-generated persuasive text exhibits linguistic differences and poses new challenges for detection. Overall, the talk presents new datasets, reasoning frameworks, and empirical insights toward more transparent and generalizable disinformation detection in the era of generative AI.

http://zil.ipipan.waw.pl/seminarium-online From Detecting Persuasion to Predicting Its Success: Persuasion Strategies as Predictive Signals (Tiziano Labruna, Fondazione Bruno Kessler)

Computational work on persuasion has largely focused on detecting rhetorical techniques in text. This talk would ask a complementary question: do those techniques tell us anything about whether an argument actually changes someone's mind? Multi-Strategy Persuasion Scoring (from Labruna, EACL 2026) is a zero-shot framework in which a large language model reasons about each of six persuasion strategies independently and produces a per-strategy score, which can then be aggregated directly or used as input to a lightweight classifier. The strategy has been tested across three datasets, proving that strategy-guided reasoning consistently outperforms both direct pairwise comparison and generic chain-of-thought baselines.

http://zil.ipipan.waw.pl/seminarium-online A Corpus of Persuasion Techniques in Slavic Languages (Jakub Piskorski, Joint Research Centre of the European Commission)

30 November 2026

PAN

AI without Borders

Mini-conference supported by the Polish Academy od Sciences (MOST PAN)  Talks in English.

http://zil.ipipan.waw.pl/seminarium-online Universalist modelling and processing of idiomaticity in the PARSEME and UniDive framework – recent developments and the state of Polish (Agata Savary, Paris-Saclay University)

Multiword expressions (MWEs), like "biały kruk" (lit. a white crow, 'a rare person or thing') or "krew kogoś zalewa" (lit. blood floods someone, 'someone get furious'), a "być komuś na rękę (lit. to be on hand to someone, 'to be convenient'), or "wchodzić w grę" (lit. to enter the play, 'to come into play'), are combinations of words which exhibit idiosyncratic behavior on the lexical, morphological, syntactic, and semantic levels. Their most outstanding feature is their semantic non-compositionality, i.e. the fact that their meaning cannot be straightforwardly deduced form the meanings of their components.

MWEs pose severe challenges in semantically-oriented NLP tasks. For instance paraphrasing sentences containing idioms is hindered by their lexical and morpho-syntactic inflexibility. Thus, when applying usual paraphrasing techniques to "krew kogoś zalewa", we often obtain incorrect paraphrases, in which the idiomatic meaning is lost, e.g. "ktoś jest zalany krwią" (someone is flooded with blood) or "krew kogoś oblewa" (lit. blood poured onto someone). One of the ways to tackle this challenge is to identify MWEs in running text and apply dedicated treatment to them.

MWE identification has been the object of many efforts, and is one of the tasks where supervised encoder-based methods still outperform more recent generative LLMs. Paraphrasing of MWEs is an understudied task, especially in multilingual contexts. These two tasks have been the subject of longstanding effort within two European COST networks: PARSEME (2013-2017) and UniDive (2022–2026), which are notably the organizers of 5 multilingual evaluation campaigns on these topics. I will summarize major challenges and findings from these campaigns and highlight future work.

http://zil.ipipan.waw.pl/seminarium-online Detecting AI generated content on surface- and idea-level (Marzena Karpińska, Simon Fraser University)

http://zil.ipipan.waw.pl/seminarium-online Can Neural Network Models Learn Logic? (Jakub Szymanik, University of Trento)

Large language models now produce chains of reasoning that look like proofs, but it remains unclear whether they have acquired the rules of inference or merely the patterns that usually accompany them. I will describe a line of joint work with Manuel Vargas Guzmán and Maciej Malicki Vargas Guzmán, M., Szymanik, J., & Malicki, M. (2024). Testing the limits of logical reasoning in neural and hybrid models. Findings of the Association for Computational Linguistics: NAACL 2024, 2267–2279. • Bertolazzi, L., Vargas Guzmán, M., Bernardi, R., Malicki, M., & Szymanik, J. (2026). Teaching small language models to learn logic through meta-learning. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL), 8049–8080. • Vargas Guzmán, M., Szymanik, J., & Malicki, M. (2026). Hybrid models for natural language reasoning: The case of syllogistic logic. Proceedings of the 23rd International Conference on Principles of Knowledge Representation and Reasoning (KR), 1143–1152. that tries to make this question precise enough to answer. We use syllogistic logic — a small, fully controlled fragment of natural language — as a testbed, and ask models not whether a conclusion holds but which premises are needed to derive it. This lets us separate two things usually conflated under "compositional generalization": extending a learned pattern to longer chains of reasoning, and recovering the underlying rules from complex cases to apply them to simpler ones. Models turn out to be markedly better at the first than the second, and their success depends systematically on the shape of the inference rather than its difficulty. I will end with two more encouraging results — that training on examples of the task itself substantially helps smaller models, and that even imperfect neural reasoners make effective assistants to a symbolic prover — and with what all this suggests about what "learning logic" could mean for a system that learns from data.

http://zil.ipipan.waw.pl/seminarium-online Faithful Explanations in the Age of Large Language Models (Mateusz Lango, Charles University in Prague)

http://zil.ipipan.waw.pl/seminarium-online AI Argues Differently: Distinctive Persuasive Patterns and Communicative Strategies of LLMs (Agnieszka Faleńska, University of Stuttgart)

http://zil.ipipan.waw.pl/seminarium-online In Search for Universal Language Similarity Metric with Application in Multilingual AI (Michał Ptaszyński, Kitami Institute of Technology)

http://zil.ipipan.waw.pl/seminarium-online Answering Questions and Temporal Reasoning in Large Language Models (Adam Jatowt, University of Innsbruck)

http://zil.ipipan.waw.pl/seminarium-online Safety Vulnerabilities in Spoken Language Models and Web Agents (Karolina Stańczak, ETH Zurich)

Please see also the talks given in 2000–2015 and 2015–2026.