Locked History Actions

Diff for "seminar"

Differences between revisions 2 and 827 (spanning 825 versions)
⇤ ← Revision 2 as of 2016-06-27 22:35:46 →
Size: 836
Comment:
← Revision 827 as of 2026-10-01 14:58:59 → ⇥
Size: 14666
Comment:
Deletions are marked like this. Additions are marked like this.
Line 3: Line 3:
= Natural Language Processing Seminar 2016–2017 = = Natural Language Processing Seminar 2026–2027 =
Line 5: Line 5:
||<style="border:0;padding:0">The NLP Seminar is organised by the [[http://nlp.ipipan.waw.pl/|Linguistic Engineering Group]] at the [[http://www.ipipan.waw.pl/en/|Institute of Computer Science]], [[http://www.pan.pl/index.php?newlang=english|Polish Academy of Sciences]] (ICS PAS). It takes place on (some) Mondays, normally at 10:15 am, in the seminar room of the ICS PAS (ul. Jana Kazimierza 5, Warszawa). ||<style="border:0;padding-left:30px">[[seminarium-archiwum|{{attachment:pl.png}}]]|| ||<style="border:0;padding-bottom:10px">The NLP Seminar is organised by the [[http://nlp.ipipan.waw.pjl/|Linguistic Engineering Group]] at the [[http://www.ipipan.waw.pl/en/|Institute of Computer Science]], [[http://www.pan.pl/index.php?newlang=english|Polish Academy of Sciences]] (ICS PAS). It takes place on (some) Mondays, usually at 10:15 am, often online – please use the link next to the presentation title. All recorded talks are available on [[https://www.youtube.com/ipipan|YouTube]]. ||<style="border:0;padding-left:30px">[[seminarium|{{attachment:seminar-archive/pl.png}}]]||
Line 7: Line 7:
||<style="border:0;padding-top:10px">Please come back in October! And now see [[http://nlp.ipipan.waw.pl/NLP-SEMINAR/previous-e.html|the talks given between 2000 and 2015]] and [[http://zil.ipipan.waw.pl/seminar|2015-16]].|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''7 September 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Varvara Magomedova''' (University of Nova Gorica)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[https://www.youtube.com/watch?v=hveN6krWBuQ|{{attachment:seminarium-archiwum/youtube.png}}]] '''[[attachment:seminarium-archiwum/2026-09-07.pdf|Veje – a future treebank of Slovenian dialectal texts]]''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Slovenian dialect research has produced a lot of literature, but not as many data sources. The lack of available syntactically annotated corpora limits research for those, who do not have access to native speakers. This talk presents the first phase of Veje – to become a Slovenian dialect treebank. The present version of the pipeline is designed to reduce the manual work. The pipeline has two modules – normalization and annotation. In the first stage, dialectal surface forms are normalized towards standard Slovenian through a sequence of corrected lexical replacements, deterministic phonological rules, Sloleks-based morphological validation, and a constrained GaMS residual corrector. The process preserves token order and alignment wherever possible and records substitutions that cannot be morphologically verified for later review. In the second stage, normalized sentences are parsed with the non-standard Slovenian CLASSLA-Stanza model, while SloBERTa identifies whatever could not be normalized and sends it to GaMS for a second check of the annotation. The resulting extended CoNLL-U representation retains the original dialect form alongside its normalized equivalent, Universal Dependencies annotation, dialect metadata, and audit information. The current output is intentionally silver-standard: native speakers and dialectologists will correct selected data into a gold subset, which will subsequently support evaluation and model fine-tuning. I will also show the web-platform developed based on user needs study held with corpora users.||

||<style="border:0;padding-top:5px;padding-bottom:5px">'''21 September 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Mirosław Koziarski''' (independent researcher)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[https://www.youtube.com/watch?v=OlNVqLLv_SE|{{attachment:seminarium-archiwum/youtube.png}}]] '''[[attachment:seminarium-archiwum/2026-09-21.pdf|The Digital Edition of the “Varsovian Dictionary”]]''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk in Polish.}}||
||<style="border:0;padding-left:30px;padding-bottom:5px">''The Dictionary of the Polish Language'', edited by Jan Karłowicz, Adam Kryński, and Władysław Niedźwiedzki and commonly known as the ''Varsovian Dictionary'', is one of the most important and extensive works of Polish lexicography. Its eight volumes, published between 1900 and 1927 and comprising almost 7,800 pages, document Polish from a wide range of periods, regions, registers, and domains. Although scans of the dictionary are available in digital libraries, its content has largely remained locked within images of printed pages.||
||<style="border:0;padding-left:30px;padding-bottom:5px">In this presentation, I will introduce the project of a digital edition of the Varsovian Dictionary, which develops the methodology and tools designed as part of my doctoral dissertation. I will discuss the successive stages of processing: source selection and analysis, image processing and segmentation, OCR—now also supported by AI-based tools—text correction and normalisation, hierarchical parsing of dictionary entries, validation, indexing, and the generation of derived data. The core component is a parser that transforms a linear transcription into a structured XML representation. It identifies several dozen types of segments, including headwords, variants, grammatical information, senses and subsenses, definitions, qualifiers, usage examples, etymologies, cross-references, and phraseological units, as well as the relations between them. Given the complexity and inconsistency of the source material, the process is iterative and combines automatic methods with manual verification.||
||<style="border:0;padding-left:30px;padding-bottom:5px">The online edition extends an earlier research prototype into a gradually expanding platform. Users can compare a formatted entry with its source transcription, raw XML, and the corresponding facsimile. Structuring the content also makes it possible to create new paths of access to the data: a corpus of examples, collections of derivative pairs, a consolidated index of abbreviations, quantitative statistics, and a matrix comparing the dictionary’s headword inventory with those of other Polish dictionaries.||
||<style="border:0;padding-left:30px;padding-bottom:15px">Finally, I will present the current state of the project, its quality-control mechanisms, and the limitations of automatic processing of historical lexicographic material. I will also outline planned developments.||

||<style="border:0;padding-top:5px;padding-bottom:5px">'''19 October 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Justyna Gromada''', '''Natalia Krawczyk''' (Orange Research)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Evaluation of Conversational Agents and Interpretable Satisfaction Modeling in Sales Dialogues''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk in Polish.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available shortly.||

||<style="border:0;padding-top:5px;padding-bottom:5px">'''22 October 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Nina Smirnova''' (GESIS – Leibniz Institute for the Social Sciences)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''The title of the talk will be made available soon''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available shortly.||

||<style="border:0;padding-top:5px;padding-bottom:5px">'''9 November 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Piotr Pęzik''' (University of Łódź)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''The title of the talk will be made available soon''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available shortly.||

||<style="border:0;padding-top:5px;padding-bottom:5px">'''19 November 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Sabina Tomkins''' (University of Michigan)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''The title of the talk will be made available soon''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available shortly.||

||<style="border:0;padding-top:5px;padding-bottom:5px">'''23 November 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Persuasion in Parliament and Beyond: Computational Analysis of Political Discourse'''||
||<style="border:0;padding-left:30px;padding-bottom:5px">Mini-conference supported by the Polish Academy od Sciences &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talks in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''!ParlaMint – Comparable and Interoperable Parliamentary Corpora''' (Tomaž Erjavec, Jožef Stefan Institute)||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Beyond the Polish Parliamentary Corpus''' (Maciej Ogrodniczuk, Institute of Computer Science, Polish Academy of Sciences)||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Leveraging Persuasion and Intent for Analysis and Reasoning-based Detection of Disinformation with Large Language Models''' (Arkadiusz Modzelewski, University of Padova)||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''From Detecting Persuasion to Predicting Its Success: Persuasion Strategies as Predictive Signals''' (Tiziano Labruna, Fondazione Bruno Kessler)||
||<style="border:0;padding-left:30px;padding-bottom:15px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''A Corpus of Persuasion Techniques in Slavic Languages''' (Jakub Piskorski, Joint Research Centre of the European Commission)||

||<style="border:0;padding-top:5px;padding-bottom:5px">'''30 November 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''AI Without Borders'''||
||<style="border:0;padding-left:30px;padding-bottom:5px">Mini-conference supported by the Polish Academy od Sciences &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talks in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Universalist modelling and processing of idiomaticity in the PARSEME and !UniDive framework – recent developments and the state of Polish''' (Agata Savary, Paris-Saclay University)||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Detecting AI generated content on surface- and idea-level''' (Marzena Karpińska, Simon Fraser University)||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Can Neural Network Models Learn Logic?''' (Jakub Szymanik, University of Trento)||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Faithful Explanations in the Age of Large Language Models''' (Mateusz Lango, Charles University in Prague)||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''AI Argues Differently: Distinctive Persuasive Patterns and Communicative Strategies of LLMs''' (Agnieszka Faleńska, University of Stuttgart)||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''In Search for Universal Language Similarity Metric with Application in Multilingual AI''' (Michał Ptaszyński, Kitami Institute of Technology)||
||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Answering Questions and Temporal Reasoning in Large Language Models''' (Adam Jatowt, University of Innsbruck)||
||<style="border:0;padding-left:30px;padding-bottom:15px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Safety Vulnerabilities in Spoken Language Models and Web Agents''' (Karolina Stańczak, ETH Zurich)||

||<style="border:0;padding-top:10px">Please see also [[http://nlp.ipipan.waw.pl/NLP-SEMINAR/previous-e.html|the talks given in 2000–2015]] and [[http://zil.ipipan.waw.pl/seminar-archive|2015–2026]].||

{{{#!wiki comment

||<style="border:0;padding-top:5px;padding-bottom:5px">'''17 November 2025''' '''(NOTE: the seminar will start at 16:00)'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Marzena Karpińska''' (Microsoft) ||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''!OneRuler: testing multilingual language models on long contexts''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk in Polish}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">In this presentation, I will look at how well language models perform when extracting information from texts of up to 128,000 tokens (approximately 100,000 words) in 26 languages, including Polish. The results of the experiments show that as the length of the context increases, the differences between languages with large and small data resources also increase. Surprisingly, even minimal changes in the command (adding the possibility that the information does not exist) cause a significant decrease in effectiveness, especially with longer texts.||


||<style="border:0;padding-top:5px;padding-bottom:5px">'''11 March 2024'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Mateusz Krubiński''' (Charles University in Prague)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Talk title will be given shortly''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in Polish.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available soon.||
}}}

Natural Language Processing Seminar 2026–2027

The NLP Seminar is organised by the Linguistic Engineering Group at the Institute of Computer Science, Polish Academy of Sciences (ICS PAS). It takes place on (some) Mondays, usually at 10:15 am, often online – please use the link next to the presentation title. All recorded talks are available on YouTube.

seminarium

7 September 2026

Varvara Magomedova (University of Nova Gorica)

https://www.youtube.com/watch?v=hveN6krWBuQ Veje – a future treebank of Slovenian dialectal texts  Talk in English.

Slovenian dialect research has produced a lot of literature, but not as many data sources. The lack of available syntactically annotated corpora limits research for those, who do not have access to native speakers. This talk presents the first phase of Veje – to become a Slovenian dialect treebank. The present version of the pipeline is designed to reduce the manual work. The pipeline has two modules – normalization and annotation. In the first stage, dialectal surface forms are normalized towards standard Slovenian through a sequence of corrected lexical replacements, deterministic phonological rules, Sloleks-based morphological validation, and a constrained GaMS residual corrector. The process preserves token order and alignment wherever possible and records substitutions that cannot be morphologically verified for later review. In the second stage, normalized sentences are parsed with the non-standard Slovenian CLASSLA-Stanza model, while SloBERTa identifies whatever could not be normalized and sends it to GaMS for a second check of the annotation. The resulting extended CoNLL-U representation retains the original dialect form alongside its normalized equivalent, Universal Dependencies annotation, dialect metadata, and audit information. The current output is intentionally silver-standard: native speakers and dialectologists will correct selected data into a gold subset, which will subsequently support evaluation and model fine-tuning. I will also show the web-platform developed based on user needs study held with corpora users.

21 September 2026

Mirosław Koziarski (independent researcher)

https://www.youtube.com/watch?v=OlNVqLLv_SE The Digital Edition of the “Varsovian Dictionary”  Talk in Polish.

The Dictionary of the Polish Language, edited by Jan Karłowicz, Adam Kryński, and Władysław Niedźwiedzki and commonly known as the Varsovian Dictionary, is one of the most important and extensive works of Polish lexicography. Its eight volumes, published between 1900 and 1927 and comprising almost 7,800 pages, document Polish from a wide range of periods, regions, registers, and domains. Although scans of the dictionary are available in digital libraries, its content has largely remained locked within images of printed pages.

In this presentation, I will introduce the project of a digital edition of the Varsovian Dictionary, which develops the methodology and tools designed as part of my doctoral dissertation. I will discuss the successive stages of processing: source selection and analysis, image processing and segmentation, OCR—now also supported by AI-based tools—text correction and normalisation, hierarchical parsing of dictionary entries, validation, indexing, and the generation of derived data. The core component is a parser that transforms a linear transcription into a structured XML representation. It identifies several dozen types of segments, including headwords, variants, grammatical information, senses and subsenses, definitions, qualifiers, usage examples, etymologies, cross-references, and phraseological units, as well as the relations between them. Given the complexity and inconsistency of the source material, the process is iterative and combines automatic methods with manual verification.

The online edition extends an earlier research prototype into a gradually expanding platform. Users can compare a formatted entry with its source transcription, raw XML, and the corresponding facsimile. Structuring the content also makes it possible to create new paths of access to the data: a corpus of examples, collections of derivative pairs, a consolidated index of abbreviations, quantitative statistics, and a matrix comparing the dictionary’s headword inventory with those of other Polish dictionaries.

Finally, I will present the current state of the project, its quality-control mechanisms, and the limitations of automatic processing of historical lexicographic material. I will also outline planned developments.

19 October 2026

Justyna Gromada, Natalia Krawczyk (Orange Research)

http://zil.ipipan.waw.pl/seminarium-online Evaluation of Conversational Agents and Interpretable Satisfaction Modeling in Sales Dialogues  Talk in Polish.

Talk summary will be made available shortly.

22 October 2026

Nina Smirnova (GESIS – Leibniz Institute for the Social Sciences)

http://zil.ipipan.waw.pl/seminarium-online The title of the talk will be made available soon  Talk in English.

Talk summary will be made available shortly.

9 November 2026

Piotr Pęzik (University of Łódź)

http://zil.ipipan.waw.pl/seminarium-online The title of the talk will be made available soon  Talk in English.

Talk summary will be made available shortly.

19 November 2026

Sabina Tomkins (University of Michigan)

http://zil.ipipan.waw.pl/seminarium-online The title of the talk will be made available soon  Talk in English.

Talk summary will be made available shortly.

23 November 2026

Persuasion in Parliament and Beyond: Computational Analysis of Political Discourse

Mini-conference supported by the Polish Academy od Sciences  Talks in English.

http://zil.ipipan.waw.pl/seminarium-online ParlaMint – Comparable and Interoperable Parliamentary Corpora (Tomaž Erjavec, Jožef Stefan Institute)

http://zil.ipipan.waw.pl/seminarium-online Beyond the Polish Parliamentary Corpus (Maciej Ogrodniczuk, Institute of Computer Science, Polish Academy of Sciences)

http://zil.ipipan.waw.pl/seminarium-online Leveraging Persuasion and Intent for Analysis and Reasoning-based Detection of Disinformation with Large Language Models (Arkadiusz Modzelewski, University of Padova)

http://zil.ipipan.waw.pl/seminarium-online From Detecting Persuasion to Predicting Its Success: Persuasion Strategies as Predictive Signals (Tiziano Labruna, Fondazione Bruno Kessler)

http://zil.ipipan.waw.pl/seminarium-online A Corpus of Persuasion Techniques in Slavic Languages (Jakub Piskorski, Joint Research Centre of the European Commission)

30 November 2026

AI Without Borders

Mini-conference supported by the Polish Academy od Sciences  Talks in English.

http://zil.ipipan.waw.pl/seminarium-online Universalist modelling and processing of idiomaticity in the PARSEME and UniDive framework – recent developments and the state of Polish (Agata Savary, Paris-Saclay University)

http://zil.ipipan.waw.pl/seminarium-online Detecting AI generated content on surface- and idea-level (Marzena Karpińska, Simon Fraser University)

http://zil.ipipan.waw.pl/seminarium-online Can Neural Network Models Learn Logic? (Jakub Szymanik, University of Trento)

http://zil.ipipan.waw.pl/seminarium-online Faithful Explanations in the Age of Large Language Models (Mateusz Lango, Charles University in Prague)

http://zil.ipipan.waw.pl/seminarium-online AI Argues Differently: Distinctive Persuasive Patterns and Communicative Strategies of LLMs (Agnieszka Faleńska, University of Stuttgart)

http://zil.ipipan.waw.pl/seminarium-online In Search for Universal Language Similarity Metric with Application in Multilingual AI (Michał Ptaszyński, Kitami Institute of Technology)

http://zil.ipipan.waw.pl/seminarium-online Answering Questions and Temporal Reasoning in Large Language Models (Adam Jatowt, University of Innsbruck)

http://zil.ipipan.waw.pl/seminarium-online Safety Vulnerabilities in Spoken Language Models and Web Agents (Karolina Stańczak, ETH Zurich)

Please see also the talks given in 2000–2015 and 2015–2026.