Locked History Actions

Diff for "seminar"

Differences between revisions 554 and 810 (spanning 256 versions)
Revision 554 as of 2023-09-09 21:11:31
Size: 3740
Comment:
Revision 810 as of 2026-08-20 08:58:55
Size: 4735
Comment:
Deletions are marked like this. Additions are marked like this.
Line 3: Line 3:
= Natural Language Processing Seminar 2023–2024 = = Natural Language Processing Seminar 2026–2027 =
Line 7: Line 7:
||<style="border:0;padding-top:5px;padding-bottom:5px">'''2 October 2023'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Agnieszka Mikołajczyk''' (!VoiceLab), '''Piotr Pęzik''' (University of Łódź / !VoiceLab)||
||<style="border:0;padding-left:30px;padding-bottom:5px">'''Talk title will be given shortly''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk delivered in Polish.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">The summary will be available soon.||

||<style="border:0;padding-top:5px;padding-bottom:5px">'''23 October 2023'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Piotr Rybak''' (Institute of Computer Science, Polish Academy of Sciences)||
||<style="border:0;padding-left:30px;padding-bottom:5px">'''Advancing Polish Question Answering: Datasets and Models''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk delivered in Polish.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">The summary will be available soon.||
||<style="border:0;padding-top:5px;padding-bottom:5px">'''7 September 2026'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Varvara Magomedova''' (University of Nova Gorica)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Veje – a future treebank of Slovenian dialectal texts''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Slovenian dialect research has produced a lot of literature, but not as many data sources. The lack of available syntactically annotated corpora limits research for those, who do not have access to native speakers. This talk presents the first phase of Veje – to become a Slovenian dialect treebank. The present version of the pipeline is designed to reduce the manual work. The pipeline has two modules – normalization and annotation. In the first stage, dialectal surface forms are normalized towards standard Slovenian through a sequence of corrected lexical replacements, deterministic phonological rules, Sloleks-based morphological validation, and a constrained GaMS residual corrector. The process preserves token order and alignment wherever possible and records substitutions that cannot be morphologically verified for later review. In the second stage, normalized sentences are parsed with the non-standard Slovenian CLASSLA-Stanza model, while SloBERTa identifies whatever could not be normalized and sends it to GaMS for a second check of the annotation. The resulting extended CoNLL-U representation retains the original dialect form alongside its normalized equivalent, Universal Dependencies annotation, dialect metadata, and audit information. The current output is intentionally silver-standard: native speakers and dialectologists will correct selected data into a gold subset, which will subsequently support evaluation and model fine-tuning. I will also show the web-platform developed based on user needs study held with corpora users.||
Line 18: Line 13:
||<style="border:0;padding-top:10px">Please see also [[http://nlp.ipipan.waw.pl/NLP-SEMINAR/previous-e.html|the talks given in 2000–2015]] and [[http://zil.ipipan.waw.pl/seminar-archive|2015–2023]].|| ||<style="border:0;padding-top:10px">Please see also [[http://nlp.ipipan.waw.pl/NLP-SEMINAR/previous-e.html|the talks given in 2000–2015]] and [[http://zil.ipipan.waw.pl/seminar-archive|2015–2026]].||
Line 22: Line 17:
||<style="border:0;padding-top:5px;padding-bottom:5px">'''2 April 2020'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Stan Matwin''' (Dalhousie University)||
||<style="border:0;padding-left:30px;padding-bottom:5px">'''Efficient training of word embeddings with a focus on negative examples''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk delivered in Polish.}} {{attachment:seminarium-archiwum/icon-en.gif|Slides in English.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">This presentation is based on our [[https://pdfs.semanticscholar.org/1f50/db5786913b43f9668f997fc4c97d9cd18730.pdf|AAAI 2018]] and [[https://aaai.org/ojs/index.php/AAAI/article/view/4683|AAAI 2019]] papers on English word embeddings. In particular, we examine the notion of “negative examples”, the unobserved or insignificant word-context co-occurrences, in spectral methods. we provide a new formulation for the word embedding problem by proposing a new intuitive objective function that perfectly justifies the use of negative examples. With the goal of efficient learning of embeddings, we propose a kernel similarity measure for the latent space that can effectively calculate the similarities in high dimensions. Moreover, we propose an approximate alternative to our algorithm using a modified Vantage Point tree and reduce the computational complexity of the algorithm with respect to the number of words in the vocabulary. We have trained various word embedding algorithms on articles of Wikipedia with 2.3 billion tokens and show that our method outperforms the state-of-the-art in most word similarity tasks by a good margin. We will round up our discussion with some general thought s about the use of embeddings in modern NLP.||
||<style="border:0;padding-top:5px;padding-bottom:5px">'''17 November 2025''' '''(NOTE: the seminar will start at 16:00)'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Marzena Karpińska''' (Microsoft) ||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''!OneRuler: testing multilingual language models on long contexts''' &#160;{{attachment:seminarium-archiwum/icon-pl.gif|Talk in Polish}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">In this presentation, I will look at how well language models perform when extracting information from texts of up to 128,000 tokens (approximately 100,000 words) in 26 languages, including Polish. The results of the experiments show that as the length of the context increases, the differences between languages with large and small data resources also increase. Surprisingly, even minimal changes in the command (adding the possibility that the information does not exist) cause a significant decrease in effectiveness, especially with longer texts.||


||<style="border:0;padding-top:5px;padding-bottom:5px">'''11 March 2024'''||
||<style="border:0;padding-left:30px;padding-bottom:0px">'''Mateusz Krubiński''' (Charles University in Prague)||
||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Talk title will be given shortly''' &#160;{{attachment:seminarium-archiwum/icon-en.gif|Talk in Polish.}}||
||<style="border:0;padding-left:30px;padding-bottom:15px">Talk summary will be made available soon.||

Natural Language Processing Seminar 2026–2027

The NLP Seminar is organised by the Linguistic Engineering Group at the Institute of Computer Science, Polish Academy of Sciences (ICS PAS). It takes place on (some) Mondays, usually at 10:15 am, often online – please use the link next to the presentation title. All recorded talks are available on YouTube.

seminarium

7 September 2026

Varvara Magomedova (University of Nova Gorica)

http://zil.ipipan.waw.pl/seminarium-online Veje – a future treebank of Slovenian dialectal texts  Talk in English.

Slovenian dialect research has produced a lot of literature, but not as many data sources. The lack of available syntactically annotated corpora limits research for those, who do not have access to native speakers. This talk presents the first phase of Veje – to become a Slovenian dialect treebank. The present version of the pipeline is designed to reduce the manual work. The pipeline has two modules – normalization and annotation. In the first stage, dialectal surface forms are normalized towards standard Slovenian through a sequence of corrected lexical replacements, deterministic phonological rules, Sloleks-based morphological validation, and a constrained GaMS residual corrector. The process preserves token order and alignment wherever possible and records substitutions that cannot be morphologically verified for later review. In the second stage, normalized sentences are parsed with the non-standard Slovenian CLASSLA-Stanza model, while SloBERTa identifies whatever could not be normalized and sends it to GaMS for a second check of the annotation. The resulting extended CoNLL-U representation retains the original dialect form alongside its normalized equivalent, Universal Dependencies annotation, dialect metadata, and audit information. The current output is intentionally silver-standard: native speakers and dialectologists will correct selected data into a gold subset, which will subsequently support evaluation and model fine-tuning. I will also show the web-platform developed based on user needs study held with corpora users.

Please see also the talks given in 2000–2015 and 2015–2026.