|
Size: 2220
Comment:
|
← Revision 1137 as of 2026-10-02 16:07:42 ⇥
Size: 28326
Comment:
|
| Deletions are marked like this. | Additions are marked like this. |
| Line 1: | Line 1: |
| ## page was renamed from seminarium-archiwum | |
| Line 3: | Line 2: |
| = Seminarium „Przetwarzanie języka naturalnego” 2016–2017 = | = Seminarium „Przetwarzanie języka naturalnego” 2026–27 = |
| Line 5: | Line 4: |
| ||<style="border:0;padding:0">Seminarium [[http://nlp.ipipan.waw.pl/|Zespołu Inżynierii Lingwistycznej]] w [[http://www.ipipan.waw.pl/|Instytucie Podstaw Informatyki]] [[http://www.pan.pl/|Polskiej Akademii Nauk]] odbywa się nieregularnie w poniedziałki zwykle o godz. 10:15 w siedzibie IPI PAN (ul. Jana Kazimierza 5, Warszawa) i ma charakter otwarty. Poszczególne referaty ogłaszane są na [[http://lists.nlp.ipipan.waw.pl/mailman/listinfo/ling|Polskiej Liście Językoznawczej]] oraz na stronie [[https://www.facebook.com/lingwistyka.komputerowa|Lingwistyka komputerowa]] na Facebooku. ||<style="border:0;padding-left:30px;">[[seminar-archive|{{attachment:en.png}}]]|| | ||<style="border:0;padding-bottom:10px">Seminarium [[http://nlp.ipipan.waw.pl/|Zespołu Inżynierii Lingwistycznej]] w [[http://www.ipipan.waw.pl/|Instytucie Podstaw Informatyki]] [[http://www.pan.pl/|Polskiej Akademii Nauk]] odbywa się średnio co 2 tygodnie, zwykle w poniedziałki o godz. 10:15 (niekiedy online – prosimy o korzystanie z linku przy tytule wystąpienia) i ma charakter otwarty. Poszczególne referaty ogłaszane są na [[http://lists.nlp.ipipan.waw.pl/mailman/listinfo/ling|Polskiej Liście Językoznawczej]] oraz na stronie [[https://www.facebook.com/lingwistyka.komputerowa|Lingwistyka komputerowa]] na Facebooku. Nagrania wystąpień dostępne są na [[https://www.youtube.com/ipipan|kanale YouTube]].||<style="border:0;padding-left:30px;">[[seminar|{{attachment:seminarium-archiwum/en.png}}]]|| |
| Line 7: | Line 6: |
| ||<style="border:0;padding:0">Obecnie trwa przerwa wakacyjna – zapraszamy na następne wystąpienia w październiku oraz do zapoznania się z [[http://nlp.ipipan.waw.pl/NLP-SEMINAR/previous-p.html|archiwum seminariów z lat 2000-2015]] oraz [[http://zil.ipipan.waw.pl/seminarium-archiwum|listą wystąpień z roku 2015-16]].|| | ||<style="border:0;padding-top:5px;padding-bottom:5px">'''7 września 2026'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Varvara Magomedova''' (Univerza v Novi Gorici)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[https://www.youtube.com/watch?v=hveN6krWBuQ|{{attachment:seminarium-archiwum/youtube.png}}]] '''[[attachment:seminarium-archiwum/2026-09-07.pdf|Veje – a future treebank of Slovenian dialectal texts]]'''  {{attachment:seminarium-archiwum/icon-en.gif|Wystąpienie w języku angielskim.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Slovenian dialect research has produced a lot of literature, but not as many data sources. The lack of available syntactically annotated corpora limits research for those, who do not have access to native speakers. This talk presents the first phase of Veje – to become a Slovenian dialect treebank. The present version of the pipeline is designed to reduce the manual work. The pipeline has two modules – normalization and annotation. In the first stage, dialectal surface forms are normalized towards standard Slovenian through a sequence of corrected lexical replacements, deterministic phonological rules, Sloleks-based morphological validation, and a constrained GaMS residual corrector. The process preserves token order and alignment wherever possible and records substitutions that cannot be morphologically verified for later review. In the second stage, normalized sentences are parsed with the non-standard Slovenian CLASSLA-Stanza model, while SloBERTa identifies whatever could not be normalized and sends it to GaMS for a second check of the annotation. The resulting extended CoNLL-U representation retains the original dialect form alongside its normalized equivalent, Universal Dependencies annotation, dialect metadata, and audit information. The current output is intentionally silver-standard: native speakers and dialectologists will correct selected data into a gold subset, which will subsequently support evaluation and model fine-tuning. I will also show the web-platform developed based on user needs study held with corpora users.|| |
| Line 9: | Line 11: |
| ##||<style="border:0;padding-top:5px;padding-bottom:5px">'''3 października 2016'''|| ##||<style="border:0;padding-left:30px;padding-bottom:0px">'''?''' (Samsung Polska)|| ##||<style="border:0;padding-left:30px;padding-bottom:5px">'''?'''  {{attachment:icon-pl.gif|Wystąpienie w języku polskim.}}|| ##||<style="border:0;padding-left:30px;padding-bottom:15px">Opis wystąpienia zostanie podany wkrótce.|| |
||<style="border:0;padding-top:5px;padding-bottom:5px">'''21 września 2026'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Mirosław Koziarski''' (niezależny badacz)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[https://www.youtube.com/watch?v=OlNVqLLv_SE|{{attachment:seminarium-archiwum/youtube.png}}]] '''[[attachment:seminarium-archiwum/2026-09-21.pdf|Cyfrowa edycja „Słownika warszawskiego”]]'''  {{attachment:seminarium-archiwum/icon-pl.gif|Wystąpienie w języku polskim.}}|| ||<style="border:0;padding-left:30px;padding-bottom:5px">''Słownik języka polskiego'' pod redakcją Jana Karłowicza, Adama Kryńskiego i Władysława Niedźwiedzkiego, znany jako ''Słownik warszawski'', należy do najważniejszych i najobszerniejszych dzieł polskiej leksykografii. Osiem tomów, opublikowanych w latach 1900–1927 i liczących łącznie niemal 7800 stron, dokumentuje polszczyznę wielu epok, regionów, rejestrów i dziedzin. Choć skany słownika są dostępne w bibliotekach cyfrowych, jego treść pozostawała dotąd w dużej mierze zamknięta w obrazie drukowanych stron.|| ||<style="border:0;padding-left:30px;padding-bottom:5px">Podczas wystąpienia przedstawię projekt cyfrowej edycji ''Słownika warszawskiego'', rozwijający metodologię i narzędzia opracowane w ramach mojej rozprawy doktorskiej. Omówię kolejne etapy przetwarzania: wybór i analizę źródła, obróbkę i segmentację obrazów, OCR — obecnie wspomagany także narzędziami opartymi na sztucznej inteligencji — korektę i normalizację tekstu, hierarchiczne parsowanie artykułów hasłowych, walidację, indeksowanie oraz generowanie danych pochodnych. Centralnym elementem procesu jest parser przekształcający liniową transkrypcję w ustrukturyzowaną reprezentację XML. Rozpoznaje on kilkadziesiąt typów segmentów, m.in. wyrazy hasłowe, warianty, informacje gramatyczne, znaczenia i podznaczenia, definicje, kwalifikatory, przykłady użycia, etymologie, odsyłacze i związki frazeologiczne, a także relacje między nimi. Ze względu na złożoność i niekonsekwencje materiału proces ma charakter iteracyjny i łączy metody automatyczne z ręczną kontrolą.|| ||<style="border:0;padding-left:30px;padding-bottom:5px">Uruchomiona edycja internetowa stanowi rozwinięcie wcześniejszej makiety badawczej w stopniowo rozbudowywaną platformę. Użytkownik może zestawić sformatowany artykuł z transkrypcją źródłową, surowym XML-em i odpowiednim fragmentem faksymile. Ustrukturyzowanie treści umożliwia ponadto tworzenie nowych ścieżek dostępu do danych: korpusu przykładów, zestawień derywatów, skonsolidowanego indeksu skróceń, statystyk oraz matrycy porównującej siatkę hasłową z innymi słownikami polszczyzny.|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Na zakończenie przedstawię aktualny stan projektu, przyjęte mechanizmy kontroli jakości i ograniczenia automatycznego przetwarzania historycznego materiału leksykograficznego, a także planowane kierunki rozwoju.|| |
| Line 14: | Line 19: |
| ##||<style="border:0;padding-top:5px;padding-bottom:5px">'''17 października 2016'''|| ##||<style="border:0;padding-left:30px;padding-bottom:0px">'''Adam Przepiórkowski, Jakub Kozakoszczak, Jan Winkowski, Daniel Ziembicki, Tadeusz Teleżyński''' (Instytut Podstaw Informatyki PAN, Uniwersytet Warszawski)|| ##||<style="border:0;padding-left:30px;padding-bottom:5px">'''Korpus sformalizowanych kroków wynikania tekstowego'''  {{attachment:icon-pl.gif|Wystąpienie w języku polskim.}}|| ##||<style="border:0;padding-left:30px;padding-bottom:15px">Opis wystąpienia zostanie podany wkrótce.|| |
||<style="border:0;padding-top:5px;padding-bottom:5px">'''19 października 2026'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Justyna Gromada''', '''Natalia Krawczyk''' (Orange Research)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Ewaluacja agentów konwersacyjnych i modelowanie satysfakcji użytkownika w dialogach sprzedażowych'''  {{attachment:seminarium-archiwum/icon-pl.gif|Wystąpienie w jęz. polskim.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Streszczenie wystąpienia zostanie podane w najbliższym czasie.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''22 października 2026'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Nina Smirnova''' (GESIS – Leibniz Institute for the Social Sciences)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Tytuł wystąpienia udostępnimy wkrótce'''  {{attachment:seminarium-archiwum/icon-en.gif|Wystąpienie w jęz. angielskim.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Streszczenie wystąpienia zostanie podane w najbliższym czasie.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''9 listopada 2026'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Piotr Pęzik''', '''Filip Żarnecki''', '''Wojciech Janowski''', '''Jakub Kwiatkowski''', '''Paweł Wilk''', '''Łukasz Stolarski''' (Uniwersytet Łódzki)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''LLM-y w przeszukiwaniu korpusów. Przykłady zastosowań agentowych i analitycznych'''  {{attachment:seminarium-archiwum/icon-pl.gif|Wystąpienie po polsku.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Streszczenie wystąpienia zostanie podane w najbliższym czasie.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''19 listopada 2026'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Sabina Tomkins''' (University of Michigan)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Tytuł wystąpienia zostanie podany wkrótce'''  {{attachment:seminarium-archiwum/icon-en.gif|Wystąpienie w jęz. angielskim.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Streszczenie wystąpienia zostanie podane w najbliższym czasie.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''23 listopada 2026'''||<rowspan=2 style="border:0;padding-left:30px;">|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Persuasion in Parliament and Beyond: Computational Analysis of Political Discourse'''|| ||<style="border:0;padding-left:30px;padding-bottom:5px">Projekt zrealizowany przy wsparciu finansowym Polskiej Akademii Nauk w ramach programu MOST PAN  {{attachment:seminarium-archiwum/icon-en.gif|Talks in English.}}|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''!ParlaMint – Comparable and Interoperable Parliamentary Corpora''' (Tomaž Erjavec, Jožef Stefan Institute)|| ||<style="border:0;padding-left:30px;padding-bottom:10px">The talk presents the [[https://www.clarin.eu/parlamint|ParlaMint corpora of parliamentary debates]] of 29 European countries and autonomous regions, covering at least the period from 2015 to 2022, and containing over 1 billion words. The corpora are uniformly encoded, contain rich metadata about their 24 thousand speakers, and are linguistically annotated, as well as machine translated to English. We present the compilation of the corpora, including the encoding infrastructure, use of !GitHub, the production of individual corpora, the common pipeline for producing their distribution, and use of CLARIN services for dissemination. We also present associated efforts, such corpora of additional countries, and the [[https://clarinsi.github.io/parlaspeech/|ParlaSpeech]] corpora. Finally, the use of the corpora and further work are discussed.|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Beyond the Polish Parliamentary Corpus''' (Maciej Ogrodniczuk, Instytut Podstaw Informatyki PAN)|| ||<style="border:0;padding-left:30px;padding-bottom:10px">The [[https://clip.ipipan.waw.pl/PPC|Polish Parliamentary Corpus]], which is also represented in ParlaMint, served as a model for a series of Polish corpora containing similar types of material, which are ideally suited to research into persuasion techniques. The presentation will introduce these datasets: the [[https://clip.ipipan.waw.pl/PRTC|Round Table Corpus]], documenting the opposition’s negotiations with the communist authorities of the Polish People’s Republic in 1989, and the Local Government Debates Corpus – a collection of transcripts from the proceedings of provincial assemblies, city councils, county councils and local councils from 2018 to 2026.|| ||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Leveraging Persuasion and Intent for Analysis and Reasoning-based Detection of Disinformation with Large Language Models''' (Arkadiusz Modzelewski, Uniwersytet Padewski)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''From Detecting Persuasion to Predicting Its Success: Persuasion Strategies as Predictive Signals''' (Tiziano Labruna, Fondazione Bruno Kessler)|| ||<style="border:0;padding-left:30px;padding-bottom:10px">Computational work on persuasion has largely focused on detecting rhetorical techniques in text. This talk would ask a complementary question: do those techniques tell us anything about whether an argument actually changes someone's mind? Multi-Strategy Persuasion Scoring (from Labruna, EACL 2026) is a zero-shot framework in which a large language model reasons about each of six persuasion strategies independently and produces a per-strategy score, which can then be aggregated directly or used as input to a lightweight classifier. The strategy has been tested across three datasets, proving that strategy-guided reasoning consistently outperforms both direct pairwise comparison and generic chain-of-thought baselines.|| ||<style="border:0;padding-left:30px;padding-bottom:15px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''A Corpus of Persuasion Techniques in Slavic Languages''' (Jakub Piskorski, Joint Research Centre of the European Commission)|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''30 listopada 2026'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''AI without Borders'''|| ||<style="border:0;padding-left:30px;padding-bottom:5px">Projekt zrealizowany przy wsparciu finansowym Polskiej Akademii Nauk w ramach programu MOST PAN  {{attachment:seminarium-archiwum/icon-en.gif|Talks in English.}}|| ||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Universalist modelling and processing of idiomaticity in the PARSEME and !UniDive framework – recent developments and the state of Polish''' (Agata Savary, Uniwersytet Paris-Saclay)|| ||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Detecting AI generated content on surface- and idea-level''' (Marzena Karpińska, Uniwersytet Simona Frasera)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Can Neural Network Models Learn Logic?''' (Jakub Szymanik, Uniwersytet w Trydencie)|| ||<style="border:0;padding-left:30px;padding-bottom:10px">Large language models now produce chains of reasoning that look like proofs, but it remains unclear whether they have acquired the rules of inference or merely the patterns that usually accompany them. I will describe a line of joint work with Manuel Vargas Guzmán and Maciej Malicki {{attachment:seminarium-archiwum/info.png|Vargas Guzmán, M., Szymanik, J., & Malicki, M. (2024). Testing the limits of logical reasoning in neural and hybrid models. Findings of the Association for Computational Linguistics: NAACL 2024, 2267–2279. • Bertolazzi, L., Vargas Guzmán, M., Bernardi, R., Malicki, M., & Szymanik, J. (2026). Teaching small language models to learn logic through meta-learning. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL), 8049–8080. • Vargas Guzmán, M., Szymanik, J., & Malicki, M. (2026). Hybrid models for natural language reasoning: The case of syllogistic logic. Proceedings of the 23rd International Conference on Principles of Knowledge Representation and Reasoning (KR), 1143–1152.}} that tries to make this question precise enough to answer. We use syllogistic logic — a small, fully controlled fragment of natural language — as a testbed, and ask models not whether a conclusion holds but which premises are needed to derive it. This lets us separate two things usually conflated under "compositional generalization": extending a learned pattern to longer chains of reasoning, and recovering the underlying rules from complex cases to apply them to simpler ones. Models turn out to be markedly better at the first than the second, and their success depends systematically on the shape of the inference rather than its difficulty. I will end with two more encouraging results — that training on examples of the task itself substantially helps smaller models, and that even imperfect neural reasoners make effective assistants to a symbolic prover — and with what all this suggests about what "learning logic" could mean for a system that learns from data.|| ||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Faithful Explanations in the Age of Large Language Models''' (Mateusz Lango, Uniwersytet Karola w Pradze)|| ||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''AI Argues Differently: Distinctive Persuasive Patterns and Communicative Strategies of LLMs''' (Agnieszka Faleńska, Uniwersytet w Stuttgarcie)|| ||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''In Search for Universal Language Similarity Metric with Application in Multilingual AI''' (Michał Ptaszyński, Kitami Institute of Technology)|| ||<style="border:0;padding-left:30px;padding-bottom:0px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Answering Questions and Temporal Reasoning in Large Language Models''' (Adam Jatowt, Uniwersytet w Innsbrucku)|| ||<style="border:0;padding-left:30px;padding-bottom:15px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Safety Vulnerabilities in Spoken Language Models and Web Agents''' (Karolina Stańczak, Politechnika Federalna w Zurychu)|| ||<style="border:0;padding-top:15px">Zapraszamy także do zapoznania się z [[http://nlp.ipipan.waw.pl/NLP-SEMINAR/previous-p.html|archiwum seminariów z lat 2000–2015]] oraz [[http://zil.ipipan.waw.pl/seminarium-archiwum|listą wystąpień z lat 2015–2026]].|| {{{#!wiki comment ||<style="border:0;padding-top:5px;padding-bottom:5px">'''17 listopada 2025'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Marzena Karpińska''' (Microsoft) || ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''!OneRuler: testowanie wielojęzycznych modeli językowych na długim kontekście'''  {{attachment:seminarium-archiwum/icon-pl.gif|Wystąpienie w języku polskim.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">W tym wystąpieniu przyjrzymy się jak dobrze modele językowe radzą sobie z wydobywaniem informacji z tekstów do 128 tysięcy tokenów (ok 100 tysięcy słów) w 26 językach, w tym po polsku. Wyniki eksperymentów wskazują, że wraz ze wzrostem długości kontekstu rosną różnice między językami o dużych i małych zasobach danych. Co zaskakujące, nawet minimalne zmiany w poleceniu (dodanie możliwości, że informacja nie istnieje) powodują znaczny spadek skuteczności, szczególnie przy dłuższych tekstach.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''7 października 2023'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Uczestnicy konkursu PolEval 2024''' || ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Planowana seria prezentacji uczestników zadań PolEvalowych'''  {{attachment:seminarium-archiwum/icon-pl.gif|Wystąpienie w języku polskim.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Lista wystąpień będzie dostępna wkrótce.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''11 marca 2024'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Mateusz Krubiński''' (Uniwersytet Karola w Pradze)|| ||<style="border:0;padding-left:30px;padding-bottom:15px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Tytuł wystąpienia podamy wkrótce'''  {{attachment:seminarium-archiwum/icon-en.gif|Wystąpienie w języku polskim.}}|| ||<style="border:0;padding-top:15px;padding-bottom:5px">'''8 stycznia 2024''' (prezentacja wyników projektu DARIAH.Lab)|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Zespół projektu DARIAH.Lab''' (Instytut Podstaw Informatyki PAN)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">'''Tytuł wystąpienia poznamy wkrótce'''  {{attachment:seminarium-archiwum/icon-pl.gif|Wystąpienie po polsku.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Streszczenie wystąpienia udostępnimy w najbliższym czasie.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''3 października 2022'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''...''' (...)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[http://zil.ipipan.waw.pl/seminarium-online|{{attachment:seminarium-archiwum/teams.png}}]] '''Tytuł wystąpienia podamy wkrótce'''  {{attachment:seminarium-archiwum/icon-en.gif|Wystąpienie w języku polskim.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Opis wystąpienia udostępnimy już niedługo.|| WOLNE TERMINY: ATLAS: Explaining abstractive summarization - Emilia Wiśnios? Albo coś z NASK-owych tematów dot. przetwarzania prawa? Czy to jest to samo? ||<style="border:0;padding-bottom:10px">'''UWAGA''': ze względu na zakaz wstępu do IPI PAN dla osób niezatrudnionych w Instytucie, w stacjonarnej części seminarium mogą brać udział tylko pracownicy IPI PAN i prelegenci (także zewnętrzni). Dla pozostałych uczestników seminarium będzie transmitowane – prosimy o korzystanie z linku przy tytule wystąpienia.|| Uczestnicy Akcji COST CA18231: Multi3Generation: Multi-task, Multilingual, Multi-modal Language Generation: – Marcin PAPRZYCKI (marcin.paprzycki@ibspan.waw.pl) – Maria GANZHA (m.ganzha@mini.pw.edu.pl) – Katarzyna WASIELEWSKA-MICHNIEWSKA (katarzyna.wasielewska@ibspan.waw.pl) ||<style="border:0;padding-top:5px;padding-bottom:5px">'''6 czerwca 2022'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Paula Czarnowska''' (University of Cambridge)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">[[https://teams.microsoft.com/l/meetup-join/19%3a06de5a6d7ed840f0a53c26bf62c9ec18%40thread.tacv2/1643554817614?context=%7b%22Tid%22%3a%220425f1d9-16b2-41e3-a01a-0c02a63d13d6%22%2c%22Oid%22%3a%22f5f2c910-5438-48a7-b9dd-683a5c3daf1e%22%7d|{{attachment:seminarium-archiwum/teams.png}}]] '''Tytuł wystąpienia podamy wkrótce'''  {{attachment:seminarium-archiwum/icon-en.gif|Wystąpienie w języku angielskim.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Opis wystąpienia udostępnimy już niedługo.|| ||<style="border:0;padding-top:5px;padding-bottom:5px">'''2 kwietnia 2020'''|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''Stan Matwin''' (Dalhousie University)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">'''Efficient training of word embeddings with a focus on negative examples'''  {{attachment:seminarium-archiwum/icon-pl.gif|Wystąpienie w języku polskim.}} {{attachment:seminarium-archiwum/icon-en.gif|Slajdy po angielsku.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">This presentation is based on our [[https://pdfs.semanticscholar.org/1f50/db5786913b43f9668f997fc4c97d9cd18730.pdf|AAAI 2018]] and [[https://aaai.org/ojs/index.php/AAAI/article/view/4683|AAAI 2019]] papers on English word embeddings. In particular, we examine the notion of “negative examples”, the unobserved or insignificant word-context co-occurrences, in spectral methods. we provide a new formulation for the word embedding problem by proposing a new intuitive objective function that perfectly justifies the use of negative examples. With the goal of efficient learning of embeddings, we propose a kernel similarity measure for the latent space that can effectively calculate the similarities in high dimensions. Moreover, we propose an approximate alternative to our algorithm using a modified Vantage Point tree and reduce the computational complexity of the algorithm with respect to the number of words in the vocabulary. We have trained various word embedding algorithms on articles of Wikipedia with 2.3 billion tokens and show that our method outperforms the state-of-the-art in most word similarity tasks by a good margin. We will round up our discussion with some general thought s about the use of embeddings in modern NLP.|| na [[https://www.youtube.com/ipipan|kanale YouTube]]. on [[https://www.youtube.com/ipipan|YouTube]]. Nowe typy: Aleksandra Gabryszak (DFKI Berlin): – https://aclanthology.org/people/a/aleksandra-gabryszak/ – https://www.researchgate.net/profile/Aleksandra-Gabryszak – miała tekst na warsztacie First Computing Social Responsibility Workshop (http://www.lrec-conf.org/proceedings/lrec2022/workshops/CSRNLP1/index.html) na LREC-u 2022: http://www.lrec-conf.org/proceedings/lrec2022/workshops/CSRNLP1/pdf/2022.csrnlp1-1.5.pdf Marcin Junczys-Dowmunt przy okazji świąt? Adam Jatowt? Piotrek Pęzik? Wrocław? Kwantyfikatory? MARCELL? Może Piotrek z Bartkiem? Umówić się z Brylską, zapytać tę od okulografii, czy to jest PJN Agnieszka Kwiatkowska – zobaczyć ten jej tekst, moze też coś opowie? Ew. Kasia Brylska, Monika Płużyczka na seminarium? Marcin Napiórkowski z Karolem? Maciej Karpiński Demenko – dawno już ich nie było; można iść po kluczu HLT Days MTAS? – NLP dla tekstów historycznych – Marcin/Witek? razem z KORBĄ, pokazać oba ręcznie znakowane korpusy i benchmarki na tagerach – maj, – może Wrocław mógłby coś pokazać? – pisałem do Maćka P. – jakieś wystąpienia PolEvalowe? Tomek Dwojak i inni z https://zpjn.wmi.amu.edu.pl/seminar/? Będzie na Data Science Summit: Using topic modeling for differentiation based on Polish parliament plus person Aleksander Nosarzewski Statistician @ Citi Artykuł o GPT napisał Mateusz Litwin: https://www.linkedin.com/in/mateusz-litwin-06b3a919/ W OpenAI jest jeszcze https://www.linkedin.com/in/jakub-pachocki/ i https://www.linkedin.com/in/szymon-sidor-98164044/ Text data can be an invaluable source of information. In particular, what, how often and in which way we talk about given subjects can tell a lot about us. Unfortunately, manual scrambling through huge text datasets can be a cumbersome task. Luckily, there is a class of unsupervised models - topic models, which can perform this task for us, with very little input from our side. I will present how to use Structural Topic Model (STM) - an enhancement over popular LDA to obtain some kind of measure of differences between given groups or agents of interest, based on an example of Polish parliamentary speeches and political parties. ||<style="border:0;padding-top:5px;padding-bottom:5px">'''12 DATA 2017''' ('''UWAGA: ''' wystąpienie odbędzie się o 13:00 w ramach [[https://ipipan.waw.pl/instytut/dzialalnosc-naukowa/seminaria/ogolnoinstytutowe|seminarium IPI PAN]])|| ||<style="border:0;padding-left:30px;padding-bottom:0px">'''OSOBA''' (AFILIACJA)|| ||<style="border:0;padding-left:30px;padding-bottom:5px">'''Tytuł zostanie udostępniony w najbliższym czasie'''  {{attachment:seminarium-archiwum/icon-pl.gif|Wystąpienie w języku polskim.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">Opis wystąpienia zostanie udostępniony wkrótce.|| ||<style="border:0;padding-left:30px;padding-bottom:5px">'''[[attachment:seminarium-archiwum/201--.pdf|...]]'''  {{attachment:seminarium-archiwum/icon-pl.gif|Wystąpienie w języku polskim.}} {{attachment:seminarium-archiwum/icon-en.gif|Slajdy w języku angielskim.}}|| ||<style="border:0;padding-left:30px;padding-bottom:15px">...|| }}} |
Seminarium „Przetwarzanie języka naturalnego” 2026–27
Seminarium Zespołu Inżynierii Lingwistycznej w Instytucie Podstaw Informatyki Polskiej Akademii Nauk odbywa się średnio co 2 tygodnie, zwykle w poniedziałki o godz. 10:15 (niekiedy online – prosimy o korzystanie z linku przy tytule wystąpienia) i ma charakter otwarty. Poszczególne referaty ogłaszane są na Polskiej Liście Językoznawczej oraz na stronie Lingwistyka komputerowa na Facebooku. Nagrania wystąpień dostępne są na kanale YouTube. |
7 września 2026 |
Varvara Magomedova (Univerza v Novi Gorici) |
Slovenian dialect research has produced a lot of literature, but not as many data sources. The lack of available syntactically annotated corpora limits research for those, who do not have access to native speakers. This talk presents the first phase of Veje – to become a Slovenian dialect treebank. The present version of the pipeline is designed to reduce the manual work. The pipeline has two modules – normalization and annotation. In the first stage, dialectal surface forms are normalized towards standard Slovenian through a sequence of corrected lexical replacements, deterministic phonological rules, Sloleks-based morphological validation, and a constrained GaMS residual corrector. The process preserves token order and alignment wherever possible and records substitutions that cannot be morphologically verified for later review. In the second stage, normalized sentences are parsed with the non-standard Slovenian CLASSLA-Stanza model, while SloBERTa identifies whatever could not be normalized and sends it to GaMS for a second check of the annotation. The resulting extended CoNLL-U representation retains the original dialect form alongside its normalized equivalent, Universal Dependencies annotation, dialect metadata, and audit information. The current output is intentionally silver-standard: native speakers and dialectologists will correct selected data into a gold subset, which will subsequently support evaluation and model fine-tuning. I will also show the web-platform developed based on user needs study held with corpora users. |
21 września 2026 |
Mirosław Koziarski (niezależny badacz) |
Słownik języka polskiego pod redakcją Jana Karłowicza, Adama Kryńskiego i Władysława Niedźwiedzkiego, znany jako Słownik warszawski, należy do najważniejszych i najobszerniejszych dzieł polskiej leksykografii. Osiem tomów, opublikowanych w latach 1900–1927 i liczących łącznie niemal 7800 stron, dokumentuje polszczyznę wielu epok, regionów, rejestrów i dziedzin. Choć skany słownika są dostępne w bibliotekach cyfrowych, jego treść pozostawała dotąd w dużej mierze zamknięta w obrazie drukowanych stron. |
Podczas wystąpienia przedstawię projekt cyfrowej edycji Słownika warszawskiego, rozwijający metodologię i narzędzia opracowane w ramach mojej rozprawy doktorskiej. Omówię kolejne etapy przetwarzania: wybór i analizę źródła, obróbkę i segmentację obrazów, OCR — obecnie wspomagany także narzędziami opartymi na sztucznej inteligencji — korektę i normalizację tekstu, hierarchiczne parsowanie artykułów hasłowych, walidację, indeksowanie oraz generowanie danych pochodnych. Centralnym elementem procesu jest parser przekształcający liniową transkrypcję w ustrukturyzowaną reprezentację XML. Rozpoznaje on kilkadziesiąt typów segmentów, m.in. wyrazy hasłowe, warianty, informacje gramatyczne, znaczenia i podznaczenia, definicje, kwalifikatory, przykłady użycia, etymologie, odsyłacze i związki frazeologiczne, a także relacje między nimi. Ze względu na złożoność i niekonsekwencje materiału proces ma charakter iteracyjny i łączy metody automatyczne z ręczną kontrolą. |
Uruchomiona edycja internetowa stanowi rozwinięcie wcześniejszej makiety badawczej w stopniowo rozbudowywaną platformę. Użytkownik może zestawić sformatowany artykuł z transkrypcją źródłową, surowym XML-em i odpowiednim fragmentem faksymile. Ustrukturyzowanie treści umożliwia ponadto tworzenie nowych ścieżek dostępu do danych: korpusu przykładów, zestawień derywatów, skonsolidowanego indeksu skróceń, statystyk oraz matrycy porównującej siatkę hasłową z innymi słownikami polszczyzny. |
Na zakończenie przedstawię aktualny stan projektu, przyjęte mechanizmy kontroli jakości i ograniczenia automatycznego przetwarzania historycznego materiału leksykograficznego, a także planowane kierunki rozwoju. |
22 października 2026 |
Nina Smirnova (GESIS – Leibniz Institute for the Social Sciences) |
Streszczenie wystąpienia zostanie podane w najbliższym czasie. |
19 listopada 2026 |
Sabina Tomkins (University of Michigan) |
Streszczenie wystąpienia zostanie podane w najbliższym czasie. |
23 listopada 2026 |
|
Persuasion in Parliament and Beyond: Computational Analysis of Political Discourse |
|
Projekt zrealizowany przy wsparciu finansowym Polskiej Akademii Nauk w ramach programu MOST PAN |
|
|
|
The talk presents the ParlaMint corpora of parliamentary debates of 29 European countries and autonomous regions, covering at least the period from 2015 to 2022, and containing over 1 billion words. The corpora are uniformly encoded, contain rich metadata about their 24 thousand speakers, and are linguistically annotated, as well as machine translated to English. We present the compilation of the corpora, including the encoding infrastructure, use of GitHub, the production of individual corpora, the common pipeline for producing their distribution, and use of CLARIN services for dissemination. We also present associated efforts, such corpora of additional countries, and the ParlaSpeech corpora. Finally, the use of the corpora and further work are discussed. |
|
|
|
The Polish Parliamentary Corpus, which is also represented in ParlaMint, served as a model for a series of Polish corpora containing similar types of material, which are ideally suited to research into persuasion techniques. The presentation will introduce these datasets: the Round Table Corpus, documenting the opposition’s negotiations with the communist authorities of the Polish People’s Republic in 1989, and the Local Government Debates Corpus – a collection of transcripts from the proceedings of provincial assemblies, city councils, county councils and local councils from 2018 to 2026. |
|
|
|
|
|
Computational work on persuasion has largely focused on detecting rhetorical techniques in text. This talk would ask a complementary question: do those techniques tell us anything about whether an argument actually changes someone's mind? Multi-Strategy Persuasion Scoring (from Labruna, EACL 2026) is a zero-shot framework in which a large language model reasons about each of six persuasion strategies independently and produces a per-strategy score, which can then be aggregated directly or used as input to a lightweight classifier. The strategy has been tested across three datasets, proving that strategy-guided reasoning consistently outperforms both direct pairwise comparison and generic chain-of-thought baselines. |
|
|
Zapraszamy także do zapoznania się z archiwum seminariów z lat 2000–2015 oraz listą wystąpień z lat 2015–2026. |




that tries to make this question precise enough to answer. We use syllogistic logic — a small, fully controlled fragment of natural language — as a testbed, and ask models not whether a conclusion holds but which premises are needed to derive it. This lets us separate two things usually conflated under "compositional generalization": extending a learned pattern to longer chains of reasoning, and recovering the underlying rules from complex cases to apply them to simpler ones. Models turn out to be markedly better at the first than the second, and their success depends systematically on the shape of the inference rather than its difficulty. I will end with two more encouraging results — that training on examples of the task itself substantially helps smaller models, and that even imperfect neural reasoners make effective assistants to a symbolic prover — and with what all this suggests about what "learning logic" could mean for a system that learns from data.