This is the way: designing and compiling LEPISZCZE, a comprehensive NLP benchmark for Polish

Łukasz Augustyniak; Kamil Tagowski; Albert Sawczyn; Denis Janiak; Roman Bartusiak; Adrian Szymczak; Marcin Wątroba; Arkadiusz Janz; Piotr Szymański; Mikołaj Morzy; Tomasz Kajdanowicz; Maciej Piasecki

System Informacji Naukowej Politechniki Poznańskiej

PL EN

Strona główna / Publikacje / This is the way: designing and compiling LEPISZCZE, a comprehensive NLP benchmark for Polish

Zgłoś uwagę

Rozdział

Pobierz BibTeX

Tytuł

This is the way: designing and compiling LEPISZCZE, a comprehensive NLP benchmark for Polish

Autorzy

Łukasz Augustyniak
Kamil Tagowski
Albert Sawczyn
Denis Janiak
Roman Bartusiak
Adrian Szymczak
Marcin Wątroba
Arkadiusz Janz
Piotr Szymański
Mikołaj Morzy (WIiT) ^{[ 1 ][ 2.3 ][ P ]}
Tomasz Kajdanowicz
Maciej Piasecki

^{[ 1 ]} Instytut Informatyki, Wydział Informatyki i Telekomunikacji, Politechnika Poznańska | ^{[ P ]} pracownik

Dyscyplina naukowa (Ustawa 2.0)

[2.3] Informatyka techniczna i telekomunikacja

Rok publikacji

2022

Typ rozdziału

rozdział w monografii naukowej / referat

Język publikacji

angielski

Słowa kluczowe

EN

nlp
benchmark
machine learning
Polish language

Streszczenie

EN The availability of compute and data to train larger and larger language models increases the demand for robust methods of benchmarking the true progress of LM training. Recent years witnessed significant progress in standardized benchmarking for English. Benchmarks such as GLUE, SuperGLUE, or KILT have become a defacto standard tools to compare large language models. Following the trend to replicate GLUE for other languages, the KLEJ benchmark1 has been released for Polish. In this paper, we evaluate the progress in benchmarking for low-resourced languages. We note that only a handful of languages have such comprehensive benchmarks. We also note the gap in the number of tasks being evaluated by benchmarks for resource-rich English/Chinese and the rest of the world. In this paper, we introduce LEPISZCZE, a new, comprehensive benchmark for Polish NLP with a large variety of tasks and high-quality operationalization of the benchmark. We design LEPISZCZE with flexibility in mind. Including new models, datasets, and tasks is as simple as possible while still offering data versioning and model tracking. In the first run of the benchmark, we test 13 experiments (task and dataset pairs) based on the five most recent LMs for Polish. We use five datasets from the Polish benchmark and add eight novel datasets. As the paper’s main contribution, apart from LEPISZCZE , we provide insights and experiences learned while creating the benchmark for Polish as the blueprint to design similar benchmarks for other low-resourced languages.

URL

https://proceedings.neurips.cc/paper_files/paper/2022/file/890b206ebb79e550f3988cb8db936f42-Paper-Datasets_and_Benchmarks.pdf

Książka

Advances in Neural Information Processing Systems 35 (NeurIPS 2022)

Zaprezentowany na

36th Conference on Neural Information Processing Systems (NeurIPS 2022), 29.11.2022 - 01.12.2023, New Orleans, United States

Tryb otwartego dostępu

witryna wydawcy

Wersja tekstu w otwartym dostępie

ostateczna wersja opublikowana

Czas udostępnienia publikacji w sposób otwarty

w momencie opublikowania