The Invalsi Benchmark: measuring Language Models Mathematical and Language understanding in Italian

Esuli, Andrea; Puccetti, Giovanni

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 2403

Change to browse by:

Computer Science > Computation and Language

Title: The Invalsi Benchmark: measuring Language Models Mathematical and Language understanding in Italian

Authors: Andrea Esuli, Giovanni Puccetti

(Submitted on 27 Mar 2024)

Abstract: While Italian is by all metrics a high resource language, currently, there are isn't a Language Model pre-trained exclusively in this language. This results in a lower number of available benchmarks to evaluate the performance of language models in Italian.
This work presents two new benchmarks to evaluate the models performance on mathematical understanding and language understanding in Italian. These benchmarks are based on real tests that are undertaken by students of age between 11 and 18 within the Italian school system and have therefore been validated by several experts in didactics and pedagogy.
To validate this dataset we evaluate the performance of 9 language models that are the best performing when writing in Italian, including our own fine-tuned models. We show that this is a challenging benchmark where current language models are bound by 60\% accuracy.
We believe that the release of this dataset paves the way for improving future models mathematical and language understanding in Italian.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2403.18697 [cs.CL]
	(or arXiv:2403.18697v1 [cs.CL] for this version)

Submission history

From: Giovanni Puccetti [view email]
[v1] Wed, 27 Mar 2024 15:46:25 GMT (103kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2403.18697

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: The Invalsi Benchmark: measuring Language Models Mathematical and Language understanding in Italian

Submission history