Musixmatch

Musixmatch is the World’s Leading Music Data Company.
Our mission is to provide data, tools and services that allows the experience of music to be enriched across the whole world.

Website · About · Developer API · GitHub · Careers


AI research

Musixmatch best-known open contribution is UmBERTo, a RoBERTa-based Italian language model trained with SentencePiece tokenization and Whole Word Masking on 70GB of deduplicated Italian text from the OSCAR corpus (~11B words). It set a new bar for Italian NLP and remains a widely used baseline.

Model Task Score
UmBERTo (CommonCrawl, cased) NER — WikiNER-ITA 92.5 F1
UmBERTo (CommonCrawl, cased) NER — ICAB-EvalITA07 87.6 F1
UmBERTo (CommonCrawl, cased) POS — UD Italian-ISDT 98.9 acc

Musixmatch/umberto-commoncrawl-cased-v1 · Musixmatch/umberto-wikipedia-uncased-v1 · source on GitHub

We also run our training and inference on AWS — read the Musixmatch × Hugging Face × AWS case study.

© Musixmatch · Bologna, Italy