Text classifiers in 51 languages that run on your CPU

2 pointsposted 8 hours ago
by nico

1 Comments

nico

8 hours ago

I built this demo after someone on HN mentioned wanting local classifiers in Polish, to use as part of their stack for a personal assitant

The demo showcases 51 classifiers (the languages from Amazon MASSIVE), all trained on the same set of 60 tasks

When you submit a text, it classifies by first getting the embeddings for the text using a multilingual sentence encoder (paraphrase-multilingual-MiniLM-L12-v2), then running the code through the logistic regression classifier for the corresponding language. The output are a set of %s of how representative the classifier thinks the text is to each class/task, and highlights the highest % one

Training all 60 tasks in 51 languages took 22 minutes of CPU time on a laptop

Accuracy ranges from 60% (Javanese, Filipino) to 86% (English), averaging 76%. Not state-of-the-art, but the classifiers are 10,000× smaller than fine-tuned alternatives and run anywhere without a GPU

Note: the demo is running on a small server, you'll get better performance (ie. lower latency on the requests) if you run it locally (pip install jeffy-classify / GitHub: github.com/nicobrenner/jeffy)