Detecting Text Language With Python and NLTK
Most of us are used to Internet search engines and social networks capabilities to show only data in certain language, for example, showing only results written in Spanish or English. To achieve that, indexed text must have been analized previously to “guess” the languange and store it together.
There are several ways to do that; probably the most easy to do is a stopwords based approach. The term “stopword” is used in natural language processing to refer words which should be filtered out from text before doing any kind of processing, commonly because this words are little or nothing usefult at all when analyzing text.
How to do that?
Ok, so we have a text whose language we want to detect depending on stopwords being used in such text. First step is to “tokenize” — convert given text to a list of “words” or “tokens” — using an approach or another depending on our requeriments: should we keep contractions or, otherwise, should we split them? we need puntuactions or want to split them off? and so on.
In this case we are going to split all punctuations into separate tokens:
nltk “wordpunct_tokenize” tokenizer
As shown, the famous quote from Mr. Wolf has been splitted and now we have “clean” words to match against stopwords list.
At this point we need stopwords for several languages and here is when NLTK comes to handy:
included languages in NLTK
Now we need to compute language probability depending on which stopwords are used:
calculate languages ratios
First we tokenize using wordpunct_tokenize function and lowercase all splitted tokens, then we walk across nltk included languages and count how many unique stopwords are seen in analyzed text to put this in “language_ratios” dictionary.
Finally, we only have to get the “key” with biggest “value”:
get most rated language
So yes, it seems this approach works fine with well written texts — those who respect grammatical rules — (and not so small ones) and is really easy to implement.
Putting it all together
If we put all the explained above into a script we have something like this:
There are others ways to “guess” language from a given text like N-Gram-Based text categorization so will see it in, probably, next post.
1. TextBlob.
Note: This solution requires internet access and Textblob is using Google Translate’s language detector by calling the API.
2. Polyglot.
Requires numpy and some arcane libraries, unlikely to get it work for Windows. (For Windows, get an appropriate versions of PyICU, Morfessor and PyCLD2 from here, then just pip install downloaded_wheel.whl .) Able to detect texts with mixed languages.
pip install polyglot
To install the dependencies, run: sudo apt-get install python-numpy libicu-dev
3. chardet
Chardet has also a feature of detecting languages if there are character bytes in range (127-255]:
pip install chardet
4. langdetect
Requires large portions of text. It uses non-deterministic approach under the hood. That means you get different results for the same text sample. Docs say you have to use following code to make it determined:
pip install langdetect
5. guess_language
Can detect very short samples by using this spell checker with dictionaries.
pip install guess_language-spirit
6. langid
langid.py provides both a module
and a command-line tool:
pip install langid
7. FastText
FastText is a text classifier, can be used to recognize 176 languages with a proper models for language classification. Download this model, then:
pip install fasttext
8. pyCLD3
pycld3 is a neural network model for language identification. This package contains the inference code and a trained model.
pip install pycld3
Have you had a look at langdetect?
![]()
And @toto_tico did a nice job in presenting the speed comparison.
Here’s a summary to complete the great answers above (as of 2021)
| Language ID software | Used by | Open Source / Model | Rule-based | Stats-based | Can train/tune |
|---|---|---|---|---|---|
| Google Translate Language Detection | TextBlob (limited usage) | ✕ | — | — | ✕ |
| Chardet | — | ✓ | ✓ | ✕ | ✕ |
| Guess Language (non-active development) | spirit-guess (updated rewrite) | ✓ | ✓ | Minimally | ✕ |
| pyCLD2 | Polyglot | ✓ | Somewhat | ✓ | Not sure |
| CLD3 | — | ✓ | ✕ | ✓ | Possibly |
| langid-py | — | ✓ | Not sure | ✓ | ✓ |
| langdetect | SpaCy-langdetect | ✓ | ✕ | ✓ | ✓ |
| FastText | What The Lang | ✓ | ✕ | ✓ | Not sure |
![]()
If you are looking for a library that is fast with long texts, polyglot and fastext are doing the best job here.
I sampled 10000 documents from a collection of dirty and random HTMLs, and here are the results:
I have noticed that a lot of the methods focus on short texts, probably because it is the hard problem to solve: if you have a lot of text, it is really easy to detect languages (e.g. one could just use a dictionary!). However, this makes it difficult to find for an easy and suitable method for long texts.
![]()
There is an issue with langdetect when it is being used for parallelization and it fails. But spacy_langdetect is a wrapper for that and you can use it for that purpose. You can use the following snippet as well:
![]()
You can use Googletrans (unofficial) a free and unlimited Google translate API for Python.
You can make as many requests as you want, there are no limits
Installation:
Language detection:
![]()
Pretrained Fast Text Model Worked Best For My Similar Needs
I arrived at your question with a very similar need. I found the most help from Rabash’s answers for my specific needs.
After experimenting to find what worked best among his recommendations, which was making sure that text files were in English in 60,000+ text files, I found that fasttext was an excellent tool for such a task.
With a little work, I had a tool that worked very fast over many files. But it could be easily modified for something like your case, because fasttext works over a list of lines easily.
My code with comments is among the answers on THIS post. I believe that you and others can easily modify this code for other specific needs.
Depending on the case, you might be interested in using one of the following methods:
Method 0: Use an API or library
Usually, there are a few problems with these libraries because some of them are not accurate for small texts, some languages are missing, are slow, require internet connection, are non-free. But generally speaking, they will suit most needs.
Method 1: Language models
A language model gives us the probability of a sequence of words. This is important because it allows us to robustly detect the language of a text, even when the text contains words in other languages (e.g.: «‘Hola’ means ‘hello’ in spanish»).
You can use N language models (one per language), to score your text. The detected language will be the language of the model that gave you the highest score.
If you want to build a simple language model for this, I’d go for 1-grams. To do this, you only need to count the number of times each word from a big text (e.g. Wikipedia Corpus in «X» language) has appeared.
Then, the probability of a word will be its frequency divided by the total number of words analyzed (sum of all frequencies).
If the text to detect is quite big, I recommend sampling N random words and then use the sum of logarithms instead of multiplications to avoid floating-point precision problems.
Method 2: Intersecting sets
An even simpler approach is to prepare N sets (one per language) with the top M most frequent words. Then intersect your text with each set. The set with the highest number of intersections will be your detected language.
Method 3: Zip compression
This more a curiosity than anything else, but here it goes. You can compress your text (e.g LZ77) and then measure the zip-distance with regards to a reference compressed text (target language). Personally, I didn’t like it because it’s slower, less accurate and less descriptive than other methods. Nevertheless, there might be interesting applications for this method. To read more: Language Trees and Zipping
4 Python libraries to detect English and Non-English language
We will discuss spacy-langdetect, Pycld2, TextBlob, and Googletrans for language detection
Are you sure that the input text data for your model is in English? Well, no one can be sure about this, as no one will read around 20k records of text data.
So, how non-English text will affect your English text trained model?
Pick any non-English text and pass it through as input to your English text trained classification model. You will come to know that the category is assigned to non-English text by the model.
If your model is dependent on one language then, other languages in your textual data should be considered as noise.
The job of the text classification model is to classify. And, it will do its job despite its input text will be in English or not.
What can we do to avoid such a situation?
Your model will not stop classifying the non-English text. So, you have to detect the non-English text and remove it from trained data and prediction data.
This process comes under the data cleaning part. Inconsistency in your data will result in a decrease in the accuracy of the model. Sometimes, multiple languages present in text data could be one of the reasons your model behaves strangely.
So, in this article, we will discuss the different python libraries which detect the language(s) of the text data.
Let’s start with the spaCy library.
1. SpaCy
You need to install the spacy-langdetect and spacy python libraries for the below code to work.
#1. Download the best-matching default model and create a shortcut link.
#2. Add LanguageDetector() function and model to NLP pipeline.
#3. Pass the text data into the pipeline for language detection.
#4 Store the detected language and accuracy in the detect_language variable.
We have tested the above library with only single language text data. What will happen when the text has multiple language sentences?
Predicted language is en, that is, English. Prediction accuracy is also low. But, this text also has a German-language sentence. And, this library did not predict that.
This library returns the detected language of longer sentences.
The spaCy is a python library used in different Natural Language Processing (NLP) tasks. It’s deploying a spacy_langdetect library model in the spaCy NLP pipeline.
2. Pycld2
What if we want to know the detected language of all sentences?
In that case, use the below code. You need to install the pycld2 python library for the below code to work.
Pycld2 python library is a python binding for the Compact Language Detect 2 (CLD2). You can explore the different functionality of Pycld2. Know about the Pycld2 here.
3. TextBlob
Let’s say you are using TextBlob for the NLP task. Then will you use spaCy or TextBlob for the language detection? I will use TextBlob and also should you.
And we also love choices.
You need to install the textblob python library for the below code to work.
TextBlob also returns the detected language of longer sentences. Use pycld2 library for multiple language sentences.
Do you know which language is represented by the “ru”?
If you don’t then, then visit this Wikipedia link. Or you can check the below list. It is the standard short-form (ISO 639–1 code) of the language used in data science.
TextBlob provides an API to perform different NLP tasks. Some applications of TextBlob are text processing, sentiment analysis, classification, spelling correction, keyword extraction, part-of-speech tagging, etc. know about TextBlob here.
4. Googletrans
Googletrans python library uses the google translate API to detect the language of text data. But this library is not reliable. So, be careful before using this library. You can consider this library as another choice for language detection.
You need to install the googletrans python library for the below code to work.
Googletrans also has the functionality for language translation. Know about Googletrans here.
Application of language detection
- Find out bias in text data based on the languages.
- You can classify the article based on the different languages.
- Language is generally associated with the region. This method helps you to classify the article based on languages.
- You can use this method in the language translation model.
- You can use it in data cleaning and data manipulation processes.
Conclusion
We should consider language detection as one of the data cleaning processes, for textual data. Internet text data are not always present in the English language.
In this article, I have explains different python libraries to detect the languages.
These libraries will help you to remove the noise from your data. It’s better to know that your data is noise-free. And you have eliminated one reason if your model doesn’t perform on new data.
Name already in use
If nothing happens, download GitHub Desktop and try again.
Launching GitHub Desktop
If nothing happens, download GitHub Desktop and try again.
Launching Xcode
If nothing happens, download Xcode and try again.
Launching Visual Studio Code
Your codespace will open once ready.
There was a problem preparing your codespace, please try again.
Latest commit
Git stats
Files
Failed to load latest commit information.
README.md
1. What does this library do?
Its task is simple: It tells you which language some text is written in. This is very useful as a preprocessing step for linguistic data in natural language processing applications such as text classification and spell checking. Other use cases, for instance, might include routing e-mails to the right geographically located customer service department, based on the e-mails’ languages.
2. Why does this library exist?
Language detection is often done as part of large machine learning frameworks or natural language processing applications. In cases where you don’t need the full-fledged functionality of those systems or don’t want to learn the ropes of those, a small flexible library comes in handy.
Python is widely used in natural language processing, so there are a couple of comprehensive open source libraries for this task, such as Google’s CLD 2 and CLD 3, langid, fastText and langdetect. Unfortunately, except for the last one they have two major drawbacks:
- Detection only works with quite lengthy text fragments. For very short text snippets such as Twitter messages, they do not provide adequate results.
- The more languages take part in the decision process, the less accurate are the detection results.
Lingua aims at eliminating these problems. She nearly does not need any configuration and yields pretty accurate results on both long and short text, even on single words and phrases. She draws on both rule-based and statistical methods but does not use any dictionaries of words. She does not need a connection to any external API or service either. Once the library has been downloaded, it can be used completely offline.
3. Which languages are supported?
Compared to other language detection libraries, Lingua’s focus is on quality over quantity, that is, getting detection right for a small set of languages first before adding new ones. Currently, the following 75 languages are supported:
- A
- Afrikaans
- Albanian
- Arabic
- Armenian
- Azerbaijani
- Basque
- Belarusian
- Bengali
- Norwegian Bokmal
- Bosnian
- Bulgarian
- Catalan
- Chinese
- Croatian
- Czech
- Danish
- Dutch
- English
- Esperanto
- Estonian
- Finnish
- French
- Ganda
- Georgian
- German
- Greek
- Gujarati
- Hebrew
- Hindi
- Hungarian
- Icelandic
- Indonesian
- Irish
- Italian
- Japanese
- Kazakh
- Korean
- Latin
- Latvian
- Lithuanian
- Macedonian
- Malay
- Maori
- Marathi
- Mongolian
- Norwegian Nynorsk
- Persian
- Polish
- Portuguese
- Punjabi
- Romanian
- Russian
- Serbian
- Shona
- Slovak
- Slovene
- Somali
- Sotho
- Spanish
- Swahili
- Swedish
- Tagalog
- Tamil
- Telugu
- Thai
- Tsonga
- Tswana
- Turkish
- Ukrainian
- Urdu
- Vietnamese
- Welsh
- Xhosa
- Yoruba
- Zulu
4. How good is it?
Lingua is able to report accuracy statistics for some bundled test data available for each supported language. The test data for each language is split into three parts:
- a list of single words with a minimum length of 5 characters
- a list of word pairs with a minimum length of 10 characters
- a list of complete grammatical sentences of various lengths
Both the language models and the test data have been created from separate documents of the Wortschatz corpora offered by Leipzig University, Germany. Data crawled from various news websites have been used for training, each corpus comprising one million sentences. For testing, corpora made of arbitrarily chosen websites have been used, each comprising ten thousand sentences. From each test corpus, a random unsorted subset of 1000 single words, 1000 word pairs and 1000 sentences has been extracted, respectively.
Given the generated test data, I have compared the detection results of Lingua, fastText, langdetect, langid, CLD 2 and CLD 3 running over the data of Lingua’s supported 75 languages. Languages that are not supported by the other detectors are simply ignored for them during the detection process.
Each of the following sections contains two plots. The bar plot shows the detailed accuracy results for each supported language. The box plot illustrates the distributions of the accuracy values for each classifier. The boxes themselves represent the areas which the middle 50 % of data lie within. Within the colored boxes, the horizontal lines mark the median of the distributions.
4.1 Single word detection


Bar plot
4.2 Word pair detection


Bar plot
4.3 Sentence detection


Bar plot
4.4 Average detection


Bar plot
4.5 Mean, median and standard deviation
The table below shows detailed statistics for each language and classifier including mean, median and standard deviation.
Open table
Language Average Single Words Word Pairs Sentences Lingua
(high accuracy mode)Lingua
(low accuracy mode)Langdetect FastText Langid CLD3 CLD2 Simplemma Lingua
(high accuracy mode)Lingua
(low accuracy mode)Langdetect FastText Langid CLD3 CLD2 Simplemma Lingua
(high accuracy mode)Lingua
(low accuracy mode)Langdetect FastText Langid CLD3 CLD2 Simplemma Lingua
(high accuracy mode)Lingua
(low accuracy mode)Langdetect FastText Langid CLD3 CLD2 Simplemma Afrikaans
78
64
67
36
30
55
55
—
57
38
38
11
1
22
13
—
80
62
65
23
10
46
56
—
96
93
98
74
80
98
96
—Albanian
88
80
79
66
65
55
65
21
68
55
53
35
33
18
18
23
95
87
84
66
63
48
77
16
100
99
100
98
98
98
99
23Arabic
99
95
97
96
91
90
67
—
97
89
93
89
84
79
19
—
99
96
98
98
90
92
82
—
100
99
100
100
98
100
99
—Armenian
100
100
—
100
94
99
100
15
100
100
—
100
83
100
100
24
100
100
—
100
99
100
100
13
100
100
—
100
100
97
100
7Azerbaijani
89
82
—
78
68
81
72
—
78
71
—
57
36
62
34
—
92
78
—
80
69
82
82
—
99
96
—
98
98
99
99
—Basque
83
74
—
71
52
62
61
—
71
56
—
44
18
33
23
—
87
76
—
70
52
62
69
—
92
91
—
100
86
92
91
—Belarusian
97
92
—
85
85
84
76
—
91
80
—
69
69
67
42
—
99
95
—
88
87
86
87
—
100
100
—
98
99
100
99
—Bengali
100
100
100
98
92
99
63
—
100
100
100
94
92
98
19
—
100
100
100
99
88
99
69
—
100
100
100
100
97
99
99
—Bokmal
58
49
—
—
13
—
—
47
38
27
—
—
3
—
—
15
59
47
—
—
12
—
—
39
75
74
—
—
23
—
—
86Bosnian
33
29
—
9
5
33
19
—
28
22
—
9
2
19
4
—
32
29
—
10
4
28
15
—
39
36
—
8
8
52
36
—Bulgarian
87
78
72
78
67
70
66
69
71
57
51
56
46
45
32
46
91
81
68
81
62
66
72
70
99
97
96
99
93
98
93
90Catalan
70
58
54
57
38
48
38
62
50
33
25
33
5
19
4
39
74
60
51
57
29
42
30
63
86
81
86
83
81
84
79
82Chinese
100
100
64
71
96
92
33
—
100
100
39
46
90
92
—
—
100
100
55
68
97
83
2
—
100
100
97
100
100
100
98
—Croatian
72
59
73
47
48
42
51
—
53
36
50
28
16
26
34
—
74
57
70
42
38
42
47
—
90
85
97
72
90
58
73
—Czech
80
71
71
76
66
64
74
52
65
54
51
58
44
39
50
35
84
72
72
79
69
65
80
45
90
87
88
92
86
88
91
75Danish
81
70
70
62
60
58
59
56
61
45
50
35
33
26
27
27
84
70
68
57
61
54
56
52
98
95
93
95
86
95
94
89Dutch
77
64
58
78
64
58
47
57
55
36
27
55
34
29
11
34
81
61
49
81
61
47
42
46
96
94
98
100
98
97
90
90English
81
62
60
96
85
54
56
69
55
29
23
90
84
22
12
38
89
62
59
98
71
44
55
71
99
96
99
100
99
97
100
98Esperanto
83
66
—
76
44
57
50
—
67
44
—
51
5
22
7
—
85
61
—
79
30
51
46
—
98
92
—
100
96
98
98
—Estonian
92
83
82
73
67
70
65
72
80
62
62
50
37
41
24
52
96
88
86
73
67
69
73
69
100
99
100
96
98
99
99
94Finnish
96
91
93
92
83
80
77
81
90
77
84
82
62
58
44
63
98
95
95
96
88
84
89
82
100
100
100
100
100
99
98
99French
89
77
75
83
71
55
46
70
74
52
48
62
42
22
12
46
95
83
78
86
74
49
48
69
99
97
99
99
98
94
80
94Ganda
92
84
—
—
—
—
61
—
80
65
—
—
—
—
23
—
95
87
—
—
—
—
62
—
100
100
—
—
—
—
99
—Georgian
100
100
—
99
99
98
100
4
100
100
—
97
97
99
100
11
100
100
—
99
100
100
100
2
100
100
—
100
100
96
100
0German
89
80
73
89
81
66
64
83
74
57
50
76
61
40
27
64
94
84
70
93
81
62
66
86
100
99
100
100
100
98
98
99Greek
100
100
100
99
100
100
100
75
100
100
100
98
100
100
100
71
100
100
100
100
100
100
100
61
100
100
100
100
100
100
100
93Gujarati
100
100
100
100
100
100
100
—
100
100
100
99
100
99
100
—
100
100
100
100
100
100
100
—
100
100
100
100
100
100
100
—Hebrew
100
100
100
100
100
—
—
—
100
100
100
99
100
—
—
—
100
100
100
100
100
—
—
—
100
100
100
100
100
—
—
—Hindi
73
33
67
87
60
58
77
5
61
11
44
74
41
34
56
2
64
20
59
88
47
45
76
4
93
67
99
99
92
95
99
11Hungarian
95
90
88
92
83
76
75
65
87
77
74
80
64
53
41
50
98
94
91
96
86
76
85
61
100
100
100
100
100
99
100
84Icelandic
93
88
—
65
66
71
66
67
83
72
—
39
33
42
26
50
97
92
—
57
66
70
73
60
100
99
—
98
99
99
99
92Indonesian
62
47
80
69
51
46
62
34
41
25
56
43
16
26
36
39
62
45
85
68
54
45
63
30
82
71
100
95
82
66
88
32Irish
91
85
—
60
63
67
66
80
82
70
—
35
28
42
29
74
94
90
—
57
64
66
78
76
96
95
—
89
97
94
92
90Italian
87
71
76
89
66
62
44
66
69
42
50
74
28
31
7
42
92
74
80
92
70
57
32
61
100
98
99
100
100
98
93
94Japanese
100
100
100
87
86
98
33
—
100
100
99
72
61
97
—
—
100
100
100
89
96
96
—
—
100
100
100
100
100
100
100
—Kazakh
92
90
—
88
80
82
77
—
80
78
—
72
67
62
43
—
97
93
—
90
78
83
88
—
99
99
—
100
96
99
99
—Korean
100
100
100
99
100
99
100
—
100
100
100
98
100
100
100
—
100
100
100
100
100
100
100
—
100
100
100
100
100
98
100
—Latin
87
73
—
50
21
62
46
70
72
49
—
24
—
44
9
50
92
76
—
41
2
58
42
66
97
93
—
85
61
83
88
93Latvian
93
87
89
82
83
75
72
43
85
75
76
65
64
51
33
37
97
90
91
83
86
77
84
32
99
97
99
97
98
98
98
59Lithuanian
95
88
87
81
80
72
70
72
87
76
71
61
58
42
30
65
98
89
91
83
85
75
82
64
100
98
100
99
99
99
99
86Macedonian
83
72
86
74
51
60
60
14
64
52
71
51
15
30
27
15
86
70
88
72
44
54
70
11
99
95
100
100
94
97
84
16Malay
30
31
—
15
11
22
18
13
25
22
—
14
2
11
9
3
36
36
—
19
9
22
22
10
28
35
—
12
22
34
23
26Maori
92
83
—
—
—
52
61
—
84
64
—
—
—
22
12
—
92
88
—
—
—
43
72
—
99
98
—
—
—
91
98
—Marathi
85
41
88
80
80
84
83
—
74
20
76
61
70
69
65
—
84
30
90
81
79
84
86
—
96
72
98
99
91
98
99
—Mongolian
97
95
—
81
86
83
78
—
92
89
—
59
68
63
43
—
99
98
—
86
90
87
92
—
99
99
—
98
99
99
100
—Nynorsk
65
52
—
29
32
—
54
30
41
25
—
8
5
—
18
9
65
49
—
18
16
—
50
24
90
81
—
61
75
—
93
55Persian
90
80
81
90
92
76
61
12
77
62
64
79
83
57
13
12
93
80
80
92
94
70
72
5
100
98
99
100
100
99
99
18Polish
95
90
89
92
89
77
75
90
86
77
75
80
73
51
38
84
98
93
93
97
93
80
87
88
100
99
100
100
100
99
99
99Portuguese
81
69
61
73
54
53
54
65
58
42
30
47
19
21
20
36
85
69
55
71
44
40
48
62
98
96
98
99
98
97
94
97Punjabi
100
100
100
100
100
100
100
—
100
100
100
99
100
99
100
—
100
100
100
100
100
100
100
—
100
100
100
100
100
100
100
—Romanian
87
72
77
64
61
53
54
61
69
49
56
38
31
24
11
44
92
74
78
60
60
48
53
53
99
94
97
95
92
88
96
86Russian
89
79
84
94
75
71
60
71
76
60
70
86
60
48
26
64
94
84
87
98
75
72
68
66
98
92
96
100
91
93
87
84Serbian
88
78
—
76
64
78
69
—
74
62
—
54
39
63
29
—
90
80
—
76
63
75
78
—
99
91
—
98
89
95
99
—Shona
90
81
—
—
—
76
65
—
77
56
—
—
—
51
24
—
95
86
—
—
—
79
71
—
100
100
—
—
—
99
99
—Slovak
84
75
75
65
68
63
71
71
64
49
51
41
40
32
38
54
90
79
76
62
66
61
76
68
99
98
98
91
97
96
99
92Slovene
82
67
73
59
63
63
48
78
61
39
47
32
33
29
8
62
87
68
72
54
61
60
42
76
98
93
98
90
95
99
92
96Somali
92
85
90
24
—
69
70
—
82
64
75
4
—
38
27
—
95
91
94
15
—
70
83
—
100
100
100
52
—
100
99
—Sotho
83
72
—
—
—
49
54
—
61
43
—
—
—
15
13
—
88
75
—
—
—
33
54
—
99
97
—
—
—
98
95
—Spanish
70
56
57
74
65
48
43
53
44
25
26
51
37
16
12
24
69
49
47
72
59
32
34
42
97
94
97
100
98
96
85
92Swahili
79
70
73
41
42
57
57
52
58
43
47
7
3
25
16
36
82
69
73
24
24
49
59
44
98
97
99
92
98
98
97
76Swedish
84
72
68
76
65
61
53
66
64
46
40
51
35
30
14
43
89
76
67
78
63
56
52
65
99
95
96
98
96
96
93
89Tagalog
78
66
76
45
42
—
50
14
52
36
50
11
2
—
9
16
82
66
77
28
26
—
44
12
98
96
99
98
98
—
95
15Tamil
100
100
100
100
100
100
100
—
100
100
100
100
100
100
100
—
100
100
100
100
100
100
100
—
100
100
100
100
100
99
100
—Telugu
100
100
100
100
100
99
100
—
100
100
100
100
100
99
100
—
100
100
100
100
100
100
100
—
100
100
100
100
100
99
100
—Thai
99
99
100
100
100
99
100
—
100
100
100
100
100
100
100
—
100
100
100
100
100
100
100
—
98
98
100
100
100
98
100
—Tsonga
83
72
—
—
—
—
61
—
63
46
—
—
—
—
19
—
88
73
—
—
—
—
68
—
98
97
—
—
—
—
97
—Tswana
82
71
—
—
—
—
56
—
62
44
—
—
—
—
17
—
86
73
—
—
—
—
57
—
98
96
—
—
—
—
94
—Turkish
94
87
82
86
67
69
66
82
84
71
63
70
50
41
30
71
98
91
84
88
67
70
71
80
100
99
100
100
84
97
97
94Ukrainian
92
86
83
91
76
81
77
75
84
75
66
78
54
62
46
67
97
92
84
94
77
83
88
68
95
93
98
100
96
98
99
91Urdu
90
80
82
63
58
61
61
—
79
65
66
40
30
39
8
—
94
78
84
50
46
53
75
—
98
96
97
99
99
92
99
—Vietnamese
91
87
93
89
86
66
63
—
79
76
81
71
65
26
—
—
94
88
98
97
93
74
90
—
99
98
100
100
100
99
100
—Welsh
91
82
85
64
49
69
72
71
78
61
69
35
11
43
34
63
96
87
88
61
39
66
85
60
99
99
99
96
95
98
98
90Xhosa
81
69
—
—
53
66
71
—
63
45
—
—
13
40
45
—
84
67
—
—
49
65
71
—
98
94
—
—
96
92
97
—Yoruba
70
62
—
8
—
15
37
—
45
33
—
1
—
5
1
—
71
61
—
1
—
11
22
—
96
93
—
21
—
28
88
—Zulu
80
71
—
—
6
63
54
—
62
45
—
—
0
35
18
—
83
72
—
—
6
63
51
—
97
95
—
—
11
92
93
—Mean
86
77
82
74
68
69
65
55
74
61
65
58
48
48
34
41
89
78
82
74
65
67
68
51
96
93
98
92
90
93
94
73Median 89.0 80.0 82.0 78.0 67.0 68.0 63.0 65.0 74.0 57.0 63.5 57.5 41.5 41.0 26.5 42.0 93.0 81.0 84.0 81.0 67.0 66.0 71.5 61.0 99.0 97.0 99.0 99.0 98.0 98.0 98.0 89.0 Standard Deviation 13.36 17.29 13.37 23.07 24.61 19.04 18.57 24.68 18.75 24.86 23.57 28.52 32.33 27.86 28.74 21.22 13.51 18.99 15.64 26.45 28.5 21.83 22.7 25.12 11.26 11.94 2.79 19.46 20.21 13.95 12.25 31.75 5. Why is it better than other libraries?
Every language detector uses a probabilistic n-gram model trained on the character distribution in some training corpus. Most libraries only use n-grams of size 3 (trigrams) which is satisfactory for detecting the language of longer text fragments consisting of multiple sentences. For short phrases or single words, however, trigrams are not enough. The shorter the input text is, the less n-grams are available. The probabilities estimated from such few n-grams are not reliable. This is why Lingua makes use of n-grams of sizes 1 up to 5 which results in much more accurate prediction of the correct language.
A second important difference is that Lingua does not only use such a statistical model, but also a rule-based engine. This engine first determines the alphabet of the input text and searches for characters which are unique in one or more languages. If exactly one language can be reliably chosen this way, the statistical model is not necessary anymore. In any case, the rule-based engine filters out languages that do not satisfy the conditions of the input text. Only then, in a second step, the probabilistic n-gram model is taken into consideration. This makes sense because loading less language models means less memory consumption and better runtime performance.
In general, it is always a good idea to restrict the set of languages to be considered in the classification process using the respective api methods. If you know beforehand that certain languages are never to occur in an input text, do not let those take part in the classifcation process. The filtering mechanism of the rule-based engine is quite good, however, filtering based on your own knowledge of the input text is always preferable.
6. Test report generation
If you want to reproduce the accuracy results above, you can generate the test reports yourself for all classifiers and languages by executing:
For each detector and language, a test report file is then written into /accuracy-reports . As an example, here is the current output of the Lingua German report: