Python как определить язык текста

от admin

Detecting Text Language With Python and NLTK

Most of us are used to Internet search engines and social networks capabilities to show only data in certain language, for example, showing only results written in Spanish or English. To achieve that, indexed text must have been analized previously to “guess” the languange and store it together.

There are several ways to do that; probably the most easy to do is a stopwords based approach. The term “stopword” is used in natural language processing to refer words which should be filtered out from text before doing any kind of processing, commonly because this words are little or nothing usefult at all when analyzing text.

How to do that?

Ok, so we have a text whose language we want to detect depending on stopwords being used in such text. First step is to “tokenize” — convert given text to a list of “words” or “tokens” — using an approach or another depending on our requeriments: should we keep contractions or, otherwise, should we split them? we need puntuactions or want to split them off? and so on.

In this case we are going to split all punctuations into separate tokens:

nltk “wordpunct_tokenize” tokenizer

As shown, the famous quote from Mr. Wolf has been splitted and now we have “clean” words to match against stopwords list.

At this point we need stopwords for several languages and here is when NLTK comes to handy:

included languages in NLTK

Now we need to compute language probability depending on which stopwords are used:

calculate languages ratios

First we tokenize using wordpunct_tokenize function and lowercase all splitted tokens, then we walk across nltk included languages and count how many unique stopwords are seen in analyzed text to put this in “language_ratios” dictionary.

Finally, we only have to get the “key” with biggest “value”:

get most rated language

So yes, it seems this approach works fine with well written texts — those who respect grammatical rules — (and not so small ones) and is really easy to implement.

Putting it all together

If we put all the explained above into a script we have something like this:

There are others ways to “guess” language from a given text like N-Gram-Based text categorization so will see it in, probably, next post.

1. TextBlob.

Note: This solution requires internet access and Textblob is using Google Translate’s language detector by calling the API.

2. Polyglot.

Requires numpy and some arcane libraries, unlikely to get it work for Windows. (For Windows, get an appropriate versions of PyICU, Morfessor and PyCLD2 from here, then just pip install downloaded_wheel.whl .) Able to detect texts with mixed languages.

pip install polyglot

To install the dependencies, run: sudo apt-get install python-numpy libicu-dev

3. chardet

Chardet has also a feature of detecting languages if there are character bytes in range (127-255]:

pip install chardet

4. langdetect

Requires large portions of text. It uses non-deterministic approach under the hood. That means you get different results for the same text sample. Docs say you have to use following code to make it determined:

pip install langdetect

5. guess_language

Can detect very short samples by using this spell checker with dictionaries.

pip install guess_language-spirit

6. langid

langid.py provides both a module

and a command-line tool:

pip install langid

7. FastText

FastText is a text classifier, can be used to recognize 176 languages with a proper models for language classification. Download this model, then:

pip install fasttext

8. pyCLD3

pycld3 is a neural network model for language identification. This package contains the inference code and a trained model.

pip install pycld3

Have you had a look at langdetect?

dheiberg's user avatar

And @toto_tico did a nice job in presenting the speed comparison.

Here’s a summary to complete the great answers above (as of 2021)

Language ID software Used by Open Source / Model Rule-based Stats-based Can train/tune
Google Translate Language Detection TextBlob (limited usage)
Chardet
Guess Language (non-active development) spirit-guess (updated rewrite) Minimally
pyCLD2 Polyglot Somewhat Not sure
CLD3 Possibly
langid-py Not sure
langdetect SpaCy-langdetect
FastText What The Lang Not sure

David Beauchemin's user avatar

If you are looking for a library that is fast with long texts, polyglot and fastext are doing the best job here.

I sampled 10000 documents from a collection of dirty and random HTMLs, and here are the results:

I have noticed that a lot of the methods focus on short texts, probably because it is the hard problem to solve: if you have a lot of text, it is really easy to detect languages (e.g. one could just use a dictionary!). However, this makes it difficult to find for an easy and suitable method for long texts.

toto_tico's user avatar

There is an issue with langdetect when it is being used for parallelization and it fails. But spacy_langdetect is a wrapper for that and you can use it for that purpose. You can use the following snippet as well:

Habib Karbasian's user avatar

You can use Googletrans (unofficial) a free and unlimited Google translate API for Python.

You can make as many requests as you want, there are no limits

Installation:

Language detection:

h3t1's user avatar

Pretrained Fast Text Model Worked Best For My Similar Needs

I arrived at your question with a very similar need. I found the most help from Rabash’s answers for my specific needs.

After experimenting to find what worked best among his recommendations, which was making sure that text files were in English in 60,000+ text files, I found that fasttext was an excellent tool for such a task.

With a little work, I had a tool that worked very fast over many files. But it could be easily modified for something like your case, because fasttext works over a list of lines easily.

My code with comments is among the answers on THIS post. I believe that you and others can easily modify this code for other specific needs.

Depending on the case, you might be interested in using one of the following methods:

Method 0: Use an API or library

Usually, there are a few problems with these libraries because some of them are not accurate for small texts, some languages are missing, are slow, require internet connection, are non-free. But generally speaking, they will suit most needs.

Method 1: Language models

A language model gives us the probability of a sequence of words. This is important because it allows us to robustly detect the language of a text, even when the text contains words in other languages (e.g.: «‘Hola’ means ‘hello’ in spanish»).

You can use N language models (one per language), to score your text. The detected language will be the language of the model that gave you the highest score.

If you want to build a simple language model for this, I’d go for 1-grams. To do this, you only need to count the number of times each word from a big text (e.g. Wikipedia Corpus in «X» language) has appeared.

Then, the probability of a word will be its frequency divided by the total number of words analyzed (sum of all frequencies).

If the text to detect is quite big, I recommend sampling N random words and then use the sum of logarithms instead of multiplications to avoid floating-point precision problems.

Method 2: Intersecting sets

An even simpler approach is to prepare N sets (one per language) with the top M most frequent words. Then intersect your text with each set. The set with the highest number of intersections will be your detected language.

Method 3: Zip compression

This more a curiosity than anything else, but here it goes. You can compress your text (e.g LZ77) and then measure the zip-distance with regards to a reference compressed text (target language). Personally, I didn’t like it because it’s slower, less accurate and less descriptive than other methods. Nevertheless, there might be interesting applications for this method. To read more: Language Trees and Zipping

4 Python libraries to detect English and Non-English language

We will discuss spacy-langdetect, Pycld2, TextBlob, and Googletrans for language detection

Are you sure that the input text data for your model is in English? Well, no one can be sure about this, as no one will read around 20k records of text data.

So, how non-English text will affect your English text trained model?

Pick any non-English text and pass it through as input to your English text trained classification model. You will come to know that the category is assigned to non-English text by the model.

If your model is dependent on one language then, other languages in your textual data should be considered as noise.

The job of the text classification model is to classify. And, it will do its job despite its input text will be in English or not.

What can we do to avoid such a situation?

Your model will not stop classifying the non-English text. So, you have to detect the non-English text and remove it from trained data and prediction data.

This process comes under the data cleaning part. Inconsistency in your data will result in a decrease in the accuracy of the model. Sometimes, multiple languages present in text data could be one of the reasons your model behaves strangely.

So, in this article, we will discuss the different python libraries which detect the language(s) of the text data.

Let’s start with the spaCy library.

1. SpaCy

You need to install the spacy-langdetect and spacy python libraries for the below code to work.

#1. Download the best-matching default model and create a shortcut link.

#2. Add LanguageDetector() function and model to NLP pipeline.

#3. Pass the text data into the pipeline for language detection.

#4 Store the detected language and accuracy in the detect_language variable.

We have tested the above library with only single language text data. What will happen when the text has multiple language sentences?

Predicted language is en, that is, English. Prediction accuracy is also low. But, this text also has a German-language sentence. And, this library did not predict that.

This library returns the detected language of longer sentences.

The spaCy is a python library used in different Natural Language Processing (NLP) tasks. It’s deploying a spacy_langdetect library model in the spaCy NLP pipeline.

2. Pycld2

What if we want to know the detected language of all sentences?

In that case, use the below code. You need to install the pycld2 python library for the below code to work.

Pycld2 python library is a python binding for the Compact Language Detect 2 (CLD2). You can explore the different functionality of Pycld2. Know about the Pycld2 here.

3. TextBlob

Let’s say you are using TextBlob for the NLP task. Then will you use spaCy or TextBlob for the language detection? I will use TextBlob and also should you.

And we also love choices.

You need to install the textblob python library for the below code to work.

TextBlob also returns the detected language of longer sentences. Use pycld2 library for multiple language sentences.

Do you know which language is represented by the “ru”?

If you don’t then, then visit this Wikipedia link. Or you can check the below list. It is the standard short-form (ISO 639–1 code) of the language used in data science.

TextBlob provides an API to perform different NLP tasks. Some applications of TextBlob are text processing, sentiment analysis, classification, spelling correction, keyword extraction, part-of-speech tagging, etc. know about TextBlob here.

4. Googletrans

Googletrans python library uses the google translate API to detect the language of text data. But this library is not reliable. So, be careful before using this library. You can consider this library as another choice for language detection.

You need to install the googletrans python library for the below code to work.

Googletrans also has the functionality for language translation. Know about Googletrans here.

Application of language detection

  1. Find out bias in text data based on the languages.
  2. You can classify the article based on the different languages.
  3. Language is generally associated with the region. This method helps you to classify the article based on languages.
  4. You can use this method in the language translation model.
  5. You can use it in data cleaning and data manipulation processes.

Conclusion

We should consider language detection as one of the data cleaning processes, for textual data. Internet text data are not always present in the English language.

In this article, I have explains different python libraries to detect the languages.

These libraries will help you to remove the noise from your data. It’s better to know that your data is noise-free. And you have eliminated one reason if your model doesn’t perform on new data.

Name already in use

If nothing happens, download GitHub Desktop and try again.

Launching GitHub Desktop

If nothing happens, download GitHub Desktop and try again.

Launching Xcode

If nothing happens, download Xcode and try again.

Launching Visual Studio Code

Your codespace will open once ready.

There was a problem preparing your codespace, please try again.

Latest commit

Git stats

Files

Failed to load latest commit information.

README.md

1. What does this library do?

Its task is simple: It tells you which language some text is written in. This is very useful as a preprocessing step for linguistic data in natural language processing applications such as text classification and spell checking. Other use cases, for instance, might include routing e-mails to the right geographically located customer service department, based on the e-mails’ languages.

2. Why does this library exist?

Language detection is often done as part of large machine learning frameworks or natural language processing applications. In cases where you don’t need the full-fledged functionality of those systems or don’t want to learn the ropes of those, a small flexible library comes in handy.

Python is widely used in natural language processing, so there are a couple of comprehensive open source libraries for this task, such as Google’s CLD 2 and CLD 3, langid, fastText and langdetect. Unfortunately, except for the last one they have two major drawbacks:

  1. Detection only works with quite lengthy text fragments. For very short text snippets such as Twitter messages, they do not provide adequate results.
  2. The more languages take part in the decision process, the less accurate are the detection results.

Lingua aims at eliminating these problems. She nearly does not need any configuration and yields pretty accurate results on both long and short text, even on single words and phrases. She draws on both rule-based and statistical methods but does not use any dictionaries of words. She does not need a connection to any external API or service either. Once the library has been downloaded, it can be used completely offline.

3. Which languages are supported?

Compared to other language detection libraries, Lingua’s focus is on quality over quantity, that is, getting detection right for a small set of languages first before adding new ones. Currently, the following 75 languages are supported:

  • A
    • Afrikaans
    • Albanian
    • Arabic
    • Armenian
    • Azerbaijani
    • Basque
    • Belarusian
    • Bengali
    • Norwegian Bokmal
    • Bosnian
    • Bulgarian
    • Catalan
    • Chinese
    • Croatian
    • Czech
    • Danish
    • Dutch
    • English
    • Esperanto
    • Estonian
    • Finnish
    • French
    • Ganda
    • Georgian
    • German
    • Greek
    • Gujarati
    • Hebrew
    • Hindi
    • Hungarian
    • Icelandic
    • Indonesian
    • Irish
    • Italian
    • Japanese
    • Kazakh
    • Korean
    • Latin
    • Latvian
    • Lithuanian
    • Macedonian
    • Malay
    • Maori
    • Marathi
    • Mongolian
    • Norwegian Nynorsk
    • Persian
    • Polish
    • Portuguese
    • Punjabi
    • Romanian
    • Russian
    • Serbian
    • Shona
    • Slovak
    • Slovene
    • Somali
    • Sotho
    • Spanish
    • Swahili
    • Swedish
    • Tagalog
    • Tamil
    • Telugu
    • Thai
    • Tsonga
    • Tswana
    • Turkish
    • Ukrainian
    • Urdu
    • Vietnamese
    • Welsh
    • Xhosa
    • Yoruba
    • Zulu

    4. How good is it?

    Lingua is able to report accuracy statistics for some bundled test data available for each supported language. The test data for each language is split into three parts:

    1. a list of single words with a minimum length of 5 characters
    2. a list of word pairs with a minimum length of 10 characters
    3. a list of complete grammatical sentences of various lengths

    Both the language models and the test data have been created from separate documents of the Wortschatz corpora offered by Leipzig University, Germany. Data crawled from various news websites have been used for training, each corpus comprising one million sentences. For testing, corpora made of arbitrarily chosen websites have been used, each comprising ten thousand sentences. From each test corpus, a random unsorted subset of 1000 single words, 1000 word pairs and 1000 sentences has been extracted, respectively.

    Given the generated test data, I have compared the detection results of Lingua, fastText, langdetect, langid, CLD 2 and CLD 3 running over the data of Lingua’s supported 75 languages. Languages that are not supported by the other detectors are simply ignored for them during the detection process.

    Each of the following sections contains two plots. The bar plot shows the detailed accuracy results for each supported language. The box plot illustrates the distributions of the accuracy values for each classifier. The boxes themselves represent the areas which the middle 50 % of data lie within. Within the colored boxes, the horizontal lines mark the median of the distributions.

    4.1 Single word detection

    Single Word Detection Performance

    Single Word Detection Performance

    Bar plot

    4.2 Word pair detection

    Word Pair Detection Performance

    Word Pair Detection Performance

    Bar plot

    4.3 Sentence detection

    Sentence Detection Performance

    Sentence Detection Performance

    Bar plot

    4.4 Average detection

    Average Detection Performance

    Average Detection Performance

    Bar plot

    4.5 Mean, median and standard deviation

    The table below shows detailed statistics for each language and classifier including mean, median and standard deviation.

    Open table

    Language Average Single Words Word Pairs Sentences
    Lingua
    (high accuracy mode)
    Lingua
    (low accuracy mode)
    Langdetect FastText Langid CLD3 CLD2 Simplemma Lingua
    (high accuracy mode)
    Lingua
    (low accuracy mode)
    Langdetect FastText Langid CLD3 CLD2 Simplemma Lingua
    (high accuracy mode)
    Lingua
    (low accuracy mode)
    Langdetect FastText Langid CLD3 CLD2 Simplemma Lingua
    (high accuracy mode)
    Lingua
    (low accuracy mode)
    Langdetect FastText Langid CLD3 CLD2 Simplemma
    Afrikaans 78 64 67 36 30 55 55 57 38 38 11 1 22 13 80 62 65 23 10 46 56 96 93 98 74 80 98 96
    Albanian 88 80 79 66 65 55 65 21 68 55 53 35 33 18 18 23 95 87 84 66 63 48 77 16 100 99 100 98 98 98 99 23
    Arabic 99 95 97 96 91 90 67 97 89 93 89 84 79 19 99 96 98 98 90 92 82 100 99 100 100 98 100 99
    Armenian 100 100 100 94 99 100 15 100 100 100 83 100 100 24 100 100 100 99 100 100 13 100 100 100 100 97 100 7
    Azerbaijani 89 82 78 68 81 72 78 71 57 36 62 34 92 78 80 69 82 82 99 96 98 98 99 99
    Basque 83 74 71 52 62 61 71 56 44 18 33 23 87 76 70 52 62 69 92 91 100 86 92 91
    Belarusian 97 92 85 85 84 76 91 80 69 69 67 42 99 95 88 87 86 87 100 100 98 99 100 99
    Bengali 100 100 100 98 92 99 63 100 100 100 94 92 98 19 100 100 100 99 88 99 69 100 100 100 100 97 99 99
    Bokmal 58 49 13 47 38 27 3 15 59 47 12 39 75 74 23 86
    Bosnian 33 29 9 5 33 19 28 22 9 2 19 4 32 29 10 4 28 15 39 36 8 8 52 36
    Bulgarian 87 78 72 78 67 70 66 69 71 57 51 56 46 45 32 46 91 81 68 81 62 66 72 70 99 97 96 99 93 98 93 90
    Catalan 70 58 54 57 38 48 38 62 50 33 25 33 5 19 4 39 74 60 51 57 29 42 30 63 86 81 86 83 81 84 79 82
    Chinese 100 100 64 71 96 92 33 100 100 39 46 90 92 100 100 55 68 97 83 2 100 100 97 100 100 100 98
    Croatian 72 59 73 47 48 42 51 53 36 50 28 16 26 34 74 57 70 42 38 42 47 90 85 97 72 90 58 73
    Czech 80 71 71 76 66 64 74 52 65 54 51 58 44 39 50 35 84 72 72 79 69 65 80 45 90 87 88 92 86 88 91 75
    Danish 81 70 70 62 60 58 59 56 61 45 50 35 33 26 27 27 84 70 68 57 61 54 56 52 98 95 93 95 86 95 94 89
    Dutch 77 64 58 78 64 58 47 57 55 36 27 55 34 29 11 34 81 61 49 81 61 47 42 46 96 94 98 100 98 97 90 90
    English 81 62 60 96 85 54 56 69 55 29 23 90 84 22 12 38 89 62 59 98 71 44 55 71 99 96 99 100 99 97 100 98
    Esperanto 83 66 76 44 57 50 67 44 51 5 22 7 85 61 79 30 51 46 98 92 100 96 98 98
    Estonian 92 83 82 73 67 70 65 72 80 62 62 50 37 41 24 52 96 88 86 73 67 69 73 69 100 99 100 96 98 99 99 94
    Finnish 96 91 93 92 83 80 77 81 90 77 84 82 62 58 44 63 98 95 95 96 88 84 89 82 100 100 100 100 100 99 98 99
    French 89 77 75 83 71 55 46 70 74 52 48 62 42 22 12 46 95 83 78 86 74 49 48 69 99 97 99 99 98 94 80 94
    Ganda 92 84 61 80 65 23 95 87 62 100 100 99
    Georgian 100 100 99 99 98 100 4 100 100 97 97 99 100 11 100 100 99 100 100 100 2 100 100 100 100 96 100 0
    German 89 80 73 89 81 66 64 83 74 57 50 76 61 40 27 64 94 84 70 93 81 62 66 86 100 99 100 100 100 98 98 99
    Greek 100 100 100 99 100 100 100 75 100 100 100 98 100 100 100 71 100 100 100 100 100 100 100 61 100 100 100 100 100 100 100 93
    Gujarati 100 100 100 100 100 100 100 100 100 100 99 100 99 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100
    Hebrew 100 100 100 100 100 100 100 100 99 100 100 100 100 100 100 100 100 100 100 100
    Hindi 73 33 67 87 60 58 77 5 61 11 44 74 41 34 56 2 64 20 59 88 47 45 76 4 93 67 99 99 92 95 99 11
    Hungarian 95 90 88 92 83 76 75 65 87 77 74 80 64 53 41 50 98 94 91 96 86 76 85 61 100 100 100 100 100 99 100 84
    Icelandic 93 88 65 66 71 66 67 83 72 39 33 42 26 50 97 92 57 66 70 73 60 100 99 98 99 99 99 92
    Indonesian 62 47 80 69 51 46 62 34 41 25 56 43 16 26 36 39 62 45 85 68 54 45 63 30 82 71 100 95 82 66 88 32
    Irish 91 85 60 63 67 66 80 82 70 35 28 42 29 74 94 90 57 64 66 78 76 96 95 89 97 94 92 90
    Italian 87 71 76 89 66 62 44 66 69 42 50 74 28 31 7 42 92 74 80 92 70 57 32 61 100 98 99 100 100 98 93 94
    Japanese 100 100 100 87 86 98 33 100 100 99 72 61 97 100 100 100 89 96 96 100 100 100 100 100 100 100
    Kazakh 92 90 88 80 82 77 80 78 72 67 62 43 97 93 90 78 83 88 99 99 100 96 99 99
    Korean 100 100 100 99 100 99 100 100 100 100 98 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 98 100
    Latin 87 73 50 21 62 46 70 72 49 24 44 9 50 92 76 41 2 58 42 66 97 93 85 61 83 88 93
    Latvian 93 87 89 82 83 75 72 43 85 75 76 65 64 51 33 37 97 90 91 83 86 77 84 32 99 97 99 97 98 98 98 59
    Lithuanian 95 88 87 81 80 72 70 72 87 76 71 61 58 42 30 65 98 89 91 83 85 75 82 64 100 98 100 99 99 99 99 86
    Macedonian 83 72 86 74 51 60 60 14 64 52 71 51 15 30 27 15 86 70 88 72 44 54 70 11 99 95 100 100 94 97 84 16
    Malay 30 31 15 11 22 18 13 25 22 14 2 11 9 3 36 36 19 9 22 22 10 28 35 12 22 34 23 26
    Maori 92 83 52 61 84 64 22 12 92 88 43 72 99 98 91 98
    Marathi 85 41 88 80 80 84 83 74 20 76 61 70 69 65 84 30 90 81 79 84 86 96 72 98 99 91 98 99
    Mongolian 97 95 81 86 83 78 92 89 59 68 63 43 99 98 86 90 87 92 99 99 98 99 99 100
    Nynorsk 65 52 29 32 54 30 41 25 8 5 18 9 65 49 18 16 50 24 90 81 61 75 93 55
    Persian 90 80 81 90 92 76 61 12 77 62 64 79 83 57 13 12 93 80 80 92 94 70 72 5 100 98 99 100 100 99 99 18
    Polish 95 90 89 92 89 77 75 90 86 77 75 80 73 51 38 84 98 93 93 97 93 80 87 88 100 99 100 100 100 99 99 99
    Portuguese 81 69 61 73 54 53 54 65 58 42 30 47 19 21 20 36 85 69 55 71 44 40 48 62 98 96 98 99 98 97 94 97
    Punjabi 100 100 100 100 100 100 100 100 100 100 99 100 99 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100
    Romanian 87 72 77 64 61 53 54 61 69 49 56 38 31 24 11 44 92 74 78 60 60 48 53 53 99 94 97 95 92 88 96 86
    Russian 89 79 84 94 75 71 60 71 76 60 70 86 60 48 26 64 94 84 87 98 75 72 68 66 98 92 96 100 91 93 87 84
    Serbian 88 78 76 64 78 69 74 62 54 39 63 29 90 80 76 63 75 78 99 91 98 89 95 99
    Shona 90 81 76 65 77 56 51 24 95 86 79 71 100 100 99 99
    Slovak 84 75 75 65 68 63 71 71 64 49 51 41 40 32 38 54 90 79 76 62 66 61 76 68 99 98 98 91 97 96 99 92
    Slovene 82 67 73 59 63 63 48 78 61 39 47 32 33 29 8 62 87 68 72 54 61 60 42 76 98 93 98 90 95 99 92 96
    Somali 92 85 90 24 69 70 82 64 75 4 38 27 95 91 94 15 70 83 100 100 100 52 100 99
    Sotho 83 72 49 54 61 43 15 13 88 75 33 54 99 97 98 95
    Spanish 70 56 57 74 65 48 43 53 44 25 26 51 37 16 12 24 69 49 47 72 59 32 34 42 97 94 97 100 98 96 85 92
    Swahili 79 70 73 41 42 57 57 52 58 43 47 7 3 25 16 36 82 69 73 24 24 49 59 44 98 97 99 92 98 98 97 76
    Swedish 84 72 68 76 65 61 53 66 64 46 40 51 35 30 14 43 89 76 67 78 63 56 52 65 99 95 96 98 96 96 93 89
    Tagalog 78 66 76 45 42 50 14 52 36 50 11 2 9 16 82 66 77 28 26 44 12 98 96 99 98 98 95 15
    Tamil 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 99 100
    Telugu 100 100 100 100 100 99 100 100 100 100 100 100 99 100 100 100 100 100 100 100 100 100 100 100 100 100 99 100
    Thai 99 99 100 100 100 99 100 100 100 100 100 100 100 100 100 100 100 100 100 100 100 98 98 100 100 100 98 100
    Tsonga 83 72 61 63 46 19 88 73 68 98 97 97
    Tswana 82 71 56 62 44 17 86 73 57 98 96 94
    Turkish 94 87 82 86 67 69 66 82 84 71 63 70 50 41 30 71 98 91 84 88 67 70 71 80 100 99 100 100 84 97 97 94
    Ukrainian 92 86 83 91 76 81 77 75 84 75 66 78 54 62 46 67 97 92 84 94 77 83 88 68 95 93 98 100 96 98 99 91
    Urdu 90 80 82 63 58 61 61 79 65 66 40 30 39 8 94 78 84 50 46 53 75 98 96 97 99 99 92 99
    Vietnamese 91 87 93 89 86 66 63 79 76 81 71 65 26 94 88 98 97 93 74 90 99 98 100 100 100 99 100
    Welsh 91 82 85 64 49 69 72 71 78 61 69 35 11 43 34 63 96 87 88 61 39 66 85 60 99 99 99 96 95 98 98 90
    Xhosa 81 69 53 66 71 63 45 13 40 45 84 67 49 65 71 98 94 96 92 97
    Yoruba 70 62 8 15 37 45 33 1 5 1 71 61 1 11 22 96 93 21 28 88
    Zulu 80 71 6 63 54 62 45 0 35 18 83 72 6 63 51 97 95 11 92 93
    Mean 86 77 82 74 68 69 65 55 74 61 65 58 48 48 34 41 89 78 82 74 65 67 68 51 96 93 98 92 90 93 94 73
    Median 89.0 80.0 82.0 78.0 67.0 68.0 63.0 65.0 74.0 57.0 63.5 57.5 41.5 41.0 26.5 42.0 93.0 81.0 84.0 81.0 67.0 66.0 71.5 61.0 99.0 97.0 99.0 99.0 98.0 98.0 98.0 89.0
    Standard Deviation 13.36 17.29 13.37 23.07 24.61 19.04 18.57 24.68 18.75 24.86 23.57 28.52 32.33 27.86 28.74 21.22 13.51 18.99 15.64 26.45 28.5 21.83 22.7 25.12 11.26 11.94 2.79 19.46 20.21 13.95 12.25 31.75

    5. Why is it better than other libraries?

    Every language detector uses a probabilistic n-gram model trained on the character distribution in some training corpus. Most libraries only use n-grams of size 3 (trigrams) which is satisfactory for detecting the language of longer text fragments consisting of multiple sentences. For short phrases or single words, however, trigrams are not enough. The shorter the input text is, the less n-grams are available. The probabilities estimated from such few n-grams are not reliable. This is why Lingua makes use of n-grams of sizes 1 up to 5 which results in much more accurate prediction of the correct language.

    A second important difference is that Lingua does not only use such a statistical model, but also a rule-based engine. This engine first determines the alphabet of the input text and searches for characters which are unique in one or more languages. If exactly one language can be reliably chosen this way, the statistical model is not necessary anymore. In any case, the rule-based engine filters out languages that do not satisfy the conditions of the input text. Only then, in a second step, the probabilistic n-gram model is taken into consideration. This makes sense because loading less language models means less memory consumption and better runtime performance.

    In general, it is always a good idea to restrict the set of languages to be considered in the classification process using the respective api methods. If you know beforehand that certain languages are never to occur in an input text, do not let those take part in the classifcation process. The filtering mechanism of the rule-based engine is quite good, however, filtering based on your own knowledge of the input text is always preferable.

    6. Test report generation

    If you want to reproduce the accuracy results above, you can generate the test reports yourself for all classifiers and languages by executing:

    For each detector and language, a test report file is then written into /accuracy-reports . As an example, here is the current output of the Lingua German report:

    Читать:
    Программа запускается только от имени администратора как исправить

Похожие статьи