When analyzing a document with Compilatio Magister or Compilatio Magister+, an “Unrecognized Languages” alert may appear when certain passages contain elements that do not match the dictionary of the detected language.
In some cases, this may indicate an attempt to alter the text in order to evade detection.
Why are some passages flagged as “Unrecognized Languages”?
The alert may appear when the document contains elements that are unusual for the language being analyzed, such as:
- words containing characters from multiple alphabets;
- invisible characters inserted into the text;
- characters or symbols that look like standard letters but come from a different character set;
- typographical changes that can affect word recognition.
For example, replacing lowercase “l” characters with uppercase “I” characters changes the actual structure of the text and can affect similarity detection, even though the difference may not be visible to the reader.
Why can these alterations affect Compilatio's analysis?
Document analysis relies in part on identifying the words contained in the text.
When words contain altered or unrecognized characters, they may not be correctly identified. This can affect how the analyzed content is interpreted and limit the relevance of the results obtained.
What should you do when an “Unrecognized Languages” alert appears?
If this alert appears in an analysis report, we recommend informing the document's author.
You can ask them to provide an unaltered version of their work, without modified characters or other elements that could interfere with text recognition. A new analysis can then be performed using the unaltered text as a reliable basis.
This article has been automatically translated. If you notice a translation error, please contact us.