Showing posts with label Multilingual taxonomies. Show all posts
Showing posts with label Multilingual taxonomies. Show all posts

Thursday, July 30, 2026

Generative AI for Creating Multilingual Taxonomies

Global Business with AI Translator photo by Roman ShashkoTaxonomies enhance the capability to identify and retrieve desired content.  Automated translation (machine translation) expands the scope of access to information to content in other languages. Applying automated translation to the contented identified using taxonomies thus enables people to find content on the subjects they want and then have it translated if needed. To retrieve such multilingual content, however, a taxonomy must also be multilingual. Concepts in the taxonomy must have names/labels in different languages: the language of the user and the languages of the content so that it can be tagged to the content.

Creating multilingual taxonomies

Fortunately, the data model upon which most large taxonomies are being built and managed supports multilingual concepts. The data model standard SKOS (Simple Knowledge Organization System) supports bilingual and multilingual taxonomies, because it models “concepts,” not “terms,” which has been referred to as “things, not strings,” and concepts can have any number of labels (strings of text) to describe them: a preferred label in each language that is displayed and any number of alternative (variant) labels in any languages, which are not displayed but match user searches and text strings for tagging.

The fact that multilingual taxonomies are simple to support technologically does not mean that they have been easy to create with translations. Whether translation is done by human translators or by machine translation software, translating a taxonomy is not as easy as translating narrative text. Both human translators and automated translation tools, (using rule-based, statistical, or neural network AI methods) look at complete sentences and not just isolated words. Words may have different meanings depending on their use in a sentence, and the context of a sentence may call for a different synonym in the translation target language.

Taxonomies with their hierarchical relationships provide context for the meaning of their concepts, but this is not the same kind of sentence-based context that machine translation programs utilize. Human translators can look at the context of the taxonomy hierarchy, but it’s a much slower task than translating sentences (which can also use machine-assisted translation tools) and requires an understanding of taxonomies.  Relying on machine translation or human translators who don’t understand the purpose or subject area of the taxonomy can lead to errors in translating taxonomies.

LLMs for generating multilingual taxonomies

Now, AI, in the form of LLMs, has become available to help generate taxonomies (which still require human review). LLMs don’t just generate terms, but they generate the correct hierarchical relationships to other concepts, multiple labels (as synonyms) for the same concept, and even definitions for concepts.

LLMs go beyond traditional machine translation and take the existing hierarchy into account when generating a translated label for a taxonomy concept. LLMs also follow specific instructions (or “prompts”) regarding the purpose and nature of the taxonomy. LLMs can also be instructed to focus on the meaning of the concept rather than on a literal translation of the preferred labels and each of the alternative labels. While the preferred labels for a concept in different languages are close translations, the alternative labels are not translations of each other but rather refer back to the concept. Even the number of alternative labels for a concept will vary by language. Finally, definitions are generated based on the concept and not as translations of existing definitions.

I explained this in a little more detail in my prior blog post Generative AI and Taxonomies for Finding Information

LLMs for generating new taxonomies in different languages

LLMs can be trained on, access, and generate text in different languages. This means that LLMs can generate entire taxonomies or parts of taxonomies in different languages from the start (with appropriate human-created prompts), without having to translate an entire existing taxonomy into another language.

Although a monolingual taxonomy does not have the same benefits of retrieving content in multiple languages as does a multilingual taxonomy, sometimes a monolingual taxonomy is desired. If a taxonomist is not available to create an entire taxonomy in the language, and a taxonomy on the subject already exists in another language, then translating a taxonomy has been a method to create a monolingual taxonomy in another language. As previously explained, translation of a taxonomy is not ideal and is prone to errors if not done by a translator-taxonomist. Using LLMs with the involvement of a subject matter expert (not necessarily a taxonomist) will generate a better taxonomy than a translation.

When using AI to build a taxonomy, it’s still best to have the top level developed manually to serve the specific use case, to create certain branches manually that are specific to an organization, and to generate incrementally those parts of the taxonomy built with AI, with the subject matter expert approving/disapproving suggestions for concepts and alternative labels at various stages. What is significant is that a person who is a combined taxonomist/linguist/subject matter expert is not needed to create a taxonomy in each language. In this way, taxonomy creation becomes easier to do globally. 

LLMs and multilingual generation in taxonomy management software

Managing a taxonomy in taxonomy management software has many benefits, especially when managing multilingual concepts. Now taxonomy management software is beginning to include support for LLMs to generate parts or all of a taxonomy, and thus support is integrated into the taxonomy creation and editing workflow. One tool, Graph Modeling from Graphwise, has now (last month) added the feature taxonomy generation in different languages to its Taxonomy Builder generative AI component. You can generate a monolingual taxonomy in any configured project language, and you can generate additional language versions of a taxonomy (creating a multilingual taxonomy) by generating concept preferred labels, alternative labels, and definitions, through LLM generations, not as direct translations. This essentially eliminates the need to translate taxonomies.

Saturday, June 28, 2025

A Multilingual Thesaurus Standard


Standards for taxonomies are of two kinds:
1) data models for interoperability and machine-readability, namely SKOS (Simple Knowledge Organization System) published by the W3C, and
2) best practices guidelines, which focus on thesauri but are relevant for taxonomies. These are ANSI/NISO Z39.19 and ISO 25964. The International Organization for Standardization will publish a revised edition of ISO 25964 Part 1: Thesauri for Information Retrieval later this year. I have been contributing to the revision as a member of its international working group

I have written before on Standards for Taxonomies, which is at a high level, and I will likely write again about the revisions in the new version of ISO 25964 Part 1, when it will be published. For now, I’d like to discuss some the specifics of defining an international standard which I have been working on recently.

Different Language Versions

The international standard is written in English, and it will be translated into other languages in the future. Since it will not be assumed to be translated into certain languages, and since the standard covers multilingual thesauri, it needs to include examples in different languages. Some of the examples within the sections of ISO 25964-1 are translated into common languages, such as French, German, and Spanish, but other languages are not included. Thus, this standard also includes an extensive table of the “tags” and “expansions” or terminology that appear in a thesaurus for 10 additional languages. Examples of tags include BT (Broader Term), NT (Narrower term), and SN (Scope note).

A German reviewer pointed out some errors in the German column of the table, which prompted me to look more carefully, and I noticed some issues in the Russian and Arabic, which are languages I had studied long ago and which are not represented by native speakers in our working group. I then sought other sources on thesauri in those languages, examples of thesauri on the web, and native-speaker experts.

As it turns out, for the specialized use of thesauri, it’s not just a matter of a translation, but what is used in the context. Scope note could have various translations in a language, as both the words “scope” and “note” can have different translations. Even, “broader,” narrower” and “related” can be translated differently. Broad can mean “wide,” and thus perhaps “superordinate” and “subordinate” are better translations in another language.

Variations and Lack of Standards

The thesaurus terminology is quite standardized in English and somewhat less so in other languages. Although the original ISO and German DIN thesaurus standards go back to 1974 and 1972 respectively, these standards have never been free and are actually rather expensive for the number of pages, unlike the ANSI/NISO standard, which has been made freely available since 1974. Thus, the free English-language standard from the United States has been more widely read and followed than the ISO standard. Creators of other standards sometimes translate from English, but inconsistently, rather than relying on a standard in their own language.

There are different reasons for such variations. Some thesaurus authors prefer to use terminology closer to English, while others prefer to user terminology that is more native, when near-synonyms exist. For example in Russian, “related” could be “assotsiativny” (similar to associative) or “rodstvenny,” and “concept” could be “kontsept,” or “ponyatiya.” There is also the matter of saving space with concise labels. While English has a single word for “broader” and“narrower,” a correct translation for the comparative requires two words, as in “more broad” or “more narrower” in other languages, such as French, Spanish, and Russian. Often the word for “more” is omitted to save space, but in other thesauri it is included for preciseness, such as inserting the word mas in Spanish. Arabic-language thesauri additionally vary in their use of tags/terminology depending on the region within the Arabic-speaking world of 22 countries.

I found the multilingual UNESCO thesaurus and UN library’s UNBIS thesaurus good sources to consult, since you can change not only the term display, but also the user interface with its tags and designation into different languages. However, these two UN-related sources are not even consistent with each other!

I suspect that in some thesauri the terminology was simply translated from English by a translator who was not familiar thesauri, rather than developed by a thesaurus specialist/taxonomist who would research the formats of other thesauri in that language.

Legacy Standards and Future Direction

Thesauri were originally developed to be presented in print, where space is an issue so short tags were created. Now thesauri are online, and two-letter tags are not needed and rarely displayed. But the new edition of the standard continues to include tags to be comprehensive and provide consistency with printed thesauri. However, it is my personal opinion that we should not invent comprehensive tags for all languages where they have not previously existed.

Should the standard be more descriptive or prescriptive? Descriptive would mean describing what is done in thesauri in existence. I looked up various thesauri online to see what tags and terminology they were using. If a certain designation is used more than another, such as the phrase used to mean “broader term,” then we could decide that is the standard for a language.

Prescriptive would mean to dictate the standard, typically based on expertise and belief in what would be best. In face of inconsistencies, the standard should be prescriptive. Being prescriptive would also mean that the latest revision of the standard should try to follow the prior edition and any previous translations of it, rather than merely following the usage practice the of leading examples of thesauri on the Web. The conclusion was to include both.

Although the distinction between terms and concepts is addressed in the current ISO thesaurus standard, the current summary table of tags addresses only “terms” and term relationships. The nuance of term versus concept was discussed at length by the working group and the conclusion was to include both concept to recognize an idea and term to be the representation of the concept itself.  Thus, the table of tags and terminology in the new version will now include Broader concept, Narrower concept and Related concept (which do not have tags). As new additions to the standard, the names for these in other languages thus need to be prescribed by the standard. Relying on thesauri published on the web, I found, results in too much inconsistency.  Official translations of the SKOS data model are a good source for this, but the translations exist for only some languages. I even looked at the German user interface of a SKOS-based taxonomy management software (PoolParty) and found yet other translations for broader, narrower, and related that were not consistent with the official German SKOS translation.

I hope the new edition of ISO 25964 Part 1: Thesauri for Information Retrieval will be read more widely and provide more consistency for thesauri.