Showing posts with label Controlled vocabularies. Show all posts
Showing posts with label Controlled vocabularies. Show all posts

Tuesday, April 30, 2024

Synonym Rings (or Search Thesaurus)

A synonym ring is a simple kind of controlled vocabulary that, as the name suggests, has controlled synonyms for concepts and nothing more. I have long included mention of synonym rings in presentations I’ve given with sections listing and describing controlled vocabulary types, and the synonym ring has appeared on diagrams illustrating comparative complexity and included features of the various controlled vocabularies, progressing from the simplest term lists to synonym rings, name authorities, taxonomies, thesauri, and finally ontologies.

However, until now, I have not gone into detail about synonym ring use and design.  

The name “synonym ring” is generally known only by taxonomists and other information professionals. It is called a “ring” because all synonyms point to each other, as in a circle or ring, rather than to a preferred term/label. Another name for it is a “search thesaurus,” although it should be clear that “thesaurus” is meant to be the Roget’s type and not the information retrieval type (similar to a taxonomy). I have also read the name “synset” but have not heard it in practice.

 

What we are talking about is a managed set of concepts, each with one or more synonyms, created specifically for supporting search, matching end-user search strings to text strings in the content being searched, for commonly searched concepts. The synonyms also match to variant names of the concept throughout the body of text that is being searched. Because the synonym ring’s purpose is to support search, it is not browsed and thus not displayed to the end users. Therefore, a preferred term or preferred label for each concept is not needed and thus not included.

Whether in a synonym ring or in another controlled vocabulary or taxonomy, “synonyms” refer to concept variants and not literal grammatical synonyms. In a controlled vocabulary, they are often phrases, not single words, and they are for things/concepts, and not all kinds of words (different parts of speech) found in a dictionary. They also don’t have to be exact synonyms, but rather sufficiently synonymous for the context of the content being searched.

Features of a synonym ring (search thesaurus)

  • It includes only concepts for which there are “synonyms,” Each concept must have at least two synonyms. If there are no synonyms for the concept, then the concept is not included in the synonym ring (in contrast to a regular controlled vocabulary). So, important concepts may be absent.
  • Synonyms are not displayed to the users, so slang, deprecated, potentially offensive terms, etc. may be included.
  • It supports searching only and not tagging. People doing manual tagging or systems doing auto-tagging will not be able to make use of the synonyms to identify the best concept to tag with. (They could utilize another taxonomy implemented in another system for tagging.)

Implementation of synonym rings

Typically, when taxonomists are called upon to design a taxonomy, they design it with synonyms (aka alternative labels, nonpreferred terms, variants, etc.) included. Thus, creating a dedicated synonym ring type of controlled vocabulary is not common, since the necessary synonyms are already included in the taxonomy. Small taxonomies may not have synonyms, though.

Search that is built into content/record management systems may support search synonyms, but this tends to be more ad hoc than as a managed controlled vocabulary. Recently I looked into the synonym support in controlled vocabularies and taxonomies in Salesforce Service Cloud. It supports the creation of “custom synonym groups,” where each group is a synonym ring of up to six synonyms per concept, but these have to be entered individually in the user interface, rather than as an imported as a list. As such, it’s not really a “controlled vocabulary” set.

Some content management systems with included taxonomies only enable synonyms as part of their standard displayed taxonomies and not as non-displayed search synonyms. Other systems, such as SharePoint support the use of synonyms for its taxonomies (managed in its Term Store) for tagging but not for searching.

Adding search synonyms in systems that support it often have it as a systems administrator feature, which is something that the technical systems administrators may do, while taxonomists, information architects and knowledge, managers may not know about it. After all, a set of synonyms is not a “taxonomy,” so taxonomist involvement may not even be considered. Thus, communication is necessary between those who advocate the need for comprehensive search synonyms and know how best to create them and those who are in a technical role for implementing them in a system.

Advantages of synonym rings

A synonym ring is relatively easy to develop. While there are nuances to creating synonyms (described below), it’s easier than creating other controlled vocabularies or taxonomies, since there is no need to worry about which term should be preferred and how to best create a hierarchy. Since it is not displayed, getting input from users is not required.

By focusing on supporting only searching and not also tagging, the task of coming up with synonyms is also simpler, since sometimes you want synonyms to support search and not tagging and sometimes for tagging and not searching (such as when the synonyms display to users) and trying to design for both scenarios in the same taxonomy is not easy.

When searching is the primary way that users access content, rather than browsing and filtering, a synonym ring may be an ideal solution. It might not make sense to go to the effort to design and create a hierarchical taxonomy for terms that users are searching on, if the goal is to simply enhance search.

A taxonomy runs the risk of being too broad or too specific, but a synonym ring never has that issue. The size of a synonym ring type of controlled vocabulary is flexible, and it can be built out gradually over time with no detriment.

Disadvantages of synonym rings

A synonym ring is not a standard controlled vocabulary type and is not supported in the SKOS (Simple Knowledge Organization System) data model standard of the World Wide Web Consortium. This is because a SKOS controlled vocabulary (including taxonomies) needs to have preferred labels for its concepts. Thus, synonym rings are not interoperable in the same way that other controlled vocabularies are. You cannot link to external synonym rings, and you cannot even import or export them easily. They are managed within a siloed system.

Since synonym rings do not support tagging, an additional tagging controlled vocabulary with synonyms, which is somewhat redundant in its subject scope, may need to be created

Creating synonyms for a synonym ring

“Synonyms” can include dictionary synonyms, synonyms for individual words withing multi-word phrases (e.g. political protests / political demonstrations), formal and colloquial names, acronyms, etc. Following is a list of example types:

  • synonyms: Cars / Automobiles
  • quasi-synonyms: Learning / Training
  • variant spellings: Email / E-mail
  • lexical variants: Selling / Sales
  • foreign language names: München / Munich
  • acronyms/spelled out: GDP / Gross domestic product
  • scientific/popular names: Neoplasms / Cancer
  • older/current names: Near East / Middle East

Care should be taken not to include synonyms that are not sufficiently equivalent or may be vague and have other usages, such as “development” (which could refer to software development, nonprofit fundraising, or something else). It depends on context, so in the example with “tools” as a synonym software would be acceptable if the content were only about technology and not include manufacturing, construction, etc.

Synonyms can be identified when doing research for concepts to include, including manual content analysis, automatic term extraction, lists of uncontrolled keyword tags, and search log reports. Search logs are especially suitable for synonym rings, since their usage is the same: user search strings. However, often searches are on single words, whose meaning is vague. For example, a search string word of “application” is too vague and not be used as a synonym. You should only take search log search strings if their meaning is clear.

Finally, developing synonyms for a synonym ring implemented in an internal content management system is not the same as developing synonyms for a public website to support web search engine optimization (SEO), for which they are also called “search synonyms.” For SEO, web search engine algorithms need to be considered, and obtaining the greatest number of visitors is the goal, even if those site visitors did not intend to come to the website. In such cases, more specific concepts (e.g. iPhone as synonym for cell phone) as “synonyms” would be fine. If website visitors do not find what they are looking for, that’s OK. By contrast, users of enterprise CMS or search system, would consider it a waste of their time if they retrieved additional content that did not match their search. Although sample user testing is not needed, search testing to check the accuracy of results should be performed.

Saturday, December 5, 2020

Differing Definitions of Ontologies

In my last blog post I discussed the different definitions and features of thesauri. Now, I will turn to the next kind of knowledge organization system in the spectrum of complexity: ontologies.

Actually, to consider an ontology as a more (or most) complex type of controlled vocabulary or knowledge organization system, after thesauri, due to additional features, is just one perspective or definition of ontologies, which is not universally shared.

When I first learned about ontologies, coming from my taxonomist perspective, I considered ontologies as merely a more complex type of taxonomy or thesaurus, characterized by customized semantic relationships between concepts (rather than merely hierarchical or associative relationships), more expressive attributes for concepts (rather than mere scope notes), and the grouping of concepts into classes to manage the semantic relationships and attribute types. In fact, I wrote in 2008 for the first edition of my book “An ontology can be considered a type of taxonomy with even more complex relationships than in a thesaurus,” which the following graphic represents.

As my understanding has evolved, I would consider this just to be one kind of understanding or definition of ontology among others.  In other words, a controlled vocabulary that has the features of semantic relationships, classes of concepts, and attributes for concepts, can be considered a kind of ontology, but there are other definitions and understanding of ontology within the field of information/knowledge management.

While we usually refer to “controlled vocabularies” as the over-arching category for these things, it is probably better to go up a further level and call an ontology a kind of “knowledge  organization system,” rather than a kind of controlled vocabulary. Controlled vocabularies are kinds of knowledge organization systems, where the emphasis is on managed terms or concepts for the purpose of tagging or categorizing and information retrieval. Ontologies, by themselves, are not necessarily for information retrieval, at least not directly. And this is one of the points of differing definitions of ontologies.

Differing definitions and perspective

There are differing definitions of the word ontology: (1) branch of philosophy that studies existence, being, becoming, and reality (Wikipedia: Ontology), and (2) a representation, formal naming, and definition of categories, entities, properties, and relations within a domain (Wikipedia: Ontology (information science)). Of course, we are interested in the second definition, although there are some connections between the two. 

The second definition, however, is already multidisciplinary, as it is a concept shared in both information science and computer science. Information scientists (including librarians, taxonomists, and knowledge managers) and computer scientists do not have different definitions of ontologies, but rather different approaches to and perspectives of ontologies and different purposes for the ontologies they create.  For computer scientists, modeling data and information helps them design a computer program to perform desired functions. For information scientists, modeling data and information makes it easier to retrieve information with complex queries. Information scientists consider an ontology as a kind of knowledge organization system, whereas computer scientists tend to consider an ontology as a form of knowledge representation.

Yet even among information scientists, who consider ontologies as knowledge organization systems and have the same objectives in developing ontologies, there are different understandings of what exactly constitutes an ontology and how it relates to other knowledge organization systems, such as taxonomies. This is due to (1) different emphasis on various ontology components, (2) the question of adherence to ontology standards, and (3) the way different ontology software tools model ontologies and their relations to taxonomies differently.

Differing understandings of ontology components

There is a shared understanding that ontologies are composed of things, their properties/attributes, and their relationships.

Ontology model example with classes, relations, and attributes
Ontology example with components: classes, relations, and attributes

However, there are differences in understand of the two kinds of “things”: classes and individuals. Classes are categories or groups of things with shared characteristics, whereas individuals are specific instances of things. This seems obvious, but if you approach ontology design from the perspective of taxonomy design it can become less certain. Is an individual the most specific concept (also called “leaf node”) in a hierarchy, or is an individual a named entity/proper noun? The definition of components of ontologies does not answer this question, because ontology structures are meant to model data, not to organize taxonomy concepts that could be either generic (common nouns)  named entities (proper nouns). Drawing the line between classes and individuals can be challenging, but whether this matters may depend on what tool you are using.

Furthermore, ontologies may have other components, such as axioms, rules, restrictions, events, and function terms, but ontologies as knowledge organization systems rarely have most of these.

Differing ontology standards or languages

In 2004 the World Wide Web Consortium (W3C) published the Web Ontology Language (OWL) specification, which is based on the Resource Description Framework (RDF), as “a Semantic Web language designed to represent rich and complex knowledge about things, groups of things, and relations between things,” which has become widely adopted. Now it is common to think that ontologies must follow OWL guidelines. But (information science) ontologies have existed before OWL, and an ontology does not have to follow OWL to be called an ontology. There are other ontology languages besides OWL, but they are not as common. To share and reuse ontologies, it is recommended to follow the OWL standard.

Differing ontology modeling software

While one could design the high-level model of an ontology in a mind-mapping tool, there would be no enforcement of standards or best practices (preventing duplications or incomplete data, etc.), and it’s difficult to scale, so dedicated ontology modeling software is recommended. However, ontology modeling/editing software does not model ontologies all in the same way.

The main difference is probably between stand-alone ontology software (such as Protégé or TopBraid Composer) and software that combines ontology with taxonomy/thesaurus development and editing (such as PoolParty, Semaphore, or Graphite). Stand-alone ontology editing software supports creating a detailed ontology as single model, thus including classes, multiple levels of subclasses, and individuals (instance concepts). In integrated software that combines taxonomy/thesaurus development with ontology development, the taxonomy or thesaurus (or multiple controlled vocabularies) is created in one space with one set of software features, and the ontology is created in another space with a different set of features. The ontology (or even just parts of it) is then applied to the taxonomy, so that concepts in the taxonomy inherit the attribute types and relationships of their associated class, and the taxonomy concepts are like individuals in the ontology. The ontology can be considered a semantic layer in the model, as the following graphic illustrates.

These two different approaches to ontology modeling thus result in different definitions of an ontology. A ontology is likely to be considered as a more complex type of knowledge organization system by users of stand-alone ontology software, whereas an ontology is likely to be considered and expressive semantic layer applied to one more taxonomies by users of integrated taxonomy/ontology software.

Ontology lite or ontology-like

When I was still considering ontologies more akin to thesauri with semantic relationships, and I expressed such views in a discussion forum, someone (whom I don’t remember), referred to this kind of ontology as “ontology lite,” since  it has features of an ontology, but does not fully follow an ontology model and standards. This is not necessarily a bad thing. Controlled vocabularies and knowledge organization systems can be considered along a continuum, and you should build what works for your situation.

Another kind of ontology-like structure is when you start linking multiple controlled vocabularies together. My initial experience with working on commercially implemented ontologies had been with such ontology-like systems, which were not actually called ontologies, at a former employer Gale. There we had controlled vocabularies (also called object classes) for subjects, persons, places events, products, companies/organizations, named works, etc., many of which had customized reciprocal relationship pairs between them (such as the relationship pair Creator/Creatby, between person names who were authors, and named works) and many customized term attributes (such as Birthdate, Death date, Birth city/state/country, Death city, state/country for persons).

I also heard this approach recently from a speaker, Ahren Lehnart, at Taxonomy Boot Camp conference, who described the linking of controlled vocabularies with related match (not equivalent match) relationships as “trending toward” creating an ontology.

 

Sunday, November 22, 2020

What it a Thesaurus and What is it Good For

It is somewhat ironic that in the domain of controlled vocabularies and knowledge organizations systems that there continue to exist differing meanings for “controlled vocabulary,” “taxonomy,” “thesaurus,” “ontology,” and “knowledge graph.” Hopefully, I have provided some clarification regarding what a taxonomy is and is not in my previous posts on taxonomy vs. classification, taxonomy vs. navigation, and when a taxonomy should not be hierarchical. Let’s turn now to thesauri.

Different meanings of thesaurus

I recently attended a webinar on taxonomies, ontologies, and knowledge graphs, in which a thesaurus was described as a set of synonyms for each identified concept in a list. This is not the right definition for this context. A set of synonyms for each of list of concepts is what we taxonomists call a “synonym ring”, and what administrators of enterprise search engines would call a “search thesaurus.” The use of the word “thesaurus” in this case refers to the dictionary-type thesaurus (as the default Thesaurus entry in Wikipedia) such as Roget’s Thesaurus, where synonyms are presented for each word. Synonyms are included to support search, by matching potential words and phrases entered by users into the search box with the words and phrases that likely occur in the text of content, so that content is not missed due to the searcher using a different synonym.

The “search thesaurus” (synonyms ring) differs from the synonym-dictionary thesaurus, however, in several ways, due to their different uses:
  • A search thesaurus includes phrases, not just single words as in a dictionary thesaurus.
  • A search thesaurus comprises concepts that are nouns, verbal nouns, or noun phrases, not just any part of speech as a dictionary may include.
  • The “synonyms” in a search thesaurus are appropriately equivalent terms that can be used interchangeably in all cases for the content repository, not synonyms that may be used in only some cases, as the dictionary suggests.

However, in the context of taxonomies/ontologies (not the context of search administration), the designation thesaurus has a significantly different meaning. Also referred to as in information thesaurus or information-retrieval thesaurus (to distinguish it from the synonym dictionary type), there is a different entry in Wikipedia for Thesaurus (Information Retrieval), which defines it as “a form of controlled vocabulary that seeks to dictate semantic manifestations of metadata in the indexing of content objects.” This is the meaning that relates to taxonomies and ontologies. More significant than the Wikipedia definition, are the published standards/guidelines for how to construct thesauri: ISO 25964 Thesauri and interoperability with other vocabularies and ANSI/NISO Z39.19-2005 (R2010) Guidelines for the Construction, Format, and Management of Monolingual Controlled Vocabularies. While the latter does not name thesauri in its title (although it did in an earlier version), it is essentially about thesauri and defines, in section 4.1 Definitions, a thesaurus: “A controlled vocabulary arranged in a known order and structured so that the various relationships among terms are displayed clearly and identified by standardized relationship indicators.

So, a thesaurus is a kind of controlled vocabulary or a kind of knowledge organization system which is quite structured and has certain standard features: terms that are noun phrases, hierarchical relationships between terms, associative (related, but not hierarchically) relationships between terms, “synonym” or variants, which are called nonpreferred terms, and scope notes on terms. Other metadata on terms is possible, and variations of hierarchical and associative relationships may also be possible.

Thesaurus usefulness

On the continuum chart of controlled vocabulary (knowledge organization system) types, a thesaurus falls between a taxonomy and an ontology in its level of complexity and support for semantics.


Controlled vocabulary types

Since both taxonomies and ontologies are recognized as useful, it would seem illogical that something that is in between should not be considered at least as a useful. A thesaurus has the benefits of supporting more semantics than a taxonomy while not being as complex as an ontology.

Even if most relationships are hierarchical, there may be times when creating an associative relationship between related subjects seems logical and would be helpful to users, such as relating between a process and agent, action and property, cause and effect, object and origins, discipline and practitioner, etc. Or it might not be subjects. For example, ecommerce may want to recommend “related” product categories, or content on activities could relate activities to products. In an expert people finder, person names can be related to subject areas of expertise, If the scope of “related” types is kept limited, then the generic associative relationships (“related term”) may suffice without getting to level of complexity of an ontology where there are multiple types of defined semantic relationships.

The added associative relationships and comprehensive inclusion of synonyms/nonpreferred terms also supports better (more comprehensive) tagging, whether manual or automated, by providing suggestions to the indexers or providing context for the auto-classification tool.

Finally, the overall structure of a thesaurus is more flexible than that of a taxonomy. A taxonomy groups concepts into categories with a limited number of top concepts (or “top terms”). A concept which has no broader and no narrower concept relationships, sometimes called an “orphan,” is considered an error in a taxonomy. In a thesaurus, on the other hand, where an over-arching hierarchical structure is not required (although may exist) and associative relationships are included, it is OK to have a concept with no broader and no narrower relationships, but at least an associative relationship. Thus, the taxonomist does not always have to force new concept into an existing hierarchy which might not be ideal.

Software for thesaurus management

Software to support the development and maintenance of thesauri has also been available for some time. (Taxobank has a historic list, not updated since 2013.) There actually is no such thing as “taxonomy” management software, because the software used to create taxonomies is really “thesaurus” management software, and the added thesaurus features, such as associative relationships, are just not utilized when creating a simple taxonomy.

As taxonomies have become more popular than thesauri, the software vendors have reflected that by having a hierarchical display (instead of alphabetical) as the default, and by marketing their solutions for taxonomies and ontologies and de-emphasizing or omitting mention of thesauri. For example, the basic core module of the PoolParty Semantic suite is appropriately named Thesaurus Server, since you can easily create thesauri with it, but the default hierarchical display suggests the use for taxonomies, whereas the website's product page says it’s for “Enterprise Taxonomy and Ontology Management.”

Thesauri today

Thesaurus design principles are applicable to both thesauri and taxonomies. Therefore, thesauri continue to be taught in library science and information science degree programs, including courses on information architecture. The book Information Architecture for the Web and Beyond (Rosenfeld, Morville, and Arango)(aka the polar bear book, due to its cover design), even in its 4th edition of 2015, devotes 20 pages, nearly half the chapter “Thesauri, Controlled Vocabularies and Metadata,” to thesauri.

The main impediment to thesauri is that the most common implementations these days, variations of off-the-shelf content management systems (CMS), usually do not support features of thesauri. Associative relationships are rarely supported. Synonyms/nonpreferred terms may be only partially supported (such as in the tagging view but not in retrieval). Thus, we tend to see thesauri implemented only in custom (home-grown) end-user systems, such as those of publishers of information retrieval databases.

Information retrieval thesauri have been around for a long time, and perhaps that is also part of the problem in their acceptance today in business and industry. People may consider thesauri as some kind of legacy knowledge organization system that was more predominant when we only had printed systems, not digital systems. It’s true that thesauri are designed to be useful in print, but their design is also adaptable and relevant to digital implementations. They can also form part of a larger system of interlinked controlled vocabularies.

This brings us to the next topic, ontologies, which can link to thesauri. Next month’s blog post will address the different meanings of ontology.

Saturday, December 28, 2019

Taxonomy Licensing Interest


Just over a year ago I had blogged on the topic of Taxonomy Licensing. I explained that usually a customized taxonomy is best, but occasionally licensing a taxonomy is a option worth considering in certain circumstances:  as a starting point to then modify, to serve as a single facet in a faceted taxonomy, or to index content from various external sources on a defined topic area for which a good taxonomy exists. There are issues, though, such as whether to right kind of taxonomy exists and whether the license permits modification of the taxonomy.

Various organizations, companies, and even individuals have created taxonomies or other controlled vocabularies, which they have made available for license.  Whether it’s worthwhile for them to promote taxonomies that are for license is uncertain. So, a year ago I created an online survey of taxonomy (or more broadly, any controlled vocabulary) licensing interest, which I announced not only on this blog, but also the blogs of taxonomy software vendors and at various conferences. The survey stayed open for about 6 months, and there were over 60 responses to most questions.  Now it is time to share those results. Although the responses are in the context of licensing controlled vocabularies, some of the questions and responses--about the taxonomy purpose, type or subject area of interest--might reflect general interest in taxonomies. (Percentages have been rounded.)

The first question asked about interest in licensing taxonomies or other controlled vocabularies. Slightly more than half of the respondents (61%) have considered licensing taxonomies, but most have not gone any further in identifying appropriate taxonomies to license. The leading reasons given not to from those respondents who said they would not likely to license a taxonomy (22 respondents out of 66), were:

  1. Custom-created taxonomies would best serve my purposes: 59%
  2. Licensed taxonomies that are modifiable and permit commercial reuse are too expensive: 14%

The leading concerns regarding licensing a taxonomy, ranked in order were the following:
  1. Difficulty finding or lack of a suitable taxonomy
  2. Difficulty integrating a licensed taxonomy into an existing taxonomy or taxonomy set
  3. Effort to modify, adapt, and/or expand a license taxonomy
  4. Licensing fee cost
  5. Features of the licensed taxonomy missing
  6. File format and implementation issues


The types of controlled vocabularies that respondents are most interested in licensing (allowing multiple responses) were:

  1. Hierarchical taxonomy: 56%
  2.  Controlled vocabulary for part of a faceted taxonomy: 55%
  3.  Ontology: 40%
  4.  Thesaurus: 35%
  5.  Name authority file (companies, places, organizations, person names, etc.): 17%
  6.  Classification scheme (such as with alpha-numeric codes): 10%


The subject areas of controlled vocabularies that respondents are most interested in licensing (allowing multiple responses) were:

  1. Business/management/enterprise functions: 36%
  2.  Information technology/computing: 30%
  3.  Industries: 28%
  4.  Company or organization names: 26%
  5.  Products/services: 23%Health/medicine: 21%
  6.  Geographic places: 21%
  7.  Engineering & design: 20%
  8.  Law & policy: 20%
  9.  Science & math: 18%
  10.  Humanities & social sciences: 13%
  11.  Occupations or job titles: 13%
 Finance was a popular write-in option under “Other.”


 
The purposes that respondents said a licensed controlled vocabulary would serve (allowing multiple responses) were:

  1. Internal content management and search & retrieval: 82%
  2. Business intelligence/market research/competitive intelligence/data analysis: 32%
  3. Expertise identification: 24%
  4. Public/website content findability – commercial: 21%
  5. Education/research: 19%
  6. Ecommerce or B2B: 18%
  7. Public/website content findability – nonprofit: 15%
  8. Public/website content findability – government: 8%


The size ranges of a controlled vocabulary that respondents said they would be interested in licensing (allowing multiple responses) were:

  1. 1,000 - 5,000 concepts: 33%
  2.  More than 10,000 concepts: 26%
  3.  500 - 1,000 concepts: 21%5,000 - 10,000 concepts: 21%
  4. Less than 100 concepts: 17%
  5.  100 - 500 concepts: 14% 
 
The formats of a controlled vocabulary that respondents said they would be interested in licensing (allowing multiple responses, especially since some of these formats are not mutually exclusive) were:
  1.  XML: 44%
  2.  Unsure:39%
  3.  Excel (xls or xlsx): 34%
  4.  SKOS: 32%
  5.  CSV: 26%
  6.  RDF: 24%
  7.  JSON: 24%
  8.  OWL: 16%
  9. Turtle: 11%
  10. Z Thes: 8%


The leading industries of respondents were:

  1. Consulting/professional services: 18% (Perhaps taxonomy consultants, like me?)
  2. Nongovernmental/nonprofit: 18% (Perhaps because licensing restrictions for commercial re-use are not an issue.)
  3. Software/Hardware/IT: 13%
  4. Manufacturing/Construction/Engineering: 10%

Additionally, 10 other individual industries were indicated with only 2-3 individual responses each.

Conclusions from the survey include:

  • Concerns around licensing are shared, and there is no dominant single concern.
  •  Hierarchical taxonomies and vocabularies for facets of faceted taxonomies are the types most of interest.
  • The subject area of greatest interest is business/management/enterprise functions.
  •  Internal content is the leading purpose for controlled licensing.
  • Size of vocabularies of interest includes all, but the mid-range dominates.
  • Industries interested in vocabulary licensing vary, and none dominates.
  • XML and CSV/Excel or the formats of greatest interest, but a significant number are unsure of format desired.