A survey of Top-Level Ontologies To inform the ontological choices for a Foundation Data Model Version 1 -- 1 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 2 Contents 1 Introduction and Purpose 3 2 Approach and contents 4 2.1 Collect candidate top-level ontologies 4 2.2 Develop assessment framework 4 2.3 Assessment of candidate top-level ontologies against the framework 5 2.4 Terminological note 5 3 Assessment framework – development basis 6 3.1 General ontological requirements 6 3.2 Overarching ontological architecture framework 8 4 Ontological commitment overview 11 4.1 General choices 11 4.2 Formal structure – horizontal and vertical 14 4.3 Universal commitments 33 5 Assessment Framework Results 37 5.1 General choices 37 5.2 Formal structure: vertical aspects 38 5.3 Formal structure: horizontal aspects 42 5.4 Universal commitments 44 6 Summary 46 Appendix A Pathway requirements for a Foundation Data Model 48 Appendix B ISO IEC 21838-1:2019 – Documenting Coverage 51 Appendix C Coverage Mapping to Assessment Framework 55 Appendix D Candidate source top-level ontologies – longlist 58 Appendix E Summary of Framework Assessment Matrix Results 63 Appendix F Selected candidate source top-level ontologies – details 68 F.1 Introduction 68 F.2 BFO – Basic Formal Ontology 68 F.3 BORO 73 F.4 CIDOC 76 F.5 CIM 79 F.6 ConML + CHARM – Conceptual Modelling Language and Cultural Heritage Abstract Reference Model 81 F.7 COSMO – COmmon Semantic MOdel 83 F.8 Cyc 84 F.9 DC – Dublin Core 85 F.10 DOLCE – Descriptive Ontology for Linguistic and Cognitive Engineering 88 F.11 EMMO 89 F.12 FIBO – Financial Industry Business Ontology 91 F.13 FrameNet 92 F.14 GFO – General Formal Ontology 94 F.15 gist 95 F.16 HQDM – High Quality Data Models 97 F.17 IDEAS – International Defence Enterprise Architecture Specification 99 F.18 IEC 62541 100 F.19 IEC 63088 100 F.20 ISO 12006-3 101 F.21 ISO 15926-2 102 F.22 KKO: KBpedia Knowledge Ontology 103 F.23 KR Ontology – Knowledge Representation Ontology 105 F.24 MarineTLO: A Top-Level Ontology for the Marine Domain 106 F. 25 MIMOSA CCOM – (Common Conceptual Object Model) 108 F.26 OWL – Web Ontology Language 110 F.27 ProtOn – PROTo ONtology 111 F.28 Schema.org 112 F.29 SENSUS 113 F.30 SKOS 113 F.31 SUMO 115 F.32 TMRM/TMDM – Topic Map Reference/Data Models 116 F.33 UFO 120 F.34 UMBEL 121 F.35 UML 121 F.36 UMLS – Unified Medical Language System 124 F.37 WordNet 125 F.38 YAMATO – Yet Another More Advanced Top-level Ontology 126 Appendix G Prior ontological commitment literature 128 Appendix H Criteria for a good scientific theory 129 Appendix I Detailed notes on 3.2.1 Basis 130 I.1 Simplicity 130 I.2 Explanatory sufficiency 130 Appendix J Ontological commitments – technical details 131 J.1 Natural language ontology – foundational ontology 131 J.2 Extensional and intensional criteria of identity 134 J.3 Indexicality 137 Appendix K Glossary 138 References 139 Acknowledgements 141 About the Construction Innovation Hub 142 -- 2 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 3 1 Introduction and Purpose The Centre for Digital Built Britain has been tasked through the Digital Framework Task Group to develop an Information Management Framework (IMF) to support the development of a National Digital Twin (NDT) as set out in “The Pathway to an Information Management Framework” (Hetherington, 2020). A key component of the IMF is a Foundation Data Model (FDM), built upon a top-level ontology (TLO), as a basis for ensuring consistent data across the NDT. This document captures the results collected from a broad survey of top-level ontologies, conducted by the IMF technical team. It focuses on the core ontological choices made in their foundations and the pragmatic engineering consequences these have on how the ontologies can be applied and further scaled. This document will provide the basis for discussions on a suitable TLO for the FDM. It is also expected that these top-level ontologies will provide a resource whose components can be harvested and adapted for inclusion in the FDM. Following the publication of this document, the programme will perform a structured assessment of the TLOs identified herein, with a view to selecting one or more TLOs that will form the kernel around which the FDM will evolve. A further report – The FDM TLO Selection Paper – will be issued to describe this process in late 2020. -- 3 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 4 2 Approach and contents The approach has three parts: 1. collect candidate top-level ontologies (2.1) 2. develop assessment framework (2.2) 3. assess candidate top-level ontologies against the framework (2.3) These are described in more detail below. A note on the terminology used in the report is contained in 2.4 and Appendix K. 2.1 Collect candidate top-level ontologies A long list of possible candidates for TLO content that might be useful for the construction of the FDM has been drawn up and reviewed. Candidate ontologies were identified both through extensive desktop research, and through the experience and domain knowledge of the expert community involved in bringing this report together. In identifying candidates, the net was thrown as wide as possible to identify as much useful content as possible. Thus, though the focus is on ontological commitment, the list includes data models that are generic in nature (ones without an explicit ontological foundation) as these are likely to have some useful ontological content. The candidates are listed in Appendix D and are available online within the IMF Developers Network on the Digital Twin Hub, www.digitaltwinhub.co.uk. 2.2 Develop assessment framework In compiling this report, a first-pass assessment framework was developed to facilitate the initial testing of the spectrum of available TLOs and other ontological models against the needs of the programme. This assessment is distinct from the future activities around the further down- selection of TLOs to a selected core for the FDM. An ontology is (according to Jonathon Lowe in The Oxford Companion to Philosophy) “the set of things whose existence is acknowledged by a particular theory or system of thought.” When we interpret a dataset, working out what the data refers to, we are acknowledging that the dataset commits to these things existing. We are committing to an ontology – making an ontological commitment. There are a variety of ways of making these commitments. The purpose of a top-level ontology is to enable us to make the choice of ontological commitments in an explicit and consistent way. There have been a series of attempts to get to grips with the kinds of choices of ontological commitments that ontologies can make. They provide a reasonable starting point but need substantial further work to provide a comprehensive framework for making choices across a broad range of commitments. We have developed a comprehensive choice component for this framework; one that is suitable for assessing information system TLOs. With this choice component in place, we can look at how these choices shape an ontology’s underlying architecture and so what it can do and how it does it. We can also characterise the candidate TLOs in terms of whether they make a choice and which they choose. -- 4 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 5 The assessment framework has three levels. 1. A general level, which looks at whether the TLO makes an ontological commitment and the strength of this commitment (see 4.1). 2. A formal level, which looks at how the formal structure of the TLO has been impacted by its ontological commitment (see 4.2). 3. A universal level, which looks at how the TLO addresses the individual universal choices (see 4.3). Further details on the basis for the assessment process are described in section 4. 2.3 Assessment of candidate top-level ontologies against the framework The assessment framework described above has been applied to the candidate TLOs. Only a small number of TLOs made their ontological commitments explicit, enabling a clear, simple assessment. In other cases, the choices could be clearly inferred from the documentation available. However, in many cases we could not determine the choice. It could be that no choice was intended, or that the choice is not documented, or not documented sufficiently clearly in the material we have reviewed or some other reason. Rather than trying to classify the exact reason for each case, where we have not noted a choice, we have marked the cell ‘not assessed’. In some cases, typically ISO standards, the documentation is not publicly available, so to make this clear we have marked these ‘not available’. The results of the assessment are stored online within the IMF Developers Network on the Digital Twin Hub, www.digitaltwinhub.co.uk. with a summary provided in section 6. In addition, a brief overview of the candidate top-level ontologies with any useful additional points is given in Appendix F. These include a graphical representation of the TLOs where one has been found. A comparison shows clearly the diversity of top-level structures. 2.4 Terminological note This is a topic that crosses multiple disciplines, including information systems, computer science, philosophy and linguistics. Confusingly, many terms are used with different senses across these disciplines and even within them. Accordingly, we will attempt, where possible, to use terms with the senses that they have in the disciplines in which they arise – to minimise any increase in the confusion and encourage cross- disciplinary consistency. There is one critical case where there is little consistency, this is a term for objects in general. These are sometimes also known as entities or things – through all three of these terms have restricted senses in various sub-disciplines. We propose to use the term ‘object’ here, rather than ‘entities’ or ‘things’ unless the context requires it as this is the term most consistently used in philosophy, and can be found with this sense as far back as the 17th century (Locke, 1975). Where needed we will qualify the term – for example, material objects. If we are using the term in a different sense, we will clearly note this. Of course, all three terms will appear in the extracts from the TLO documentation. Other terms are defined in the glossary in Appendix K. -- 5 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 6 3 Assessment framework – development basis The development of the assessment framework is driven by two broad considerations: general ontological requirements (3.1); and the requirement for an overarching ontological architecture (3.2) These are described below. 3.1 General ontological requirements The general requirements provide a context for the framework and comprise: 1. the need for an ontological framework (3.1.1); 2. how the need for an ontological framework translates into making choices of ontological commitments (3.1.2 and 3.1.3); 3. the requirements that arise from the lack of prior knowledge (3.1.4); and 4. the need for consistent independent (federated) development (3.1.5). 3.1.1 Real world ontology framework If one wants to share data from different systems, then one needs to have something like a common framework within which to share it. When the data is in this common framework, its meaning needs to be clear and unambiguous to the systems sharing it. This is often called semantic interoperability. For example, it needs to be clear and unambiguous whether data items (for example, rows on a table) from two systems are referring to the same object or different objects in the ‘real world’ – for example the UNICLASS Code ‘Ac_05_50_91 – Timber sourcing’ is marked as mapping to NBS Code ‘45-60-90/340 – Timber procurement’. To implement this systematically, one needs to be clear and unambiguous about what these objects in the real world are. This involves knowing the ontology, in other words, knowing the set of objects that the common framework assumes exist. 3.1.2 Choices of ontological commitments Unfortunately, when one starts to look closely, it is neither clear nor unambiguous exactly what the objects in the real world are. Ontologically, there are a variety of ways that one can take the real world to be. However, for our assessment, these can be crystallised into a small number of focussed choices – called ontological commitments – which build up into an integrated ontological architecture. A key purpose of this paper is to provide a framework for understanding the range and nature of the ontological commitments and apply this to the collected top-level ontologies. Thus, providing the groundwork for the choice of an appropriate ontological architecture. -- 6 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 7 3.1.3 Implicit and explicit choices Understandably, most of the datasets currently available are not clear about their ontological architecture, which of these ontological commitments have been made – their choices, such as they are, are implicit. In practice, datasets often make these choices implicitly – choosing, without realising, one way in one area and another way in another area. This point is often made in philosophy textbooks, see, for example, (Lowe, 1998). The assessment framework gives a clear picture of the range of these choices. With this in hand, when selecting or developing a top-level ontology, one can be clear which choices are made (and which left to chance); and so have some idea how this will be, in turn, reflected in the data structures and data that are implemented. 3.1.4 Lack of prior knowledge Usually the developers of a system of systems (which will include behavioural, societal and human elements) have prior knowledge of some of the systems that will use the common framework. However, often other requirements will arise as new systems are added to the common framework – of which it will have no prior knowledge. Hence, the framework needs to be sufficiently expressive to accommodate them. More specifically, care needs to be taken not to adopt ontological commitments which unnecessarily restrict its ability to express meanings that probably will occur in data from new source systems. For example, the commitment choices include whether to restrict the types to first order (where types cannot have types as instances). One cannot just assume that as the current data set only has first order types, then one can restrict oneself to these – one also needs some confidence that a requirement for higher order types will not emerge in the future. In this case, there is then a requirement to be sensitive to how a choice of ontological commitment might restrict useful expressivity. 3.1.5 Consistent independent (federated) development Many systems for connecting systems, like the NDT, have a hub and spoke structure where spoke systems map their data into the central hub system. It is likely that these mappings will be done independently. In many cases, the content from different systems will overlap. Where this happens, the mappings produced should be equivalent. Adopting an ontological approach is a big step towards achieving this because it provides an independent basis for establishing identity between the systems. Fine tuning the choice of ontological commitments to ensure a clear notion of what is referred to is a good further step. In this case, there is then a requirement to be sensitive to how a choice of ontological commitment can be clearer about what is referred to and so give rise to equivalent independent mappings. -- 7 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 8 3.2 Overarching ontological architecture framework As mentioned earlier, there have been several attempts to get to grips with the kinds of choices of ontological commitments that TLOs should make; these are listed in Appendix G. These attempts provide a good starting point and we refer to them when this is useful, usually in the technical appendices. However, all these lists are partial and, in some cases, not based upon sufficient familiarity with the relevant research. Furthermore, none of them provide an over-arching organising structure; one that provides a framework for understanding and assessing choices across a range of commitments. We develop this framework in 3.2.1 below. 3.2.1 Basis There needs to be a clear and solid underlying basis for the framework. There is an established set of criteria for assessing what a good ontology is, based upon what makes a good scientific theory – listed in Appendix H. One of these criteria, simplicity, provides us with a good basis for a broad assessment of the architecture and is broadly outlined below, with more technical detail in Appendix I. Simplicity can be thought of as having two aspects; structural and ontological. Where structural or syntactic simplicity is roughly concerned with the shape of the organising structure, ontological simplicity is roughly concerned with the number of objects. For structural (syntactic) simplicity we look at the characteristic ways the ontological commitments shape the organising structure (see section 4). For ontological simplicity we do some more analysis to establish broad brush accounting principles (as set out in 3.2.1.1 in 3.2.1.3 below). 3.2.1.1 Accounting for ontological simplicity Ontological simplicity is usually associated with a number of characteristics including parsimony, explanatory sufficiency and fruitfulness. Parsimony is usually characterised as Ockham’s Razor – objects are not to be multiplied beyond necessity. There is less well-known, but with an equally long history, principle of explanatory sufficiency – “the variety of entities should not be rashly diminished” (Kant, 1964). Parsimony and explanatory sufficiency, taken together, imply a kind of ontological economy, which aims for explanatory sufficiency with the minimum number of entities. We look at how to account for this first and then return to fruitfulness (see 3.2.1.3). Simple counting of objects does not seem intuitively correct way of accounting. If Ann claims that the damage to my carrot patch was caused by exactly 100 rabbits, and Ben claims it was done by 101 rabbits, then it is hard to feel that Ben’s theory is any way less economical than Ann’s, or that Ben has multiplied entities in a way that calls for concern. If Ann now claims the damage to my carrot patch was caused by 5 rabbits (one type and five individuals) and Ben claims it was caused by 1 deer and 3 rabbits (two types and four individuals) – then despite both claims involving six objects, Ben’s seems more complex. In the literature, this is seen as arising from a distinction between qualitative parsimony (roughly, the number of types of object) and quantitative parsimony (roughly, the number of individual objects). In the research, most people claim qualitative parsimony matters and quantitative parsimony is less relevant – making a distinction between the making of the commitment and its cost. If Ann now claims the damage was caused by 12 rabbits (one type and 12 individuals) and Ben claims it was caused by 1 deer and 3 rabbits (two types and four individuals), then Ann’s claim is more qualitatively parsimonious (one versus two types) but less quantitatively parsimonious (twelve versus four individuals) than Ben’s. The suggestion is that Ann’s claim is more relevantly economic than Ben’s. The claim that qualitative parsimony matters more than quantitative parsimony resonates for the design and maintenance information systems. For example, function or object point -- 8 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 9 analysis (e.g. FiSMA: ISO/IEC 29881 or IFPUG: ISO/IEC 20926:2009) measures are based upon qualitative (type) rather than quantitative (individual) counts. However, as has been noted, this qualitative-quantitative distinction seems too simplistic – for example, not taking account of algorithmic complexity. 3.2.1.2 The laser There is a revised approach that seems to capture some of the relevant complexity. This uses a distinction between fundamental and derived objects and updates Occam’s Razor with what Jonathan Schaffer (Schaffer, 2015) calls the laser – “do not multiply fundamental objects without necessity”. He illustrates the difference between the razor and the laser with this example. Imagine Esther posits a fundamental theory with 100 types of fundamental particle. Her theory is predictively excellent and is adopted by the scientific community. Then Feng comes along and—in a moment of genius—builds on Esther’s work to discover a deeper fundamental theory with 10 types of fundamental string, which in varying combinations make up Esther’s 100 types of particle. This looks like a paradigm case of scientific progress in which a deeper, more unified, and more elegant theory replaces a shallower, less unified, and less elegant theory. However, under razor accounting, both the number of particles and strings are counted and therefore Feng’s theory has 10 more objects and so should be replaced with Esther’s. Under laser accounting though, 100 fundamental objects have been replaced by 10 – so Feng has made an improvement. Here again we have a distinction between making a commitment and its cost. We make a commitment to both fundamental and derived objects, but the cost of derived objects is significantly less than that of fundamental objects. As Schaffer notes, what emerges from this approach is a general pressure towards a permissive and abundant view of what there is, coupled with a restrictive and sparse view of what is fundamental. As he notes, classical mereology (the relations of parts to wholes) and pure set theory (where the only sets, well-determined collections of objects, under consideration are those whose members are also sets) come out as paradigms of methodological virtue, for making so much from so little. This suggests a preference for, what has been called, plenitude – not placing unnecessary constraints on what can exist; if it is possible for something to exist, then it does. Both classical mereology and (impure) set theory exhibit this. Simplifying a little, in classical mereology, given any two objects, their fusion exists – in set theory, their set exists. Where many of the candidate TLOs make explicit their mereological position, they chose classical mereology. However, where they make explicit their position on types, only a significant minority adopt a position of plenitude. Schaffer suggests a principle to capture this, the Ontological Bang for the Buck principle: optimally balance minimization of fundamental objects with maximization of derivative objects, especially useful ones. -- 9 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 10 3.2.1.3 Fruitfulness In 3.2.1.1 above, fruitfulness was mentioned as being associated with simplicity. The examples provided above show, derivative objects are part of what makes a package of fundamental objects fruitful. In other words, they show that these fundamental objects can be used to produce something useful. However, as discussed, there is a need to be sensitive to both cost and benefits. If two very similar theories had roughly the same cost in terms of fundamental objects, but one had a large commitment to many useless entities but the other did not – and they were similar in all other relevant respects, this seems like overgeneration. The additional useless plenitude is more like profligacy or promiscuity – it is not fruitfulness. This gives us a ‘useful’ basis for assessing the TLOs. -- 10 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 11 4 Ontological commitment overview Our overview of the framework for ontological commitments is divided into three parts: 1. Section 4.1 looks at the general choices TLOs make on whether and what kind of overall ontological commitment to make; 2. Section 4.2 looks at the overall formal structure; and 3. Section 4.3 considers the individual core commitments that lead to that structure. More detailed technical notes are given in Appendix J. 4.1 General choices The general choices track the ontological approach chosen by the top-level ontologies. They firstly note whether the TLO has chosen to make ontological commitments or not. They then note whether the ontological commitment is lightweight or heavyweight (see 4.1.2). Finally, they note what they have chosen to make the subject of their ontological commitments; natural language or the (foundational) real world (this is discussed in 4.1.3). Our survey includes examples of TLOs making all these choices and this range provides useful examples to compare and contrast as well as a comprehensive range of components that could be useful in developing a TLO. 4.1.1 Ontologically committed: ontological or generic The top-level ontologies longlist was compiled to include any data models that might have content useful for the construction of a top-level ontology. Hence, one of the key conditions for inclusion is that the model must be sufficiently general to include content that might be useful. There are cases where the TLO specifies a data structure with no intended ontological commitment; to deliberately leave open how the data is modelled. A classic indication of this is where the modeller can validly choose which data type to use in a model (whether something is modelled as an entity or attribute) based upon, typically, performance requirements. These TLOs are classified as generic; Topic Maps and Schema.org are examples of this. One consequence of this choice is that these top- level ‘ontologies’ are not able to harness the interoperability benefits of adopting ontological commitments mapping to the real world (discussed above). 4.1.2 High or low ontological commitment Where TLOs are ontologically committed, the analysis reveals that some have explicitly committed to most, if not all, of the choices whereas others have only committed to a few. This gives us a good basis for distinguishing between the heavyweight TLOs that are highly committed and the lightweight TLOs that are only committed to a few. -- 11 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 12 4.1.3 Subject: appearance or reality: natural language or foundational ontology One can broadly classify top-level ontologies into two kinds by their subject matter. The subject matter can be what a community implicitly accepts when using a language – a natural language ontology. This will take the surface structure of the language, for example the distinction between nouns and verbs (or the words it uses), as a window on the ontology. Or it can be an ontology of what ‘really’ exists according to science (and philosophy) – a foundational ontology. This is suspicious of the surface structure of the language as it has often turned out to be a false friend. For example, the English language classes tomatoes as a vegetable, but this has not persuaded botanists to stop classifying them as ‘really’ a fruit. The natural language ontology may include merely conceived objects as well as those that happen to be actual – whereas the foundational ontology should only include actual ones, as well as some infrastructure to help assure that they are actual. The natural language ontology may focus on the linguistic structure of the language or on the concepts implied by the language. See Appendix J for a more detailed background. Some TLOs on our longlist have explicitly stated their aspirations to one or other kind of subject matter – and provided the appropriate infrastructure. Clear examples are DOLCE as a natural language ontology; BFO and BORO as foundational ontologies. Some TLOs with no stated aspirations are clearly focussed on language and its linguistic infrastructure (nouns, verbs, etc.) and thereby categorised natural language ontologies: Wordnet and FrameNet are good examples. Sometimes it is difficult to make the case for a classification; where there is neither a clear statement of intent nor clear infrastructure for one or other approach. These have not been classified. Where one’s focus is on the language used in a community, then, other things being equal, a natural language ontology makes a better fit. If the focus is on reality, then a model of what really exists makes more sense. However, both kinds of ontology should be considered as useful sources for components for a TLO. 4.1.4 Categorical There is a long tradition of categorical ontologies, where the types of the ontology are meant to be comprehensive, covering all types of thing that can exist – or, at the very least, a broad swathe; Aristotle and Kant’s Categories are historic examples of this. In principle, one would expect a TLO to be categorical. However, in practice, the comprehensiveness is often limited. In some cases, the scope is limited to a broad family of domains – such as MIMOSA’s focus on asset information for machinery and systems. In other cases, the TLO adopts a cautious position and its ontology is explicitly left open-ended allowing for extensions that involve new top-level categories, so making the TLO technically non-categorical – BFO is an example of this. -- 12 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 13 4.1.5 General classifications In summary, the general level has the following classifications: Figure 1 provides a visual summary of this. category type choice general ontologically committed ontological or generic general commitment level high or low (heavyweight or lightweight) general subject foundational or natural language general categorical yes or no Figure 1 – General classification of the TLOs -- 13 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 14 4.2 Formal structure – horizontal and vertical Many, if not most, of the ontological choices leave their mark on the formal structure, the ontological architecture, of the TLOs in characteristic ways; this section is about two ways we use these marks to classify them – which we tag vertical and horizontal aspects. Three core hierarchical relations – each usually visualised upwards – provide a backbone to the TLOs. The ontological choices shape these upwards (vertical) structures in various ways – and we use these ways to characterise the impact of the choices on the TLOs on the ontological architecture. If one looks in more detail at one of these hierarchies – the super-sub-type hierarchy – then one can see a repeating pattern of stratification across the hierarchy (horizontal) that mark particular choices. We use these to determine whether the TLO has made a particular choice. These two ways of looking at the formal structure (the ontological architecture) – tagged vertical and horizontal aspects – map neatly onto the simplicity basis introduced above. There we divided simplicity broadly into structural and ontological – roughly the shape of the organising structure and the number of objects. The vertical aspect deals with the structure, and so structural simplicity, of the various hierarchies; roughly their shape up and down. The horizontal aspect deals with the broad ontological choices that can introduce a division across the hierarchy (horizontal stratification) – which impacts the ontological simplicity – as they increase the number of objects. Taking a broad-brush view, this section of the framework separates the vertical and horizontal aspects of the formal hierarchies. In practical terms, these characterise the formal structures that arise in the ontological architecture from the various ontological choices. Together these form the backbone of the ontological architecture upon which the flesh of the ontology is built. Here we identify the broad structures and outline how their component ontological commitments fit into this structure. In the next section, we look inwards into the specific details of the commitments – rather than outwards at their impact of the structure. -- 14 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 15 4.2.1 Vertical aspect – varieties of hierarchies There is a core of basic ontological hierarchical relations that are typically found in top-level ontologies; whole-part, type-instance and super-sub-type (they go by various names, these are the ones we adopt in this paper – Table 1 lists some alternatives with examples). There is debate about whether some of these are fundamental (for example, super-sub-type can be defined in terms of type-instance). There is also debate whether these all belong to the same family of relations or are distinct types. Whatever the outcome of these debates, as noted earlier, in practice these hierarchies are a key part of the backbone of the ontological architecture. One powerful way they do this is through their formal structure. Here we look at the formal properties of hierarchies and how these apply to the three relations – to see how they, together, help shape the ontological architecture. We outline the properties below and consider their relevance to the three relations. These three relations normally manifest as hierarchies; in other words, they have the structure of a partially ordered set (or in the case of type-instance, it’s cover relation, as it is not transitive). They are standardly represented in an obvious way in Hasse diagrams (sometimes known as upward diagrams) as a directed acyclic graph – nodes connected by arrows that have no cycles – see Figures 1 to 6 below. We use these diagrams to show the formal structures we are examining. The relations have a conventional direction, given in Table 2. Table 1 – Hierarchal relations – terms Adopted term Alternative terms Examples whole-part part of This building has a whole-part relation to my front door (my front door is part of this building) type-instance instantiation, class-member, member, instance Building has a type-instance relation to this building (this building is an instance/member of the type building – this building is a building) super-sub- type generalisation, subsumption, super-type, sub-type Opening has a super-sub-type relation to door and window (door and window are sub-types of opening – doors and windows are openings) -- 15 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 16 There are some TLOs – such as Entity-Attribute- Relation (the original Chen version) – where there are no super-sub-type relations. This is often found in the physical implementation – SQL being a clear example. Further, it is often the case that whole-part relations are not explicitly marked, in other words, separated out from other relations. Typically, TLOs will have a range of choices on how they constrain the three hierarchies – 4.2.1.1 to 4.2.1.8 identify the relevant choices which we use below to analyse the TLOs. Often, as noted earlier, the underlying question is whether these constraints breach an explanatory sufficiency (plenitude) principle and so unnecessarily limit expressiveness. As always, there is often various factors in play, so the decision is not clear cut. 4.2.1.1 Parent-child-arity In hierarchies, the number of parents a node can have is the parent-arity, the number of children the child-arity (note that the number of parents may differ from the number of ancestors). In some TLOs some of the relations have their parent-arity or child-arity limited to one. This changes the structure from lattice-like to tree-like (see Figure 2). Table 2 – Conventional directions relation upwards direction whole-part part-to-whole type-instance instance-to-type super-sub-type subtype-to-supertype Figure 2 – Parent-child-arity structures in Hasse diagram format -- 16 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 17 The way these constraints are typically applied to the three relations is outlined in Table 3; which shows that the relevant choices are for type-instance and super-sub-type. In the object- oriented modelling community, these are known respectively as single or multiple classification and single or multiple inheritance. Constraining the hierarchy to a single parent is prima facie parsimonious. However, it also seems prima facie explanatorily insufficient. Why should a type such as mare not have female and horse as its supertypes? Why should an individual such as Donald Trump not be an instance of the types ‘human being’ and ‘biologically male’? These choices for single parent structures look likely to be less than optimal unless there are other factors counting in their favour. 4.2.1.2 Super-sub-type – transitivity We focus here on the transitivity of the super- sub-type relation. Type-instance is generally considered not to be transitive and whole-part to be transitive. The super-sub-type relation can be found in the early logic, in, for example, Aristotle’s syllogisms – in the assertion that ‘Every S is P’ (Every Human is an Animal). This can be translated into ‘S is a sub-type of P’ (Human is a sub-type of Animal). Its explicit recognition as a relation came with the nineteenth century mathematization of logic by Boole and others. Implicit visualisations of it as containment appear in Euler and Venn circles. However, it is more commonly visualised now as a hierarchy diagram – where the links represent instances of the sub-type relation. The majority of the graphic representations of the TLOs (in Appendix F) are super-sub-type hierarchy diagrams. However, some care needs to be taken when interpreting these diagrams – as they only show the cover relation (parents and children with no ancestors or descendants). The traditional semantics of super-sub-type is transitive: if every B is A and every C is B, then it seems clear that every C is A. So, ancestors or descendants are automatically included (though not shown in the diagrams). In some TLOs, their version of super-sub-type is not transitive – ancestors or descendants are not automatically included. Typically, only the links shown in the hierarchy are deemed to exist. The UML TLO is an example of this, which is most likely driven by implementation rather than semantic concerns. This choice should be made explicit. Table 3 – General parent-child-arity relation general direction-arity general parent-arity general child-arity whole-part always unconstrained always unconstrained type-instance single or unconstrained always unconstrained super-sub-type single or unconstrained always unconstrained -- 17 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 18 A similar kind of interpretation situation occurs when the super-sub-type hierarchy diagram only shows the relevant types. This is commonplace where the TLO supports extensional types. In this case, where a super-sub-type hierarchy diagram shows B as a sub-type of A, one cannot automatically infer that B is a child sub-type of A – as there may be intervening sub-types. An extensional TLO allows any collection of objects to be a type. If there is more than one instance of A that is not an instance of B – for example c and d in Figure 3, then there is a type C which has the members of B plus c as its members. C is not identical to either A or B as it has c but does not have d as a member. C is however a super-type of B and a sub-type of A This hopefully illustrates how ontological commitments impact upon the interpretation of these diagrams. This choice is dealt with under the formal generation section. 4.2.1.3 Boundedness Hierarchies as ordered relations might or might not have a top or bottom. As shown in Figure 2, the hierarchy is bounded if all of the maximal paths terminate, and unbounded if any maximal path does not terminate, though, as the diagram also shows, some may. The hierarchy is upwards bounded if all the maximal paths terminate upwards, and the set of terminating nodes are the top elements of the hierarchy. The hierarchy is downwards bounded if all the maximal paths terminate downwards, and the set of terminating nodes are the bottom elements of the hierarchy. These four options permute into four possible configurations. One of them, upwards and downwards bounded, opens up the possibility of further constraining the hierarchy to a finite number of levels. The interesting hierarchical relation for us, in the top-level ontologies we have reviewed, is type-instances. The first interesting case for us is firstly, whether type-instances is downwards bounded – the left-most case in Figure 4. Figure 3 – Intervening subtypes -- 18 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 19 Type-instance downwards boundedness is associated with the universals-particulars division that goes back to the Ancient Greek Aristotle, and beyond; where universals have instances, but particulars do not (another way of defining bottom). All seriously ontologically committed top-level ontologies make this division. Some of the generic ontologies have meta-models that place no constraints on the hierarchy. An example would be OWL’s use of punning; it does not identify a bottom level, so it is always possible to extend downwards. The Topic Map Reference Model is similarly unconstrained. A pragmatic argument for this lack of constraint is that it is too onerous to build the bound into the model at design time, and that it is more useful to let the users at runtime decide on whether or not to extend the hierarchy down a level. This then places the onus on the users to ensure the quality of the boundary, to ensure, for example, that something that is clearly an individual, such as Donald Trump, has no instances. Figure 4 – Possible boundedness options in Hasse diagram format Figure 5 – Fixed level boundedness -- 19 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 20 Then if type-instance is also, upwards bounded – the left-most case in Figure 5; there is a choice as to whether it is finitely bounded to a fixed number of levels – and if so, to how many levels. If TLOs are so constrained, they are often fixed to either three or two levels. Figure 5 has examples of two and three levels. OMG’s Meta Object Facility (MOF) is a case whether there are four. The way these constraints are typically applied to the type-instance relations is outlined in Table 4; the boundedness choices for the other two relations (whole-part and super-sub-type) are not sufficiently interesting to make it to the framework. relation downwards bounded fixed finitely bounded fixed number of levels type-instance can be unbounded or bounded fixed finitely bounded or not often two or three, but can be more Table 4 – Type-instance boundedness options 4.2.1.4 Intransitive vertical stratification Type-instance is intransitive – for example, if a is an instance of type b and b is an instance of type c – it does not follow that a is an instance of type c. This allows us to distinguish between cases where a node’s descendants could have its ancestors as parents – and where it does not (for transitive relations, it is always the case). This affects the structure and is visible in the hierarchy’s Hasse diagram as illustrated by Figure 6. Also, stratified hierarchies can be ranked – also illustrated in Figure 6. Another way of characterising this is as a distinction between hierarchies where each new rank can only be constructed, or based upon, the components of the previous rank (stratified) – and ones that can be constructed from all earlier ranks (unstratified). Of course, in cases of two levels, the previous rank is all earlier ranks, so it is a limit case of stratified. The standard technical terms for these are stratified and unstratified respectively – so we use them. This vertical stratification is a different sense of stratified from the one used in horizontal stratification; it is worth paying attention to the different senses. -- 20 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 21 Ontologies that are defined using meta-models and meta-meta-models, such as UML and MOF, are typically stratified. Extensional ontologies, such as those based upon BORO and IDEAS, are usually unstratified. This distinction is much discussed in the mathematics of set and type theory, where sets are unstratified and types stratified. The stratified approach, by some measures, is less structurally simple as it involves more restrictions. Also, as one can see from the Figure 6, the stratified approach is ontologically parsimonious and the unstratified approach plenitudinous. The key question is which provides more relevant expressiveness. This naturally leads to questions about what motivates the stratification restriction – there does not seem to be a good ontological answer for this. 4.2.1.5 Formal generation Ontology models are typically built through the careful manual addition of references to objects in the model. However, this is not the only way the structure in the model is created. Some top-level ontologies include algorithms for automatically adding new objects to the model. This is known as formal generation, as there are formal (algorithmic) rules for the generation. If we want a rounded picture of the structure, we need to consider this formal generation as well. There are a variety of types of algorithm that can be adopted. For our broad-brush picture, we just consider two core cases here: fusion and complement, the first an upwards (parent) generation the second a downwards (child) generation. This is enough to give a measure of the generative approach. Fusion is where one is given two objects of the right kind and then one can infer the existence of their parent fusion. Classic cases are mereological fusion and the pairing axiom in set theory – a whole that consists exactly of two or more particulars. Complement is where, when one is given a parent and one of its children, one can infer the existence of another child that is the rest of the parent. As Figure 7 shows, the formal generation produces both the object and its hierarchical relation(s). Figure 6 – Ranks: vertically stratified and unstratified -- 21 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 22 As Table 5 – formally generative options shows, there are choices for both these modes of generation for all three relations except for type-instance and complement. To see why type-instance is an exception, consider a singleton type A = {a}. It has the single instance a, but there is no complement of a. Figure 7 – Two kinds of formal generation relations formally generative fusion complement whole-part yes or no yes or no type-instance yes or no typically, no super-sub-type yes or no yes or no Table 5 – formally generative options Adopting formal generation for each of the three relations can be seen as examples of plenitude; all possible applications of the rule are automatically allowed. But this raises questions of whether there is overgeneration (see discussion on Basis in 3.2.1 above). Could these be examples of profligacy and promiscuity? Given the inter-related nature of TLOs, this assessment needs to be done in the context of all the choices made by the TLO. However, a prima facie case for them can be made on the basis of their close association with the two standard examples of ontological economy, classical mereology and set theory. -- 22 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 23 4.2.1.6 Relation class-ness The concept of first- and second-class objects was introduced by Christopher Strachey in the 1960s (Strachey, 2000). A second-class object is one that is not given the same ‘rights’ as other objects, a first-class object has the same rights as other objects. The particular right we are considering here is whether an object can be an instance of a type. If the three core relations are first-class objects, then they can be. Though the details differ slightly for the three relations, in each case, this allows for a (type-instance) link in our Hasse diagrams that starts with a link – as shown in Figure 8. The resultant graph structure is known as a hyper-graph. Whole-part relations are usually first-class. Type-instance and super-sub-type are often second-class but can be first-class Figure 8 – First-class relations Hasse diagram Imposing second-class-ness on these relations is a restriction on plenitude – as it introduces a block on the existence of possible objects. Hence, recognising all three relations as first-class would be an example of plenitude; all possible applications of the type-building rule are automatically allowed. The question then arises whether this plenitude veers into profligacy and promiscuity. It turns out that the standard pattern for classification, as used by Linnaeus in his taxonomy, needs these first-class objects (see Formalization of the classification pattern (Partridge, 2016) and Business Objects (Partridge, 1996). Table 6 – Relation class relation class whole-part usually first type-instance often second, but can be first super-sub-type often second, but can be first -- 23 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 24 4.2.1.7 Vertical structures These vertical classifications are shown below. Figure 9 provides a visual summary of this. type relation characteristic choice parent-arity type-instance single or unconstrained parent-arity super-sub-type single or unconstrained boundedness type-instance downwards bounded or unbounded boundedness type-instance fixed finite levels fixed or not fixed boundedness type-instance number of fixed levels [a number] (vertical) stratification type-instance stratified or unstratified formal generation whole-part fusion yes or no formal generation whole-part complement yes or no formal generation type-instance fusion yes or no formal generation super-sub-type fusion yes or no formal generation super-sub-type complement yes or no relation class-ness type-instance first- or second-class relation class-ness super-sub-type first- or second-class -- 24 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 25 4.2.1.8 Other vertical structures There are a number of other vertical structures that have not been included as they are less relevant. These include: • connectedness • restricted single type-instance parent-arity – single classification. Connectiveness It is not necessarily the case that any two nodes in the graph are connected. If some nodes are not connected, then the graph is disconnected. In this case, the disconnected graph can be divided into connected graphs. See the section below on possibilia for an application. While the connectedness of the structure is important, there is no overall pattern in these core relations that allows us to broadly characterise the owning TLO. Single classification There are some groups of types where one would expect the instances to belong to only one type – in other words, the types partition their instances. Quantities are an example. We would not expect something to have two masses, to both weigh 5 kg and 10 kg – it weighs one or the other. Similarly, for qualities, we would not expect an object to be both coloured and transparent at the same time, it has to be one or the other. This kind of restriction is common in TLOs, but again there is no overall pattern that allows us to broadly characterise the owning TLO. Figure 9 – Visual summary of the vertical aspects -- 25 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 26 4.2.2 Horizontal aspects: stratification versus unification There is a group of fundamental choices that impact the ontological architecture which involves whether or not to make a distinction. If one chooses not to make the distinction, one only introduces a single type. If one chooses to make the distinction, one introduces two types; one for each alternative. The choice boils down to whether to horizontally stratify or unify. One can describe choosing to make the distinction as ‘separating one potentially unified type into two’, creating a horizontal stratification in the hierarchy – and not making the distinction, ‘unifying the potentially separated two types into one’. These choices are perhaps best explained by looking at the specific cases (see 4.2.2.1 to 4.2.2.8). We only consider the major cases relevant to our review of the TLO candidates. We focused on identifying the formal choice – whether to horizontally stratify or unify – and leave the other aspects driving the decision to the more detailed description later in the report (see 4.3). We have mostly described the choices from a unifying perspective, they could equally well have been described from a stratifying perspective. 4.2.2.1 Spacetime We start with one familiar from 20th century physics. Prior to then, it was assumed that the spatial geometry of the universe was independent of one-dimensional time. There were two related but independent types, spatial regions (regions of space) and temporal regions (regions of time). The work of Einstein and Minkowski introduced the idea – which became accepted – that space and time could be fused into spacetime. Ontologically, this can be seen as unifying spatial regions and temporal regions into spatio-temporal regions whose instances are regions of spacetime. Figure 10 provides some examples from the TLOs. BFO is an interesting case as it hyper-separates, it separates but keeps the unifying type. This raises interesting ontological accounting questions about whether this is overgeneration (as so profligate) or interesting plenitude. -- 26 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 27 While space, time and spacetime may appear to be familiar notions for interpreting data, it turns out to be a tricky area to tie down formally. It takes some study to develop a clear idea of what the spacetime stratification choice here implies. We can illustrate the choice simply as between a 1D time plus 3D space or a 4D spacetime – as shown in Figure 11 (based upon Figure 1 in (Gilmore, 2016)) Figure 10 – Separating spacetime – TLO examples Figure 11 – Separation and unity of instances -- 27 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 28 In many cases, as here, the stratification (or unification) of the types also implies a separation or unity of their instances. To appreciate the consequences of the choices, Figure 11 shows how three (simple) instances of locations end up under the two regimes – this is recapitulated in Table 7. Table 7 – Three locations location 1D time plus 3D space 4D spacetime instant (in time) a (simple) point on the 1D timeline a (complex) horizontal slice in spacetime a point of space a (simple) point in 3D space a (complex) vertical line in spacetime a spacetime point A point in 3D space located at a point in 1D time a (simple) point in 4D space An under-appreciated consequence of the separation is that 3D space is multiply located – it is located as a whole at each instant of 1D time. Typically, a TLO will choose to stratify or unify the types and so separate or unify the instances. However, it can attempt to do both. As noted earlier, the TLO BFO provides us with an example. It opts to include both spacetime as well as space and time. This raises interesting semantic redundancies analogous to data redundancy. And so, a requirement that the spaces and times need to be coordinated with spacetimes – one can regard this as a kind of semantic or ontological denormalization analogous to database denormalization. 4.2.2.2 Locations People often talk of physical objects and their locations, where physical objects occupy their locations, suggesting two related types; objects and locations (let’s leave the decision whether the location is spatial, temporal or spatiotemporal to the previous choice). For example, “today your car is parked in the same place as mine was yesterday” could be regarded as a location which was occupied by my car (a physical object) yesterday and your car today. There is a debate going back to Newton and Leibnitz in the 17th century as to whether location is absolute or relative. If it is relative, then location is clearly fundamentally different from physical objects – which aren’t. However, if it is absolute, a kind of substance, then this opens the possibility that one could unify objects and their locations as fundamentally the same, technically known as supersubstantivalism. If one does not have cases of interpenetration (see 4.3.3) then this resolves the oddity where physical objects exactly occupy a single location throughout their life – unifying eliminates this double- counting. If there is interpenetration, then two objects may collapse to the same location – which may have unintended consequences. After the unification, the physical object and its location, two kinds of substance, are replaced by a single supersubstantival object. One has a broad choice between separating or unifying physical objects and locations. -- 28 of 143 -- A survey of Top-Level Ontologies constructioninnovationhub.org.uk 29 4.2.2.3 Properties In language, there is a distinction between nouns and adjectives; between rose and red. This has been taken as an indication of a more fundamental distinction, between what are known as substances and the properties or qualities they bear, for example, where a red rose would have a rose substance that bears a red property/quality. However, in other contexts, such as a Venn or Euler diagrams, we would be happy to have overlapping circles for roses and red objects – with a single icon for each red rose in the overlap. This implies there are instances of the type object that belonged to both the lower level types rose and red. So, here is a choice between stratifying to substances as the bearers of properties or unifying to objects. 4.2.2.4 Endurants Philosophers have noted that people say some types of object, such as stones and chairs, exist; whereas other types, events, occur or happen or take place. It is suggested that this marks a fundamental distinction between continuants and occurrents – where occurrents (that occur) are the events that happen to continuants (that exist). There is a competing view that there is no fundamental difference between the two, rather a different perspective on the same object – this has been labelled perdurantist. A classic example (for perdurantists) is glaciers. From a day to day perspective they are solid, unmoving material objects – they exist and so could be classified as continuants. From a geological perspective, glaciers flow – they are events like the flowing of a river – so could be classified as occurrents. A common continuant/occurrent stance, is that there are two glaciers; the existing continuant and the flowing occurrent. A perdurantist stance would be there is a single perdurants object which can be looked at from two perspectives. 4.2.2.5 Immaterial Philosophers have suggested that the hole inside a doughnut is different from the doughnut. The doughnut is composed of stuff whereas the hole is not composed anything – it is defined by the doughnut. They suggest making a stratification where the doughnut is a material object and its hole is an immaterial object dependent upon the material doughnut. A unifying stance would not regard this distinction as fundamental and not recognise material and immaterial are fundamental types in its ontology. The hole has a spatial extent that contains matter, though this matter may well change over time. In a sense it is ‘immaterial’ what matter is in the hole, but this does not, by itself
BORO Publications
A survey of Top-Level Ontologies:
To inform the ontological choices for a Foundation Data Model
17 November 2020Published in Centre for Digital Built Britain, Report, version 1, Cambridge, 2020
Overview
The Centre for Digital Built Britain has been tasked through the Digital Framework Task Group to develop an Information Management Framework (IMF) to support the development of a National Digital Twin (NDT) as set out in “The Pathway to an Information Management Framework” (Hetherington, 2020). A key component of the IMF is a Foundation Data Model (FDM), built upon a top-level ontology (TLO), as a basis for ensuring consistent data across the NDT.
This document captures the results collected from a broad survey of top-level ontologies, conducted by the IMF technical team. It focuses on the core ontological choices made in their foundations and the pragmatic engineering consequences these have on how the ontologies can be applied and further scaled. This document will provide the basis for discussions on a suitable TLO for the FDM. It is also expected that these top-level ontologies will provide a resource whose components can be harvested and adapted for inclusion in the FDM.
Following the publication of this document, the programme will perform a structured assessment of the TLOs identified herein, with a view to selecting one or more TLOs that will form the kernel around which the FDM will evolve. A further report – The FDM TLO Selection Paper – will be issued to describe this process in late 2020.