Extending the design space of ontologization practices: Using bCLEARer as an example STIDS 2024 Semantic Technology for Intelligence, Defense, and Security Wednesday October 23rd, George Mason University, USA Chris Partridge, Andrew Mitchell, Sergio de Cesare, John Beverley STIDS 2024 Two-dimensional analysis Baselining current practices AaE – Ask an Expert TDC – Top-Down Classification Comparison bCLEARer Product and process Digitalization process factorization Setting the right context - evolution bCLEARer accounting example Summary 2 Abstract Our aim in this paper is to suggest that the design space for the ontologization process is richer than current practice would suggest. And so that it is possible to open the space to a range of radically new practices. This consciously builds upon the notion that engineering processes as well as products need to be designed. We provide evidence for the new practices from our work over the last three decades with an outlier methodology, bCLEARer. We also provide some contextual scaffolding for a perspective that we have found we needed to better understand the nature of these new practices. This is an evolutionary perspective which sees digitalization (the evolutionary emergence of computing technologies) as part of the latest step in a long evolutionary trail of information transitions. This reframes ontologization as a tool for exploiting the emerging opportunities offered by digitalization. 3 Two-dimensional analysis Design Space – in two dimensions 5 metadata schema data brain speech writing printing computer levels of generality early-stage process - levels of digitalization two-dimensional design space - exploitation levels of generality One can broadly divide information into levels of generality. From a syntactic ‘data’ perspective, these are the natural levels: metadata, schema and data. For our purposes here, these are a good enough rough proxy for the semantic levels to make the substitution fair for our broad classification top, the most general - categories, middle – universals/types and bottom level - the most specific – typically particulars. We need the syntactic perspective we often start with ‘raw’ structured information where this is simple to identify. 6 levels of digitalization These are the levels on the journey to digitalization Looks at the evolutionary steps on the journey to digitalization Very broadly a journey that goes: from brains to speech to writing to printing and then computing We call this ‘levels of digitalization’ 7 brains speech computing printing writing Information pathway An information pathway refers to the flow and transformation of data or knowledge from its source to its intended destination. It can occur within a system (such as a human brain, an organization, or software application) or between multiple systems. It also occurs in an ontologization process We look at the way the ontologization process engages in the information pathway with the two-dimensional levels over time. 8 Two-dimensional analysis 9 metadata schema data brain speech writing printing computer levels of generality early stage process - levels of digitalization exploited unexploited two-dimensional design space - exploitation When we review current practices, we see that not all the space is exploited. NOTE: for levels of digitalization we focus on the early stages of the process Exploiting the new space involves a ‘shift left’ 10 Smith, L. (2001). "Shift-left testing". Dr. Dobb’s Journal, 26(9), 56–ff. Bahrs , P. (2014). Shifting Left - Approach and Practices. https://www.slideshare.net/Urbancode/shift-left Firesmith , D. (2015). "Four types of shift left testing". Podcast, Software Engineering Institute Website, September. Exploiting the space involves shifting the introduction of data (generality) and computing (digitalization) to as early in the project as feasible. Baselining current practices Work with two baseline methodologies The ‘Ask-an-Expert’ (AaE) approach based upon OntoCommons report D.4.2 - ‘ Methodological framework for ontology management ’, documents a multiplicity of methodologies, including: LOT, METHONTOLOGY, On-To-Knowledge, DILIGENT, NeOn, RapidOWL, SAMOD and AMOD together these provide many good examples of the ‘ask-an-expert’ (AaE) approach, which has its roots in Artificial Intelligence (AI) and knowledge representation. The ‘Top-Down-Classification’ (TDC) approach based upon R. Arp, B. Smith, and A. D. Spear, Building ontologies with Basic Formal Ontology provides a good, clear example of the approach, with a clear summary of how it aims to construct an ontology approach has roots in biological classification and philosophy 12 Two-dimensional analysis – general legend legend unexploited exploited exploitation TDC (Top-Down Classification AaE (Ask an Expert) methodology bCLEARer levels digitalization generality AaE – Ask an Expert The ‘Ask-an-Expert’ ( AaE ) approach Process is a rationalist armchair exercise in the sense that there is little empirical content. input for the process is domain experts: “ The goal of the ontology implementation activity is to build the ontology using a formal language, based on the ontological requirements identified by the domain experts .” [p. 28] an underlying focus on natural language (from a levels of digitalization perspective, speech) “If domain experts have no knowledge about ontology data generation and querying, we recommend writing the requirements in the form of natural language sentences.” p. 22] Revealing comment Says can build an interim concept model suggesting “ diagraming tools such as MS Visio or draw.io, as well as non-digital tools as pen and paper or a blackboard ” may be used to build this AaE - Two-dimensional analysis 16 metadata schema data levels of digitalization metadata schema data brain speech writing printing computer levels of generality early-stage process - levels of digitalization time levels of generality AaE AaE exploited AaE unexploited AaE computer printing writing speech brain two-dimensional design space - exploitation TDC – Top-Down Classification TDC – Top-Down Classification p. 49 “The terms in an ontology are the linguistic expressions used in the ontology to represent the world, and drawn as nearly as possible from the standard terminologies used by human experts in the corresponding discipline.” p. 5 pre-computer computer Radical descoping of data ISO 21838-1 standard: domain ontology: section 3.18: “ontology (3.14) whose terms (3.7) represent classes (3.2) or types and, optionally, certain particulars (3.3) (called distinguished individuals‘) in some domain (3.17)” and “Some ontologies also allow terms representing certain privileged particulars (referred to as ‘distinguished individuals’), such as ‘the actual world’, ‘spacetime’, or (in an ontology of US law) ‘the US Supreme Court’”. The standard recognizes that the domain includes particulars, as it is defined in section 3.17 thus: “collection of entities (3.1) of interest to a certain community or discipline” Note “‘Entities of interest’ can include both particulars and classes or types.” The standard assumes (without explanation) that unless something is a special ‘distinguished’ particular, it is excluded from domain ontologies. 19 TDC – Two-dimensional analysis 20 metadata schema data levels of digitalization metadata schema data brain speech writing printing computer levels of generality early-stage process - levels of digitalization time levels of generality TDC TDC exploited TDC unexploited two-dimensional design space - exploitation computer printing writing speech brain TDC Comparison levels of digitalization – TDC and AaE Top-Down Classification Ask An Expert levels of digitalization time computer printing writing speech brain TDC AaE From a ‘levels of digitalization’ perspective, the two approaches are similar. They only engage with computerised information towards the end of the process. computer printing writing speech brain levels of generality – TDC and AaE Top-Down Classification Ask An Expert metadata schema data time levels of generality TDC metadata schema data AaE From a ‘levels of generality’ perspective, the two approaches have differences and similarities. Top-Down Classification engages with metadata - Ask An Expert does not. Neither engages data. bCLEARer What is BORO? 1. Introduction The Business Object Reference Ontology BORO has chosen to adopt a closer integration with philosophy than other ontologies in the information systems domain … also, unlike them, it emerged from and was developed in commercial projects rather than in academia BORO includes a foundational (or upper) ontology and a closely intertwined methodology for information systems (IS) re-engineering (Partridge, 1996), hence the term BORO refers to both the ontology and the methodology . BORO was originally conceived in the late 1980s to address a particular need for a solid legacy re-engineering process and then evolved to address a wider need for developing enterprise systems in a ‘better way’; in other words, in a way that … enable[ed] higher levels of reuse and, as a consequence, capable of reducing the effort and cost of (re-)developing, maintaining and interoperating enterprise systems. It was eventually publicly documented in (Partridge, 1996) de Cesare, S. and Partridge, C. 2016. BORO as a Foundation to Enterprise Ontology. Journal of Information Systems. 30 (2), pp. 83-112. https://doi.org/10.2308/isys-51428 Partridge, C. 1996. Business Objects: Re-Engineering for Re-Use, Butterworth-Heinemann. 25 style.visibility style.visibility style.visibility What is BORO? BORO has two closely intertwined components BORO Foundational Ontology a foundational (or upper) ontology bCLEARer a methodology systematically mining (re-engineering) the semantics from information systems The two frameworks validate and inform each other top-down BORO Foundational Ontology guides the bottom-up framework bottom-up bCLEARer framework validates the whole model 26 BORO FO - top-down framework bCLEARer - bottom-up framework Components deployed across various exploitation routes transform assess build has the application already been deployed? is there a commercial application available in the market? has it already been selected? standard needs to be implemented in application? sustain target requires a standard? requirement semantics composition mereology classification external identifiers content standardise configure 27 The user perspective 28 information ‘data’ pipeline bCLEARer’s five stages {5C22544A-7EE6-4342-B048-85BDC9FD1C3A} stages Collect Collect the datasets Establish the broad scope of the process Load Select the data in scope Translate the data into the cells Evolve Reveal the underlying semantics Mine the ontology Assimilate Merge the run into the full model Establish a single integrated model Reuse Export into applications and (re-)use 29 collect load evolve assimilate reuse Multiple inputs, multiple runs, integrated 30 a repeated sequence of automated processes: a scalable way to systematically improve semantic maturity data 2 increasing semantic maturity reuse reuse Foundational ontology collect data 1 data 3 collect collect integrated into a single foundational ontology Schematic view of one bCLEARer engine 31 Increase maturity in pragmatic steps 32 increasing semantic maturity evolve Entity repository project 1 project 2 project 3 entification O-O repository object-orientation Ontological repository ontologisation (or ontologification) style.visibility style.visibility style.visibility style.visibility bCLEARer – Two-dimensional analysis computer printing writing speech brain metadata schema data levels of digitalization metadata schema data brain speech writing printing computer levels of generality early-stage process - levels of digitalization time levels of generality bCLEARer bCLEARer exploited two-dimensional design space - exploitation bCLEARer bCLEARer exploits the full range of the two-dimensional design space Product and Process A standard for processes: Not for the ontologization process An example of this is provided by the main standard, ISO 21838-1:2021 – Information technology: Top-level ontologies (TLO) – Requirements While this references the ontology of processes, the standard makes no mention of the ontologization process itself. Hence, unsurprisingly, the standards based upon it do not mention the ontologization process either. input process output COFFEE BEANS GRINDING BREWING Need to engineer both the product and the process often-quoted engineering dictum that: “the quality of the process determines the quality of the product” More than just product and process Of course, more than just process, product (or input, process, output) Porter, Michael E., "Competitive Advantage". 1985 Digitalisation Process Factorisation Factorization of digitalisation 39 unstructured structured implicit (syntax) explicit (syntax) implicit (semantics) explicit (semantics) *ontologization surface-*computerization *computerization deep-*computerization digitalisation *computerization Ordering the factored components AaE and TDC process bCLEARer Pipeline *ontologization formalization surface-*computerization deep-*computerization *ontologization formalization = deep-*computerization?? Here is the task of making explicit what had been tacit, and precise what had been vague; of exposing and resolving paradoxes, smoothing kinks, lopping off vestigial growths, clearing ontological slums. … There is no such cosmic exile. He cannot study and revise the fundamental conceptual scheme of science and common sense without having some conceptual scheme, whether the same or another no less in need of philosophical scrutiny, in which to work. He can scrutinize and improve the system from within, appealing to coherence and simplicity; but this is the theoretician’s method generally. True, no experiment may be expected to settle an ontological issue; but this is only because such issues are connected with surface irritations in such multifarious ways, through such a maze of intervening theory. Quine, Willard Van Orman. Word and Object . 41 bCLEARer accounting example A data driven adaptation Pacioli Ruffon Transaction view from nowhere Pacioli view Ruffon View Parties Pacioli Ruffon Purchase Pacioli Ruffon Sale Owner (Owner) Counterparty Owner (Owner) Counterparty {5940675A-B579-460E-94D1-54222C63F5DA} Cash 240 Ducats … … {5940675A-B579-460E-94D1-54222C63F5DA} Cash 240 Ducats … … Purchase or sale? Debit or credit? Example from ‘ Thoroughly Modern Accounting: Shifting to a De Re Conceptual Pattern for Debits and Credits ’. (ER 2018). data debit credit T-Ledger Format {5940675A-B579-460E-94D1-54222C63F5DA} Account Title DEBIT CREDIT (left side) THEBES ATHENS Uphill Downhill N ATHENS THEBES Uphill Downhill 0 400 300 200 100 400 300 200 100 Directional terms Example from Aristotle’s Physics: “… the road from Thebes to Athens is the same as the road from Athens to Thebes; for things need not be identical in all respects because they are the same in some, but only if they are identical in what they actually are—in a word if they are not two the road from Thebes to Athens contrasts with the road from Athens to Thebes.” The description ‘uphill’ and downhill’ depends upon (is indexed to) the direction you are travelling. The purchase and sale (and debit and credit) are similarly dependent upon (indexed to) the party who own the ledger. They are a single transaction when considered on their own, but different when considered relative to one of the parties to the exchange. View from nowhere 45 SPACE TIME 240 Ducats Fra Luca Pacioli Phillip Ruffon 240 Ducats Owned by Fra Luca Pacioli Stage 240 Ducats Owned by Phillip Ruffon Stage Transaction The two parties engage in a transaction. When an agreement has been reached, the transaction is complete, and the ducats change ownership. owned by owned by 25 th April 1492 Original space-time map Figure E.3: £10,000 state change event space-time map from Partridge, Chris. 1996. Business Objects: Re-Engineering for Re-Use . Parent Company Subsidiary 1 Subsidiary 2 Lateral Upstream Downstream Indexing raises accounting issues When there is a transaction, each party records the transaction from its perspective as if it is transacting with a third party. Later when the parent company accounts are consolidated, the parent company’s books are required to view the same transactions as if they were internal and therefore have a net-zero effect on the group’s assets and liabilities. As each account’s perspective is incompatible with the other, this creates an overhead of off-book calculations. More generally, where there is a need for a cross-organisational viewpoint (for example, supply chain management) or for interoperability between organisational units the single de se perspective becomes unwieldy. Pacioli Ruffon data Venezia S.r.l . The bCLEARer process PHAS AAS ZAS Domains Source Systems holding company accounting Target System 1 Other 2 NAS Target System 4 Target System 5 Target System 6 Target Systems BORO Foundation Ontology P recise U nambiguous R eusable E xtensible C ollect L oad E volve A ssimilate R euse single company accounting {5C22544A-7EE6-4342-B048-85BDC9FD1C3A} PHAS Peak Holdings Accounting System Peak Holdings Ltd Holding Company AAS Acme Accounting System Acme Ltd Subsidiary ZAS Zenith Accounting System Zenith Inc Subsidiary NAS New Accounting System --- Consolidated Described in more detail in the paper The same integrated datasets can be reused in multiple applications Integrated Dataset 3 Integrated Dataset 2 Integrated Dataset 1 style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility style.visibility Setting the right context ontologization as a tool for exploiting the emerging opportunities offered by digitalization Setting the right context - evolution For foundational issues, setting the right context can have a bigger impact on success than the quality of the problem-solving processes. Ontologization needs some contextual scaffolding to provide the perspective that enables us to better understand the scope and nature of these new practices identify leverage points (Lamarckian targeting) places within a system where a small change can lead to significant, long-term improvements Ontologization is an essential part of a much wider more pervasive phenomena, the latest information evolutionary step – digitalization (the emergence of computing technologies). From a short-term perspective, this recapitulates the relatively recent steps of printing and writing. From a short-term perspective, this fits into a long evolutionary trail of information transitions that spans life on earth. Within such perspectives, ontologization can be understood as a tool for exploiting the emerging opportunities offered by digitalization. 50 Summary Summary Highlighted a couple of key points from the paper Look at paper for both more points, and these points in more depth Focused on point 1: The design space for the ontologization process is richer than current practice would suggest Using the two-dimensional perspective (generality and digitization) Opportunities along both dimensions Generality – data inclusivity Digitalisation – early adoption (in the process) bCLEARer provides an example of how this can be done Indicated point 2: Setting the right context can be the key to success 52 questions 54
BORO Publications
Extending the design space of ontologization practices: Using bCLEARer as an example
22 October 2024Presented at STIDS 2024, Twelfth International Conference on Semantic Technology for Intelligence, Defense, and Security, 22-23 October, Woodbridge VA, USA
Overview
Our aim in this paper is to outline how the design space for the ontologization process is richer than current practice would suggest. We point out that engineering processes as well as products need to be designed – and identify some components of the design. We investigate the possibility of designing a range of radically new practices, providing examples of the new practices from our work over the last three decades with an outlier methodology, bCLEARer. We also suggest that setting an evolutionary context for ontologization helps one to better understand the nature of these new practices and provides the conceptual scaffolding that shapes fertile processes. Where this evolutionary perspective positions digitalization (the evolutionary emergence of computing technologies) as the latest step in a long evolutionary trail of information transitions. This reframes ontologization as a strategic tool for leveraging the emerging opportunities offered by digitalization.