Broadening Ontologization Design: Embracing Data Pipeline Strategies Chris Partridge BORO Solutions Ltd University of Westminster London, UK 0000-0003-2631-1627 Andrew Mitchell BORO Solutions Ltd University of Westminster London, UK 0000-0001-9131-722X Sergio de Cesare Westminster Business School University of Westminster London, UK 0000-0002-2559-0567 John Beverley Department of Philosophy University at Buffalo Buffalo, USA 0000-0002-1118-1738 Abstract—Our aim in this paper is to outline how the design space for the ontologization process is broader than current practice would suggest. We point out that engineering processes as well as products need to be designed – and identify some components of the design. We investigate the possibility of designing a range of radically new practices implemented as data pipelines, providing examples of the new practices from our work over the last three decades with an outlier methodology, bCLEARer. We also suggest that setting an evolutionary context for ontologization helps one to better understand the nature of these new practices and provides the conceptual scaffolding that shapes fertile processes. Where this evolutionary perspective positions digitalization (the evolutionary emergence of computing technologies) as the latest step in a long evolutionary trail of information transitions. This reframes ontologization as a strategic tool for leveraging the emerging opportunities offered by digitalization. Keywords—ontologization design space, data pipelines, bCLEARer methodology, ontologization methodologies, computerization, digitalization, information transmission, information evolution I. INTRODUCTION In ontology engineering there is, in theory at the very least, a tight coupling between ontologies and ontologization, the process that produces them. Our aim in this paper is to suggest that the design space for the ontologization process is wider than a look at many of the current methodologies would indicate. To illustrate this at a general level, we partition the space along two dimensions: the levels of generality and digitalization. Fig. 1 shows how this partitioned space is currently exploited – with a focus on the early stages of the process – exposing the areas that are not being exploited. Fig. 1. Two-dimensional design space - exploitation This suggests that the space is broader than current practices would indicate. That it is possible to open the space up to a range of potentially radical, new practices based upon data pipelines. We hope that by outlining some of these practices here we will make the case for broadening the design space and encourage the community to adopt a wider range of practices. In large part, the evidence for this is derived from our work over the last three decades with an outlier methodology, bCLEARer, which provides a useful example of these data pipeline practices. We raise the engineering point that processes as well as products should be designed – and note a design poverty for formal ontologization processes relative to the final ontology products. We provide a factorization of the wider digitization process into which we believe formal ontologization fits. This separates computerization and ontologization concerns and we raise questions about how these components should be ordered. It is recognized that for foundational issues, setting the right context can have a bigger impact on success than the quality of the problem-solving processes. This is the case for ontologization which needs contextual scaffolding to provide the perspective that enables one to better understand the scope and nature of these new practices. Specifically, that one should see ontologization as an essential part of a much wider more pervasive phenomena, the latest information evolutionary step – digitalization (the emergence of computing technologies). From a short-term perspective, this recapitulates the relatively recent steps of printing and writing. From a long-term perspective, this fits into an evolutionary trail of information transitions that spans life on earth. Within such perspectives, ontologization can be understood as a tool for exploiting the emerging opportunities offered by digitalization. A. Structure of the paper In the next section, we provide a broad picture of the ontologization process and introduce two current mainstream ontologization methodologies to act as a baseline for comparisons. In the third section we introduce our example outlier data pipeline-based methodology, bCLEARer. In the fourth section we build the contextual scaffolding, firstly situating digitalization in an evolutionary perspective, then situating ontologization within digitization. In the fifth section, we situate bCLEARer in this evolutionary perspective. In the sixth section, we illustrate from within the evolutionary perspective some of the outlier design choices that bCLEARer has made. -- 1 of 18 -- II. A BROAD PICTURE OF THE ONTOLOGIZATION PROCESS In ontology engineering one would expect a tight coupling between ontologies and ontologization, the process that produces them. One can characterize this as a product-process distinction, a fundamental concept in traditional engineering. The engineering mindset differentiates the product (the output) and the process (the method or system used to produce the output) and expects both to be engineered. And, as part of the engineering, to both be quality managed, hence quality assurance (process) and quality control (product). A. Top-level ontologies and ontologization In top-level ontology (engineering) work, the process is often an invisible relation of the product. An example of this is provided by the main standard, ISO 21838-1:2021 – Information technology: Top-level ontologies (TLO) – Requirements [1]. While this references the ontology of processes, the standard makes no mention of the ontologization process itself. Hence, unsurprisingly, the standards based upon it do not mention the ontologization process either. Some top-level ontologies have associated documentation for the ontologization process. The top-level Descriptive Ontology for Linguistic and Cognitive Engineering (DOLCE) has a related analysis tool OntoClean [2], but this falls far short of an ontologization process. The Basic Formal Ontology (BFO) has a book [3] on the ontologization process that assumes the BFO top-level ontology. We look at this text in more detail below. The BORO Foundational Ontology has a closely intertwined bCLEARer ontologization process described in [4] and [5]. There are a variety of domain level ontologization processes that we discuss later in this paper. B. The case for engineering the ontologization process The importance of engineering the process is reflected in the often-quoted dictum that: the quality of the process determines the quality of the product. For a historical background to this, from the wider history of innovation, see Mokyr’s The Past and the Future of Innovation [6] or his A Culture of Growth [7]. He argues that history shows that technological progress cannot rely on artisanal skills alone, it needs to be supplemented with formal and systematic (that is, engineered) knowledge. Within engineering, this idea was explored, analyzed and championed in manufacturing in the second half of the 20th century by quality management pioneers such as W. Edwards Deming [8] and Joseph Juran [9]. This led to a rich variety of designs including movements such as Total Quality Management and its successors Lean Manufacturing, and Six Sigma. These developed a fertile range of ways of managing manufacturing processes. An example is the Plan-Do-Check-Act (PDCA) Cycle used to design, implement, and refine processes on a small scale – which fits well with the Kaizen philosophy of continuous improvement, where processes are regularly reviewed and improved incrementally. This has spread to some other domains. For example, one can see Kaizen-like principles being used in the Agile software development methodology. Within the ontology engineering community, one does not find a comparatively rich selection of designs and range of ways of managing the ontologization processes – a kind of process design poverty. This is despite a few interesting innovative examples such as the ROBOT tool (https://robot.obolibrary.org/extract.html). Especially for top- level ontologies, there appears to be more ‘theory’ for, and so more attention on, the design of the final product – ‘the ‘ontology’ – than the process – ‘ontologization’ – that produces it. From an engineering perspective, this imbalance looks unhealthy. One could argue that this poverty arises from the process being relatively new and under-researched, unlike, for example, top-level ontology which can build upon a rich heritage. 1) Process design poverty in logic Interestingly, a similar poverty of design process has been pointed out in logic, which is a key part of the last stages of the ‘ontologization’ process. Novaes [10] makes a product-process distinction for logic, distinguishing the formal product from its formalization process, noting an almost exclusive focus on the former in contemporary logic: “As a discipline, logic is arguably constituted of two main sub-projects: formal theories of argument validity on the basis of a small number of patterns, and theories of how to reduce the multiplicity of arguments in non-logical, informal contexts to the small number of patterns whose validity is systematically studied (i.e. theories of formalization). Regrettably, we now tend to view logic ‘proper’ exclusively as what falls under the first sub- project, to the neglect of the second, equally important sub- project.” She discusses two historical theories of argument formalization, from Aristotle and Medieval Logic that have more balance. Both “illustrate this two-fold nature of logic, containing in particular illuminating reflections on how to formalize arguments (i.e. the second sub-project).” She suggests reflecting on these should lead to a broader conceptualization of what it means to formalize. Given how much the ontology (engineering) product builds on formal logic – inheriting many of its (cultural) practices – this may contribute to the poverty in ontology engineering. This suggests that developments in the formalization process could be recruited by and enrich ontologization’s approach to formalization. C. Comparing different ontologization processes In this section, we briefly look at some current mainstream methodologies that guide the ontologization process to provide a basis for comparison with the data pipeline approach exemplified by bCLEARer. This gives us a rough benchmark on common practices. A caveat: we do not claim that this selection reflects all the work that is happening in this area. Rather, we are aiming for examples that lend themselves to our broad comparison. In this section we restrict ourselves to ontologization to help make a clear comparison. This is even though, as we touch upon later from a bCLEARer data pipeline perspective, there are interesting features in the methodologies guiding the processes in other software related domains, such as: -- 2 of 18 -- • Waterfall model: clear top-down separation of concerns. • Agile: flexible, responsive, efficient, iteration • DevOps (and DataOps): automation into a data pipeline to improve and shorten life cycles. There is a reasonably rich literature on ontology methodologies, including [11], [12], [13], [14], [15], [16], [17]. We roughly divide these into two broad camps, which we have colloquially labelled: ‘Ask-an-Expert’ (AaE) and ‘Top-Down- Classification’ (TDC). We have selected a representative document for each camp: For AaE, OntoCommons report D.4.2 [18] and for TDC, Building ontologies with Basic Formal Ontology [3]. One aspect of these methodologies we inspect is the information pathway they create, the flow or movement of information through the stages of the overall process. We specifically explore how this interacts with the two dimensions of the design space. Firstly, the level of generality dimension which, for ease of understanding, we introduce from a data perspective as the metadata, schema and data levels. This is a simplification as it is about the syntax of the implementation, whereas generality is also a semantic matter. However, there is a good enough rough match between syntax and semantics here to make the substitution fair for our broad classification. Secondly, the levels on the journey to digitalization dimension [19], [20]. This looks at the evolutionary steps on the journey to digitalization. Very broadly a journey that goes from brains to speech to writing to printing and then computing. We call this the levels of digitalization. The results of the inspection are in the earlier Fig. 1. 1) The ‘Ask-an-Expert’ approach We selected the OntoCommons report D.4.2 [18] as our basis for AaE. It suits our purposes as it not only describes its own approach (the LOT methodology) but documents other similar approaches (including Grüninger & Fox [13], METHONTOLOGY, On-To-Knowledge, DILIGENT, NeOn, RapidOWL, SAMOD and AMOD). Together these provide many good examples of the ‘Ask-an-Expert’ (AaE) approach, which has its roots in Artificial Intelligence (AI) and knowledge representation. This process is largely a rationalist armchair exercise – in the sense that there is little empirical content. The input for the process is domain experts – as this quote illustrates: “The goal of the ontology implementation activity is to build the ontology using a formal language, based on the ontological requirements identified by the domain experts.” [13, p. 28] Across all the approaches reviewed, there is a similar information pathway from a level of digitalization perspective. In the early stages there is an underlying focus on natural language (from a levels of digitalization view, speech), sometimes organized into (natural language) competency questions [13] – as this quote illustrates: “If domain experts have no knowledge about ontology data generation and querying, we recommend writing the requirements in the form of natural language sentences.” [13, p. 22] The methodology’s input to the information pathway is the brains of experts via speech into documented (unstructured) natural language. Then the methodology broadly separates concerns [21]: separating the confirmation of content from its formalization – and chooses to address the first concern before the second. We will revisit this point later, but it is important to note that this separation and ordering choice assumes that reaching content agreement prior to the formalization process won’t negatively impact the final product. In this design architecture, the first stage is a confirmation of content which uses mostly (unstructured) natural language which is organized and agreed as a statement of the requirements of the ontology. The second stage takes the natural language and formalizes them. The early pathway is not always or entirely natural language, as there is a mention of the possibility of using more structured information in the shape of a “tabular technique” using “3 types of tables: Concepts, Relations, Attributes”. Formalization (structured information) only really enters the process in the later stages of the pathway in ontology implementation, after the requirements (expressed in natural language) are collected. The paper notes that there is optionally a conceptualization stage, where an interim concept model based upon the requirements may be built. Interestingly, it suggests that “diagraming tools such as MS Visio or draw.io, as well as non- digital tools as pen and paper or a blackboard” may be used to build this. 2) The ‘Top-Down-Classification’ approach We take Building ontologies with Basic Formal Ontology [3] as the baseline for the ‘Top-Down-Classification’ (TDC) approach. This provides a clear example with a concise summary of how it aims to construct an ontology (this shows why it deserves the top-down-classification nickname). This process is also largely a rationalist armchair exercise, one that has roots in biological classification and philosophy. It is a common approach to developing top-level ontologies in Information Systems (IS). “Overview of the Domain Ontology Design Process Ontology is a top-down approach to the problem of electronically managing scientific information. This means that the ontologist begins with theoretical considerations of a very general nature on the basis of the assumption that keeping track of more specific information (for example, about specific organs, genes, or diseases) requires getting the very general scientific framework underlying this information right, and doing so in a systematic and coherent fashion. It is only when this has been done that the detailed terminological content of a specific science such as cell biology or immunology can be encoded in such a way as to ensure widespread accessibility and usability.” [3, p. 49] This informal view is then structured into a step-by-step process in a table – see below. Table 3.1 An outline of the steps to be followed in designing a domain -- 3 of 18 -- ontology 1. Demarcate the subject matter of the ontology. 2. Gather information: identify the general terms used in existing ontologies and in standard textbooks; analyze to remove redundancies. 3. Order these terms in a hierarchy of the more and less general ones. 4. Regiment the result in order to ensure: a. logical, philosophical, and scientific coherence, b. coherence and compatibility with neighboring ontologies, and c. human understandability, especially through the formulation of human-readable definitions. 5. Formalize the regimented representational artifact in a computer usable language in such a way that the result can be implemented in some computable framework. [3, p. 50] From this table, we can pull out a level of digitalization perspective along the information pathway. The first four stages work with unstructured natural language. Though the terms ‘order’ and ‘regiment’, at steps 3 and 4, suggest some structure in the information, it is only at step 5 that the information is formalized, and so fully structured data enters the process. So here as well, the information pathway to the ontology starts with brains then via speech or directly into text is documented (unstructured) natural language. In the process, there is a similar reliance upon human experts to justify choices, see: “The terms in an ontology are the linguistic expressions used in the ontology to represent the world, and drawn as nearly as possible from the standard terminologies used by human experts in the corresponding discipline.” [3, p. 5] As an aside, it is often not recognized that the terms themselves, as inscriptions or utterances, are also elements of the domain that can usefully be represented in the ontology 3) Process Comparison One can make a rough assessment of the engineering maturity of these methodologies. As the quotes above hint at, they are currently collections of "ad hoc rules" with simple heuristics. There is no background context to act as a foundation to guide the engineering of the process design – certainly no common context. Hence, they are, from an engineering design perspective, at an early stage of development. There is still plenty of scope for them to undergo the kind of serious engineering re-design Deming and Juran undertook for manufacturing. Both approaches have several features in common, ones that differentiate them from data pipeline approaches, such as bCLEARer. From the perspective of digitalization, in both cases, their early processes for establishing the base ontology focus their efforts on working with unstructured (pre-digitalization) natural language related to human understandability, with less focus on machine understandability. In both approaches the ontologization happens before the formalization. The (unstructured) ontological information is captured and regimented in natural language first and then formalized. One can broadly divide information into levels of generality. From a syntactic ‘data’ perspective, these are the natural levels: metadata, schema and data. For our purposes here, these are a good enough rough proxy for the semantic levels; top, the most general, middle and the most specific bottom level – typically particulars. We can use this perspective to show how the two approaches showcased differ and agree. They differ in their approach to the metadata level. The ‘top-down-classification’ approach works by framing the middle level in terms of the relevant top-level structure – so the middle schema level is framed by this metadata. The ‘ask-an-expert’ approach works explicitly at the schema level – focusing on the domain. In principle, there is no reason why the ‘ask-an-expert’ approach could not start with the metadata level, or the ‘top-down-classification’ approach could not ignore the metadata level and work at the schema level. The two approaches agree on their approach to the bottom data level. They both ignore it (a significant omission, we return to below). One can relate the design of the digitalization and generality features. If one chooses to design a process to work with humans and unstructured information, one needs to recognize that one is building in scaling constraints that block the processing of large amounts of information. One way around this is, of course, to work with the data level indirectly through the schema level. Data pipeline approaches take a different route and aim to automate the process so removing these scaling constraints. III. BCLEARER – HISTORY In this section, we provide more context with a brief background history of the data pipeline approach, bCLEARer and its place in BORO (an acronym for ‘Business Object Reference Ontology’). BORO’s development and deployment started in the late 1980s. This early work is described in Business Objects [4]. BORO’s focus was then, and is now, on enterprise modelling; more specifically, it aims to provide the tools to salvage and reuse the semantics from a range of enterprise systems building a single ontology with a common foundation in a consistent and coherent manner. BORO was originally developed to address a particular need for a solid legacy re-engineering process. This naturally led to the development of a methodology for re-engineering existing information systems, currently named bCLEARer – where the capital letters are an acronym for Collect, Load, Evolve, Assimilate and Reuse. This was co-developed with a closely intertwined top-level ontology (the BORO Foundational Ontology). Hence, the term BORO on its own can refer to either of, or both, the mining methodology and the ontology. Our focus here is on the bCLEARer methodology which is used to systematically unearth reusable and generalized ontological business patterns from existing data. Most of these patterns were developed for enterprises and successfully applied in commercial projects within the financial, defense, and energy industries. bCLEARer has evolved organically over the last three decades both in response to the evolutionary pressures of experience as well as exploiting the opportunities provided by evolving digital technology. An early version of the methodology is described in [4] – with a detailed description in -- 4 of 18 -- Part 6. At the time this was developed, the late 80s and early 90s, the technology support was immature, so while the core process was systematized it was not fully automated. Over the last decade, as appropriate technology has emerged, the core process has been fully automated into a data pipeline. Later versions of this are described in several places, including [22]. There are also open-source examples on GitHub (https://github.com/boro- alpha). bCLEARer (and its associated top-level ontology) have, over the years, been configured to exploit a variety of situations ranging from its original legacy system migration to application migration to developing requirements and quality controlling existing systems. A common feature of all these projects has been the initial collection of one or more datasets (where this may include both structured and unstructured data – though structured data is preferred) and its regimented evolution to a more digitally aligned state. IV. SITUATING DIGITALIZATION AS INFORMATION EVOLUTION We have established that (engineers have learnt that) the quality of the final product depends upon the engineering quality of the design of the process. We have also suggested that the mainstream ontologization processes are very lightly engineered with a weak background context. This indicates that there is an opportunity to develop a more engineered ontologization process. What is less clear is what form this engineering should take. Over the last three decades, the evolution of bCLEARer’s practices was initially driven by experience. This pragmatic, experiential approach is supported by many including Aristotle [23] who said, “for the things we have to learn before we can do them, we learn by doing them”. Reflecting upon the process has always been a central part of the practice. However, in the last decade, as the practice has matured, questions about the broader context have naturally arisen and this has led to a much better understanding of how the practice should be engineered. This better understanding has merged in large part from a recognition that for foundational issues, setting the right context can have a bigger impact on improvements than the quality of the problem-solving processes. This is not a new idea; it is already established in many fields. In the context of professional practice, Donald A. Schön’s “The Reflective Practitioner,” [24] raised the concept of problem setting as a crucial part of having a successful outcome. In design thinking, David Kelley emphasizes the importance of problem framing in his book "The Art of Innovation" [25] explaining that reframing the problem often opens the door to more creative and effective solutions. In systems thinking, Peter Senge’s “The Fifth Discipline” [26] differentiates between problem identification and problem solving. For him effective problem solving involves identifying leverage points – places within a system where a small change can lead to significant, long-term improvements. In each of these cases, the initial stage involves developing a clear understanding of the underlying causes and interconnections of the entire complex system. Problem solving then involves developing interventions that address the root causes. In the case of ontologization there is a ready-made context that can provide the perspective needed – this is evolutionary theory. If one positions ontologization as an essential part of a much wider more pervasive phenomena, the latest information evolutionary step – digitalization (the emergence of computing technologies) then a new picture emerges. From a short-term perspective, this recapitulates the relatively recent steps of printing and writing. From a long-term perspective, this fits into an evolutionary trail of information transitions that spans life on earth. Within such perspectives, ontologization can be understood as a tool for exploiting the emerging opportunities offered by digitalization. We can then position the bCLEARer practices into the overall evolutionary context. This leads, in turn, to a clearer picture of the specific evolutionary pressures within the practice. From the start, bCLEARer has been framed in terms of information engineering, evolution and revolution. The narrative in [4] was the evolution of information paradigms. For much of its early life the focus of bCLEARer’s evolution has been pragmatic practice adapting to the evolutionary pressures presented by actual ontologization with only a modicum of reflection upon the nature of the process. In the last decade or so there has been more reflection on what these practices might reveal. We are reaching a conclusion that the best way to understand the design of the process is to make the originally evolutional framing much richer – to firstly more clearly frame the process as digitalization and secondly to show the digitalization as part of a wider trend that frames evolution in terms of information. In the next section, we set up a broad picture of this wider trend of evolving information. In the section after that we situate digitalization as one transition in that evolution. 1) Situating digitalization as information innovation Digitalization is an information transition – and it turns out information transition is a ubiquitous pattern which can be used to frame the whole of macro-evolution. This provides a reassuringly broad context where digitalization is the latest in a long history of information transitions. Maynard Smith and Szathmáry [27], [28], [29] suggest that macro-evolution can be characterized as a series of information transitions and that one can frame the whole of macro-evolution in these terms: “… that evolution depends on changes in the information that is passed between generations, and that there have been ‘major transitions’ in the way that information is stored and transmitted, starting with the origin of the first replicating molecules and ending with the origin of language.” And these changes in information transmission, the passing of “information … between generations”, are central to evolution. Maynard Smith and Szathmáry provide a table of seven major transitions, where the third is the "genetic code" and the seventh is "language". Each transition not only transforms life but also transforms the way life evolves – and so, in a sense, evolution evolves through transformations in information transmission. They also suggest that each of these transitions has accelerated and expanded evolution enabling more complex entities to emerge quicker. -- 5 of 18 -- Jablonka and Lamb [30] expanded this framework saying: “… we argue that information transmitted by non-genetic means has played a key role in the major transitions, and that new and modified ways of transmitting non-DNA information resulted from them.” More specifically, they argue that: “The evolution of a nervous system not only changed the way that information was transmitted between cells and profoundly altered the nature of the individuals in which it was present, it also led to a new type of heredity—social and cultural heredity—based on the transmission of behaviorally acquired information.” This new type of heredity enables even faster, more flexible, more complex evolution. One key feature is that the social and cultural heredity is not (like genetic heredity) necessarily dependent upon (genetic) life cycles, so it can give rise to adaptations which easily spread through a population within a life cycle. This social and cultural adaptation is orders of magnitude faster than genetic evolution and, particularly in changing environments, faster adaptive evolution is more successful. Human culture is a good example of this. In Evolution in four dimensions: genetic, epigenetic, behavioral, and symbolic variation in the history of life [31], Jablonka and Lamb, describe in Chapter 9 – Lamarckism Evolving: The Evolution of the Educated Guess how new types of non-genetic heredity – behavioral and symbolic inheritance – enable a new directed evolution. They label this, understandably, Lamarckian. “… the variation on which natural selection acts is not always random in origin or blind to function: new heritable variation can arise in response to the conditions of life. Variation is often targeted, in the sense that it preferentially affects functions or activities that can make organisms better adapted to the environment in which they live. Variation is also constructed, in the sense that, whatever their origin, which variants are inherited and what final form they assume depend on various “filtering” and “editing” processes that occur before and during transmission.” a) Evolutionary transitions in information transmission This Lamarckian ‘targeting’ and ‘constructing’ enables further evolutionary transitions in information transmission. Obvious examples from symbolic evolution are the emergence of writing and printing information technologies. These involve fundamental “changes in the way information is stored and transmitted” where writing involved changes in structure and printing changes in economics. These both clearly led to further innovations in information evolution. Ironically, this reveals CRISPR technology, where DNA is selectively modified, as a Lamarckian genetic evolution. While the transitions clearly involve the use of new technology, closer examination (see, for example, Ong [32] and Olson [33]) reveals they depended upon the coevolution of human behavior and external information technologies. Where both the emergence and exploitation of the technology depends upon intertwined coevolution with human behavior. There is a similar intertwined evolution of behavior and technology so far in the emergence of computing technology. This current transition is often broadly called, in the context of enterprise processes, ‘digitalization’ – which includes ‘digitization’, the process of converting information, data, or physical objects into a digital format, readable by computers. b) Domestication as an example of co-evolution A more familiar example may help us to appreciate the nature of coevolution – the domestication of plants and animals. This has been studied as a distinctive coevolutionary relationship between domesticator and domesticate in a range of research [34], [35]. In this domain, it is easy to see that both parties (the domesticators and domesticates) coevolve in the sense of contributing to the relationship. Zeder [34, p. 3191] describes domestication as a: “… relationship in which one organism assumes a significant degree of influence over the reproduction and care of another organism in order to secure a more predictable supply of a resource of interest.” Our relationship with the new digital forms of information technology can be described in a similar way. One where we domesticate our computer systems controlling their breeding. c) A new ‘digital’ form of information transmission In The Selfish Gene, Dawkins [36] introduced an idea. He distinguished between genes as “replicators” that pass on copies of themselves through generations and organisms as “vehicles” or “survival machines” constructed by genes to survive in the environment and so ensure their continued replication. One could adopt a ‘promiscuous ontology’, one that regards computer systems as a form of life subject to evolution [37], [38], [39]. If so, then computers can easily be seen as “vehicles” for the information they carry and replicate as well as things we domesticate and breed. Furthermore, the emergence of these individuals then passes the Smith and Szathmáry test in so far as it radically ‘changes the way information is stored and transmitted’ directly between these individuals – and with humans. One could be less adventurous and instead see the computer systems as an extension of humans [40]. In this view, computer systems are replicators rather than vehicles – they are part of the apparatus transmitting information rather than “survival machines” in a computer ecosystem. Even on this view, they pass the Smith and Szathmáry test. So, either way, one can see them as providing an opportunity for an information transition [20]. 2) Narrower context – Lamarckian choices for co-evolution If, as looks likely, history repeats itself and this digital transition follows the pattern of most previous transitions, then it is likely to involve the coevolution of human behavior and digital technology. The energy enterprises currently devote to their digitalization efforts show an intuitive understanding that some kind of directed effort needs to be made. From our evolutionary perspective, we can see this as our culture starting to co-evolve with the new technology – looking to construct the Lamarckian variations that will exploit this opportunity. What is less clear is which Lamarckian constructed variations are likely -- 6 of 18 -- to lead to significant success – in the language of evolution to be able to exploit the natural selection pressures well. Maybe history has a clue. A common historical narrative is that technology drives change, that it is the emergence of a new technology that initiates the associated cultural change. Careful study reveals a more interrelated pattern of co-evolution between technology and culture. Where cultural change often prepares the ground for technological innovation and then feeds off it and feeds further innovation. Olson [33] provides a relevant example, explaining how cultural developments in Western Europe from the 9th century onwards played a key role in the invention of moveable type in the 15th century. This technological innovation then laid the ground for developments in Western science in the 16th and 17th centuries. So maybe the coevolution of culture and digital technology evolution can give us some clues on where to target Lamarckian variations. It is well-accepted that work in logic and mathematics laid the foundation for computing. If we look at the culture of this work, then we can see some trends that help us to target variations to help the co-evolution. Well before the introduction of digital computers, Frege [41] uses the analogy of a microscope and the eye to explain how his formal language compares with ordinary informal language, noting that it provided a superior sharpness of resolution. Carnap [42] talks about ‘rational reconstruction’ (rationale Nachkonstrucktion). Quine [43] says that one doesn’t merely clarify commitments that are already implicit in unregimented language; rather that one often creates new commitments by regimenting. Quine notes that paying attention to the ontological commitment often leads to radical ‘foreign’ differences [44, pp. 9–10]: “Ontological concern is not a correction of lay thought and practice; it is foreign to the lay culture, though an outgrowth of it” Adding “There is room for choice, and one chooses with a view to simplicity in one’s overall system of the world.” Lewis [45, pp. 133–5] following the theme of ‘outgrowth’, argues that the differences are a result of taking the lay common sense seriously, by trying to make it simpler and consistent. 3) The ontologization value chain The bCLEARer methodology is an example of data pipeline approach, which is the topic of this paper. It has, through experience refined these historical intuitions, developing a view of the digitalization process as a network of transformation processes in a data pipeline. One way to characterize this is as a value chain, where each transformation adds value. We briefly outline what a value chain is below and then describe the factorization of digitalization into component transformations. a) Recruiting the value chain view In manufacturing, Porter’s value chain [46] provides a useful tool for broadly characterizing processes as a system of transformations. The system is provided with inputs which feed into a network of transformation processes. This network feeds into the output. The characterization is recursive. Each transformation process can be seen as a sub-system with its own value chain. We recruited a lightweight version of this tool to characterize ontologization. Under this view, at the broadest level, the ontologization process starts with pre-ontologization information and is transformed using an ontologization process into an ontology. It adds value by transforming the pre- ontologization information into a formal ontology. This reveals an information pathway that starts with the pre-ontologization inputs, undergoes transformations and is output as a formal ontology. Different methodologies have different intermediate transformations and so different information pathways. While the inputs and outputs remain similar, the value chain transformations differ. b) Factoring digitalization into *computerized and *ontologized We firstly factorize digitalization into two types of digital transitions that it has found useful to target (and construct). These are *computerization and *ontologization. We use the ‘*’ prefix convention to indicate our specialized use of the term and differentiate it from the many other senses in which it is used. *Computerization is the process of converting relatively unstructured information into formally structured data. It implies something more than the digitization mentioned earlier, which just aims at bare computer readability. Formally structured data refers to information that is organized into a highly defined and predictable form, typically within a fixed schema or format. So, a scan of an engineering drawing in, say, PDF format would be digitized but not *computerized, as there is no direct way for the computer to read the components of the drawing. Whereas an engineering drawing in a CAD format, such as native DWG, would be *computerized, as the information in the drawing is explicit in its structure and can be read directly by a computer. *Computerization is intended to be a pragmatic distinction and while there are borderline cases, there are also cases that clearly fall into the pre-*computerization and *computerization camps. *Ontologization is the process of converting relatively semantically unorganized information into information organized into a common ontological structure. Typically, the information is used in a domain, and there is a level of semantic precision needed for it to be fit for purpose. The *ontologization organizes the information into a common ontological structure that is sufficiently fine-grained to capture the requisite semantic precision. One way of characterizing *ontologization is that it develops an explicit picture of ontological commitment [47], [48]. There is a long tradition of seeing this process as a transformation that reveals a deeper structure. Currently, most information systems being digitized have no precise explicit ontological commitment. So, in practice, the *ontologization is often a regimentation [43] and rational reconstruction [42] of what the ontological commitment would be given some preferred top-level ontology. In bCLEARer’s case, the top-level ontology is the BORO Foundational Ontology [5]. c) *Computerization – transitioning from implicit to explicit formal structure The bCLEARer methodology has further identified a factorization of the *computerization transition into two sub- transitions: surface-*computerization and deep- *computerization, corresponding to two levels of *computerization. -- 7 of 18 -- The process of transforming unstructured pre-*computerized information, giving it a highly defined and predictable form is sufficient to make it surface-*computerized. Much data in data stores is in this state today. It may, and often in practice does, have significant implicit formal structure. In some cases, making the structure implicit may be deliberate, as part of the process of improving the performance of a system. For our purposes we want to make the structure explicit to facilitate the *ontologization. We do this in the process of deep- *computerization. Uncovering the underlying implicit form of surface- *computerized information requires a degree of ethnographic hermeneutics – one needs to be able to interpret, to understand, its implicit structure from its perspective. The deep- *computerization transformation aims to expose this and, as far as feasible, make the underlying infrastructure transparently clear. A simple example of surface-*computerized information would be SQL table schemas and their associated data without the foreign keys noted. This meets the criteria for being *computerized – it has a fixed format. However, when the deep- *computerization adds the foreign keys, one can appreciate that the pre-deep-*computerization information had implicit structure that was not explicitly visible to a computer reading the data. In other words, it was only surface-*computerized. There are existing software techniques that work in this space, that one can build upon. These include refactoring [49] and clean coding [50]. Both are bodies of practices for restructuring existing code, altering its internal structure to improve it, without changing its external behavior. The restructuring can be recruited to reveal the deeper structure. From an ethnographical perspective, deep-*computerization (and maybe surface-*computerization too) is a kind of ‘surfacing’. As Star notes in The Ethnography of Infrastructure [51] the details are technical and “excursions into this aspect of information infrastructure can be stiflingly boring”. This means that large parts of the infrastructure are often invisible, in the sense that one doesn’t pay attention to them. So, one of the challenges is training oneself to see, and so surface, the invisible structure – a practice with similarities to Bowker’s [52] “infra- structural inversion”, which foregrounds the backstage operational elements. As with the other factorizations, this is intended to be a pragmatic distinction where most cases clearly fall into one or other camp – but with some borderline cases. One common borderline case is data cleansing. This includes technical matters such as resolving encoding issues and the treatment of whitespaces (which we find are both still common) as well as keying and spelling errors. While these might degrade the quality of the surface-*computerized information, they do not seem sufficiently grave to undermine its *computerized status. And fixing them does not obviously qualify as immediately revealing implicit structure – though if they are not fixed, they can hide structure. For pragmatic reasons, we take fixing these to be part of the deep-*computerization process. d) Inter-process dependency Obviously, there is an order to the surface- and deep- *computerization process. One surface-*computerizes information before *deep-computerizing it. Theoretically, at least, the *computerization and *ontologization processes would seem to be sufficiently independent that one could undertake either one without the other – implying that there is a choice in which to do before the other. However, the bCLEARer experience is that there are strong pragmatic reasons for undertaking the full *computerized transition before undertaking the *ontologization transition [48], [20]. Our experience has been that the formalization process inherent in *computerization is best done with raw unaltered data, straight from the operational ‘wild’. This is because we found that in cases where the *ontologization process was carried out on pre-*computerization information, it often obfuscated structure that *computerization needed – making the overall process significantly harder. Hence, in bCLEARer we see *ontologization as primarily a process for refining already *computerized data. V. BCLEARER’S BROAD STRUCTURE The bCLEARer process has been modularized (see [53, App. B], [54]) into a component architectural pattern. We describe this in the first section. In the earlier comparison of current methodologies, we assessed them relative to two levels: generalization and digitalization. We now describe how bCLEARer addresses these levels in the second and third sections below. In the final, fourth, section we look at whether it is better to surface-*computerize (using bCLEARer) in vitro or in vivo. 1) bCLEARer’s Pipeline Component Architecture Framework The bCLEARer process has a pipeline (pipe-and-filter) architecture [54], a prevalent approach for data transformation. This architecture consists of a sequence of processing components, arranged so that the output of each component is the input of the next one creating a ‘flow’. The pipeline architecture has, as the 'pipe-and-filter' name suggests, a series of pipe and filter components, where pipes pass data to and from filters that transform the data — the pipeline flow. The architecture can be nested, in that filters can encapsulate a sub- pipeline process. This generic architectural pattern is refined into a more constrained pattern for bCLEARer’s more specific needs. It must include the components of the ontologization process in a structure where the specific arrangement of components can be dictated by the needs of the project and this arrangement can flexibly evolve over time, potentially into a radically different shape. Typically, it is divided into three broad levels: 1. thin slices – which typically correspond to ways of dividing the domain and the dataset [55] 2. bCLEARer stages – the stages that correspond to a particular type of transformation -- 8 of 18 -- 3. bUnits level – the filters within a single bCLEARer stage, the base filters are called bUnits. a) The bCLEARer stage types While the contents of the individual thin slices and bUnits level vary from project to project depending upon their needs, as well as evolving over time, the bCLEARer stage types are a more stable architectural feature. The design of these types is motivated by the ‘separation of concerns’ [21] principle – where each type deals with a different kind of transformation. This builds upon the factorization discussed above. The five stage types are Collect, Load, Evolve, Assimilate and Reuse (whose initials contribute to the acronym bCLEARer). Collect is the stage at which a dataset enters the pipeline. Collect stores the dataset and ensures it is not changed. There is no transformation at this stage. This provides a fixed baseline for tracking. Larger datasets are divided into chunks, to be consumed one chunk at a time. The Load stage receives the dataset from the Collect stage. The first thing it does is establish the identity of the contents to facilitate tracking and tracing. The Load stage is responsible for ensuring that the data passed onto the next Evolve stage is *computerized – at least surface-*computerized. If the dataset comes from an operational application system that uses an enterprise database, the data will probably be sufficiently structured and so need no *computerization transformation. If it is unstructured text, for example a PDF text document, it will be pre-*computerized, and so need transforming. The Load stage undertakes the minimal amount of transformation to *computerize it, in effect it surface-*computerizes it. Where this is required, the project will need to decide on the output format to use. In our projects, we usually make the target structure simple tables. The Evolve stage assumes its input data is (at least) surface- *computerized. It is responsible for digitalizing this input data. This is done in two major sub-stages. First it deep-*computerizes the data and then *ontologizes it. Typically, the very first exercise in the deep-*computerize stage is to check whether the data needs cleaning, and if so, clean it. When the data comes from several systems, it normally makes sense as part of the deep-*computerization stage to integrate the data across systems into a common format, as far as possible, after firstly transforming the data from each system on its own. When the deep-*computerization is complete, the *ontologization can start. This is guided by a minimal foundation, the BORO Seed – for an example of a relatively recent minimal seed see Top-Level Categories [56]. A full digitalization project will include both *computerization and *ontologization. But pragmatic considerations may dictate that this is done in phases – and the early phases may only go so far along the digitalization journey. For example, undertaking deep-*computerization and delaying *ontologization to a later stage. The Assimilate stage assumes its input data is ‘evolved’ – so both locally *computerized and *ontologized. It is responsible for assimilating this into a common cross-project model. The assimilated model is then ready for use in future Assimilate stages. The Reuse stage assumes its input data is assimilated. It is responsible for translating this data back into a format usable by the targeted operational systems. b) Managing micro-coevolution To some extent, the discussions about factorization and components shift focus away from the micro-coevolution that takes place. The bCLEARer journey typically involves evolutionary adaptations simultaneously on two fronts: • Information Evolution: Adaptation of information throughout its journey. • Journey Evolution: Adaptation of the journey itself to emerging requirements, accelerating the information's evolution. The whole process supports both these adaptations: identifying and accommodating significant changes in both the information and its digital journey. A key element is adaptive resilience: maintaining stability and efficiency of the factorization and components amidst continuous change. 2) A bCLEARer example A concrete example of how *computerization and *ontologization are deployed in the first three bCLEARer stages might help to make some of these points clearer. Let us say we have a legacy migration project that encompasses intercompany accounting systems. We have three source systems: PHAS (Peak Holdings Accounting System) from Peak Holdings Ltd, AAS (Acme Accounting System) from Acme Ltd and ZAS (Zenith Accounting System) from Zenith Inc., where the latter two companies are owned by Peak Holdings. For simplicity, assume all three systems are being migrated to a new COTS system – NAS (New Accounting System). Assume there is also a desire to take this opportunity to harmonize the accounting practices across the three systems. Using bCLEARer provides a systematic approach to not only integrating the data but also exposing digitalization opportunities. a) Collect stage The Collect stage will initially take a snapshot of the full dataset from the systems for the initial development of the pipeline. Later snapshots will be taken as required. As these datasets come from operational systems, they have a holistic coherence and consistency that we need to make sure persists through the bCLEARer process. a) Load stage For simplicity, assume the PHAS, AAS and ZAS systems all use relational databases. The datasets are already surface- *computerized so there are only a few further specific *computerization tasks needed at this stage. The first task in bCLEARer Load is always to mark the identity of the data, to provide a baseline for tracking and tracing. We give the systems, tables, columns and rows identities – and, where necessary, the cells as well. This is typically done with a cryptographic hash function. We also identify and inherit from the source systems the queries that can be used to check for coherence and consistency. These would include standard reports such as, in this case, the -- 9 of 18 -- account ledger, balance sheet and profit and loss reports. We typically run and hash the figures in the reports so we can easily run a simple automated binary comparison check. a) Early Evolve stage – *deep-computerization The early Evolve *deep-computerization stage is approached with an ethnographic mindset, interpreting and understanding the dataset’s implicit structure from its own perspective – aiming not to introduce any biases. This opens the possibility for a multiplicity of syntactic changes (adaptations) – data cleansing being one type. b) Early Evolve design pattern – unification of types The Evolve stage focuses on deep-*computerized. We have developed a range of design patterns to facilitate this stage. One useful design pattern simplifies the handling of data formats. There is no restriction of the format in which the dataset comes in at the Collect stage. It could be in XML, JSON or SQL or a combination of these or other formats. However, for the ontologization process these specific implementation data formats are an irrelevancy so can be filtered out. The higher levels of (syntactic) generality, the metadata and schema can be mapped into the data, which removes the dependency on any specific implemented format. We call this mapping the unification of types [37], [57]. This enables us to choose for this stretch of the pipeline a data format that suits the work we want to do – and build common code for this. It also greatly simplifies making the metadata explicit – as what was built implicitly into the collect data form can now be made explicit. In the case of the three systems, which use relational databases, the tables and columns are shifted into the data – unifying the schema and the data. At the same time, the balance sheet, profit and loss and other queries are adapted to the unified data structures and used to test the relevant semantics are preserved. There is no gap between the migration to the unified structures and testing with the queries. When we have unified the types, we have standard tools to graph-visualize the data. We do this at each major stage along the pipeline. We have found (and it is well-recognized) that it is a good way to handle large quantities of data. a) Early Evolve stage – syntactic integration Typically, different systems implement what is clearly the same information in different structures, sometimes very different structures. This creates opportunities for syntactic integration. For example, the format for the chart of accounts and postings is likely to vary between the three systems. At this stage, we take the opportunity to make simple changes that harmonize the information in the three systems, taking care to respect the perspective of the individual systems. a) Later Evolve design pattern – *ontologization Once the opportunities for deep*computerization have been exhausted, if appropriate we move to the later Evolve stage and start the *ontologization. However, there may well be situations where it makes sense to delay this until some future project. *Ontologization typically involves making semantic adaptations. We have *ontologized accounting systems before and seen the kind of semantic adaptations that emerge, see [58], [59], [60], [61], [62]. One adaptation these identify – see [62] – is the shift from de se perspectival accounting to de re ’objective’ accounting. We would expect this adaptation to emerge here as well. The way it would emerge is as follows. The top ontology would provide criteria for identity. When these are applied to individual intercompany transactions, this will provide the basis for recognizing where transaction and accounts are the same. However, under current accounting conventions these will be marked with opposing debit and credit properties. Transactions and account balances that are marked in one system as debits will be marked as credit in the other system. It turns out that whether these are tagged as debit or credit is subjective and depends upon which company’s perspective is taken. One then recognizes debit and credit as a relational property between the transaction and the company. In more practical terms, it means that the form of the data is changed. All the identical original accounts and transactions are merged into new ones – and a debit/credit relation between the company and them are established. These changes start in the data and are propagated into the schema. At the same time, the queries are amended (evolved) to take account of the new structure – and tested to ensure they can reproduce the figures in the original reports. They both confirm the consistency of the new structure and help to ensure the adaptation is preserved along the pipeline. One can recognize this as an empirical exercise where the changes emerge from the data. I
BORO Publications
Broadening Ontologization Design:
Embracing Data Pipeline Strategies
22 October 2024Presented at STIDS 2024, Twelfth International Conference on Semantic Technology for Intelligence, Defense, and Security, 22-23 October, Woodbridge VA, USA
Overview
Our aim in this paper is to outline how the design space for the ontologization process is richer than current practice would suggest. We point out that engineering processes as well as products need to be designed – and identify some components of the design. We investigate the possibility of designing a range of radically new practices, providing examples of the new practices from our work over the last three decades with an outlier methodology, bCLEARer. We also suggest that setting an evolutionary context for ontologization helps one to better understand the nature of these new practices and provides the conceptual scaffolding that shapes fertile processes. Where this evolutionary perspective positions digitalization (the evolutionary emergence of computing technologies) as the latest step in a long evolutionary trail of information transitions. This reframes ontologization as a strategic tool for leveraging the emerging opportunities offered by digitalization.