As at 12th May 2025 The bCLEARer Pipeline Architecture Framework eManual www.BOROSolutions.net -- 1 of 229 -- The bCLEARer Pipeline Architecture Framework ii Preface BORO’s experience shows that the ongoing evolution of radical innovative practices like bCLEARer depend upon adequate levels of inspectability. We have found that living documentation is a critical— though sometimes costly—enabler of this inspectability. • At a day-to-day, hygiene level, we have found that adequate documentation reduces vagueness and inconsistency, streamlining knowledge-sharing and supporting smoother scaling. • Strategically, we have found that it provides an informed foundation for the regular reflection that guides development. This helps us to clearly spot and assess successes and gaps, as well as avoiding dead ends. A while ago, a client engaged us to crystallise our then current working documentation into an eManual for their bCLEARer programme. What follows is a sanitized version of that manual, capturing the architecture as it existed then. Though bCLEARer continues to evolve, the core principles in this eManual remain relevant. We share this to support anyone running bCLEARer projects. Use it freely—it’s offered as is. -- 2 of 229 -- The bCLEARer Pipeline Architecture Framework iii Note This document is a PDF rendition of an HTML eManual. For those interested in the original, it can be downloaded from here: https://borosolutions.net/bclearer-pipeline-architecture-framework-emanual -- 3 of 229 -- The bCLEARer Pipeline Architecture Framework 1 • • • • • • • • • • • • • • • • • • • • • 1 Getting Started This bCLEARer Pipeline Architecture Framework Manual is one of a series of manuals that aim to get you started with bCLEARer. The first section provides an introduction. bCLEARer - an introduction (see page 2) provides a general introduction to bCLEARer. Its sub-section - bCLEARer's Pipeline Architecture Framework - an introduction (see page 4) - gives an introduction to bCLEARer’s pipeline architecture framework - the topic of this manual. The subsequent sections - shown in the contents listed below - describe the framework. Contents bCLEARer - an introduction (see page 2) Pipeline or pipe-and-filter architecture (see page 5) Nested gated pipeline architecture (see page 15) Architectural nesting breakdown (see page 21) Core design practices and patterns (see page 34) Appendices (see page 39) eManuals bibliography (see page 115) Bibliography (see page 117) Acknowledgements (see page 154) Appendices The manual contains these appendices: Appendix - The Standard 'Pipeline' or 'Pipe-and-Filter' Architecture (see page 40) Appendix - Aggregated Single Source of Truth Principle (see page 46) Appendix - Glossary of Major Terms (see page 47) Appendix - Design Patterns and Anti-patterns (see page 49) Appendix - Immutability and Idempotence Principles (see page 50) Appendix - Loose Coupling and Tight Cohesion Principle (see page 54) Appendix - Pipeline Architecture Iconography (see page 57) Appendix - Resilient Governance - The Right Balance of Principles and Rules (see page 69) Appendix - Separation of Concerns Principle (see page 75) Appendix - Single-Transformation (Responsibility) Principle (STP) (see page 78) Appendix - Historic Pipeline Examples (see page 79) Appendix - Further Related Topics (see page 99) -- 4 of 229 -- The bCLEARer Pipeline Architecture Framework 2 1.1 bCLEARer - an introduction 1.1.1 bCLEARer bCLEARer stands at the forefront of digital transformation, championing an evolutionary approach to harnessing digitization and digitalization opportunities. It guides information on a transformative journey, curating its evolution into fitter forms, ones more suited for computing, that deliver increased value. To accomplish this, bCLEARer has evolved an architecture framework for semantic data pipelines, along with a methodology for engineering these pipelines. This manual explains the framework. 1.1.2 Its Digital Transformation Journey This journey features two pivotal stages: Digitization: Extracting and converting information from existing sources into accessible, computer-readable formats for further transformation. Digitalization: Mining, extracting (making the implicit explicit), refining and evolving information to boost its utility and efficiency, enhancing its 'fitness' for digital ecosystems. 1.1.3 As Directed Evolution The bCLEARer journey can be regarded as a form of directed evolution for information. One that mimics the processes of natural selection, steering information toward a fitter, improved state. This steering is enabled by the improved visibility not only of the transformations, but also the patterns of incoherence and inconsistency in the data. This transparency enables easy identification and correction of issues, fostering trust and reliability in the process. 1.1.4 Rapid, Radical, Resilient Evolutionary Changes The bCLEARer journey typically involves rapid, often radical, (though resilient) evolutionary adaptations at scale. These occur simultaneously on two fronts: Information Evolution: Adaptation of information throughout its journey. Journey Evolution: Adaptation of the journey itself to emerging requirements, accelerating the information's evolution. The whole process supports this rapid evolutionary adaptation: identifying and accommodating significant changes in both the information and its digital journey. A key element is adaptive resilience: maintaining stability and efficiency amidst continuous change. 1.1.5 Deployability and Adaptability bCLEARer is deployable consistently and efficiently across a wide range of domains in programmes and projects of varying size and shape. This adaptability is underpinned by flexible, scalable, and maintainable processes. -- 5 of 229 -- The bCLEARer Pipeline Architecture Framework 3 1.1.6 Core Practices and Patterns Over three decades, bCLEARer has established core practices and patterns to consistently meet the challenges of change and variety, ensuring effective implementation. (for more details on bCLEARer see BORO, BORO Foundation and bCLEARer --- A very brief introduction (see page 116)). -- 6 of 229 -- The bCLEARer Pipeline Architecture Framework 4 1. 2. 3. 1.1.7 bCLEARer's Pipeline Architecture Framework - an introduction The bCLEARer pipeline architecture, as part of the bCLEARer environment, has evolved along with it as a resilient system, also striking a good balance between stability and flexibility. The pipeline architecture core practices and patterns help to deliver the right level of flexible stability - providing the stability within which the system can flex. The architectural framework encapsulates these practices and patterns. The pipeline architecture practices and patterns are vital to giving bCLEARer a firm foundation. These have been distilled into the bCLEARer Framework - Architecture - Pipeline (bCFAP). This report delves into the current bCFAP, highlighting its pivotal role in digital transformation. This document describes the four primary elements of this framework (with more detailed material in the Appendices (see page 39)). The initial three sections focus on the core aspects of the bCLEARer architecture: Pipeline Architecture: The first section describes the pipeline (or pipe-and-filter) architecture a prevalent approach for data transformation. (Additional background material is in the appendix: Appendix - The Standard 'Pipeline' or 'Pipe-and-Filter' Architecture (see page 40)). Nested Gated Architecture: The second section explores the nested gated architecture, which enhances scalability through its nesting structure and ensures process transparency through gating mechanisms. Three-Core-Level Nesting Architecture: The third section examines the breakdown into broad three core level architecture of the nesting. In this, the mid-level is structured to reflect bCLEARer’s five stages of evolution: Collect, Load, Evolve, Assimilate and Reuse (BORO, BORO Foundation and bCLEARer --- A very brief introduction (see page 116) gives more details of these). It also describes the two extended levels of nesting that can be added to the core levels. The fourth, and concluding, section presents the general design practices and patterns adopted in the architecture. Detail, including more technical matter and reference material, is relegated to Appendices (see page 39). This includes a glossary of the main terms (Appendix - Glossary of Major Terms (see page 47)) and a reference iconography (Appendix - Pipeline Architecture Iconography (see page 57)). -- 7 of 229 -- The bCLEARer Pipeline Architecture Framework 5 1.2 Pipeline or pipe-and-filter architecture 1.2.1 Introduction As is often noted in the literature, data transformation systems typically have a ‘pipeline’ or ‘pipe-and-filter’ architecture (see, for example, the extract in Appendix - The Standard 'Pipeline' or 'Pipe-and-Filter' Architecture (see page 40)). ‘Pipe-and-filter' is the original name for this architecture, though ‘pipeline’ is a more common name nowadays. The bCLEARer approach is a data transformation process and so, unsurprisingly, is implemented with a pipeline architecture – giving rise to a bCLEARer pipeline. In this section, we look at the pipeline architecture, including how it can be nested. In the next section, we look at the pipeline architecture’s core hierarchical, nesting structure and how larger pipelines are typically given an extended, modular, nesting structure to facilitate management. -- 8 of 229 -- The bCLEARer Pipeline Architecture Framework 6 • • 1.2.2 Visualising the pipeline flow This architecture consists of a sequence of processing components, arranged so that the output of each component is the input of the next one creating a ‘flow’. A simple visualisation of a pipeline flow is given below: simple pipeline structure The pipeline architecture has, as the 'pipe-and-filter' name suggests, a series of pipe and filter components, where pipes pass data to and from filters that transform the passed data — the pipeline flow. As well as the general inter-filter pipes, there are two specific types of pipes, the start and end pipes; start: a data source (pipe) to feed the pipeline and end: a data target (pipe) to persist the transformed data. As illustrated in the figure above, these pipes are typically adorned with a dataset collection icon (as shown below). -- 9 of 229 -- The bCLEARer Pipeline Architecture Framework 7 1.2.3 Process time Conventionally, the pipeline flow - pipes passing data and filters transforming it - is shown across the diagram from left to right. This is made explicit in some diagrams with the use of a ‘process time’ arrow, as shown below. -- 10 of 229 -- The bCLEARer Pipeline Architecture Framework 8 1.2.4 Evolutionary time As noted earlier, the bCLEARer flow itself will typically evolve over time, sometimes radically, to accommodate emerging requirements. Pipes and filters may be added and removed - or changed. Sometimes one needs be able to represent this evolution in diagrams. This is typically represented as a series of snapshots of the pipeline (process) arranged down the page with an arrow from top to bottom showing the direction of evolution - a pro-forma example of this is in the figure below. The simplest, ubiquitous kind of evolution is the iteration and extension of the bCLEARer stage level - shown below. 🔗 Similar simple iteration and extension will happen at the other two levels of nesting, as the project unfolds. These are examples of run time evolution. There is also design time evolution. For examples of gate design evolution over time see figures in Nested pipeline's data stage gates (see page 18). -- 11 of 229 -- The bCLEARer Pipeline Architecture Framework 9 1.2.5 Pipeline components The pipeline has two core components, 'filters' and 'pipes' (including the data source and data target pipes). 1.2.5.1 Filters Filters are components that transform ('filter') data that is received as an input via pipe connectors. The icon for a filter is shown below: filter icon 1.2.5.2 Pipes Pipes are the connectors for filters. The role of a pipe is to pass messages, or information, to and from filters. The flow is unidirectional, and, when needed in an implementation, the data is persisted until the filter processes it. The icon for a pipe is shown below: pipe icon 1.2.5.2.1 Pipes adorned with data icon As noted earlier, there is a data source pipe at the start of the pipeline to feed it and a data target pipe at the end of the pipeline to persist the transformed data. And these are usually adorned with a dataset collection icon. All pipes, not just the start and end pipes, transport data. Optionally, this point can be highlighted through the use of a data icon (in this case, the dataset collection icon) on any pipe in the diagram including those pipes inside the pipelines — as shown below for pipes 2 and 3. -- 12 of 229 -- The bCLEARer Pipeline Architecture Framework 10 1.2.5.3 Filters and their pipes Filters always have at least one input pipe and one output pipe, as shown in the following diagram: filter and its pipes 1.2.5.3.1 Acyclic pipeline flow In this pipeline architecture, a filter’s pipe cannot flow to itself, either directly or indirectly. The flow is acyclic — with no cycles. This does not inhibit reuse, as the same processing may be reused in different filters. 1.2.5.3.2 Multiple pipes to and from different filters A filter can have several input pipes and several output pipes, as shown in the following diagram: filter with multiple input and output pipes 1.2.5.3.2.1 Merging and splitting pipelines Filters with multiple input and output pipes can be organised into pipeline flows that split and then merge — as shown below. -- 13 of 229 -- The bCLEARer Pipeline Architecture Framework 11 1.2.5.3.2.2 Multiple pipes between the same filters None of a filter's pipes should flow to the same destination — where this happens, the pipes should be encapsulated — as shown in the diagram below. 🔗 1.2.5.3.2.3 Pipes: single or multiple filter inputs Typically a pipe will have a single filter output fed by a single filter input. However, there can be cases where it is important to record that the same data is fed from one filter to many other filters. In these cases, the pipe is shown with a single filter output feeding multiple filter inputs. Pro-forma examples of these two cases are in the figure below. -- 14 of 229 -- The bCLEARer Pipeline Architecture Framework 12 -- 15 of 229 -- The bCLEARer Pipeline Architecture Framework 13 1.2.6 Nesting Any subset of filters in a pipeline forms a sub-pipeline, however, it can be useful to distinguish between pipelines that are connected and disconnected — see below. This makes it simple to organise a sub-pipeline into a nested pipeline — as shown below. One can view the nesting in a diagram — as shown below. 1.2.6.1 Nesting — pipe encapsulation Where a nesting is going to encapsulate a number of filters, one consequence may be that a filter outside the nesting, whose pipes previously fed into multiple filters, now feeds into (or is fed from) a single encapsulated filter. In this case, the rule 'Multiple pipes between the same Filters' mentioned above comes into play, and the pipes need to be encapsulated to ensure in the nested pipeline there are not multiple pipes between the same (nested) filters — as shown in the figure below (as an aside, note the sub-pipeline is technically disconnected). -- 16 of 229 -- The bCLEARer Pipeline Architecture Framework 14 We can see the original multiple pipes in the nesting diagram — as shown in the figure below. -- 17 of 229 -- The bCLEARer Pipeline Architecture Framework 15 1.3 Nested gated pipeline architecture 1.3.1 Introduction bCLEARer pipelines have a hierarchical, nesting structure. Within this pipeline architecture, the bCLEARer pipeline has a gated nesting structure. In the first part of this section, we look at how the pipelines nest. In the second part, we look at the gating strategies. The appendix subsection - Gate Object Accounting (see page 112) - explains the use of this architecture in object accounting. -- 18 of 229 -- The bCLEARer Pipeline Architecture Framework 16 1.3.2 The bCLEARer pipeline's multi-level nesting structure How the bCLEARer pipeline nests is described in this section. 1.3.2.1 Multi-level nesting As noted earlier, within a pipeline architecture, one can encapsulate sub-pipelines as filters (with pipes). These encapsulated components can themselves be nested, creating a multi-level nesting (breakdown) structure in the pipeline. One can view the encapsulation structure in a nesting diagram — as shown below. The final nesting is of all the filters into a single overall-pipeline-level filter. From the perspective of this top level final pipeline, the nesting is a series of decompositions. These can be a fixed series of decompositions into levels, where each decomposition takes the units at one level and divides them into units at the next level. This breaks down the overall process into levels, with base (undivided) units at the bottom level — as shown in the figure below. -- 19 of 229 -- The bCLEARer Pipeline Architecture Framework 17 levelled breakdown 1.3.2.2 bCLEARer core nesting architectural levels As discussed in the next section, Architectural nesting breakdown (see page 21), the bCLEARer pipeline’s nesting has an architecture. It is initially composed of three core levels based on the type of filter: levels descriptions pipeline the top, root, level of nesting for the bCLEARer pipeline - there is only a single root pipeline filter at this level. It is the upper boundary of the pipeline architecture. stage a single boundary layer (with no nesting) that contains the filters for the stages of the bCLEARer digital journey. bUnit the leaf nodes in the nesting structure, the bUnit filters. As this is a hierarchical tree structure, they belong to one, and only one, bCLEARer stage. It is the lower boundary of the pipeline architecture. 🔗 (see page 225) Between these levels the nesting can be extended to accommodate further modularisation. -- 20 of 229 -- The bCLEARer Pipeline Architecture Framework 18 1.3.3 Nested pipeline's data stage gates The nested pipeline data stage gates are described in this section. 1.3.3.1 The data stage gates Where appropriate, the nested pipelines are designed with their input and output pipes as data stage gates. Typically, bCLEARer stage and thin slice pipelines have these gates. This dual (input and output) gate design for nested pipelines enables the original and transformed data to be inspected and compared. To assist with this, the design of an output gate's data also aims to clean it sufficiently to provide a SSOT (Appendix - Aggregated Single Source of Truth Principle (see page 46)) snapshot of its state at this stage of its digital journey. Where an input gate is an output gate of the previous process, it will also be in a SSOT snapshot. And in so doing, one can make the journey's transformations visible by comparing snapshots. In the bCLEARer process, these stage-gates are not decision points (as they are in some waterfall processes); though where data fails an inspection, it may be held back if required. So the focus is on inspection and not decision in an agile iterative process. Establishing data gates for nested pipelines, especially SSOT data gates, is a very cost effective way of creating inspection points that can greatly simplify improving and maintaining quality. It helps, for example, in the identifying of the source of problems: finding the first gate at which a problem appears, isolates between which gates it arose. One of the drivers in the design of the nesting structure is building a hierarchy of levels with structured gaps between gates that enable one to drill down into processes so finding problem data is easier. 1.3.3.2 Designing a gate Hence, gates are designed into the process. The design involves two layers of nesting. Given a pipeline that needs to be gated at one end, one needs to create a pipe that aggregates all the data flowing through that end of the pipeline. One way of doing this is creating a nesting that encapsulates all the filters with pipes that travel outside the pipeline. This will encapsulate the outgoing pipes into a single (aggregated) pipe, an example of this is given in the figures below. In this figure, the filters are just encapsulated. This results in the unaggregated output of two pipes — pipe 5 and pipe 6. -- 21 of 229 -- The bCLEARer Pipeline Architecture Framework 19 In this figure, the pipes are encapsulated first and then the filters. This results in the aggregated output of a single pipe — pipe 5&6. This diagram includes the gate icon, as shown below. -- 22 of 229 -- The bCLEARer Pipeline Architecture Framework 20 The data icon is typically adorned with a gate icon if the data is gated, as shown below. As usual, we can look at a nesting diagram to see the lower level structure hidden in the simple pipeline diagram - as shown in the figure below. -- 23 of 229 -- The bCLEARer Pipeline Architecture Framework 21 1.4 Architectural nesting breakdown 1.4.1 Introduction In the previous section, we noted that bCLEARer pipeline architecture has a hierarchical, levelled nesting structure. In this section, we look in more detail at this. The levels are an essential part of the stable bCLEARer infrastructure, the environment, within which the transformations evolve. Sometimes it is conceptually cleaner to describe the hierarchy from the bottom up starting from the base unit. In practice, the design often starts with the overall scope and decomposes this into base units. We follow this design practice in this section, where we describe this nesting by level from the bCLEARer Pipeline down. 1.4.2 Three core levels The bCLEARer pipeline’s primary nesting is based upon three core levels - described in the table below. levels descriptions pipeline the top, root, level of nesting for the bCLEARer pipeline - there is only a single root pipeline filter at this level. It is the upper boundary of the pipeline architecture. stage a single boundary layer (with no nesting) that contains the filters for the stages of the bCLEARer digital journey. bUnit the leaf nodes in the nesting structure, the bUnit filters. As this is a hierarchical tree structure, they belong to one, and only one, bCLEARer stage. It is the lower boundary of the pipeline architecture. 🔗 (see page 225) This can be visualised as a three-level breakdown. -- 24 of 229 -- The bCLEARer Pipeline Architecture Framework 22 🔗 1.4.3 Two extended levels The pipeline can be extended between these core levels with two further secondary nesting levels - described in the table below. levels descriptions thin slice where there are a significant number of sequences of bCLEARer stages in a pipeline, inserting this level between the pipeline and the stages enables the stages to be grouped into a more hierarchical modular structure. sub-stage where a bCLEARer stage contains a significant number of bUnits, inserting this level between the stages and the bUNits enables the bUnits to be grouped into a more hierarchical modular structure. 🔗 (see page 226) 1.4.4 Level type faceting The types of level in the pipeline nesting architecture can be classified under the three facets described in the table below. -- 25 of 229 -- The bCLEARer Pipeline Architecture Framework 23 Facet Elements Description kind core mandatory core levels in the architecture that set the framework extension optional extension levels that enable modularisation in the architecture modality mandatory a mandatory level in the pipeline nesting architecture optional an optional level in the pipeline nesting architecture nestability boundary slice a single level - with no nesting - in the pipeline nesting architecture nestable nesting is allowed - but not mandatory - in the pipeline nesting architecture 🔗 (see page 224) The types of level are classified under these facets as shown in the table below. level types kind modality nestability pipelines core mandatory boundary (single) thin slices extension optional nestable stages core mandatory boundary (single) sub-stages extension optional nestable bUnits core mandatory boundary (single) 🔗 (see page 223) -- 26 of 229 -- The bCLEARer Pipeline Architecture Framework 24 • • • 1.4.5 Core levels The bCLEARer pipeline’s primary nesting is based upon the three core levels described in this section. The next section describes the two extended levels. Pipelines level (see page 25) Stages level (see page 26) bUnits level (see page 27) -- 27 of 229 -- The bCLEARer Pipeline Architecture Framework 25 1.4.5.1 Pipelines level The pipelines level is the top, root, filter of the bCLEARer pipeline. It is the upper boundary of the pipeline architecture. Every bCLEARer pipeline has this root filter. -- 28 of 229 -- The bCLEARer Pipeline Architecture Framework 26 1.4.5.2 Stages level The (bCLEARer) stages are a single boundary layer (with no nesting) that contains the filters for the stages of the bCLEARer digital journey. bCLEARer provides a clear structure of five stages for the digital journey. These bCLEARer stages are sequential — and so the pipeline should flow in a sequence that reflects the journey. The sequence is well established (see first figure below). 🔗 The bCLEARer stages are usually gated — as shown in the figure below. 🔗 Sometime the pipeline involves manual work. In bCLEARer stage diagrams we usually mark the stages that involve manual work, using an icon — this is visible in the Load stage in the diagram above. The manual icon is shown below. manual icon As manual work interferes with running, scaling and costs, the aim is, as far as possible, to automate this work. Where it cannot be automated, it is often a good idea to move the manual work to the early stages in the sequence, typically restrict it to the Load stage. -- 29 of 229 -- The bCLEARer Pipeline Architecture Framework 27 1.4.5.3 bUnits level The bUnit filters are the leaf nodes in the nesting structure. Together, they form the lower boundary of the pipeline architecture. As this is a hierarchical tree structure, they belong to one, and only one, bCLEARer stage. 1.4.5.3.1 Base unit of transformation In the bUnit pipeline architecture, the bUnit pipes are the base units of identity and difference, and the bUnit filters the base units of transformation. In this architecture, the pipes transport data, the data is not transformed — so it is immutable stage of the data through the flow. The flow can then be seen as a sequence of bUnit immutable stages, where any transformation is located in the filters linking the stages. 1.4.5.3.1.1 Finer-grained level of datasets The pipes in the base bUnits work at the finer-grained level of datasets rather than dataset collections. In diagrams, the bUnit pipes are adorned with the dataset icon, as shown below. 1.4.5.3.2 Stage gates Where the bCLEARer stages are gated (as they usually are), the bUnits need to be designed to accommodate the gates — this is shown as start and end bUnits in the figure below, as are the finer-grained datasets. 🔗 Sometimes, unavoidably, a bUnit will be manual. This is marked using the manual icon — there is an example of this in the diagram above. -- 30 of 229 -- The bCLEARer Pipeline Architecture Framework 28 • • 1.4.6 Extended levels The pipeline can be extended between these core levels with two further extended levels, described in this section. Thin slices level (see page 29) Sub-stages level (see page 33) -- 31 of 229 -- The bCLEARer Pipeline Architecture Framework 29 1.4.6.1 Thin slices level Where the breakdown from the pipeline to the (bCLEARer) stages involves a significant number of stages, then it makes sense to modularise these stages into a nesting hierarchy of sub-pipeline filters. This is done in the thin slice extended level and built by nesting the stages in a hierarchy of thin slices. 1.4.6.1.1 Thin slice pipeline decomposition When designing the nesting structure for the thin slice level, it often makes sense to start with the overall bCLEARer pipeline and make a series of decompositions down until one reaches its boundary level of stages. A pro-forma example of a first stage decomposition of a bCLEARer pipeline into thin slices (filters), where the flow is ordered by pipes is shown in the figure below. 🔗 1.4.6.1.1.1 Thin slice nesting Where needed a thin slice can be further broken down, creating multiple layers of thin slice nesting. The breakdowns are ordered using pipes, reflecting the dependencies between the chunks - as shown in the pro- forma nesting diagram in the figure below. -- 32 of 229 -- The bCLEARer Pipeline Architecture Framework 30 🔗 1.4.6.1.2 Domain-based thin slicing It is not uncommon for the bCLEARer pipeline to cover more than one domain and in these cases the thin slice decompositions are often driven by domain considerations. For example, it is a common practice when working with multiple domains, to start bCLEARing each domain independently before merging them, especially so if the domains are sizeable. We illustrate this in the figure below, which involves two systems: GSAP and Sigraph. The first two thin slices transforms the two systems in isolation. The third thin slice merges and transforms the data from the two initial thin slices. Systems and their sub-systems often form natural boundaries for thin slices. Gates are typically placed at the start and end of each thin slice – as shown in the figure below using the dataset collection icon. 🔗 1.4.6.1.3 Intra-domain thin slicing With any sizeable domain, it often makes sense to divide it into sub-pipelines based upon smaller intra- domain chunks. Where there are dependencies between the chunks, these need to be recognised in the bCLEARer pipeline flow. -- 33 of 229 -- The bCLEARer Pipeline Architecture Framework 31 We illustrate this in the figure below, which again involves two systems: GSAP and Sigraph. This time we consider chunks within the system domains. Again, gates are typically placed at the start and end of each thin slice. 🔗 This figure shows the merging of data within systems as well as across systems. This is not unusual, often a significant amount of the bCLEARer pipeline deals with the merging of data from different sources - sometimes relating to the same domain or systems, other times across different domains or systems. 1.4.6.1.4 Thin slice bCLEARer sequences The leaves of a thin slice nesting hierarchy are bCLEARer stage (Stages level (see page 26) ) filters. Within the thin slice, these are typically ordered by pipes into a bCLEARer sequence - as shown in the example below. 🔗 Where there are multiple thin slices these are ordered by pipes into a series of sequences, where one sequence follows another — see example below. -- 34 of 229 -- The bCLEARer Pipeline Architecture Framework 32 🔗 1.4.6.1.5 Thin slice hierarchy evolution Also, the thin slice pipeline will typically evolve during a project — as shown in the example below. 🔗 -- 35 of 229 -- The bCLEARer Pipeline Architecture Framework 33 1.4.6.2 Sub-stages level Where a bCLEARer stage contains a significant number of bUnits, the sub-stages level is introduced to group the filters into a more hierarchical modular structure. 🔗 -- 36 of 229 -- The bCLEARer Pipeline Architecture Framework 34 1.5 Core design practices and patterns 1.5.1 Introduction This section deals with the general design of the bCLEARer pipeline to facilitate transparency and resilience. The bCLEARer pipeline is designed as a sequence of nested filters - the pipeline flow. It is important that this flow is not only transparent, open to inspection, but also resilient in the face of change. Hence we discuss the general design principles that will encourage this. -- 37 of 229 -- The bCLEARer Pipeline Architecture Framework 35 • • 1.5.2 Design principles We have found it useful to design the bCLEARer pipeline with a view to reducing the cost of evolving the pipeline - especially radically evolving - to take advantage of the opportunities that emerge on the journey of discovery. There are many well-known principles and techniques to achieve this. For example, focusing on building in modularity not only increases maintainability and reusability, but also improves refactoring and extensibility. These are well-documented, and we aim to point to these resources here and in the associated appendices. However, our focus here is on the less well-documented goal of facilitating transparency. One core central motivation for this is that delivering transparent transformations, and so enabling easy inspection, is a solid basis for improving the value of, as well as enhancing the trust in, the transformations. The aim is to systematise this inspection so reducing the effort needed to develop and maintain it during the rapid evolution of the pipeline. In the following sections, we aim to highlight here with some examples ways we can use the many well- known principles to better achieve our more focused goals. We start with a look at the general approach to decomposition, and then look at two principles: separation of transformation concerns immutability and idempotence There is a significant literature on these principles, so here we outline them and give references for further reading in the associated appendices. -- 38 of 229 -- The bCLEARer Pipeline Architecture Framework 36 1.5.3 General approach to decomposition In previous sections, we have visualised the modular structure of pipeline flow in terms of a flowchart to provide a picture of its structure. However, we need a different approach to decomposing it into filters. To reinforce this point, consider David L. Parnas’s classic Communications of the ACM paper (Parnas 1972 Decomposing Systems into Modules) (see page 139) where he makes this point, proposing an alternative to simple flowchart decomposition saying: “We have tried to demonstrate by these examples that it is almost always incorrect to begin the decomposition of a system into modules on the basis of a flowchart. We propose instead that one begins with a list of difficult design decisions or design decisions which are likely to change. Each module is then designed to hide such a decision from the others.” In the case of bCLEARer, the design decisions should consider, at the least, the various different kinds of information transformation. At the lowest level, the breakdown should end up with bUnits that deal with a single transformation - bUnits should not handle more than one transformation. This could be regarded as a "single-transformation principle" - alluding to the single-responsibility principle. For more details see: Appendix - Single-Transformation (Responsibility) Principle (STP) (see page 78). -- 39 of 229 -- The bCLEARer Pipeline Architecture Framework 37 1.5.4 Principle: separation of transformation concerns The separation of concerns (described in more detail in the Appendix (Appendix - Separation of Concerns Principle (see page 75))) is similar in many respects to single-transformation principle (see: Appendix - Single- Transformation (Responsibility) Principle (STP) (see page 78))). From a concerns perspective, the different types of transformation can be seen as different concerns, and so should be separated. As a first stage, this leads to base bUnits with a single concern. This enables cleaner tracing of transformations, as it also separates out the contributors to the transformation. This can lead to a significant number of base bUnits. When organising these base bUnits into larger modules, it is useful to consider principles of cohesion and coupling (there are more details on this in Appendix - Loose Coupling and Tight Cohesion Principle (see page 54)). Cohesion is a measure of functional closeness - and a good modular design will have high cohesion where the functional closeness is high inside the modules and low outside. Coupling is a measure of interdependence. A good modular design will have low coupling between modules and higher coupling inside the module. This is illustrated visually in the figure below. -- 40 of 229 -- The bCLEARer Pipeline Architecture Framework 38 1.5.5 Principle: immutability and idempotence Immutability and idempotence are about how transformations are handled. Immutable data’s content cannot be modified after it is created. An idempotent process can be run multiple times with the same input without changing the output. The pipeline architecture promotes idempotent processes (filters) and immutable datasets. We can illustrate this by comparing interactive and batch processing. The data updated by interactive processing is typically not immutable - the process changes its content. Whereas the data created by batch processing - the common pipeline process - is typically immutable. This is shown in the graphic below. So, adopting immutable data and idempotent processes (filters) is a natural choice for a pipeline architecture. And adopting it has several advantages. The immutability simplifies auditability, making transformations easier to track. The idempotence makes rerunning the processes safer and simpler. For more on these topics, see the appendix: Appendix - Immutability and Idempotence Principles (see page 50). -- 41 of 229 -- The bCLEARer Pipeline Architecture Framework 39 • • • • • • • • • • • • 1.6 Appendices Appendix - The Standard 'Pipeline' or 'Pipe-and-Filter' Architecture (see page 40) Appendix - Aggregated Single Source of Truth Principle (see page 46) Appendix - Glossary of Major Terms (see page 47) Appendix - Design Patterns and Anti-patterns (see page 49) Appendix - Immutability and Idempotence Principles (see page 50) Appendix - Loose Coupling and Tight Cohesion Principle (see page 54) Appendix - Pipeline Architecture Iconography (see page 57) Appendix - Resilient Governance - The Right Balance of Principles and Rules (see page 69) Appendix - Separation of Concerns Principle (see page 75) Appendix - Single-Transformation (Responsibility) Principle (STP) (see page 78) Appendix - Historic Pipeline Examples (see page 79) Appendix - Further Related Topics (see page 99) -- 42 of 229 -- The bCLEARer Pipeline Architecture Framework 40 • • • 1.6.1 Appendix - The Standard 'Pipeline' or 'Pipe-and-Filter' Architecture 'Pipe-and-filter' is a software architecture pattern. The pattern is used to break down a complex monolithic process into individual and reusable components. Doing so can improve performance, scalability, and reusability by allowing task elements that perform the processing to be deployed and scaled independently. The name "pipeline" comes from a rough analogy with physical plumbing in that a pipeline usually allows information to flow in only one direction, like water often flows in a pipe. This appendix discusses the standard 'Pipeline' or 'Pipe-and-Filter' Architecture in the following three sections: The standard 'Pipeline' or 'Pipe-and-Filter' Architecture - History (see page 41) The standard 'Pipeline' or 'Pipe-and-Filter' Architecture - Pipe-and-filters breakdown approach (see page 42) The standard 'Pipeline' or 'Pipe-and-Filter' Architecture - Standard Textbook Example (see page 44) -- 43 of 229 -- The bCLEARer Pipeline Architecture Framework 41 1. 2. 3. 1.6.1.1 The standard 'Pipeline' or 'Pipe-and-Filter' Architecture - History Doug McIlroy introduced the ‘pipe-and-filter’ architecture in Unix in 1972. (Ritchie 1984 The Evolution of the Unix Time-Sharing System) (see page 150) gives a detailed description of this. It subsequently became known as ‘pipeline’ architecture. Since then it has become an accepted architecture or architectural pattern and documented in the standard textbooks; an example is given at the end of the section. 1.6.1.1.1 Compliers Particular applications often have a more detailed architecture. For example, compilers typically have a three stage pipeline architecture: front-end: parses input language into an intermediate language middle: performs transformations in the intermediate language back-end: translates the intermediate language into the output language The use of an intermediate language allows for reusable architectural components. -- 44 of 229 -- The bCLEARer Pipeline Architecture Framework 42 1.6.1.2 The standard 'Pipeline' or 'Pipe-and-Filter' Architecture - Pipe-and-filters breakdown approach A straightforward approach to implementing an application is to perform each processing task in a module. This results in an inflexible, monolithic architecture, which is likely to reduce the opportunities for reusing the code and so enhancing the value of refactoring and optimizing it. The diagram below illustrates this monolithic approach. The two modules were designed separately and the code that is closely coupled to its module. A pipe-and filter architecture breaks down the processing for each stream into a set of separate filters, each performing a single task. The filters are then combined into a pipeline. This not only avoids code duplication, but makes it easy to remove or replace or integrate additional filters, if the requirements change. This diagram below shows a solution that's implemented with pipes and filters: -- 45 of 229 -- The bCLEARer Pipeline Architecture Framework 43 • • • • • 1.6.1.2.1 Recognised benefits As the diagrams above help to illustrate, there are a number of recognised benefits of using this pattern: Ensures loose and flexible coupling of pipes and filters. Loose coupling allows filters to be changed without modifications to other filters. Conductive to parallel processing. Filters can be treated as black boxes. Users of the system don’t need to know the logic behind the working of each filter. Re-usability. Each filter can be called and used over and over again. -- 46 of 229 -- The bCLEARer Pipeline Architecture Framework 44 1.6.1.3 The standard 'Pipeline' or 'Pipe-and-Filter' Architecture - Standard Textbook Example (Bass 2012 Software Architecture in Practice) (see page 121) is one of the standard textbooks on software architecture. It describes architectures in terms of patterns. This extract provides the ‘standard’ view of pipe- and-filter architecture. The standard pipe and filter pattern Pipe-and-Filter Pattern Context: Many systems are required to transform streams of discrete data items, from input to output. Many types of transformations occur repeatedly in practice, and so it is desirable to create these as independent, reusable parts. Problem: Such systems need to be divided into reusable, loosely coupled components with simple, generic interaction mechanisms. In this way they can be flexibly combined with each other. The components, being generic and loosely coupled, are easily reused. The components, being independent, can execute in parallel. Solution: The pattern of interaction in the pipe-and-filter pattern is characterized by successive transformations of streams of data. Data arrives at a filter's input port(s), is transformed, and then is passed via its output port(s) through a pipe to the next filter. A single filter can consume data from, or produce data to, one or more ports. Data transformation systems are typically structured as pipes and filters, with each filter responsible for one part of the overall transformation of the input data. The independent processing at each step supports reuse, parallelization, and simplified reasoning about overall behaviour. Often such systems constitute the front end of signal-processing applications. These systems receive sensor data at a set of initial filters; each of these filters compresses the data and performs initial processing (such as smoothing). Downstream filters reduce the data further and do synthesis across data derived from different sensors. The final filter typically passes its data to an application, for example providing input to modelling or visualization tools Key terms terms descriptions Overview Data is transformed from a system's external inputs to its external outputs through a series of transformations performed by its filters connected by pipes. -- 47 of 229 -- The bCLEARer Pipeline Architecture Framework 45 terms descriptions Elements Filter, which is a component that transforms data read on its input port(s) to data written on its output port(s). Filters can execute concurrently with each other. Filters can incrementally transform data; that is, they can start producing output as soon as they start processing input. Important characteristics include processing rates, input/ output data formats, and the transformation executed by the filter. Pipe, which is a connector that conveys data from a filter's output port(s) to another filter's input port(s). A pipe has a single source for its input and a single target for its output. A pipe preserves the sequence of data items, and it does not alter the data passing through. Important characteristics include buffer size, protocol of interaction, transmission speed, and format of the data that passes through a pipe. Relations The attachment relation associates the output of filters with the input of pipes and vice versa. Constraints Pipes connect filter output ports to filter input ports. Connected filters must agree on the type of data being passed along the connecting pipe. Specializations of the pattern may restrict the association of components to an acyclic graph or a linear sequence, sometimes called a pipeline. Other specializations may prescribe that components have certain named ports, such as the stdln, stdout, and stderr ports of UNIX filters. Weaknesses The pipe-and-filter pattern is typically not a good choice for an interactive system. Having large numbers of independent filters can add substantial amounts of computational overhead. Pipe-and-filter systems may not be appropriate for long-running computations. 🔗 (see page 123) -- 48 of 229 -- The bCLEARer Pipeline Architecture Framework 46 1.6.2 Appendix - Aggregated Single Source of Truth Principle The single source of truth (SSOT) principle is also called single point of truth (SPOT) principle. 1.6.2.1 Two versions of the principle There are syntactic and semantic (or ontological) versions of this principle. In the syntactic version, this is the practice of structuring information models and associated data schemas such that every data element is mastered in only one place. In the semantic (or ontological) version, there is a further assumption that objects in the real world are only represented once by a single data element. 1.6.2.2 Aggregated single source of truth The need for an aggregated SSOT arises when there are multiple systems. While each of the systems may have implemented a SSOT policy, new requirements emerge when the information in the systems are aggregated. From a syntactic perspective, one can often identify when data elements in different systems are (in some sense) the same - and set up processes to aggregate them. From a semantic (ontological) perspective, there is a requirement to aggregate (syntactically) different data elements that represent the same object in the real world. Identifying where the same real world object is being represented is often greatly simplified by adopting a top ontology. 1.6.2.3 BORO and bCLEARer’s approach bCLEARer projects often work with multiple systems - in fact, this is encouraged as the use of multiple systems increases the quality of the resulting ontology. bCLEARer’s integrated use of the BORO Top Ontology greatly simplifies semantic aggregated SSOT. Its integrated use of object-identity enables transparent inspection of the aggregation process. 🔗 (see page 46) -- 49 of 229 -- The bCLEARer Pipeline Architecture Framework 47 1.6.3 Appendix - Glossary of Major Terms 1.6.3.1 Pipeline architecture terms descriptions bCLEARer pipeline A nested pipeline that has thin slices, bCLEARer stages and bUnits as its decomposition (nesting) levels. filter See pipe-and-filter. gate (gated pipeline) Also known as data stage gate or information gate, this is a pipe that has been designed to enable all the original and transformed data to be inspected and compared. nest (nested pipeline) A pipeline with collapsed connected sub-pipelines. pipe See pipe-and-filter. pipe-and-filter Pipe-and-filter is a sequence of processing components (filters), arranged so that the output of each component is the input of the next one, connected by pipes, creating a ‘flow’. This is also known as a pipeline. pipeline See pipe-and-filter. sub-pipeline Any subset of filters in a pipeline. pipeline topology The network structure of the relations between pipes and filters in the pipeline 1.6.3.2 Nested levels terms descriptions 3 core level architecture The bCLEARer pipeline’s 3 core level architecture pipeline level The top level of the 3 core level architecture. A single, root filter. -- 50 of 229 -- The bCLEARer Pipeline Architecture Framework 48 terms descriptions (bCLEARer) stage level The middle level of the 3 level architecture. bUnit level The bottom level of the 3 level architecture. thin slice level One of the two extended levels to the 3 core level architecture. sub-stage level One of the two extended levels to the 3 core level architecture. bCLEARer stage The base units (filters) of the bCLEARer stage level of the bCLEARer Pipeline. Each unit is one of bCLEARer’s five stages of evolution. In sequence these are: Collect, Load, Evolve, Assimilate and Reuse. bCLEARer (stage) sequence The composite units (filters) of the bCLEARer stage level that contains some or all of the bCLEARer stage filters in sequence (see bCLEARer stage). bUnit The base units (filters) of the bUnits level of the nested bCLEARer Pipeline, and so the lowest, bottom units of the pipeline. thin slice Composite unit (filters) that contains either other thin slices or stages. sub-stage Composite unit (filters) that contains either other sub-stages or bUnits. 🔗 (see page 47) -- 51 of 229 -- The bCLEARer Pipeline Architecture Framework 49 • • 1.6.4 Appendix - Design Patterns and Anti-patterns As Wikipedia (https://en.wikipedia.org/wiki/Design_pattern) notes, a design pattern 'is the re-usable form of a solution to a design problem'. Each pattern describes a problem that occurs over and over again in our environment, and then describes the core of the solution to that problem, in such a way that you can use this solution a million times over, without ever doing it the same way twice. The range of situations in which a pattern can be used is called its context. In software engineering, a software design pattern is a general, reusable solution to a commonly occurring problem within a given context in software design. It is not a finished design that can be transformed directly into source or machine code. Rather, it is a description or template for how to solve a problem that can be used in many different situations. Design patterns are formalized best practices that the programmer can use to solve common problems when designing an application or system. 1.6.4.1 Anti-pattern An anti-pattern in software engineering, project management, and business processes is a common response to a recurring problem that is usually ineffective and risks being highly counterproductive. As opposed to a bad practice An anti-pattern is a commonly-used process, structure or pattern of action that, despite initially appearing to be an appropriate and effective response to a problem, has more bad consequences than good ones. Another solution exists to the problem the anti-pattern is attempting to address. This solution is documented, repeatable, and proven to be effective where the anti-pattern is not. 🔗 (see page 49) -- 52 of 229 -- The bCLEARer Pipeline Architecture Framework 50 • • • 1.6.5 Appendix - Immutability and Idempotence Principles Immutability and idempotence are two inter-related software principles. They both deal with how data transformations are handled: immutability is concerned with the data content and idempotence is concerned with its processing. Both the immutability and idempotence principles play significant roles in software engineering, providing guidance on designing more predictable, fault-tolerant, and maintainable systems. It is recognised that together they contribute to creating robust architectures, especially in environments where consistency, reliability, and performance are paramount. This appendix has three subsections, linked below. The first discusses immutability and the second idempotence. The final section explains how these principles are used within bCLEARer (and BORO) projects: Immutability principle (see page 51) Idempotence principle (see page 52) Immutability and idempotence in bCLEARer (and BORO) (see page 53) -- 53 of 229 -- The bCLEARer Pipeline Architecture Framework 51 1.6.5.1 Immutability principle The immutability principle applies to objects, typically in object-oriented (OO) and functional programming, but can easily be extended to ‘database objects’. The principle is that the object’s state is immutable - it does not change. An example of immutability in computation would be a constant - such as the value of pi. Constants do not change during their lifetime in the pipeline. One key benefit of immutability is simpler auditability. Having immutable data gives the content a clear identity making possible easy tracking through the processing it is involved in. Immutability reduces temporal coupling which has various downstream benefits in software design. One of these benefits is that it enables multithreading by promoting thread safety - programs that have a lower coupling on when they need to execute with respect to each other may be more safely multi-threaded. Usually we think immutable objects can be immutable simpliciter, globally immutable. However, there is a weaker notion of immutable, where objects are immutable with respect to a process. Repeated applications of this process to the immutable object must produce the same result. However, there may be other processes that change the state of the object. An object is globally immutable if there is no process that changes its state. The notion of immutable needs a corresponding notion of (state) change and (object) creation. The creation (transaction) of an object may be an intricate affair, During the creation its attributes (its state) may change. The major attributes may be created as a first stage and the minor ones subsequently: this is not considered state change. The extent of the immutability can be extended. This is easier to see with data objects. So not just a row (corresponding to an object) can be immutable, but the whole table can be immutable. -- 54 of 229 -- The bCLEARer Pipeline Architecture Framework 52 1.6.5.2 Idempotence principle The idempotence in software engineering is a property of a process (operation), where the process (operation) can be applied multiple times without changing the result beyond the initial application. This principle is crucial for designing robust, fault-tolerant systems, particularly in distributed systems, RESTful APIs, and database transactions. The principle of idempotence has a rich history in computer science and software engineering, with its roots in mathematics and logic. Understanding its evolution provides insight into its broad applicability and importance across different areas of technology. The concept of idempotence originates from mathematics, specifically algebra, where an operation is idempotent if applying it multiple times has the same effect as applying it once. For example, multiplying a number by one or taking the absolute value of a value are idempotent operations. The term was introduced by the mathematician Benjamin Peirce in the context of elements of algebras that remain invariant when raised to a positive integer power, and literally means "(the quality of having) the same power", from (idem + potence = same + power). In formal logic and set theory, the concept of idempotence has been applied to operations and functions. It was recognized that certain operations, when applied multiple times, do not change the outcome beyond the initial application. This concept was naturally extended into computer science, particularly in the development of algorithms and data processing techniques. As we discuss in the next section (Immutability and idempotence in bCLEARer (and BORO) (see page 53)), bCLEARer bUnits are analogous to idempotent operations and functions. Idempotence has also influenced software design patterns and practices. It underpins many aspects of software development, including error handling, state management, and the implementation of idempotent functions and services. This ensures that systems are more resilient to failures and can recover more gracefully from errors. Idempotence and immutability are related. Where an object is immutable with respect to a process, then the the process is idempotent with respect to the object. Repeated applications of the process to the immutable object must produce the same result. For instance, the program for calculating the area of a circle that treated pi as an (immutable) constant would not change the value of pi. The operation (calculating area) is idempotent (concerning pi). -- 55 of 229 -- The bCLEARer Pipeline Architecture Framework 53 1.6.5.3 Immutability and idempotence in bCLEARer (and BORO) It is generally accepted that, following idempotence and immutability in system design reduces the costs associated with complex systems, leading to easier maintainability, easier development, less costly refactoring, easier auditability. It also reduces the risks associated with complex systems leading to safer processing of data. Understandably, given these benefits bCLEARer adopts these principles. The resulting design of the bCLEARer pipeline has a “batch processing” structure, where data is input to a process and different data is created by the process - the input data is never updated. This ensures idempotence. As noted in the section, Immutability principle (see page 51), one needs to be clear of the extent of the creation transaction. In the bCLEARer pipeline, the creation transactions occur in the base bUnits. After a dataset is output by the base unit, it is immutable. This is an example of immutability extended to tables (mentioned in the earlier section, Immutability principle (see page 51)). -- 56 of 229 -- The bCLEARer Pipeline Architecture Framework 54 1.6.6 Appendix - Loose Coupling and Tight Cohesion Principle Loose Coupling and Tight Cohesion is one of the guiding principles of bCLEARer (see Appendix - Resilient Governance - The Right Balance of Principles and Rules (see page 69)). 1.6.6.1 History Larry Constantine is considered to be the father of these concepts. It was first publicly described in his paper with Stevens and Myers: (Stevens 1974 Structured Design) (see page 152). It was originally described for procedural programming modules. It was later extended to cover object oriented programming classes. 1.6.6.2 What are cohesion and coupling? The terms typically apply to the design of modules for software. Cohesion is a measure of how closely related parts of a module are functionally. One (rough) way of thinking of this is that a module with parts that work together for the same goal has high cohesion. Coupling is a measure of modules' (inter)dependance with one another. One (rough) way of thinking of this is that modules that are loosely coupled need each other less to function. This is shown graphically below. 🔗 This table summarises these points. cohesion coupling Cohesion is typically an intra-module concept. Coupling is typically an inter-module concept. Cohesion is a measure for the relationships
Overview
bCLEARer stands at the forefront of digital transformation, championing an evolutionary approach to harnessing digitization and digitalization opportunities. It guides information on a transformative journey, curating its evolution into fitter forms, ones more suited for computing, that deliver increased value.
To accomplish this, bCLEARer has evolved an architecture framework for semantic data pipelines, along with a methodology for engineering these pipelines.
A while ago, a client engaged us to crystallise our then current working documentation on the architecture framework into an eManual for their bCLEARer programme. What follows is a sanitized version of that manual, capturing the architecture as it existed then. Though bCLEARer continues to evolve, the core principles in this eManual remain relevant.