Ontology-Driven Information Systems Engineering Workshop (ODISE) CAiSE 2011 20 th June 2011 Chris Partridge www.borosolutions.co.uk Underlying Motivation 2 www.gartner.com/it/page.jsp?id=1513614 The ‘How?’ question How to deploy ontology in Enterprise Software ? Two approaches: Speculating Mining The ‘Where?’ Question If ontology is to have more influence, where should it expand? Spend on Enterprise Software ($254bn) is orders of magnitude bigger than Semantic Web A small slice of Enterprise Software will be bigger than a very large slice of Semantic Web style.visibility ppt_x ppt_y Ontology Mining versus Ontology Speculation When we embed the building of an ontology into an information system development or maintenance process, then the question arises as to how one should construct the content of the ontology. One of the choices is whether the construction process should focus on the mining of the ontology from existing resources or should be the result of speculation (‘starting with a blank sheet of paper’). I present some arguments for choosing mining over speculation and then look at the implications this has for application modernisation. 3 Two approaches to building ontology content 4 Introspect – look inside one’s head – at the representation of the domain there. AND. Ask an SME to introspect. Start with a blank sheet of paper. Unconstrained by existing (external) representations (path dependency). Thinking outside the box. Rationalism Inspect – look outside one’s head at pre-existing representations of the domain. AND. Look at the domain itself. Build on existing foundations. Not re-inventing the wheel. Stand on the shoulders of giants . Empiricism Speculation Mining style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y Plan Set the context Domain Ontology Context Application Development Context Argue for mining Transparent vision Formal models 5 Context “ When we embed the building of an ontology into an information system development or maintenance process ” Domain Ontology Context Application Development Context 6 Domain Ontology Context Need to be refine what is meant by “the building of an ontology” How? Stand back and look at information modelling (Ontology can be seen as a kind of information modelling) Enables us to separate two key concerns 7 Separating two concerns 8 What is in the domain? How do we represent it? The content developer perspective: Work out what is in the domain and then represent it. Direction of focus is domain-to-model. The ontology tool developer perspective: Work out how to build a model whose icons can represent things in the domain. Direction of focus is model-to-domain. Domain (represented) Model (representation) style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y Leads to two ways of defining an ontology 9 Domain (represented) Model (representation) The ontology is the domain. An ontology is "the set of things whose existence is acknowledged by a particular theory or system of thought." (E. J. Lowe, The Oxford Companion to Philosophy) The limit case: The question ontology asks can be stated in three words ‘ What is there?’ and the answer in one ‘ everything’. Not only that, “everyone will accept this answer as true” though “there remains room for disagreement over cases.” (Quine “On What There Is ”) The ontology is the model. An ontology is an explicit specification of a conceptualization . A body of formally represented knowledge is based on a conceptualization ( Genesereth & Nilsson, 1987 ) A conceptualization is an abstract, simplified view of the world that we wish to represent for some purpose. Every knowledge base, knowledge-based system, or knowledge-level agent is committed to some conceptualization, explicitly or implicitly . (Tom Gruber) style.visibility style.visibility style.visibility Concern dependency From the development perspective the direction of fit is domain to model: Answers to the question ‘What is in the domain?’ raise requirements for the model/representation In other words, if it is in the domain, then the model needs to be able to represent it. Examples: Does the domain contain: type-instance relations? If so, we need to represent them. John’s car is a Mini (a type-instance relation) types whose instances are also types? If so, we need to represent them. Minis are makes of cars (see Business Objects (Partridge 1996) Answers to the ‘How do we represent it?’ question depend upon the answers to the ‘What is in the domain?’ question So we cannot answer the representation question without answering the domain question 10 Domain Ontology C ontext Broad focus is on the ‘What is in the domain?’ question A domain ontology can (should?) be embedded in a top ontology. This gives us two levels: Top ontology Domain ontology Specific focus on the domain ontology 11 Application Development Context Need to refine what we mean by “an information system development or maintenance process” Two aspects to clarify: Shift from greenfield to brownfield Application layer 12 Increasing Automation 13 Increasing coverage Emerging automation: occupying the automatable space style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y From greenfield to brownfield 15 Greenfield Brownfield A problem space needing the deployment of software applications where there are NO existing (legacy) software applications. These typically develop from scratch, on the basis of a "clean sheet of paper ". A problem space needing the deployment of new software applications where there are pre-existing (legacy) software applications. There may be an opportunity to harvest the investment in the pre-existing application systems. A classic example of an enterprise legacy application: SABRE (Semi-Automatic Business Research Environment), a computer reservation system initially developed by American Airlines in the late 1950s – going live in 1960. The development cost at that stage was $ 40 million ( about $350 million in 2000 dollars) http://en.wikipedia.org/wiki/Sabre_(computer_system) style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y Which application renewal approach? 16 Re-development Modernisation The re-development of a legacy system on a modern technology platform. It aims to avoid the out-of-date architecture inherent in the legacy application. The reengineering, conversion or porting of a legacy system to a modern technology platform. It aims to retain and extend the value of the legacy investment through migration to new platforms In a brownfield site, (at least) two possible approaches to renewing the system. Application Layer Brings the two previous contexts together Argue that: ‘Domain Ontology’ is an application layer. In a brownfield situation An application modernisation option is to migrate the domain ontology layer 17 Information (architecture) Framework 18 Based upon a separation of concerns (similar to OMG’s MDA divisions) Layered concerns, where each concern builds upon all previous concerns. Domain Ontology Layer 19 Domain Ontology Layer style.visibility ppt_x ppt_y Opportunity for migration 20 Legacy system creates an application modernisation opportunity At the domain layer level: Speculate = Re-development Mine = Modernise Our route 21 Fundamental argument is ‘do not reinvent the whee l ’ Large complex systems have significant investment in them Significant proportion of that is in the domain layer Domain layer typically is not changing So, why re-develop it from scratch? Possibly even ´re-inventing the square wheel ´, where the re-development project does not have sufficient budget to recreate the complexity of the legacy system domain layer style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y No brainer 22 Which would you choose? Complication: there are no layers 23 If you are going to mine, then you need to have (or build) the tools to extract the domain layer from the legacy system. Speculation seems more attractive. In most legacy systems, there are no layers! Updated route 24 Why bother mining, when … ? Our goal: Answer: What is in the domain? Isn’t this easy to answer? Can’t we see the objects in the domain directly? Isn’t it just a question of looking or remembering? 25 Can’t we see the objects in the domain directly? 26 Now... that should clear up a few things around here Everyone knows that if you put 10 data modellers to work on the model for a fixed domain, you will end up with 12 different models. And none of these will really be a model of the real world. Certainly the modellers will not agree which one is. How do we explain this? style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y style.visibility ppt_x ppt_y Explaining the situation Another way of asking the question: When I look at a domain, is my vision (the process of seeing) transparent? If it is transparent then it does not get in the way of seeing the domain – I see the domain directly Why in some cases is it transparent and in others not? Look in the literature, we find the answer Transparent vision is constructed When we get used to it, it feels natural – but it is constructed Naïve view is that it is just natural (Larson is mocking this) Different visions for different purposes: E.g. Medical transparent vision – it would look obvious to a trained doctor Has the ontological community/profession constructed a transparent ontological vision? 27 Transparent vision: a well researched area Modern discussion: (E.g.) Goodwin, Charles (1996). Transparent Vision. Interaction and Grammar. E. Ochs, E. A. Schegloff and S. Thompson. p. 370-404. Not a new issue Modern discussion of the history of this topic. Edward Craig, The Mind of God and The Works of Man (1987) Often assumed that it is natural – this is often challenged. For example, David Hume, A Treatise of Human Nature (1740) challenges it The ‘Image of God’ Doctrine: Man is made in God’s image, and that although human beings are far less perfect than God, human minds and God’s mind are the same kinds of thing. ‘The Insight Ideal’ is derived from the ‘Image of God’ Doctrine: God in his goodness endowed human beings with faculties that enable them, in principle, to gain knowledge of the world he created for them. It is totally taken for granted that ‘the universe was in principle intellectually transparent …’ (Craig 1987, 38) 28 Developing transparent ontological vision Looks like transparent vision (perception) is constructed not natural A profession can develop a transparent vision – e.g. doctors may see the same medical things Is the ‘professional ontologists ’ vision transparent? If we take the data modellers case seriously, then plainly not What does this mean for ontological domain modelling? We need to work out how to: construct the transparent professional vision first stage, make the community realise that their vision is not transparent Note. In a sense, developing the ontology and developing transparent vision is the same exercise. 29 Speculate or mine Negative argument: If the community’s ontological vision were transparent, then speculation would reveal the domain objects directly There would be little incentive to mine As it is not, then speculation is not the obvious alternative Deciding factor is whether speculation or mining is the best way to develop an ontological transparent vision 30 Formal models One factor that sets apart the application and the ontological domain model from other everyday ‘visions’ is that it needs to be formal This also sets apart mathematical ‘visions’ Mathematics is the exemplar of formality 31 Formalising properly takes work Major examples: Mathematical formal contradictions: Frege Basic Law V and Russell’s Paradox Quine The Burali-Forti Paradox in the First Edition of Quine’s ML If we are speculating, then when building an ontology we have to bear the cost of formalising our pre-formal notions If we are mining, and what we are mining is an already formalised (and tested) structure, and our mining preserves the formality, then we do not have to bear the cost of formalising 32 Stages in formalising In the history of mathematics, a pattern of two stage formalising emerges: Finding a formal solution, Making the solution elegant. 33 Succession of proofs (1) It is an article of faith among mathematicians that after a new theorem is discovered, other simpler proofs of it will be given until a definitive one is found. A cursory inspection of the history of mathematics seems to confirm the mathematician’s faith. The first proof of a great many theorems is needlessly complicated. “Nobody blames a mathematician if the first proof of a new theorem is clumsy”, said Paul Erdős . It takes a long time, from a few decades to centuries, before the facts that are hidden in the first proof are understood , as mathematicians informally say. This gradual bringing out of the significance of a new discovery takes the appearance of a succession of proofs, each one simpler than the preceding. New and simpler versions of a theorem will stop appearing when the facts are finally understood . p.146 , The Phenomenology of Mathematical Proof, Gian -Carlo Rota 34 Succession of proofs (2) ... Beauty is seldom associated with pioneering work in mathematics. The first proof of a difficult theorem is seldom beautiful, and its discoverer is never blamed for his or her awkwardness: it usually takes generations of mathematicians to refine the first proof of a theorem to the point where the theorem will be viewed as beautiful. The first proof of the ergodic theorem, a lengthy and clumsy proof, was given by G.D. Birkhoff in 1932; the definitive one-page proof was obtained by Garsia in 1965. In the intervening period, one hundred-odd other proofs were given, including one by the present speaker. p. 178 GIAN-CARLO ROTA - THE PHENOMENOLOGY OF MATHEMATICAL BEAUTY 35 Useful analogies Useful analogies with formalisation in mathematics in building ontologies Formalisation takes work In practice, this work is done in stages: First work out a formal proof, Then refine it ( understand it) In our context, If we mine: the legacy system has done the work of the first stage If we speculate: we need to do both stages 36 Summary Given: A decision to build or use a domain ontology, and A situation where there is a brownfield site. Then: There are good reasons to think a mining approach may bring better results than a speculative approach Overall ‘don’t reinvent the (square) wheel’ argument One needs to construct an ontologically transparent vision for the domain (because it does not exist yet) A big challenge is producing a formalised representation of the domain One can reduce the cost of this by mining / harvesting the formal model from the legacy system BUT … One needs mining (harvesting) processes and tools that are not too costly and preserve the good features of the legacy system So, this is a promising area of research!!! 37
BORO Publications
Ontology Mining versus Ontology Speculation
19 June 2011Presented at Keynote - ODISE III 2011, 3rd International Workshop on Ontology-Driven Information Systems Engineering, colocated with CAiSE 2011, 20-24 June 2011, London, UK
Overview
When we embed the building of an ontology into an information system development or maintenance process, then the question arises as to how one should construct the content of the ontology. One of the choices is whether the construction process should focus on the mining of the ontology from existing resources or should be the result of speculation (‘starting with a blank sheet of paper’). I present some arguments for choosing mining over speculation and then look at the implications this has for legacy modernisation.
