# The End of Faith is Stupid Sam Harris is the Dumbest Philosopher of All-Time October 1, 2023 Greg Coppola # Introduction Sam Harris’ incredible career of sophistry has visited many new heights over the years. But, it all began with his first exercise in confusion: ****************The End of Faith**************** (2004). In this book, Harris argues that *faith* is ***bad***. Actually, *****faith*****, is a necessary thing. In this article, we look at: 1. how goofy Sam Harris defines *****faith***** as “belief without evidence”, a classic New Atheist-style *********straw man********* 2. how *********faith*********, defined as “willingness to act on a theory despite 100% certainty”, is absolutely necessary Sam Harris’ identification of *****faith***** with a non-sensical straw man (belief without evidence), obscures the necessary role that faith plays in our every day lives. One ***may*** operate in the world without faith in ***God***. But, everyone must have faith in *********something*********. This is because, as Hume told us, no amount of experience with the past can ever entitle us to predict the future, a fact proven by Wolpert, Macready et al. relatively recently, in the 1990’s. # Harris’ Definition of Faith Sam Harris is, as a general rule, *****never***** clear about his definitions. This is probably calculated: the less clear he makes himself, the harder he is to expose as an intellectual fraud. Overall, in patchy allusions to the idea of *****faith***** in the book ************End of Faith************, we can reconstruct the following definition: - **Sam Harris’ definition of *faith*** - **faith** is *belief* without *evidence* Let us look at some quotes from Harris’ ************End of Faith************ to illustrate that this is indeed his definition. In the following passage, Harris literally equates faith with “an act of knowledge that has a low degree of evidence”: > The faith that I am calling into question is precisely the gesture that Tillich himself decried as “an act of knowledge that has a low degree of evidence.” (Sam Harris, End of Faith) > In the following, passage he notes that faith is “the idea that belief can be sanctified by something other than evidence”, in other words reaffirming that faith is belief without evidence: > The concessions we have made to religious faith—to the idea that belief can be sanctified by something other than evidence—have rendered us unable to name, much less address, one of the most pervasive causes of conflict in our world. (Sam Harris, End of Faith) > In the following passage, Harris equates belief with “unjustified belief”: > [T]he truth is that religious faith is simply unjustified belief in matters of ultimate concern. (Sam Harris, End of Faith) > Finally, “faith is what credulity becomes” when it “achieves escape velocity” from “terrestrial discourse”: > Faith is what credulity becomes when it finally achieves escape velocity from the constraints of terrestrial discourse—constraints like reasonableness, internal coherence, civility, and candor. (Sam Harris, End of Faith) > # The Necessary Role of Faith We have seen that Sam Harris’ definition of “faith” is “belief without evidence”, a classic Harris “straw man” characterization of a position, so that he has something simple that his tiny mind can conclusively “demolish” in an rgument. In fact, as we will see, there is a *********necessary********* definition of faith: - **necessary definition of “faith”** - the willingness to ***act*** without 100% certainty in one’s *belief model* This is the definition that would probably empirically fit most believers in religion, because, while they intuit the existence of a God, and must act on these beliefs, including in ways that might cause strife or loss, even though they cannot be 100% sure of the existence of God. However, if Sam Harris were to say it this way, where faith is the willingness to act in the face of doubt, he would obviously lose the argument on the spot. This is because every successful person knows that they have had to act on *****faith***** at some point in their careers. # Faith Among the Great Philosophers In general, the fact that we can never have 100% certainty in any propositions has been a major theme in philosophy, and, we include, mathematics. ## Descartes According to the legendary philosophy expositor Bryan Magee, the concept of *modern philosophy*, as apart from *ancient* or *medieval* philosophy, is defined in terms of the life span of René Descartes (1596—1650), who was the first “modern” philosopher, initiating the “modern” period. Such is the importance of Descartes to Western Philosophy. Overall, Descartes himself believed in God, as we will see. And, we will see that Sam Harris evidently never learned the lessons of In his classic treatise ****Meditations**** (in full, *Meditations on First Philosophy in which the existence of God and the immortality of the soul are demonstrated*), Descartes famously sought to investigate exactly what could be known without **doubt**. Descartes resolved to only believe something if it *********could not be doubted********* at all. This led to his famous ***************method of doubt***************. The goal is to figure out which believes ************************cannot be doubted at all************************. And, this turns out to be a very small number of beliefs, ************************************much smaller than those needed to operate in the world************************************. He considered increasingly severe thought experiments, in which his experiences could increasingly not be trusted, culminating in an experiment in which he supposed himself interacting with a “malicious, powerful, cunning demon”, who manipulates his sense perceptions—this is like the plot of the movie **********The Matrix********** (1999)—so that Descartes would believe he was living a life completely different than where his body really was. Descartes concludes that, no matter what he does, he ************cannot prove************ that his existence is not actually just a clever ruse played by the evil demon. In more modern parlance, we can never prove that we are not living in the plot of a movie like **********The Matrix********** (1999). Descartes’ famous conclusion in ***************The Meditations***************, actually turns out to be quite limited: he concludes that one thing he ***********cannot*********** doubt is that he “thinks” (or “cognizes”, from the Latin ******cogito******). Thus, the most famous phrase in the history of Western Philosophy: “I think therefore I am”, or, in Latin, “cogito ergo sum”. Basically anything beyond the assumptions that “I think”, “I cogitate”, or “I am” requires, according to Descartes, assumptions that ****************could be doubted****************. In other words, on questions like whether or not reality actually exists (versus being a The Matrix-style illusion), or whether the laws of physics will be the same “tomorrow” as they are today, we cannot *****prove***** that the universe exists and is orderly beyond any reasonable doubt. In other words, we cannot prove with 100% certainty that we really exist, or that the universe will continue to operate the same way tomorrow. In other words yet again, in order to act in the world, we must have *****faith***** in the truth of propositions that one could, if one were determined, doubt, and no amount of evidence can ever allay this doubt. Overall, Descartes solution to the problem that basic operation in the world requires faith is that he invokes he attempts to prove that our universe is run by a benevolent God (see, *Descartes' Philosophy*, Bernard Williams & Bryan Magee, 1987, [https://www.youtube.com/watch?v=ba9-sCXRC-E](https://www.youtube.com/watch?v=ba9-sCXRC-E)). With the benefit of hindsight in 2023, we do not believe that Descartes arguments for God are the most persuasive at this point in time. However, one has to appreciate the progress that Descartes concretely made on this issue. He proved that: - Descartes’ lasting contribution on faith and doubt 1. showed the limited nature of knowledge that cannot be doubted at all 2. showed that many beliefs that we do need for operation in the world require some kind of *****faith***** 3. realized that, one way or another, the belief that God is *****benevolent*****, or at least not *********malicious*********, is the, if not a, primary way by which we can hope to believe in the rationality of the universe ## Hume The next “great philosopher” to fit into this story is David Hume (1711—1776), one of the great contributors to our understanding of the philosophy of **********empiricism**********, i.e., the study of the “outside world”. In his famous **************************************Enquiry Concerning Human Understanding************************************** (1748), Hume stated a thesis that is so basic but also so counter-intuitive: no amount of experience with the past ****ever**** entitles us to predict the future with 100% confidence. In Hume’s words: > Even after the observation of the frequent or constant conjunction of objects, we have no reason to draw any inference concerning any object beyond those of which we have had experience. (Hume, *An Enquiry Concerning Human Understanding*, 1748) > This insight was formalized using mathematics and computer science as the “no free lunch” theorems of Wolpert and Macready, circa 1997. In other words, no amount of experience with the past entitles us to predict the future with 100% certainty. The laws of nature could completely change tomorrow or at any time. Thus, we need *****faith***** to believe that the laws of physics will continue operating. ## Probability Theory Similarly, though it is not clear to us which author would be associated with the idea, the field of ******************probability theory****************** recommends to practitioners never assign 0% probability to an event. There are several technical reasons for this that relate to one another. First of all, if an event had 0% probability according to a model, then observing that event in the data would result in infinite ***surprise***, and it is considered incoherent to have infinite surprise. Another related reason is that, in a language model, if some sub-event has 0% probability, then all derivations for that document might have 0% probability, and so there would be no way to choose between derivations, even though some are clearly better or worse when *low*, rather than ****zero****, probabilities are used for unseen events. Since no contingent event can ever have 0% probability, therefore no contingent event can ever have 100% probability—its complement would have 0%. Thus, we see that we can never assign an event 100% according to probability theory. Thus, the probability that the laws of physics will be the same tomorrow as they are today can never be 100%. Thus, we need *****faith*****. # The Necessary Role of Faith (Reprise) At the end of the day, we can never have 100% certainty in any theory. This includes the theory of God. But, it is a fundamental fact about all theories, and, indeed, even about the fact that the universe behaves according to a pattern, or that science is even applicable at all. In other words: it requires *****faith***** to believe that science will be applicable in the future. It requires *****faith***** in order to act in the world: faith that the universe will behave according to similar rules to what it has done in the past. There is, once again, ****************************no way to prove beyond doubt**************************** that the universe ****will**** continue to follow a pattern. Thus, we need ***faith***. This is what Descartes, Hume and Wolpert all told us years ago. Sam Harris seems not even to be aware of this issue.#34,973,830text# What is Philosophy? A Four-Question Formulation Greg Coppola September 27, 2023 # Introduction Traditionally, **Western Philosophy** is defined or introduced as an **enumerated list** of seemingly arbitrary subtopics. For example, *********Wikipedia********* defines philosophy as: > **Philosophy** (*love of wisdom* in [ancient Greek](https://en.wikipedia.org/wiki/Ancient_Greek)) is a systematic study of general and fundamental questions concerning topics like [existence](https://en.wikipedia.org/wiki/Existence), [reason](https://en.wikipedia.org/wiki/Reason), [knowledge](https://en.wikipedia.org/wiki/Knowledge), [value](https://en.wikipedia.org/wiki/Value_(ethics_and_social_sciences)), [mind](https://en.wikipedia.org/wiki/Mind), and [language](https://en.wikipedia.org/wiki/Language). (Wikipedia, [https://en.wikipedia.org/wiki/Philosophy](https://en.wikipedia.org/wiki/Philosophy)) > When viewed as this list of seemingly arbitrary topics it becomes difficult to answer the mystery of the popularity of philosophy: - **the mystery of the popularity of philosophy** - *popularity* - philosophy is evidently very popular with, and meaningful to, thinking people, and even the masses - *mystery* - why would this seemingly arbitrary, obscure list of academic topics be so relevant for so many people? We believe the answer can be found by distilling the entire enterprise of philosophy down to its essence as: to ******orient****** oneself in the ******world******. In particular, we view philosophy as the giving of answers to ****four**** simple questions. # The Subtopics of Philosophy Since philosophy is defined using an enumerated list of subtopics, everyone can give their own list. We highlight the following areas as among the most important in what has traditionally been called “philosophy”: - **metaphysics** what is the nature of reality? - **epistemology** how do we know what we know? - **ethics** - a study of moral principles - **aesthetics** - the study of the beautiful - **logic** - study of valid arguments - **political philosophy** - the study of political systems - **philosophy of mind** - the question of consciousness - **philosophy of language** - the study of how language - **philosophy of science** - the definition and justification of the scientific method # Enumerative Definition and its Drawbacks Variously called an **********************enumerative definition********************** or a ***************list definition***************, this conception of philosophy defines the field as an ********arbitrary-seeming list******** of topics. One is not clear on how we would add more subtopics, or when it would be possible to add. It would seem to be a matter of opinion which topics get added, and it would not be clear whose opinion that would be. For example, suppose we were to define a ********restaurant******** by a list of three examples: 1) the cheap fast-food restaurant *McDonald’s*, 2) the mid-market steak house *******The Keg*******, 3) the New York three-star Michelin restaurant *****Per Se*****. From this list of items, can we say whether a new addition to the list would definitely be a restaurant? Would a “food truck” be a restaurant? What about a home where dinner is served? What about a “bed and breakfast”? Fuzzy “I’ll know it when I see it”-style definitions create a lack of clarity. # **Four-Question Formulation** ## The Four Questions I believe it is easier to appreciate the central role that *philosophy* plays in the lives of so many people if we view philosophy as the project to answer the following *four questions*: 1. *what* is **happening**? - this is the focus of *******science******* 2. *how* do we **know**? - this is the focus of *************the philosophy of science************* 3. *what* do we **value**? - this is the focus of **********aesthetics********** - note - this is the ******only****** step where we access our “emotions” - the rest of the questions are purely logical, and free of emotion 4. *what* will we **do**? - this is the focus of ***************decision theory*************** and ******ethics****** When viewed as this list of topics, we see clearly that every person, whether a professor, a business person, or a blue collar worker, must answer these questions for themselves. Thus, philosophy is seen as the formalization into a science of something that common sense people do all the time: orient themselves in their environment. ## List of Subtopics Revisited Let’s review the classic subtopics of philosophy, to see which of the four questions each subfield is meant to address: - **metaphysics** what is the nature of reality? (what is happening?, how do we know?) - **epistemology** how do we know what we know? (how do we know?) - **ethics** - a study of moral principles (what do we value?) - **aesthetics** - the study of the beautiful (what do we value?) - **logic** - study of valid arguments (how do we know?) - **political philosophy** - the study of political systems (what should we do?) - **philosophy of mind** - the question of consciousness (what is happening?, how do we know?) - **philosophy of language** - the study of how language works (what is happening?, how do we know?) - **philosophy of science** - the definition and justification of the scientific method (how do we know?) # Explaining the **Importance of Philosophy** When philosophy is viewed not as a random list of topics, but as this set of four foundational questions, it is easy to see why every educated person prides themselves on knowing at least *some* philosophy, even if they aren’t the next Socrates. This is because the four questions allow us to orient ourselves in the world by determining: what is possible, what we value, and what to do.#34,727,367text# What is First Principles Thinking? Elon Musk and the Vienna Circle Greg Coppola September 27, 2023 # The Concept of First Principles Thinking The notion of *first principles thinking* has been popularized by Elon Musk (1971—present), the CEO of *X* and *****Tesla*****. About his own conception of the term, he said: > First principles is kind of a physics way of looking at the world. You boil things down to the fundamental truths and say, “What are we sure is true? Or, as sure as possible is true?” And then reason up from there. (Elon Musk, *Innomind*, [https://www.youtube.com/watch?v=NV3sBlRgzTI](https://www.youtube.com/watch?v=NV3sBlRgzTI)) > However, Elon is assuming that *******physics******* itself is a sufficiently low level of analysis to be considered “first principles”. We do not assume this. We want to explain the scientific method itself. This is a philosophy even more ****meta**** and so more close to “first” than even physics itself. We must go all the way to the most basic level: information theory. # Vienna Circle According to the Vienna Circle, *********************meaningful statements********************* were based on two subtypes: - **meaningful statements** for the **Vienna Circle** (circa 1920’s) - ********************scientific statements******************** - claims about *****************empirical reality***************** - subject to the *****************scientific method***************** - claims that are ***********falsifiable*********** - **********************************************mathematical statements********************************************** - claims about virtual objects in a virtual space - a claim might be either proven true, proven false, or not proven either way - the truth or falsity of this system The Vienna Circle sought to expunge from the University the influence of such allegedly “unproductive” philosophies as that of the Existentialists, like Martin Heidegger (1889—1976), who famously said ********************the nothing nothings********************. There was never any clear meaning to the statement “the nothing nothings”, and this was what the Vienna Circle were criticizing. The Vienna Circle believed that it was unproductive to spend time on questions which were not “meaningful”. # The Definition of First Principles For us, we propose the following definition of first principles: - **components of first principles thinking** - reasoning according to either - *mathematics* - things that are true by definition - *science* - correctly following the scientific method, to the best extent that it is formalized at the time - in 2023, we believe that the philosophy of science must incorporate the mathematics of information theory, especially ****************************minimum description length**************************** and *********************Kolmogorov complexity********************* In other words, we are proposing the following identity: - ***first principles thinking and Vienna Circle identity*** - the definition of first principles thinking is tied to the conclusions of the Vienna Circle - in other words - the ********meaningful statements********, according to the Vienna Circle, were math and science - the *************************first principles thinking************************* disciplines, according to us, are math and science # Implications In order to fit with this definition of “first principles thinking”, a theoretical statement must either be true: 1. because it mathematically follows 2. because it is a probabilistic statement about empirically observable data An example of something which does ***not*** follow from first principles thinking is the **********************************infallibility of the Catholic Pope**********************************. The traditional rationale for this would be something like: - **rationale for the infallibility of the Catholic Pope** - Jesus said to Peter, “on this rock I will build my church” (Matthew 16:18) - this statement is interpreted as meaning that Peter is infallible - the historical line of Popes form an “apostolic succession”, in which the infallibility is passed between the outgoing and incoming Pope This is not a first principles statement according to our definition because it neither 1) follows from definitions, nor 2) is an attempt to do empirical science.#34,726,613text# The Yom Kippur Lectures Debunking Three Fallacies of Richard Dawkins September 25, 2023 Greg Coppola # Introduction This is a collection of three essays written over the Yom Kippur fast of 2023. Each one addresses a different fallacy in thinking from Richard Dawkins’ book ****************The God Delusion****************. These three fallacies are, we believe, the strongest scientific (or pseudo-scientific) arguments to come from the New Atheists. Each lecture is written from the perspective that ******************information theory****************** is ****the**** way to interpret the philosophy of science in 2023. This is 1) given the incredible success of ************************artificial intelligence************************, which we view as a branch of information theory, but also 2) because of the elegance of fundamental ********accuracy******** of information theory. In particular, we refer to Kolmogorov’s principle of **minimum description length**, recently discussed at length by Ilya Sutskever in ********************************An Observation on Generalization******************************** (Simons Institute, 2023, YouTube). This specifies that *the optimal model for a data set* is the one which can represent that data using the smallest total disk space, through a combination of 1) compressing the data as small as possible, and 2) the model being as small as possible. # Dawkins’ Boeing 747 Argument In ****************The God Delusion****************, Richard Dawkins introduces the notion of an *****Ultimate Boeing 747*****, as a metaphor for God. The name is a reference to Fred Hoyle’s argument that the likelihood of Darwin’s theory of undirected evolution would be comparable to that of a “hurricane blowing through a junkyard and assembling a Boeing 747”. In other words, the argument is that undirected evolution is unlikely, because it intuitively *feels* as as unlikely as a plane assembled by chance, which does not happen. ## Dawkins’ ***************Ultimate Boeing 747*************** This is how Dawkins explained his view it at his 2018 ********ideacity******** speech: > The essence of the argument that I want to put is that God Himself would have to be *a forteriori* even more complicated and improbable an entity [than the human race] if he is to do what is expected of him, which is to create life or to create the universe, let alone listen to prayers and forgive sins. The essence of my argument is that God is the *ultimate* Boeing 747 and that any attempt to deploy the argument from improbability—which is actually the main argument that any theist uses in favor of theism—any attempt to deploy the argument from improbability backfires—shoots itself in the foot because it can be turned on itself, it can be turned on God. (”God and Science”, *ideacity*, YouTube, September 18, 2018) > We summarize Dawkins’ argument as follows: - ***Fallacious* A*ssumptions*** - **Modeling Fallacy** - God is more “complicated” than humans, and so our explanation of God must be even more *improbable* than an explanation of humans. - **Explanation Fallacy** - In order to invoke a theory of God, we must first *******explain******* God. - ***Fallacious Conclusions*** - **God Must Lose** - Any theory of God is “a forteriori even more complicated and improbable” than any theory of humans alone. We will examine Dawkins’ two fallacious assumptions in turn, to see why this argument falls apart. ## Modeling Fallacy in Dawkins’ Boeing Argument A preliminary problem is that Dawkins is confusing different notions (or *definitions*) of the word *probability*. This kind of sloppy use of language always leads to confusion. In human language, many times two different meanings are shared by the same word. The classic example is a ****bank****, which can refer variously to a “river bank”, or to a “financial institution”. Needless to say, an argument about properties of “river banks” does not apply to a discussion of “financial institutions”. The fact that these two objects are both called “banks” is only an accident of the ambiguity and sloppiness of natural human languages, and a careful thinker will separate the two. Similarly, when skeptics of undirected evolution criticize the apparently impossible ***********probability*********** of undirected evolution, we mean that we have the specific Darwin-style model of the process of “random” evolution, and we have modeled this in the apparently unique obvious way suggested by information theory: using the **********************random walk********************** model. And, *****************within that model*****************, we have computed the probability of evolution as *impossibly* ***low***. In the case of a Creator or God, ***********in contrast***********, we ******do not****** have either a specific proposal of what **********generated********** God. This is different than the Darwinians, because they crucially ***are*** presenting a theory of how humans came to be, and the Darwinian process has a clear information-theoretic modeling. Since people who doubt evolution are ***not*** proposing a mechanism that generated God (and, they don’t need to in order to *****doubt***** random evolution), it does not make sense to assess the “probability” of God, because we have no proposed mechanism. The difference can be summarized thus: - **Darwin’s specific proposal of ***undirected*** evolution** - undirected evolution’s proponents ***are*** proposing a specific model family that generated the human race: *undirected evolution* - there is an obvious modeling of this using ******************information theory****************** - the *probability* of Darwin’s proposal, according to this modeling, is **************impossibly low************** - **we *do not* have a proposed model for God** - skepticism of undirected evolution is simply the ********negation********, or ********doubting********, of a specific theory - doubting a theory (like random evolution) does ***not*** require us to provide a competing positive theory - given the lack of a theory of how God was generated, we cannot speak of the “probability” of God existing To summarize, is incorrect to compare “probabilities” for the two cases, because only in the case of Darwin’s theory of undirected evolution do we actually have a concrete proposal for which we can estimate a probability. No one has proposed, nor needs to propose, a theory of how God was generated, so we can’t discuss the “probability of God”. This is what we call the ********************************modeling fallacy******************************** in Dawkins’ argument. ## Explanation Fallacy in Dawkins’ Boeing Argument The second aspect of Dawkins’ confusion is when he assumes that we even need to ****explain**** an “intelligent designer” (or any entity) in order to appeal to one in our scientific theory. This is a major reasoning ****fail**** on Dawkins’ part. It seems Dawkins is ignorant of one of the most famous quips in all of science, when Isaac Newton wrote ********************hypotheses non fingo********************. Let’s consider this story and its implications right now. ### Hypotheses Non Fingo One of the most memorable moments in the history of science was circa 1713 CE, when all-time great physicist Isaac Newton responded to real and hypothetical critics of his then-new theory of gravitation with one of the most famous quotes in the history of science: “*hypotheses non fingo*”. Translated to English, Newton’s response means, “I feign no hypotheses”. That is, Newton’s theory of gravitation explained the motion of the planets and of falling bodies on earth as a result of a newly postulated force Newton called *********************universal gravitation*********************. However, skeptics of Newton’s proposed “action at a distance” wanted Newton to also go on to explain **************the force of Gravity itself**************. Newton’s classic response, *********************I feign no hypotheses*********************, expressed a deep scientific truth that Newton had evidently already grasped: giving an explanation at one level of analysis, does not require giving an explanation at ***all*** explanations of analysis. This is something we now see as true in light of the assumption that ****information theory**** contains the ultimate definition of science, even though Newton did not have access to information theory, which was created around 1948, led by Claude Shannon. This is a general truth about ***all*** theories, including the question of Intelligent Design of the human race. That is, if the concept of an “intelligence” in design better explains and predicts the data, then this is a better theory than its alternative, and we do not have to “explain” the intelligence to refer to it. ### Information Theory Perspective **Information Theory** We reiterate our fundamental thesis that ************information theory************ is the ultimate refinement and expression of science. In particular, we want to appeal to the principle of **************************minimum description length**************************. As recently alluded to by Ilya Sutskever in his talk ********************************An Observation on Generalization******************************** (2023, YouTube), in the context of **********prediction********** and ***********compression***********, the overall quality of a theory is a sum of the following: - the *overall disk size* of a theory according to **************************Minimum Description Length************************** 1. the length of the data, compressed using the theory 2. the length of the theory itself Between two theories, the **better one** is the one with the *****lower***** overall disk size. **Why Newton was Correct to say *Hypotheses Non Fingo*** Newton’s position, recall, was that he did not need to *****feign***** a hypothesis about what *********************caused gravity itself*********************. He intuitively understood—i.e., long before the advent of information theory—that his theory of gravitation had added value (tremendous value, in fact) as it was, because it allowed us to predict many new things that couldn’t be predicted before. Or, in other words, his theory allowed us to better “compress” the past, because now we can exploit the patterns to store it in “compressed” form. Thus, his theory added value, even though he could not *******explain******* gravitation. In fact, we still can’t full *******explain******* gravitation to our total satisfaction, but this does ***not*** mean we cannot benefit from or use the theory of gravitation, as far as it works. # Dawkins’ “God of the Gaps” Another common Dawkins fallacy is when he refers to the ***************God of the Gaps***************. His formulation in ****************The God Delusion**************** is typical: > Searching for particular examples of irreducible complexity is a fundamentally unscientific way to proceed: a special case of arguing from present ignorance. It appeals to the same faulty logic as 'the God of the Gaps' strategy condemned by the theologian Dietrich Bonhoeffer. Creationists eagerly seek a gap in present-day knowledge or understanding. If an apparent gap is found, it is *assumed* that God, by default, must fill it. What worries thoughtful theologians such as Bonhoeffer is that gaps shrink as science advances, and God is threatened with eventually having nothing to do and nowhere to hide. (Dawkins, The God Delusion) > Summarizing his argument: - ***Fallacious Assumptions*** - **Modeling Fallacy** - Dawkins incorrectly equates gaps in the empirical *fossil record* with gaps in our *knowledge* - **Explanation Fallacy** - Dawkins incorrectly assets that positing a “God of gaps” is a “fallacy”. - Actually, we should use ******************information theory****************** and **************minimum description length************** as our only yard sticks, not these antiquated, pre-information theoretic notions. - ***Fallacious Conclusion*** - it is wrong to discount evolution due to gaps in the fossil record ## Modeling Fallacy in Dawkins’ Gaps Argument Dawkins’ first error is that he thinks that he is equating **************gaps in the fossil record************** with *************gaps in our knowledge*************. The fact is that we ********************do have knowledge******************** of the fossil record and that fossil record ********does not show transitions******** (Darwin’s Doubt, Stephen Meyer, 2013) We **do** have the knowledge. And, the knowledge is that there *are* gaps in the fossil record. These gaps are mis-predictions for the theory of evolution. This is the difference between knowing that the answer is “no” to a yes-no question, versus just not knowing the answer. Dawkins is confusing two meanings of the word “gap”. A gap in knowledge is not the same as a gap in the fossil record. These are simply two ********************************************different concepts with different definition********************************************, that share an English word form. But individual word forms are understood to be *****ambiguous***** by careful linguists, and someone who understands language does not confuse different meanings for a word. ## Explanation Fallacy in Dawkins’ Gaps Argument In information theory, once again, the best theory is described by the principle of ***********************minimum description length***********************. That is, the best theory is the one which minimizes the total disk space required to compress a data set, including both compressed data and the program. The phrase “God of the gaps” does not appear anywhere in the field of information theory. There is no reason to call “God of the gaps” a “fallacy” in 2023. This is an out-dated notion, and we should only measure by modern, quantifiable, computable information-theoretic metrics. # Russell’s Tea Pot Argument A famous argument among modern atheists is an argument about the **********burden of proof********** when it comes to religion due to the philosopher and mathematician Bertrand Russell (1872—1970) : > Many orthodox people speak as though it were the business of skeptics to disprove received dogmas rather than of dogmatists to prove them. This is, of course, a mistake. If I were to suggest that between the Earth and Mars there is a china teapot revolving about the sun in an elliptical orbit, nobody would be able to disprove my assertion provided I were careful to add that the teapot is too small to be revealed even by our most powerful telescopes. But if I were to go on to say that, since my assertion cannot be disproved, it is intolerable presumption on the part of human reason to doubt it, I should rightly be thought to be talking nonsense. (Bertrand Russell, Is There a God?) > To summarize, Russell’s argument is: - ***Russell’s Tea Pot Argument*** - dogmatic believers have proposed to believe in a God - the burden of proof is on them to prove that there definitely **is** a God, rather than on the atheist to ********disprove******** God - this is analogous to the question of a tea pot for which there are no evidence ## Misleading Nature of Metaphor The first thing that we have to understand is that this metaphor is inappropriate because of string theory: - ***string theory proposes more dimensions*** - says that there are dimensions beyond our observable 3-dimensional world - we assume that we cannot observe these dimensions - therefore, we have ***************no expectations*************** about what is going on beyond the 3-dimensional world The central hypothesis of believers in mysticism is: - ***mystics say God lives in higher dimensions*** - God and other intelligent spirits are hypothesized to live in other dimensions, beyond the 3-dimensional Therefore, this is a misleading metaphor because: - ***misleading expectations in the tea pot metaphor*** - we **do** have expectations about how tea pots look, and what it would look like to find or not find a teapot in orbit on the sun - we ******do not****** have any expectations of what happens in other dimensions than our 3-dimensional observations - a situation where we have expectations is completely different than a situation where we do not have expectations - therefore, it is misleading to use a situation we ************can conceive************ and ************************do have expectations for************************ as a metaphor where we don’t ## Incorrectness of the Analogy Aside from the misleading nature of the analogy, it is also just flat out incorrect to use this analogy, at least in terms of the question of *intelligent design*. The notion of intelligent design, which is in other words just a rejection of Darwin’s undirected evolution, is the appeal to the idea of an *********************intelligent cause********************* to explain the universe and human life, because of a judgment that Darwin-style random evolution is not possible. In other words, we are reasoning from a phenomenon, to an explanation. We are not proposing an unobservable phenomenon. Since the structure of the scenario is not the same, the analogy completely breaks down: the entire point of an analogy that the two related scenarios have an *********analogous********* structure.#34,607,540textInscription #26768639#26,768,639webp**A Perfect Empirical Truth Predicate Requires Omniscience** Gregory Francis Coppola Apocalypse Book Club August 22, 2023 # Introduction This paper casts the informal, and colloquially used concept of “**truth**” in a formal way, by introducing the concept of an **empirical truth predicate**. We show that the existence a **perfect empirical** **truth** **predicate** requires the existence of an **omniscient entity**. In colloquial language, an *****************omniscient entity***************** would correspond to what people think of as “******God******”. We thus understand why it is that, **post-modernists**, who reject God, also reject truth. # Empirical Truth Predicate ## Truth Predicate A **truth predicate** is a **boolean function** over **sentences** ******************************in ********************************logical language**. A **language** is a possibly infinite **set** of **sentences**, with each *sentence* either clearly **in** or **out** of the *language*. A **logical language** is one in which **implication** (i.e., logical inference) is defined, so that from a set of ****premises****, we can clearly define the ************consequences************ of these premises. ## Universe State An **************************************************empirical truth predicate************************************************** defines defines a truth predicate over what “**actually exists**”. This presumes that there is something called “reality”, which is isomorphic to a state machine. That is, at any given time index, there is some ************************universe state************************, which can be described in some language. Presumably the various states are related to each other using a set of *****************state transitions*****************, but we don’t actually need to even take a position on this. ## Indexed State We assume that the entire history of the universe can be recorded as a set of indexed states, which are pairs of 1) a **time index**, and 2) a **Turing-representable state**. Presumably this set would be finite, but we don’t use the assumption that it is finite, so the same discussion would work if somehow there were an infinite number of states. # Aspects of an Empirical Truth Predicate ## Finite Beings We contrast two **archetypes** of beings, which recur over different **************aspects************** of beings. **Finite beings** are **************limited************** in various aspects. This is compared with the archetype of an ************omniscient being************, which is unlimited with respect to the given characteristics. ## Perception All **finite beings** have ************************************limited perception************************************. That is, even assuming that there is a describable **universe state** underlying any time index, limited perception beings do not have the ability to perceive this universe state directly. Instead, the perception of *************finite beings************* is mediated by **************************finite senses**************************. It is known, for example, that there are more electromagnetic waves, for example X-rays or gamma rays, that humans cannot detect at all. There is no bound on the number of aspects of universal state that limited humans cannot detect. ## Interpretation Every **individual** has their own **individual language**. The collection of *individual languages* in a **community** make up a **community language**, which would be analogous to English or Spanish, etc. However, only the *individual languages* are actually instantiated in the physical universe, and the concept of a ******************community language****************** does not actually exist in the world. In order to define truth over an individual sentence, we need to actually “**mind read**” the person or community who said it. This means that the interpreter must be aligned with the what the interpreter “would have intended”. This would imply, for one thing, that, no matter how much new information the individual receives, they will never learn anything which contradicts this answer. ## Memory In order to be able to define a complete truth predicate, the interpreter must have an accurate recording of the ******unmediated universe state****** at any given time index. Humans have fallible **memories**. That is, even among the fallible sense perceptions of humans, they still forget most of what they once perceived. # Perfect Truth and Fallible Entities A **perfect empirical truth predicate** is an *empirical truth predicate* that has: - unmediated perception - perfect interpretation - perfect memory This is analogous to what we colloquially call the “*capital T Truth*”. ## Fallible Entities A ******************************fallible entity****************************** has one or more of the following: - fallible interpretation - fallible perception - fallible memory The fallibility of ***any*** of these three properties would make a being unable to provide a perfect truth predicate. # Results ## Combination of Fallible Entities Proposition: - A collection of fallible entities is still fallible. There is no guarantee that two fallible quantities, when added, will add up to the infinite quantity that would be needed to be infallible, whether for **************interpretation**************, ******perception******, or ******memory******. ## Truth Requires Omniscience Proposition: - A perfect empirical truth predicate requires there to be an interpretation entity with properties: - infallible interpretation - infallible perception - infallible memory If this were not so, and one of these properties were limited, then the truth predicate would not be perfect. A being which is infallible on all three dimensions we will call **********************omniscient**********************. # Discussion ## Omniscience is God In colloquial usage, the only being that would be “omniscient” would be God. Thus, we can say that, “for truth to exist, we have to have a God.” ## Why Post-Modernists Reject Truth Post-modernists *often* reject: - that there exists an “actual” state of the universe Post-modernists, when they are ****************atheists****************, which is *******usually*******, reject: - that there exists an omniscient being As we saw, “capital T Truth” requires both of these things.#25,991,632textBest Refutation of Sam Harris' "Moral Landscape" Greg Coppola Apocalypse Book Club Aug 14, 2023 https://www.youtube.com/watch?v=-bBKrE_oanY Automatic Transcript 0:00 this is apocalypse book club I want to 0:02 talk about Sam Harris today you've got 0:04 people who are passionate and dedicated 0:06 to what they believe and then you get 0:07 Sam Harris basically okay Tim Harris is 0:10 someone that I think has set philosophy 0:12 back for all of his fans and even for 0:14 everybody that's engaging with him for a 0:16 variety of reasons I think the first 0:18 reason is that he's trying to invent his 0:20 own style which I would call like the 0:23 imaginary scenario style of philosophy 0:25 dial up the deadliness of the pathogen 0:28 bodies of kids are being stacked up in 0:30 parks and we have a vaccine that 0:32 actually works and then we've got RFK Jr 0:34 saying maybe you don't want to get the 0:36 jab he says imagine a scenario and 0:38 children and children are dying being 0:40 part of the street now RFK is saying 0:42 don't do it no no no RFK was talking 0:45 about now you made up a fake scenario in 0:48 your own mind and then criticized RFK uh 0:51 I mean I don't know if he's the first 0:52 one to ever use it but like he seems to 0:54 think that this is like what 0:55 philosophers do or he thinks this is an 0:57 advance based on like what we previously 0:58 had but it's neither there you know you 1:00 don't just create these imaginary 1:02 scenarios and I want to explain what I 1:03 mean and everybody has been like kind of 1:05 ragging on Sam Harris lately Dave Smith 1:07 he said imagine if everything was 1:08 completely different then things would 1:10 be different well I think what's strange 1:12 is the mental gymnastics he has to go 1:14 through to create a scenario where the 1:16 the world that he wants is correct which 1:18 is good because I think if you think 1:19 about it it's like how do we actually do 1:21 peer review one way is that there's like 1:23 three peer reviewers these magical peer 1:25 reviewers at journals that now even 1:27 people like Richard Dawkins and 1:29 peterborghin are like saying you know 1:31 there actually are problems with peer 1:33 review or they'll if they're 1:34 knowledgeable and they're an academic 1:36 they'll talk about literature the 1:39 peer-reviewed literature 1:41 and that peer-reviewed literature has 1:43 been what I've written about before it's 1:45 been ideal laundered so you get a bunch 1:47 of people who are ideologues together uh 1:49 who already they start with their 1:51 conclusion first and they work backward 1:52 from the conclusion the exact opposite 1:55 of science so that Weinstein says 1:56 there's problems with peer review 1:57 everybody uh John lacun said like the AI 2:00 guy he said who invented like deep 2:02 learning partly he was always critical 2:04 of the peer review process because the 2:06 peer review process turned down deep 2:07 learning in the early days and so it's 2:10 basically like if we're not going to do 2:12 peer review the traditional way which is 2:14 kind of good because I think it's just 2:15 like very slow then the question is how 2:17 do we do peer review as an internet 2:19 community and I think ideally when 2:22 someone was proved wrong they would 2:23 admit the point but Sam Harris won't do 2:25 that if he's not going to hit the point 2:26 he could at least go and do interviews 2:28 or podcasts with people where they 2:30 actually discuss whether or not it's 2:32 true what he says for example the one 2:35 thing that I'm always on Sam Harris 2:36 about is that he said that we can get 2:37 values from science we can get values 2:39 from Facts yeah well I should preface 2:41 this by saying this is a very 2:43 controversial opinion uh I think no 2:46 that's why I bring it up 2:49 I don't think it should be but it is and 2:51 it's been part of modern philosophy from 2:53 like since Hume which was like 300 years 2:55 ago that you can't get an ought from an 2:57 is which is like a linguistic argument 2:59 and it's a good argument and I would 3:01 update it using modern logic that's 3:03 going to be the contribution that I want 3:04 to make in my book AI philosophy so we 3:07 want to update hume's argument there's 3:08 more things we could say we could bring 3:09 in decision Theory we could bring in 3:11 artificial intelligence perspectives and 3:12 all of these would contradict what Sam 3:14 Harris is saying and so there's ways we 3:16 can modernize hume's argument that you 3:18 can't get an ought from an is but that's 3:19 just like deepening the Insight that he 3:21 already had so Hume is one of the major 3:23 philosophers of the English language and 3:25 yeah it's like definitely one of the 3:26 greats there's not that many people 3:27 really that have been great philosophers 3:29 in the last 2 500 years I actually think 3:31 it's less than 25 I would say there's 3:32 probably been less than one great 3:34 philosopher every Century so in some 3:36 sense we don't like it doesn't make 3:37 sense for everyone to try to be a 3:38 philosopher at once like not at least 3:40 not a great philosopher everyone needs 3:41 some philosophy and then like there's 3:43 like ways to apply things in different 3:44 scenarios and ways to mix things like 3:46 Tai Lopez is kind of a great philosopher 3:48 because he's like philosopher of life 3:50 and he's mixing all these different 3:51 perspectives and he does have new stuff 3:53 controversial take you should try to 3:55 make your body 3:56 look specifically whatever is most 3:59 attractive to the opposite sex but these 4:01 kind of fundamental insights that come 4:03 from people like Hume they don't come 4:04 very often and so for Sam Harris who 4:08 disagree with Hume you have to like it's 4:10 he's one of you know he's one of the 4:12 standards I think I think courtesy of 4:14 Hume 4:15 we have drawn the lesson in wet in in 4:18 Western philosophy and science that 4:21 facts and values are two different 4:24 classes of thing he's one of the people 4:27 in the Canon you can't just like dismiss 4:29 him and say well humans said this but I 4:30 don't think so you have to give a reason 4:31 this is like 101 okay so if you think 4:33 Hume was wrong that's what he said the 4:34 other day with Bernard Coast he was 4:36 invited for some kind of after dinner 4:38 event with Sam Harris where I guess like 4:40 the idea is like you know talk about 4:41 some ideas which is in theory a fun 4:43 thing but the problem is Sam Harris 4:45 doesn't know what he's talking about and 4:46 vinod khosla who's a VC he's an engineer 4:48 an entrepreneur and stuff even he knows 4:50 the issue better than Sam Harris and you 4:52 know Sam Harris just like Rambles on 4:54 every time of a note asks a question and 4:56 then he doesn't answer the question of 4:57 no coastla the entrepreneur VC not 5:00 full-time philosopher but no Costa who 5:02 understands the issue better keeps 5:03 trying to like prompt Sam Harris and 5:05 prompt Sam Harris and he keeps missing 5:06 the point that's crazy but one thing 5:07 that Sam Harris said in the middle of 5:09 this um interview was like he says okay 5:12 Hume said you can't get an odd from an 5:14 is which is true he's like but I don't 5:15 think we need to put play these language 5:16 games I I think I think that the concept 5:19 of should the linking of morality and 5:21 questions of right and wrong and good 5:23 and evil to questions about should and 5:25 ought is a 5:27 a language game I don't think we have to 5:29 play so basically that's how he argues 5:31 away what humans saying is if we don't 5:32 need to play these language games and 5:34 then that's it and then he just starts 5:35 talking about like random stuff and it's 5:37 like I think Sam Harris fans 5:38 unfortunately like that they don't 5:40 realize that Sam Harris is just rambling 5:42 on like they think that because Sam 5:43 Harris talks for 20 minutes that he's 5:44 answering the question and this is bad 5:46 same here shouldn't be teaching people 5:47 this it's like philosophers reach 5:49 conclusions about important issues you 5:51 know what I mean and Sam Harris it's 5:53 like he just talks and talks and he 5:54 likes to have these conversations that 5:56 don't go anywhere you know he just likes 5:57 to see his like face on people's 5:59 podcasts and he likes to be invited to 6:01 dinner parties and stuff like that but 6:02 it's not really what philosophy is about 6:03 philosophy is about answering questions 6:05 like in a concrete way like Hume did 6:07 when he said there's no op from an is 6:09 and um yeah I mean like I think you know 6:11 you could add a bunch of math to this 6:13 argument but the intuition of what Hume 6:15 is saying is like basically there's like 6:16 what is and what we should do about it 6:18 and so if you have a like science which 6:21 is a database of like facts that are 6:23 have the verb is or does like for 6:25 example the sun does rise in East or you 6:28 could say the sun is something that 6:29 rises in the East so is and does I would 6:32 say those are like the main verbs or 6:33 probably does right like um it probably 6:35 does rain during the rainy season it 6:38 probably doesn't rain not in the rain 6:39 right is and does are the main verbs of 6:42 science now when you want to talk about 6:43 action or making a decision now all of a 6:45 sudden you're talking about what should 6:47 we do in such and such a situation or a 6:49 context given a situation or a context 6:51 what should we do so basically if you 6:53 have a database of sentences where the 6:55 verb is is or does I think that's why 6:57 Hume just says is but like is and does 6:59 are similar is or does that's like 7:00 science is like a database of facts or 7:02 potential facts or theories all of which 7:04 have the verb is or does so now when you 7:06 want to make a statement of like we 7:07 should do x y z you can't you can't have 7:10 a conclusion that says we should do 7:12 something if all of the sentences in 7:15 your database are using the verbs is or 7:17 does and so Sam Harris wants to just 7:19 brush aside Hume who first of all is 7:21 Right second of all is one of the great 7:23 philosophers of all time and third of 7:25 all has always been recognized as being 7:27 right under issue and so if you were 7:29 going to deconstruct Hume you'd have to 7:31 do it in a very serious way you don't 7:32 just say you don't just brush it off and 7:33 that's basically what Sam Harris does do 7:35 is he says well Hume said you can't get 7:37 an OP from an is but I don't think we 7:38 need to play this language game I I 7:40 think I think that the concept of should 7:42 the linking of morality and questions of 7:44 right and wrong and good and evil to 7:46 questions about should and ought is a 7:50 a language game I don't think we have to 7:51 play right so I think if we just leave 7:54 that aside that's what he says we don't 7:55 have to play this language game now the 7:57 word language game is from wickenstein 7:59 and it refers to the idea that when we 8:02 are using language in our everyday lives 8:05 it's kind of like we're playing a game 8:06 and so then language is kind of like a 8:09 tool and this leads to a whole 8:10 interpretation of language and the word 8:12 language game the way the wittenstein 8:13 said that's fine there's nothing wrong 8:14 with the word language game the problem 8:15 is same here is just here's the word 8:17 language game he uses it in a way that 8:19 it's not supposed to be used like he's 8:20 not referring to a language game the way 8:23 that it is of like what are we using 8:25 language as a tool to do so according to 8:27 wittenstein who introduced the word 8:29 language game language game refers to a 8:32 form of human activity in which language 8:33 is used each language game involves a 8:35 set of rules and conventions that govern 8:37 how language is employed within a 8:39 particular context so the word language 8:40 game does make sense but not the way Sam 8:42 Harris is using it he's using it as 8:44 though like basically just in a very 8:45 sophistry kind of a way like everything 8:47 that he does like I don't even know 8:48 exactly what he means but is a language 8:50 game we don't have to play I I think I 8:53 think that the concept of should the 8:55 linking of morality and questions of 8:56 right and wrong and good and evil to 8:58 questions about should and ought is a 9:02 a language game I don't think we have to 9:03 play I mean I you know that's what he 9:05 always does is he says these things that 9:06 sound smart but they don't really like 9:08 mean a precise thing so if by The 9:10 Language game you mean like the fact 9:12 that we're doing science like in 9:13 philosophy like philosophy and science 9:15 are a language game is that the game we 9:17 don't need to play like philosophy and 9:19 science like isn't that your whole game 9:20 like it's like I don't know like what I 9:22 really say about this it's a crazy 9:24 sentence it's a language game we don't 9:26 need to play like it's a language game 9:27 that we do need to play because we want 9:30 to form sentences that have the word 9:32 should in them so how do we not play 9:34 this language game Sam Harris but even 9:36 so I just I think the main thing to say 9:37 about this sentence is that it's just 9:39 it's self history because it doesn't 9:40 really mean anything and he's just 9:41 taking a word that sounds cool and 9:43 putting in a sentence and this is why I 9:45 think Michael schellenberger I'm with 9:46 Brett Weinstein the other day called Sam 9:48 Harris childlike you know Sam has a very 9:51 wrong core world view 9:55 um that is sort of astonishing in its 9:58 Simplicity and childlike nature because 10:00 it's like he just takes the word 10:01 language game that he heard somewhere 10:03 and it sounded cool and then he says it 10:05 on stage and he's like I'll say the word 10:06 language game but language game means 10:08 something is a 10:10 a language game I don't think we have to 10:12 play it means a precise thing according 10:13 to Wittgenstein or whoever you mean but 10:15 the point is first of all he doesn't 10:16 define Language game he just says like 10:19 language game you know and I don't even 10:21 know if he's trying to be sophist 10:22 sophistry or like to use sofa Street or 10:25 if he just doesn't even know what he's 10:26 doing or both I think it's probably both 10:27 but I I think the first thing I really 10:30 want to say is like you cannot dismiss 10:32 Hume with this nonsensical argument Hume 10:34 is right he's a landmark in the field 10:36 and this argument doesn't make any sense 10:37 okay so then Sam Harris like blabbers on 10:40 and on and he probably talks for like 20 10:42 minutes in this interview and I saw him 10:44 on with Lex Friedman complaining that 10:45 people take him out of contact Sam 10:47 Harris is concerned that people take him 10:48 out of context but it's like it's 10:50 because you talk for 20 minutes and you 10:51 don't answer the question and like it's 10:53 well known even in pop culture it's 10:55 almost like funny how many people are 10:56 remarking on this now but Sam Harris 10:58 talks in these like word salad style of 11:00 arguments where he just talks for like 11:01 20 minutes Mark Andreessen he called it 11:03 word salad he said that Sam Harris 11:05 creates a chain of hypotheticals with no 11:08 data attached to it he just like it's 11:10 like imagine this imagine it's like the 11:12 John Lennon song Imagine like that's 11:13 what Sam Harris does he's just like well 11:15 imagine this imagine that and the 11:18 problem with this imaginary scenario 11:19 style is that like yes philosophers do 11:22 imagine assumptions and then examine the 11:24 conclusions but the assumptions have to 11:26 be relevant and so that's what's in pool 11:28 and Dave Smith and even Tristan Tate the 11:32 philosopher Tristan Tate they were all 11:33 getting on Sam Harris for this recently 11:35 because he has this imaginary scenario 11:37 about vaccines where like the death rate 11:39 is higher than it is and even I think 11:41 the effective rate is higher than it is 11:43 and if both there was way more death and 11:46 children were dying and the vaccine was 11:48 perfectly effective then it would make 11:49 sense for everyone to take dial up the 11:51 deadliness of the pathogen bodies of 11:53 kids are being stacked up in parks and 11:55 we have a vaccine that actually works 11:57 and then we've got RFK Jr saying maybe 11:59 you don't want to get the jab he says 12:00 imagine a scenario and children and 12:03 children are dying pot on the street now 12:05 RFK is saying don't do it no no no RFK 12:08 was talking about now you may up a fake 12:11 scenario in your own mind and then 12:13 criticize the RFK well that's not the 12:15 scenario that we're in so you know like 12:17 the the point of philosophy is to find 12:19 assumptions that are relevant and show 12:21 that those assumptions lead to 12:22 conclusions that are relevant and that's 12:24 not what Sam Harris is doing but 12:25 basically then okay so he so he's 12:27 dismissed Hume with a flimsy argument 12:29 that's like major error that would be a 12:31 fail for a paper right there this was 12:33 peer review but let's keep going with 12:34 the interview and then Sam Harris trying 12:37 to refute Hume this 300 year old result 12:39 says yes we can derive morality or 12:42 values from science because all we need 12:44 to do is imagine the word there's that 12:47 word again imagine imagine the worst 12:49 possible outcome for everyone we want to 12:52 avoid that so this is the Crux of his 12:53 argument after having not really 12:55 dismissed Hume this is basically what 12:56 he's come up with so this is wrong so 12:59 many ways first of all I think it's 13:01 wrong because he said he was going to 13:03 derive values from Facts from science 13:05 but he hasn't derived he hasn't used 13:07 science in this this is a purely 13:08 philosophical argument and the whole 13:09 point of the moral landscape was how 13:12 science can determine human value so 13:13 that's a big problem he's kind of like 13:16 changed his argument because it was 13:17 supposed to be science that we're going 13:18 to use science to determine values now 13:20 I'm not saying that we don't use science 13:21 to determine what we do but the thing is 13:23 we use science and our values to 13:25 determine what we do we don't get the 13:26 values from science so the first problem 13:28 is he didn't use science that was his 13:29 big thing is that science was going to 13:30 determine the values we were going to 13:31 get the values from the world but he's 13:33 not getting it from the world he's just 13:34 using a philosophical argument that 13:35 would be true in any world it doesn't 13:37 depend on our actual world so he's not 13:38 using science to different values when 13:39 he says imagine a world or just imagine 13:41 the worst outcome for everyone the 13:43 second is this is not complete you know 13:45 like what do we do even if you were to 13:47 Grant the rest of this makes sense which 13:48 it doesn't we already know that we don't 13:50 want the best the worst possible world 13:51 for everyone but within that space of 13:53 things that are not the worst possible 13:54 world for everyone there's still an 13:55 infinite number of things we need to 13:57 make decisions about and this Theory 13:58 doesn't cover that finally I guess this 14:00 is related to the first point that I 14:01 made he is introducing values he's 14:03 saying we don't want the worst possible 14:04 outcome for urban but why not did you 14:06 get that from science no he's not 14:07 appealing to science he's appealing to 14:08 his own philosophy so he's dispersing 14:10 himself in a third way related to the 14:12 first because he actually just is 14:13 introducing values which is we don't 14:15 want the worst possible outcome for 14:16 everybody well that's a value and that 14:18 didn't come from science so you 14:19 disproved yourself again so basically in 14:21 summary Sam Harris proposed to change a 14:24 result by Hume without explaining why he 14:25 was wrong Hume is right and then his own 14:28 explanation of what he's trying to do 14:29 doesn't make any sense because he said 14:31 we want to avoid the worst possible 14:32 outcome for everyone well first of all 14:33 we already know that we don't need like 14:35 a big brain philosopher to tell us we 14:37 don't want the worst possible world for 14:38 everybody nobody even talked about 14:39 creating the worst possible world for 14:40 everybody obviously somebody's trying to 14:42 everybody the default is that everyone's 14:44 trying to maximize their own world and 14:45 then the Christian assumption or the 14:47 sort of higher order love based 14:49 assumption is that we want to do the 14:50 best for everybody so the sort of 14:52 choices are between doing the best for 14:53 yourself and doing the best for your 14:55 community doing the best for the entire 14:57 universe or just doing the best for 14:58 yourself nobody it's not even like a 15:00 relevant question whether we would want 15:02 to do the worst possible thing for 15:03 everybody but even if we did that's 15:05 still a value so I think the main thing 15:07 that Sam Harris has contributed is that 15:09 he's draw drawn attention to this 15:11 question so we can consider it a new can 15:13 you get values from science and since 15:14 the answer is no then the question is 15:16 where do the values come from and the 15:17 two answers are there either subjective 15:19 and we make up our own values or they're 15:21 objective and we have to seek the values 15:22 outside of ourselves but they can't come 15:24 from science Sam Harris didn't explain 15:26 what the pro what the problem with Hume 15:27 was he didn't offer his own concurrent 15:29 Theory a theory that he offered is 15:31 actually incoherent so he's wrong about 15:33 this but more generally I think the 15:34 important point is how does the internet 15:36 bring someone to Justice when they just 15:38 want to admit that they're wrong like a 15:39 gentleman or a gentlewoman or a gentle 15:41 person Sam Harris won't be a gentleman 15:42 and just say you know I put out this 15:44 Theory it's called the moral landscape 15:45 and it was actually completely off sorry 15:47 about that my mistake let's move on but 15:49 he won't do that and until he does that 15:51 I think you know it is appropriate to 15:53 mock him because it's like Sam Harris 15:54 fans who are a little on the slower side 15:56 they're just watching his videos and 15:57 being like oh wow this is what 15:59 philosophy is but this is not what 16:00 philosophy is so I think it's important 16:02 that some way we use certain language to 16:04 communicate the Sam Harris that you know 16:06 he's wrong and he should admit that he's 16:07 wrong and until he admits that he's 16:09 wrong I agree that the internet should 16:10 be putting pressure on Sam Harris and um 16:13 yeah until he till he admits his mistake#24,053,996text\pdfoutput=1 \documentclass[11pt]{article} \usepackage{amsmath} \usepackage{EMNLP2023} \usepackage{times} \usepackage{latexsym} \usepackage[T1]{fontenc} \usepackage[utf8]{inputenc} \usepackage{microtype} \usepackage{inconsolata} \title{Symbolic Logic is Also Needed} \author{Greg Coppola \\ {\em Apocalypse Like Right Now} \\ \texttt{[email protected]}} \begin{document} \maketitle \begin{abstract} ChatGPT has shown that, surprisingly, a GPT language model can learn from unsupervised text something approximating a logical model of accurate world knowledge. But, ChatGPT has the well-known ability to ``hallucinate'', and give facts or cite sources that do not exist (i.e., were not in the training data). Also, ChatGPT has the limitation that it is not able to reason in a principled way, empirically making reasoning mistakes, and also not being theoretically connected to a complete or consistent proof system. We propose to address both of these using a unified solution paradigm we call {\em SymbolicGPT}, characterized by: 1) the incorporation of {\em discrete logical forms} into a generative pre-trained transformer model, 2) the postulation as latent variables of syntactic parses and logical forms, and 3) a probabilistic model that can assign probabilities to discrete symbolic logical statements, based on a knowledge base of probabilistic discrete logical statements. While the model proposed is only theoretical, in that we do not have results from running the model, this paper does contribute several examples of empirical data of ChatGPT's performance in important relevant cases. \end{abstract} \section{Introduction} \citep{vaswani:attention-is-all-you-need:2017}, which introduced the popular {\em transformer} architecture, used the famously memified title {\em Attention Is All You Need}. The literal meaning of this intuitively humorous phrase is that, given attention, one can dispense with ``complex recurrent or convolutional neural networks.'' However, implicit in the idea that this title is a joke, is that, at some level of analysis, eventually attention will {\em not} be all that is needed. We analyze that the problems with ChatGPT as it currently stands are: 1) it hallucinates, and 2) it cannot maintain a logically coherent worldview. We propose that what is also needed at this time is a probabilistic knowledge base expressed in {\em symbolic logic}. \section{Analysis of ChatGPT} \subsection{The Ability to Build a World Model and Converse} The popularity of the ChatGPT application is due to the fact that ChatGPT allows better access to information than any other alternative, implying both a useful world model, and a useful interface. It has been suggested that ChatGPT is passing the ``Turing Test'' \citep{bengio:eye-on-ai:2023}. \subsection{The Problem of Hallucination} The problem of {\em hallucination} is a central one for ChatGPT. \subsubsection{An Experiment Showing Hallucination} Here is example from the preparation of this paper. Consider that in response to the question {\em what are some citations for a logical probabilistic database?}, ChatGPT responded with a citation that seems to not exist: \begin{quote} Kersting, K., \& De Raedt, L. (2007). Probabilistic logic programming. In Proceedings of the 20th International Joint Conference on Artificial Intelligence (IJCAI), 1623-1630. \end{quote} It seems according to {\em Google Search} and {\em Google Scholar} that this article does not exist, but there is a similar existing source: \begin{quote} {\em Probabilistic inductive logic programming}, Luc De Raedt, Kristian Kersting, 2008, Probabilistic inductive logic programming: theory and applications, 1-27, Springer Berlin Heidelberg. \end{quote} In other words, the title is {\em Probabilistic inductive logic programming}, not {\em Probabilistic logic programming}, and apparently the journal, publisher, pages and year are wrong. This is a typical example of hallucination, where ChatGPT returns a result similar to something that exists, but not exactly or actually something that exists. The fact that ChatGPT's results cannot be relied on is a major barrier to being used more widely. \subsubsection{Sutskever's Take on Hallucination} In {\em Forbes}, OpenAI Chief Scientist Ilya Sutskever recently gave his thoughts on the hallucination issue: \begin{quote} Now let's talk about the limitations. {\em It is indeed the case that these neural networks have a tendency to hallucinate}. That's because a language model is great for learning about the world, but it is a little bit less great for producing good outputs. And there are various technical reasons for that. There are technical reasons why a language model is much better at learning about the world, learning incredible representations of ideas, of concepts, of people, of processes that exist, but its outputs aren't quite as good as one would hope, or rather as good as they could be. \citep{sutskever:forbes:2023}. \end{quote} So, clearly OpenAI views hallucination as a primary limitation of ChatGPT. Sutskever theorizes that the ChatGPT model {\em is} good at {\em learning} from the world, but that the problem is with the {\em outputs}, and goes on to mention that {\em reinforcement learning} can be used to better control the output. \subsubsection{An Experiment on ChatGPT's Conception of Truth} We believe that the transformer model does not have a clear distinction between true and false. In support of this hypothesis, consider ChatGPT's response in the following: \begin{quote} {\bf paper author}: how does a gpt model represent a difference between "true" and "false"? {\bf ChatGPT}: {\em A GPT model does not inherently understand the concepts of "true" and "false" in the same way humans do}. Instead, it learns statistical patterns from large amounts of text data to generate outputs that are likely to be coherent and grammatically correct given the input prompt. \end{quote} \subsubsection{Aristotle's Law of the Excluded Middle} Aristotle's law of the excluded middle states that: \begin{itemize} \item Aristotle's Law of Excluded Middle: For any proposition P, either P or not-P is true. \item Probabilistic Version: For any proposition $P$, $p(P) + p(\neg P) = 1$. \end{itemize} There is no clear analog to this in the transformer, either in practice or in theory, and even ChatGPT does believe there is an analog. The generative GPT model has a fixed number of continuous parameters. This intuitively suggests something that needs fixing. \subsubsection{Discrete Line Between Known and Unknown} In 2002, Donald Rumsfeld famously popularized a distinction between {\em known unknowns}, and {\em unknown unknowns}. That is, there are three kinds of states for a database: \begin{itemize} \item {\em known knowns}: both the question and the answer are known \item {\em known unknowns}: the question is known but the answer is unknown \item {\em unknown unknowns}: a relevant factor exists, but neither was it known nor was even the question to ask known \end{itemize} We observe that ChatGPT does not make this distinction. It seems only to have parameters that help it predict the next character, without clearly demarcating these distinctions between known and unknown. \subsection{The Problem of a Logically Consistent Worldview} Evidently, a the key-value mechanism of attention is able to encode some kind of functional relationships, that, when expanded large enough, can learn a representation of the ``world" that is good enough to store information, and make it available again for consumers on a wide array of tasks. ChatGPT can answer logical questions, and shows in much behavior signs of logical reasoning. \subsubsection{Hinton Stresses the Problem of a Logically Consistent Worldview} Recently, in an appearance on {\em CBS Mornings}, Geoff Hinton discussed the problem of creating a {\em logically consistent worldview}: \begin{quote} We're at a transition point now where ChatGPT is this kind of idiot savant, and it also doesn't really understand about truth. It's being trained on lots of inconsistent data it's trying to predict what someone will say next on the web. And people have different opinions and it [ChatGPT] has to have a kind of blend of all these opinions so that it can model what anybody might say. It's very different from a person who tries to have a consistent world view. Particularly, if you want to act in the world, it's good to have a consistent world view. And, I think what's good one thing that's going to happen is we're going to move towards systems that can understand different world views and can understand that okay if you have this world view then this is the answer and if you have this other world view then that's the answer. \citep{hinton:cbs-mornings:2023} \end{quote} \subsubsection{Aspects of a Logically Consistent Worldview} The ability to hold a ``logically consistent'' worldview includes the following competencies: \begin{itemize} \item the ability to expand a set of premises $P$ to some $P' \supset P$ using a consistent set of {\em inference rules} \item the ability to detect a contradiction in a set $P$ of sentences \item the ability to partition a contradictory set of sentences $P$ into maximally consistent sets of sentences, $P_1, ..., P_N$ \item the ability to compare two theories $P$ and $Q$, e.g., and determine which is a better fit for the data \item the ability to explain a piece of data based on multiple consistent worldviews, i.e., to compute the probability of an observation under multiple different world views \end{itemize} \subsubsection{An Experiment in Reasoning with ChatGPT} There is no formal guarantee that the generative word model's explanatory parameters amount to logical inference. Indeed, consider the following interaction we recorded: \begin{quote} \item {\bf paper author}: if Alice is shorter than Bob and Bob is shorter than Charlie, then who is the tallest and why? \item {\bf answer}: Based on the information given, it is not possible to determine who is the tallest among Alice, Bob, and Charlie. All we know is that Alice is shorter than Bob, and Bob is shorter than Charlie. However, we do not have any information about the actual heights of any of the three individuals, so we cannot make any conclusions about who is the tallest. For example, it is possible that Alice is 5 feet tall, Bob is 5'5" tall, and Charlie is 6 feet tall. In this case, Charlie would be the tallest. Alternatively, it is possible that Alice is 6 feet tall, Bob is 5'10" tall, and Charlie is 5'5" tall. In this case, Alice would be the tallest. Therefore, without additional information about their heights, we cannot determine who is the tallest among Alice, Bob, and Charlie. \end{quote} To analyze this mistake, we can first of all say that the answer should have been that {\em Charlie} is the tallest. This true because the notion of being taller/shorter is transitive, so from the definitions of words and the two relations given, we can conclude that Charlie is the tallest. ChatGPT also suggests a possible problem scenario in which Alice is 6 feet tall, but Bob is 5'10", contradicting the assumption that {\em Alice is shorter than Bob}. Clearly, ChatGPT is not thinking in terms of rigorously consistent logical structures. We believe this is related to the fact that there is no clear distinction between true and false, in the generative pre-trained model. And, because there is no proof system inherent in the transformer that guarantees consistency. There is no mechanism for detecting a logical contradiction. \section{A Symbolic Probability Database} \subsection{The Concept of a Discrete Probabilistic Database} We can refer generally to a {\em discrete logical database}, as a set of pairs $\left\{(z, p) \in (\ell, [0,1]) \right\}$, where each $z$ is a statement in $\ell$ and each $p$ is its associated probability. One example of such a network would be a {\em Bayesian Network} \citep{koller:propabilistic-graphical-models:2009}. A Bayesian Network could satisfy all of the properties we need of evaluating and learning from sentences in $\ell$, but for the problem of {\em universal quantification}, i.e., sentences of the form $\forall x\ f(x) \rightarrow g(x)$. More generally we can look to various theories of {\em logic programming} \citep{deraedt:probabilistic-inductive-logic-programming:2008}, which do attempt to address the quantification issue. There is no clear sense of unity or standardization so far, it seems, in theories of logical inference like there is in neural network inference. \subsection{Motivation for a Discrete Probabilistic Database} We propose that by making the knowledge database discrete, we will avoid the problem of hallucinations, because we have two clearly stored dimensions: \begin{itemize} \item discrete database membership: for a given sentence $z$, either $(z, p) \in K$ for some $p$, or it is not. We can thus detect whether a fact has actually been seen, and incorporated, before. \item ability to evaluate likelihood: for a given sentence $z$ that is in $K$, we can evaluate whether its associated probability $p$ is high or low \end{itemize} Given this information correctly constructed, hallucination could be avoided, because unseen things are not in this database. \subsection{Requirements of a Logical Database} In order to incorporate a discrete probabilistic model into our generative story, we must assume a system that has the following properties: \begin{itemize} \item given a model $\Phi$, can efficiently assign a probability $P_\Phi(z)$ for any $z \in \ell$ \item given a database of logical facts $\left\{z_n\right\}_{n \in U}$ of all facts in the entire universe $U$, we can call $\Phi = \textsc {ReTrain}(\left\{z_n\right\}_{n \in W})$ \end{itemize} Presumably, we would want to be able to make ``on-line'' updates to the model, and it will not be practical to re-estimate all parameters all of the time. However, we want to leave $\Phi$ an abstract object, and so avoid considerations of how updates could be made online, to avoid restraining the model class prematurely, and leave optimizations for future work. \section{Syntax and Semantics} \subsection{Traditional Conception of Linguistics} The traditional conception of linguistics was that a layer called {\em syntax} served to intermediate between the two logical interfaces of {\em logical form}, in which reasoning is done, and {\em surface form}, corresponding to what is written or spoken \citep{sausure:course-in-general-linguistics:1916,chomsky:minimalist-program:2014,steedman:syntactic-process:2001}. Thus, the fact that ChatGPT does not rely on syntactic structures is perhaps the most surprising feature, and has on this grounds drawn criticism from Chomsky \citep{chomsky:debunking-the-great-ai-lie:2023}. We believe that the fact that linguistics has been seen as involving a layer of syntactic analysis is a major kind of evidence that the syntactic layer has relevance. It is an amazing achievement by ChatGPT that they can build a world knowledge model by avoiding this layer. However, the ever-present intuition that syntactic analysis is relevant suggests that in the future, syntax must make a comeback in natural language processing. \subsection{Mapping Logical Form to Surface Form} Let $z_n$ be a logical form, $x_n$ be a paragraph of surface tokens, and $y_n$ be a parse that mediates between them. For a given paragraph $x_n$, the set of candidate parses $C(x_n)$ will grow more than exponentially in the length of $x_n$. However, in practice this space can be pruned using, if necessary using a $k$-best pruning from an existing parsing model \citep{charniak:coarse-to-fine:2005,huang:forest-reranking:2008}. We will assume that a parse $y$ uniquely determines a logical form $z$, so that we can write that the set of candidate parses for a sentence $x$ is $C(x)$ a set of pairs $(y, z)$, where $y$ is a parse that yields $x$, and $z$ is the unique logical form determined by $y$. \section{Generating Text with Latent Logical Forms} \subsection{ChatGPT's Generative Story} Let ${\bf x} = \left[x_n\right]_{n=1}^N$ be a document of $N$ paragraphs, each of length $W$ tokens. A ChatGPT-style language model \cite{radford:improving-language-understanding-by-generative:2018} models the probability of a document as: \[ p({\bf x}) = \Pi_{i=1}^n p(x_n \mid x_{n-1}) \] That is, in the generative story, one paragraph is created based only on a previous one, with no hypothesized latent state of any kind. \subsection{High-Level Flow of a Latent Variable Generative Story} We want to create a generative story, which we call {\em SymbolicGPT}, in which, to create $x_n$ based on $x_{n-1}$: \begin{itemize} \item with probability $p(z_n \mid x_{n-1}; \theta, \Phi)$, choose a logical form $z_n$, based on (encoding vectors from) the previous paragraph $x_{n-1}$, the neural parameters $\theta$, and the prior likelihood of articulating the logical concept $z_n$ according to logical database $\Phi$ \item with probability $p(y_n \mid x_{n-1}, z_n; \theta)$, choose a syntactic analysis $y_n$, based on $z_n$, (encoding vectors from) the previous paragraph $x_{n-1}$, and the neural parameters $\theta$ \item the output text $x_n$ is determined uniquely by $y_n$ \end{itemize} \subsection{Factorization of Logical Forms} We suppose that the $x_n$ are {\em paragraphs}, which is to say sequences of sentences. We assume that we receive the document or corpus segmented into paragraphs $x_n$, but that each $x_n$ contains within it logically separable sentences $[s_1, ..., s_{S_{x_n}}]$, but that tokenization internal to $x_n$ is to be done internal to the model we are describing. The idea is that since each $x_n$ represents a sequence of logically separable sentences, we can use the distribution over sentence interpretations from one logical sentence within $x_n$, to mutually help inform the interpretation of another logically separable sentence within $x_n$. That is, in the generative story, we include a function $p(z_n \mid x_{n-1})$, rather than $p(z_n \mid z_{n-1}, x_{n-1})$. That is, $z_n$ does {\em not} depend on $z_{n-1}$, to avoid a blowing up in complexity, both conceptual and run-time, of the computational graph. \subsection{Generating the Logical Form} First, we choose a logical form $z_n$, based on $x_{n-1}$, $\theta$, and $\Phi$: \[ p(z_n \mid x_{n-1}, \Phi, \theta) \] One straightforward way to factor this is as a product of the syntactic likelihood of producing $z_n$ (as a syntactic object), in the given context: \[ p(z_n \mid x_{n-1}, \theta) \] And then the logical likelihood of saying $z_n$ at all, according to $\Phi$: \[ p(z_n \mid \Phi) \] \subsection{Generating the Parse} We next choose the parse $y_n$ in the normal transformer way: \[ p(y_n \mid z_n, x_{n-1}; \theta) \] There are a number of different parse formalisms (CFG parsing, unlabeled dependency parsing, labeled dependency parsing), and we are not currently proposing to choose between them. The only requirement of the parsing formalism is that it must map a surface form $x_n$ to some $z_n$ that we can reason with. One can use a discriminative parser to constrain the space, and this would create a parse in a bottom-up fashion. Once the space is pruned and structured, the generative story can be told in a top-down fashion, like in the original parsing models \citep{collins:a-new-statistical-parser:1996,charniak:statistical-parsing-with-a-cfg:1997}. \subsection{Generating the Text} The text $x_n$, we reiterate, is a part of the parse $y_n$. \section{Parameter Estimation} The problem of parameter estimation in the ChatGPT GPT model is the problem of estimating what we are calling $\theta$, the continuously valued parameters of the neural network part of the model. In our case, we must also now estimate $\Phi$, which is the discrete logical model. And, we must hypothesize latent syntactic/logical parses. \subsection{Expectation Maximization} The first problem to be solved is that the logical forms are {\em not} part of the given data, and so must be modeled as {\em latent variables}. The ChatGPT transformer does not use latent variables in the way that they were understood in the context of part-of-speech tagging or syntactic parsing, and this greatly simplifies the training. However, in order to realize the traditional vision of syntax as mapping logical forms to surface form \citep{montague:universal-grammar:1970,chomsky:minimalist-program:2014,steedman:syntactic-process:2000}, we will have to involve a logical representation, which is not present in the data, and therefore is latent. In order to train with latent variables, we can use {\em expectation maximization} \citep{dempster:maximum-likelihood:1977}. In the expectation step, we assign distributions to possible parses $y_n$, each of which implies an associated logical representations $z_n$. In the {\em maximization} step, we optimize $\theta$ and $\phi$. Boot-strapping this process will perhaps be difficult. One option is to start by boot-strapping from a high resource language, e.g., English. A syntactic parser can be discriminatively trained based on a GPT generative model, the way that few shot learning happens today. If logical world knowledge can be encoded based on English, then other languages can map onto this logical space. \subsection{Heterogenous Training} The neural network part of the model can be trained using back-propagation. It is not clear exactly how the logical model would be trained, or whether back-propagation would be appropriate. We have specified that we require the abstract interface $\Phi = \textsc {ReTrain}(\left\{z_n\right\}_{n \in W})$ to retrain $\Phi$, but how this would be done is difficult future work. However, in order to allow the most generality, we can train the neural network part in a heterogeneous way compared to the logical model. The predictions of the logical model can be treated as exogenous features, and back-propagation can work through the exclusively neural part. \subsection{The Logical Model} The fundamental difficulty in estimating a model using symbolic logic is the question of, what are the symbols? In other words, the problem for the Bayesian network formulation is that we must identify both a graphical structure $G$, not known a priori, and a probability model $\Phi$ over these. The question of how this would work must be left to future work. However, for a fixed directed Bayesian network, the problem of parameter estimation is trivial, and just involves counting. \section{Discussion} We are not able to evaluate this model at the present time. We only record the logic and the mathematics in order to receive feedback on the equations. However, we reiterate the benefits we expect to see from this model. Hallucinations can hopefully be removed from the use of a discrete membership database, where a fact is either in the database or not, and then if it does exist, it either has high or low probability. The discrete logical database will allow for the creation of logical reasoning, and an ability to detect whether things actually contradict. \section{Conclusion} We have introduced {\em SymbolicGPT}, a new algorithm that proposes to integrate a symbolic logic grammar into a ChatGPT-style generative pre-trained transformer, with the goal of eliminating hallucinations, and creating a logically consistent worldview. This new architecture requires the postulation of a latent syntactic form, which can be learned through expectation-maximization. Of course, this is only a theoretical proposal, and not an empirical result. The obvious future work is to instantiate the parameters of this theory to get it actually working. A narrower future work would be to simply refine the logical reasoning engine. \bibliography{anthology,custom} \bibliographystyle{acl_natbib} \end{document}#5,144,592text\pdfoutput=1 \documentclass[11pt]{article} \usepackage{amsmath} \usepackage{EMNLP2023} \usepackage{times} \usepackage{latexsym} \usepackage[T1]{fontenc} \usepackage[utf8]{inputenc} \usepackage{microtype} \usepackage{inconsolata} \title{Symbolic Logic and Syntax in a Generative Model of Text} \author{Greg Coppola \\ {\em Apocalypse Like Right Now} \\ \texttt{[email protected]}} \begin{document} \maketitle \section{A Symbolic Probability Database} \subsection{The Concept of a Discrete Probabilistic Database} We can refer generally to a {\em discrete logical database}, as a set of pairs $\left\{(z, p) \in (\ell, [0,1]) \right\}$, where each $z$ is a statement in $\ell$ and each $p$ is its associated probability. One example of such a network would be a {\em Bayesian Network} \citep{koller:propabilistic-graphical-models:2009}. A Bayesian Network could satisfy all of the properties we need of evaluating and learning from sentences in $\ell$, but for the problem of {\em universal quantification}, i.e., sentences of the form $\forall x\ f(x) \rightarrow g(x)$. More generally we can look to various theories of {\em logic programming} \citep{deraedt:probabilistic-inductive-logic-programming:2008}, which do attempt to address the quantification issue. There is no clear sense of unity or standardization so far, it seems, in theories of logical inference like there is in neural network inference. \subsection{Motivation for a Discrete Probabilistic Database} We propose that by making the knowledge database discrete, we will avoid the problem of hallucinations, because we have two clearly stored dimensions: \begin{itemize} \item discrete database membership: for a given sentence $z$, either $(z, p) \in K$ for some $p$, or it is not. We can thus detect whether a fact has actually been seen, and incorporated, before. \item ability to evaluate likelihood: for a given sentence $z$ that is in $K$, we can evaluate whether its associated probability $p$ is high or low \end{itemize} Given this information correctly constructed, hallucination could be avoided, because unseen things are not in this database. \subsection{Requirements of a Logical Database} In order to incorporate a discrete probabilistic model into our generative story, we must assume a system that has the following properties: \begin{itemize} \item given a model $\Phi$, can efficiently assign a probability $P_\Phi(z)$ for any $z \in \ell$ \item given a database of logical facts $\left\{z_n\right\}_{n \in U}$ of all facts in the entire universe $U$, we can call $\Phi = \textsc {ReTrain}(\left\{z_n\right\}_{n \in W})$ \end{itemize} Presumably, we would want to be able to make ``on-line'' updates to the model, and it will not be practical to re-estimate all parameters all of the time. However, we want to leave $\Phi$ an abstract object, and so avoid considerations of how updates could be made online, to avoid restraining the model class prematurely, and leave optimizations for future work. \section{Syntax and Semantics} Let $z_n$ be a logical form, $x_n$ be a paragraph of surface tokens, and $y_n$ be a parse that mediates between them. For a given paragraph $x_n$, the set of candidate parses $C(x_n)$ will grow more than exponentially in the length of $x_n$. However, in practice this space can be pruned using, if necessary using a $k$-best pruning from an existing parsing model \citep{charniak:coarse-to-fine:2005,huang:forest-reranking:2008}. We will assume that a parse $y$ uniquely determines a logical form $z$, so that we can write that the set of candidate parses for a sentence $x$ is $C(x)$ a set of pairs $(y, z)$, where $y$ is a parse that yields $x$, and $z$ is the unique logical form determined by $y$. \section{Generating Text with Latent Logical Forms} \subsection{ChatGPT's Generative Story} Let ${\bf x} = \left[x_n\right]_{n=1}^N$ be a document of $N$ paragraphs, each of length $W$ tokens. A ChatGPT-style language model \cite{radford:improving-language-understanding-by-generative:2018} models the probability of a document as: \[ p({\bf x}) = \Pi_{i=1}^n p(x_n \mid x_{n-1}) \] That is, in the generative story, one paragraph is created based only on a previous one, with no hypothesized latent state of any kind. \subsection{High-Level Flow of a Latent Variable Generative Story} We want to create a generative story, which we call {\em SymbolicGPT}, in which, to create $x_n$ based on $x_{n-1}$: \begin{itemize} \item with probability $p(z_n \mid x_{n-1}; \theta, \Phi)$, choose a logical form $z_n$, based on (encoding vectors from) the previous paragraph $x_{n-1}$, the neural parameters $\theta$, and the prior likelihood of articulating the logical concept $z_n$ according to logical database $\Phi$ \item with probability $p(y_n \mid x_{n-1}, z_n; \theta)$, choose a syntactic analysis $y_n$, based on $z_n$, (encoding vectors from) the previous paragraph $x_{n-1}$, and the neural parameters $\theta$ \item the output text $x_n$ is determined uniquely by $y_n$ \end{itemize} \subsection{Factorization of Logical Forms} We suppose that the $x_n$ are {\em paragraphs}, which is to say sequences of sentences. We assume that we receive the document or corpus segmented into paragraphs $x_n$, but that each $x_n$ contains within it logically separable sentences $[s_1, ..., s_{S_{x_n}}]$, but that tokenization internal to $x_n$ is to be done internal to the model we are describing. The idea is that since each $x_n$ represents a sequence of logically separable sentences, we can use the distribution over sentence interpretations from one logical sentence within $x_n$, to mutually help inform the interpretation of another logically separable sentence within $x_n$. That is, in the generative story, we include a function $p(z_n \mid x_{n-1})$, rather than $p(z_n \mid z_{n-1}, x_{n-1})$. That is, $z_n$ does {\em not} depend on $z_{n-1}$, to avoid a blowing up in complexity, both conceptual and run-time, of the computational graph. \subsection{Generating the Logical Form} First, we choose a logical form $z_n$, based on $x_{n-1}$, $\theta$, and $\Phi$: \[ p(z_n \mid x_{n-1}, \Phi, \theta) \] One straightforward way to factor this is as a product of the syntactic likelihood of producing $z_n$ (as a syntactic object), in the given context: \[ p(z_n \mid x_{n-1}, \theta) \] And then the logical likelihood of saying $z_n$ at all, according to $\Phi$: \[ p(z_n \mid \Phi) \] \subsection{Generating the Parse} We next choose the parse $y_n$ in the normal transformer way: \[ p(y_n \mid z_n, x_{n-1}; \theta) \] There are a number of different parse formalisms (CFG parsing, unlabeled dependency parsing, labeled dependency parsing), and we are not currently proposing to choose between them. The only requirement of the parsing formalism is that it must map a surface form $x_n$ to some $z_n$ that we can reason with. One can use a discriminative parser to constrain the space, and this would create a parse in a bottom-up fashion. Once the space is pruned and structured, the generative story can be told in a top-down fashion, like in the original parsing models \citep{collins:a-new-statistical-parser:1996,charniak:statistical-parsing-with-a-cfg:1997}. \subsection{Generating the Text} The text $x_n$, we reiterate, is a part of the parse $y_n$. \section{Parameter Estimation} The problem of parameter estimation in the ChatGPT GPT model is the problem of estimating what we are calling $\theta$, the continuously valued parameters of the neural network part of the model. In our case, we must also now estimate $\Phi$, which is the discrete logical model. And, we must hypothesize latent syntactic/logical parses. \subsection{Expectation Maximization} The first problem to be solved is that the logical forms are {\em not} part of the given data, and so must be modeled as {\em latent variables}. The ChatGPT transformer does not use latent variables in the way that they were understood in the context of part-of-speech tagging or syntactic parsing, and this greatly simplifies the training. However, in order to realize the traditional vision of syntax as mapping logical forms to surface form \citep{montague:universal-grammar:1970,chomsky:minimalist-program:2014,steedman:syntactic-process:2000}, we will have to involve a logical representation, which is not present in the data, and therefore is latent. In order to train with latent variables, we can use {\em expectation maximization} \citep{dempster:maximum-likelihood:1977}. In the expectation step, we assign distributions to possible parses $y_n$, each of which implies an associated logical representations $z_n$. In the {\em maximization} step, we optimize $\theta$ and $\phi$. Boot-strapping this process will perhaps be difficult. One option is to start by boot-strapping from a high resource language, e.g., English. A syntactic parser can be discriminatively trained based on a GPT generative model, the way that few shot learning happens today. If logical world knowledge can be encoded based on English, then other languages can map onto this logical space. \subsection{Heterogenous Training} The neural network part of the model can be trained using back-propagation. It is not clear exactly how the logical model would be trained, or whether back-propagation would be appropriate. We have specified that we require the abstract interface $\Phi = \textsc {ReTrain}(\left\{z_n\right\}_{n \in W})$ to retrain $\Phi$, but how this would be done is difficult future work. However, in order to allow the most generality, we can train the neural network part in a heterogeneous way compared to the logical model. The predictions of the logical model can be treated as exogenous features, and back-propagation can work through the exclusively neural part. \subsection{The Logical Model} The fundamental difficulty in estimating a model using symbolic logic is the question of, what are the symbols? In other words, the problem for the Bayesian network formulation is that we must identify both a graphical structure $G$, not known a priori, and a probability model $\Phi$ over these. The question of how this would work must be left to future work. However, for a fixed directed Bayesian network, the problem of parameter estimation is trivial, and just involves counting. \bibliography{anthology,custom} \bibliographystyle{acl_natbib} \end{document}#5,109,059text# May The Force Be With You **Remarks on Artificial Intelligence for May 4, 2023** May 4, 2023 Gregory Francis Coppola *Apocalypse Like Right Now* # Introduction This is a collection of essays on the topics of 1) *artificial intelligence* and 2) *human intelligence*, published in honor of May 4, 2023. *****May 4***** is a day when *********Star Wars********* fans traditionally wish one another “may the *fourth* be with you”, in a humorous reference to the Star Wars line “may the *****force***** be with you”. The *****force***** in Star Wars is a whimsical metaphor for the “energy” or “intelligence” that reportedly pervades the universe. # The Turing Test is Passed ## Background The Turing Test was first proposed by Alan Turing in his 1950 paper *Computing Machinery and Intelligence*. Only 20 years ago, when I started as an undergraduate in Computer Science at The University of Waterloo, it seemed that the passing of the Turing test would be several generations away. 10 years ago, I was finishing my PhD in artificial intelligence at The University of Edinburgh, and would have put the passing of the Turing test 30-50 years away. In 2023, the milestone has apparently been reached. My, has time flown. ## The Turing Test is so Passed Right Now Many technical reports and much anecdotal experience has suggested that the Turing test is effectively now “passed”. For example, A. I. thought leader Joshua Bengio recently said: > Now, why did I sign [i.e., a letter to pause large A. I. experiments]? Like why now? Like maybe last year I wouldn’t have signed this letter. It’s because now we reached the threshold. The threshold is the Turing test, meaning that we have systems that we can dialogue with, and we can’t be sure if this is coming from a machine or a human. (*Yoshua Bengio: Pausing More Powerful AI Models and His Work on World Models, Eye on AI, Youtube*, April 12, 2023) > ## Implications ### A Time for Celebration We believe that the passing of the Turing Test is a major event that should, first of all, be celebrated more consciously. There has been remarkably little celebration of this fact either in the field of artificial intelligence. Researchers are too busy nit-picking, and the public is too busy worrying about re-training for new careers, for there to have ever been any acknowledgement that a human achievement on par with the trip to the moon has now been achieved. ### A Time for Reflection Geoff Hinton has recently left his job at Google, in part, he said, to spend his time writing and speaking on the ramifications of the artificial intelligence, saying that we may be very close to when computers are smarter than people (*AI 'godfather' quits Google over dangers of Artificial Intelligence - BBC News*, *BBC News*, *Youtube*, May 2, 2023). We completely agree that this is a necessary time for thought about the implications of artificial intelligence for humanity. # The Optimal Philosophy of Science ## The Historical Problem of Philosophy Philosophy, in its modern sense, refers to the use of logic to answer questions, such as: 1. what we can know 2. how we can know it 3. what to do about what we know ## Science is the Answer to Most of Philosophy Of these three questions, the resolutions of two of them are found in *******science*******. That is: - solved by science - what can we know - how can we know it - not solved by science - what to do about what we know In other words, the only part of what was traditionally called philosophy, once the empirical parts have been ceded to science is: - remaining task for philosophy - decide **********what to do********** about what the world that *******science******* describes - decide **********how to do********** science ## The Philosophy of Science is Necessary for Science The questions of ******************what can be known****************** through science and **********how to do********** science are inextricably linked, and the study of these questions has for many years been known as the “philosophy of science”. The question of *****************what can be known***************** has been a question for philosophy, especially Plato, Aristotle and Kant, and remains so. The question of ***********************how to go about knowing it*********************** was initially part of philosophy, but has now been adapted and refined to a more precise point by artificial intelligence. That is, an artificial intelligence program can only do science if its algorithms, and by extension, its equations, are accurate. Empirically, it is not easy to build a program that passes the Turing Test. Thus, for a program to finally do it, this implies that the equations must be very accurate. Thus, the question of ***********how to know*********** what science can know is now more developed in the field of artificial intelligence than it is in philosophy. ### An Artificial Intelligence Perspective Doing science is analogous to the **********generative********** phase of the GPT model. That is, the best model is the statistical model that best compresses the data, and at the same time has the smallest models size. It has been remarked that there is a model size beyond which it does not help to train a ChatGPT-style model. In other words, even in the “biggest” data sets now, one must consider model size. ### Minimum Description Length The rule of ***************Ockham’s Razor*************** says that the “simplest theory that fits the facts” is best. But, this always left two questions: 1. how many observations does a theory need to “predict” in order to justify itself 2. how to evaluate the “simplicity” of a theory, to compare two theories for “simplicity” Information theory gave an interesting perspective on this question. According to the principle of **************************minimum description length**************************: 1. the goal is to minimize the sum of 1. the size of the model, expressed as a computer program 2. the compressed size of the data, given the model This gives an answer to the problem of Ockham’s razor, which is ***how*** to measure the complexity of the theory versus the observations predicted. ## Specific Applications in 2023 ### Atheism vs. Agnosticism On the question of whether or not we have a creator, atheists propose the default position is that they categorically **do not believe** in a creator, because ******they have not been convinced of one******. A popular atheist who has made this argument is Sam Harris: > It’s [i.e., atheism] is not even the assertion that there is no God. It’s just a failure to be convinced of any of the Gods on offer. It’s just like not believing in Zeus. (AD Harris/Murray/Peterson Discussion: London, Jordan B. Peterson, Youtube, September, 14, 2018) > In other words, Harris believes that a rational default is to start out atheist—the positive assertion that there *is no* creator—until evidence comes in to argue positively ***for*** a creator. But the principle of minimum description length would advise us to stay neutral on a question, until we are ready to use it to predict data, because, in order to justify using up theory space, we would have to gain back the cost of the theory in predictions. But, ahead of seeing any data, we don’t have any reason to take a position, because there is no prediction to make. ### The Complexity of the Theory of God Richard Dawkins argues that, from a scientific perspective, a theory including God must always lose to a theory that does not contain God, because God is “infinitely complex”, and, by Ockham’s Razor, simpler theories should win. From the perspective of minimum description length, we can say that Dawkins’ error is: - Dawkins’ error - we only need to pay for the ******************theory that we use****************** according to minimum description length - we don’t get penalized for the inherent complexity in the object itself ### The Complexity of Infinite Universes In order to explain the “finely tuned universe”, atheists will often resort to the argument of “infinite universes”. Assuming that the entire universe must be represented somewhere in order to model it, then the theoretical complexity of infinite universes, in terms of minimum description length, is infinite. Since the God hypothesis is finite in complexity, God is a more parsimonious explanation for the finely tuned universe than infinite universes. # Ego, Id, and Subsystem Analysis ## Freud’s Framework Sigmund Freud famously proposed a partition model of the human psyche. For our purposes, the two primary systems that Freud has identified are: - the **id**, is described variously as - The unconscious and primitive part of the human psyche - The unconscious and primitive part of the human psyche - The source of instinctual drives and impulses - The part of the psyche that operates on the *pleasure principle* - The reservoir of repressed or socially unacceptable desires and emotions - the ***ego***, is described variously as - The conscious part of the psyche that mediates between the demands of the id and the constraints of reality - The part of the psyche that operates on the *reality principle* - The sense of self or personal identity that develops as a result of interactions with the external world - The seat of rational thought and decision-making Freud also discusses the *********super-ego*********, which represents the desires of one’s community internally. ### Creativity in Freud’s Framework It is unclear where **********creativity********** is located in Freud’s framework, there is controversy as to where it comes from. We believe ***creativity*** is associated with the id, or the inner child, a theory consistent with the following passage: > The creative writer does the same as the child at play. He creates a > > > world of phantasy which he takes very seriously—that is, which he in- > > vests with large amounts of emotion—while separating it sharply from > > reality. (*Creative Writers and Day-Dreaming*, Freud, 1908) > ## Criticism of Freud ### Recognition for Freud’s Achievement Sigmund Freud’s distinction between ***ego*** and **id**, and his delineation of the ****subconscious**** is one of those ideas that is at first original, but then so immediately pervasive, that people forget it was ever original. It is one of the greatest contributions by a single thinker in the history of Western thought. It is for its centrality that we want to build on Freud’s system. ### Limitations of Freud’s System Freud’s system, however, has some limitations, as it naturally would given that Freud’s ****The Ego and the Id**** was published in 1923. This is for the following reasons: - subjective data sets - a lot of Freud’s ideas came evidently from inspection of his own thoughts - Freud’s thinking was done under the influence of controlled substances - especially cocaine - informal data sets - Freud was not able to capture detailed psychometric data about his clients - limited data sets - Freud’s data was limited to a small amount of data because - he was dealing with clients, of which one can interview a limited number - his practice was related to a specific time and place, circa 1923 Vienna - no mathematical or computational language - Freud did not have access to the kinds of mathematical language and computational paradigms that we have today - the Ego and the Id was published in 1923 - while Alan Turings *On Computable Numbers, with an Application to the Entscheidungsproblem*, which introduced the “Turing machine”, was published in 1936 ## Concrete Open Questions for the Freudian Framework Freud’s approach leaves open the following questions: - how many subsystems are there? - how do we justify a delineation of the subsystems - what kind of data is applicable? ## Freud from an Artificial Intelligence Perspective It remains difficult to do neuroscientific experiments to identify the potential boundary between the ego and the id. Thus, we propose to clarify Freud’s subsystems analysis using computer science. This can take two forms: - computationally rigorous subsystems models of the human brain - use concepts from computer science like - function - argument - memory - use analogies to existing computer systems - like databases - refer to A.I. models as they exist today, in the vein of ChatGPT - draw analogies between A. I. systems and humans - in order to provide a unified account Concretely, we propose: - theses on Freud and artificial intelligence - artificial intelligence helps us understand Freud - Freud helps us understand artificial intelligence ## The Ego and the Id in ChatGPT Open AI (*Improved Language Understanding by Generative Pre-Training*) has proposed a two step process by which they leverage the large generative language models: 1. pre-train a large language model on text 2. train a discriminative model to do a task based on the generative model To us, there is a striking analogy between 1) Freud’s ego-id distinction, and 2) the distinction between the generative model and the discriminative model. In other words, the analogy is: 1. **generative model, id** 1. analogous to the **Id** 2. the source of “generative” power 3. “generates” the dataset 4. has a model of the world 5. must accurately model the universe for maximum effect 2. **discriminative model, ego** 1. analogous to the ***ego*** 2. uses “discriminative” methods in order to predict an answer given a context 3. does not have a general model of the world ### The Tension Between Knowledge and Relationships In a sense, it makes sense that there would be a natural cleavage site between these: 1. the need for knowledge - an organism surviving in the world needs to know as much about the world as possible 2. the need for social survival - in order to survive, a social animal must survive ********socially******** - when an entire group detaches from reality, then it is dangerous for the group In order to account for both, it seems necessary that the generative “id” and the discriminative “ego” would be as they are, and not the other way around. That is: - **generative** - in order to actually “generate” the (joint distribution of) the data set, one must have an accurate model of the world - if the generative model were forced to “say things that aren’t true” in the generative modeling stage, this can only reduce the effectiveness of compression of the data - because, by assumption, the most compressive statement was that which was chosen in an unrestricted setting - **discriminative** - does not need to ever “generate” data - does not need to actually learn a model of the universe - only needs to focus on tasks that “have a beneficial outcome” # A. I. Can be Unbiased ## The Concept of Bias The etymology of "bias" can be traced back to the Old French word "biais," which originally meant "oblique." This word likely comes from the Old Provençal "biais," meaning "sideways" or "slanting." Over time, "bias" came to refer to a diagonal line or cut in cloth, which was used to create a particular effect or pattern. The modern usage of "bias" in the sense of a mental or emotional inclination is thought to have originated in the 16th century. It was first used in the context of archery, where a bias was a weight added to one side of an arrow to make it curve in flight. From there, the term came to be used more broadly to refer to any influence that causes something to deviate from its expected or normal course. In modern literature terms, "bias" generally refers to a systematic error or distortion in research or data analysis that arises from factors such as flawed study design, measurement errors, or cultural or social biases. ## Reality versus Ideology In modern terms, when people refer to bias, one way to distinguish between different uses of the word are: 1. deviation from the data 1. this means the the model does not accurately reflect reality or the data - usually, when a model does this, it is out of an interest to “say the politically correct” thing, when the data contradicts this 2. deviation from ideology - this is when conclusions contradict with the prevailing ideology, also called “political correctness” - this is called “algorithmic fairness” In this section, we are interested bias as “deviation from the data”. ## Two Sources of Disagreement among Humans For immediate purposes, we say that an *agent* is an object that can think, speak and act. A *disagreement* arises whenever two different agents espouse contradictory statements. That is, for some statement $T$, agent $A$ espouses $T$ and some other agent $B$ espouses $\neg T$. Disagreements arise for two conceivable reasons: - different values - different agents can have competing interests - they prefer to be in different situations - the tendency thus grows to create a story which differs from reality in ways that are perceived to help the agent - or else, one has a bias away from believing things that could hurt the agent, when there is uncertainty - different interpretation of the data - differential ******access****** to data - different agents have access to different data - information is valuable and so often kept secret - differential *ability to process* data - not every organization or person is equally good at information processing - some people have higher processing capacity in their neurons - some organization have higher processing capacity in their data centers To summarize, we have identified two reasons that disagreements over facts occur: - different values - in which there are **********incentives********** for people to internally or externally view the world in a particular way, to support their own interests - different reads on the data - here, people’s incentives may be aligned, but they are unable to come to an agreement ## Bias as a Prior One notion of “bias” that is interesting, but that we are also not focused on is the use of an informative Bayesian prior. We will assume that we have access to all of the data “in one go”, and so do not use an informative prior. ## Optimal Algorithm of Science Suppose there is an optimal algorithm of science $A$ for a data set. That is, for a data set $D$, if we run $A$ on $D$, then the result will be the “best theory possible” given the data $D$. In practice, we believe that the optimal algorithm for science would be some instantiation of the ***************************minimum description length*************************** principle. At this point, in any imaginable sense, the “optimal” algorithm is the one that best compresses the data. There are strong theoretical guarantees for this, as well as empirical success of the idea of “prediction as compression”. ## Why a Computer Can (Mostly) Overcome Scientific Bias ### Review Let us review the reasons that a person can form a biased (incorrect) view of the world: - incomplete data - limited agents always have incomplete data - however, some people have access to more information than others - limited processing ability - until now, no human or organization has had enough processing capability to understand the entire world - values that conflict with the truth, and so makes it difficult for a person to see the “truth” - a person might be incentivized to disagree with the “scientific” truth ### Why a Computer can be Unbiased Let us consider the case of an unlimited-size computer, operating on a huge data set including 1) all public data, and 2) a large collection of private data: - incomplete data - suppose one has access to all of the data of the entire human race - obviously, there are certain questions that cannot be answered, because no one has the data - but, for many questions, those questions can be answered from the data - in such a case, the sum total of human data **is** enough to answer such questions - so, the limited data is not a problem, for an A. I. tool that can read all human recorded data - limited processing ability - an unbounded A. I. agent does not have a limited processing ability - thus, limited processing ability is not a problem for an unbounded A. I. agent - conflicts of interest that prevent the discovery of the truth - we have supposed in the last section that there is an “optimal” algorithm for science - whether there is a unique such algorithm or not, experience has shown that already algorithms exist which - approach the true data distribution in practice - approach the true data distribution in principle - also, algorithms for science are based on the notion of “compression”, which is neutral on all actual empirical questions, and leaves the conclusions up to the data, in a defined way - thus, we believe that an algorithm can be unbiased in terms of its “values” ## Conclusion If, by “unbiased” we mean that an algorithm accurately represents the world and/or our data about the world, we believe that a computer with 1) infinite computational resources, 2) unbounded access to human data, 3) no value-driven bias, but instead an “unbiased algorithm”, can actually be “unbiased”. # Morals Must be Programmed into A.I. The philosopher Sam Harris has tried to revive the idea that morality can be gotten from science. In other words, he wants to revive the idea that an “ought” can be gotten from an “is”. It has traditionally been thought impossible for an “is” to imply an “ought”, although we must hasten to add some historical references. In this essay, we will show from a scientific perspective in the case of a ChatGPT-style model, why values (and therefore policies) *******do not******* come from the data, but come from an exogenous step. ## Perspectives on Facts Versus Values ### Decision Theory According to decision theory, the expected utility of an action is: - results for an action - let $R(a)$ be the set of all *******results******* that can result from taking action $a$ - expected utility of action $a$ - $u_{\theta,v}(a) = \Sigma_{r \in R(a)}\ p_\theta(r|a)\cdot v(r)$ - choice of action - out of possible actions $a \in D$, choose the action that maximizes $u(a)$ Thus, we see that the probability model $p_\theta$ is completely separate from the value function $v$. ### Generative Pre-Trained Models A generatively pre-trained model with a discriminative second stage can be thought of as follows: - science doing, generative part - the generative part, with parameters $\theta$ - *builds* a ***********generative*********** model of the outside world - by building a generative model of the data - task-oriented discriminative part - discriminative part, with parameters $\beta$ - *****uses***** the generative model to accomplish tasks - does not do science - implicitly represents values ### Conclusions From both perspectives, that of decision theory and that of generative pre-training, we find the same story. That is, the part of the program that scientifically models the world, and the part of the program that uses the world model to accomplish tasks, are separate. In the case of decision theory, we see that the probability model is different from the values that drive decision making. In the case of a generative pre-trained model, we see that the generative model, which models the outside world, is different from the discriminative model, which uses the world knowledge to accomplish tasks. ## Conclusions for Humans Thus, we see, from an artificial intelligence perspective, why: - statements confirmed from A. I. perspective - facts do not determine values - scientific conclusions about what **is** do not imply **********what ought to be********** # Logic is Also Needed ## The Inherent Hilarity of the Meme Title *************************Attention is all you Need************************* The paper that launched the modern GPT revolution is ***Attention is All You Need*** (2017, Google). While “attention” (or, in general the transformer) was “all that was needed” in order to pass the Turing Test, it seems that, more will also be needed. In particular, experience has shown that the transformer models do not currently reason logically with arbitrary precision, and they also hallucinate. Thus, though the meme title caused enjoyment, it is presumably implicit in the assumption that “attention is all you need” is funny, that attention would eventually not be “all you needed”. We suggest that what else is needed is explicitly *symbolic* *****logic*****. ## The Problem of a Logically Consistent Worldview ### Hinton Notes the Problem of Logical Consistency Viewed from the perspective that we want full AGI, the problem with the current ChatGPT model is that it does not definitely have a “logically consistent” world view. This point was recently raised by Geoff Hinton: > People have different opinions and it [ChatGPT] has to have a kind of > > > blend of all these opinions so that it can model what anybody might > say. It's very different from a person who tries to have a consistent world view > … if you want to act in the world um it's good to have a consistent world > view (Geoff Hinton, *"Godfather of artificial intelligence" talks impact and potential of AI*, CBS Mornings, March 25, 2023) > ### Criteria of “Logical Consistency” The ability to hold a “logically consistent” worldview includes the following competencies: - the ability to expand a set of premises $P$ to some $P' \supset P$ using a consistent set of ***************inference rules*************** - the ability to detect a contradiction in a set $P$ of sentences - the ability to partition a contradictory set of sentences $P$ into maximally consistent sets of sentences, $P_1, ..., P_N$ - the ability to compare two theories $P$ and $Q$ - e.g., and determine which is a better fit for the data - the ability to explain a piece of data based on multiple consistent worldviews - i.e., to compute the probability of an observation under multiple different world views ### The Conspicuous Absence of Syntactic Analysis in ChatGPT The traditional view of parsing was that it mapped between logical and phonetic interfaces. The concept of there being a syntactic “parse” for the sentence goes back to Chomsky (1957, ********************Syntactic Structures********************). The concept of the syntactic parse having as its byproduct a semantic parse in the tradition of “semantics” goes back to Montague (1970, English as a Formal Language). Alternatively, it goes back to the “categorial grammars of Ajdukiewicz (1935) and Bar Hillel (1953). The ***********prima facie*********** most surprising thing about the ChatGPT model is that it eschews the concept of explicit structural analysis, one of the most prevalent notions in pre-RNN NLP. It makes sense, in a sense, that the resolution to the problem of logical consistency would be to return to the notion of structural analysis, resolving two problems at once: - we will have **logical consistency** in the model - we will be using **syntactic structural analysis**, as was always intuited to be needed ## Logic in Natural Language ### Types of Logic In order to do logic, there are a variety of formalisms, each of which capturing a different aspect of the representation of language: - propositional logic - allows statements like $A \cup B \rightarrow A$ - first-order logic - adds quantification, e.g., $\forall x\ p(x) \rightarrow q(x)$ - higher-order logic - allows the ability to compare ****sets****, which allows quantifiers like ****many**** or ****most**** - intensional logic - allows for *****names***** of sentences to be used as arguments - that is, a predicate can take the **idea** of a sentences as its argument, rather than a truth value ### Bayesian Networks Probabilistic inference in the context of logic can be represented using *****************Bayesian networks***************** (Probabilistic Graphical Models: Principles and Techniques, Koller and Fridman, 2009). Bayesian networks only natively support ********************propositional logic.******************** However, first-order logic can be simulated by including templates over propositions, e.g., $\forall x\ p(x) \rightarrow q(x)$ can be instantiated as $p(a) \rightarrow q(a)$. The relationship between sets that might have in some cases necessitated a switch to higher-order logic are now captured in the probabilistic relationships between premise and conclusion represented in the Bayesian network. Intensional logic much the same as quantified logic, but the arguments can either be propositions or “names” of propositions. While there remain details to be worked out, we believe that a basic inferential backbone similar to a Bayesian network is capable of representing the knowledge needed to compute probabilities over logical forms. Such a representation can also facilitate “back-propagation”, though we are unsure at present what kind of schedule or scheme this would be done on. ## The Logical GPT Thesis ### The Need for Logical Inference The first point that we want to note is that: - the need for actual logical inference - perfect logical inference is necessary to have a logically consistent worldview - transformers on their not logically equivalent to logical reasoning - there are empirical differences between the two - the two are not equivalent by construction - for a bounded length of transformer, one can find a logical puzzle that cannot be represented by that transformer Thus, the first conclusion is: - **we need to incorporate symbolic logic into the ChatGPT-style model** ### New Opportunity with ChatGPT The dream of creating a large database of logical knowledge dates back at least to the Cyc project (Lenat, 1984). The problem with that project, and with any project looking to encode world knowledge, is that *********************knowledge acquisition********************* becomes a bottleneck. This is because it was assumed that world knowledge would be ****************manually encoded****************, and manual encoding is a difficult process, especially because it requires skill and training to even encode such knowledge. That is, it was always found to be impossibly hard to: - assumed too hard 1. manually encode all of this knowledge 2. acquire it in an unsupervised way However, ChatGPT does now, it seems have the ability to encode world knowledge: - crucial new observation - knowledge can be acquired in an unsupervised way - even if this is being recorded in vector space, rather than discrete space right now Thus, the primary inhibitor to building a Cyc-style system has now been removed: - new potential with ChatGPT - ChatGPT **is** able to acquire knowledge in an unsupervised way - all that remains is to encode it in a discretely logical way ### Summary In summary, we have seen that: 1. ChatGPT needs logical to do reasoning 2. ChatGPT removes the previous bottleneck for encoding logical knowledge, which was the knowledge acquisition bottleneck ## Proposed Solution Our proposed solution style is to create a model in which: - every sentence is **************generated from************** a *logical representation* - the logical representation for sentence $x_n$ is $z_n$ - this logical representation is a structured ***************hidden variable***************, not present in the linguistic data - the logical representation can be scored for its probabilistic likelihood using a discrete logical model $\phi$ - the mapping from *****logic***** to ***************sentence tokens*************** is done using a ***************syntactic parse*************** - the parse we denote using $y$ - it is this syntactic parse that is used to generate the sentence - the sentence $x_n$ is a part of syntactic parse $y_n$ - the set of valid parses for a sentence $x$ is denoted $C(x)$ - this is the set of all parses such that $reduce(y) = x$ - the model is **********generative********** - but at the expense of introducing latent variables into the ChatGPT model - the generative story now includes a factor for the most likely interpretation according to a world-model ## The Generative Story of Logical GPT ### Notation for the ChatGPT-Style Model We consider the task of the generation of a text one “sequence” $x_n$ , at a time, and that the breaking of a document into sub-sequences (sentences) is given. That is, a document is a sequence of sequences: - $Document=[x_1, ..., x_n, ..., x_{N}]$ Given such an input, we then generate the probability of the document as: - sequence-language model objective - $p(Document) = \Pi_{i=1}^n p(x_n | x_{n-1}, x_{n-2}, ...)$ Each $x_n$ is actually a sequence of tokens $[t_{n,1}, ..., t_{n, m}]$, where each $x_{n,1}$ is a word or character in the language. In other words, $x_n$is a sequence of words, a sequence of characters, or a sequence of multi-character phrase fragments (byte-pair encoding). ### Generative Story for a Sentence in ChatGPT In a sequence-to-sequence transformer, the goal is to generate a target sequence based on a source sequence. This is done by first encoding the source sequence into a fixed-length vector representation using a multi-layer transformer encoder. The encoder takes in the source sequence and outputs a vector representation that captures the meaning of the entire sequence. Once the source sequence has been encoded, a decoder is used to generate the target sequence. The decoder is also a multi-layer transformer, but it is structured differently from the encoder. At each step of the decoding process, the decoder takes in the current target sequence prefix, the encoded source sequence, and an attention mask that tells the decoder which parts of the encoded source sequence to attend to. Using this information, the decoder generates the next token in the target sequence by computing a probability distribution over all possible tokens and selecting the most likely one. The decoder then adds this token to the target sequence prefix and repeats the process until the entire target sequence has been generated. ### Desiderata for a Generative Story with Logical Forms Suppose that we want to use logical forms to generate sentence in a new “more logical” version of GPT. What then are the possible **********desiderata********** of the overall generative story? We mention two desiderata and their rationales: - the model is required to use the logical form $y_i$ in a non-trivial way in generating the sentence $x_i$ - the logical probability of seeing the form $y_i$ in the given logical context ($y_{n-1}, y_{n-2}, ...$) can be evaluated against a discrete logical knowledge database $\phi$ ### Adding Logical Forms and Syntax to the Generative Story Suppose that we can segment a document into sequences of tokens corresponding to ***********logical sentences, $[x_1, ..., x_n, ..., x_N]$.* We select some logical language $\ell$ and some parse formalism $\rho$. The model parameters are: - real-valued neural network model $\theta$ - governs the syntactic parse - parameters used for neural networks, attention, embedding, etc. - discrete object Bayesian network style probabilistic model $\phi$ - parameters that score the logical sentences Consider the following objects: - $x_n$ is the $n$’th sentence - this is the **********input data********** that we are modeling - $h_n$ is a vector-space encoding of the context up to $x_{n-1}$ - $z_n$ is a *logical* parse for $x_n$ - this is a *********************hypothesized quantity********************* - $z_n$ is a statement in the logical language $\ell$ - $z_n$ can be a conjunction of separate clauses, even optionally encoding context - e.g., **************speaker is Bob************** and **********************Bob says it is raining********************** instead of just ************it is raining************ - $z_n$ can be evaluated against a ********************symbolic logic model******************** $\phi$, and the past logical forms (short term memory) $z_{n-1}, z_{n-2}, ...$. - that is we can define - $p(z_n|x_{n-1}, x_{n-2}, ...) =$ $p(z_n|z_{n-1}, z_{n-2}, ..., \phi)$ - $y_n$ is a syntactic parse - each parse $y_n$ implies a logical form $z_n$ and a sentence $x_n$ - in other words both $z_n$ and $x_n$ are recoverable from $y_n$ - in other words - there is a function $f_z$ such that for every parse $y$, $f_z(y) = z$ for some $z$ - there is a function $f_x$ such that for every parse $y$, $f_x(y) = x$ for some $x$ - this parse maps some logical form $z_n$ to the sentence $x_n$ over a sequence of *parse steps* $y_n=[s_1, ..., s_{len(x_n)}]$ - we assume that the number of steps to parse the sentence $x_n$ is $len(x_n)$, the length in words (or characters) of $x_n$ - that is, every parse for a given sentence has the same length - the length of the parse is always the length of the sentence - this way, we do not have to worry about how to score parses of different lengths - the probability of the parse steps is evaluated the probability of the parse $y_n$ is - $p(y_n|x_{n-1}, x_{n-2}, ...) = \Pi_{i=1}^{len(x_n)}p(s_i | s_{i-1}, s_{i-2}, ..., x_{n-1}, x_{n-2}, ..., \theta)$ - for any sentence $x$ the set $C(x)$ is the set of pairs $(y, z)$ such that $y$ is a parse for $x$, and $z$ is the resulting logical interpretation - the parse $y$ and its associated logical form $z$ are inherently paired because they are inextricably linked, i.e., there is no separate step mapping the parse to its logical form, cf. > The second error [i.e., in Chomskyan linguistics] lies in viewing Surface Structure as a level of representation at all, rather than viewing it (as > > > computational linguists tend to) as no more than a trace of the algorithm that delivers the representation that we are > > really interested in, namely the interpretation. (Steedman, The Syntactic Process, 2000, p. 3) > Then, the modeling of a sentence $x_n$ is given as: - $p(x_n|x_{n-1}, x_{n-2}, ...) = \Sigma_{(y, z)\in C(x)}\ [p(z|z_{n-1}, z_{n-2}, ..., \phi) p(y|x_{n-1}, x_{n-2}, ...)]$ ### The Set **$C(x)$** For the sentence $x$, the set $C(x)$ is the set of pairs $(y, x)$ such that the sentence of $y$ is $x$ and the logical form of $y$ is $z$. To simplify logic, the number of parse steps must either be equal to $len(x)$, or else an even multiple of this. We suppose that the parse is generated according to labeled dependency parsing (McDonald, Pereira, Nivre, Collins, Eisner). The number of labels is up to the theory of language embedded in $\rho$, but assume it is some integral $K_\rho$. The semantics is uniquely determined by the *labeled* dependency parse (Steedman, 2000). ## Parameter Estimation The parameters can be learned using ************************expectation maximization************************ (Dempster, Laird, and Rubin, 1977). That is, we alternate in steps between: - E-step (expectation step) - choosing the most likely distribution over $y_n$ for each $x_n$ - M-step (maximization step) - update the parameters $\theta$ and $\phi$ based on the distribution over $y_n$ ## Conclusion We have proposed the following new opportunity: - new opportunity - ChatGPT shows that world knowledge can be learned directly from data - the question now is if we can use this ability to encode a ********discrete******** knowledge base We have proposed a model that: - introduces a logical parse as part of the generative story - assigns a probability to the ************logical form************ being conveyed using a *****************************discrete logical probabilistic database***************************** The benefits will be: - solve the problem of hallucination - something is either a discrete fact in the database or it is not - enable the maintenance of a “logically consistent worldview”#3,660,676text