Source: Papert, S. (1980). Artificial Intelligence and General Development Mechanisms: The Role of Artificial Intelligence in Psychology. In M. Piattelli-Palmarini (Ed.), Language and Learning: The Debate Between Jean Piaget and Noam Chomsky (pp. 89-106). Harvard University Press.
CHAPTER THREE
Artificial Intelligence and General Developmental Mechanisms
Chomsky has repeatedly stated (see Chapters 1 and 2) that he does not expect to see a “general learning theory” being produced and proving both informationally rich and refutable or confirmable by experiment. Cellérier has argued against Chomsky’s negative heuristic, contending not only that a general learning theory can indeed exist, but that it is already available, at least as a first good approximation. Marvin Minsky and Seymour Papert of the Laboratory of Artificial Intelligence at MIT are credited with this accomplishment.
In this chapter, Papert draws a clear sketch of the principles that have guided his and Minsky’s experiments on computer simulations of higher cognitive functions. The visible result of their work is a device known (after Rosenblatt) as a perceptron, which is capable of formulating simple hypotheses from regular exposure to raw data and then testing these tentative hypotheses against further relevant data under a suitable reinforcement schedule fed back into the machine by the experimenter. The advantage of an artificial rudimentary “learning subject like a perceptron is free access to all its constitutive parts and to all its constitutive relations, that is, the complete description of its initial state ().
Papert challenges the innatist hypothesis put forward by Chomsky and, in the course of the ensuing discussion, by Fodor, presenting the perceptron’s “discovery” of Euler’s theorem. When asked to analyze irregular patterns and to assign numerosity or to evaluate the overall degree of connectivity of such patterns, the machine performs perfectly well by “learning” a rule (Euler’s theorem of the total curvature of simply connected figures). Now this rule was not a built-in property of the perceptron, Papert affirms, and therefore not part of its initial state (). What is innate (that is, what is part of the wiring diagram of the machine) is only a general, nonspecific ability to calculate local primitives (for example, square versus non-square, round versus elongated, and so forth), not Euler’s theorem, nor connectivity or numerosity as predicates ascribable to the whole pattern.
Chomsky and Fodor are charged with committing a fallacy, the “innatist fallacy.” Of such a crime they will plead innocent again and again, in spite of its being charged by Cellérier, Papert, Toulmin, and Putnam. The fallacy, to put it in simplified terms, is to ascribe to the brain what is only an abstract property of its operations. As Putnam says (see Chapter 17), if the geography of the White Mountains can be said to be in my mind, and therefore somehow represented in my brain, that does not imply that the geography of the White Mountains is itself a description of my brain. And in Mehler’s terms, if eye color is innate, that does not imply that the genes are themselves colored. In the following chapters, less provocative and more subtle formulations of the alleged innatist fallacy will be presented and discussed. The contours of agreements and disagreements on general learning theory and innate cognitive structures will become clearer and increasingly more sophisticated.
The Role of Artificial Intelligence in Psychology
By Seymour Papert
I will now try to show some ways in which artificial intelligence (AI) is able to elucidate a debate in theoretical psychology such as the present one between Chomsky and Piaget. As a first step I want to guard against common misunderstandings of the role of AI in relation to traditional psychology.
The most important of these misunderstandings sees AI as methodologically similar to behaviorism, in that both are sometimes presented as seeking greater rigor by limiting the domain of psychological investigation: behaviorism by excluding reference to mentalism, AI by refusing theoretical models that cannot be simulated on a computer. This view of AI as a limitative theory is completely false, as is the associated belief that AI imposes “simple” models on complex phenomena. Quite on the contrary, AI seeks to be expansive by adding new and powerful kinds of mechanisms to the theoretical repertory of the psychologist, and the models it constructs are very much more complex than those of traditional psychologies. The perception of the models as “simple” can often be traced back to ignorance of computer languages, which leads people to confuse English-language descriptions of AI theories with the theories themselves.
AI has thrusts in three directions which relate to controversies in psychology and linguistics about the property of innateness. Two of these are easy to state: The first tends to reduce the set of structures considered to be innate by showing how they could be acquired through the operation of more powerful developmental mechanisms. The second thrust goes in the opposite direction; paradoxically, by understanding very powerful developmental mechanisms one is able to see how certain structures totally unsuspected by traditional psychology could, if they were innately present, play the role of seeds for the growth of mental functioning.
These two thrusts find a direct translation in terms of the present debate. I believe that Chomsky is biased toward perceiving certain syntactic structures as “unlearnable” because his underlying paradigm of the process of learning is too simple, too restricted. If the only learning processes were those he seems to recognize, these syntactic structures might indeed have to be innate! Allowing a richer set of learning methods certainly does not in itself prove that the principles in question are, in fact, “learned”; proving the existence of such learning would at the very least require a rather complex empirical investigation. But I will argue that a stronger learning theory shifts the balance of plausibility toward conjecturing that syntactic structures are “learned,” and so toward adopting a research strategy that looks for how this could happen rather than following a strategy of assuming “innateness.”
As an integral part of my argument, I will give some hints (the most I can do in a short exposition) of kinds of structures that computational thinking leads us to recognize as being almost certainly necessary to the operation of so complex a computer as a brain. These structures are not specifically linguistic and so bring us in contact with another, perhaps more fundamental aspect of the Chomsky-Piaget debate: the extent to which the operation of language shares common structures with the rest of the intellectual system. Piaget finds this sharing to be extensive, in contrast to Chomsky’s “organicist” tendency to see mental functions as more sharply distinct, organized into organs of the mind (to use his analogy) much as the heart, the liver, and so on are distinct organs of the body. The contribution of AI to this area is to show how computational primitives are fundamentally important to all the mental functions: by seeing all the “mental organs” as computational processes, we are inclined to see them as less fundamentally different from one another than are hearts and livers.
The highly metaphorical flavor of these last remarks leads to the third, and more subtle, thrust of AI: to prune psychological language of prescientific metaphors by developing a more precise terminology, conceptual framework, and, indeed, a new set of “metaphors” in the form of well-understood situations against which general ideas can be tested for intellectual coherence. To understand this, consider the reactions one might have to Chomsky’s assertion that such things as the “SSC” (specified subject condition), the “abstract notion of ‘subset,'” or “bound anaphor” are properties of the genetically determined initial state of the newborn. One could argue about the assertion’s truth or plausibility; but it might be more fruitful to ask how metaphorical these assertions really are: what does it mean to say that the notion of “subject” or “anaphor” is part of the infant’s initial state?
Easy answers lead to conceptual difficulties. Is it being said, quite literally, that these notions in the form we know them in adults are in some sense (what sense?) present in the newborn? Or is it being said that some precursor of the notion is present in the newborn—something of an unspecified kind that will grow as a seed into, for example, the notion of “bound anaphor”? The literal choice is certainly hard to swallow without a theoretical framework to explicate (among other puzzles) what could be meant by grammar in a prelinguistic subject. On the other hand, the “seed” version is in itself a very weak, almost trivial, statement with which no one would quarrel and which certainly does not justify Chomsky’s strident criticisms of Piaget, AI, and the like.
The same dilemma appears in more sophisticated informational reformulations of the question as a kind of determinism: can one deduce from the initial state of the baby the fact that his adult language will exhibit the SSC? If one means this in the strict sense that allows us to deduce the property “blue eyes” from the presence of certain genes, it becomes very hard to believe. If, on the other hand, we weaken it to make it credible, it becomes difficult to know exactly what is being asserted. (This line of argument is certainly not alien to Chomsky’s way of thinking; he used it most persuasively in his critique of Skinner on language.)
The first of some more technical comments, to which I now turn, is intended mainly to illustrate more sharply the need for greater clarity about what it means to be innate. I will do this by describing an automaton, a machine that we understand quite thoroughly, and asking questions about what is and what is not innate in the machine. If the question is unclear even in this “toy” situation, how much more clarification does it need in the complex situation of human development?
The machine in question is called a perceptron. Its structure is quite simple: It has a retina on which pictures can be projected, and the purpose of the machine is to recognize whether or not this image has a certain property called “the predicate” (for example, is a square). It has a large number of submechanisms, each of which can compute the answer, to be expressed as “yes” or “no,” to any well-defined question about a tiny region of the retina. Collectively these submechanisms (called the “local function”) cover the whole retina, but none of them has any global knowledge. In particular, none of them can see the whole figure.
There is also a central organ, which has access to the answers given by the local mechanisms; but this organ is constrained to a particular simple algorithm to generate the global decision (for example, is it a square?) from the local functions, namely, the “linear threshold decision function”: the binary yes/no outputs of the local functions are represented as 1 and 0, respectively; the machine forms a weighted sum of these numbers using weighting coefficients characteristic of the particular perceptron; it then makes its decision according to whether the sum comes out to be more than or less than a certain quantity called the “threshold.”
Finally, the perceptron is equipped with a “learning mechanism,” which works like this: when the machine says “that’s a square” it will be told whether it is right or wrong, and if wrong, will use this feedback to alter its weighting coefficients (in a manner whose details are not relevant here).
What can a perceptron learn? The answer is not always immediately obvious, either from an examination of its “innate structure” or from simple experiments. It is easy enough to see that a perceptron can learn to distinguish between dichotomies such as a square versus triangular; in this case the local functions recognize the presence or absence of at least one angle that is not a right angle, and so the hypothesis square can be eliminated on local grounds. But there are cases in which it is much harder to see whether the global decision is reducible to such local ones. The classic example is the predicate is connected. The picture on the retina is assumed to be a black figure on a white ground, and the question for the perceptron is whether the black figure is made of one or several pieces. Is such a decision reducible to local observations? Intuition still says that a perceptron should not be able to tell whether there is only one blob or several. But a deep mathematical theorem by Euler can be adapted to show that the perceptron can learn any predicate like the number of blobs is less than .
Let us now imagine an investigator who does not know the Euler theorem and happens to be concerned with whether blob-numerosity (in the sense of the predicate just mentioned) is innate in the perceptron. One can easily imagine such a person being very puzzled when shown the wiring diagram of the perceptron. He would see nothing there which (to his mind) even remotely resembles numerosity. He might conclude that something must be missing from our wiring diagram (so one of the cautioning morals of the story is that one has to be very careful about conclusions like this). But the more important, if more subtle, conclusion is that even with full knowledge (of the wiring diagram and of the mathematics), it is not at all clear whether one ought to say that numerosity is innate. In some senses of “innate” it is, and in other senses it is not. The conclusion for me is that we need a much more carefully elaborated theoretical framework within which to formulate the real questions that lie behind formulations such as, “Does the subject have the notion
?” or “Is
a property of the initial state of this subject?” or “Is
innate?”
One of the purposes of such a theoretical framework would be to replace notions such as “notion” by something more technical. It is quite sobering to take stock of how much of the language used today in serious discussions of the mind is not very different from that of Aristotle. Of course, Chomsky himself has contributed enormously to changing this state of affairs within linguistics; but this makes it all the more paradoxical that his psychological metalanguage is so pretechnical. Let me develop the point by contrasting the way Chomsky and Piaget formulate a question about the development of what is popularly called a “notion.”
A naive perception of Piaget sees him as saying: nothing is innate, everything emerges from development. Of course, this is absurd if pushed undialectically to the limit, but what he does teach us is this: if you make a list of structures and notions and rules (or whatever you call them) found in adult intelligence, and if you ask which of them is innate, the answer will be none. The point behind the apparent contradiction is simple enough: everything has a developmental history through which it emerges from other, very different things. Whatever it is that is innate, we can at least be sure that it is not (and probably does not even resemble) any discrete part of the adult mind. The principle might seem quite obvious; but whereas it is exemplified in a deep and subtle way in Piaget’s total work, it is violated by the very form of Chomsky’s suggestion that developed entities such as “the specified subject condition (SSC)” or “the notion of bound anaphor” (these are his actual words) might be “properties” of the initial state. In fact one can see a large part of Piaget’s work as the search for intermediate entities, which can play the role of precursors of the structures we find in the adult, or even the child of any particular age. Thus the question “Is the notion of ‘number’ innate?” is displaced by seeing how it grows out of various precursors whose existence was not recognized by Aristotle or any other pre-Piagetian psychologists.
The technical spirit of Piaget’s work is reflected in the fact that the abstract noun number occurs much more prominently in the title of his book about it than in the text: the pretechnical concept “number” now serves only as an indicator of a direction for research whose actual dimension cannot be specified by the pretechnical intuition of “number.” I argue (and this has some irony in the context of the present discussion) that the power of Piaget’s contribution would scarcely be altered if the intermediate objects he has discovered are proved to be “innate”; the hard and deep work was discovering the precursors.
When we transpose the observation to the problems raised by Chomsky, the observation turns into the suggestion that the hard work here will be discovering the precursors out of which the SSC and similar structures emerge. And then we will have to understand the precursors of those precursors and, by a longer or shorter chain of genesis, eventually arrive at the properties of the initial state. This is what developmental studies are about, it is a long, arduous, and technical path. It seems to me very remarkable that Chomsky should, by contrast, take a position that essentially declares: I, Chomsky, cannot see how the SSC can be learned, so all reasonable men should conclude that it is innate.
But what leads Chomsky to this position? If I claim that he is in an absurd position, the onus is clearly on me to explain this mystery. So I turn next to sketch my theory of Chomsky. The theory pivots around a claim that Chomsky, despite his well-known criticism of Skinner and of behaviorism in general, has remained wedded to a behavioristic position in one crucial area, namely in his model of learning. I will develop this point by sketching a taxonomy of developmental theories in which Skinner, Chomsky, Piaget, and Simon will be explicitly located. As a first step I recall, with slight modification, Chomsky’s notation. The baby is born in state and eventually comes to state
.
and
are deep structures and thus not directly visible. It is consequently possible, in principle, for different observers to argue that behaviorists typically underestimate the complexity of
. One can classify developmental theories according to how much of the complexity of the human mind they attribute to
and to
, and to the component of mental function responsible for development. For example, Skinner believes that it is possible to explain the passage from
to
by a very simple general developmental mechanism (GDM). Both Chomsky and Piaget have used forceful arguments to persuade us that
is certainly more complex than Skinner believes, and probably too complex for a simple GDM to guide its growth from the kind of
Skinner would admit. The difficulty facing Skinner is a mismatch between his simple
and the putatively complex
.
There are in principle three (nonexclusive) reactions to this apparent mismatch: One can postulate that is more complex than Skinner thought, which is the route taken by Chomsky; one can postulate, as Piaget does, that the GDM must be more powerful than Skinner believed; finally, one could deny the mismatch, for example by arguing as Herbert Simon does that
really is structurally simple after all and thus that
and the GDM can both be simple as well (although, of course, Simon’s GDM and
are radically different from Skinner’s).
I will devote the rest of my discussion to developing the idea that Chomsky seems to be tacitly committed to a very Skinnerian kind of GDM and illustrating some ways in which Piaget’s developmental mechanisms and also the emerging viewpoint of AI can be understood as more powerful mathematical and epistemological principles rather than more powerful biological mechanisms. In this regard, it is appropriate to look at Chomsky’s exact words (see Chapter 1):
Again, it cannot be imagined that the language learner is taught these facts or the relevant principles. No one ever makes mistakes to be corrected. As in the case of the structure-dependent principle, passive observation of a person’s total performance might not enable us to determine whether the principles are in fact being observed (just as experience would not suffice, normally, to provide this information to the language learner), though “experiment” will quickly reveal that this is so. The only rational conclusion is that the SSC and the relevant abstract notion of “subject” and “bound anaphor” are properties of , that is, part of
.
The point of the quotation is to draw attention to the underlying model of the learning process, which has the following characteristics of typically behaviorist models: the emphasis on teaching, either explicitly or by correction of mistakes; the apparent rejection of “experiment” as applicable to the child’s learning of language; and the assumption that what is being learned is the thing itself (the adult structure) rather than some developmentally deep structural precursor. Since the point is central, let me put it differently. Of course the child does not learn the SSC by being taught in any direct sense; this is not the way any fundamental structures or skills are ever acquired. And of course the child does not learn the SSC by passive observation. Piaget has removed any lingering tendencies people might have had to see the child as ever engaging in “passive observation.” So I would have thought that the “only rational conclusion” from Chomsky’s set of statements is that the child discovered the SSC through “experiments.” If asked where the ability to do experiments came from, I would be a little more inclined to grant that this is a property of . We certainly know that recognizable precursors of components of the complex activities involved in experimenting are visible in very young babies.
But of course, it is not sufficient to say that the child learns by doing experiments. We need to understand the conceptual framework within which such experiments can be conceived and interpreted. So again and again, we come back to the central problem of discovering the intermediate intellectual structures.
Piaget’s contribution to this problem is much deeper than any particular structural analysis he makes of any particular aspect or stage of intellectual ability. One can disagree with his specific analyses of stages or aspects (indeed, he himself often does) but still appreciate the very original paradigm he has given to developmental psychology. To emphasize an analogy with Chomsky, we can describe this as the concept of deep structure in the developmental domain. I myself like Piaget’s phrase: “There is no genesis without structures, there are no structures without genesis” (pas de genèse sans structures, pas de structures sans genèse). You cannot understand the genesis of number by trying to trace surface forms such as actual uses of numbers themselves. One has to look for a system of deeper structures, of which Piaget’s “grouping” (groupement) should be seen for purposes of the present discussion as an example to illustrate in a very general way what kind of entity such a structure might be. But one will not find intelligible structures without focusing from the outset on the problem of genesis.
I will make this idea more concrete by stepping back and looking at a formulation of what I think should be the key problem of learning theory: to reduce the sense of miracle induced by the power of the mind both to think and to grow. Euler’s theorem did this for the hypothetical student of perceptrons in my little computational parable. I would like now to sketch a real example, for which credit goes to Piaget and to Bourbaki. We will see at the same time how epistemological and mathematical insights meld with psychological/developmental thinking.
I am sure that the Bourbaki group did not think of their theory of structures as a contribution to the theory of learning, nor do textbooks of psychology include Bourbaki in their list of learning theorists. Yet if the task of the theory of learning is to reduce the apparent mystery of the acts of learning, Bourbaki made a significant contribution by developing a perception of mathematics which happens to make it appear much more learnable than it did before. This is not mere appearance: certain insights into the nature of number reveal aspects of its structure that play a significant role in all processes of learning. One (out of many) of these could be called the principle of factoring: structures that can be seen as being composed of simpler but still significant parts (substructures) tend to be more learnable. One could order foundational theories of arithmetic along a learnability dimension: number “à la Peano” or recursive function theory stands at the unlearnable end; number “à la Russell/Whitehead” is more learnable, number “à la Bourbaki/Piaget” is very much more learnable, and AI extensions of this last perception go still further.
One can argue about whether Piaget’s theory of the antecedent structures of informal mathematical thinking is exactly the same as Bourbaki’s theory of formal mathematics, but the significant circumstance is already there in the fact that they are close enough for such questions to be asked. And for our present purpose, the point of all this discussion is to illustrate a kind of explication of learning that falls outside the behaviorist GDM and seems not to be considered by Chomsky. Of course the example is from mathematics rather than linguistics, but the question is whether similar structural analyses will reduce the apparent degree of unlearnability of properties of language.
Grounds for optimism are to be found in the recent tendency to intermarriage between AI and styles of linguistic theory related to systemic and case grammars, and to a much tighter interconnection between syntactic structures and cognitive (some would say “semantic”) structures. Since I can discuss only one example, I choose in the interests of simplicity the very general one that Chomsky calls “the structure-dependent property of linguistic rules.” One possible answer to the question of how the child learns so easily to develop structure-dependent rules of language is that he has already come to do this outside of language; another possible answer is that the nature of computational processes dictates that rules will be structure-dependent. Now I grant that either of these conjectures is compatible with some sense of innateness of structure dependence. But it would not be a linguistic rule that is innate but some more general precursor, and this in itself is a significant shift away from Chomsky’s position toward a more Piagetian one.
At last we come to the final link in my chain: the emergence within AI of a structure dependence general enough to induce language, visual perception, common sense reasoning, and, indeed, all information processing in men and in machines. A technical account of such a theory is impossible to give here and, indeed, has not been described in detail as a topic in itself. But it is a pervasive theme of research in AI and readers can find more detail in other sources.
Discussion
Chomsky: Papert claims that I don’t see any way of explaining the final state in terms of a behaviorist type of general developmental mechanism, and therefore I conclude it is innate. What I actually say is that I don’t see any way of explaining the resulting final state in terms of any proposed general developmental mechanism that has been suggested by artificial intelligence, sensorimotor mechanisms, or anything else, and that remains precisely true. At that point Papert objects, quite incorrectly, that somehow the burden of proof is on me to go beyond saying that I see no way of doing something. Of course, that’s not the case: if someone wants to show that some general developmental mechanisms exist (I don’t know of any), then he has to do exactly what Rosenblatt (who invented the perceptron) did, namely, he has to propose the mechanism; if you think that sensorimotor mechanisms are going to play a role, if you can give a precise characterization of them, I will be delighted to study their mathematical properties and see whether the scope of those mechanisms includes the case in question. Similarly, you are mistaken if you think that the mechanisms here relate in some way to the acquisition of structure-dependent rules—they don’t at all. The fact that you can describe or even discover hierarchical organization tells you nothing about whether the rule in question should observe a hierarchical structure in this case or should observe the property leftmost in this case. If anyone thinks there is such a general developmental mechanism, fine, propose it, make it explicit, then I and others will investigate it to see whether there is any relation whatsoever, be it direct or metaphorical, to the concrete problem of attaining the final state. Since that has not been done in artificial intelligence or in studies of sensorimotor intelligence anywhere to my knowledge, I can’t do what I did, for example, twenty years ago when the theory of finite-state Markov sources was proposed. This is an example similar to that of the perceptron: a specific theory was proposed about the nature of grammatical structures, and it was then possible to investigate its mathematical properties and to demonstrate that no conceivable realization of that structure, no matter how much time you allowed, could in fact have certain properties that languages have. So once somebody had proposed the theory of finite-state Markov sources as a generalized superdevelopment of the most complicated behaviorist theory of chaining that had ever been invented—in fact a theory even including Lashley-style rhythmic properties—then it was possible to study that theory and in this case to demonstrate that the theory, and correspondingly any of its subtheories, were demonstrably false. Now, when theories are not proposed, of course all you can say is: “I don’t see any conceivable way in which these ideas relate to a given consequence”; but I don’t accept that the burden goes beyond—that the burden there has to fall on the person who believes that the principles really exist, and if somebody can present, for example, a theory of sensorimotor constructions in a form explicit enough to investigate the possibility of obtaining the “specified subject condition” in those terms, I would be delighted to investigate it similarly in this case. As for the general theory of structure, that is exactly what generative grammar has been concerned with for twenty-five years: the whole complicated array of structures beginning, let’s say, with finite-state automata, various types of context-free or context-sensitive grammars, and various subdivisions of these theories of transformational grammar—these are all theories of proliferating systems of structures devised for the problem of trying to locate this particular structure, language, in that system. So there can’t be any controversy about the legitimacy of that attempt; in fact, that is what all the work in formal linguistics has been about for a certain number of years.
Papert: What I am trying to do is to situate Chomsky and Piaget in terms of a wider context. It could be put in this way: although it is true that transformational grammar looks for a certain kind of structure, it looks toward linguistic examples as its main source for formulating the kinds of structures that it wants. As opposed to this there are others, including specialists in artificial intelligence and Piaget, who look at the structure of mathematics itself for examples from which you might draw a general theory of structures. This is a trend that is going on now, and if one wants to make sense of these debates and different points of view, one has to have a wider perception of the goal that all these researchers are pursuing.
Chomsky: The theory of structures developed in generative grammar was within a particular mathematical theory, namely, recursive function theory—that is, the theory of subrecursive hierarchies is precisely what offered the framework for investigating the different types of linguistic structures. Now, maybe we shouldn’t have looked at the theory of subrecursive hierarchies but at some other mathematical theory; but the idea of looking at mathematical theories for ideas about the class of structures, that is relevant…
Papert: No, that is not the point: the idea is not to look to mathematics for mathematical theories of structure, it is to look at the structure of mathematics in the same way that you look at the structure of language, as the object of the theory of structures. For example, it is very important in Piaget’s way of thinking, and for the plausibility of Piaget’s explanation of the development of mathematical thinking in children, that something like the deep discoveries made by the Bourbaki group be true. Bourbaki, by revealing certain things about the structure of mathematics that were not quite clear before, altered the apparent degree of complexity involved in learning mathematics. Since then we have had a quite different perception of the evolution of mathematics, either historically or in the individual.
Chomsky: I think all this is really, as Toulmin said previously, a red herring. There is no such thing as the structure of mathematics: you can find within mathematics all sorts of different types of structures which you can study historically, non-historically, or whatever. But the theory of subrecursive hierarchies barely existed as a structure of mathematics twenty-five years ago, and therefore you couldn’t look for detailed suggestions there as to the relevant types of structures; it provided at most a general framework. As for what Bourbaki investigated, algebraic structures or topological structures, fine, if those are relevant to this issue, let’s see how; let’s see how the investigation of topological structures, for example, relates to the problem of the existence of structure-dependent rules. But I will repeat the statement that Papert does not like: that I don’t see any possible conceivable connection between them; if you think that that statement is wrong, then show me even the most remote connection between them.
Papert: The connection between Bourbaki and structure-dependent rules in language cannot be immediate. But since you deny any conceivable, even remote, connection between them, let me repeat one of the connections I have stressed. Arithmetic as axiomatized by Peano would seem quite as unlearnable as any rule of language as axiomatized by Chomsky. Along comes Bourbaki/Piaget, and the situation changes in arithmetic. Peano was not wrong as a formal mathematician; he was wrong as a developmentalist. Perhaps the linguistic correctness of your theories of language does not alter the possibility that you are wrong as a developmentalist.
Chomsky: That was not my question. My question has to do with the point that you originally raised, namely, that somehow the kind of investigation conducted by Bourbaki leads to a theory of structures that relates to the question of why there are structure-dependent rules in grammar, and I say there is no conceivable connection there, although I think there is some connection with another part of mathematics, for example, the theory of subrecursive hierarchies. Now, I’m not wedded to that theory: if you find some connection of the theory of topological spaces, for instance, to the problem of structure-dependent rules, fine, but it’s pointless just to say that there might be such a connection.
Atran: It seems to me that this interest in the Bourbakian structures follows from some very peculiar assumptions and leads to some rather dubious conclusions. According to Piaget’s doctrine that he calls “genetic epistemology,” and I assume this is what Papert is really referring to, each advance made toward a foundation theory in mathematics is, by that very fact, a more comprehensive theory of cognition. Moreover, if a mathematical model is developed which is not meant to be a foundation theory it is still, by definition, part of the structural movement. This must be the case if, as Piaget claims, “logico-mathematical knowledge does not succeed in detaching us from reality or the world of objects,” and if every logico-mathematical deduction represents a “theory of the physical phenomena after the event.”
Now it is quite possible that the Bourbaki effort to analyze mathematically significant notions in terms of certain historical precedents or of structures that are basically heuristic provides some account of the thinking involved in the construction of certain mathematical systems. If so, this represents a contribution to that part of the mind concerned with the construction and interpretation of mathematical systems. Biologists would then be justified in beginning a search for the neurophysiological bases of these systems in a region of the brain specified by the complex genetic program of human beings.
For Piaget, however, a Bourbakian structure like the group is today the foundation of algebra; it is, according to Piaget, already being used in an important way in physics, and very likely the day will come when it acquires a central role in biology as well. Such “facts” as these thus seem to suggest to Piaget that the mother structures of the Bourbaki group correspond to coordinations that are necessary to all intellectual activity. But even though groups were the first general algebraic structures to be investigated, they are no more basic or essential to algebra than other general structures. Each kind of algebraic structure captures some interesting set of uniformities characteristic of a related collection of mathematical systems; however, there may be as many “basic” or general structures as there are interesting uniformities among any of the various systems that already exist or that may be created in the future.
Perhaps certain groups, with specific properties, will be useful in describing certain areas of physics, biology, or cognition. Most of the significant theories of science, however, do not depend in any nontrivial way on groups or on any of the other Bourbakian structures. Consequently, Piaget attempts to increase the importance of the group by providing it with a significance that lies outside the concerns of algebra. Using metaphors drawn from the equilibrium theory of gases and the regulatory and structural function of genes, he proposes that the structural coherence of the group is maintained by autoregulation; mathematical notions of closure, identity, and transformation are equated with self-maintenance, reversibility of processes, and formation. For the working mathematician, this definition of the group would appear quite arbitrary and can be completely ignored. Such ad hoc efforts to generalize precise notions by analogy can only lead to trivialities.
It must be admitted, nevertheless, that the general ad hoc approach characteristic of genetic epistemology may occasionally lead to interesting results. Because structuralism sets itself the task of constructing formal languages as an empirical program, there is a constant effort to try out the whole range of mathematical structures on various domains of cognition. This process could conceivably lead to an interesting explanation of phenomena—this may be the case, for instance, in the apparent success of a grouplike structure, called “groupement,” in accounting for displacements in a perceptual centration. Usually the attempt to extend important results to other domains, such as language or mathematics itself, only serves to obscure the significance of the original findings. The fact that some success is possible with such an approach, however, means that a refutation of the premises of structuralism does not necessarily hold for particular consequences. But the reasoning that those premises entail often makes it difficult to separate the significant from the trivial.
In Seymour Papert’s presentation we have read some statements that deserve to be singled out, because much of the following discussion will revolve around them: “If you make a list of structures and notions and rules (or whatever you call them) found in adult intelligence, and if you ask which of them is innate, the answer will be none… The hard and deep work was discovering the precursors. When we transpose the observation to the problems raised by Chomsky, the observation turns into the suggestion that the hard work here will be discovering the precursors out of which [linguistic] structures emerge. And then we will have to understand the precursors of those precursors and, by a longer or shorter chain of genesis, eventually arrive at the properties of the initial state. This is what developmental studies are about.”
This is the core of the compromise between Chomsky’s innatism and Piaget’s constructivism as forwarded by its proponents. Let’s first adopt an innatist view with regard to nonspecific, undifferentiated, multipurpose primitives or precursors; then let’s embrace a strictly constructivist view to account for the more complex cognitive or linguistic structures arising out of these precursors in the course of development. We will see that such a “division of labor,” as Cellérier has called it, between innatism and constructivism has appealed to many, including Monod (as witnessed by his reply to Fodor in Chapter 6). Now, granted this, options can further diverge as to the innateness of the developmental pathways themselves. Piaget rejects the hypothesis that the sequence of stages representing successive integration-plus-differentiation of primitives is itself the product of some genetically determined program, whereas Monod explicitly endorses such a view.
Chomsky and Fodor are more trenchant, however, and their disagreement blocks the argument at a much earlier state; they maintain that cognitive development is not an “enriching” process at all, but rather consists of a progressive specialization, channeled by the environment. In other words, the task of any developmental theory in the field of cognition and language, as seen by Chomsky and Fodor, is not to account for a stepwise construction of more powerful and specific structures out of raw primitives (be it through trial and error or hill climbing) but rather to account for the organism’s inborn predisposition to select quickly and without mistake a specific working hypothesis (structure dependence, SSC, understood subject equivalence, systematic ambiguity, and so forth). Learning is, to them, plugging in the right device at the right time, using the structure of the incoming information only as a test for highly specific filters that are already built in.
Chomsky explains in the next chapter why such an “extreme” innatism appears to him plausible—indeed, the most plausible hypothesis that is compatible with a wealth of linguistic, psychological and neurophysiological data. Fodor will present his own analogous point of view in Chapter 6.