Initial States and Steady States

Source: Chomsky, N. (1980). Initial States and Steady States: The Linguistic Approach. In M. Piattelli-Palmarini (Ed.), Language and Learning: The Debate Between Jean Piaget and Noam Chomsky (pp. 107-130). Harvard University Press. 

CHAPTER FOUR
Initial States and Steady States

The present chapter consists basically of Chomsky’s plea for “the Cartesian mind.” His presupposition is that cognitive and linguistic structures are in principle “explicable,” though as yet “unexplained,” in terms of the expression of a genetic program universal to the human species. His discussion in this chapter is so clear and exhaustive that any external comment would be, at best, redundant. However, Chomsky does mention a most revealing experiment in neurophysiology that is pivotal to his arguments and that may be unknown to the general reader. I assume that a brief account of Hubel’s and Wiesel’s findings may be welcome to those who, unlike most participants in the meeting, are not familiar with it. The reasons why these findings are apt to support Chomsky’s model of the mind and why they reinforce the notion of learning as “selective stabilization of functioning synapses” (adopting the felicitous expression of Changeux and Danchin)* also need to be made explicit.

Towards the middle of the sixties, neurophysiologists such as Jerry Lettvin, Humberto Maturana, David Hubel, and Thorstein Wiesel developed an ingenious technique for determining how single nerve cells in an animal’s visual system respond to specific patterns (horizontal lines, vertical lines,

*See J.-P. Changeux, P. Courrège, and A. Danchin, in Proceedings of the National Academy of Science U.S.A. 70:2974-2978, 1973; and J.-P. Changeux and A. Danchin, “Apprendre par stabilisation sélective de synapses en cours de développement,” in L’Unité de l’homme, ed. E. Morin and M. Piattelli-Palmarini (Paris: Editions de Seuil, 1974).

sharp angles, moving spots) in the visual field. The interesting result was that even within a few hours after birth, particular neurons were specifically preset to react (that is, to send a train of electrical impulses) only to a well-specified class of visual stimuli. For example, neurons that are innately preset to react when a horizontal stripe is presented before a newborn kitten will remain silent when the same animal is exposed to vertical lines, and vice versa. Even more interestingly, Hubel and Wiesel were the first to demonstrate that a given class of specialized neurons can become totally inert if the kitten is brought up, from birth onward, in a “deprived” optical environment (for example, a tall cylinder painted with only vertical or only horizontal stripes and nothing else). Only the class of neurons corresponding to the pattern effectively shown by that environment will remain active, whereas all other optical neurons will become inactive because the corresponding synapses will degenerate. Such loss of visual competence can be irreversible. These and numerous other experiments of the same kind are directly relevant to the central topic of this debate because they have demonstrated, or at least made plausible, the following inferences: (1) highly specific perceptual filters are already present at birth (evidence for innatism); (2) these filters shape the visual world into orderly geometrical figures (the Cartesian mind); (3) the geometrical figures without (structures of the external world) elicit specific responses within (orderly patterns of neuronal activity) but do not determine the form of the response (evidence against “assimilation” and in favor of a simple triggering of preset mechanisms).

Therefore: (1) experience of the surrounding environment results in a selection of “relevant” stimuli according to discriminative criteria which “are there” beforehand, that is, which are innate (learning as a selective process); (2) the actualization of a given cognitive structure is made at the expense of other competing possible structures which become irretrievably lost in the course of development (fixation of synapses by actual functioning and degeneration of the inactive ones). This last point has been widely developed by Changeux and Danchin in their theory of selective stabilization of functioning synapses and by Jacques Mehler in his theory of “knowing by unlearning.”* On the basis of these premises, however

* J. Mehler, “Connaître par désapprentissage,” in L’Unité de l’homme.

sketchy and imprecise, the reader may now be in a better position to grasp Chomsky’s arguments and their far-reaching consequences.

The Linguistic Approach
Noam Chomsky

I will just refer to my paper (see Chapter 1) and bring out some of the general points I was trying to make, leaving the details to the paper. The general approach I’m taking seems to me straightforward and unsophisticated, but nevertheless correct.

Let’s start by noting that we belong to a certain species, and we can assume uniformity, idealizing away from variation. This species has what we can call an initial state, that is, a state prior to experience, fixed for the species and called So. We discover through investigation that in particular cognitive domains, for example in the domain of language but not only there, the individual goes through a series of states and reaches what is in effect a steady state, which is a state that doesn’t change very much except in marginal respects. In the case of language, it seems that the steady state is invariably attained about the time of puberty. We can then ask ourselves what is the nature of the steady state attained and what must have been the character of the initial state for that steady state to be attained, given the nature of the existing experience. One needn’t be interested in this question, but I am interested in it. The most interesting question from this point of view is what is the nature of the initial state, that is, what does human nature consist of in this respect? Evidently experience is required to attain the steady state, so we can think then of the initial state as being in effect a function that maps experience onto the steady state. We can express the idea in this way without losing any generality, and this formulation permits any degree of interaction as well as complexity, new devices at later stages, and so forth.

This function, which maps experience onto the steady state, the function that in effect characterizes the initial state, can be quite properly thought of as a learning theory, that is, this function is simply the learning theory for human language. As I mentioned in my paper, I don’t see any particular reason to believe that there exists such a thing as a learning theory; it seems to me a very odd idea that such a thing should exist, rather as if there existed something like a growth theory for organs. Of course there is a level at which a growth theory does exist, namely, cellular biology, but I don’t see why there should be one above that level.

Similarly, if we study the development of cognitive structures I think that we will discover separate learning theories; no doubt they will have some properties in common, but I assume that they will be quite specific to particular cognitive domains and particular species, or at least if this is not the case it will be surprising and it will be necessary to demonstrate it. The common assumption to the contrary, that is, that a general learning theory does exist, seems to me dubious, unargued, and without any empirical support or plausibility at the moment.

Let’s look at this particular question and ask what is the nature of the learning theory for human language, that is, what is the acquisition function that maps experience onto the final state, or to put it differently, what is the nature of the initial state So (these are all different ways of formulating the same question).

One can make a variety of different proposals on this point, and through the history of the last several hundred years one can find a number of general attempts to approach the question I put in these terms. For example, one might argue that there are general developmental mechanisms (GDMs) which constitute in effect the initial state and which therefore constitute the function that maps experience onto the state attained. If I understand the ideas put forth about constructions of sensorimotor intelligence, they essentially fall into this category; similarly, Papert’s remarks about GDMs fall into this category as well (see Chapter 3). These I take to be proposals about the properties of the initial state, namely, that the initial state is characterized by a system of GDMs which will give rise to the final state, given experience. Now, if someone presents such a system it can be investigated; if someone proposes a GDM in a specific enough form so that it is possible to draw out consequences from the proposal, then we can investigate the adequacy of such a system. Prior to the presentation of any such GDM, I am afraid I have to resort to the comment that Papert dislikes so much, namely, that in the absence of any concrete proposals that can be investigated, I see no plausibility to the suggestion.

There are people who have actually made such proposals. For example, Patrick Suppes at Stanford has presented a very explicit and interesting proposal based on the Suppes-Estes statistical sampling theory, which gives a general developmental model for attaining a final state on the basis of experience.1 This is a proposal that one can investigate, and that is precisely what Suppes did: it was clear and explicit enough to investigate, and he showed that this system could attain in the limit a finite-state Markovian system which produces symbols from left to right. That is interesting because we know that such a system is inadequate for language, that is, we know that the systems that can in principle be attained within this model are not the systems that are attained in fact by humans; therefore we have an explicit refutation of this theory of developmental mechanisms. To say that we refute the theory is to make a positive comment about it, that is, this theory was presented in a clear enough way so that it was possible to determine whether or not it is correct, or at least on the verge of being correct. It is a merit of a theory to be proved false. Proposals that do not allow such a determination, or the determination of whether or not evidence bears on them, do not have that merit. The proposal that I just mentioned is the most sophisticated suggestion that I know of regarding general developmental mechanisms.

If you go back in history there are other proposals. For example, Hume had some specific proposals which are much denigrated these days, although I don’t know why—I think they are good because they are concrete. He listed the mechanisms that he felt belonged to what I am calling the initial state: contiguity, resemblance, perceived cause and effect, and he made an explicit claim that the systems of knowledge that we have can be attained on the basis of these general developmental mechanisms. This is a moderately explicit theory it was developed further in subsequent associationist theories, and it can be submitted to investigation, but it is so evidently false that there is no point in pursuing it. But at least it was concrete enough to investigate.

Descriptive linguists of the 1940s made very explicit proposals about what one could think of as developmental mechanisms (which they didn’t regard as such), namely, the so-called discovery procedures for grammars, which became rather sophisticated. One could investigate these proposals, and one could prove again that they could not in principle lead to the steady state attained and that is interesting.

If you want my judgment about this whole issue, of course you cannot demonstrate that general developmental mechanisms do not exist—that’s impossible-but I think that the efforts that have been undertaken have in fact proved abortive, and I don’t think there is a particular reason to believe that this is a plausible way to proceed.

There is historically another general approach to this whole question, which in the modern period is traceable back to Descartes. He makes the following kind of observation: Suppose I draw a triangle on the blackboard; Descartes says that anybody looking at that figure will see it as a distorted triangle. The question is, according to Descartes, why is that true-why do we regard that figure as a distorted triangle and not as a precise example of what it is? That sounds like a somewhat silly question, but as with a lot of silly questions, it is in a way rather deep. Why do we see this thing as a distorted triangle? Descartes’ answer is that the nature of the mind is such that regular geometrical figures are simply produced as models for the interpretation of experience; the mind just has that character, and if you were to develop a science of the mind, you would develop it in terms of such a geometry. Hence we are compelled to regard any figure that we see as a distorted regular geometrical figure because that is simply the way that the mind functions. We have here another kind of approach to S0; what it says is that So has certain specific structural conditions that are imposed on any system that is acquired. Of course in this case S0 and the steady state Ss are presumably identical for Descartes, but that doesn’t matter; one can also talk about systems in which, perhaps, the initial conditions simply involve regular geometrical figures, and some pattern of experience leads to the attainment of a more specific state (I don’t know if that is true in this case).

As far as the Cartesian type of approach is concerned, I think that is the right one. One can then go ahead and investigate it in various ways one can investigate it through psychological means, and one can look for neurological mechanisms. In this particular case, it seems rather plausible to me to interpret the very exciting work of the last ten or fifteen years on the visual cortex as presenting a neurological basis for a kind of Cartesian theory of mind. Thus the experiments of Hubel and Wiesel, experiments which lead to a picture of the visual cortex as carrying out primitive analyses in terms of line, angle, motion, and so on, produce a concrete physical system which has many of the properties of the postulated Cartesian mind and offers perhaps the beginning of an explanation as to why we see that object as a distorted triangle. This approach seems quite reasonable to me.

Coming back to language, if human beings were mice we would know exactly how to proceed to investigate any particular hypothesis at this point; however, since direct and intrusive methods are obviously excluded for humans, we have to use another type of approach. The natural way to proceed, if we are trying to determine the nature of So, is to try to find some property of the steady state that is minimally affected by experience, a property for which E (experience) is reduced as close to zero as possible. Of course, in order to demonstrate that there is no relevant experience with respect to some property of language, we really would have to have a complete record of a person’s experience—a job that would be totally boring; there is no empirical problem in getting this information, but nobody in his right mind would try to do it. So what we can try to do is to find properties for which it is very implausible to assume that everyone has had relevant experience. In my paper (see Chapter 1), I gave a number of examples of these, and I had thought that I might be able to talk now about some of the more complicated ones, but after the previous discussion I thought that it might be a good idea just to talk about the simplest example and to try to bring out the logic of that.

The most simple example that I can conceive, an almost trivial one, is the following: Consider sentences like “The man is tall,” for which we form corresponding questions: “Is the man tall?” by moving the word “is.” From examples of this kind we can make certain inductive inferences, which enable us to ask questions of a more complicated kind. Consider, for example, cases like “The man who is tall is sad.” From simple examples of this kind, we might make two different types of induction: we might assume that what is going on here is that the leftmost occurrence of the word “is” is moved to the front, and pursuing that hypothesis we would get “Is the man who tall is sad?”; or we might construct the hypothesis that the first occurrence of “is” that follows the subject of the sentence is moved to the left, in which case we would get “Is the man who is tall sad?”. Clearly, we make the second inference (there are, of course, other possibilities, but let’s limit ourselves to these two). One is an induction over the property leftmost, the other an induction over the property follows the first noun-phrase. What we see is that people make the second type of induction. Why? This is not a trivial question. I am very dubious about notions of absolute simplicity or the general theory of simplicity, and so on, but if someone wanted to propose such notions (which I do not), he would certainly say that leftmost is a simpler property than follows the first noun-phrase, for a very obvious reason, namely, that the property leftmost is completely definable in terms of the physical symbols themselves. To understand the property leftmost, all you have to know is the physical symbols and their order, whereas to know the property follows the first noun-phrase, a certain abstract mental processing has to intervene which tells you that this thing (the man who is tall) is a unit of a particular type. Of course, there need be no physical demarcation; this unit doesn’t have a space after it or anything like that there is a continuum of sound. “Follows” is a concrete physical notion, but the concept of “first noun-phrase” or “subject” is a highly abstract notion, so complex that in fact nobody knows how to describe it properly. Thus, supposing that there is a hierarchy of inductive processes, if a scientist from Mars were to look at examples like these and were to ask the question that any scientist would raise, namely, how to go on to the next case, he would of course first take the property leftmost, not the property follows the subject. Indeed, in any kind of hierarchy of inductive processes, the former is the obvious property to examine first, because it is a concrete physical property that doesn’t involve subtle, complex, unknown kinds of abstract mental processing. It is certainly true that children never make mistakes about this kind of thing: no child ever tries the Martian hypothesis first, then is told that is not the way it works and subsequently goes to the other hypothesis. The process of language learning does not resemble the way in which a scientist would study the question, by trying the elementary inductive processes first and then going on if this fails to more complex assumptions. This raises the question of why induction makes use of this very complex property and not the very simple one.

I said in my paper that from passive observation of what people are doing we could hardly know whether they are using this hypothesis or that one, because in fact the more complex cases that distinguish the hypotheses rarely arise, you can easily live your whole life without ever producing a relevant example to show that you are using one hypothesis rather than the other

*Editor’s note: For a detailed counterargument see Putnam’s paper in Chapter 14 of the present volume, Chomsky’s rejoinder to the critique expressed by Putnam is to be found in his “Discussion of Putnam’s Comments” in Chapter 15.

one. The examples cited are the only kind for which the hypotheses differ, and you can go over a vast amount of data of experience without ever finding such a case. Thus in many cases, the Martian scientist could not know by passive observation whether the subject is using the first hypothesis or the second one. The proposals that were previously attributed to me were described in the wrong terms: I did not say only that the child could not generally determine by passive observation whether one or the other hypothesis is true; what I said is that we could not know by passive observation whether in fact a person is using one hypothesis or the other because the evidence is essentially not available, just as for the most part it is not available to the child in terms of direct linguistic experience. Nevertheless, the fact is that the inductive operation does proceed without error, and even without trial, to the second hypothesis. How can this be explained?

First of all, I will mention some approaches that don’t offer a valid explanation. Compare Papert’s work on the recognition of the “arch” pattern. We can construct a program that will deal with this structure in terms of groups of elements in a hierarchical structure; we can also construct a program that deals with this as a linear array in sequence. Since we can do it either way, we can in principle carry out the induction either way. We can describe the linguistic structure in either way: We can describe it as a linear array of words, and in this case we use the property leftmost as the basis for induction, or we can describe it as a hierarchical structure, in which case we use the property follows the subject if that is the way the hierarchical system works. We learn nothing from the fact that it is possible to construct programs to look at things hierarchically, because it is also possible to construct programs to look at things linearly. The issue is, in this case, do we look at the sentences in a linear or a hierarchical manner in order to carry out the induction? The existence of a program that will do one or the other is irrelevant to the question of explaining how the inductive step is taken.

A more relevant question would be this: in other cases do people in their inductive behavior regard systems as hierarchically or linearly structured? Unfortunately, this question doesn’t get us anywhere because they do both. There are cases in which people deal with properties like leftmost (they may regard an array of elements as linear and consider the physical arrangement of the elements), whereas there are other cases where people take into account all kinds of hierarchical structures in visual space or whatever. What we have to ask is what is the property in the initial state So that forces us, in this specific linguistic case, always to go to the hierarchical abstract rule and always to neglect the more elementary linear physical rule?

Several answers have been proposed to this question: the right one, I think, is the one which is implicit in the theory of transformational grammar, which in effect asserts that there is a notation available for describing linguistic rules that does not permit the formulation of the property leftmost. In fact, to formulate the property leftmost within the framework of that notation, you would have to use quantifiers in structural descriptions of rules, and that is not permitted. Alternatively, in terms of that notation, to describe the property leftmost would be extremely complex (involving quantifiers), whereas to describe the property follows the first noun-phrase would be very trivial.*

*Editor’s note: This assumption is challenged by Putnam in his paper (Chapter 14) and developed further by Chomsky in his reply to Putnam.

Of course these are properties of the specific theory and the notation it provides, not general properties of systems of representation. It is a property of one very concrete specific theory that within it, to formulate the property leftmost requires the use of quantifiers, whereas the property follows the first noun-phrase can be formulated without the use of quantifiers. So there is a very specific theory of representations in terms of which follows the first noun-phrase is a more elementary property than leftmost; but that happens to be a property of this specific concrete theory and not a consequence of any general theory of representations and structures. Of course this property has many consequences elsewhere; it has vast consequences for grammar, where, applied to other linguistic structures, the use of the category leftmost (that is, rules involving the use of quantifiers) should always be less accessible than properties like follows the first noun-phrase. This hypothesis is one that is rich in empirical consequences and to my knowledge true.

Just as Descartes is saying that we perceive figures in two-dimensional visual space in terms of regular geometrical figures and distortions of them because that is in the nature of the mind, what I am suggesting here is that the very explicit and specific theory that makes the formulation of one property far more complex than the formulation of another is a characteristic of the mind. Perhaps somebody will find the physical basis for it someday. I would assume that we can regard ourselves as talking about the brain, in the same sense that we regard Descartes’ theory of perception as talking about the brain, but now we do it at an abstract level because we have no analogue to the Hubel-Wiesel mechanisms.*

Discussion

Wilden: A point of information, in order to understand a further dimension of the example you are using: What happens to the meaning of the sentence if you add a “not” after the second “is” and then make the same transformation as before? You will see then that the transformation no longer produces the same kind of result. The reason for this difference appears to stem from hierarchical functions in the sentence, and specifically from the hierarchical situation of syntactic negation (that is, “not” as distinct from “no”) in linguistic utterances.

Chomsky: What you get is “Isn’t the man tall?”

Wilden: That is not the point of my question. Let me rephrase it. You have set up a transformation based on moving the expression “is” so as to produce a question. In your example, “The man who is tall is sad” is transformed into: “Is the man who is tall sad?” What I am suggesting is that if you insert “not” into the example and then make what appears on the surface to be the very same transformation, you do not obtain the same results. Let the original sentence be: “The man who is tall is not sad.” Move the “is not” to create a question, and the result is: “Isn’t the man who is tall sad?” By means of what appears to be a simple syntactic reordering, the semantic structure of the sentence has changed. The question is why? Why has the semantic level changed?

Chomsky: That is true, but the semantics has nothing to do with it!

Wilden: The semantics has everything to do with it! The object of my question was precisely to bring out this point: Just what is the relationship between syntactic structure and semantic structure? And why do we find these apparent interferences of one with the other?

*Editor’s note: See the editorial comments at the beginning of this chapter.

Chomsky: Semantic considerations arise in a different connection. We are talking here about a system of rules that gives us the surface forms of sentences, and the theory in question says that the surface forms of sentences are given by, let’s say, derivations, each derivation being a sequence of abstract representations of phrase structure; let’s call them K1 to Kn, each of these being a representation of the phrase structure of the sentence, and the move from K1 to K1+1 taking place by means of one of the transformations that is expressible in our notation. The last of these phrase structure representations is what is sometimes called the surface structure, while the first one is sometimes called the deep structure. Now, the particular rule that I was talking about is one that gives you the surface structure from a slightly more abstract representation of what I present here as a sentence. The question you are asking is what is the meaning of this whole thing.

Wilden: No, I am not concerned with “the meaning of the whole thing.” The actual meaning of this sentence is of no concern to me at all. I am not asking questions about the meaning of this or of any other particular sentence. I am asking a question about relationships. In this example, my question is concerned with what happens to your theoretical explanation of transformations when one begins with an example which includes an expression in this case “not”—which is clearly of a different level of logical types (to quote Russell) from that of all of the other expressions in the sentence. That is what I meant by my earlier reference to hierarchies.

Chomsky: The answer to that question is that if you introduce the word “not,” then the whole unit, for reasons that I won’t bother explaining here, goes to the front.

Wilden: And the sentence changes its meaning.

Chomsky: Yes, but that has no relevance at all until we turn to the next question that one can ask, namely, how are derivations associated with meanings?

Wilden: If I may be permitted to insert a clarification at this point, it seems to me that our obvious inability to agree on what this question is actually about is the result of our speaking different dialects in the discourse of science. You are speaking in linguistic terms; my question was posed in the metalinguistic terms of semiotic communication. I leave the matter there and let Chomsky continue.

Chomsky: The next question one can ask is how derivations are related to meanings through the surface structure. There is one type of semantic relations, sometimes called case relations or thematic relations (things like agent, instrument, and so on), which seems to have to do with the initial (deep) structure, but in every other respect, and certainly with respect to notions like negation and, in fact, all logical structure, meaning is directly determined by the surface structure. Correspondingly, we would expect the meaning of those sentences to differ under this theory, because the meaning would depend on the configuration of the surface structure, which happens to have “not” in different places. You can find very striking cases where the meaning changes, depending on the application of transformations. Take the sentence “Many arrows didn’t hit the target” and compare it with its passive: “The target wasn’t hit by many arrows”; these sentences have totally different meanings, totally different truth conditions, and so forth. This follows from a very reasonable theory of surface structure interpretation of logical particles. So the question is a perfectly legitimate one, but it has no bearing on the issue.

Wilden: It does have some bearing. Unfortunately, your example about the arrows, although perfectly legitimate in itself, is not of the same type as the example produced by my insertion of “not” into the sentence you began with. In other words, in your example of negation in the active and passive forms of the sentence about the arrows and the target, the expression “not” does not undergo any change in function after the transformation, whereas in my example “not” does change function, and the sentence changes its meaning because of it. To put the issue as simply as possible: When “The man who is tall is not sad” becomes, by the simple process of positional change, “Isn’t the man who is tall sad?” then the negative declaration in the original-which in Jakobson’s terms refers only to the level of the code, and not to the communicational circuit between senders and receivers not only becomes a question, but also introduces a shifter. The shifter indicates the relationship of the sender to the message and thence to the receiver, as well as to the code. And the shifter in the question is the word “not.” How is it that what appears to be a simple syntactic shift can result in the introduction of not just a different meaning, but what is in effect a new semantic dimension into the sentence? The “not” has changed its characteristics and its function. What I am now wondering is whether, in such examples, we are not in fact faced with a hierarchy of logical determination in communication (and not only in language), a relationship between levels of communication and a relationship between the message and its environment—in other words, a relationship or relationships which are not taken into account by the way in which you are explaining these forms of transformation. That is not an easy question, and I know that we won’t find a simple answer to it.

Chomsky: It’s a good question; I think the answer is no, for the reason I have just described, namely, that insofar as logical particles are concerned (like negation, quantifiers, anaphora, scope and so on), the meaning is a function of the surface structure, including the form of the sentence and its actual bracketing into phrases. The two sentences “Isn’t the man who is tall sad?” and “The man who is tall isn’t sad” have a different physical shape, and different physical arrangements, and from these physical shapes and arrangements we can infer the meanings, getting the differentiated meanings that you talked about. The point is that there are interactions among order, logical particles, and quantifiers, all of which have to do with surface structure, and the more elements of this type that you introduce into a sentence, the more the meaning is likely to change, depending on the transformations that take place, as a result of the simple fact that the meaning is directly determined from the surface structure.

Returning to my paper, at this point I went on and gave more complex examples; I will mention one and will give an analogue in French. One of the examples had to do with what I call the specified subject condition; briefly, that has to do with the following kind of configuration. Suppose we consider something like: “We expect John to like each other”; why can’t we say sentences like that? We can, however, say “We expect each other to win.” Why doesn’t the first sentence mean “Each of us expects John to like the others”? It could mean that-the meaning is perfectly sensible-so why doesn’t the sentence have this meaning? There is nothing semantically incorrect about it. Let’s take another sentence: “We like each other’; this sentence means “Each of us likes the others.” Correspondingly, why doesn’t the sentence “We expect John to like each other,” by induction, mean “Each of us expects John to like the others”? This sentence is fine, but “We expect John to like each other” doesn’t mean that; somehow, the induction, the generalization is blocked. The generalization that says “replace we… each other by each of us … the others” is blocked in this case. Again, this is the kind of thing that nobody is taught, so once more the question arises, why don’t we make the natural generalization? I think that the answer lies in a very general condition of rules which says roughly that if an embedded sentence has a subject, only the subject is accessible to the rule. This sentence has a subject, John; therefore only that subject is accessible to any rule, and particularly the rule of reciprocal interpretation. Therefore, I think that just as it is reasonable to assume that the specific notation which gave the appropriate hierarchy to the inductive possibilities is simply part of S0, correspondingly it is perfectly reasonable to assume that the specified subject condition is part of S0. If somebody believes that this condition or anything remotely like it arises from general developmental mechanisms, all he has to do is present the mechanism that leads to the condition that if that embedded sentence has a subject, then only the subject is accessible to the rule.

Papert: What do you mean exactly by “accessible to the rule”?

Chomsky: Consider a rule that relates two phrases for example, the rule for reciprocal interpretation takes a plural noun-phrase and it takes the phrase each other, which can be anywhere in the sentence-and it says that these two elements are related by the principle of reciprocal interpretation, which essentially gives the meaning each of these… the others. Now, here the rule cannot apply to the noun-phrases we and each other in “We expect John to like each other”; that is, each other is not accessible to the reciprocal rule. Why? Because the embedded clause “John to like each other” has a subject, John, and the principle says that if the sentence has a subject, only the subject is accessible to the rule.

Papert: I think Piaget has often proposed such principles.

Chomsky: I will simply assert that there is no proposal in the literature from which one can deductively conclude that the specified subject condition operates in English. If you have an example to the contrary I’d like to study the deduction; but frankly, I don’t expect to see one.

Returning to my paper, what I then went on to point out is that really the issue is far more complex than this, and I gave the following example: take the sentence “John seems to the men to like Bill.” Now let’s try “John seems to the men to like each other.” The first one is correct; the second is not. Why doesn’t the second sentence mean “John seems to each of the men to like the others”? Now the sentence “John seems to each of the men to like the others” is a perfectly fine sentence in English so why don’t we carry out the induction from “We like each other” to “John seems to the men to like each other” to give the meaning “John seems to each of the men to like the others”? What blocks that induction? Notice that the specified subject condition (SSC) as so far formulated doesn’t block it because there is no subject in this case, and therefore the SSC does not block the induction, at least so it appears. Nevertheless, this example looks very much like the other, so we may argue that if the SSC was correctly formulated it ought to cover this case.

A natural way to give a correct formulation is to use a concept of traditional grammar, namely, the concept of understood subject. If you think about the following sentence: “John seems to the men to like Bill,” traditional grammar would say that the word like does have a subject, that is, an understood subject, namely John, since it is John who is doing the liking. This concept of understood subject is very complex and very intricate and depends on all sorts of complex semantic properties of words, but it is a very important concept. What apparently is the case is that the SSC does not care whether the subject is physically present or whether it is only mentally present. Let us say rather that at one of the levels of this abstract system of representations it is physically present, in the computational systems which, according to our theoretical hypothesis, the mind is employing. Thus at a certain level the subject is really there, and this fact turns out to be crucial.

What apparently is the case, judging from examples like this (and there are many others), is that the rule is not permitted to apply to an element if there is a subject either physically or mentally present in the sentence in which that element appears, that is the real principle. Notice that there is nothing logically necessary about that principle, English would be a perfectly fine language if it didn’t have that principle, only in English you would then say “John seems to the men to like each other.” There would be no problems about communication; all the properties would be acceptable. It just wouldn’t be a human language; it could be the language of another organism, namely, an organism whose initial state would not include this principle. One could conceive of an organism exactly like humans, but minus the specified subject condition, and it would talk with a fine language which it could use for all possible purposes. In fact, we could observe that organism for a very long time and not even know that it is not using the SSC, because these questions generally don’t even arise. Similarly, physicists could observe the natural world for a long time and not know that certain physically interesting properties have ever been realized that is why people do experiments. Likewise, if you were simply to observe the flow of experience, to make a movie of people talking for instance, you might not know whether they are using the SSC at all, and certainly not if they are using it in the special sense where mentally present subjects act like physically present subjects. Yet the fact is that everyone, without ever committing an error, uses that condition even if they have had no relevant experience whatsoever. We could prove this by taking a total record of a person’s experience and showing that he never produced any such examples or had any specific training concerning them, but I don’t think that is a worthwhile experiment because it is obvious what the result would be however, it is an experiment that someone could carry out if he doesn’t believe what I am saying. What is obvious, I mean, is that there would never be any relevant experience. Someone who thinks the contrary is true should study children and discover the moment when they are presented with experience that tells them that mentally present subjects act like physically present subjects. This is an answerable question. It could be that children do try to say things like “John seems to the men to like each other,” but they are told they can’t do that and then they conclude in some way that mentally present subjects are not like physically present subjects, but surely that doesn’t happen.

Let’s turn to something that illustrates the operation of this principle in a totally different domain, in a different language -French in this case-in which apparently you can say things like “J’ai laissé Jean manger X” or “J’ai laissé manger X à Jean.” If you have a quantifier, for example “tout” instead of X, then there is a rule that says that “tout” can be moved, which ought to give these two sentences: “J’ai tout laissé Jean manger” or “J’ai tout laissé manger à Jean.” According to my sources, the second sentence is correct and the first one isn’t, why is there that distinction? In the first case there is a subject in the embedded sentence; therefore, by the SSC, only the subject is accessible to a rule that moves an element, and thus you can’t move the word “tout” just as you couldn’t interpret the reciprocal phrase “each other,” when there was a subject. In the second case there is no subject, but rather something that functions as a subject in a prepositional phrase (“à Jean”); since there is no subject, the SSC doesn’t apply, therefore “tout” is accessible to the rule, and therefore it moves. So it is perhaps the same condition, but operating in a totally different domain where its effects are quite different from the earlier case I discussed.

There are a lot of cases in which that same, quite abstract condition has empirical consequences as to what is possible or is not in language. To the extent that that condition leads to empirical consequences, it is very easy to prove that it is false, if that is the case; but my point about that condition is that it is a highly specific one which, as far as I know, has no analogue in any other cognitive domain. There is no general principle of cognitive structure from which one can deduce that this principle holds for language. For this reason, I think it extremely reasonable to postulate that we have here a highly specific linguistic mechanism specific to this cognitive domain, which is part of S0.

One can go on with many other examples of this kind, but all these examples are misleading in one important respect (and in rereading my paper I noticed that it is misleading in that respect), namely, what I gave there looks like a list of properties that belong to S0. What I failed to say, and should have said, is that this list of properties forms a highly integrated theory. It may not seem so when they are just listed as examples, but in fact they all fall together in a very natural way; they don’t logically follow from one another, but they are so close in form when you actually formulate them correctly that they really flow from a kind of common concept, an integrated theory of what the system is like. This seems to me exactly what we should hope to discover: that there is in the general initial cognitive state a subsystem (that we are calling S0 for language) which has a specific integrated character and which in effect is the genetic program for a specific organ (here it is the program for the specific organ which is human language). It is evidently not possible now to spell it out in terms of nucleotides, although I don’t see why someone couldn’t do it, in principle. Likewise, it seems to me that someone who did not know the laws of physics could have hoped in the seventeenth century that the theory of perception could be spelled out in terms of properties of the brain. Similarly, we are in a position like that of Descartes with respect to visual perception. We can say what the genetic program must look like (of course this a scientific and not a mathematical “must” we are dealing with a hypothesis about reality) but we cannot yet say what the genetic program is-which does not mean that we could not in principle say what it is. One has to make a sharp distinction between notions like “inexplicable” and notions like “unexplained.” At the moment there is no explanation, in terms of the biological structure of the organism, for the genetic program for this particular human language, and of course that is true of any other organ as well. To say that there is no explanation at the moment means, to me, that there is no set of principles by which we can deductively conclude this or that. There is no explanation at the moment for the fact that the heart is what it is, or the liver. That is not to say that it is inexplicable. It is possible that the principles are actually known but we don’t know how to draw the conclusions because it is too complicated.

Now, if that is true for the liver, why shouldn’t it be true for the brain? The liver, I am told by neuroanatomists, has about the same number of cells as the brain, but the brain is far more complex as an organ. So why should we expect that the genetic components of the brain should be easier to account for or to deal with than the genetic components of the liver, which are also unexplained but not inexplicable? The answer is that there is no reason at all to believe that. When people say, as I think Toulmin is saying (see Chapter 13) if I read him correctly, that to assume these properties in S0 is to attribute to it a structure so specific that the burden on evolutionary theory is too heavy-if that is asserted, I simply deny it. First of all, I don’t think that the structure assumed here, that is, the SSC, is more complex than the structure we presume to be true of the liver; in fact, if we were really to explain in detail the functioning of the liver, I’m sure that we would have properties at least as complex as the SSC-even far more complex-yet nobody is troubled by the fact that we leave it to the theory of evolution to explain why the liver is what it is. We don’t assume that the child learns to have a liver, and similarly, I don’t see the reason why much simpler properties should impose an unsupportable burden on the theory of evolution, quite the contrary, even if we could advance this study to the point where we really had a much more complex theory, and I think we are ultimately going to discover a far more complex theory of what the initial state must be. A theory of the brain should surpass in complexity a theory of the liver. We assume that the liver is a much less complex organ than the brain, despite the fact that it has approximately the same number of cells.

As a final remark, it seems to me that no burden of an unacceptable nature is imposed on the theory of evolution in this case-quite the contrary. It is possible, of course, that my argument may be wrong, since it is not a demonstrative argument, but I think there is plenty of evidence for it. Furthermore, I think it’s the kind of thing we ought to expect to find. We ought to expect to find that highly specialized structures, like the general mammalian structure for the visual cortex which leads to “Cartesian” properties, or the very specific structures of particular cognitive organs like the brain, are just the reflection of a genetic program. This explains why we have these highly complex and remarkable systems that develop in a curious fashion along patterns of inductive inference which seem very strange from the point of view of scientific induction but which nevertheless are uniform, very specific, and highly complex in the results to which they lead.

Dütting: I have difficulty in accepting your explanation of the SSC. It seems that in the two sentences, “Each of the men likes the others” and “The men like each other,” the meaning is different. Why do you need the SSC condition, if it is a semantic problem?

Chomsky: What I said was that the two sentences are synonymous. That’s a good first approximation, but still somewhat inexact. What I should have said is that the two sentences have a fixed relationship in meaning which comes very close to synonymy. Let’s call the relationship in meaning between the two “R.” The relation R holds between the two sentences, which is to say that the pair of phrases “The men… each other” can be interchanged with the pair “Each of the men… the other(s),” preserving the relation R over a very large domain. However, there are cases where this is not true, in certain contexts. Now, why is this so? That is the problem that has to be explained, namely, why does the pattern of generalization break down in this case? The answer that I suggested to this problem is the SSC. There is also a difference between “the other” and “the others.” So let’s say “the other” is related by R to “each other” and “the others” by R’ to “each other.” We can then ask the same question at a slightly more complex level: Why is neither the relation R nor the relation R’ preserved?

Suppose we agree that the SSC is the reason why the generalization breaks down. Then a paradox presents itself: namely, the sentence “Each of the men expected John to like the other(s)” is perfectly correct, but the sentence “The men expected John to like each other” is not. The reason is the SSC. However, assuming that the rule relating “the men” to “each other” is blocked in this case, how do you explain why the rule relating “each of the men” and “the others” is not blocked in the other case? Why are the two cases different? Why isn’t the SSC refuted by the perfectly good sentence “Each of the men expected John to like the others”? The reason is that another reciprocal rule applies in this case which says that “the others” is not the same as “the rest of the men.” Some rule is applying why isn’t it blocked?

This is an important question; its answer hinges on the very great difference between “each other” and “the others.” “The others” is not related by bound anaphora, a rule of sentence grammar, to “the men.” (This becomes clear if you notice that “the others” can begin a sentence, whereas “each other” cannot; for example, I can say, “Some of the men are happy. The others are sad,” but I can’t say “Some of the men are happy. Each other are sad.”)

Premack: Chomsky’s first example, “The man who is tall,” reminds me of a logically parallel case from a simpler system concerning a chimpanzee and the rule for pluralization that she apparently learned. In Sarah’s language the plural is formed by adding a piece of plastic which we call the plural particle or “pl.” In training her, we contrasted singular forms such as “apple is red” with plural forms such as “apple banana is pl red,” where “is + pl” = “are.” Although Sarah learned the training sentences and went on to pass the usual transfer tests involving new sentences, it dawned on us that there might be other tests she could not pass. For from the training we had given her, she could have learned either of two rules, only one of which was correct. She could have learned to use a physical feature comparable to the physical feature in Chomsky’s example, namely, applying a plural particle whenever there are two words to the left of “is.” Or she could have induced a rule based not on a physical feature but on a grammatical one, namely, applying a plural particle when the subject is plural. When we tested her with sentences of the kind, “red apple? fruit” versus “orange grape? fruit,” we found that she answered correctly, replacing “?” with “is” in the first case, and with “is pl” in the second. Thus the chimpanzee induced a rule based on a grammatical feature, not a physical one, despite the fact that the physical feature might be considered simpler than the grammatical one, and the training given the animal was equally compatible with both alternatives.

Sperber: I would like to make two remarks that refer directly to Chomsky’s discussion. First, there is no final state for all the domains of knowledge. For example, encyclopedic knowledge accrues throughout life. Second, Chomsky suggests that he can construct a special theory of learning for each domain of knowledge. I do not disagree with this point, but I would like to know how Chomsky would determine what constitutes a truly autonomous domain under the jurisdiction of a proper theory of learning.

Chomsky: As far as the first question is concerned, of course that is true: encyclopedic knowledge does not stop growing. Of course, that is true for language as well; for example, you keep learning more and more words, but the system doesn’t fundamentally change. We simply don’t know whether that is true for what you call encyclopedic knowledge. It might very well be that up to a certain point, whatever that may be, people are developing systems for organizing information, and that after that point they are just adding information to the systems. (I don’t know if that’s true, but I wouldn’t be surprised if that were the case.)

About the second point: how do we decide when we have identified an autonomous domain? Well, we can’t decide a priori; we can only decide when we have an organized system. When we come across an integrated system with special properties and internal interconnections, and so forth, then we can reasonably postulate that is a system. Of course that is an idealization, just as to talk about the heart is an idealization. Why do we say that the heart is an organ? In principle, we can put it in an entirely different way, and in fact, that idealization is misleading at a certain stage of investigation, where we want to talk about the integration of this system into other richer systems, some mental, some physical, and so on. So it is a legitimate but at some point misleading idealization, like other idealizations that demarcate subsystems for investigation. As for other cognitive domains, I think there are other areas where one can make a guess at least. Let’s take music, for example, which to all appearances constitutes a system, or mathematics. I don’t have enough information in these areas, but my feeling is that if one looks at the history of mathematics, what one sees is that for very obscure reasons, as a byproduct of something else, there developed in the course of human evolution this truly weird capacity to deal with abstract properties of the number system. Now, this is a capacity that has no selective advantages in itself; it may be a byproduct of something that does. But in any event, at some stage in human evolution a kind of dual capacity emerged to deal abstractly with problems of three-dimensional space and properties of the number system, and out of that amazing capacity developed 2,000 or 2,500 years of mathematics which, in a certain sense, finally led to a kind of consummation at the beginning of the twentieth century. At that point, roughly speaking, I think it would be fair to say that classical mathematics came close to being exhausted in the sense that the problems of number theory and of classical analysis were just too difficult for any mere human being; if they hadn’t been, Gauss would have solved them. What happened at that point is that classical analysis and number theory became in a way the special province of an exotic group of people who, from an evolutionary point of view, were freaks, special geniuses who emerged. This domain was just too difficult for the working mathematician to work on. The result was that a number of other branches of mathematics developed, perhaps because classical mathematics became too difficult. This is, of course, a vast oversimplification, but I think it is a rough account of what happened, and what it seems to me to suggest is that a specific human cognitive capacity, which one could investigate, was pursued nearly to its limits in the development of classical mathematics and was then more or less exhausted.

Like all theoretical models, the “Cartesian mind” postulated by Chomsky is supported by some constitutive metaphors whose analysis can be quite illuminating. Chomsky describes the process that is responsible for the attainment of the final steady state Ss starting from the initial state S0 as a “mapping function.” S0 is in effect “a function which maps experience onto the steady state.” If one keeps in mind the perceptual mechanisms “à la Hubel-Wiesel” (also summarized in my comments at the beginning of this chapter), it is clear why the notion of mapping is so important for Chomsky. Mapping means imposing a lawful point-to-point correspondence between patterns belonging to two different spaces (the “real” space and the “image” space or map). Experience is said to be “mapped” onto the steady state, therefore to be actively processed by the subject according to rules which are imposed by the subject, are universal, and are susceptible to be made explicit. In contrast to the constitutive metaphors of the Piagetian system, the Chomskian notion of mapping is projective and not introjес- tive; it is conceived as a complete system of rules available to the newborn baby, and not as a stepwise construction through trial and error.

In accordance with the neurophysiological discovery of preset filters, the environment is given a structure, not explored to derive structures from it. The environment channels the mapping from S0 toward Ss; it does not instruct the subject on how to do it. According to Chomsky, the subject is “exposed” to relevant information and the underlying structures are then “revealed” by the action of the environment. The predominance of optical metaphors (projective and photographic) is in perfect harmony with the rationalistic character of his basic conceptions. As we have seen in the Introduction, the rationalistic outlook of Chomsky differs from the “vitalistic” outlook of Piaget. Nowhere else in the book is Chomsky’s basic attitude so completely and analytically expounded as in this chapter. The full significance of the following discussions, including the one with Putnam in Part II, will not be appreciated if Chomsky’s positions, as expressed here, are not kept constantly in mind.

Scroll to Top