MY CRITIQUE OF NATIVISM
Is an anti-nativist stance logical?
I should say at the outset that Chomsky and all the other nativists implicitly acknowledge that they could be wrong. This follows from the fact that they defend nativism from an empirical standpoint, which means to say that they infer their thesis from the ostensible fact that all the evidence seems to point in that direction. To illustrate this further, it is important to consider the distinction between analytic and synthetic statements.
If the nativist’s hypothesis were an analytic statement, then they would not even need to provide evidence for their case; we would be forced to accept it based simply on logic. One cannot call one’s thesis an empirical one, and still insist that it is the only logical one. In fact, this is what distinguishes an analytic statement from a synthetic one.
An analytic statement is a statement that merely paraphrases the subject of the sentence; in other words, there is no additional information that is provided by the predicate that was not to be found in the subject. Real-world knowledge plays a minimal role in this regard, since all one needs to establish the veracity of the statement is to have just a rudimentary understanding of the terms in the statement.
Consider an analytic statement like Bachelors are unmarried males. What makes this an analytic statement is:
1) The fact that there is no additional information about the world; the predicate says nothing more than what is in the subject, and
2) If we were to negate this sentence, it would lead to a logical contradiction; in other words, we cannot imagine a possible world where the converse of this sentence can be true without entailing a logical contradiction. It is meaningless to say Bachelors are married males or Bachelors are unmarried females.
A synthetic statement, on the other hand, tells us something about the subject of the sentence which is over and above a mere definition/paraphrase. We have to also know something about the state of affairs in the real world in order to ascertain the veracity of a particular statement. For example, with the statement There is a bag of potatoes in the pantry, one of the truth-conditions that would make this sentence true is if there is actually a bag of potatoes in the relevant pantry. To take one more example: Freddie Mercury is a singer. For this sentence to be true, there would have to be a person by the name of Freddie Mercury, and he would have to be singer; if those conditions are met, then this sentence would be true.
What makes these sentences different from analytic sentences is the mere fact that we can imagine a possible world where the converse is true, without there being any logical contradiction. So even though it does not make sense to imagine a world where the statement Bachelors are married males/unmarried females is true, it would make perfect sense to imagine a world where There is no bag of potatoes in the pantry, or where Freddie Mercury is not a singer are sentences that happen to be true. So when I utter a synthetic sentence, like There is a computer in my room, what I am also saying, ipso facto, is that there is a possibility that there may not be a computer in my room; there is no contradiction in terms in saying There is no computer in my room.
So when the nativists say There is a Language Acquisition Device [LAD] in our minds which enable us to learn language, it is a statement of the non-analytic kind. We know this because they deem it necessary to quote empirical evidence for their stance, and conclude that all the evidence points towards their theory. If this statement were analytic, it would make no more sense to try and show evidence that the statement Bachelors are unmarried males is true.
It seems quite obvious that the nativist’s hypothesis is an empirical one, and that a statement like Humans possess an LAD is a non-analytic statement. However, some scholars, like Jerry Fodor, do actually try and argue that their manifesto is analytically based. Even Pinker sometimes confounds the analytic-synthetic distinction when he argues that ‘mentalese’ (the language of thought) cannot not exist, and that if it did not, we could not possibly learn a language, translate between languages, etc. Despite proposing that it is logically untenable to not postulate mentalese, he still dedicates an entire chapter in Pinker to argue for its viability. Once again, let me reiterate, statements that are analytically true need neither empirical verification nor empirical argument.
So despite Fodor arguing analytically, and despite Pinker blurring analytic-synthetic arguments, it is otherwise evident that the nativist manifesto is a synthetic one, which means that they acknowledge that there is no contradiction in the converse being true, viz. There is no LAD in our minds that enables language learning to take place.
As Stephen Cowley (in a paper published in 2001)tells us, Pinker, joining the other nativists, merely assumes his “radical, nativist theory is the only alternative to behaviourism”, and “that the work of Vygotsky, Piaget, Stern, Trevarthen, Bruner, Wertsch, Cole and Nelson [and Sampson] – to name a few – lends little understanding to language development.” This blinkered view is obviously unjustified.
When I asked what his thoughts on Educating Eve were, Pinker dismissively told me that he has not looked “the Sampson book” because he was too busy writing other books! Either Pinker is being deliberately blinkered here, or he is saying that The Language Instinct needed to be updated anyway, but when I asked him if the latter was the case, he said: “That’s not what I’m saying at all!”. It is evident, then, since this is not the case, that criticisms should not be ignored.
The problem with prematurely accepting nativism
Research necessarily cannot be neutral and objective, because it is based on observations; and one’s observations are always theory-laden. When this is acknowledged at the outset, and one’s assumptions are made clear, then meaningful scientific inquiry can be conducted. However, when a theory is accepted prematurely as axiomatic truth, it can lead one to force interpretations onto data that one is not even aware of (as I believe Pinker is guilty of doing) [as in Pinker’s biased representation of the SLI data mentioned above]. Many red herrings and pseudo-problems may present themselves. A pseudo-problem is one that arises solely as a consequent of one’s theoretical assumptions, which may never have even been an issue had one been working in another framework. One such pseudo-problem arises as a result of Chomsky drawing a distinction between competence and performance. It is said that to postulate an innate predisposition for language in the mind is “an unnecessary hypothesis forced upon one by the previous postulation of the concept of competence”, to quote David Crystal. If the notion of competence had not been proposed to begin with, there would be no need to rephrase the question of “How do we acquire language?” as “How do we acquire competence?” The danger is that of “imposing too much grammatical structure on the mind of the child: it is very easy to get into the habit of talking about certain structures as being innately present, the only reason for wanting to do so being that such structures have an important role to play in the analysis of adult grammar.
For Popper, scientific progress is based on problem-shifts. These problem-shifts can be progressive, or degenerative. A progressive problem-shift occurs when we update our hypotheses in light of new evidence, without tenaciously holding on to our theoretical constructs. A degenerating problem-shift can occur when one maintains a particular thesis, despite evidence to the contrary. I contend that this is the key problem with nativism. The nativists may be correct in their claims, but at the moment there is plenty of evidence, or other problems in terms of methodology, that run counter to what they claim, and this is what the central thesis of this dissertation is.
Chomsky
Response to the “speed of acquisition” thesis
As scientists, one of the first things we need to isolate is potential falsifiers. If any hypothesis does not have falsifiable status, it cannot possibly have scientific status. With regards to the speed of acquisition thesis, it is not at all clear what state of affairs would contradict this thesis: how slow would acquisition have to be if we did not have an LAD which Chomsky postulates? It is not at all clear that having an elaborate linguistic structure hard-wired somewhere in our brains would yield a different time-frame than the theory of language-learning outlined by Sampson.
Aside from this problem, there is the fact that Chomsky compares two very different things when drawing an analogy between language acquisition and other kinds of learning. A more apt analogy would be to compare meta-linguistic skills to that of the physics student, and to compare our unconscious linguistic knowledge to that of certain expectations regarding the laws of physics.
To illustrate: any competent L1 speaker of English will agree that John said that he was sick is an ambiguous sentence, and yet the ambiguity disappears when you say He said that John was sick. Very few people will be able to explain why this is so, because their knowledge regarding these things is largely unconscious; the linguist, on the other hand, ought to have no problem in explaining it in terms anaphor binding. Likewise, most people would “know” that once you burn a match-stick it would not be possible to somehow un-burn it, such that it reverts to being the way it was before it was burnt. Once again, most people would not be able to give an intellectually satisfying explanation as to why this is so. The physicist, on the other hand, ought to have no problem in explaining this phenomenon in terms of the third law of thermodynamics, which states that entropy can only increase with time.
One cannot compare tacit knowledge of language with overt knowledge of physics. A more commensurable analogy would be to compare meta-linguistic knowledge (how to draw phrase-structure trees, how to depict things like question-formation via transformational rules…etc.) with knowledge of physics (how to calculate the gravitational pull between two bodies of mass… etc.), neither of which could possibly be said to be innate. Tacit grammatical knowledge would be more commensurable with certain tacit expectations like object-permanence, time running ‘forward’ not ‘backward’, etc.
Sometimes Chomsky argues that it would take an indefinite period of time to learn language if not for our LADs, and it would therefore not even make sense to talk about language acquisition. I cannot see the force of this argument, since Chomsky seems to be treating what ought to be an obviously empirical hypothesis analytically, which as we have seen is something that simply cannot be done (cf. section 5.1.). Note, this is not the same as having innate dispositions. Chomsky’s LAD is far more specific, and makes a much stronger claim than Popper’s belief that we are innately predisposed to learn certain things. A disposition, on the other hand, is not restricted to language. As mentioned at the outset, the fact that humans are genetically endowed with certain faculties is not at all in dispute – it is a question of how much is innate. An innate disposition is not the same as having an elaborate, abstract, and very specific LAD hard-wired into our minds.
Response to the “stages of acquisition” thesis
Firstly, it is not necessarily a fact that language is acquired in the same stages all over the world. From a Popperian point of view, the sensible thing to do would be to try and find instances where language was not learnt in these stages, where the babbling stage was skipped for some reason, etc., but this has not been done. Of course, from a practical point of view an experiment of this sort would take about three years at least, but if the nativist researcher really cared about the veracity of his claims, it would indeed be a worthy endeavour; after all, many theorists have spent numerous years trying to teach primates to use sign language – unsuccessfully.
Pinker makes the claim that each preceding stage is necessary in order to progress onto the next stage. Rosenbek et al., in a book entitled APHASIA: A CLINICAL APPROACH (published back in 1989) document instances of babies who were forced to skip their “babbling stage” due to problems like severe tonsillitis, yet acquired language without a problem; this shows Pinker’s claim to be counter-factual.
Nevertheless, back to the point: if instances that are not in consonance with one’s hypothesis are found, we would have to update our hypothesis; if not, we can tentatively assume that the hypothesis holds true. David Crystal, though without any references or data, claims that the stages of acquisition are not universal, and therefore cannot be part of our biological endowment. Even Fromkin and Rodman, in an introductory textbook for Linguistics freshman, say that “Observations of children in different areas of the world reveal that the stages are similar, possibly universal […]” (my italics). Why do these evidently pro-nativist authors use the word ‘similar’, not ‘same’, and why are the stages ‘possibly universal’? This seems to be a tacit concession that this strand of their argument is flawed. Victoria Fromkin, by the way, was a student of Chomsky’s, and has always been a sycophantic follower his – according to Dr. N. Thwala (former head of the School of Literature and Language Studies at the University of the Witwatersrand), who studied with her and knew her personally.
Nonetheless, even it was true that these stages occur universally, it still does not make a good case for nativism; let us consider a parallel case:
Karate all over the world is learnt in a specific order: when learning to do karate, the beginner starts with basics, then incorporates these basic moves into a pre-set sequence of movements (called a kata), then moves on to basic three-step sparring, then to one-step semi-free sparring, then finally to free-sparring. This is literally true of every style of karate found around the globe.
Now, given that there are so many different varieties of karate all over the world, all started by different karate masters, why would each and every style follow the same steps when learning to do karate? I was quite intrigued to learn that even those who practise in the Kyokushin-kai school of karate learn it in essentially the same way. This is interesting because this style of karate is quite different from the others in the sense that it is a full-contact sport (also called “knock-down karate”), and therefore it is a lot less about an art form, and more about brute strength. Nonetheless, it still conforms to the general schema of karate-learning.
Also, let us take this one step further. Imagine a world where, for some reason, every style of karate was wiped off the planet. If people were to develop a fighting technique analogous to what karate was/would be, it is unlikely that it would be learnt in any order other than the (progressive) order that has always been. Even though this is speculative, the fact still remains that the various styles of karate all over the world conform to the aforementioned schema, and if it mattered to us and we wanted to explain this pattern, how would we do so? How would we explain the fact that this happens, seemingly universally? Would it make sense to say that we would have to be either at a loss or be willing to postulate a massive coincidence, if we are not prepared to accept some sort of innate karate gene?
Of course not, since there are other more sensible reasons for karate having to be taught and therefore learnt in this given order: quite simply, the respective steps are progressive, which means that the first step forms the basis for the next step, and so on. It does not make sense to go on to learning something more advanced when the simpler, more basic steps have not been mastered. Students of karate cannot learn to do flying kicks from their first lesson for obvious reasons. It is for the very same reason that a freshman would not be able to produce a qualitatively adequate Ph.D thesis, even if he were given adequate time and resources. A two-year old child cannot speak with the proficiency of an adult because he has not learnt how to – he has not had sufficient interaction with speakers of the language. Is this not a more reasonable explanation than the idea that the child cannot speak proficiently yet because his LAD has not yet instigated his genetic blueprint to move from the initial state to the steady state?
Why, then, is it a surprise that children do not speak fluently from the beginning? Simply because they need to master the basic steps first, and as with everything else, the fact that there are basic steps, which progress in a specific order, should really not be a surprise. One could object to this argument on the grounds that karate is consciously taught, whereas language is unconsciously acquired. Anecdotal evidence shows that children are actually taught to certain degree. Pinker acknowledges this as part of the explanation as to why hearing children of deaf parents do not acquire language efficiently by watching TV and listening to the radio. They need to interact with real, live speakers, because speech per se will not suffice without the concomitant actions, according to Pinker. (This is also untrue, though, as I have a niece who now speaks fluent Hindi solely from watching Bollywood movies – both her parents speak English to her, and she has never been exposed to Hindi in a real-world context).
Anyway, in various works, like his book EMPIRICAL LINGUISTICS, Sampson presents data showing that language learning can be construed as a graded series of language lessons. If this is true, it has to be acknowledged that language learning is not as passive and ‘unconscious’ as is commonly claimed.
It is therefore evident that since we do not think it viable to postulate a karate gene because karate is learnt in uniform stages across the world, we are, by the same logic, also not obliged to accept the claim that there is some kind of unfolding of a genetic blueprint because language is learnt in the same stages across the world (barring the fact that even that is contentious).
Response to the “critical period” thesis
The critical period thesis can be contended on many grounds. Firstly, let us briefly consider what one would need to do in order to test this hypothesis. There would need to be two individuals, of the same age, and with identical genes; ideally, identical twins would suit us best. Then, one would need to serve as a control ‘group’, and the other would need to serve as the experimental ‘group’. We need to then keep them apart in identical environments for the same amount of time, and manipulate the relevant variable, that of linguistic input. In other words, the control would have a normal upbringing, and the experiment would be normal, aside from the fact that nobody would be speaking to him. Of course we need to decide beforehand when this alleged critical period will end, so if we take it to be the age of ten, then we would, after ten years, put the twins in the same environment, and see if the experimental ‘group’ is able to acquire language competently. How long we give him to do so will depend upon his progress, but obviously with a maximum of ten years. If he is not able to acquire language, we may conclude that it may be because he has passed his critical period for language acquisition; if he does manage to acquire language, it could either mean that the critical period ends later than ten, or that there is no such thing as a “critical period”, in which case more research needs to be done.
The problem with this is blatantly clear: no ethics committee will allow research of this kind to be done. Aside from the fact that it is almost impossible to test this properly, there would still be problems even if we did conduct an experiment of this kind. We could not preclude other factors, like the fact that the experimental ‘group’ may feel somewhat neglected for not being talked to, and may even be rather traumatised because his/her brother/sister was talked to, and can converse so proficiently. This could lead to resentment, lack of motivation, etc.; the possibilities are endless, and an experiment of this kind is therefore both impractical and unethical.
Despite the fact that there are so many problems associated with this, the nativists use the example of Genie as evidence for the existence of a critical period as if the other factors did not even need to be considered. Obviously, any human being going through an experience like this would end up very traumatised, and therefore cognitively impaired in more ways than one. Despite this, Pinker and the other nativists ignore this fact; in fact, Derek Bickerton even says that there is no compelling evidence that Genie suffered any “lasting damage”. If this were true, it would eliminate some of the confounding variables, but most rational people would agree that it is not. As mentioned earlier, Genie grew up under rather traumatising circumstances, and such prolonged abuse would affect her development, cognitive and otherwise, in an adverse manner.
It is a well-known fact that learning generally slows down the older you get, and from a biological (or evolutionary) point of view, it would make sense that this would happen around the age of puberty. Up until puberty, we learn the general life-skills necessary for survival, and at around the age of thirteen, we are ready for autonomy. The human body is ready to bear children, and according to the psychologist Kohlberg, even our moral standards, by which the rest of our lives are governed, are established around this age, etc. It is not the case that these things are unchangeable, it just makes sense to establish certain things as standard so as to have a basis to live your life, reproduce, and teach your offspring to survive. Why should language be any different? I would find it difficult to leave a live human being out in the snow for polar bears to consume; an Eskimo does not – it is common practice to do so with invalids as they compromise the mobility (and therefore the survival) of the tribe. If I were raised as an Eskimo, and understood the reasoning behind it, I would accept this without any qualms.
My point is simply that what one was raised with is what has become part of one’s norms, and unless there is a good reason to change the status quo, it should not be changed. If learning a new language is perceived as pointless, it will obviously not be assimilated easily into one’s schema, and vice versa. Hence, motivation, or lack of it, is a more obvious explanation as to why adults do or do not learn language as efficiently as children do.
It is a fact that the ability to learn in general slows down with age, and that we are particularly predisposed to learn things faster up until puberty. For example, it is generally accepted that if a child starts karate before puberty he learns faster, and progresses more quickly than children who start later on. Some contend that this is because younger children (i.e. pre-puberty) are more supple, and generally more enthusiastic. Now, if this actually were the case, as it seems to be from casual observation, would we want to postulate a “critical period” for learning karate? Of course, the very thought is ludicrous. As Pinker acknowledges, brain metabolism is particularly high during these years. Human life is divided into five stages: infancy, toddlerhood, childhood, adolescence, and adulthood. During the first three stages, humans are generally helpless and in need of care. During adolescence, humans develop autonomy, and could quite easily survive on their own. During this stage, humans also peak in all their faculties: their sexual desire, metabolism, curiosity (intellectual and otherwise), etc., is at an all-time high. From a biological point of view, adolescents are designed to break away from dependence on their parents, find a mate and start a family of their own. Hence, it is neither surprising nor a coincidence that children around this age feel an irresistible urge to rebel against their parents and other authority figures. Aside from the prima facie plausibility of this theory, other leading biologists also agree.
Now, given that this is the case, it makes sense that we would be predisposed to learning things faster up until this stage of life. If we think back to the days prior to civilisation, where formal schooling, a tertiary education, etc., would not have existed, humans probably did break away from their parents and form their own families at around this age. In fact, some contend that that is the very reason vestiges of initiation rites exist in many of the world’s practices; and older traditions less influenced by the dawn of civilised man still practise initiatory rites around this time for the sole purpose of allowing adolescents sexual passage, and in so doing allowing them to break away from the family and form their own. The details are not important, but the point is simply that human beings as a result would have had to learn all the necessary life-skills in order to survive. This would include being able to identify potential predators, being able to identify members of rival clans, being able to hunt effectively, being able to recognise and work with members of one’s own clan in order to attain a common goal otherwise not attainable, etc. Failure to develop these skills would result in failure of that given clan/family to survive. Hence, natural selection would have to see to it that they are able to pick up these skills quickly, and able to put them to use at the appropriate time. As we have seen, this ‘appropriate time’ is around the age of puberty at the onset of adolescence. It so happens that the acquisition of language is one of these skills, but that is not surprising given the profound survival advantage it gives.
Why would this general predisposition drastically slow down after puberty? Simply because it is energy consuming. It makes sense to not have this kind of predisposition for the rest of one’s life because having such a high metabolic rate would compromise functioning in other arenas, thereby stopping one from carrying out certain tasks effectively and ultimately compromising survival.
This explanation seems more plausible than postulating a critical period unique for language acquisition. Neither Chomsky nor the other nativists give us convincing reasons to believe that this phenomenon occurs in order to facilitate the process of language acquisition per se. Hence, the “critical period” thesis is not tenable.
Response to the “poverty of input” thesis
Chomsky takes the instance of yes-no question formation in an embedded clause to be confirming evidence for his poverty of stimulus thesis, but this is clearly an example of language being structure-dependent, and therefore becomes a question of whether structure-dependency is a matter of innateness or not. This issue will be discussed in more detail later.
As for the other evidence Chomsky cites in support of this particular hypothesis, Sampson gives numerous examples of documented studies which indicate that Chomsky’s claim that children are not corrected on grammatical issues, but matters of ethics, is quite simply false. I clearly remember as a very young child being overtly corrected on grammatical matters; at this very moment I know for a fact that my nephews and nieces are almost always corrected when they utter grammatically deviant statements (I obviously would not know about the times when I am absent, except to take the parents’ word that they “have to learn how to speak properly”). Aside from this, I am continually told by my students, year after year, that their “little sister” or their “nephew from Cape Town” was always rigorously corrected whenever he/she spoke in a deviant manner.
This is, of course, rebutting anecdotal evidence with even more anecdotal evidence. The critic would obviously say that if I have a problem with Chomsky’s use of anecdotal evidence, it is hypocritical to use such evidence myself. Of course, this is a fair criticism. However, Chomsky is the one making a novel claim, and hence the burden of providing systematic, empirical data lies on him. If his method of putting forth his argument is flawed, then the conclusions drawn from that argument should also be flawed, or at least not taken at face-value – or as axiomatic truth, which is what Chomskyan tenets have become. Nevertheless, I would predict that if a systematic study had to be done on parental feedback, it would show that parents do, in general, care whether their children’s utterances are grammatical or not. This is based on the assumption that all the anecdotal claims, like the ones mentioned above, are veracious and worthy of investigation. Indeed those who have done studies of this kind show Chomsky to be either distorting the facts, or simply mistaken. Aside from Sampson’s 2000 paper, which documents research of this kind from various corpora, other scholars have so as well. Academics like Roger Brown and William Labov have been documenting child-parent interaction since the early 1970’s, confirming the common-sense view that parents do indeed systematically correct their children when they speak in a deviant manner.
One could object to this by saying that the point is not whether children make mistakes sometimes, or whether they are corrected by caretakers sometimes, but why children never make certain mistakes despite never being told not to. This is a valid point, and will be duly dealt with below.
Perhaps the main argument that grammar must arise in the individual from some special-purpose device, genetically coded and neurobiologically expressed, is that grammar is too arbitrary, subtle and quirky to arise otherwise. But if the influence on language acquisition is not only the language the infant hears, but all of narrative imagining, including all of the systems from which narrative imagining recruits, there would actually be an overabundance of sources for subtleties and quirks – without having to postulate a special device to introduce them, as per Mark Turner’s theories in books like THE LITERARY MIND.
Hence, the assumption Chomsky and the other nativists make in terms of poverty of input is somewhat inaccurate.
Response to the “convergence amongst grammars” thesis
Aside from the fact that it is indeed somewhat counter-intuitive to think that a rather well-read, well-educated, and rather eloquent orator has a level of language proficiency tantamount to that of a man with, for example, no formal education, Chomsky would say that the disparity lies in their performance, not their competence. In fact, it is not at all clear-cut what Chomsky means by these terms, since he uses them differently at different times.
Sometimes he uses the term ‘performance’ to mean what a person actually does, with ‘competence’ being whatever mental abilities enable him to do what he does. At other times, he uses ‘performance’ to include perceptual strategies and psychological processing abilities. A set of perceptual strategies is not part of what one does in terms of actual linguistics behaviour, and therefore should be considered as part of one’s of one’s competence. In the latter sense of performance, processing abilities and perceptual strategies are taken to be part of a human being’s general mental abilities, rather than being part of some particular language, like Zulu or Afrikaans. The ability to speak this particular language is referred to as competence. But then, Chomsky also speaks of particular languages having “performance rules”; for example, he speaks of free word order as being determined by “rules of performance”. Here, part of the grammar of some language, then, would fall under the rubric of performance, not competence. Now competence is defined as covering certain language-particular rules. So when Chomsky says that linguistics ought to concern itself with the study of competence, not performance, it is not even clear what he means; if free word order is determined by “rules of performance”, does it mean that the study of it is outside the domain of Chomskyan linguistics?
Nevertheless, let us ignore the problems inherent in Chomsky’s equivocal use of terminology. The question still remains: how would we measure the language proficiency (call it what you like) of a sample of members in a given community in a way that would show Chomsky’s assumptions to be correct?
To show that speakers in a given community, regardless of formal education, social background, etc., would be equal in terms of their language proficiency, what would we need to do? We would need to take a sample representative of a certain community and show them to be of equal linguistic proficiency. We would need to show that aspects of their grammar converge on points we have no viable reason to believe originates in environmental speech input. If Labov’s work is anything to go by, it is quite evident that native speakers differ systematically in their judgements on various aspects of grammar. If we take this to be performative, and therefore not indicative of competence, perhaps we could observe naturalistically the members of our speech community, and based on the data we gather draw conclusions on their various language proficiencies. But to do this would be to revert to Bloomfieldian linguistics, according to which the linguist makes generalisations from the corpus. As we know, Chomsky would be quick to remind us that this only allows us to make generalisations about performance.
As is evident, if one had to analyse the performance levels of various speakers in a community, most people would expect that there would be various levels of language proficiency based on numerous factors. To claim that this does not count as viable counter-evidence is little more than an objection to rationality, and a surreptitious way of immunising the hypothesis from refutation (some may notice that I am using Popperian terminology here).
If the relevant methods of investigating this in a scientific manner are not going to count, then we need not treat this as a scientific hypothesis, and therefore need not take it seriously.
Response to the “argument from universals”
Chomsky’s main argument for rationalism is by and large based on the existence of complex linguistic universals. Chomsky always cites examples of putative universals from transformational grammar, but just about every other theory of grammar that has been proposed has incorporated claims for extremely complex and sophisticated linguistic universals. Transformational grammar is typically seen as rationalistic, whereas structuralism is typically seen as behaviouristic, and as an upshot claim made from a structuralist’s perspective are stigmatised and not given due consideration. According to Lakoff (Parrett, 1974, pp. 170-171), Chomsky’s theory about the organisation of language, deep structures, transformations, etc., is consistent with strict empiricism, whilst structuralist linguistic theories about the organisation of language are consistent with rationalism. Hence, one doing research on language universals is not necessarily obliged to accept the other axioms that go with the Chomskyan school of thought. Pinker and Chomsky tend to mistakenly polarise theorists along these lines. Pinker, for example, places immense emphasis on the work of Joseph Greenberg, whose research was on universals of word order. Greenberg was not particularly interested in whether such universals reflected some kind of innate linguistic knowledge. Indeed it would be rather strange if he was. In Pinker quotes Greenberg’s observation that “in languages with prepositions, the genitive almost always follows the governing noun, while in languages with postpositions it almost always precedes” [my italics]. The phrase “almost always” shows that not all languages adhere to this restriction, which means that it is not, strictly speaking, a linguistic universal. This in turns means that knowledge of this kind cannot be part of our biological endowment. A universal like the one quoted above applies statistically, with exceptions. Norwegian and Kitharaka, for example, are obvious exceptions to this rule– they have both prepositions and postpositions (cf. section on “Principles and Parameters” above). One could also argue that the English word AGO is a posposition as well. Hence, the fact that Greenberg did not use his findings to bolster Chomsky’s line of thought (that it is universal, therefore innate) is not surprising; why Pinker makes that very claim in The Language Instinct is very surprising indeed. In explaining why, one would admittedly have to speculate, and perhaps Pinker and Chomsky feel the need to polarise theorists into either the “pro-Chomsky” school or the “anti-Chomsky” school.
Before we go on, I would like to allude to the logic of argument structure, in order to help us understand why Chomsky and his followers may be mistaken in inferring nativism from hierarchical structuring.
There are many ways to show an argument to be flawed. One of the ways to show a particular conclusion to be a non sequitur is to show a parallel argument with the same logical form, where the premises are true and the conclusion not necessarily true. For example, the following syllogism has both true premises and a true conclusion:
Pigs breathe
Animals breathe
____________
Pigs are animals
(Which reads: Pigs breathe; animals breathe; therefore pigs are animals). We know that this argument is not valid, despite both the premises and conclusion being true, because we can easily think of an argument that espouses the same logical form, has premises we all agree to be true, and yet leads to a conclusion that we would all agree to be not true:
Dogs are four-legged
Cats are four-legged
______________
Cats are dogs
So because this is an invalid logical form, we are not necessarily obliged to accept the veracity of any argument that makes use of this kind of form.
What makes an argument valid is the fact that the conclusion must necessarily follow from the premises. Hence, if it is a valid argument, and if we agree that the premises are true, then we are obliged to accept the conclusion as being true. Whether the premises are true or not is a separate issue, but the fact remains that if they were true, the conclusion would be true.
We could put the argument Chomsky postulates into a logical form as follows:
If all known natural languages make use of hierarchical structuring, then this must be as a result of biological endowment
All known languages natural languages make use of hierarchical structuring
This is a result of biological endowment
Chomsky gives further supporting arguments, like the fact that it is logically possible to have a language that is structure-independent, yet this does not occur in the real world, analogous to the fact that it is logically possible that humans could have four arms, and the fact that we have two shows something about the nature of our genetic make-up.
If we isolate the logical form of the above-mentioned argument, it looks like this:
p -> q
p
____
q
This is a valid logical form, which means that if we accept the conditional p à q as being true in this world, then we ought to accept the rest of the argument. So the way to show this argument to be flawed would be to think of a similar conditional, and show it to be not necessarily true due to the fact that it leads to a conclusion that seems at odds with reality – and uncontroversially so.
So, before concluding that we do not need to concede that structure-dependency is necessarily part of our genetic make-up, let us consider a few non-linguistic instances of hierarchical organisation.
The following arguments are adapted from a similar argument put forth by Sampson in Educating Eve. The reader may think I am over-doing it by using so many examples to illustrate essentially the same point, but I feel this is necessary since this is by far the most important and convincing argument that Chomsky puts forth for his case, and therefore needs to be argued against in quite some detail.
Imagine there are two computer programmers. Programmer One (henceforth P1) makes use of modules, such that each sub-part of the program can be saved and run autonomously, and yet work as one complex program when the entire program is completed. Programmer Two (henceforth P2) does not make use of modules, but still produces a program of equal complexity to that of P1.
P2’s program is in no way different to that of P1’s, insofar as the end-product is concerned. The main difference is that P1 works on his program bit by bit, making sure each constituent sub-program runs as it is supposed to, and is saved onto the hard-drive; only then does he go on to write the next sub-program, etc., until the entire program is completed. If there happens to be something wrong with the program, it will be easier to found out where the fault lies by re-checking each sub-program. P2, on the other hand, writes out his entire program from beginning to end, then saves and runs it. If P2’s program contains a bug, he would have to go through the entire program to see where the fault lies.
Now, imagine there were two rooms; one containing a group of ten P1-type programmers, and the other containing a group of ten P2-type programmers. Imagine the two groups are having a competition of some sort, in terms of which the room that produces the largest number of completed programs in a day does not self-detonate. In other words, the group that survives would be the one that works the most productively. The problem, however, is that the power supply keeps shorting out at unpredictable times, causing the computers in both rooms to shut down. Imagine also that only executionable programs can be saved; if a program were not able to run autonomously then it would be impossible to save it. When the power supply begins to run again, the programmers in both rooms resume their respective tasks.
Obviously, this poses a serious problem for the P2 programmers. When they resume their programming, they would have to start all incompleted programs from scratch.
If we were to imagine a world like this, it is evident that P1-type programmers have more of a chance of surviving. The thing that sets P1 programmers apart from P2 programmers is the fact that P1 programmers make use of modularity. Aside from the fact that our P1 programmers stand a better chance not self-detonating, there are also other advantages of using modularity in computer programming. To name just one of the many: if there were a problem with any given program, a P1 programmer would be able to find the glitch rather easily: he would just need to run each module separately and see which one does not run properly. This would narrow the scope of his troubleshooting very substantially. A P2 programmer, on the other hand, would need to peruse the entire program from start to finish in the hope of finding the glitch. Hence, it just makes more sense that programmers adopt the P1-type modus operandi.
Another kind of “modularity” is to be found in the human body. We have different parts in our body which function in different ways, and yet work together as a whole. For example, we have a separate circulatory system, separate respiratory system, etc. If something were to go wrong with the human body, it is therefore easier to diagnose and treat illness; if something were to go wrong with the heart, it would make sense to also check the condition of the lungs, since they work together. It would not make sense to pay particular attention to, say, your left toe when treating a heart problem since there is no obvious structural/functional link between the two. Imagine if the human body were not somehow divided into constituent parts with separate functions. Even though this may not be easy to imagine, it is evident that if something did go wrong, it would be that much more difficult to diagnose any given illness since we would not be sure where to look. Normally, if something were wrong with the visual system, the doctor would check the actual eye-ball first, then failing to find any flaw he may examine the nerves connected to eyes, along with the occipital lobe of the brain, etc., because they all function together to constitute the visual system. If the body did not function in this manner, we would have to check everything from A-Z in order to determine the problem.
Aside from computer programs and physical organisms, hierarchical structuring can also be found in social institutions, like universities and schools. Let us briefly consider how an educational institution like a school is hierarchically structured.
Amongst the learners there are the various grades. Amongst the staff members there are post level 1 teachers (PL1), level 2 teachers (PL2), and level 3 teachers PL3. Post-level 1 teachers are teachers who do not belong to management, post-level 2 teachers are the HOD’s, and post-level 3 ‘teachers’ are the posts that the vice-principal and the principal hold.
(I now request you to imagine a hierarchical tree structure which illustrates the components mentioned above; I regret that I am not able to reproduce the one I have drawn on this blog...)
With a structure of this kind, let us consider what follows. Firstly, there are clear dominance relationships, with the principal presiding over the school. If the principal decides to, say, have a civvies day this coming Friday, he may do so simply by informing the others. If one of the HOD’s decides to do so, they can only get their wish granted with the permission of the vice-principal and principal, who have the authority to disagree and therefore not grant the wish.
If a student wanted to change some aspect of school policy, he would have difficulty doing so. If he were to follow protocol, the first thing he would need to do is go to his class-teacher, who would then approach his HOD, who would go to the vice-principal, who would go to the principal. Students are not allowed to go directly the principal or vice-principal, even though this is possible in theory. They have to follow protocol.
The decision(s) (pertaining, of course, to the running of the school) made by those higher up affect all those lower down, yet the decisions made by those lower down can only affect those still lower down, and then to can only be implemented if those higher up are in agreement with them. Since those right at the bottom of the structure have no nodes to dominate, they have virtually no say in the running of the institution.
Note, once again, that schools do not necessarily have to be like this. It is quite possible, from a logical point of view, that schools need not employ structures of this kind; for example, it is quite possible to imagine a school which does not categorise its staff members into different levels. Either each staff member works autonomously, or all the staff members form a committee where each member has equal power. Given that this is a logical possibility, how do we explain the fact that this kind of structuring (or non-structuring) never occurs? The fact that schools throughout the world, which have had no significant connection to each other, make use of this kind of structuring requires explanation.
Most people would agree that schools have this kind of structure due to the fact that it makes school-management easier. Any other system would lead to veritable anarchy.
Now imagine I tried to explain this ‘phenomenon’ along the following lines:
The universal structuring of schools throughout the world occurs due to the fact that there must exist an “innate universal school structure-dependency gene” (IUSSDG, consider this an analogue to Chomsky’s LAD), which has to be a result of pre-programming in our genetic make-up, one which predisposes us to arrange institutions of this kind in a structure-dependent manner. The most obvious objection to this claim would be that not all schools are exactly the same all over the world. This quasi-contradiction can easily be explained in terms of parameters; we have a universal schema regarding the outlay of schools, which can be adapted to various circumstances. If a school needs four deputy principles, then our parameters would have to be set in a relevant manner to accommodate this. (We would have to ignore the fact that an instinct, by definition, is something which is not adaptable, and therefore not variable; the webs that a black-widow spins are the same all over the world, and always has been since the discovery of the black widow).
How would one refute my hypothetical IUSSDG? It is indeed not easy to think of a possible scenario that would disprove my purport. Whenever an alleged counter-example presents itself, we could very easily just put that down to a variable parameter. In other words, it is just about impossible to prove this thesis wrong.
The parallels with language is quite evident, so given that it would be somewhat odd to accept this argument, why would we think that the same argument holds true for language?
As a final non-linguistic instance of (contingent) hierarchical structuring, let us consider the following scenario:
Imagine a freshman linguistics programme in which there are two courses taught in the second semester concurrently, with the course codes LING114 and LING110, each course being taught by two lecturers: one in the first block and one in the second.
Once again, I ask you to imagine the following set-up in the form of tree diagram. This would show two courses, further subdivided into the two blocks they are spread over, with the various course requirements the students need to fulfil in order to pass the course. Once again, it is necessary to point out that this tree structure would need to be much bigger, with many more nodes and further sub-divisions to accurately capture the various details of an actual course that was taught in this manner.
Let us consider what follows from this sort of set-up:
Notice that the person who teaches either course during block 1 is at liberty to do as he pleases for his block. He can add as many nodes as he likes (in fact, the number of nodes that can be added are theoretically infinite) by giving more tests, more assignments, projects, etc. Why is he at liberty to do so? Simply because he dominates the that node in the tree. What he does in his block, by the same logic, will also not affect what happens in the block 2 portion of the tree; this because it is governed not by B1, but by B2, and therefore has no say in that arena.
Also, let us say that the LING110 B1 lecturer has an urgent family crisis to attend to in England, and therefore will not be present to teach for most of his block. Upon consultation, the head of department decides to ask the B2 lecturer to just swap her course around with the B1 lecturer. After the necessary parties agree to this, and the necessary administrative work is sorted out, the B1 and B2 courses are duly swapped around. What would swapping the courses around entail? One of the obvious upshots of this kind of arrangement was that all B2’s powers, in terms of deciding his course content, assessment criteria, etc., cannot be somehow ‘left behind’ because he is changing positions on our tree. It is indeed logically possible that because B1 and B2 are now going to be transposed, some of B1’s decision-making powers may be left behind, giving B2 more power. Even though this seems like a somewhat outlandish speculation, the fact that it is a logical possibility, and universally does not occur, needs explaining. Why would it necessarily follow that if B1 and B2 are swapped around, both the B1 and the B2 lecturer would retain all their powers? Well, because it is more pragmatic to do this, as doing it any other way would cause unnecessary complications.
B1 and B2 cannot gain or lose any powers in virtue of the fact that they occur on the same level in the tree. If B1 were moved up the tree, to where LING110 stands, for example (in other words, if he became the course coordinator), he would take with him all the powers he had in his previous position, coupled with the additional power and responsibility that comes with being in a higher position. If it were the other way round, i.e. if the LING110 course coordinator had to swap with the B1 lecturer, he would have to dispense with some of his powers in virtue of the fact that he is now moving down on our tree structure.
This latter principle cannot fail to remind one of the “pied-piping principle” in syntax, which is defined by Andrew Radford as
a process by which a moved constituent drags with it all the properties,
along with its set of features, when it moves to another position on the
same level.
The stipulation regarding items moving only on the same level is unique to syntax since an NP obviously cannot somehow be ‘promoted’ to the S node in the same way it would be in our linguistics tree structure.
The point, however, is that the principles which are involved in cases like this have nothing to do with elaborate abstractions of innate knowledge, but follow from the fact of hierarchical structuring. To take this point further, let us assume someone notices that this phenomenon applies to all university courses taught all over the world, and say it matters enough to him to try and explain this seemingly universal characteristic. How amenable would we be to his explanation along the following lines?
Logically speaking, it is possible that this kind of “course structure” need not occur; it would not be contradictory in any manner to imagine a course that does not make use of this, or to imagine a school where the teachers are not categorised into different levels, etc., so the fact that no school, or no university course, makes use of linear organisation is something that ought to be in need of explanation.
Once again, the oddity of this argument accentuates the fact that it actually does not need any explanation, which shows that hierarchical structuring is actually something quite natural; something that we would expect to occur.
One could object to this argument by saying that the structure of language is something unconscious, whereas the instances mentioned above are a result of conscious decision-making. Let us consider this objection for a moment. Did we really decide that everything must be structure-dependent? If we look around us, and consider the nature of the world, we would notice that just about everything is, in some way or the other, hierarchically arranged; organisms, physical structures, societies, institutions, etc., all make use of hierarchical structuring. Even amongst groups of friends, sooner or later it comes to be that some members, or one member, is given leadership status, and the others fall into some kind of rank. It is hard to imagine that some time back, at the dawn of the human era, people made a conscious decision to start arranging things, like (the nuclear family structure) in a hierarchical manner. Given that nature evolved that way, why should it be so surprising that social institutions naturally evolved that way? And that when the human species evolved more acute cognitive structures, our knowledge (linguistic and otherwise), also evolved in a hierarchical fashion? If social institutions did not evolve naturally, then there should have been a time in human history when societies quite comfortably dispensed with ranks of this sort. As far as we know, this has never been the case. Some of the most ancient texts known to mankind tell us that society cannot be ‘rankless’; we have to categorise its members using some kind of hierarchy (cf. Plato’s Republic and the ancient Hindu Scripture The Bhagavad Gita, for example, which explains how a utopian society ought to be ranked). Maybe the texts are inaccurate, but then we need to ask ourselves what other evidence there is. Certainly in the modern-day world it would be hard to even imagine what this would be like. A rankless society, or a school where the staff members are all ‘equal’ is certainly a “logical possibility” (in much the same way a language that is structure-independent is logically possible), but its universal non-occurrence needs explaining. I contend that societies, like organisms and languages, naturally evolved this way because natural selection dictated that it just makes more evolutionary sense for things to adhere to a structure of this kind. It is true that we can consciously design school systems (and we can consciously design languages too; otherwise there would be no such thing as artificial languages and computer languages, for example), but all systems that evolve naturally have to adhere to this kind of structuring. School systems did not start off as homogenous masses of teachers and students, but must have started with some kind of ranking, which became formalised later on. It is only by deliberate intervention that, say, a school system can dispense with this kind of arrangement, and then too it would be difficult imagine how it would function effectively. Likewise, it is only by deliberate intervention that one can construct a language that does not make use of the principles of hierarchical inclusion relationships. If languages are allowed to ‘evolve naturally’, they would also have to adhere to this sort of structuring.
Nevertheless, if we are prepared to state that hierarchical structuring is one of the most important features that languages have in common, let us consider the following scenario:
Imagine a botanist is particularly interested in the study of weeping willows, and wants to write a paper on “What weeping willows have in common”. When he presents the thesis of his paper to his fellow botanists, they ask him what he has found that makes all willows common. “Well”, he replies, “they are all hierarchically structured.”
What is so odd about this scenario is that this so-called novel claim is actually quite a weak truism. We would expect something more concrete, like the number of leaves produced on average, etc. To extend our parallel to language, when trying to see what it is that languages have in common, we would expect more concrete examples, instead of something so abstract. It is not for lack of trying, as the nativists have been trying to pin these down for a long time, without much success (cf. examples discussed earlier).
Hitherto, the only concrete universal that seems to have survived refutation is that “all languages have nouns and verbs”, to quote Sampson. Surely this is not because we are innately predisposed to have these categories (whilst the others have to be learned), but because every human society needs to name things, and describe actions.
As an aside, Pinker gives numerous concrete instances of principles that follow from structure-dependency, to which Sampson duly presents blatant counter-examples. When I presented this counter-example at a conference, organised by the South African Applied Linguistics Association, I was told by a certain pro-nativist syntactician, Dr Jochen Zeller, now the head of Linguistics at the University of Kwazulu Natal, that whether this is a true counter-example or not depends on one’s representation. Hence, even though one may present a counter-example which is obviously grammatical, yet contrary to what an innate constraint would allow, the evidence can simply be ignored in light of the fact that one can “represent” the structure in a different way; obviously one which fits your theory. The sceptic will not concede that there are independent reasons for, say, adopting X-bar theory, and Pinker gives “no grounds for assuming that the universals posited have extra-theoretical reality”, to quote Cowley. Unfortunately, many of the nativists argue along unscientific terms like this, which makes it quite difficult to have any kind of intellectually meaningful debate with them.
Also, many other facts about language follow from the fact that languages are hierarchically structured. The fact that we do not make use of rules like “To form questions reverse the order of words”, as Pinker would say, follows from the fact that a sentence is not a linear string of words, but an organised tree-structure; the obvious upshot of this being that we can only manipulate the sentence in a particular way. The same reasoning applies to the fact that we cannot use the “middle” of something to form questions: “To turn a statement into a question, take the word in the middle of the sentence, or, if it has an even number of words, the word closest to the left of the middle”. Again, this is because a sentence is not a linear string of words. It is an organised tree-structure. Once we understand this, it is easy to see why picking out the middle car from a string of cars is not the same as picking out the middle word from a sentence, and more importantly, we also need to appreciate the fact that we are not necessarily obliged to accept that this is something specific to language that is hard-wired into our brains. Of course the human mind is responsible for arranging things in a hierarchical manner, but it is the same human brain that insists upon hierarchical organisation for just about everything else.
Why, then, should we believe that this is something unique to language?
In fact, to even think that it makes sense to ask why it is we do not have a rule like “to form a sentence reverse the order of words” is somewhat odd. This is certainly a logical possibility, but only in the very strict sense in that there is no contradiction in terms in asking that question. The fact of the matter is that it is practically impossible to manipulate a sentence in that way because the words we see on the piece of paper are not the psychologically relevant parts. The parts lower down on the tree cannot be taken into consideration when doing something that requires higher levels of cognition. To illustrate, imagine a friend comes over to your house for dinner (say you have a fairly new house), and he asks you for directions to the kitchen. Tacitly, what he is actually saying is that he wants you to tell him what he needs to do in order to get to the kitchen. Imagine you gave him directions along these lines: your brain needs to send a signal to your relevant muscles, and you need to wake up. Then you need to face the door, and after you begin to walk forward, continue until you get to a green door. When you get to the green door, turn the handle…and so on.
Now, after you give him directions, and he eventually follows it to his destination, you join him in the kitchen. There you ask him: Can you re-trace your steps back to where we were by reversing the directions I gave you? He most certainly can, from a logical point of view, but he will most certainly have a lot of trouble doing so.
What is odd about this? It is the fact that we are mixing high-level and low-level bits of information. Regarding the directions to the kitchen, the low-level bits are psychologically irrelevant (which leg goes first, the fact that you need to turn the handle, etc.), and what matters is the high-level bits: which passage to walk down, which door to turn in to…etc. Although it is logical to think otherwise, it just does not make practical sense.
We are doing exactly this when expecting the “reverse word-order rule” to be a viable alternative; we are mixing high-level (manipulation of a sentence) with low-level bits (manipulation of the actual words). The restriction, then, is not so much logical as psychological.
Philip Carr, in a paper he published in 2003, says that there is no empirical evidence offered to substantiate the claim that the child learning language is unable to exploit the capacities available in other domains like association, induction, conditioning, hypothesis formation and testing, generalisation, and so on. Nativists tend to hedge this obviously counter-factual claim by saying that there are some areas where this may play a restricted role, but in absence of any evidence, why should this be seen as anything other than a cop-out?
Carr then goes on to say that “One possible middle-way approach is to argue that analogical generalisation is available in syntax, but that the capacity for forming analogical generalisation is constrained by the Principle of Structure Dependence, so that sets of expressions are taken to be analogous only if they are structurally analogous. It then remains to be established whether Structure Dependence is truly an innately endowed, specifically linguistic, principle. But even if it is, the possibility arises that the domain-specific capacity to form analogical generalisations operates in tandem with that principle”.
As we have already seen, this is not something which is unique to language, and we have no reason to assume that structure dependency is innate.
Regarding the “Principles and Parameters” branch of UG, the following criticism can be levied:
Any putative universal that turns out to be untrue can simply be dismissed as “representing a new option on a parameter”, to quote Cowley again. This is certainly anything but scientific in the Popperian sense.
In addition to this, an explanation of language acquisition from a P&P framework necessitates the (ostensibly) innocuous simplification of instantaneity. This means that we should assume that the acquisition of language takes place in one fell swoop, as Chomsky puts it: from the initial state to the steady state, and that there is little variability thereafter. Chomsky believes this kind of simplification is not only innocuous, but also necessary for the purposes of meaningful scientific inquiry, in much the same way as one would, say, assume a runway to be frictionless when doing some calculations in kinetics. This in itself is problematic, in light of the fact that it not only precludes certain versions of learning theory by fiat, it also simply ignores phenomena which could lead to certain insightful revelations apropos language acquisition.
This approach also makes false predictions. For example, a phrase-structure rule like
PP -> P NP*
would predict that this particular language makes use of prepositions only, whereas another language may have a rule like
PP -> NP* P
which would make use of postpositions only. For these reasons, Chomsky would say that this strand of his argument should rather forcefully entice us to concede that we are endowed with a profound abstract structure that is genetically inherited. However, there are languages that have both prepositions and postpositions, like Norwegian, as I mentioned earlier. Pinker says that if these parameters were a matter of binarity, it could help explain the mystery of language acquisition, but if there are other instances like this, then the mystery is still a mystery, and the “principles and parameters” approach does not suffice as an explanation. In fact, many of the principles that are postulated as such turn out not to be principles as such, but contingent properties that happen to occur in many languages. Even (dialectal) native English speakers have apparent non-English constructions, like verb-final sentences.
Chomsky does admit that research in this area is still quite rudimentary, and that we are not sure how accurate this way of viewing things is. Hence, nativism, insofar as it depends on “principles and parameters” as an explanation, is not tenable. It may also be worth noting that Chomsky uses this theory to try and explain why languages all over the world are so variable. If this approach is shown to be flawed, then the nativists are left wanting with regards to explaining how it is that an alleged instinct is so variable, in much the same way as marriage customs are variable from culture to culture.
Also, if intuitions regarding grammaticality are to be explained by the violations of these P&P’s, then contrasting intuitions (even amongst native speakers) regarding grammaticality need to be explained. Carr points out the same problem I always encounter when I ask a group of students to make judgements on grammaticality; for example, when shown an “ungrammatical” sentence, “there is always a substantial minority of native speakers who find such expressions well-formed”. Once again, a better explanation is to be found in a paper by my former mentor Stefan Ploch on Link Phonology, published in 2001, who debunks the grammaticality hypothesis (GH) by explaining it in terms of attestation. The GH can be defined as the idea that a native speaker can take any well-formed input and judge whether it is grammatical or not. Since this is not true, it is evident that attestation provides a better explanation for this phenomenon. Sampson also opens his book, EMPIRICAL LINGUISTICS, explaining why the GH is methodologically flawed.
Response to “Descartes’ problem”
At numerous points, Chomsky says things like:
Descartes himself devoted little attention to language and
I will make no attempt to characterise Cartesian Linguistics
as it saw itself.
And a little later on, he points out that it is indeed a well-known fact that “Descartes made only scant reference to language”.
In fact, the phrase “Cartesian Linguistics” was actually coined by Chomsky, and from the above-mentioned quotes, one once again wonders why Chomsky would attribute this position to Descartes, and, on a related point, why Chomsky would invoke his name in the first place.
It is indeed well known in philosophical circles that Descartes did not place much emphasis on the apparent problem Chomsky would like to ascribe to him. In his Meditations, Descartes was actually trying to make another point altogether. He was chiefly concerned with showing that humans are the only ones who possess a soul, and he even takes it further by saying that non-humans are not only lacking in a soul, but are also no more important than mere automata. Descartes wanted us to share his belief that a computer and a dog are actually equal in God’s eyes because they both do not possess a soul. One of his arguments leading to that conclusion was the fact that humans have language, and animals do not. Even though a parrot can utter what we hear as words, a computer can also be programmed to do so; with the parrot as with the computer, there is no understanding taking place. That was what Descartes was arguing, and we can see that this had nothing to do with creativity or how it lends more viability to linguistic nativism. (Descartes was actually building up to a rather odd theistic argument purportedly proving the existence of God. His conclusion is that God must exist because otherwise we would not be able to cognise his existence. We are only able to do so because God wills us to! Most philosophers are actually not convinced by his argument, but the details are not relevant here.)
George Lakoff also concurs that there has never been such a thing as ‘Cartesian linguistics’. Chomsky claims that Cartesian rationalism gave birth to a linguistic theory like transformational grammar in its essential respects. He bases this claim on Arnauld’s and Lancelot’s Grammaire Générale et Raisonné, published in 1660. This was a series of grammars by Lancelot, the most extensive of which was his Latin grammar. Chomsky “never checked out his Latin grammar”, but Robin Lakoff did, and she published her findings in the journal Language. What R. Lakoff discovered was that in the introduction Lancelot credited all of his findings to Sanctius, a Spanish grammarian, whose work antedated Descartes by half a century. If this is true, then what Chomsky called “Cartesian Linguistics” had nothing to do with Descartes, but came directly from an earlier Spanish tradition. In Lakoff’s words, what is “equally embarrassing” for Chomsky is the fact that the theories of Sanctius, and the Port Royal grammarians generally, differ from the theory of transformational grammar in a crucial way: they do not acknowledge the existence of a syntactic deep structure, but assume that syntax is based on meaning and thought. This is a position which Chomsky vehemently opposes.
Aside from this, using the fact that language makes infinite use of finite media as evidence for innateness is not logical in any case. According to Lakoff, Chomsky’s use of the word ‘creative’ is “very strange”. There is nothing in transformational grammar that accounts for human creativity or that even pretends to. All that transformational grammar does, is provide a recursive mechanism for generating sentences. There is nothing ‘creative’ about this, since in principle it is not any different from constructing a computer program to do mathematics. The program could perform an infinity of mathematical operations, “but no one would say that it accounted for mathematical creativity”.
This is evident since we can easily think of instances where finite media are used to produce what could be of a potentially infinite variety. For example, the number of letters in the English alphabet is twenty-six, yet these can be manipulated and combined in an infinite variety of ways. The number of keys playable on a keyboard is also finite, yet these keys can be combined to play a potentially infinite variety of songs. As a young child I used to wonder whether song-writers are going to have a problem sometime in future because all the songs would already have been written, but once the fact that song-writing also makes infinite use of finite media was understood, this is no longer a concern. (Chomsky actually does believe that, for example, the number of art forms are limited, and may indeed one day have reached “saturation”!)
A musician who has been playing the piano for a while, and is rather impressively piano-literate in terms of being able to distinguish between the various extant genres, would be able to tell which genre a given piece of music belongs to. Sometimes he may not be sure for whatever reason, but chances are he would be able to accurately categorise various piano pieces based on the ‘rules’ that go with composing a particular type of musical score. Note that despite being rule-bound, in a sense, the composer can still produce a potentially infinite variety of musical scores, and recognise a particular score as part of an accepted genre, part of another genre, or not like any genre accepted according to the current standards of musicianship.
This is not unlike the native speaker’s ability to tell apart grammatical sentences, ungrammatical sentences, and sentences which they are not quite sure of. The only reason our pianist-composer is so skilled at telling the various genres apart is because he knows the rules that need to be complied with in order to fit into a particular paradigm. Likewise, speakers of a given language know the rules of their language, and they use this knowledge as a criterion to judge whether or not sentences fit in with their idiolect (cf. below, where I discuss Ploch and Jensen’s notion of “attestation” in more detail, which provides an alternative explanation for Chomsky’s ‘grammaticality hypothesis’).
Now, what has this got to do with being born with an elaborate linguistic structure hard-wired into our genotype? If we are not inclined to attribute specific innate knowledge to our pianist-composer, why attribute similar knowledge when it comes to language since these cases are actually the same in principle? One could object to this criticism by claiming that the rules of writing, piano-playing, etc. are consciously taught, whereas the rules of spoken language are not taught.
Whether rules are taught or not is itself a claim that should not be controversial. Sampson, for example, argues that parent-child interaction is analogous to a graded series of language lessons. Regardless, if we claim that rules are not language rules are not learnt, but part of our biological endowment, then why do different languages have different rules, and why do they differ even from dialect to dialect? The “principles and parameters” strand of Chomsky’s theory was postulated to answer precisely this objection, and as we have seen above, the “principles and parameters” theory is flawed. Before accepting the claim that actual linguistic rules are innate, we would have to ask what other evidence there is to lead us to accept this conclusion. Surely the burden of proof lies on the one making the claim, and until such evidence is forthcoming, we are not obliged to accept that rules are not formed from interaction with the environment, but that we are born knowing all the rules of language.
This strand of the argument, then, also breaks down.
Response to Pinker’s The Language Instinct
Phonotactic Constraints
The claim that English has rules of this kind seems quite false. I am told that there is an electric blanket brand in the UK called “fnug”. Sampson tells us that there is a houseplant called “vriesia”, and that he knows an Englishman by the name of Srawley, who sees his name as quite a natural, English name. I know of at least three adult monolingual English speakers (all of whom were privileged to have a tertiary education) who until very recently thought that the word “gnome” was pronounced as it was spelt. There is an L1 speaker of English who, until recently resided in an ashram in the south of Johannesburg – having passed away a while back. She happened to be of Scottish descent, and attended Victoria Girls’ High School in Cape Town. This is relevant because her school happened to be a very upper-class school for “whites only”. Long into her adulthood, and therefore long past her putative ‘critical period’, she joined a Hindu monastery, and as a result, learnt to say things like [k∧rma], and says it quite naturally. The “gnome” and “karma” instances are interesting because the former has an illicit onset, whereas the latter has an illicit post-vocalic r.
I myself am a monolingual speaker of English. However, I do not aspirate my (syllable initial) bilabial stops. Hence, I found my freshman phonetics classes very frustrating, when John Rennison would insist that if we put our hands in front of our mouths, we would FEEL the difference between pin and spin, where the latter is supposed to make us manifest a burst of air. Simon Donnelly was just as confused when I pointed out that I do not ‘round my lips’ when I say the sh part of milkshake.
Their confusion stems from the fact that all this is meant to be “innate”, you see...
Of course it is true that English may be uncomfortable “in general” with sequences of this kind, but it makes more sense to explain this intuitive discomfort regarding certain phonotactic forms in terms of learning, exposure, and attestation, rather than innate machinery.
Infant Prodigies
It would indeed take too long to test for every possible variable when a child is trying to make sense of the world around him, like whether the person you are talking to has a higher social status than you, and whether the previous word happened to be ‘and’. But “if you innately know that these last two variables cannot be relevant, then you will be innately incapable of growing up as a speaker of Japanese and [Biblical] Hebrew, respectively”! In these languages, social status and the position of the word ‘and’ matter in these languages.
It looks like Sampson may have been unfair here, because one can say that he is just taking two examples which he knows languages make use of, presenting counter-examples as “fruitless possibilit[ies]”, and saying “See, Pinker is wrong.” Sampson’s point is simply this: how are we to know what counts as a “fruitless possibility”? How do we know that some language somewhere does not consider whether a statement was uttered indoors or outdoors as a grammatical variable? Of the thousands of languages that are found in the world, how do we know? Given that all sorts of seemingly strange variables matter, like the isihlonipho sabafazi practice amongst tribal Xhosa speakers, according to which the syllables in the names of the male members of a woman’s husband’s family have to be avoided when they are in each other’s company. Even if a language were found that makes use of the indoor-outdoor variable, the nativists would just say, ‘oh’, and try to pin other things down as “fruitless possibilities”, whilst maintaining that we have an innate disposition to know not to look for these things. If we are to continue to assert that there is this kind of innate restriction, surely it ought to be because the evidence points us in that direction; my point here is simply that the evidence does not point in that direction.
Regarding the claims mentioned in the previous section made by Karen Stromswold, who said that children never make mistakes by way of false analogy like:
I like → He likes; I can → *He cans
because our innate knowledge precludes that possibility. If this were true, then how is the following construction, taken from sixteenth century Shakespearean English, possible:
I like → Thou likest; I can → Thou canst
Perhaps Shakespearean English was not spoken in this way, perhaps this is just a written convention. The fact of the matter is we could never know for sure. My point in mentioning this is simply to accentuate the fact that a construction of this kind is not as logically impossible as Pinker would like it to be. There may also be many modern-day languages that also work like this. A representative sample has not been adequately analysed. If a language like German is shown to do exactly what Stromswold predicts to be impossible, what does that show? It would show that there is no innate restriction of that kind. Indeed the crux of this thesis centres around the fact that every specifically innate constraint of this kind applies usually only to English, which means it cannot be part of our biological endowment.
A more plausible explanation would be that some aspects of grammar seem to be easier to learn than others, but it is “not obvious what distinguishes the easier bits from the hard bits”, quoting Sampson. What is evident is that it is far less plausible to explain phenomena of this kind by appealing to innate restrictions.
Second Language Learning and the Critical Period
The ‘critical period’ thesis has already been discussded earlier, and it seems more obviously false in light of the fact that there is nothing obvious that precludes someone past the putative critical period from learning language with equal proficiency to that of a native speaker. (Let us ignore the fact that neither Chomsky nor Pinker gives us any meaningful criterion as to what actually constitutes “proficient”).
How do we explain the fact that there was a student I once taught who did linguistics 114, who started learning French at the age of seventeen (in high school). Upon hearing that she was selected as an exchange student to go to France for a year, she started attending French courses, reading books on French, etc. By the time she had to go overseas, she was speaking French rather fluently, so much so that most Frenchmen she met did mistake her for a fellow citizen. If the facts are exactly as she explained them to me (I of course can only assume veracity), then Pinker is wrong in drawing the conclusion that he does. Newport and Johnson must have had a rather idiosyncratic notion of “motivation”, and either way they fail to explain the facts.
Hence, Pinker’s spin on this variation of the critical period is as unamenable to the facts as Chomsky’s.
The Language Mutants from Essex
A group of medical practitioners did a follow-up study on these so-called “language mutants” mentioned by Pinker. Firstly, they found that there is no consensus that SLI even exists, as those afflicted with it were everything but normal in other cognitive domains. Their IQ scores were substantially below the average, indicating that it is a more general impairment; this fact is not mentioned by Pinker.
Carr quotes from a subsequent work where he gives further details of Pinker’s inaccuracy. For example, “they do not cite the average verbal IQ of 75.” He also mentions that in Neil Smith’s pro-Chomskyan review, he “acknowledges that there are alternative interpretations of the data”. Once again, this is a cop-out, and Carr rightly points out, this is not an issue of interpretation, but whether SLI sufferers do or do not exhibit low IQ’s, and the aforementioned paper provides the relevant evidence. It might be normal to argue about the interpretation of data, but it certainly is “not normal practice to simply fail to discuss falsifying data”.
He then concludes, on a more general note, that “To the extent that Chomskyans fail to mention or discuss empirical evidence bearing on their claims, they do not deserve to have taken seriously their claim to be engaged in scientific investigation”.
This clearly shows that this deficiency was not unique to language, but instead indicative of a more general impairment. In fact, “it is not even clear that the term is more than a vague syndrome-type cover term for a range of deficits”, and “we are thereby forced to the general conclusion that Specific Language Impairment is not very specific” (quoting Carr again).
The Structure of Words
Firstly, regarding the restriction on inflexional-derivational suffixes, David Crystal expressed the preference of drawing a distinction between a “linguist” and a “linguistician”, such that a “linguist” should refer to a student of the discipline linguistics, whereas a linguistician should refer to someone who is involved in the discipline professionally. If someone had to encounter this distinction without any knowledge of the conventional usage of the term “linguist”, one would accept the terminology without any qualms, especially since a new-comer to the discipline is bombarded with so much jargon that he would expect to come across hitherto unfamiliar words. This example is a clear instance of what Pinker insists should “sound ridiculous” – we have an innate restriction which states that a derivational suffix like –ian can only attach itself to roots, and never to stems. According to Pinker, linguistician should also be an impossible construction which “sounds ridiculous”, since –ian attaches to the non-root stem linguist + ic. Would someone reading this accept it as an English word, or perceive it as something which violates our innate knowledge of word structure? Crystal certainly had no qualms about using it, and when I first read it, it did not even strike me as odd.
Later Pinker says that the restriction is that derivational suffixes cannot attach themselves to inflected roots. This rules out the counter-examples just mentioned, but Sampson refers to a film called Heathers, about a group of girls all called Heather, and asks us to imagine this film having some sort of cult following. Would there be anything wrong in describing this as Heathersian? Or think of a hypothetical spoof of the film, in the same way one could see Austin Powers as having a rather James Bondian paradigm, one could see this spoof as having a… Heathersian paradigm. I do not think this sounds ridiculous at all.
It might be true that we as English speakers prefer not to add derivational suffixes to inflexional suffixes, but there is a much better explanation than appealing to instinctive knowledge of word structure – one that provides an explanation based on culture and history. The –ian suffix is Latin, and words that use this suffix in English are taken from Latin. Later, –ian words were coined that had not existed in Latin, but the people who coined these words tried to make them sound as Latin as possible. Thinking in terms of Latin, it is not clear what it would mean to add an –ian suffix to a word inflected with
–s. Hence, there was no possibility of this occurring, given that they were using Latin as their model. Now things have changed. Sampson says that when he went to university in the 1960’s, he had to provide proof of having passed a substantial Latin exam, otherwise his application would not have even been considered. When my father was in school, he was forced to take Latin right up to matric. Now, Latin is no longer seen as the language of the educated. In light of this fact, it seems viable that more words like Heathersian would be accepted, as they would not have to be based on the Latin model.
Pinker happens to be writing at a time when the study of classical languages is fairly recent, yet obsolete enough for most people to not know the exact reason why there are these regularities amongst –ian words in English. Hence, it seems quite evident that this is not the case in this instance.
Regarding the way compounds work in terms of headedness and percolation conduits, Pinker uses this explanation to apply to more than one phenomenon. He talks about how irregular verbs trigger “blocking” mechanisms which preclude the application of the rule. The way I understand it, this predicts that either a rule will be generated, or it will be blocked; if it is blocked, then it cannot be generated. For example, consider the incorrect past tense of the verb sing. This will be perceived as a morphologically complex word: sing+ed. When a child encounters the correct past tense, sang, which now has to be rote-learned as a listeme, it “blocks” the application of the “add -ed” rule to this word, since sang and singed both have the same semantic content. Therefore, whereas the past tense may have been taken to be singed, it will now be over-ridden (or blocked) by the new word.
However, this does not happen. There is even a clear correlation between exposure to the irregular form and correct usage of it. Instead of saying that we have a “blocking mechanism” which precludes rules from applying, it makes far more sense to say that this is a result of exposure to stimuli from the environment, which, for example, can be explained by the notion of attestation (Ploch, 2001). Attestation refers to the fact that any form can be tested by a native speaker, and that phonological forms can be added and subtracted from the speaker’s set of attestations. Forms are “more normal” if they are retrieved from long-term memory and put into short-term memory. If this is not done, then the form will be “less normal”; this may occur either because the form is not retrieved, not retrieved properly, or because the form may be entirely foreign altogether. This explanation, besides sounding more plausible from an a fortiori point of view, is also more in consonance with common sense. It seems rather unlikely that this could be explained by a binary mechanism that either triggers a rule or not; to interpret the phenomenon in this way seems incoherent.
Also, with regards to headed compounds, the same criticism applies. Pinker explains this phenomenon in terms of “percolation conduits”, which either allow a rule to “filter” to the top of the morphological node or not. The determining factor here is that of headedness. If the compound has a head, then the rules that apply to that head percolate up the morphological tree and apply to the compound as a whole. If it is headless, then the default rule applies to the compound. How, then, does one account for the fact that some native speakers may say either at different times, and sometimes not even be sure of which one to use? To explain this, Pinker now says it is a matter of exposure, and whether the speaker perceives the compound as having a head or not; and some may not be sure! Surely speaking about innate “percolation conduits” and blocking mechanisms is far less plausible than saying that we employ a more general learning strategy. Clearly then, we are once again not obliged to take this as support for nativism.
The third instance may seem like a more convincing case, but upon analysis is just as vacuous. Despite having tried to do everything to save Gordon’s hypothesis from refutation, it seems like there is still a problem with his claim. The Brown and LOB corpus, features a natural-speech sentence which reads as follows,
…the smaller European carriers, who have in the past been strong
opponents of fares-cutting airlines…
The compound we see here is fares-cutting, which is a compound which has an analytically complex word as the first part of the word; this is precisely what Pinker and Gordon would predict to not be possible. According to Sampson, there are more instances of this usage in the UK than in the USA, where Gordon would have been working from, but this only accentuates the fact that it cannot be an innate restriction which guides the formation of these compounds, unless we are prepared to concede that there is some viable explanation regarding this aberration.
In addition to this quite blatant refutation, the validity of Gordon’s experiment can be questioned for other reasons. Firstly, his sample was based on a fairly homogenous group of children, all of whom were L1 English speakers, and all of whom were Americans. To draw the conclusion that these restrictions are universal to all human beings may be a bit premature. The reliability of this experiment may also be questioned, since there were also no follow-up tests done; it is at least possible that six months later some of the children in his sample may find the word “rats-eater” a well-attested form.
On the Origin of Language
Pinker aims to show that we possess a database embodying knowledge of X-bar syntax, traces, etc., and this must have evolved as a result of Darwinian natural selection. Evolutionary processes could have given rise to a language instinct, but Cowley
rightly tells us that “the very idea is muddled”. If this LAD is as technical and detailed as Pinker would require it to be, we would have to ask how an N-bar could count as useful both for the brain and in terms of evolutionary fitness. Why would “networks that represent formal categories” arise though natural selection? As will be mentioned below, Pinker also plays down the fact that evolution works by selecting inheritable variations. As will also be explained below, he “treats natural selection as if it were goal driven”, which is blatantly incorrect.
Mark Turner, in the book I referred to earlier, THE LITERARY MIND, proposes a rather interesting alternative explanation to the origin of language, an explanation which runs counter to the widely accepted Chomskyan claim that it is a result of genetic mutation (or natural selection for Pinker); the spontaneous development of a “grammar organ”. Before we go on to explain Turner’s theory, we would first need to understand some of the basic premises upon which his theory is based.
Chomsky was very emphatic on the point that syntax must necessarily be one the cognitive primes of the human mind. By that he means that syntax is primary in the sense that it takes precedence over other levels of language structure. In fact, Chomsky even goes so far as to call us “syntactic creatures”, arguing along Aristotelian lines. Turner does not agree with this interpretation. In fact, he is not the only one who proposes that syntax cannot be as primary as Chomsky claims it to be. The psycholinguist, Joseph Kess, for example, explains the fact that when we recall something which was said to us, we tend to focus on getting the meaning across instead of the syntactic structure. In fact, it is almost never the case that we reproduce the exact syntax when relaying information. This shows that semantics takes precedence over syntax, a standpoint to which Chomsky has always been vehemently opposed. As an aside, this is one of the reasons George Lakoff divorced himself from the Chomskyan school in the first place, and Pinker also acknowledges this fact in his latest book, THE STUFF OF THOUGHT.
Likewise, Turner does not endorse Chomsky’s hypothesis, but instead suggests that parable constitutes the primary unit of language – not so much as a structural precedent, but certainly as a temporal and, more importantly, a cognitive one.
Turner defines parable as “the projection of story”. This is a much broader definition than the one conventionally used in the English language, where the word it refers to an allegory of some kind in order to illustrate a moral, and much narrower than its Latin-Greek origin, which refers to a comparison of some kind. Turner uses his definition to refer to a more general instrument of everyday thought that shows up everywhere, “from telling time to reading Proust”.
Turner wishes to address a fundamental misconception in the minds of the general educated public, viz that the everyday mind has little to do with literature. Although literary texts may be creations of art, the instruments of thought used to invent and interpret them are basic to everyday thought. The mental instrument which Turner refers to as “narrative story” is something which is actually basic to human thinking.
The bulk of Turner’s book is in fact dedicated to illustrating how the processes we go through when analyzing parables, are exactly those we use in our everyday lives. For example, the use of mixed metaphors, which Turner refers to as “blending”, is explained as confounding two disparate things, be it physical attributes, mutually exclusive character traits, etc. For example, Turner tells of a story where a donkey concocts a rather sophisticated plan to help his friend the ox. Both the ox and donkey discuss this plan in detail before it is implemented. The details of the actual story are not important. The point is that there is an element of anthropomorphism here, since these otherwise “a-linguistic” creatures are speaking rather eloquently. This is contradictory, since animals cannot communicate with that degree of complexity. Also, donkeys almost certainly do not have the cognitive capacity to plot and plan with any degree of complexity, nor do oxen for that matter; this is also contradictory. Hence, we call this process blending, since we are combining otherwise disparate traits.
This example may be literary in nature, but Turner’s point is that the processes involved in understanding the subtleties of this parable are something everybody goes through everyday. To understand a sentence like
Roosevelt accomplished a great deal in his first one hundred days,
but Clinton has accomplished by comparison little
we must build two mental spaces and an intricate comparison between them. We can understand this as a conceptual metaphor according to which accomplishment is travelling along a path; hence, we could understand this more clearly by paraphrasing the above statement as
Roosevelt covered a lot of ground during his first one hundred days,
and Clinton covered comparatively little during his.
So in one blended space, Roosevelt is moving along a path where there are locations, and reaching this location is considered the accomplishment of a goal. In the other blended space, Clinton is starting to move along a similar path, just at a slower rate and reaching fewer locations.
When we say, then, that Clinton should have hit the ground running, we are saying that Clinton should have accomplished in his first one hundred days what Roosevelt achieved in his. Hence, Clinton has failed to “keep pace with” Roosevelt. Now, to “keep pace with” requires a conventional blend that has both presidents competing along the same track, but this is something that we would construct and use entirely unconsciously. We can force this blend into consciousness, so to speak, by saying that “Clinton was in a race with the ghost of Roosevelt”. But we also know unconsciously that this is an imposed construct, and therefore there is a temporal constraint, so there would be something anomalous in saying that “So far, Roosevelt has succeeded in keeping well ahead of Clinton”. The problem with this is that we could only say something like this when it pertains to a real race, not a parabolic blend like this one. However, it is not even as simple as that, since we could actually envisage a possible world where a blend of this sort is viable. Imagine there was a passage in Roosevelt’s diary in which he claims to set a record accomplishment during his first one hundred days of office; and part of his entry states that he will endeavour to set a record for all presidents, past and future. With these premises in mind, it now actually does make sense to say something like “So far, Roosevelt has succeeded in keeping well ahead of Clinton”.
Turner’s point here is that processes like blending are just that, cognitive processes which are not exclusively the product or the property of the literary mind, but of the everyday mind. Even contemporary neuroscience is acknowledging this fact, because at the most basic levels of perception and cognition, understanding and memory, “blending is fundamental”.
Scientific models of thought conventionally start with what they take to be basic, on the tacit assumption that scientists must do first things first. However, it is a well-accepted fact that common sense, or what we perceive with our rather limited senses, may not necessarily reflect reality. Leeching on to a common-sense paradigm is analogous to clinging to a Newtonian view of the cosmos. It is not implausible that the concepts behind such models are wrong. Indeed, even a rather perfunctory read through the various discoveries in the physical sciences would confirm this, especially the findings and subsequent implications of quantum mechanics.
It is not implausible, then, that something like imaginative blending and integration are basic; an explanation that cannot handle the dynamics involved in a child pointing to a balloon and saying This is my imagination dog, cannot hope to explain the most ‘basic’ concepts like dog, or the meaning of The dog has four legs.
What Turner refers to as the processes of The Literary Mind, are usually considered not only different, but secondary to the processes of the everyday mind. On the contrary, processes that we have always considered literary are at the foundation of the everyday mind. Literary processes like blending make the functioning of the everyday mind possible. Turner then goes on to use this idea to explain the origin of language.
It is evident that human beings have the mental capacities Turner refers to as “parable”. Now, considering Ockam’s Razor, is it necessary to add something new to parable in order to explain the linguistic mind? Do we need to make the additional hypothesis that special autonomous instructions arose in the genetic material for building an autonomous black box that does the entire job? If parable gives us what we need for a satisfactory explanation, the answer would have to be negative. Cognitive mechanisms whose existence we must grant independently of any analysis of grammar can account for the origin of grammar. The linguistic mind is “a consequence and subcategory of the literary mind”. Let us now consider Turner’s explanation of this hypothesis in more detail.
Stories have structure that human vocal sound per se does not have. Stories have objects and events, actors and movements, viewpoint and focus, image schemas and force dynamics, etc. Hence, parable takes structure from story, which is thereby superimposed on to voice; the structure it creates is what we call grammar, and sentences come from story by way of parable.
Parable draws on the full range of cognitive processes involved in story. This would include spatiality, motor capacities, our sense organs, perceptual categorization and other basic cognitive instruments. Parable draws on all of this structure to create structure for vocal sound. Grammar, built from this structure, coheres with it.
For Turner, grammar would have arisen in a community that already had parable. The members of that community used parable to project structure from story to create rudimentary structure for vocal sound.
Turner asks us to consider an analogy not unlike my example on p. 45, but to illustrate a different point:
Once again, imagine a community of people who have trained themselves in rudimentary martial arts, and suppose all members of the community have it. No genetic specialization in martial arts exist, but depends on pre-existing capacities like muscle control, balance, vision, etc. Once the skill is acquired, it seems like second nature.
Now suppose an infant is born into this community with a little genetic structure (strong bone structure, good vision, etc.) which would help direct these pre-existing capacities to learning this particular form of martial arts. The members train him as normal, but he has secret edge. And if the community is structured so as to confer reproductive advantage onto those who are more proficient at martial arts, then the community provides an environment of evolutionary adaptiveness for the genetic change. Here, each increment of further genetic specialization brings an increment of reproductive advantage. However, in this scenario, martial artistry arose without genetic specialization for it.
Along the same lines, now imagine a community of people who use parable to create rudimentary grammatical structure for vocal sound. Everyone in this community develops story and projection, has voice, receives training from his parents, and is assimilated into the work of creating grammar through parable. Now suppose a special infant is born with just a little genetic structure that helps him project story onto voice. The members (or his parents) train him as normal, but the child has a secret edge. If the community is structured so that greater facility with grammar confers reproductive advantage, then the community provides an environment of evolutionary adaptiveness for genetic change. This situation could plausibly give rise to a kind of genetic arms race in which each increment or further genetic specialization brings an increment in relative reproductive advantage. However, in this scenario, grammar arose without a genetic instruction for grammar – it arose by parable.
Let us consider a basic example of how story is projected to create grammar. Consider the small spatial story in which Chomsky throws a stone and Pinker flips a coin. These are both instances of the same basic abstract story.
This abstract story has certain kinds of structure. One kind of structure it possesses is distinction of certain elements. There is Chomsky, the act of throwing, and the perception of a stone; conceptually, we distinguish these three elements.
These elements have category structure. Chomsky, for example, would be placed in the category of animate agents; throwing is placed into the category of events; stone could be placed into the category of objects.
The story also has combinatorial structure. The distinguished elements of the story include Chomsky, the stone, the causal relationship between Chomsky and the movement of the stone, the event shape of the throwing, and our temporal viewpoint with respect to the throwing; all of which are combined as simultaneous: the act of throwing involves all of them at once. More importantly, this combination also has hierarchical structure, as it must in order to explain its universal occurrence.
So the abstract story has certain kinds of structure: reliable distinction of elements, distribution of elements into categories, simultaneous combination, hierarchy, etc. Other basic stories also show recursive structure: If Lakoff catches the stone that Chomsky threw, then one story (Chomsky throws a stone) feeds into a second story (Lakoff catches the stone).
Vocal sound per se does not have this structure. The elements of the story have structure, but the actual sound is merely a continuous stream which can be divided in any number of ways. The elements of the story also have reliable hierarchical structure. The elements are joined but conceptually distinct – this is something which the sound does not mirror. The elements of the story also have category structure which the sound does not mirror. If Chomsky throws, Pinker pushes, Jackendoff tosses, etc., then Chomsky, Pinker and Jackendoff all belong to a certain category in light of that, which is not at all evident from its phonological form. Also, the causal structure has nothing to do with the causal structure of the vocal sound. The temporal structure of the sound is always linear, whereas the temporal structure of this story involves highly complex simultaneity. So story and phonological form are two different things. Story structure is projected to create structure for vocal sound, which the latter does not intrinsically have.
The distinction of elements in the abstract story is projected to make “Chomsky”, “throws” and “a stone” precisely distinct not as sound but as grammatical elements. The category structure in the abstract story is projected to vocal sound to fit “Chomsky” and “a stone” into the same grammatical category: NP. Once again, the phonological form gives us no reliable way of putting them into categories. The different roles of Chomsky and the stone are projected to give these different grammatical relations (subject vs. object) and different semantic roles (agent vs. patient). The structure of temporal foci and viewpoint in the story is then projected to give the sentence tense.
For Turner, then, grammar is not the beginning point of language; parable is. Grammar arises from conceptual operations. Rudimentary grammar is a repertoire of related grammatical constructions established through parable. The backbone of any language consists of grammatical constructions that arise by projection from basic abstract stories.
Turner then goes on to explain how rudimentary language can be extended by other processes like metaphorical blending, resulting in the vast and intricate phenomenon we call language.
Turner’s view of language as arising through parable would allow for diversity of languages and grammars. Details of narrative structure vary from culture to culture and even person to person; what is universal is not all the specific narrative structures but rather stability of basic abstract stories. All cultures have stable repertoires of basic abstract stories. Some of them vary from culture to culture. Projection is widely variable in the actual structures it projects and the ways in which it projects them. Even when two different languages project the same basic abstract story, thereby giving rise to similar grammatical constructions, the details will often be strikingly different. In Latin, word-order does not matter; in English it does. Yet in both languages narrative relations project to create grammatical structure; it just so happens that in Latin it is a grammatical structure of case-endings in nouns, whereas in English the result is a grammatical structure of word-order.
This explains the diversity and universality of linguistic phenomena in a way that the Principles and Parameters theory fails to. Turner’s hypothesis that grammar arose by parable is not at odds with Sampson’s theory. What puts this view at odds with the Chomskyan conception that grammar arose because there evolved, with no help from natural selection, and no help from pre-existing capacities, a species-specific, modular LAD that was specialized for grammar, just as the heart evolved from natural selection and became specialized for its functions.
From a Popperian point of view, a hypothesis of this sort is not tenable. It trades Ockam’s Razor for God’s Magic Hat. Against all odds, it makes the most cosmic and all-embracing extra hypothesis imaginable so as to solve everything at once.
This is a problem which Pinker actually acknowledges. Hence, he tries and explain the origin of language by natural selection. This may seem prima facie to put Pinker even further away from Turner than Chomsky, but if we look carefully at their practical attempts to explain how natural selection could have produced a genetically coded mental organ, Pinker actually implicitly embraces an account of language origins in which grammar begins from meaning. Pinker & Bloom say: “Language is a complex system of many parts, each tailored to mapping a characteristic kind of semantic or pragmatic function onto a characteristic kind of symbol sequence.”
Mapping is the critical term and concept in this assertion. It is usually the critical concept in any explanation of grammar as “encoding” something else, “signalling” something else, “mapping” certain structures, etc. However, Pinker does not have anything to say about the role of mapping. Mapping (which is similar to what Turner calls “projection”) is a mental competence. It does not come for free in an explanation. It is actually the principle process to be explained. To speak of mapping is to make a theoretical commitment to a mental capacity for projection; Pinker and Bloom give various examples of this kind of mapping throughout their paper.
Their explanation depends upon the existence of a mental capacity to project one kind of structure onto something entirely different. Pinker assumes that the mental capacity for projection precedes grammar, but underplays the importance and complexity of this mental capacity by perfunctorily alluding to it; this mental capacity is what principally needs explaining in an account of grammar.
Chomsky’s argument is weak because it asks us to accept an almost inconceivably unlikely event in the absence of any evidence for that event. Pinker’s argument is weak in a different way: it skips briefly and vaguely over its central step. For natural selection to be responsible for the origin of grammar, we must have two events: some genetic structure must arise to result in a trait of minimal grammar, and this trait must occur in an environment in which it confers reproductive advantage. Of course, this environment cannot be one where a grammatical community already exists. This first event is not hard to imagine, even though there is no compelling evidence for it. Let us now consider the second step, regarding the conferring of reproductive advantage.
Parable is a basic capacity of human beings. In a community where grammar arose by parable, grammatical speech would be a highly useful cognitive and cultural art. In this community, language would certainly confer reproductive advantage, making genetic specialization for language adaptive. In Turner’s scenario, the origin of (rudimentary) grammar happens before any genetic specialization for grammar. The grammatical expressions produced by the lone first genetically grammatical person are parsed at least in part as grammar by members of the community who have no special genetic instruction for doing so, but use parable to do so.
This is where Pinker faces a challenge. What survival advantage would the first person with language have? We know that natural selection is blind, and is only concerned with the here and now). The scenario to which Pinker is committed is not viable in light of this, whereas Turner’s is. Grammatical processing might assist this lone person’s memory, reasoning, or imagination, and so be adaptive in this very indirect way. But Pinker is precluded from proffering an explanation along these lines since he advocates the notion that grammar is modular and therefore autonomous from other cognitive faculties. Pinker would also not be able to offer a scenario in which members of the community (who have not evolved the relevant genetic instruction) use parable, together with other cognitive processes, to recognize and parse the grammatical utterances of our hypothetical first lone linguistic being.
Pinker provides no alternative to the scenario in which the adaptiveness of genetic specialization for grammar depends upon the presence of a grammatical community. All of his speculations concerning reproductive advantage depend upon a community of speakers with rudimentary grammar. Hence, there is no way in which a solitary grammatical person would have a reproductive advantage in a community whose other members are a-linguistic. The utterances directed at him would be a-grammatical; his own grammatical utterances would have no audience; communication would not be impossible, but there is no additional advantage conferred by the grammatical component.
Pinker is actually not unaware of this problem, and he tries and explain away the problem by suggesting that “One possible answer is that any such mutation is likely to be shared by individuals who are genetically related. Because much communication is among kin, a linguistic mutant will be understood by some relatives…”
This is nothing more than a cop-out, since it is not true that such a mutation is likely to be shared by individuals who are genetically related. Consider once again the first genetically grammatical person. None of his ancestors or siblings is genetically grammatical, ipso facto. If the genetic material expressed in his grammatical competence arose by mutation from his parents’ genetic material, it is very unlikely that his kinsmen would have that mutation too. If it arose because error-free sexual recombination of his parents’ genetic material finally put together the right package, then later siblings would receive quite different genetic packages, especially if they do not have the same parents (i.e., a different father or mother). There is also the important difficulty that even if a later sibling had the right package, the first genetically grammatical infant would nonetheless still live in “grammatical isolation”, so to speak. Conventional Mendelian genetics clearly state that mutations which have survival value are passed on to succeeding generations, not to siblings; in other words, mutations conducive to increased chances of survival are passed on vertically, not horizontally. According to Cowley, who concurs with Turner, what needs to be explained is what “reproductive benefits [are] accrued to individuals who…drop (or insert) pronouns”, and unless this is accounted for, it remains much more likely that talking co-evolved with “the neurophysiological mechanisms and cognitive skills typical of primate social behaviour”.
If evolution could think ahead, it would certainly see that producing a genetically grammatical community would be enormously useful to the members of that community; and if that were the case, Pinker’s theory would have some viability. But just about every exponent of Darwinian evolution agrees that evolution cannot think ahead. Chomsky is right in saying that we can tell the story as we like, but he is wrong in saying that we ought to not try and provide one, and that we should look at these “stories” in light of what our scientific paradigms and evaluate them logically. A good natural selection story for the origin of language must show that the first genetically grammatical human being had an immediate reproductive advantage. If he was born into a community whose members had, by virtue of parable, a minimal capacity for grammar, then the immediate reproductive advantage is obvious. Pinker is obliged to show immediate benefit to the first genetically grammatical person who is born into a community whose other members have no persons endowed with grammatical competence, and who lack the ability to recognize even the existence of grammar, let alone parse it.
It is my contention that Pinker has known his position to be untenable all along, and finally realised that he would have to radically revise his ideas if he were to maintain his standing in the scientific world. As mentioned, he tacitly rejects many of his previous ideas by embracing the fundamental tenets of Cognitive Linguistics. He also no longer refers to himself as a psycholinguist (in THE LANGUAGE INSTINCT, he specifically defined the word to refer to himself! – “...a psychologist by training , who studies how people understand, produce or learn language”), but actually explicitly says that he does not consider himself a linguist: he’s a cognitive psychologist. Calling himself a cognitive linguist would not be ... very convenient, and just about any linguist would know why.
Anyway, that’s a topic for another discussion.
My point was that the view that rudimentary grammar arises from parable provides an appropriate environment of evolutionary adaptiveness of the sort that any good natural selection story must have.
Admittedly, Turner’s account does not preclude the view that humans living now have some genetic specialization for grammar, but as already stated, this must be decided independently on empirical grounds.
Be that as it may, Turner’s story reverses the conventional Chomskyan view that out of syntactic phrase structures one builds up language. With story, projection, and their combination in parable, we have a cognitive basis from which language can originate. Language follows from these mental capacities as a consequence.
Converging evidence leads us to believe that claims about natural selection “provide no reason to hypothesise that neural networks embody special, language-specific algorithm-like forms”, according to Cowley.
Linguistic Relativity
This strand of the argument is dominated by both straw-man and ad hominem arguments in a way that the other arguments are not.
When I first learnt about Pinker’s version of this theory, even as a freshman, I felt very uncomfortable with what I heard as it was quite evident that there was some propaganda at play. When I ventured to ask Dr Norine Berenz, the lecturer, if Whorf really believed in determinism the way Pinker portrayed it, she said that she has not checked Whorf’s original works, but that I should, and then conceded that Pinker is certainly guilty of rhetoric bordering on propaganda. I made a resolution then to cross-check Pinker, and having done so, it clear that there is once again blatant misrepresentation.
Regarding Pinker’s ad hominem argumentation, Sapir is conceded to be a brilliant linguist, whilst Whorf is (merely) "an inspector for the Hartford Fire Insurance Companyand an amateur scholar of Native American languages" (my italics). One plausible reason why Pinker would think these facts relevant is to make us believe that Whorf’s academic reputation is questionable, and therefore we should not take him seriously. He then goes on to say that "the more you examine Whorf's arguments, the less sense they make", followed by an analysis of Whorf's "empty gasoline drums", an example which invents facts and includes information which Whorf did NOT himself state, such as the worker tossing a cigarette into an “empty” drum filled with gasoline [sic] because his language led him to believe he could do so without any harm. This is an inaccurate representation of what Whorf said because he only talked about behavior in general around the full and empty drums. Whorf merely pointed out, based on his experience as a fire insurance investigator, that calling drums “empty” led workers to believe that they were, which is why they would see nothing wrong with tossing a cigarette butt into them, and Whorf merely uses this fact as a starting point to ask the question: do behavioural patterns reflect linguistic patterns; when put this way, it certainly does not seem like an outlandishly vacuous endeavour. In a discussion of Hopi time, we find the most virulent attack of all: "No one is really sure how Whorf came up with his outlandish claims, but his limited, badly analyzed sample of Hopi speech and his long-time leanings toward mysticism must have contributed" (my italics). Once again, these “outlandish” claims are inaccurate, making Pinker’s claim outlandish, and mentioning Whorf’s alleged association with mysticism is irrelevant since we should be focusing on whether his actual arguments make sense empirically and logically, using an accepted scientific paradigm, like that of Popper’s. Whorf does point out that our own notions of flowing time and static space are “mystical” to the Hopi, in whose language time disappears and space is altered, so that it is no longer the homogeneous and instantaneous timeless space of our supposed intuition or of classical Newtonian mechanics. Whorf’s actual field work was not only with Hopi, but he also researched Mayan writing systems, and did semantic analyses on language data collected by other linguists. His examples are often from Hopi and Shawnee, and he observed that Hopi seems to have little reference to time, and then merely posed the question: could it be that different, yet equally valid, representations of time are possible? In fact, there is nothing outlandish about a claim of this sort, and it is something worth thinking about, since quantum physicists have been telling us for almost a century that our particular cultural notion of time is a linguistic construct, and anyone who reads popular science books, like Stephen Hawking’s A BRIEF HISTORY OF TIME would know this. Of course this fact does not make Whorf infallible, since everything is subject to empirical verification, but certainly dismissing Whorf’s thesis as outlandish prima facie should not be deemed acceptable in academia.
Pinker refers to “The linguistic determinism hypothesis closely linked to the names Edward Sapir and Benjamin Whorf"; and the use of passive voice is very telling here, since we cannot help but wonder who the deleted agent is. Either it is Pinker himself, or the camp of academics representing nativism, who essentially created, developed and promulgated this strawman “Linguistic Determinism” argument in the first place. Whorf did formulate a principle of linguistic relativity, but neither he nor Sapir ever formulated the Sapir-Whorf Hypothesis they are so indelibly linked with in Pinker’s book. Who is Pinker fighting here, when even Whorf and Whorfians agree that linguistic determinism is wrong? Actually, it is not implausible to assume that the distinction between Relativity and Determinism was something put forth by nativists like Pinker, since no distinction of that sort is to be found in the original writings of either Sapir or Whorf. Pinker does not understand that Whorf generally argued from a systems perspective inherited from quantum physics, where his arguments make sense, not from the Newtonian perspective, where monocausal determinist arguments make sense. In fact, Whorf would have been writing at a time when the implications of Einstein’s theories of relativity were just beginning to astound the academic world. Finally, let's examine Pinker's colour words and Hopi time arguments. In showing the obvious absurdity of language having anything whatsoever to do with influencing perception, he contrasts the way physicists and physiologists look at colour: while to the former colour is a continuous wavelength dimension without our familiar delineations, to the latter it is a matter of three kinds of cones in the eyes wired to neurons, etc. "No matter how influential language might be, it would seem preposterous to a physiologist that it could reach down into the retina and rewire the ganglion cells." Once again, this is a strawman argument: who actually said that language reaches down into the retina and rewires the ganglion cells? I e-mailed Pinker about the obvious absurdity of this claim, inviting him to respond with either some references or justification, but he never did. Granted, maybe he was too busy, or maybe he just did not want to reply, but it is worth noting that he replied quite readily to my earlier more sycophantic and adulatory e-mails! The very claim is so preposterous that one wonders why argumentation along these lines is still acceptable.
Friday, January 2, 2009
My views on the NATURE-NURTURE debate – Part 4
THE CASE FOR NATIVISM
Chomsky’s key arguments are as follows:
Speed of acquisition
Critical period
Poverty of input
Convergence amongst grammars
Language universals
Pinker bolsters Chomsky’s paradigm in many of his works, but The Language Instinct and Words and Rules will be considered in more detail than others since this is where the former outlines his case and shows us where he stand on this debate.
I will now explain each of these in turn:
Speed of acquisition
The fact that children acquire language so effortlessly and so quickly, despite its obvious complexity, shows that they must be innately disposed to do so. Other forms of learning are slow and protracted (cf. the freshman’s struggle in learning things like physics).
Stages of acquisition
Children learn language in the same stages all over the world, starting with the babbling stage, moving on to the one-word stage, then the two-word stage, then the telegraphic stage, and finally culminating in fluent speech.
Why would a child belonging to some tribe in the Amazon acquire his language in the same stages as those of an Italian child in Milan? In fact, children all over the world seem to follow the same pattern when learning a language. These stages occur in this order not because they are required to by the laws of logic, but because it is a matter of the unfolding of a genetic blueprint. If we do not accept this as the only viable explanation of the facts, we would either be forced to concede that we do not know, or have to postulate a rather surprising coincidence. If we accept that language is no more ‘learnt’ in this order any more than we “learn to have two arms”, it at least provides us with an intellectually satisfying explanation of the facts.
Critical period
Then there is the critical period thesis, which claims that there is a period in one’s life when one is particularly disposed to learning a language, after which language learning becomes so much more difficult. Chomsky insists that the “functioning of the language capacity is […] optimal at a certain ‘critical period’ of intellectual development”; just think of the mother-tongue English speaker trying to learn French at, say, university. According to Susan Curtiss, “there are two versions of Lenneberg’s critical period hypothesis”: the strong version and the weak version. The strong version says that it is not possible at all to acquire language naturally after puberty; the weak version says that one cannot acquire language with equal proficiency to that of a native speaker after puberty. As will be discussed in a subsequent section, the strong version is not testable in practice. Hence, the case of Genie is used to lend support to the viability of the weak version. (Nonetheless, I contend that both the strong and weak versions are flawed.)
Genie was the victim of an abusive father, who locked her up in a room upstairs. She was chained to a toilet seat, and not allowed to leave the room. They passed her food through a little flap at the bottom of the door. She literally had to live in a room for all her childhood. The neighbours reported something strange in the household, and after investigation, a social worker eventually discovered Genie.
Upon her discovery, the father shot himself dead, and her mother claimed that she was continually beaten into submission to follow his orders, and therefore could not do anything about her quandary.
Without going into the maudlin details, this story is relevant because when they discovered Genie she could do little more than feed herself. She walked in a very animal-like fashion, had no sense of decorum, often grunted as if it was normal, and she could not speak or understand language. After intensive therapy, she was able to control her outbursts, stopped grunting, learnt how to walk properly…etc. However, up till today, she has never learnt how to speak with the efficiency of a native speaker. Despite some of the world’s best psycholinguists and speech therapists working on her, she was only able to utter very basic, rudimentary sentences. The question is: why did she (inter alia) learn how to walk properly, but could not acquire language proficiently? The nativist’s answer: because she had passed her critical period for language acquisition.
Poverty of input
Chomsky claims that the input children get from their linguistic environment is far from ideal, and children are generally not corrected on grammatical issues - parents are more interested in things like ethics and truth-content. Even though overt correction is not necessary, it is still interesting that even young children never utter sentences which are syntactically deviant, even though they may over-generalise rules. Chomsky would say that this is because we have the syntax hard-wired into our minds.
The point is that children must infer general rules about language after hearing a few existential examples, despite not being told to do so, and despite the fact that even these examples are not ideal.
Aside from the anecdotal evidence that Chomsky presents, he always points out children’s knowledge of yes-no question formation in complex sentences. He explains that if we consider a sentence like
The boy who is at the shop is naughty
there are two ‘hypotheses’ a child could construct when trying to work out the rule, let us call these hypothesis 1 and hypothesis 2:
Move the first verb of the sentence to the beginning when forming a yes-no question
Move the verb that occurs in the matrix sentence (as opposed to the embedded clause) to the beginning
Of course, the correct rule would be hypothesis number 2, which would produce the grammatical sentence:
Is the boy who is at the shop naughty?
However, the point Chomsky is trying to make is that no child would ever utter a question using hypothesis number 1:
*Is the boy who at the shop is naughty?
If the child was going on what he heard from people around him speaking, there is nothing in the speech input that prevents him from distinguishing between the two hypotheses. So, given that the application of hypothesis number 1 is a logical possibility, and given that number 2 is also a logical possibility, what precludes the child from forming questions at least some of the time in the manner above, using hypothesis 1?
If the child is using the “fragmentary input” from the environment as a basis for his hypothesis formation, there is absolutely nothing that stops him from doing so. Logically, then, Chomsky would have us conclude that the child is not using the environment, but comes with his own innate linguistic structure hard-wired into his mind. This might seem prima facie implausible for someone using his common-sense, but how would you otherwise explain the above-mentioned facts?
If one considers the structure of the affirmative sentence, it would look like this:
S
NP VP
S
NP VP
PP
NP
The boy who is at the shop is naughty
As this structure clearly indicates, there is a relative clause which is embedded in the matrix sentence. Now, when forming a question, the verb that needs to move to the beginning is the verb that occurs as part of the VP that is higher up in the tree structure, i.e., part of the main VP. For that reason, the only way to form a question from this statement would be to move the main verb to the beginning, not the one that happens to appear first in the linear order of words, because it is much lower down on the structure due to its being part of the subordinate clause.
Because the transformation can only occur in this way, it would result in a structure like:
S
NP VP
S
NP VP
PP
NP
Is The boy who is at the shop (trace) naughty
Without knowledge of the structure, there is no way a child could know this. Given that children are not taught about the structure of language, we would have to explain how it is that children seem to always make use of principles like this.
If we do not concede that we have this elaborate linguistic structure hard-wired from birth, then we would not be able to explain why many aspects of even very young children’s linguistic competence is “grossly underdetermined by the fragmentary input from the environment”.
Convergence amongst grammars
People in a community all speak their mother tongue with equal proficiency, regardless of educational background; a ten-year-old child speaks with the same grammatical complexity as an educated adult. Given all the logically possible ways a language could develop, it is rather surprising that linguistic competence would not be as disparate as one would expect if there were no innate mechanisms constraining the development of grammar.
Language universals
In this section I will be discussing universals insofar as they relate to structure-dependency in language. I have not included the other kinds of universals that Chomsky refers to because the evidence for them is so weak that they are not even worth considering seriously. To illustrate what I mean by this, I will provide just two examples:
Colour Terms - The claim that all languages segment the colour spectrum, such that we have different words referring to different colours, is something that needs to be explained, since that is not the way things are in nature; the various colours simply ‘blend’.
Hence, how does one account for the fact that some languages have “one word for both blue and yellow”?
Related to Chomsky’s point on colour terms, Berlin and Kay conducted a survey, published in 1969, of ninety-eight different languages, and noticed certain patterns regarding their use of basic colour terms. By basic, they mean terms that are morphologically simplex (which precludes terms like bluish), non-compound terms, and terms that are not ‘shades’ of another colour ( thereby precluding terms like scarlet, which can be considered a shade of red), and words with restricted applications (like blond, which only refers to a male’s hair-colour).
According to Berlin and Kay, languages differ in the number of terms they utilise, but not otherwise. They made the startling claim that languages follow a universal schema. This schema can be illustrated as follows:
black pink
red green blue brown purple
white yellow orange
grey
What Berlin and Kay claimed is that if a language has two colour terms, they would be black and white. If they had three terms, they would be black, white and red. If a language had four terms, they would be black, white, red, and either green or yellow, etc. In languages which have more than seven terms, the schema breaks down, but the interesting thing is that in languages with seven or less terms, the schema applies universally.
According to Geoffrey Sampson, these results are only impressive if taken at face-value. Firstly, the results were not gathered by Berlin and Kay themselves, but by their students at the University of California. Each student was to find a language, and find out as much as they could about its colour terminology as part of their course work. Perhaps it is for this reason that some of their findings were inaccurate. For example, there are four basic colour terms in Homeric Greek, one of which is glaukos, which is actually translated as ‘gleaming’, with no implications of colour. Later on it came to mean something like what the English derivate ‘glaucous’ means: bluish-greenish-grey. However, they needed a word for black and duly translated it as such, despite the fact that there was a standard word for black in ancient Greek, viz. melas, and in the Homeric texts, it is the most common colour term.
With regards to loan words, Berlin and Kay include them when it fits their schema, and exclude them when it does not. What ought to happen is that they should decide at the outset whether or not it is viable to accept loan words as part of their data, and stick to that. However, they exclude the Chinese loans from Korean, but include the same Chinese loans in Vietnamese. This is because the latter are more amenable to their theory; the former are not.
Nevertheless, these are just some of the problems with the treatment of Berlin and Kay’s theory of basic colour terms. Sampson’s critique is more comprehensive, but the point I am trying to make is that this theory is not at all as water-tight as commonly assumed.
Proper nouns –Chomsky claims that there is “no logical necessity” for names to meet any condition of spatio-temporal contiguity, but the fact that they do is most certainly “non-trivial”, and in need of explanation. For example, the fact that no language makes use of a word like LIMB to refer to a quadruped’s four legs as if they were a single object (as in “The cat’s LIMB is grey”, referring to its four legs) needs explanation, because there is no logical reason for this restraint.
Regarding spatio-temporal contiguity, consider a noun phrase like The Three Sisters, with reference to the constellation of three linear stars, which are actually millions of kilometres apart. Also, French has a word, rouage, which means a singular object comprising the wheels of a vehicle. These are counter-examples which clearly show that these restrictions are actually not restrictions at all.
Examples of this kind are numerous, but due to the fact that just about every alleged universal has been refuted, or as Jackendoff once euphemised it, is “in need of explanation”; hence, these concrete examples will not be dealt with any further. What are more interesting are the abstract properties that languages have in common, or more specifically the fact that languages make use of hierarchical organisation, i.e., structure-dependency. This is interesting because:
- It is quite possible, logically, for languages to be structure-independent, and one can easily imagine a possible world where this is indeed the case.
- It is a matter of fact that all human languages so far known to us use structure-dependency, and this needs to be explained.
Chomsky says that if a Martian scientist were to come to earth to study our language, he would conclude that Earthlings speak a single language. The point he is trying to make (Chomsky, not the Martian scientist) is that all the languages in the world have a similar underlying structure, despite superficial disparities. The languages of the world are not as diverse as they could be because they share properties which they are not logically required to share. The fact that these findings are empirical claims is what makes this strand of Chomsky’s argument so convincing.
Many contend that the argument from universals is one of the strongest supporting the nativist’s paradigm. In fact, Chomsky has often remarked that the goal of modern day linguistics is to find out what these principles are; sometimes this is referred to as “the initial state of the Language Acquisition Device” (which is what Universal Grammar is) by Chomsky, or more simply the “language instinct”. Chomsky’s argument is by and large syntactically motivated.
It may be worth noting that in his earlier works, Sampson was quite content to agree with Chomsky’s inference of innateness from the fact that languages make use of tree-structuring. In a book called LIBERTY AND LANGUAGE, Sampson he says that “until someone offers a concrete and demonstrably preferable alternative, we will do well to believe the standard explanation… [and] I believe that Chomsky has quite adequately made his case for genetic inheritance…” (p. 24 – my italics). Of course, Sampson has now rejected even that, but we will deal with Sampson’s critique in a subsequent section.
Chomsky’s universals are syntactically motivated. The syntactic universals that Chomsky has pointed out have to do with the hierarchical organisation of sentences. A sentence consists of a main clause within which will typically be included subordinate clauses of various types. A clause has a verb at its core. These clauses are made up of phrases, which have a head and optional modifiers. Hence, an NP has a noun as its head, which may be modified by an article and adjectives. In other words, any sentence includes a hierarchy of elements of different sizes, groups of which are nested within elements at the next higher level of the hierarchy, and each of which is decomposable into a group of elements at the next lower level.
Chomsky illustrates how this hierarchical property of language is not logical by contrasting it with fictitious languages that are structure-independent. So, let us consider question-transformations, for example. A question is formed from a statement by moving a certain word to the beginning, so The men John wanted to see are in the kitchen becomes Are the men John wanted to see in the kitchen? The word to be moved (are) is chosen by reference to its position in the hierarchical structure: it is always the main verb of the main clause of the sentence. (This is discussed in more detail below when I talk about Sampson’s response to the ‘poverty of input’ argument). The point is that it would be perfectly possible to imagine languages whose questions were formed according to some principle that does not incorporate structure-dependency. Such a language might have, for example, as a rule, something like “To turn a statement into a question, front the third word in the statement.” No language does this, even though it is a logical possibility. There are countless other hypothetical rules that can be constructed, like “to form a question, reverse the order of words”, but we will not state them all explicitly. The point is that all the (known) languages of the world make use of hierarchical inclusion relationships, and that needs to be explained. This is a contingent fact that does not follow as a logical necessity.
Ironically, Sampson himself, in his earlier days, once conceded that this discovery is tantamount to discovering something like: the number of thorns on a rose bush is always a multiple of seven. This is an empirical finding, and needs to be explained. Why would one not agree with someone who says: the fact that every rose bush all over the world, regardless of the environment in which it grew, always displays this property? Does it not surpass belief to think that it could be due to mere coincidence? Hence, he accepted Chomsky’s idea of innate linguistic knowledge on that basis.
Upon analysis, we do indeed find that the various languages display certain universal features which follow as a consequence of structure-dependency. In addition to that of the question-transformation rule, let us consider another example, given by Chomsky:
a) John said that he was sick
b) He said that John was sick
In a), there is an ambiguity; it is not clear whether ‘he’ refers to John, or somebody else. In b), by merely inter-changing ‘John’ with the pronoun ‘he’, we know that ‘he’ cannot refer to John, but has to refer to some outside entity instead. One gets a better idea of why this happens if we consider the structure of these sentences. a) would look like this:
[John] [said that he was sick]
whereas b) would look like this:
[John [said [that [he [was sick]]]]]
The reason is that in a) ‘he’ can refer to the NP John because ‘John’ is “higher up” in the phrase structure tree than ‘he’. In b), ‘He’ is part of the main NP, making ‘John’ part of the NP that is lower down (hierarchically) on the phrase structure tree; and anaphors cannot be bound to a referent that is lower down in the tree. (The whole point of phrase-structure trees, by the way, is to show hierarchical inclusion relationships. Given that even four-year-old children know these things, knowledge of this kind must be part of our genetic make-up, since children are clearly not taught things like this. It should come as no surprise then that every native speaker of English intuitively knows that b) is dis-ambiguated (though most would not know why), given that we have these phrase-structure rules pre-programmed into our minds. If we believe that language is only acquired from the environment in which the child is in, “it seems inevitable that major aspects of verbal behavior will remain a mystery”.
(Unfortunately I cannot illustrate the tree here, as this would have made the point more clear.)
Chomsky thinks there are many other principles which constrain the structure of language, and that the goal of contemporary linguistics is to discover exactly what the nature of these constraints are. Some tentative principles which he postulates are the Extended Projection Principle (EPP), and the principle of full interpretation. The above-mentioned example is an instance of the binding principle.
Chomsky says that it is not the case that structure-independent principles are too complex; in fact they are cognitively quite simple. “Rather, what seems to be going on is that the messages humans need to transmit to one another through the medium of language are messages that come in certain fixed forms, and these forms to a large extent dictate the syntactic patterning of human languages”. These invariant syntactic properties are determined by the mechanisms of genetic inheritance.
The obvious, common-sensical objection to the thesis that we have a Universal Grammar is: how do we explain the fact that not all languages are the same? Chomsky accommodates this disparity into his theory by explaining it in terms of “Principles and Parameters” (henceforth P&P).
This theory claims that the child is able to learn language so effortlessly because he is born with a set of principles, together with a “list” of (unset) language-specific parameters. When these parameters are set, these, together with the common principles that form part of our biological endowment, constitute what Chomsky calls “knowledge of language”. So we are born with rules governing sentence structure, like (to oversimplify) the following phrase-structure rule:
S -> {NP,VP}
The comma indicates variable order. When the child is about four years old, he realises that the language he is exposed to, language X, has the NP-first construction, and the genetically pre-programmed Language Acquisition Device crystallises this rule to read:
S -> NP VP
In language Y, however, the child realises that his language employs the VP-first construction, and now the genetically pre-programmed Language Acquisition Device crystallises this rule to read:
S -> VP NP
Hence, these two children may grow up speaking mutually unintelligible languages.
Obviously these are not the only rules that get crystallised during this process of so-called “parameter setting”, but we will leave it at that for the sake of simplicity.
This is the nativist’s explanation as to why language can be both innate, and variant across the human race. We will see below that this theory has problems which need to be addressed.
Nativists are wont to ask why we cannot form questions by interchanging the first and second words in a sentence, by reversing the order of words, or by picking out the middle word in a sentence.
What is wrong with turning the sentence
John ate an apple
into a question by saying
Apple an ate John?
Why does this feel so very unnatural? As pointed out, this is because language always makes use of structure-dependency. But if it is possible in principle for a language to have a rule like “Reverse the words in a sentence to form a question”, the fact that it never occurs needs explaining. This mystery can be explained if we assume that this restriction is hard-wired into our genes, and therefore restricted by our biology. If we do not accept this, then we cannot explain why the above-mentioned rule is so unnatural, and we would have to be at a loss in terms of being able to explain why rules of this type are so uncomfortable for us to utilise.
Descartes’ Problem
One can outline what Chomsky refers to as “Descartes’ problem” (cf. his book CARTESIAN LINGUISTICS) as follows:
How is it that we as native speakers are able to understand and produce a potentially infinite number of sentences, despite very often hearing/producing utterances that one has never before encountered in his life, or as Chomsky puts it, never before in the history of the universe. This Chomsky refers to as the creative aspect of language use (aka “Descartes’ problem” – so-called because it was Descartes, according to Chomsky, who alluded this problem).
In addition to this, Chomsky emphasises the fact that speakers are not able to express this variety of linguistic utterances due to being endowed with the relevant physiological apparatus (vocal tract, etc.), because a parrot is able to do the same, and he contends that the physiological mechanisms pertaining to the manifestation of sound are not necessarily in any way inferior to those of a human beings.
What, then, distinguishes a parrot’s ‘speech’ from human speech? Chomsky tells us that it is because there is something unique to the human mind which neither parrots nor any other (non-human) animal possesses. This ‘something’ is none other than our inborn Language Acquisition Device.
Aside from the fact that our innate rules give us the machinery to be ‘creative’, humans also have the ability to tell which sentences are grammatical, and which are not. How are children, faced with countless possibilities as to how to string sentences together, able to actually make sense of what they hear? If children were going on what they hear, they would need some way of knowing which sentences are acceptable, and which are not, and given that the possibilities are endless, how would children be able to cope with this mammoth task of learning the rules of language? Once again, Chomsky says that this would not be a mystery if we simply had a genetic predisposition to do so.
That pretty much sums up Chomsky’s case for his theory. His other writings serve only to supplement the points made above, with minor variations.
Let us now turn to the work of Steven Pinker, one of the leading exponents of neo-Chomskyan nativism. Pinker’s point of view will be taken primarily from his book THE LANGUAGE INSTINCT, and less so from a later book called WORDS AND RULES.
The Language Instinct
Here now are some of the key arguments found in Pinker, followed by my criticisms:
Phonotactic Constraints
Pinker claims that languages have rigid rules about how sounds are assembled into words. Many introductory texts call these “phonotactic constraints”. If you had to ask a native speaker of English about the grammaticality of certain words, he would somehow always know whether they are acceptable words or not. For example, words like “plaft” and “blick” are unequivocally judged as possible yet contingently non-existent words. Words like “sram” are unequivocally judged as impossible, and necessarily non-existent. This is because “plaft” and “blick” fit in with the existing sound system of English, whereas “sram” does not; it has an illicit onset sequence of “sr”. These constraints are part of our LAD, which allow us to set the parameters for our language at some relatively early stage of our linguistic development.
Infant Prodigies
Pinker says that children have to know what to expect when acquiring language, otherwise they would take forever to learn their mother-tongue. How would they know what to test for? How would they know what a verb’s inflexion depends on? Quoting Pinker, Sampson says that logically speaking, “an inflection could depend on whether the third word in the sentence referred to a reddish or a bluish object; whether the sentence was being uttered indoors or outdoors, and billions of other fruitless possibilities that a grammatically unfettered child would have to test for”.
Pinker then goes on to say that children never make the mistakes we would expect them to make if they learnt language from scratch. Pinker also refers to one of Stromswold’s Stromswold’s studies, based on the speech of pre-schoolers. She found out that even though there were false but logical analogies that could have been made by the children, they were never tempted to give in to these mistakes. Why does the child not commit a false analogy like He cans go based on I like going → He likes going; I can go → ???
Stromswold reported that the pre-schoolers never made a single false analogy. It is not the case that children never make mistakes when acquiring language. Pinker himself quotes the example of a father trying to correct his daughter from saying other one spoon to saying the other spoon, but to no avail. So what does Pinker want us to conclude from Stromswold’s work? It seems as if he is saying that the rule placeholders, such as ‘one’, can be followed by the noun for which it stands is a rule found in some languages, and since the child has not quite set her parameters to English yet, she utters possible, but (contingently) non-occurring utterances. However, no child will try out the idea that “can” will inflect to agree with its subject if it is in the third person; our innate knowledge precludes this kind of option, therefore children never commit this error.
Second Language Learning and the Critical Period
Pinker mentions Johnson and Newport’s research on Korean and Chinese immigrants to the USA. They found that the later the immigrants arrived, the less proficient they were at detecting grammatical errors. They also found that the poor immigrants’ attitude towards the language, and their motivation, had nothing to do with it; their age when they arrived had everything to do with it.
Pinker also mentions that despite the fact that Meryl Streep is renowned in the US for her accents, her English accent is “considered rather awful” in England, and that her Australian accent does not “go down too well in Australia either”. This is the case because she is trying to speak like a native speaker when the rules for her L1 have already crystallised, and she will therefore never be able to speak as proficiently as a native speaker.
The Language Mutants from Essex
One seemingly compelling case is that of one family who exhibited an anomaly known as SLI (Specific Language Impairment), whereby certain linguistic faculties are impaired, and their cognitive functioning in other areas were normal; the fact that it was carried over within the family shows that there certainly is a genetic element involved. Despite the fact that many of the ones who had SLI grew up in a “linguistically normal” environment, they still exhibited the anomalies that SLI sufferers have. Pinker also reminds us that they are cognitively normal in other domains.
The Structure of Words
Pinker claims that the structure of words also lends evidence to nativism in the sense that children’s minds seem to be designed with the logic of word-structure built in.
His first example is that of a pattern whereby the derivational affix can attach only to roots, not to stems (i.e. roots with at least one affix attached). When we try to attach a derivational affix to a stem, we get something like Darwinsian, which could refer to the works of Charles Darwin and his grandfather, Erasmus Darwin, whose work Charles drew profound inspiration from. However, this would sound ridiculous, not because it is illogical, but because there seems to be some sort of innate restriction on what affix can go where in a word. Later Pinker says that the restriction is that derivational suffixes cannot attach themselves to inflected roots.
Pinker also discusses headless compound words, and draws conclusions which provide further evidence for his case. An example of a headed compound would be barman, with the –man part constituting the head. Hence, a “barman” is a kind of man, viz. one that works in bars. Headless compounds are words like “walkman”. Here, -man cannot be the head because a “walkman” is not a kind of man. (Nor is it a kind of walk, but the rightmost morpheme is usually defined as the head of a compound).
Pinker’s explanation is that when a compound has a head, all the information regarding the head gets percolated up to the top-most node of the (morphological) tree, including the rules. When there is a headless compound, it is not clear where the information must come from, so the “percolation conduits” get blocked, and the default rule applies. That is the reason why we say “sabre-tooths” and “still-lifes”.
Other research cited by Pinker includes Peter Gordon’s experiment with 3-year-olds, commonly referred to as the “Jabba experiment” (Gordon, 1986). He showed them a puppet and said: this is Jabba. Jabba eats mice. Jabba is a ___? And the children reply: Mice-eater. Thereafter, he asks: Jabba also eats rats. Jabba is also a ___? And the children reply: Rat-eater.
Now, Gordon asks, why do they use the plural “mice”, but not the plural “rats”? Given that not a single child ever used the inflected plural as the initial constituent of the compound, this could indicate that there is some sort of rule that they are applying. It is unlikely that they were told that this is so, as Motherese contains few or no compounds, especially of that kind; and it would be preposterous to assume that someone told them not to include inflected words in their compounds. Why do they not generalise the rule: “create a compound by putting two words together”? Why should there be a restriction of this sort? The nativist would explain it as follows: if there were an innate rule of this kind to facilitate language acquisition, the child would already “know” some of the things that cannot be said, like including an inflected word in the first part of a compound word, which is obviously why they avoid constructions of this sort.
The results were interesting because there seemed to be a consistent restriction in terms of what kind of word was allowed to be in the first part of a compound word. “Mice-eater”, for example, was an acceptable form to them; their answer to the first question would be just that. However, they never answered “rats-eater” to the second question; they always answered “rat-eater”. This is interesting because if one goes solely on logic, then “rats-eater” ought not to be precluded, since “mice” and “rats” are both plural forms, and both (equally) morphologically complex.
The question, then, that Gordon posed is: what exactly makes forms like “rats-eater” ungrammatical? The reason he postulated was that the form “rats” is clearly derived from a rule: rat + -s. The word “mice”, however, is not; despite its morphological complexity, the word is learnt as a listeme, and therefore treated as such by the lexicon. Hence, “mice” in “mice-eater” is treated by the phonology as a simplex word.
This, Gordon claims, is an innate constraint placed on us by UG, and we have fairly good reason to assume this, given that children as young as three consistently do it.
There are problems with this claim, however, since there are languages which do exactly that which is claimed by Gordon to be universally excluded by an innate constraint. German, for example, makes use of regular plural forms in compounds. Consider the following set of data in the tables below:
English compound
German translation
1) Mice-eater
Mäusefresser
2) Rat-eater
Rattenfresser
3) Butterfly-eater
Schmetterling(s)fresserª
4) Insect-eater
Insektenfresser
5) Cat-eater
Katzenfresser
6) Flower-eater
Blumenfresser
7) People-eater
Menschenfresser
8) Tree-eater
Bäumefresser
9) Sheep-eater
Schaf(s)fresser
10) Fish-eater
Fischfresser
In German, -e(n) represents the plural suffix. To clarify, let us look at the singular/plural forms of the (German) words used in the data above:
Singular
Plural
1) Maus
Mäuse
2) Ratte
Ratten
3) Schmetterling
Schmetterlinge
4) Insekt
Insekten
5) Katte
Katzen
6) Blume
Blumen
7) Mensch
Menschen
8) Baum
Bäume
9) Schaf
Schafe
10) Fisch
Fische
In examples 9) and 10) [schaf(s)fresser and fischfresser], the singular form was preferred by my German informant; in examples 1) and 4) [mäusefresser and insektenfresser], both the singular and plural were judged to be acceptable forms, such that both mausfresser and insektfresser would not count as ungrammatical, though mäusefresser and insektenfresser are the more preferred of the two.
What is interesting about this is that examples 4-8 make use of the plural form in exactly the way Gordon would predict to be impossible, since here we have rule-derived words found in the position they should not be found in.
These alleged counter-examples can be dealt with if it can be shown that the phonology treats these morphologically complex words as simplex; in other words, the phonology is blind to the morphological complexity of these forms. Now we know that irregular plurals are certainly treated by the phonology in the same way as simplex listemes, but if we can show that this is true of not only irregulars, but other complex forms as well, maybe we can save Gordon’s thesis from these alleged counter-examples.
Jonathan Kaye, a phonologist from SOAS, has given us some novel insight into this issue of word structure. In discussing Kaye’s views, we will see why words cannot always be classified as either simplex or complex a binary fashion. Hence, let us now turn to Kaye's discussion of this issue (his theory is based on Government Phonology), so as to gain a better understanding of the matter at hand. (Note: the data here is once again based on English):
We cannot assume that the phonology is per se blind to morphological complexity, since the [mz] (from "dreams") is clearly a pseudo-cluster, and deemed both morphologically and phonologically complex.
It is clear, then, that at least some morphological structure is visible to the phonology. Kaye draws a distinction between what he calls analytic morphology and non-analytic morphology. Analytic morphology is morphology wherein the phonology is able to recognise a complex domain structure; non-analytic morphology is not able to. Let us first consider analytic morphology (which is visible to the phonology).
Analytic Morphology
As mentioned, the [mz] cluster from “dreams” is invariably complex. What does this tell us about morphological penetration into phonology?
Let us firstly assume that only morphological domains can be represented in the phonology. A compound like “black-board”, then, will have the following structure:
a) [[black] [board]]
Since there are three pairs of brackets, it reflects three phonological domains. These brackets represent how the phonological string is to be processed. To give an exact definition, we can use the CONCAT and φ (apply phonology) functions, so “blackboard” can be shown as
b) φ (CONCAT (φ(black) φ(board)))
To translate this pseudo-formula into plain language: apply phonology to “black”, then to “board”, concatenate the result, and then apply phonology to that string. So the brackets we see in example a) are actually not part of the phonological representation, but actually outline domains which are arguments to functions.
So this illustrates one type of structure that involves morphologically complex forms. It has the schema [[A][B]]. This is the case for most English compounds.
There is a second structure, involving two morphemes, but only two domains. The structure is [[A] B], and is interpreted as follows:
φ (CONCAT (φ(A), B))
To translate, we do phonology on A, and concatenate the result with B; then do phonology on the result. This is the case with English (regular) inflexional morphology. For example, consider the past tense of the word “seep”. Its structure is [[seep] ed]. With regards to Government Phonology, we could also illustrate this structure in more detail:
[[O N O N] O N]
x x x x x x x
s i: p d
Note that empty nuclei are to be found at the end of each domain, and are licensed to be silent by virtue of their domain-final position. Also, the licensing principle states that all positions in a phonological domain need to be licensed, unless that position is the head of that particular expression. This is actually why the suffix does not form a domain by itself. We see here that the suffix consists of two positions: an onset position and a nuclear position, both being licensed. If “-ed” were a domain, it is obvious that it would violate the licensing principle, which states that a phonological domain must have at least one unlicensed position – its head. The onset is licensed because, like all onsets, it is licensed by the nucleus to its right, which is itself p-licensed, i.e., licensed to be silent .
Analytic morphology also has a rather interesting property in that the distribution of empty nuclei is very restricted; in English, it is almost always found at the end of domains. Hence, this provides us with a fairly reliable parsing cue, since a p-licensed empty nucleus signals the end of a domain, so if it occurs within a word it shows that it is not a morphologically simplex one. So the pseudo-cluster [mz], from “dreams” is so-called because:
The vowel length is maintained (which would not be possible had it been an authentic cluster), and [m] and [z] are not homorganic, which again is not possible for true clusters. It is for this reason that much of analytic morphology is phonologically parsable.
So hitherto, we have considered two structures:
[[A][B]]
[[A] B],
and we have seen how morphology impacts on phonology in the form of these two domains.
Let us now turn to non-analytic morphology (which is invisible to the phonology).
Non-analytic Morphology
If morphological structure was present in the form of domains, then to that extent morphology can have an impact on phonology. These domains have the effect of respecting the integrity of the internal domains. By this, Kaye means that when you append an affix, for example, the actual pronunciation of the root/stem does not change. Consider the “-ing” suffix, whose form can be shown as [[V]-ing]. To pronounce this, you just pronounce the verb on its own and attach the suffix. This procedure does not apply to all forms of morphology. For example, take the words “parenthood” and “parental”; the respective suffixes interact with the phonology in different ways, as is evident from the variable pronunciation of the aforementioned words. In fact, the “-al” type of morphology is invisible to the phonology. In other words, the phonology reacts as if there were no morphology at all. Hence, the phonology would treat a word like “parental” in the same way it would treat a mono-morphemic word like “agenda”. So we can characterise the internal structure of a form like “parental” as:
[AB]
It is this kind of morphology that is referred to as non-analytic. And since the only effect morphology can have on phonology is the presence of internal domains, it follows that this kind of morphology can have no effect on phonology, since there are no internal domains. To further accentuate the disparity between analytic and non-analytic morphology, let us look at the morpheme “un-” (which is analytic), and the morpheme “in-” (which is non-analytic).
“un-” does not change its form, regardless of what consonant follows it. This property follows from its analytic morphology. Since the prefix “n” is not adjacent to the following onset, there is no restriction in terms of phonotactics. A word like “unreal”, then, can be represented as
[[un Ø] [ri:l Ø]],
where Ø represents the empty nucleus, whereby the first of which separates the two domains. The first empty nucleus provides us with a parsing cue, along with the nucleus in the following syllable, such that we are able to perceive the prefix as a separate structure. So a string segment like “nØr” is phonologically parsable, and one can analyse the form as “un-real”.
With the prefix “in-”, the situation is different. Appending “in-” to a stem must yield a well-formed phonological domain on its own; this is because non-analytical morphology, by definition, has no internal domains. A non-analytic combination of morphemes is interpreted as follows:
φ (concat (A,B))
You concatenate the two strings and then apply phonology to the result. So a form like “in-rational” is not well-formed because “nr” is not a possible sequence. That is why the “n” has to be dropped, resulting in the form “irrational”.
So non-analytic morphology is invisible to the phonology. There are no internal domains, or any other phonological indication of morphological complexity. Non-analytic forms have the same phonological properties as any simplex form. It is reasonable to assume, then, that any form manifesting non-analytic morphology is listed in the lexicon.
Irregular forms that exhibit morphological complexity are not phonologically parsable. This is because there are no phonological hints of their complex morphological structure. It is for this reason that irregular morphology is always non-analytic.
So with regards to the kind of compounds we are looking at, compounds can be classified into four groups, viz.:
Group 1 – Analytic, singular (example, tree-eater)
Group 2 – Analytic, plural (example, rats-eater)
Group 3 – Non-analytic, singular (example, linguistics-programme)
Group 4 – Non-analytic, plural (example, mice-eater)
In the groups above, ‘analytic’ refers to the compound itself, and ‘singular/plural’ refers to the modifying noun.
The relevance of his thesis is quite evident in that it may explain why the counter-examples to Gordon’s experiment are only pseudo-counter-examples. German, for example, does make use of compound words that contain a plural form. This seems, prima facie, a blatant refutation of Gordon’s claim. However, Kaye offers us some novel way to save the Gordon-hypothesis from refutation. My point here, however, is that Gordon’s argument as it stands is untenable, but can be saved in light of Kaye’s explanation, as the plural compound words have been shown to belong to our “Group 4” type words. However, German is peculiar in that the language has more irregular forms than regular ones, which leads one to the rather quaint conclusion that the ‘irregular’ plural seems to be more regular than the ‘regular’ plural inflexional suffix. Hence, the explanation for this phenomenon has nothing to do with genetic inheritance, but can be explained by appealing to the logic of word structure. Whether that is innate is a separate question, and Gordon would have to explain why languages do it differently. One could easily do so with a parameters-like explanation, but as I will explain later, this kind of approach is unscientific, and therefore vacuous. If there are other explanations that explain the same data more consistently and more logically, why should we accept the claim that this is unequivocally a result of innate constraints? For the Nativists, they insist that innateness is the only logical explanation for the phenomenon Gordon refers to; Kaye has shown otherwise. This concludes my discussion of Kaye.
A form like “mice-eater” is acceptable because the phonology is blind to its morphological complexity. A form may also not be irregular, and still be an acceptable part of the compound, as in the plural German compound words found in the data above. As Kaye has established, complex forms with no internal domains are treated by the phonology as if they were simplex. If German can be shown to exhibit non-analytic morphology with regards to using plural modifying nouns, they do not constitute true counter-examples; or rather, Gordon’s original hypothesis can be easily adapted to accommodate these facts without actually compromising the crux of his original thesis, which could now state that the restriction applies only to constituents that exhibit analytic morphology. For example, in a compound like rattenfresser, even though ratten is the (regular) plural of ratte, it actually is non-analytic, which means that the phonology is blind to the morphological complexity of the modifying noun.
On the Origin of Language
Chomsky is of the opinion that we would have to remain agnostic as to how language came in to being. In fact, Chomsky goes so far as to insist that we can “tell the story as we like”, and that Darwinian natural selection will not suffice as an explanation since there are no precedents anywhere in the natural world for language. Language, according to Chomsky, lacks the messiness we would expect of an accumulation of accidents made good by evolutionary ‘tinkering’. Characterised by beauty bordering on perfection, it cannot have evolved in the normal biological way. Pinker disagrees with this claim. In paper with Bloom (published in 1990) as well as in The Language Instinct, it is argued that natural selection is the only viable explanation for the origin of language. Pinker draws the analogy between language and spider webs; Pinker & Bloom say that language “is a topic like echolocation in bats or stereopsis in monkeys”.
Pinker hypothesises that there must have been a mutation of some kind in our genetic make-up, which happened to be useful to that particular individual. This trait would have been passed on, as natural selection retains traits which maximise the chances of survival, therefore the grammar gene would have been spread to become what it is today. Pinker is scanty on the details, and he himself concedes that this is speculative, but goes on say that the survival benefits are very obvious, and therefore it is quite plausible for a scenario of this kind to take place. How else would humans have worked together in groups to hunt wild-life far more powerful than themselves?
The hypothesis that grammar is historically the exclusive product of a special-purpose genetically provided mental organ for grammar rests upon a negative argument: we have not managed to explain how grammar arises through general processes. Hence, we must postulate the existence of a special genetic code that is entirely responsible for building in the brain a UG organ, which will someday be located and understood in the way the retina or the heart has been located and understood.
Linguistic Relativity
We are introduced on p. 57 of the THE LANGUAGE INSTINCT to "the famous Sapir-Whorf hypothesis of linguistic determinism", named after Edward Sapir and Benjamin Lee Whorf. According to Pinker, this is the notion that language determines the way we think, and the way in which we perceive the world. He illustrates this notion with the misunderstood statement that Eskimos have hundreds of words for snow, and therefore they somehow perceive snow differently to how English speakers would. This does not make sense, since it is actually not true that Eskimos have so many words for snow, and this does not take into consideration the fact that English has over a dozen words for snow. Clearly, it does not make sense to stipulate that Eskimos have different perceptual mechanisms just because they have a different language. He also mentions that just because some languages have less colour terms than, say, English, it does not follow that speakers of these languages do not perceive colour in the same way we do. The point Pinker is making here is simply that all human beings have the same cognitive faculties, and that we would all perceive the world in the same way, even though our linguistic conventions in naming that world may differ. In fact, Pinker notes, no matter how influential language may be, the idea that language can “reach down into the retina and rewire the ganglion cells” is preposterous.
ª The bracketed “s” is referred to as a “fugen-s ” [gapping/bridging s], so-called because it is there to serve as a bridge between the two constituents, and does not in itself carry any semantic significance.
Chomsky’s key arguments are as follows:
Speed of acquisition
Critical period
Poverty of input
Convergence amongst grammars
Language universals
Pinker bolsters Chomsky’s paradigm in many of his works, but The Language Instinct and Words and Rules will be considered in more detail than others since this is where the former outlines his case and shows us where he stand on this debate.
I will now explain each of these in turn:
Speed of acquisition
The fact that children acquire language so effortlessly and so quickly, despite its obvious complexity, shows that they must be innately disposed to do so. Other forms of learning are slow and protracted (cf. the freshman’s struggle in learning things like physics).
Stages of acquisition
Children learn language in the same stages all over the world, starting with the babbling stage, moving on to the one-word stage, then the two-word stage, then the telegraphic stage, and finally culminating in fluent speech.
Why would a child belonging to some tribe in the Amazon acquire his language in the same stages as those of an Italian child in Milan? In fact, children all over the world seem to follow the same pattern when learning a language. These stages occur in this order not because they are required to by the laws of logic, but because it is a matter of the unfolding of a genetic blueprint. If we do not accept this as the only viable explanation of the facts, we would either be forced to concede that we do not know, or have to postulate a rather surprising coincidence. If we accept that language is no more ‘learnt’ in this order any more than we “learn to have two arms”, it at least provides us with an intellectually satisfying explanation of the facts.
Critical period
Then there is the critical period thesis, which claims that there is a period in one’s life when one is particularly disposed to learning a language, after which language learning becomes so much more difficult. Chomsky insists that the “functioning of the language capacity is […] optimal at a certain ‘critical period’ of intellectual development”; just think of the mother-tongue English speaker trying to learn French at, say, university. According to Susan Curtiss, “there are two versions of Lenneberg’s critical period hypothesis”: the strong version and the weak version. The strong version says that it is not possible at all to acquire language naturally after puberty; the weak version says that one cannot acquire language with equal proficiency to that of a native speaker after puberty. As will be discussed in a subsequent section, the strong version is not testable in practice. Hence, the case of Genie is used to lend support to the viability of the weak version. (Nonetheless, I contend that both the strong and weak versions are flawed.)
Genie was the victim of an abusive father, who locked her up in a room upstairs. She was chained to a toilet seat, and not allowed to leave the room. They passed her food through a little flap at the bottom of the door. She literally had to live in a room for all her childhood. The neighbours reported something strange in the household, and after investigation, a social worker eventually discovered Genie.
Upon her discovery, the father shot himself dead, and her mother claimed that she was continually beaten into submission to follow his orders, and therefore could not do anything about her quandary.
Without going into the maudlin details, this story is relevant because when they discovered Genie she could do little more than feed herself. She walked in a very animal-like fashion, had no sense of decorum, often grunted as if it was normal, and she could not speak or understand language. After intensive therapy, she was able to control her outbursts, stopped grunting, learnt how to walk properly…etc. However, up till today, she has never learnt how to speak with the efficiency of a native speaker. Despite some of the world’s best psycholinguists and speech therapists working on her, she was only able to utter very basic, rudimentary sentences. The question is: why did she (inter alia) learn how to walk properly, but could not acquire language proficiently? The nativist’s answer: because she had passed her critical period for language acquisition.
Poverty of input
Chomsky claims that the input children get from their linguistic environment is far from ideal, and children are generally not corrected on grammatical issues - parents are more interested in things like ethics and truth-content. Even though overt correction is not necessary, it is still interesting that even young children never utter sentences which are syntactically deviant, even though they may over-generalise rules. Chomsky would say that this is because we have the syntax hard-wired into our minds.
The point is that children must infer general rules about language after hearing a few existential examples, despite not being told to do so, and despite the fact that even these examples are not ideal.
Aside from the anecdotal evidence that Chomsky presents, he always points out children’s knowledge of yes-no question formation in complex sentences. He explains that if we consider a sentence like
The boy who is at the shop is naughty
there are two ‘hypotheses’ a child could construct when trying to work out the rule, let us call these hypothesis 1 and hypothesis 2:
Move the first verb of the sentence to the beginning when forming a yes-no question
Move the verb that occurs in the matrix sentence (as opposed to the embedded clause) to the beginning
Of course, the correct rule would be hypothesis number 2, which would produce the grammatical sentence:
Is the boy who is at the shop naughty?
However, the point Chomsky is trying to make is that no child would ever utter a question using hypothesis number 1:
*Is the boy who at the shop is naughty?
If the child was going on what he heard from people around him speaking, there is nothing in the speech input that prevents him from distinguishing between the two hypotheses. So, given that the application of hypothesis number 1 is a logical possibility, and given that number 2 is also a logical possibility, what precludes the child from forming questions at least some of the time in the manner above, using hypothesis 1?
If the child is using the “fragmentary input” from the environment as a basis for his hypothesis formation, there is absolutely nothing that stops him from doing so. Logically, then, Chomsky would have us conclude that the child is not using the environment, but comes with his own innate linguistic structure hard-wired into his mind. This might seem prima facie implausible for someone using his common-sense, but how would you otherwise explain the above-mentioned facts?
If one considers the structure of the affirmative sentence, it would look like this:
S
NP VP
S
NP VP
PP
NP
The boy who is at the shop is naughty
As this structure clearly indicates, there is a relative clause which is embedded in the matrix sentence. Now, when forming a question, the verb that needs to move to the beginning is the verb that occurs as part of the VP that is higher up in the tree structure, i.e., part of the main VP. For that reason, the only way to form a question from this statement would be to move the main verb to the beginning, not the one that happens to appear first in the linear order of words, because it is much lower down on the structure due to its being part of the subordinate clause.
Because the transformation can only occur in this way, it would result in a structure like:
S
NP VP
S
NP VP
PP
NP
Is The boy who is at the shop (trace) naughty
Without knowledge of the structure, there is no way a child could know this. Given that children are not taught about the structure of language, we would have to explain how it is that children seem to always make use of principles like this.
If we do not concede that we have this elaborate linguistic structure hard-wired from birth, then we would not be able to explain why many aspects of even very young children’s linguistic competence is “grossly underdetermined by the fragmentary input from the environment”.
Convergence amongst grammars
People in a community all speak their mother tongue with equal proficiency, regardless of educational background; a ten-year-old child speaks with the same grammatical complexity as an educated adult. Given all the logically possible ways a language could develop, it is rather surprising that linguistic competence would not be as disparate as one would expect if there were no innate mechanisms constraining the development of grammar.
Language universals
In this section I will be discussing universals insofar as they relate to structure-dependency in language. I have not included the other kinds of universals that Chomsky refers to because the evidence for them is so weak that they are not even worth considering seriously. To illustrate what I mean by this, I will provide just two examples:
Colour Terms - The claim that all languages segment the colour spectrum, such that we have different words referring to different colours, is something that needs to be explained, since that is not the way things are in nature; the various colours simply ‘blend’.
Hence, how does one account for the fact that some languages have “one word for both blue and yellow”?
Related to Chomsky’s point on colour terms, Berlin and Kay conducted a survey, published in 1969, of ninety-eight different languages, and noticed certain patterns regarding their use of basic colour terms. By basic, they mean terms that are morphologically simplex (which precludes terms like bluish), non-compound terms, and terms that are not ‘shades’ of another colour ( thereby precluding terms like scarlet, which can be considered a shade of red), and words with restricted applications (like blond, which only refers to a male’s hair-colour).
According to Berlin and Kay, languages differ in the number of terms they utilise, but not otherwise. They made the startling claim that languages follow a universal schema. This schema can be illustrated as follows:
black pink
red green blue brown purple
white yellow orange
grey
What Berlin and Kay claimed is that if a language has two colour terms, they would be black and white. If they had three terms, they would be black, white and red. If a language had four terms, they would be black, white, red, and either green or yellow, etc. In languages which have more than seven terms, the schema breaks down, but the interesting thing is that in languages with seven or less terms, the schema applies universally.
According to Geoffrey Sampson, these results are only impressive if taken at face-value. Firstly, the results were not gathered by Berlin and Kay themselves, but by their students at the University of California. Each student was to find a language, and find out as much as they could about its colour terminology as part of their course work. Perhaps it is for this reason that some of their findings were inaccurate. For example, there are four basic colour terms in Homeric Greek, one of which is glaukos, which is actually translated as ‘gleaming’, with no implications of colour. Later on it came to mean something like what the English derivate ‘glaucous’ means: bluish-greenish-grey. However, they needed a word for black and duly translated it as such, despite the fact that there was a standard word for black in ancient Greek, viz. melas, and in the Homeric texts, it is the most common colour term.
With regards to loan words, Berlin and Kay include them when it fits their schema, and exclude them when it does not. What ought to happen is that they should decide at the outset whether or not it is viable to accept loan words as part of their data, and stick to that. However, they exclude the Chinese loans from Korean, but include the same Chinese loans in Vietnamese. This is because the latter are more amenable to their theory; the former are not.
Nevertheless, these are just some of the problems with the treatment of Berlin and Kay’s theory of basic colour terms. Sampson’s critique is more comprehensive, but the point I am trying to make is that this theory is not at all as water-tight as commonly assumed.
Proper nouns –Chomsky claims that there is “no logical necessity” for names to meet any condition of spatio-temporal contiguity, but the fact that they do is most certainly “non-trivial”, and in need of explanation. For example, the fact that no language makes use of a word like LIMB to refer to a quadruped’s four legs as if they were a single object (as in “The cat’s LIMB is grey”, referring to its four legs) needs explanation, because there is no logical reason for this restraint.
Regarding spatio-temporal contiguity, consider a noun phrase like The Three Sisters, with reference to the constellation of three linear stars, which are actually millions of kilometres apart. Also, French has a word, rouage, which means a singular object comprising the wheels of a vehicle. These are counter-examples which clearly show that these restrictions are actually not restrictions at all.
Examples of this kind are numerous, but due to the fact that just about every alleged universal has been refuted, or as Jackendoff once euphemised it, is “in need of explanation”; hence, these concrete examples will not be dealt with any further. What are more interesting are the abstract properties that languages have in common, or more specifically the fact that languages make use of hierarchical organisation, i.e., structure-dependency. This is interesting because:
- It is quite possible, logically, for languages to be structure-independent, and one can easily imagine a possible world where this is indeed the case.
- It is a matter of fact that all human languages so far known to us use structure-dependency, and this needs to be explained.
Chomsky says that if a Martian scientist were to come to earth to study our language, he would conclude that Earthlings speak a single language. The point he is trying to make (Chomsky, not the Martian scientist) is that all the languages in the world have a similar underlying structure, despite superficial disparities. The languages of the world are not as diverse as they could be because they share properties which they are not logically required to share. The fact that these findings are empirical claims is what makes this strand of Chomsky’s argument so convincing.
Many contend that the argument from universals is one of the strongest supporting the nativist’s paradigm. In fact, Chomsky has often remarked that the goal of modern day linguistics is to find out what these principles are; sometimes this is referred to as “the initial state of the Language Acquisition Device” (which is what Universal Grammar is) by Chomsky, or more simply the “language instinct”. Chomsky’s argument is by and large syntactically motivated.
It may be worth noting that in his earlier works, Sampson was quite content to agree with Chomsky’s inference of innateness from the fact that languages make use of tree-structuring. In a book called LIBERTY AND LANGUAGE, Sampson he says that “until someone offers a concrete and demonstrably preferable alternative, we will do well to believe the standard explanation… [and] I believe that Chomsky has quite adequately made his case for genetic inheritance…” (p. 24 – my italics). Of course, Sampson has now rejected even that, but we will deal with Sampson’s critique in a subsequent section.
Chomsky’s universals are syntactically motivated. The syntactic universals that Chomsky has pointed out have to do with the hierarchical organisation of sentences. A sentence consists of a main clause within which will typically be included subordinate clauses of various types. A clause has a verb at its core. These clauses are made up of phrases, which have a head and optional modifiers. Hence, an NP has a noun as its head, which may be modified by an article and adjectives. In other words, any sentence includes a hierarchy of elements of different sizes, groups of which are nested within elements at the next higher level of the hierarchy, and each of which is decomposable into a group of elements at the next lower level.
Chomsky illustrates how this hierarchical property of language is not logical by contrasting it with fictitious languages that are structure-independent. So, let us consider question-transformations, for example. A question is formed from a statement by moving a certain word to the beginning, so The men John wanted to see are in the kitchen becomes Are the men John wanted to see in the kitchen? The word to be moved (are) is chosen by reference to its position in the hierarchical structure: it is always the main verb of the main clause of the sentence. (This is discussed in more detail below when I talk about Sampson’s response to the ‘poverty of input’ argument). The point is that it would be perfectly possible to imagine languages whose questions were formed according to some principle that does not incorporate structure-dependency. Such a language might have, for example, as a rule, something like “To turn a statement into a question, front the third word in the statement.” No language does this, even though it is a logical possibility. There are countless other hypothetical rules that can be constructed, like “to form a question, reverse the order of words”, but we will not state them all explicitly. The point is that all the (known) languages of the world make use of hierarchical inclusion relationships, and that needs to be explained. This is a contingent fact that does not follow as a logical necessity.
Ironically, Sampson himself, in his earlier days, once conceded that this discovery is tantamount to discovering something like: the number of thorns on a rose bush is always a multiple of seven. This is an empirical finding, and needs to be explained. Why would one not agree with someone who says: the fact that every rose bush all over the world, regardless of the environment in which it grew, always displays this property? Does it not surpass belief to think that it could be due to mere coincidence? Hence, he accepted Chomsky’s idea of innate linguistic knowledge on that basis.
Upon analysis, we do indeed find that the various languages display certain universal features which follow as a consequence of structure-dependency. In addition to that of the question-transformation rule, let us consider another example, given by Chomsky:
a) John said that he was sick
b) He said that John was sick
In a), there is an ambiguity; it is not clear whether ‘he’ refers to John, or somebody else. In b), by merely inter-changing ‘John’ with the pronoun ‘he’, we know that ‘he’ cannot refer to John, but has to refer to some outside entity instead. One gets a better idea of why this happens if we consider the structure of these sentences. a) would look like this:
[John] [said that he was sick]
whereas b) would look like this:
[John [said [that [he [was sick]]]]]
The reason is that in a) ‘he’ can refer to the NP John because ‘John’ is “higher up” in the phrase structure tree than ‘he’. In b), ‘He’ is part of the main NP, making ‘John’ part of the NP that is lower down (hierarchically) on the phrase structure tree; and anaphors cannot be bound to a referent that is lower down in the tree. (The whole point of phrase-structure trees, by the way, is to show hierarchical inclusion relationships. Given that even four-year-old children know these things, knowledge of this kind must be part of our genetic make-up, since children are clearly not taught things like this. It should come as no surprise then that every native speaker of English intuitively knows that b) is dis-ambiguated (though most would not know why), given that we have these phrase-structure rules pre-programmed into our minds. If we believe that language is only acquired from the environment in which the child is in, “it seems inevitable that major aspects of verbal behavior will remain a mystery”.
(Unfortunately I cannot illustrate the tree here, as this would have made the point more clear.)
Chomsky thinks there are many other principles which constrain the structure of language, and that the goal of contemporary linguistics is to discover exactly what the nature of these constraints are. Some tentative principles which he postulates are the Extended Projection Principle (EPP), and the principle of full interpretation. The above-mentioned example is an instance of the binding principle.
Chomsky says that it is not the case that structure-independent principles are too complex; in fact they are cognitively quite simple. “Rather, what seems to be going on is that the messages humans need to transmit to one another through the medium of language are messages that come in certain fixed forms, and these forms to a large extent dictate the syntactic patterning of human languages”. These invariant syntactic properties are determined by the mechanisms of genetic inheritance.
The obvious, common-sensical objection to the thesis that we have a Universal Grammar is: how do we explain the fact that not all languages are the same? Chomsky accommodates this disparity into his theory by explaining it in terms of “Principles and Parameters” (henceforth P&P).
This theory claims that the child is able to learn language so effortlessly because he is born with a set of principles, together with a “list” of (unset) language-specific parameters. When these parameters are set, these, together with the common principles that form part of our biological endowment, constitute what Chomsky calls “knowledge of language”. So we are born with rules governing sentence structure, like (to oversimplify) the following phrase-structure rule:
S -> {NP,VP}
The comma indicates variable order. When the child is about four years old, he realises that the language he is exposed to, language X, has the NP-first construction, and the genetically pre-programmed Language Acquisition Device crystallises this rule to read:
S -> NP VP
In language Y, however, the child realises that his language employs the VP-first construction, and now the genetically pre-programmed Language Acquisition Device crystallises this rule to read:
S -> VP NP
Hence, these two children may grow up speaking mutually unintelligible languages.
Obviously these are not the only rules that get crystallised during this process of so-called “parameter setting”, but we will leave it at that for the sake of simplicity.
This is the nativist’s explanation as to why language can be both innate, and variant across the human race. We will see below that this theory has problems which need to be addressed.
Nativists are wont to ask why we cannot form questions by interchanging the first and second words in a sentence, by reversing the order of words, or by picking out the middle word in a sentence.
What is wrong with turning the sentence
John ate an apple
into a question by saying
Apple an ate John?
Why does this feel so very unnatural? As pointed out, this is because language always makes use of structure-dependency. But if it is possible in principle for a language to have a rule like “Reverse the words in a sentence to form a question”, the fact that it never occurs needs explaining. This mystery can be explained if we assume that this restriction is hard-wired into our genes, and therefore restricted by our biology. If we do not accept this, then we cannot explain why the above-mentioned rule is so unnatural, and we would have to be at a loss in terms of being able to explain why rules of this type are so uncomfortable for us to utilise.
Descartes’ Problem
One can outline what Chomsky refers to as “Descartes’ problem” (cf. his book CARTESIAN LINGUISTICS) as follows:
How is it that we as native speakers are able to understand and produce a potentially infinite number of sentences, despite very often hearing/producing utterances that one has never before encountered in his life, or as Chomsky puts it, never before in the history of the universe. This Chomsky refers to as the creative aspect of language use (aka “Descartes’ problem” – so-called because it was Descartes, according to Chomsky, who alluded this problem).
In addition to this, Chomsky emphasises the fact that speakers are not able to express this variety of linguistic utterances due to being endowed with the relevant physiological apparatus (vocal tract, etc.), because a parrot is able to do the same, and he contends that the physiological mechanisms pertaining to the manifestation of sound are not necessarily in any way inferior to those of a human beings.
What, then, distinguishes a parrot’s ‘speech’ from human speech? Chomsky tells us that it is because there is something unique to the human mind which neither parrots nor any other (non-human) animal possesses. This ‘something’ is none other than our inborn Language Acquisition Device.
Aside from the fact that our innate rules give us the machinery to be ‘creative’, humans also have the ability to tell which sentences are grammatical, and which are not. How are children, faced with countless possibilities as to how to string sentences together, able to actually make sense of what they hear? If children were going on what they hear, they would need some way of knowing which sentences are acceptable, and which are not, and given that the possibilities are endless, how would children be able to cope with this mammoth task of learning the rules of language? Once again, Chomsky says that this would not be a mystery if we simply had a genetic predisposition to do so.
That pretty much sums up Chomsky’s case for his theory. His other writings serve only to supplement the points made above, with minor variations.
Let us now turn to the work of Steven Pinker, one of the leading exponents of neo-Chomskyan nativism. Pinker’s point of view will be taken primarily from his book THE LANGUAGE INSTINCT, and less so from a later book called WORDS AND RULES.
The Language Instinct
Here now are some of the key arguments found in Pinker, followed by my criticisms:
Phonotactic Constraints
Pinker claims that languages have rigid rules about how sounds are assembled into words. Many introductory texts call these “phonotactic constraints”. If you had to ask a native speaker of English about the grammaticality of certain words, he would somehow always know whether they are acceptable words or not. For example, words like “plaft” and “blick” are unequivocally judged as possible yet contingently non-existent words. Words like “sram” are unequivocally judged as impossible, and necessarily non-existent. This is because “plaft” and “blick” fit in with the existing sound system of English, whereas “sram” does not; it has an illicit onset sequence of “sr”. These constraints are part of our LAD, which allow us to set the parameters for our language at some relatively early stage of our linguistic development.
Infant Prodigies
Pinker says that children have to know what to expect when acquiring language, otherwise they would take forever to learn their mother-tongue. How would they know what to test for? How would they know what a verb’s inflexion depends on? Quoting Pinker, Sampson says that logically speaking, “an inflection could depend on whether the third word in the sentence referred to a reddish or a bluish object; whether the sentence was being uttered indoors or outdoors, and billions of other fruitless possibilities that a grammatically unfettered child would have to test for”.
Pinker then goes on to say that children never make the mistakes we would expect them to make if they learnt language from scratch. Pinker also refers to one of Stromswold’s Stromswold’s studies, based on the speech of pre-schoolers. She found out that even though there were false but logical analogies that could have been made by the children, they were never tempted to give in to these mistakes. Why does the child not commit a false analogy like He cans go based on I like going → He likes going; I can go → ???
Stromswold reported that the pre-schoolers never made a single false analogy. It is not the case that children never make mistakes when acquiring language. Pinker himself quotes the example of a father trying to correct his daughter from saying other one spoon to saying the other spoon, but to no avail. So what does Pinker want us to conclude from Stromswold’s work? It seems as if he is saying that the rule placeholders, such as ‘one’, can be followed by the noun for which it stands is a rule found in some languages, and since the child has not quite set her parameters to English yet, she utters possible, but (contingently) non-occurring utterances. However, no child will try out the idea that “can” will inflect to agree with its subject if it is in the third person; our innate knowledge precludes this kind of option, therefore children never commit this error.
Second Language Learning and the Critical Period
Pinker mentions Johnson and Newport’s research on Korean and Chinese immigrants to the USA. They found that the later the immigrants arrived, the less proficient they were at detecting grammatical errors. They also found that the poor immigrants’ attitude towards the language, and their motivation, had nothing to do with it; their age when they arrived had everything to do with it.
Pinker also mentions that despite the fact that Meryl Streep is renowned in the US for her accents, her English accent is “considered rather awful” in England, and that her Australian accent does not “go down too well in Australia either”. This is the case because she is trying to speak like a native speaker when the rules for her L1 have already crystallised, and she will therefore never be able to speak as proficiently as a native speaker.
The Language Mutants from Essex
One seemingly compelling case is that of one family who exhibited an anomaly known as SLI (Specific Language Impairment), whereby certain linguistic faculties are impaired, and their cognitive functioning in other areas were normal; the fact that it was carried over within the family shows that there certainly is a genetic element involved. Despite the fact that many of the ones who had SLI grew up in a “linguistically normal” environment, they still exhibited the anomalies that SLI sufferers have. Pinker also reminds us that they are cognitively normal in other domains.
The Structure of Words
Pinker claims that the structure of words also lends evidence to nativism in the sense that children’s minds seem to be designed with the logic of word-structure built in.
His first example is that of a pattern whereby the derivational affix can attach only to roots, not to stems (i.e. roots with at least one affix attached). When we try to attach a derivational affix to a stem, we get something like Darwinsian, which could refer to the works of Charles Darwin and his grandfather, Erasmus Darwin, whose work Charles drew profound inspiration from. However, this would sound ridiculous, not because it is illogical, but because there seems to be some sort of innate restriction on what affix can go where in a word. Later Pinker says that the restriction is that derivational suffixes cannot attach themselves to inflected roots.
Pinker also discusses headless compound words, and draws conclusions which provide further evidence for his case. An example of a headed compound would be barman, with the –man part constituting the head. Hence, a “barman” is a kind of man, viz. one that works in bars. Headless compounds are words like “walkman”. Here, -man cannot be the head because a “walkman” is not a kind of man. (Nor is it a kind of walk, but the rightmost morpheme is usually defined as the head of a compound).
Pinker’s explanation is that when a compound has a head, all the information regarding the head gets percolated up to the top-most node of the (morphological) tree, including the rules. When there is a headless compound, it is not clear where the information must come from, so the “percolation conduits” get blocked, and the default rule applies. That is the reason why we say “sabre-tooths” and “still-lifes”.
Other research cited by Pinker includes Peter Gordon’s experiment with 3-year-olds, commonly referred to as the “Jabba experiment” (Gordon, 1986). He showed them a puppet and said: this is Jabba. Jabba eats mice. Jabba is a ___? And the children reply: Mice-eater. Thereafter, he asks: Jabba also eats rats. Jabba is also a ___? And the children reply: Rat-eater.
Now, Gordon asks, why do they use the plural “mice”, but not the plural “rats”? Given that not a single child ever used the inflected plural as the initial constituent of the compound, this could indicate that there is some sort of rule that they are applying. It is unlikely that they were told that this is so, as Motherese contains few or no compounds, especially of that kind; and it would be preposterous to assume that someone told them not to include inflected words in their compounds. Why do they not generalise the rule: “create a compound by putting two words together”? Why should there be a restriction of this sort? The nativist would explain it as follows: if there were an innate rule of this kind to facilitate language acquisition, the child would already “know” some of the things that cannot be said, like including an inflected word in the first part of a compound word, which is obviously why they avoid constructions of this sort.
The results were interesting because there seemed to be a consistent restriction in terms of what kind of word was allowed to be in the first part of a compound word. “Mice-eater”, for example, was an acceptable form to them; their answer to the first question would be just that. However, they never answered “rats-eater” to the second question; they always answered “rat-eater”. This is interesting because if one goes solely on logic, then “rats-eater” ought not to be precluded, since “mice” and “rats” are both plural forms, and both (equally) morphologically complex.
The question, then, that Gordon posed is: what exactly makes forms like “rats-eater” ungrammatical? The reason he postulated was that the form “rats” is clearly derived from a rule: rat + -s. The word “mice”, however, is not; despite its morphological complexity, the word is learnt as a listeme, and therefore treated as such by the lexicon. Hence, “mice” in “mice-eater” is treated by the phonology as a simplex word.
This, Gordon claims, is an innate constraint placed on us by UG, and we have fairly good reason to assume this, given that children as young as three consistently do it.
There are problems with this claim, however, since there are languages which do exactly that which is claimed by Gordon to be universally excluded by an innate constraint. German, for example, makes use of regular plural forms in compounds. Consider the following set of data in the tables below:
English compound
German translation
1) Mice-eater
Mäusefresser
2) Rat-eater
Rattenfresser
3) Butterfly-eater
Schmetterling(s)fresserª
4) Insect-eater
Insektenfresser
5) Cat-eater
Katzenfresser
6) Flower-eater
Blumenfresser
7) People-eater
Menschenfresser
8) Tree-eater
Bäumefresser
9) Sheep-eater
Schaf(s)fresser
10) Fish-eater
Fischfresser
In German, -e(n) represents the plural suffix. To clarify, let us look at the singular/plural forms of the (German) words used in the data above:
Singular
Plural
1) Maus
Mäuse
2) Ratte
Ratten
3) Schmetterling
Schmetterlinge
4) Insekt
Insekten
5) Katte
Katzen
6) Blume
Blumen
7) Mensch
Menschen
8) Baum
Bäume
9) Schaf
Schafe
10) Fisch
Fische
In examples 9) and 10) [schaf(s)fresser and fischfresser], the singular form was preferred by my German informant; in examples 1) and 4) [mäusefresser and insektenfresser], both the singular and plural were judged to be acceptable forms, such that both mausfresser and insektfresser would not count as ungrammatical, though mäusefresser and insektenfresser are the more preferred of the two.
What is interesting about this is that examples 4-8 make use of the plural form in exactly the way Gordon would predict to be impossible, since here we have rule-derived words found in the position they should not be found in.
These alleged counter-examples can be dealt with if it can be shown that the phonology treats these morphologically complex words as simplex; in other words, the phonology is blind to the morphological complexity of these forms. Now we know that irregular plurals are certainly treated by the phonology in the same way as simplex listemes, but if we can show that this is true of not only irregulars, but other complex forms as well, maybe we can save Gordon’s thesis from these alleged counter-examples.
Jonathan Kaye, a phonologist from SOAS, has given us some novel insight into this issue of word structure. In discussing Kaye’s views, we will see why words cannot always be classified as either simplex or complex a binary fashion. Hence, let us now turn to Kaye's discussion of this issue (his theory is based on Government Phonology), so as to gain a better understanding of the matter at hand. (Note: the data here is once again based on English):
We cannot assume that the phonology is per se blind to morphological complexity, since the [mz] (from "dreams") is clearly a pseudo-cluster, and deemed both morphologically and phonologically complex.
It is clear, then, that at least some morphological structure is visible to the phonology. Kaye draws a distinction between what he calls analytic morphology and non-analytic morphology. Analytic morphology is morphology wherein the phonology is able to recognise a complex domain structure; non-analytic morphology is not able to. Let us first consider analytic morphology (which is visible to the phonology).
Analytic Morphology
As mentioned, the [mz] cluster from “dreams” is invariably complex. What does this tell us about morphological penetration into phonology?
Let us firstly assume that only morphological domains can be represented in the phonology. A compound like “black-board”, then, will have the following structure:
a) [[black] [board]]
Since there are three pairs of brackets, it reflects three phonological domains. These brackets represent how the phonological string is to be processed. To give an exact definition, we can use the CONCAT and φ (apply phonology) functions, so “blackboard” can be shown as
b) φ (CONCAT (φ(black) φ(board)))
To translate this pseudo-formula into plain language: apply phonology to “black”, then to “board”, concatenate the result, and then apply phonology to that string. So the brackets we see in example a) are actually not part of the phonological representation, but actually outline domains which are arguments to functions.
So this illustrates one type of structure that involves morphologically complex forms. It has the schema [[A][B]]. This is the case for most English compounds.
There is a second structure, involving two morphemes, but only two domains. The structure is [[A] B], and is interpreted as follows:
φ (CONCAT (φ(A), B))
To translate, we do phonology on A, and concatenate the result with B; then do phonology on the result. This is the case with English (regular) inflexional morphology. For example, consider the past tense of the word “seep”. Its structure is [[seep] ed]. With regards to Government Phonology, we could also illustrate this structure in more detail:
[[O N O N] O N]
x x x x x x x
s i: p d
Note that empty nuclei are to be found at the end of each domain, and are licensed to be silent by virtue of their domain-final position. Also, the licensing principle states that all positions in a phonological domain need to be licensed, unless that position is the head of that particular expression. This is actually why the suffix does not form a domain by itself. We see here that the suffix consists of two positions: an onset position and a nuclear position, both being licensed. If “-ed” were a domain, it is obvious that it would violate the licensing principle, which states that a phonological domain must have at least one unlicensed position – its head. The onset is licensed because, like all onsets, it is licensed by the nucleus to its right, which is itself p-licensed, i.e., licensed to be silent .
Analytic morphology also has a rather interesting property in that the distribution of empty nuclei is very restricted; in English, it is almost always found at the end of domains. Hence, this provides us with a fairly reliable parsing cue, since a p-licensed empty nucleus signals the end of a domain, so if it occurs within a word it shows that it is not a morphologically simplex one. So the pseudo-cluster [mz], from “dreams” is so-called because:
The vowel length is maintained (which would not be possible had it been an authentic cluster), and [m] and [z] are not homorganic, which again is not possible for true clusters. It is for this reason that much of analytic morphology is phonologically parsable.
So hitherto, we have considered two structures:
[[A][B]]
[[A] B],
and we have seen how morphology impacts on phonology in the form of these two domains.
Let us now turn to non-analytic morphology (which is invisible to the phonology).
Non-analytic Morphology
If morphological structure was present in the form of domains, then to that extent morphology can have an impact on phonology. These domains have the effect of respecting the integrity of the internal domains. By this, Kaye means that when you append an affix, for example, the actual pronunciation of the root/stem does not change. Consider the “-ing” suffix, whose form can be shown as [[V]-ing]. To pronounce this, you just pronounce the verb on its own and attach the suffix. This procedure does not apply to all forms of morphology. For example, take the words “parenthood” and “parental”; the respective suffixes interact with the phonology in different ways, as is evident from the variable pronunciation of the aforementioned words. In fact, the “-al” type of morphology is invisible to the phonology. In other words, the phonology reacts as if there were no morphology at all. Hence, the phonology would treat a word like “parental” in the same way it would treat a mono-morphemic word like “agenda”. So we can characterise the internal structure of a form like “parental” as:
[AB]
It is this kind of morphology that is referred to as non-analytic. And since the only effect morphology can have on phonology is the presence of internal domains, it follows that this kind of morphology can have no effect on phonology, since there are no internal domains. To further accentuate the disparity between analytic and non-analytic morphology, let us look at the morpheme “un-” (which is analytic), and the morpheme “in-” (which is non-analytic).
“un-” does not change its form, regardless of what consonant follows it. This property follows from its analytic morphology. Since the prefix “n” is not adjacent to the following onset, there is no restriction in terms of phonotactics. A word like “unreal”, then, can be represented as
[[un Ø] [ri:l Ø]],
where Ø represents the empty nucleus, whereby the first of which separates the two domains. The first empty nucleus provides us with a parsing cue, along with the nucleus in the following syllable, such that we are able to perceive the prefix as a separate structure. So a string segment like “nØr” is phonologically parsable, and one can analyse the form as “un-real”.
With the prefix “in-”, the situation is different. Appending “in-” to a stem must yield a well-formed phonological domain on its own; this is because non-analytical morphology, by definition, has no internal domains. A non-analytic combination of morphemes is interpreted as follows:
φ (concat (A,B))
You concatenate the two strings and then apply phonology to the result. So a form like “in-rational” is not well-formed because “nr” is not a possible sequence. That is why the “n” has to be dropped, resulting in the form “irrational”.
So non-analytic morphology is invisible to the phonology. There are no internal domains, or any other phonological indication of morphological complexity. Non-analytic forms have the same phonological properties as any simplex form. It is reasonable to assume, then, that any form manifesting non-analytic morphology is listed in the lexicon.
Irregular forms that exhibit morphological complexity are not phonologically parsable. This is because there are no phonological hints of their complex morphological structure. It is for this reason that irregular morphology is always non-analytic.
So with regards to the kind of compounds we are looking at, compounds can be classified into four groups, viz.:
Group 1 – Analytic, singular (example, tree-eater)
Group 2 – Analytic, plural (example, rats-eater)
Group 3 – Non-analytic, singular (example, linguistics-programme)
Group 4 – Non-analytic, plural (example, mice-eater)
In the groups above, ‘analytic’ refers to the compound itself, and ‘singular/plural’ refers to the modifying noun.
The relevance of his thesis is quite evident in that it may explain why the counter-examples to Gordon’s experiment are only pseudo-counter-examples. German, for example, does make use of compound words that contain a plural form. This seems, prima facie, a blatant refutation of Gordon’s claim. However, Kaye offers us some novel way to save the Gordon-hypothesis from refutation. My point here, however, is that Gordon’s argument as it stands is untenable, but can be saved in light of Kaye’s explanation, as the plural compound words have been shown to belong to our “Group 4” type words. However, German is peculiar in that the language has more irregular forms than regular ones, which leads one to the rather quaint conclusion that the ‘irregular’ plural seems to be more regular than the ‘regular’ plural inflexional suffix. Hence, the explanation for this phenomenon has nothing to do with genetic inheritance, but can be explained by appealing to the logic of word structure. Whether that is innate is a separate question, and Gordon would have to explain why languages do it differently. One could easily do so with a parameters-like explanation, but as I will explain later, this kind of approach is unscientific, and therefore vacuous. If there are other explanations that explain the same data more consistently and more logically, why should we accept the claim that this is unequivocally a result of innate constraints? For the Nativists, they insist that innateness is the only logical explanation for the phenomenon Gordon refers to; Kaye has shown otherwise. This concludes my discussion of Kaye.
A form like “mice-eater” is acceptable because the phonology is blind to its morphological complexity. A form may also not be irregular, and still be an acceptable part of the compound, as in the plural German compound words found in the data above. As Kaye has established, complex forms with no internal domains are treated by the phonology as if they were simplex. If German can be shown to exhibit non-analytic morphology with regards to using plural modifying nouns, they do not constitute true counter-examples; or rather, Gordon’s original hypothesis can be easily adapted to accommodate these facts without actually compromising the crux of his original thesis, which could now state that the restriction applies only to constituents that exhibit analytic morphology. For example, in a compound like rattenfresser, even though ratten is the (regular) plural of ratte, it actually is non-analytic, which means that the phonology is blind to the morphological complexity of the modifying noun.
On the Origin of Language
Chomsky is of the opinion that we would have to remain agnostic as to how language came in to being. In fact, Chomsky goes so far as to insist that we can “tell the story as we like”, and that Darwinian natural selection will not suffice as an explanation since there are no precedents anywhere in the natural world for language. Language, according to Chomsky, lacks the messiness we would expect of an accumulation of accidents made good by evolutionary ‘tinkering’. Characterised by beauty bordering on perfection, it cannot have evolved in the normal biological way. Pinker disagrees with this claim. In paper with Bloom (published in 1990) as well as in The Language Instinct, it is argued that natural selection is the only viable explanation for the origin of language. Pinker draws the analogy between language and spider webs; Pinker & Bloom say that language “is a topic like echolocation in bats or stereopsis in monkeys”.
Pinker hypothesises that there must have been a mutation of some kind in our genetic make-up, which happened to be useful to that particular individual. This trait would have been passed on, as natural selection retains traits which maximise the chances of survival, therefore the grammar gene would have been spread to become what it is today. Pinker is scanty on the details, and he himself concedes that this is speculative, but goes on say that the survival benefits are very obvious, and therefore it is quite plausible for a scenario of this kind to take place. How else would humans have worked together in groups to hunt wild-life far more powerful than themselves?
The hypothesis that grammar is historically the exclusive product of a special-purpose genetically provided mental organ for grammar rests upon a negative argument: we have not managed to explain how grammar arises through general processes. Hence, we must postulate the existence of a special genetic code that is entirely responsible for building in the brain a UG organ, which will someday be located and understood in the way the retina or the heart has been located and understood.
Linguistic Relativity
We are introduced on p. 57 of the THE LANGUAGE INSTINCT to "the famous Sapir-Whorf hypothesis of linguistic determinism", named after Edward Sapir and Benjamin Lee Whorf. According to Pinker, this is the notion that language determines the way we think, and the way in which we perceive the world. He illustrates this notion with the misunderstood statement that Eskimos have hundreds of words for snow, and therefore they somehow perceive snow differently to how English speakers would. This does not make sense, since it is actually not true that Eskimos have so many words for snow, and this does not take into consideration the fact that English has over a dozen words for snow. Clearly, it does not make sense to stipulate that Eskimos have different perceptual mechanisms just because they have a different language. He also mentions that just because some languages have less colour terms than, say, English, it does not follow that speakers of these languages do not perceive colour in the same way we do. The point Pinker is making here is simply that all human beings have the same cognitive faculties, and that we would all perceive the world in the same way, even though our linguistic conventions in naming that world may differ. In fact, Pinker notes, no matter how influential language may be, the idea that language can “reach down into the retina and rewire the ganglion cells” is preposterous.
ª The bracketed “s” is referred to as a “fugen-s ” [gapping/bridging s], so-called because it is there to serve as a bridge between the two constituents, and does not in itself carry any semantic significance.
My views on the NATURE-NURTURE debate - Part 3
WHAT DO BOTH CAMPS AGREE ON, AND WHERE DOES THE DISAGREEMENT COME FROM?
It is evident, then, that both camps agree that knowledge of language has to be a matter of degree. To quote Chomsky:
A preliminary observation is that the term “innateness hypothesis”
is generally used by critics rather than advocates of the position to
which it refers. I have never used the term, because it can only mislead.
Every “theory of learning” that is even worth considering incorporates
an innateness hypothesis. Thus, Hume’s theory proposes specific innate
structures of mind and seeks to account for all human knowledge on the
basis of these structures, even postulating unconscious and innate knowledge.
The question is not whether learning presupposes innate structure – of course
it does; that has never been in doubt – but rather what these innate structures
are in particular domains.
Elsewhere Chomsky says that
Any non-vacuous theory of language, whether rationalist, empiricist,
behaviorist, or whatever, must put forth an innateness hypothesis. The
interesting questions have to do with the precise character of this hypothesis,
not its existence as a necessary component of any non-vacuous theory.
What transpires from an exchange of this kind is that there are theorists who consider themselves either empiricists (like Sampson) or rationalists (like Chomsky and Pinker), and that it not a matter of contention whether there are any innate mechanisms, but rather what these innate mechanisms constitute.
It seems to be quite impossible to define rationalism and empiricism in a testable way, and that categorically distinguishes them. Due to these problems, the historical difference between the nativist and the empiricist will be ignored. When these terms are used here, “nativist” will be used to refer to Chomsky’s (and Pinker’s) view, whereas “empiricist” will be used to refer to Sampson’s.
It is evident, then, that both camps agree that knowledge of language has to be a matter of degree. To quote Chomsky:
A preliminary observation is that the term “innateness hypothesis”
is generally used by critics rather than advocates of the position to
which it refers. I have never used the term, because it can only mislead.
Every “theory of learning” that is even worth considering incorporates
an innateness hypothesis. Thus, Hume’s theory proposes specific innate
structures of mind and seeks to account for all human knowledge on the
basis of these structures, even postulating unconscious and innate knowledge.
The question is not whether learning presupposes innate structure – of course
it does; that has never been in doubt – but rather what these innate structures
are in particular domains.
Elsewhere Chomsky says that
Any non-vacuous theory of language, whether rationalist, empiricist,
behaviorist, or whatever, must put forth an innateness hypothesis. The
interesting questions have to do with the precise character of this hypothesis,
not its existence as a necessary component of any non-vacuous theory.
What transpires from an exchange of this kind is that there are theorists who consider themselves either empiricists (like Sampson) or rationalists (like Chomsky and Pinker), and that it not a matter of contention whether there are any innate mechanisms, but rather what these innate mechanisms constitute.
It seems to be quite impossible to define rationalism and empiricism in a testable way, and that categorically distinguishes them. Due to these problems, the historical difference between the nativist and the empiricist will be ignored. When these terms are used here, “nativist” will be used to refer to Chomsky’s (and Pinker’s) view, whereas “empiricist” will be used to refer to Sampson’s.
Thursday, January 1, 2009
My views on the NATURE-NATURE debate - part 2
Part 2 – PHILOSOPHICAL BACKGROUND
Concerning the theory of knowledge, there are two prominent schools of thought which go by the names of rationalism and empiricism. Empiricism can be described as a meta-theoretical school which illustrates how knowledge is acquired (from sensory experience) via inductive reasoning. Innate mental mechanisms are not denied, but restricted to domains for inductive generalization, and hence contribute nothing to knowledge content per se. Rationalism is the name of the opposing meta-theoretical paradigm. The actual schema of our knowledge is dictated by our innate cognitive structures; the content of our knowledge is to be found in the mind prior to sensory experience. The rationalist, then, defends nativism: the view that certain perceptual and conceptual capacities are innate. Chomskyan linguistic theory cannot be empiricist because it allows for unobservable grammatical properties to be stated as part of the speaker’s internalized linguistic competence. Also, empiricism cannot be equated with rationalism simply because they postulate innate mechanisms that enable the subject to apply procedures of hypothesis formation and inductive generalization.
These capacities may, however, lie dormant until the appropriate conditions for their manifestation arise; and the innate principles which determine the nature of thought and experience may be applied at an unconscious level of mind. Even the principles “of language and natural logic” are said to be known unconsciously at birth, and are “in large measure a precondition for language acquisition”, to quote Chomsky.
One of my freshman philosophy lecturers, Darrel Moellendorf, once told me that Chomsky sees himself as the heir to the rationalist’s throne. Indeed, he [Chomsky, not Darrel!] explicitly accepts their doctrine of innate ideas. Despite the great many languages spoken in the world, Chomsky says that they sufficiently resemble each other syntactically to suggest that there must be a universal schema, which he refers to as Universal Grammar (or UG). These are determined by the genetic constraints in the human mind. These innate constraints set the pattern for all experience, and fix the rules for the formation of meaningful sentences, and also explain why languages are readily translatable into one another. But I will return to Chomsky’s theory in more detail later.
The issue of innate knowledge vs empirical knowledge (with regards to language learning) is clearly a matter of degree. If all knowledge was derived from sensory experience, then a parrot that grew up in the same environment as a human child (and, ideally, went through the same experiences) should eventually acquire a level of language proficiency tantamount to that of the child. This clearly does not happen, even though the parrot is physiologically able to produce sounds, as Chomsky (quoting Descartes) points out in his famous book, Cartesian Linguistics, published in 1966. On the other hand, if all the details of our linguistic knowledge were hard-wired into our genes, then all humans would speak one language throughout the world. Chomsky believes that the environment does play a role in language acquisition, but it is only there to manifest the dormant linguistic principles that are latent in the acquiring child’s mind, and to set the language specific parameters. On the other hand, the empiricist holds that there are no domain specific structures extant at birth: their innate machinery is restricted to hypothesis formation and inductive reasoning, and this may well be a faculty unique to human beings.
Both Pinker and Chomsky are guilty of arguing against straw-man versions of empiricism, based on the long out-dated school of behaviourism. B. F. Skinner, for example, was of the opinion that we did not even have to postulate innate mechanisms of any kind. A living organism is a living organism. We do not have any reason to believe that there is something other than environmental conditioning that distinguishes them. If a pigeon is able to play ping-pong, he reasoned, why must we believe that there is something about the organism (other than the obvious phenotype) which categorically distinguishes it from a human being? I do not think any theorist, who also calls himself an empiricist, would agree that Skinner accurately represents what empiricism purports. Linguists tend to think so because in Chomsky wrote a scathing and influential review in 1959, where Skinner, who happens to be the leading exponent of classical Behaviourism, was the victim of Chomsky's venom; the review was of Skinner’s book entitled Verbal Behavior, published in 1957. He had been wise enough not to take issue with, say, the school of child psychology pioneered in the Soviet Union by Lev Vygotsky, or the subtle and fruitful approach adopted by the Swiss developmental psychologist Jean Piaget. I contend that Chomsky chose his target deliberately here, as he characteristically does – this is why you never find Chomsky making any references to those who oppose his theories. He is well aware of Popper, for example, but does not apply his standards of testability; when I asked him what he thought of the nascent field of Cognitive Linguistics, he curtly said that he didn’t “know much” about it! Does he really, or is it the case that if Cognitive Linguistics were to take off, that would be the end of Generative Grammar as we know it? In acknowledging this, Chomsky would have to either concede that his empire is crumbling, or he would have to radically review his theory, neither of which are viable alternatives for him. I also think this is the reason why he is writing less and less about linguistics, and more about politics of late. And this is also why I think Steven Pinker has suddenly hit the brakes and made a hundred and eighty degree turn in his latest book, THE STUFF OF THOUGHT. However, that’s a topic for another discussion, so let me stop digressing now and get back to the topic at hand...
My point here is simply that despite major differences with psychoanalysis, these psychologists like Vygotsky and Piaget had echoed Freud in taking as read that humans, like other animals, must have deep-rooted instincts of some relevance to a study of the mind. Chomsky, however, refrained from acknowledging the existence of such scientists. By singling out behaviourism for attack and ignoring everything else, he succeeded in arranging the battleground to suit his own needs.
One gets the impression that Chomsky always conceives of behaviourism in the image of Skinner. In fact, Quine’s behaviourism, for example, only remotely resembles what is usually referred to as behaviourism. Quine did not advocate the view that things can or should be explained in terms of correlations between stimuli and subsequent responses to these stimuli. Quine is even quoted as saying, in his contribution to the symposium on linguistic rationalism and innateness: “It may well turn out that processes are involved that are very unlike the classical processes of reinforcement and extinction of responses. This would be no refutation of behaviourism, in a philosophically significant sense of the term; for I see no interest in restricting the term ‘behaviourism’ to a specific psychological schematism of conditioned response”.
Not only are Chomsky and Pinker wont to capitalize on the fact that most academics are reluctant to peruse in great detail the relevant source texts, but pro-Chomskyans also cover up for untenable statements make by Chomsky in his earlier writings. In an understandably irritated tone, Peter Carr writes in his 2003 paper on Chomskyan rationalism and its problems, where he refers to Smith, who claims that we are all mistaken in attributing the notion that language is a set of well-formed sentences to Chomsky, since that was apparently put forth by Chomsky as a view that he was attacking. Carr then asks us to consider what Chomsky could have meant when he said that “From now on, I will consider language to be a set of (finite or infinite) sentences”, and furthermore when Chomsky said in a later work that “A fully adequate grammar must assign to each of an infinite range of sentences a structural description indicating how this sentence is understood by the ideal speaker-hearer”. Carr then sarcastically concludes that “If this amounts to outlining a position to be attacked, it certainly does not read that way”!
Sampson is not guilty of the rhetoric that is so rife in many of the nativist's manifestos. He says that he has tried to represent the nativist’s case fairly and without distortion, and I think he has indeed succeeded. Sampson concedes that Chomsky was quite justified in his criticism of Skinner. However, to quote Sampson: “to treat Skinner’s unreasonable theories as representative of the centuries-old tradition of empiricist thought is a travesty”, and to expect the world to believe in innate knowledge because “some half forgotten psychology professor did not believe in minds is a bit rich”.
Concerning the theory of knowledge, there are two prominent schools of thought which go by the names of rationalism and empiricism. Empiricism can be described as a meta-theoretical school which illustrates how knowledge is acquired (from sensory experience) via inductive reasoning. Innate mental mechanisms are not denied, but restricted to domains for inductive generalization, and hence contribute nothing to knowledge content per se. Rationalism is the name of the opposing meta-theoretical paradigm. The actual schema of our knowledge is dictated by our innate cognitive structures; the content of our knowledge is to be found in the mind prior to sensory experience. The rationalist, then, defends nativism: the view that certain perceptual and conceptual capacities are innate. Chomskyan linguistic theory cannot be empiricist because it allows for unobservable grammatical properties to be stated as part of the speaker’s internalized linguistic competence. Also, empiricism cannot be equated with rationalism simply because they postulate innate mechanisms that enable the subject to apply procedures of hypothesis formation and inductive generalization.
These capacities may, however, lie dormant until the appropriate conditions for their manifestation arise; and the innate principles which determine the nature of thought and experience may be applied at an unconscious level of mind. Even the principles “of language and natural logic” are said to be known unconsciously at birth, and are “in large measure a precondition for language acquisition”, to quote Chomsky.
One of my freshman philosophy lecturers, Darrel Moellendorf, once told me that Chomsky sees himself as the heir to the rationalist’s throne. Indeed, he [Chomsky, not Darrel!] explicitly accepts their doctrine of innate ideas. Despite the great many languages spoken in the world, Chomsky says that they sufficiently resemble each other syntactically to suggest that there must be a universal schema, which he refers to as Universal Grammar (or UG). These are determined by the genetic constraints in the human mind. These innate constraints set the pattern for all experience, and fix the rules for the formation of meaningful sentences, and also explain why languages are readily translatable into one another. But I will return to Chomsky’s theory in more detail later.
The issue of innate knowledge vs empirical knowledge (with regards to language learning) is clearly a matter of degree. If all knowledge was derived from sensory experience, then a parrot that grew up in the same environment as a human child (and, ideally, went through the same experiences) should eventually acquire a level of language proficiency tantamount to that of the child. This clearly does not happen, even though the parrot is physiologically able to produce sounds, as Chomsky (quoting Descartes) points out in his famous book, Cartesian Linguistics, published in 1966. On the other hand, if all the details of our linguistic knowledge were hard-wired into our genes, then all humans would speak one language throughout the world. Chomsky believes that the environment does play a role in language acquisition, but it is only there to manifest the dormant linguistic principles that are latent in the acquiring child’s mind, and to set the language specific parameters. On the other hand, the empiricist holds that there are no domain specific structures extant at birth: their innate machinery is restricted to hypothesis formation and inductive reasoning, and this may well be a faculty unique to human beings.
Both Pinker and Chomsky are guilty of arguing against straw-man versions of empiricism, based on the long out-dated school of behaviourism. B. F. Skinner, for example, was of the opinion that we did not even have to postulate innate mechanisms of any kind. A living organism is a living organism. We do not have any reason to believe that there is something other than environmental conditioning that distinguishes them. If a pigeon is able to play ping-pong, he reasoned, why must we believe that there is something about the organism (other than the obvious phenotype) which categorically distinguishes it from a human being? I do not think any theorist, who also calls himself an empiricist, would agree that Skinner accurately represents what empiricism purports. Linguists tend to think so because in Chomsky wrote a scathing and influential review in 1959, where Skinner, who happens to be the leading exponent of classical Behaviourism, was the victim of Chomsky's venom; the review was of Skinner’s book entitled Verbal Behavior, published in 1957. He had been wise enough not to take issue with, say, the school of child psychology pioneered in the Soviet Union by Lev Vygotsky, or the subtle and fruitful approach adopted by the Swiss developmental psychologist Jean Piaget. I contend that Chomsky chose his target deliberately here, as he characteristically does – this is why you never find Chomsky making any references to those who oppose his theories. He is well aware of Popper, for example, but does not apply his standards of testability; when I asked him what he thought of the nascent field of Cognitive Linguistics, he curtly said that he didn’t “know much” about it! Does he really, or is it the case that if Cognitive Linguistics were to take off, that would be the end of Generative Grammar as we know it? In acknowledging this, Chomsky would have to either concede that his empire is crumbling, or he would have to radically review his theory, neither of which are viable alternatives for him. I also think this is the reason why he is writing less and less about linguistics, and more about politics of late. And this is also why I think Steven Pinker has suddenly hit the brakes and made a hundred and eighty degree turn in his latest book, THE STUFF OF THOUGHT. However, that’s a topic for another discussion, so let me stop digressing now and get back to the topic at hand...
My point here is simply that despite major differences with psychoanalysis, these psychologists like Vygotsky and Piaget had echoed Freud in taking as read that humans, like other animals, must have deep-rooted instincts of some relevance to a study of the mind. Chomsky, however, refrained from acknowledging the existence of such scientists. By singling out behaviourism for attack and ignoring everything else, he succeeded in arranging the battleground to suit his own needs.
One gets the impression that Chomsky always conceives of behaviourism in the image of Skinner. In fact, Quine’s behaviourism, for example, only remotely resembles what is usually referred to as behaviourism. Quine did not advocate the view that things can or should be explained in terms of correlations between stimuli and subsequent responses to these stimuli. Quine is even quoted as saying, in his contribution to the symposium on linguistic rationalism and innateness: “It may well turn out that processes are involved that are very unlike the classical processes of reinforcement and extinction of responses. This would be no refutation of behaviourism, in a philosophically significant sense of the term; for I see no interest in restricting the term ‘behaviourism’ to a specific psychological schematism of conditioned response”.
Not only are Chomsky and Pinker wont to capitalize on the fact that most academics are reluctant to peruse in great detail the relevant source texts, but pro-Chomskyans also cover up for untenable statements make by Chomsky in his earlier writings. In an understandably irritated tone, Peter Carr writes in his 2003 paper on Chomskyan rationalism and its problems, where he refers to Smith, who claims that we are all mistaken in attributing the notion that language is a set of well-formed sentences to Chomsky, since that was apparently put forth by Chomsky as a view that he was attacking. Carr then asks us to consider what Chomsky could have meant when he said that “From now on, I will consider language to be a set of (finite or infinite) sentences”, and furthermore when Chomsky said in a later work that “A fully adequate grammar must assign to each of an infinite range of sentences a structural description indicating how this sentence is understood by the ideal speaker-hearer”. Carr then sarcastically concludes that “If this amounts to outlining a position to be attacked, it certainly does not read that way”!
Sampson is not guilty of the rhetoric that is so rife in many of the nativist's manifestos. He says that he has tried to represent the nativist’s case fairly and without distortion, and I think he has indeed succeeded. Sampson concedes that Chomsky was quite justified in his criticism of Skinner. However, to quote Sampson: “to treat Skinner’s unreasonable theories as representative of the centuries-old tradition of empiricist thought is a travesty”, and to expect the world to believe in innate knowledge because “some half forgotten psychology professor did not believe in minds is a bit rich”.
My views on the NATURE-NATURE debate - part 1
Part 1 - INTRODUCTION
I have decided to post my views on this rather contentious topic as there are some interesting developments in the field. I will have to deal with this in parts because the topic is so very vast that there is just too much to write, summarise and post in one go.
The aim of this is to investigate the viability of the Chomskyan paradigm in light of Geoffrey Sampson’s account, as outlined in his Educating Eve. After looking at both arguments, I will conclude that Sampson does rightly criticise the nativists, whose complacency can indeed impede scientific progress. I will do this by critically evaluating some evidence proposed to bolster the Chomskyan account, in addition to that which Sampson offers.
The method here is essentially a negative one. The nativists claim that there is evidence that language acquisition occurs in a special way, which distinguishes it from other kinds of learning. However, this evidence will be shown to be flawed in some way or the other.
Sampson’s paradigm is indeed very useful in that it precludes the kind of complacency that is so dominant in the intellectual climes of today. Nativism needs to be looked at with a more critical eye; we should not take their purports as axiomatic truths. The flaws to be found in Chomsky’s theory are a problem regarding the modus operandi of research done; for example, question begging and looking for confirming evidence are meta-theoretical flaws which permeate much of the research done in the nativist framework, and need to be remedied if scientific progress is to be made. I will illustrate these meta-theoretical short-comings with specific examples from Chomsky and Pinker, and I will provide suggestions as to how the status quo can be improved.
Sampson’s representation of Popper can also be called into question, as Popper did indeed believe in innate linguistic mechanisms. This fact does not damage Sampson’s thesis though, so I will deal with this disparity only at the end.
I have decided to post my views on this rather contentious topic as there are some interesting developments in the field. I will have to deal with this in parts because the topic is so very vast that there is just too much to write, summarise and post in one go.
The aim of this is to investigate the viability of the Chomskyan paradigm in light of Geoffrey Sampson’s account, as outlined in his Educating Eve. After looking at both arguments, I will conclude that Sampson does rightly criticise the nativists, whose complacency can indeed impede scientific progress. I will do this by critically evaluating some evidence proposed to bolster the Chomskyan account, in addition to that which Sampson offers.
The method here is essentially a negative one. The nativists claim that there is evidence that language acquisition occurs in a special way, which distinguishes it from other kinds of learning. However, this evidence will be shown to be flawed in some way or the other.
Sampson’s paradigm is indeed very useful in that it precludes the kind of complacency that is so dominant in the intellectual climes of today. Nativism needs to be looked at with a more critical eye; we should not take their purports as axiomatic truths. The flaws to be found in Chomsky’s theory are a problem regarding the modus operandi of research done; for example, question begging and looking for confirming evidence are meta-theoretical flaws which permeate much of the research done in the nativist framework, and need to be remedied if scientific progress is to be made. I will illustrate these meta-theoretical short-comings with specific examples from Chomsky and Pinker, and I will provide suggestions as to how the status quo can be improved.
Sampson’s representation of Popper can also be called into question, as Popper did indeed believe in innate linguistic mechanisms. This fact does not damage Sampson’s thesis though, so I will deal with this disparity only at the end.
Sunday, August 17, 2008
A Comparative Analysis of Generative and Cognitive Grammar in light of Popper’s meta-theory
Introduction
This proposal aims to compare two opposing accounts of how language and the brain interact; these go by the names of generative linguistics and cognitive linguistics. I will outline the basic assumptions common to most mainstream adherents of these fields, together with the conclusions they are obliged to accept in light of their respective theoretical commitments. After assessing various aspects of these theories, they will be compared using the standard criteria of what ought to constitute a scientific theory, and thereby conclude that the cognitive approach seems to be more empirically and logically plausible for various reasons.
Generative Linguistics
The school of linguistics known as Generative Linguistics (henceforth GL) was started by numerous collaboraters, many of whom have since either abandoned the enterprise as initially outlined, or gone on to expand the theory in novel ways. Those who have abandoned the ship include thinkers like George Lakoff and Haj Ross, whereas those who expanded the theory in ways not originally envisioned include Ray Jackendoff, Steven Pinker and Jerry Fodor.
The central figure since the inception of the GL enterprise in the late 1950’s has been Noam Chomsky, and his position regarding the nature of language, how it ought to be studied, and its relation to human nature, has remained fundamentally unchanged over the years.
The main arguments within the GL framework have come from Noam Chomsky and Steven Pinker, and therefore their work will be drawn upon more heavily than others in outlining the said theory. Others, who have accepted the core assumptions and methodologies, will be alluded to as well, but will not constitute the core of this thesis.
The key assumptions that pervade this approach are as follows:
Language is innate, in the sense that it is part of an autonomous module (a mental organ) in the mind that grows in the child the same way your limbs grow.
Language learning or acquisition are therefore misnomers; language growth is a more appropriate term.
The physical manifestation of language (our linguistic performance) is only to be studied insofar as it provides clues to the language in our minds (our linguistic competence). We should therefore assume that a speech community is homogenous, and language should be studied as if it were part of such an idealized community.
For this reason, we are to draw a distinction between surface structure and deep structure linguistic forms.
Language is a discrete combinatorial system, and this fact allows us to use language in novel ways, which they refer to as creativity.
Language is at the core a syntactic phenomenon, which employs various finite rules which enable the user to produce a potentially infinite number of utterances. This process works like an abstract algorithm, and therefore functions independent of context and meaning. Hence, semantics should not only be seen as secondary, but as epiphenomenal.
As a genetically based module in the mind, language must have evolved at a single point in our history, either by Darwinian natural selection (Pinker) or by some kind of small genetic mutation, which somehow interacted with the mind’s computational principles (Chomsky), which gave rise to language.
Language has no direct influence on our thoughts, as thought and language are two separate processes which function separately in the mind.
Since language is part of our biological endowment, there must be a Universal Grammar (UG) present in all languages, which manifests as a result of being part of our minds at birth, providing the scaffolding that enables language to grow in the mind.
Just as our physical organs interact on some level with other organs, they are nonetheless structurally and functionally independent of them. Language, being a mental organ, should also be expected to be understood independently; language should not be intertwined with other aspects of our mental life, any more than the heart is intertwined with the functioning of the liver.
Language is understood in light of Rationalism, whereby thought is believed to be conscious and literal.
Generative Linguists are vehemently opposed to Empiricism and behaviourism.
The human mind is a uniform, modular organ.
These principles are certainly not exhaustive, but I have selected those which I seem to be the most definitive.
There are certain things that follow from this approach to the study of language and mind. Generally, accepting these premises obliges one to accept the conclusions which follow. Regarding the study of language, this includes:
Having to concede that second language learning is not at all related to the growth (sic) of your first language.
Language teaching is not related to the primary goal of linguistics.
Socio-linguistics, prescriptive grammar, speech therapy, etc. are not of primary concern to the linguist.
Reading and writing not part of language; they are cultural artefacts just like painting and dancing.
Documenting corpora is a pointless exercise, since the range of potential utterances are infinite. The goal should be to try and discover what the universal, finite rules are which all languages employ. The goal of linguistics as a whole has to be redefined, and those who pursue corpus linguistics are wasting their time.
Language is not for communication. Language appeared in the mind somehow, and we discovered that it could also be used for communication. Where it came from, and the real reason it manifested, are unknown.
Language is to be seen as part of the physical world, and needs to be studied as such.
In addition, GL uses some words in deviant ways. For example:
- creative refers to the simple act of speaking
- language is no longer seen as something spoken by a community; it is more accurately seen as an underlying mental competence, which grows in the mind and is the same in every human being. Recently Chomsky has taken to distinguishing between i-language, e-language, and Language, which adds to the confusion
- mind is merely a synonym for brain, but seen more as aspects of the brain which we know must be there
- the use of the word grow to refer to the development of a first language in the child; if the child is exposed to three languages as a young child, then three languages will “grow” in him. Generative linguists do not see this as odd, and they would be quite comfortable with an analogy along the following lines: if a child grows up gathering rocks and hunting, his hands will be stronger and more versatile; likewise, if a child grows up hearing different languages, his innate language faculty grows to be more versatile.
Also, it is not a coincidence that people like Pinker tend more towards a conservative outlook, whereas people like Lakoff tend more towards a progressive outlook. The idea that we are all the same at the core, that our linguistic and cultural disparities are just superficial, that our minds come fully equipped with genetic information, etc. means we do not need a government to see to our individual needs, since we are capable of using our innate potential to make a success of our lives. It is this mind-set that justifies the right-wing extremist view that the world should be run under one government, and any country or nation which questions it simply does not realize that we are all the same. This is the reason why democracy needs to be superimposed on Middle Eastern countries, whether they like it or not. Saudi Arabians, who are perfectly happy with their king, simply do not realize that American democracy is the best way of running a country.
Why democracy? Or why America’s brand of democracy? Well that is just a coincidence arising from the fact that they happen to be the first country to realize this. Countries like Dubai and Saudi Arabia just do not know what is good for them. One may wonder why, if “self-evident” civil liberties are innate, do they not know this; the answer: if a child is raised with his hands tied behind his back, and that is all he knows, he will not realize that they were meant to be free until some benevolent person (or institution) unties his hands, and forces the society to ban such practices. One of the conservative arguments justifying the war in Iraq goes along these lines.
This mind-set is the very same one which insists that some of the rules which pertain to American English must be part and parcel of human nature. Not only should this apply too every other dialect of English, and to the some six thousand other languages in the world, but every human being is born with this UG. If it does not appear that way, it is only because that principle must not have a surface manifestation – or maybe that principle is actually a parameter. In fact, at one point, if you did not accept this, you could hardly expect to find a job or get a research grant in the USA.
I am certainly not insinuating that the conservatives in Washington read The Blank Slate, got over excited, and invaded Iraq in an attempt to spread democracy and freedom in the Middle East, knowing now that the notion of universal liberty can be justified. I am merely drawing a parallel between these ways of thinking, and the methodologies they employ in reasoning. Just as you get different forms of democracy in the world, the principles underlying a democratic government are the same, just as with language.
Pinker refutes this by saying that political equity simply means being treated equally, not being literally equal in the biological sense. In fact, Pinker does not believe that we are all equal. In addition to his conviction that men and women have different skills and abilities (though he is careful not to hierarchialise this), Pinker also believes that there could be a racial correlate regarding our skills and abilities. During an interview with John Brockman, he said that science may one day discover that certain race groups are better at certain things than other things, which may be something we are not ready for (interview available at http://www.edge.org/q2007/q07_index.html). The interview was based on the theme Dangerous Ideas, and Pinker based his statement on a book by Daniel Dennett called Darwin’s Dangerous Idea. After pointing out how Darwin was ostracized, and how the theory shook the very foundations of the church, he says that the next shock to us will be the said discovery. He was very careful not to spell out the details, but I doubt he was implying that we may discover that Black Africans are more intellectual, and that White American males are better at manual labour.
As an aside, it is a well-known fact that the HSRC used to publish very convincing statistics during the apartheid days ‘proving’, for example, that Black people are biologically designed for intellectually inferior tasks.
The various forms of evidence cited to support these views will be discussed and criticized scientifically. I contend that the ideas on which GL is based are methodologically and empirically flawed. The same criticism applies to other paradigms which utilize the same lines of reasoning, like the conservative politicians.
These upshots and underlying assumptions are not accepted universally by linguistics scholars for various reasons. Scholars like Geoffrey Sampson rejects every argument and every logical implication of this approach; he has argued very effectively against GL in his various works over the years, like in his book Educating Eve, and various papers like Popperian Language Acquisition Undefeated, Minds in Uniform, Revisiting the ‘poverty of stimulus’ argument, etc.
His approach is a negative one, where the nativists say they have evidence and arguments for an innate language module in the brain, and Sampson points out that their evidence and arguments are flawed. These will be outlined here, but the problems with Sampson’s own account will be highlighted. His representation of Popper is not accurate, his alternative account of language acquisition and human nature is simplistic and reductionistic, and there seems to be a disparity between his academic work and his political views. For example, he got into trouble whilst an MP for saying that Asians are better at Maths because of their genes; in Minds in Uniform he takes issue with forcing a little country to abolish advocating the death sentence for speaking out against their country’s leader (cultural relativism).
Sampson is however a liberal, and his alternative theory can be supplemented by what has come to be known as Cognitive Linguistics (henceforth CL). As an aside, Sampson, like Chomsky, admits that he does not know anything about CL.
CL as a whole was born as a reaction to what is still referred to as “mainstream linguistics”, ie. Generative Linguistics. The question I will ask here is whether this reaction was justified, as is entailed rejecting all the key assumptions upon which GL rests, and in fact virtually all of Western philosophy as well, if you accept the import of George Lakoff’s and Mark Johnson’s Philosophy in the Flesh. For such a paradigm shift to be justified, the following has to be clearly demonstrated:
Each theory will have to be independently evaluated in light of evidence; refuting evidence should be taken seriously
Each theory will also need to be assessed in terms of other methods of argumentation employed to justify their stance
Where there is a deviation from the norm, either by way of reinventing terminology, postulating counter-intuitive hypotheses, or postulating more hypotheses than necessary, it needs to be adequately justified
The implications of each theory, both within its own field and in cognate disciplines, need to be assessed as well
Aspects of each theory which may be metaphysical in nature, need to be assessed in terms of its relative explanatory power
Are there aspects outside of the theory and evidence which influences the theory, like ad homonim considerations, socio-political factors, etc.?
These are the standard criteria used when deciding which theory is a better one. It is only a coincidence that I will use the works of Karl Popper in doing so, as he is one of the few philosophers of science to systematically outline what the logic of scientific inquiry entails, though few would disagree with his approach. Had I used Einstein, Newton, Feynman, Penrose, Putnam, or any other serious scientific thinker, the criteria would be the same.
The evidence and methodology used by the GL school specifically will be assessed. If their hypotheses make predictions, I will test these predictions, and see how the theory handles counter-evidence. If the theory can be always adapted to assimilate counter-evidence, then it is problematic. If the theory can construct ad hoc hypotheses without compromising testability, it would be better. If new concepts and terms are invented to account for this, then these new concepts have to be justified in light of the evidence, not just to protect the theory. If counter-evidence is ignored all together, then the theory is not an empirically responsible one, and therefore not scientific.
Before doing this, we need to understand a little about the alternative theory, which we will turn to now:
Cognitive Linguistics
CL was born as a reaction to GL, after exponents like Lakoff initially tried to expand Chomsky’s theory to include semantics, which at the time went by the name Generative Semantics. As more things came to be known, the idea of trying to incorporate all evidence within a GL paradigm became more and more difficult, and a ‘breakaway faction’ ensued, resulting in what was documented by Harris as The Linguistics Wars. Lakoff was a student of Chomsky’s, and the latter seemed to take this as a betrayal. After Lakoff’s criticisms, Chomsky is on record as saying that Lakoff does not understand linguistics very well. Today, Chomsky claims that he knows nothing about Cognitive Linguistics!
Anyway, while Lakoff is one of the exponents, CL cannot be attributed to a single ‘founding father’. It is a collaborative movement which is always evolving, though there are theorists who are prominent within the field. As a result, there are various branches of CL, and not every linguist who calls himself a CL has to accept all the assumptions that others espouse. Hence, delineating criteria which distinguish the field can lead to some fuzziness.
Regardless, these are some of the key assumptions shared by CL as I understand it:
Language may or may not be innate. It remains to be seen how much is innate if so. Regardless, evidence thus far shows that language is part and parcel of other higher cognitive functions, and is neither an autonomous mental organ, nor does it grow – it is learnt via interaction with the environment and the mind, using all its capacities.
Language as expressed and transcribed is one our most important sources of data, and should be studied as such without assuming that each speaker is homogenous, and that his language does not change in any way.
The idea of every sentence having a deep structure is thereby also rejected, since the reasons for postulating it in the first place are not clear.
Language is at the core a semantic phenomenon which cannot be divorced from context. It makes evolutionary sense that parable and narrative evolved first, followed by other aspects of language. Otherwise, nothing precludes a computer from simulating human language.
Language could not have evolved the same way our physical organs evolved, since that assumes that natural selection has a vision to the future; if it evolved with other cognitive faculties and manifested initially as parable, it would have immediate survival value.
Language and thought are inextricably linked, so much so that without language our level of consciousness would be reduced.
Though there is an innate basis for language, it has to be learnt via hypothesis formation. Hence, the will be commonalities constrained in trivial ways by our biology, but nothing to suggest that there is a UG hard-wired into our brains.
Language is an abstract cognitive faculty, which is neither functionally nor cognitively separate from the rest of our other cognitive faculties. This predicts that damage to our language centres in the brain would affect other things as well, and that language proficiency is conducive to clear thought (and of course, that linguistic proficiency varies from individual to individual)
Language is by and large figurative (or metaphorical, in the larger sense of the word), and thought is mostly unconscious.
Empiricism and behaviourism are not rejected by fiat, but aspects of either are assessed on their merit.
Language and mind are not necessarily part of the material world. While not saying so explicitly saying so, they leave a gap in their paradigm to cater for the possibility that they might be good candidates as possible denizens of Popper’s World 3.
Once again, these assumptions are not exhaustive, and the evidence will be duly scrutinized, as with GL. As mentioned, this school was born as a reaction to mainstream linguistics, and as such will be expected to handle the evidence more responsibly, as they claim, should be commensurable with common sense, and the rejection of things like the primacy of syntax will be looked at.
Stylistic issues like the novel use terminology will be assessed as either bad style or bad science, since CL scholars do not do this. The question will be asked: are there any theoretical implications leading to pseudo-problems[1] as a result of this? If so, would the problem still be there if these additional assumptions were not made? (For example, the concept of creativity in the conventional sense does not exist anymore, for generativists – though this is more of a semantic reductionism, which is the converse…)
To avoid redundancy, the entailments of these assumptions will not be reiterated as they are pretty much the converse of the opposing paradigm, as expected.
Regarding politics, CL seem to have a less euro-centric view of the world, and are progressive in their thinking. They do not feel the need to create a hierarchy, either in the academic world or the political world. This is not only because of George Lakoff’s now defunct progressive think tank, The Rockridge Institute, but because there are explicit links between the two (unlike GL, where the links are implied and never stated explicitly; in Chomsky’s case, he explicitly denies the link).
As mentioned, there are parallels between conservative political arguments and the way GL argue. The correlation does not entail causation, as any statistician would confirm. This merely shows a parallel in thinking, with possible causation. For example, the reason gay marriage is opposed is argued for as follows:
If they allow this, then what is next? What’s to stop from condoning bestiality, and then allowing them to get married!
Or their argument against raising the minimum wage:
If you give them a two dollar raise, what is to stop the unions from demanding a hundred dollar raise the next time round!
For some reason, it has to be one extreme or the other. Gay marriage not creating a terrible problem in South Africa is not considered as counter-evidence. A colleague told me: Wait and see what happens in five years time! When I pointed out that this has been happening for five years, he did not feel the need to reevaluate his position, but without reason. Likewise, Pinker would tell you that you either accept that there is a UG, with built in x-bars and island constraints, hard-wired into our minds (and in fact many other alleged things not related to language, as he explains in The Blank Slate, How the Mind Works, and The Stuff of Thought), or that we believe that the mind is a blank slate no different from a rock – he says we must either believe there is a human nature (and therefore accept his argument in its entirety), or totally reject the idea of the existence of human nature. He is careful to quote John Locke as advocating the “blank slate” idea, knowing that few people have read his works[2], but then adds that Locke does not deserve the blame for the idea that most intellectuals today believe that our minds are empty buckets and that people are hapless victims of their environment. As with the “gay marriage” argument, the alternative does not follow; likewise, you can reject GL and still believe that there is such a thing as human nature. Likewise, Chomsky has stated that the existence of UG is unquestionable, unless you believe in magic; clearly, it is quite tenable to reject both given the evidence.
Also, as with conservative politics, there always seem to be hidden, ulterior motives. For example, in Thinking Points[3], the authors quote various right-wing journals and magazines as stating before the “war” in Iraq that:
- Iraq’s oil reserves are open for the taking, and that profits should go to American companies since Iraqis are unable to handle it
- A puppet government is what is required to make this possible
- Iraq is an ideal military base for an attack on Iran
- An invasion can be justified by villainising Saddam Hussein, and using 9/11 as an excuse
- The move is justified since the Iraqi people have nothing anyway, and a semblance of democracy would make them happy; however, Americans should do what maximizes profit for themselves…
Despite being published in dozens of dated pre-war sources, they would all deny such an agenda. They pretend it was to find WMD’s, and then to remove a brutal dictator. Only in light of the above would we understand why America never invaded to remove Robert Mugabe from power, also a brutal dictator, and certainly worse in many ways. Likewise, Chomsky was out to make a name for himself. His ulterior motive was exactly that. His representation of Empiricism is very inaccurate, as is his representation of behaviourism. When I took Verbal Behavior out of the Wits Library in 2002, it had NEVER been taken out before, despite being a very old book. The spine was not even cracked. What I found was a very sophisticated and well-written account of language learning, not very different from Piaget’s account. Skinner never believed that a parrot could learn language, because of syntactic and contextual nuances. One of the reasons he proposed was a lack of attention span. The point is, though Skinner’s assumptions may be questioned, Chomsky’s famous review was not accurate, and most people do not know this because they trust his authority and veracity. In the past, Chomsky used give many lectures on his ideas on language, preceded by polemical remonstrations against the Vietnam war; many contend that the latter drew the crowds, not the former. When I wrote to Chomsky, asking some rather sycophantic questions, he replied in the form of two hand-written letters, signed. He also used to reply to every email I sent. When I raised some critical questions, his replies got more terse. My last question was regarding CL, and his one line reply was: I don’t know much about it.
Pinker also feigns ignorance about Sampson’s Educating Eve, which demolishes his (and Chomsky’s) arguments very systematically and logically. Strangely, he uses some of the very arguments used in that book in his subsequent book. He claims that he was too busy with his subsequent projects to “dwell on critiques of old ones” (personal communication). He also uses Lakoff’s theory of metaphor in How the Mind Works, saying that many researchers, including himself, have discovered it; this was not long after pointing out in very vicious and inaccurate review that the very idea is wrong.
I cannot help but draw parallels between Bush’s rejection of Scott McCelland, former White House spokesperson, after he wrote a critical book on the Bush administration, and Chomky’s rejection of his disciples in the early 1970’s when all they did was try and improve the theory. Lakoff also alleges that many of his ideas were adopted by Chomsky, without credit.
These parallels may or may not be coincidental, but this idea will be explored in more detail to ascertain its viability. At the moment I conjecture that the methodology used by CL drives one to a progressive mode of thought, and GL drives one towards doing what the conservatives do.
Conclusion
It is thereby clear that GL and CL are mutually exclusive, and the conventional hybrid solution would therefore not be a viable cop-out. Both theories aim to provide a comprehensive and holistic analysis of language and mind, together with implications of how to go about doing it.
Hitherto, it seems like CL has the upper hand; its alleged empirical accountability together with the relevant political and intellectual implications seems more palatable. GL needs to adopt a less dogmatic approach, as the assumptions guiding it mostly precede the research and therefore the evidence, which is not the way it is meant to be.
It remains to be seen whether this remains true in light of a more in-depth analysis or not.
____________________________________________________________________
[1] This is a problem which follows only as a result of the assumptions inherent in your particular theory.
[2] And in fact Locke did indeed believe in innateness, which Pinker very well knows.
[3] A progressive handbook, available from the Rockridge Institute’s web-site: www.rockridge.com.
This proposal aims to compare two opposing accounts of how language and the brain interact; these go by the names of generative linguistics and cognitive linguistics. I will outline the basic assumptions common to most mainstream adherents of these fields, together with the conclusions they are obliged to accept in light of their respective theoretical commitments. After assessing various aspects of these theories, they will be compared using the standard criteria of what ought to constitute a scientific theory, and thereby conclude that the cognitive approach seems to be more empirically and logically plausible for various reasons.
Generative Linguistics
The school of linguistics known as Generative Linguistics (henceforth GL) was started by numerous collaboraters, many of whom have since either abandoned the enterprise as initially outlined, or gone on to expand the theory in novel ways. Those who have abandoned the ship include thinkers like George Lakoff and Haj Ross, whereas those who expanded the theory in ways not originally envisioned include Ray Jackendoff, Steven Pinker and Jerry Fodor.
The central figure since the inception of the GL enterprise in the late 1950’s has been Noam Chomsky, and his position regarding the nature of language, how it ought to be studied, and its relation to human nature, has remained fundamentally unchanged over the years.
The main arguments within the GL framework have come from Noam Chomsky and Steven Pinker, and therefore their work will be drawn upon more heavily than others in outlining the said theory. Others, who have accepted the core assumptions and methodologies, will be alluded to as well, but will not constitute the core of this thesis.
The key assumptions that pervade this approach are as follows:
Language is innate, in the sense that it is part of an autonomous module (a mental organ) in the mind that grows in the child the same way your limbs grow.
Language learning or acquisition are therefore misnomers; language growth is a more appropriate term.
The physical manifestation of language (our linguistic performance) is only to be studied insofar as it provides clues to the language in our minds (our linguistic competence). We should therefore assume that a speech community is homogenous, and language should be studied as if it were part of such an idealized community.
For this reason, we are to draw a distinction between surface structure and deep structure linguistic forms.
Language is a discrete combinatorial system, and this fact allows us to use language in novel ways, which they refer to as creativity.
Language is at the core a syntactic phenomenon, which employs various finite rules which enable the user to produce a potentially infinite number of utterances. This process works like an abstract algorithm, and therefore functions independent of context and meaning. Hence, semantics should not only be seen as secondary, but as epiphenomenal.
As a genetically based module in the mind, language must have evolved at a single point in our history, either by Darwinian natural selection (Pinker) or by some kind of small genetic mutation, which somehow interacted with the mind’s computational principles (Chomsky), which gave rise to language.
Language has no direct influence on our thoughts, as thought and language are two separate processes which function separately in the mind.
Since language is part of our biological endowment, there must be a Universal Grammar (UG) present in all languages, which manifests as a result of being part of our minds at birth, providing the scaffolding that enables language to grow in the mind.
Just as our physical organs interact on some level with other organs, they are nonetheless structurally and functionally independent of them. Language, being a mental organ, should also be expected to be understood independently; language should not be intertwined with other aspects of our mental life, any more than the heart is intertwined with the functioning of the liver.
Language is understood in light of Rationalism, whereby thought is believed to be conscious and literal.
Generative Linguists are vehemently opposed to Empiricism and behaviourism.
The human mind is a uniform, modular organ.
These principles are certainly not exhaustive, but I have selected those which I seem to be the most definitive.
There are certain things that follow from this approach to the study of language and mind. Generally, accepting these premises obliges one to accept the conclusions which follow. Regarding the study of language, this includes:
Having to concede that second language learning is not at all related to the growth (sic) of your first language.
Language teaching is not related to the primary goal of linguistics.
Socio-linguistics, prescriptive grammar, speech therapy, etc. are not of primary concern to the linguist.
Reading and writing not part of language; they are cultural artefacts just like painting and dancing.
Documenting corpora is a pointless exercise, since the range of potential utterances are infinite. The goal should be to try and discover what the universal, finite rules are which all languages employ. The goal of linguistics as a whole has to be redefined, and those who pursue corpus linguistics are wasting their time.
Language is not for communication. Language appeared in the mind somehow, and we discovered that it could also be used for communication. Where it came from, and the real reason it manifested, are unknown.
Language is to be seen as part of the physical world, and needs to be studied as such.
In addition, GL uses some words in deviant ways. For example:
- creative refers to the simple act of speaking
- language is no longer seen as something spoken by a community; it is more accurately seen as an underlying mental competence, which grows in the mind and is the same in every human being. Recently Chomsky has taken to distinguishing between i-language, e-language, and Language, which adds to the confusion
- mind is merely a synonym for brain, but seen more as aspects of the brain which we know must be there
- the use of the word grow to refer to the development of a first language in the child; if the child is exposed to three languages as a young child, then three languages will “grow” in him. Generative linguists do not see this as odd, and they would be quite comfortable with an analogy along the following lines: if a child grows up gathering rocks and hunting, his hands will be stronger and more versatile; likewise, if a child grows up hearing different languages, his innate language faculty grows to be more versatile.
Also, it is not a coincidence that people like Pinker tend more towards a conservative outlook, whereas people like Lakoff tend more towards a progressive outlook. The idea that we are all the same at the core, that our linguistic and cultural disparities are just superficial, that our minds come fully equipped with genetic information, etc. means we do not need a government to see to our individual needs, since we are capable of using our innate potential to make a success of our lives. It is this mind-set that justifies the right-wing extremist view that the world should be run under one government, and any country or nation which questions it simply does not realize that we are all the same. This is the reason why democracy needs to be superimposed on Middle Eastern countries, whether they like it or not. Saudi Arabians, who are perfectly happy with their king, simply do not realize that American democracy is the best way of running a country.
Why democracy? Or why America’s brand of democracy? Well that is just a coincidence arising from the fact that they happen to be the first country to realize this. Countries like Dubai and Saudi Arabia just do not know what is good for them. One may wonder why, if “self-evident” civil liberties are innate, do they not know this; the answer: if a child is raised with his hands tied behind his back, and that is all he knows, he will not realize that they were meant to be free until some benevolent person (or institution) unties his hands, and forces the society to ban such practices. One of the conservative arguments justifying the war in Iraq goes along these lines.
This mind-set is the very same one which insists that some of the rules which pertain to American English must be part and parcel of human nature. Not only should this apply too every other dialect of English, and to the some six thousand other languages in the world, but every human being is born with this UG. If it does not appear that way, it is only because that principle must not have a surface manifestation – or maybe that principle is actually a parameter. In fact, at one point, if you did not accept this, you could hardly expect to find a job or get a research grant in the USA.
I am certainly not insinuating that the conservatives in Washington read The Blank Slate, got over excited, and invaded Iraq in an attempt to spread democracy and freedom in the Middle East, knowing now that the notion of universal liberty can be justified. I am merely drawing a parallel between these ways of thinking, and the methodologies they employ in reasoning. Just as you get different forms of democracy in the world, the principles underlying a democratic government are the same, just as with language.
Pinker refutes this by saying that political equity simply means being treated equally, not being literally equal in the biological sense. In fact, Pinker does not believe that we are all equal. In addition to his conviction that men and women have different skills and abilities (though he is careful not to hierarchialise this), Pinker also believes that there could be a racial correlate regarding our skills and abilities. During an interview with John Brockman, he said that science may one day discover that certain race groups are better at certain things than other things, which may be something we are not ready for (interview available at http://www.edge.org/q2007/q07_index.html). The interview was based on the theme Dangerous Ideas, and Pinker based his statement on a book by Daniel Dennett called Darwin’s Dangerous Idea. After pointing out how Darwin was ostracized, and how the theory shook the very foundations of the church, he says that the next shock to us will be the said discovery. He was very careful not to spell out the details, but I doubt he was implying that we may discover that Black Africans are more intellectual, and that White American males are better at manual labour.
As an aside, it is a well-known fact that the HSRC used to publish very convincing statistics during the apartheid days ‘proving’, for example, that Black people are biologically designed for intellectually inferior tasks.
The various forms of evidence cited to support these views will be discussed and criticized scientifically. I contend that the ideas on which GL is based are methodologically and empirically flawed. The same criticism applies to other paradigms which utilize the same lines of reasoning, like the conservative politicians.
These upshots and underlying assumptions are not accepted universally by linguistics scholars for various reasons. Scholars like Geoffrey Sampson rejects every argument and every logical implication of this approach; he has argued very effectively against GL in his various works over the years, like in his book Educating Eve, and various papers like Popperian Language Acquisition Undefeated, Minds in Uniform, Revisiting the ‘poverty of stimulus’ argument, etc.
His approach is a negative one, where the nativists say they have evidence and arguments for an innate language module in the brain, and Sampson points out that their evidence and arguments are flawed. These will be outlined here, but the problems with Sampson’s own account will be highlighted. His representation of Popper is not accurate, his alternative account of language acquisition and human nature is simplistic and reductionistic, and there seems to be a disparity between his academic work and his political views. For example, he got into trouble whilst an MP for saying that Asians are better at Maths because of their genes; in Minds in Uniform he takes issue with forcing a little country to abolish advocating the death sentence for speaking out against their country’s leader (cultural relativism).
Sampson is however a liberal, and his alternative theory can be supplemented by what has come to be known as Cognitive Linguistics (henceforth CL). As an aside, Sampson, like Chomsky, admits that he does not know anything about CL.
CL as a whole was born as a reaction to what is still referred to as “mainstream linguistics”, ie. Generative Linguistics. The question I will ask here is whether this reaction was justified, as is entailed rejecting all the key assumptions upon which GL rests, and in fact virtually all of Western philosophy as well, if you accept the import of George Lakoff’s and Mark Johnson’s Philosophy in the Flesh. For such a paradigm shift to be justified, the following has to be clearly demonstrated:
Each theory will have to be independently evaluated in light of evidence; refuting evidence should be taken seriously
Each theory will also need to be assessed in terms of other methods of argumentation employed to justify their stance
Where there is a deviation from the norm, either by way of reinventing terminology, postulating counter-intuitive hypotheses, or postulating more hypotheses than necessary, it needs to be adequately justified
The implications of each theory, both within its own field and in cognate disciplines, need to be assessed as well
Aspects of each theory which may be metaphysical in nature, need to be assessed in terms of its relative explanatory power
Are there aspects outside of the theory and evidence which influences the theory, like ad homonim considerations, socio-political factors, etc.?
These are the standard criteria used when deciding which theory is a better one. It is only a coincidence that I will use the works of Karl Popper in doing so, as he is one of the few philosophers of science to systematically outline what the logic of scientific inquiry entails, though few would disagree with his approach. Had I used Einstein, Newton, Feynman, Penrose, Putnam, or any other serious scientific thinker, the criteria would be the same.
The evidence and methodology used by the GL school specifically will be assessed. If their hypotheses make predictions, I will test these predictions, and see how the theory handles counter-evidence. If the theory can be always adapted to assimilate counter-evidence, then it is problematic. If the theory can construct ad hoc hypotheses without compromising testability, it would be better. If new concepts and terms are invented to account for this, then these new concepts have to be justified in light of the evidence, not just to protect the theory. If counter-evidence is ignored all together, then the theory is not an empirically responsible one, and therefore not scientific.
Before doing this, we need to understand a little about the alternative theory, which we will turn to now:
Cognitive Linguistics
CL was born as a reaction to GL, after exponents like Lakoff initially tried to expand Chomsky’s theory to include semantics, which at the time went by the name Generative Semantics. As more things came to be known, the idea of trying to incorporate all evidence within a GL paradigm became more and more difficult, and a ‘breakaway faction’ ensued, resulting in what was documented by Harris as The Linguistics Wars. Lakoff was a student of Chomsky’s, and the latter seemed to take this as a betrayal. After Lakoff’s criticisms, Chomsky is on record as saying that Lakoff does not understand linguistics very well. Today, Chomsky claims that he knows nothing about Cognitive Linguistics!
Anyway, while Lakoff is one of the exponents, CL cannot be attributed to a single ‘founding father’. It is a collaborative movement which is always evolving, though there are theorists who are prominent within the field. As a result, there are various branches of CL, and not every linguist who calls himself a CL has to accept all the assumptions that others espouse. Hence, delineating criteria which distinguish the field can lead to some fuzziness.
Regardless, these are some of the key assumptions shared by CL as I understand it:
Language may or may not be innate. It remains to be seen how much is innate if so. Regardless, evidence thus far shows that language is part and parcel of other higher cognitive functions, and is neither an autonomous mental organ, nor does it grow – it is learnt via interaction with the environment and the mind, using all its capacities.
Language as expressed and transcribed is one our most important sources of data, and should be studied as such without assuming that each speaker is homogenous, and that his language does not change in any way.
The idea of every sentence having a deep structure is thereby also rejected, since the reasons for postulating it in the first place are not clear.
Language is at the core a semantic phenomenon which cannot be divorced from context. It makes evolutionary sense that parable and narrative evolved first, followed by other aspects of language. Otherwise, nothing precludes a computer from simulating human language.
Language could not have evolved the same way our physical organs evolved, since that assumes that natural selection has a vision to the future; if it evolved with other cognitive faculties and manifested initially as parable, it would have immediate survival value.
Language and thought are inextricably linked, so much so that without language our level of consciousness would be reduced.
Though there is an innate basis for language, it has to be learnt via hypothesis formation. Hence, the will be commonalities constrained in trivial ways by our biology, but nothing to suggest that there is a UG hard-wired into our brains.
Language is an abstract cognitive faculty, which is neither functionally nor cognitively separate from the rest of our other cognitive faculties. This predicts that damage to our language centres in the brain would affect other things as well, and that language proficiency is conducive to clear thought (and of course, that linguistic proficiency varies from individual to individual)
Language is by and large figurative (or metaphorical, in the larger sense of the word), and thought is mostly unconscious.
Empiricism and behaviourism are not rejected by fiat, but aspects of either are assessed on their merit.
Language and mind are not necessarily part of the material world. While not saying so explicitly saying so, they leave a gap in their paradigm to cater for the possibility that they might be good candidates as possible denizens of Popper’s World 3.
Once again, these assumptions are not exhaustive, and the evidence will be duly scrutinized, as with GL. As mentioned, this school was born as a reaction to mainstream linguistics, and as such will be expected to handle the evidence more responsibly, as they claim, should be commensurable with common sense, and the rejection of things like the primacy of syntax will be looked at.
Stylistic issues like the novel use terminology will be assessed as either bad style or bad science, since CL scholars do not do this. The question will be asked: are there any theoretical implications leading to pseudo-problems[1] as a result of this? If so, would the problem still be there if these additional assumptions were not made? (For example, the concept of creativity in the conventional sense does not exist anymore, for generativists – though this is more of a semantic reductionism, which is the converse…)
To avoid redundancy, the entailments of these assumptions will not be reiterated as they are pretty much the converse of the opposing paradigm, as expected.
Regarding politics, CL seem to have a less euro-centric view of the world, and are progressive in their thinking. They do not feel the need to create a hierarchy, either in the academic world or the political world. This is not only because of George Lakoff’s now defunct progressive think tank, The Rockridge Institute, but because there are explicit links between the two (unlike GL, where the links are implied and never stated explicitly; in Chomsky’s case, he explicitly denies the link).
As mentioned, there are parallels between conservative political arguments and the way GL argue. The correlation does not entail causation, as any statistician would confirm. This merely shows a parallel in thinking, with possible causation. For example, the reason gay marriage is opposed is argued for as follows:
If they allow this, then what is next? What’s to stop from condoning bestiality, and then allowing them to get married!
Or their argument against raising the minimum wage:
If you give them a two dollar raise, what is to stop the unions from demanding a hundred dollar raise the next time round!
For some reason, it has to be one extreme or the other. Gay marriage not creating a terrible problem in South Africa is not considered as counter-evidence. A colleague told me: Wait and see what happens in five years time! When I pointed out that this has been happening for five years, he did not feel the need to reevaluate his position, but without reason. Likewise, Pinker would tell you that you either accept that there is a UG, with built in x-bars and island constraints, hard-wired into our minds (and in fact many other alleged things not related to language, as he explains in The Blank Slate, How the Mind Works, and The Stuff of Thought), or that we believe that the mind is a blank slate no different from a rock – he says we must either believe there is a human nature (and therefore accept his argument in its entirety), or totally reject the idea of the existence of human nature. He is careful to quote John Locke as advocating the “blank slate” idea, knowing that few people have read his works[2], but then adds that Locke does not deserve the blame for the idea that most intellectuals today believe that our minds are empty buckets and that people are hapless victims of their environment. As with the “gay marriage” argument, the alternative does not follow; likewise, you can reject GL and still believe that there is such a thing as human nature. Likewise, Chomsky has stated that the existence of UG is unquestionable, unless you believe in magic; clearly, it is quite tenable to reject both given the evidence.
Also, as with conservative politics, there always seem to be hidden, ulterior motives. For example, in Thinking Points[3], the authors quote various right-wing journals and magazines as stating before the “war” in Iraq that:
- Iraq’s oil reserves are open for the taking, and that profits should go to American companies since Iraqis are unable to handle it
- A puppet government is what is required to make this possible
- Iraq is an ideal military base for an attack on Iran
- An invasion can be justified by villainising Saddam Hussein, and using 9/11 as an excuse
- The move is justified since the Iraqi people have nothing anyway, and a semblance of democracy would make them happy; however, Americans should do what maximizes profit for themselves…
Despite being published in dozens of dated pre-war sources, they would all deny such an agenda. They pretend it was to find WMD’s, and then to remove a brutal dictator. Only in light of the above would we understand why America never invaded to remove Robert Mugabe from power, also a brutal dictator, and certainly worse in many ways. Likewise, Chomsky was out to make a name for himself. His ulterior motive was exactly that. His representation of Empiricism is very inaccurate, as is his representation of behaviourism. When I took Verbal Behavior out of the Wits Library in 2002, it had NEVER been taken out before, despite being a very old book. The spine was not even cracked. What I found was a very sophisticated and well-written account of language learning, not very different from Piaget’s account. Skinner never believed that a parrot could learn language, because of syntactic and contextual nuances. One of the reasons he proposed was a lack of attention span. The point is, though Skinner’s assumptions may be questioned, Chomsky’s famous review was not accurate, and most people do not know this because they trust his authority and veracity. In the past, Chomsky used give many lectures on his ideas on language, preceded by polemical remonstrations against the Vietnam war; many contend that the latter drew the crowds, not the former. When I wrote to Chomsky, asking some rather sycophantic questions, he replied in the form of two hand-written letters, signed. He also used to reply to every email I sent. When I raised some critical questions, his replies got more terse. My last question was regarding CL, and his one line reply was: I don’t know much about it.
Pinker also feigns ignorance about Sampson’s Educating Eve, which demolishes his (and Chomsky’s) arguments very systematically and logically. Strangely, he uses some of the very arguments used in that book in his subsequent book. He claims that he was too busy with his subsequent projects to “dwell on critiques of old ones” (personal communication). He also uses Lakoff’s theory of metaphor in How the Mind Works, saying that many researchers, including himself, have discovered it; this was not long after pointing out in very vicious and inaccurate review that the very idea is wrong.
I cannot help but draw parallels between Bush’s rejection of Scott McCelland, former White House spokesperson, after he wrote a critical book on the Bush administration, and Chomky’s rejection of his disciples in the early 1970’s when all they did was try and improve the theory. Lakoff also alleges that many of his ideas were adopted by Chomsky, without credit.
These parallels may or may not be coincidental, but this idea will be explored in more detail to ascertain its viability. At the moment I conjecture that the methodology used by CL drives one to a progressive mode of thought, and GL drives one towards doing what the conservatives do.
Conclusion
It is thereby clear that GL and CL are mutually exclusive, and the conventional hybrid solution would therefore not be a viable cop-out. Both theories aim to provide a comprehensive and holistic analysis of language and mind, together with implications of how to go about doing it.
Hitherto, it seems like CL has the upper hand; its alleged empirical accountability together with the relevant political and intellectual implications seems more palatable. GL needs to adopt a less dogmatic approach, as the assumptions guiding it mostly precede the research and therefore the evidence, which is not the way it is meant to be.
It remains to be seen whether this remains true in light of a more in-depth analysis or not.
____________________________________________________________________
[1] This is a problem which follows only as a result of the assumptions inherent in your particular theory.
[2] And in fact Locke did indeed believe in innateness, which Pinker very well knows.
[3] A progressive handbook, available from the Rockridge Institute’s web-site: www.rockridge.com.
Subscribe to:
Posts (Atom)