4

Collective Intentionality

There is no thinking apart from common standards of correctness and relevance, which relate what I do think to what anyone ought to think. The contrast between “I” and “anyone” is essential to rational thought.

—WILFRID SELLARS, PHILOSOPHY AND THE SCIENTIFIC IMAGE OF MAN

A modern human society may be characterized in two dimensions. The first is its synchronic social organization: the coordinated social interactions that make it a society in the first place. Early human individuals, as we have seen, coordinated in acts of collaborative foraging with specific others from a loosely structured pool of collaborators. But now with modern humans we need to scale up to much larger social groups with much more complex social organization, that is to say, to fully cultural organization. Modern humans became cultural beings by identifying with their specific cultural group and creating with groupmates various kinds of cultural conventions, norms, and institutions built not on personal but on cultural common ground. They thus became thoroughly group-minded individuals.

The second dimension of a modern human society is its diachronic transmission of skills and knowledge across generations. Social transmission of one sort or another was almost certainly important in the lives of early humans (as it is in some apes’ lives), as difficult-to-invent, tool-based, subsistence activities became ever more complex and important for survival. But now with modern humans we need to scale up to full-fledged cultural transmission supporting cumulative cultural evolution. This required that modern humans not just acquire instrumental actions by observing others, as did early humans, but actively conform to the behavior and norms of the group, and even enforce conformity on others through teaching and social norm enforcement.

The combination of these changes in the two dimensions of human sociality created some totally new cultural realities. The transformative process was conventionalization, which has both a coordinative component, as individuals implicitly “agree” to do something in a consistent way (everyone wants to do it this way as long as everyone else does, too), and a transmitive component, as this way of doing things sets a precedent to be copied by others who want to coordinate as well. The result is what we may call cultural practices, in which individuals, in effect, coordinate with the entire cultural group via collectively known cultural conventions, norms, and institutions. In communication this means, of course, linguistic conventions, which serve their coordinative function because, and only because, they exist as “agreements” in the group’s cultural common ground.

In terms of thinking, early humans imagined the world in order to manipulate it in thought via perspectival cognitive representations, socially recursive inferences, and social self-monitoring—which prepared them to coordinate with other specific individuals. But group-minded and linguistically competent modern humans had to be prepared to coordinate with anyone from the group, with some kind of generic other. This meant that modern human individuals came to imagine the world in order to manipulate it in thought via “objective” representations (anyone’s perspective), reflective inferences connected by reasons (compelling to anyone), and normative self-governance so as to coordinate with the group’s (anyone’s) normative expectations. And these group-minded ways of operating and thinking were not present just in specific ad hoc collaborative interactions of the moment; rather, because of the way that modern humans became competent members of a cultural group during ontogeny, they created a permanent imprint in the human mind-set.

So once more let us look, first, at the new forms of collaboration evident in human cultural organization, then at the new forms of conventional linguistic communication for coordinating cultural life, and then at the resulting new forms of agent-neutral, normatively governed thinking that cultural life demanded.

The Emergence of Culture

A number of animal species, from whales to capuchin monkeys, engage in one or another form of social transmission, requiring some form of social learning. The most cultural of nonhuman animals are undoubtedly the great apes, especially chimpanzees and orangutans. Observations in the wild have documented for these two species a relatively large number of population-specific behaviors that persist in the group over time and that very likely involve social learning (Whiten et al., 1999; van Schaik et al., 2003). Experimental studies have also demonstrated some skills of social learning in these two species, for example, in learning to use novel tools, that very likely are at work in generating their cultural patterns in the wild (see Whiten, 2010, for a review).

But great ape culture is not human culture. Tomasello (2011) characterizes great ape culture as mainly “exploitive,” as individuals socially learn from others who may not even know they are being watched. Modern human culture, in contrast, is fundamentally cooperative, as adults actively teach children, altruistically, and children actively conform to adults, as a way of fitting in cooperatively with the cultural group. The hypothesis is that this cooperative form of culture was made possible by the intermediate step of early humans’ highly cooperative lifeways and how this transformed great ape social learning into truly cultural learning. Teaching borrows its basic structure from cooperative communication in which we inform others of things helpfully, and conformity is imitation fortified by the desire to coordinate with the normative expectations of the group. Modern humans did not start from scratch but started from early human cooperation. Human culture is early human cooperation writ large.

Group Identification

The small-scale, ad hoc collaborative foraging characteristic of early humans was a stable adaptive strategy—for a while. In the hypothesis of Tomasello et al. (2012), it was destabilized by two, essentially demographic, factors.

The first factor was competition with other humans. This meant that a loose pool of collaborators had to turn into a proper social group in order to protect their way of life from invaders. A loose social grouping of early humans was under pressure to transform into a coherent collaborative group with joint goals aimed at group survival (each group member needing the others as collaborative partners for both foraging and fighting) and division-of-labor roles toward this end. As with early humans’ smaller-scale collaborations, this meant that group members were motivated to help one another, as they were all now clearly interdependent with one another at all times: “we” must together compete with and protect ourselves from “them.” Individuals thus began to understand themselves as members of a particular social group with a particular group identity—a culture—based on a we-intentionality encompassing the entire group.

The second factor was increasing population size. As human populations grew, they tended to split into smaller groupings, leading to so-called tribal organization in which a number of different social groupings were still a single supergroup or “culture.” This meant that recognizing others from our own cultural group became far from trivial—and of course we needed to ensure that they could recognize us as well. Such recognition in both directions was important because only members of our cultural group can be counted on to share our skills and values and so be good and trustworthy collaborative partners. Contemporary humans have many diverse ways of marking group identity, but one can imagine that the original ways were mainly behavioral: people who talk like us, prepare food like us, and net fish in the conventional way—that is, those who share our cultural practices—are very likely members of our cultural group.

And so early humans’ skills of imitation became modern humans’ active conformity, both to coordinate activities more effectively with in-group strangers and to display group identity so that others will choose me as a knowledgeable and trustworthy partner. Teaching others to do things, perhaps especially one’s children, became a good way to assist their functioning in the group and to ensure even more conformity in the process. Teaching and conformity then led to cumulative cultural evolution characterized by the “ratchet effect” (Tomasello et al., 1993; Tennie et al., 2009; Dean et al., 2012) in which modifications of a cultural practice stayed in the population rather faithfully until some individual invented some new and improved technique, which was then taught and conformed to until some still newer innovation ratcheted things up again. Tomasello (2011) argues that great ape societies do not display the ratchet effect or cumulative cultural evolution because their social learning is fundamentally exploitative and not cooperatively structured in the human way via teaching and conformity, which constitute the ratchet that prevents individuals from slipping backward.

The new sense of group identity characteristic of modern humans was thus extended not just in space to in-group strangers but also in time to ancestors and descendants in the group: this is the way “we” have always done things; it is part of who “we” are. As cultural practices were handed down across generations cooperatively—adults altruistically teaching and youngsters trusting and even conforming—the resulting cumulative effect was that the “we” became an enduring culture to which we (past, present, and future) are all committed (just as early humans were committed to their ongoing, small-scale collaborations). Human populations thus became more than a loosely structured pool of collaborators; they become self-identified cultures with their own “histories.” Once again, precisely when this all happened is not crucial to our story, but the first clear signs of distinct human cultures appear with Homo sapiens sapiens, that is, modern humans, beginning at the earliest some 200,000 years ago.

That humans do indeed think of their group as a “we” of interdependent individuals—that humans identify with their group—is a well-established psychological fact. Most fundamentally, humans have a marked in-group/out-group psychology that is, in all likelihood, unique to the species. Much research shows that humans favor their in-group in all kinds of ways, and they care about their reputation more in their in-group than in any out-groups as well (Engelmann et al., in press). Moreover, they think of others from other groups not just as strangers, as do apes and as did early humans, but as members of specific out-groups with alien, often despised ways. Perhaps the most striking phenomenon of group identity is collective guilt, shame, and pride. Individuals feel guilty, ashamed, and/or proud when an individual of their group does something noteworthy in basically the same way that they would if they themselves had done the deed (Bennett and Sani, 2008). In the contemporary world, one sees such group identity and collective guilt, shame, and pride quite clearly in struggles over ethnic identity, linguistic identity, collective responsibility, and so forth—and even in such frivolous phenomena as fan support of sports teams. As far as we know, great apes do not have, and early humans did not have, this sense of group identity at all.

The proposal is thus that with increasing population sizes and competition among humans, the members of human groups began to think of themselves and their groupmates (known and unknown, present and past) as participants in one big, interdependent, collaborative activity aimed at surviving and thriving in competition with other human groups. Group members were identified most readily by specific cultural practices, and so teaching and conformity to the group’s lifeways became a critical part of the process. These new forms of group-mindedness led to what we may call the collectivization of human social life, as embodied in group-wide cultural conventions, norms, and institutions—which transformed, one more time, the way that humans think.

Conventional Cultural Practices

Group identification means that human groups each have their own set of conventional cultural practices. Conventional cultural practices are things that “we” do, that we all know in cultural common ground that we do, and that we all expect one another in cultural common ground to do in appropriate circumstances. Thus, in an open barter food market in which a conventionalized set of measurements for coordination is in place, if I show up with my honey in unconventional containers, no other traders will know what to do with me and my undetermined quantities of honey. With conventional cultural practices, deviations are not punished per se; they are simply left on the outside looking in. And there are some conventions that one cannot opt out of: one can wear this clothing, or that clothing, or nothing at all, but whatever one wears, it is a cultural choice that will either conform to or violate the expectations of others in the group.

Unlike the second-personal common ground that early human individuals created with one another as they engaged in collaborative activities, the common ground at this point is what Clark (1996) calls cultural common ground: things that we all in the group know that we all know even if we did not experience them together as individuals. Indeed, Chwe (2003) argues that the main function of the public events of a culture is to make sure that such things as the chief ’s coronation or his daughter’s marriage ceremony become public knowledge: a part of the cultural common ground that everyone can count on everyone else knowing, which no one can plausibly deny knowing, and knowledge of which serves as a shibboleth of group membership. Interestingly, children as young as two years of age are already tuned into cultural common ground. Thus, Liebal et al. (2013) had two- and three-year-old children meet a novel adult (clearly from their group). This in-group stranger then asked them sincerely, “Who is that?” while they looked together at a Santa Claus toy and a toy the child had just made before the adult entered. Children answered by naming the newly created toy, as even children this young know that no one in the culture, not even someone they have never before met, needs to ask who is Santa Claus. (In a second condition they named Santa Claus if the stranger asked for the name of a toy she seemed to recognize.) Children in this same age range also expect that in-group strangers will know the conventional name of an object but not a novel, arbitrary fact about that same object (Diesendruck et al., 2010).

Some conventional cultural practices are the product of explicit agreement. But this is not how things got started; a social contract theory of the origins of social conventions would presuppose many of the things it needed to explain, such as advanced communication skills in which to make the agreement. Lewis (1969) thus proposed another way to get started. We begin with a coordination problem, say, what time to show up to go group fishing every day at our new camp. Let us say that by chance we depart on the first day at midday (just because that is when enough people for the task have congregated). Assuming no advanced communication skills, what do we do the next day? Following Schelling (1960), Lewis (1969) proposes that we search for anything to single out one time from all the others on this second day, and a natural way of doing that for humans seems to be “precedent”—we do what we did before (what worked for us before) and show up at midday again. And so we habituate, and new participants just imitate and conform to us. Anyone who does not conform simply does not participate.

But with the possibility of communication, we may also teach the convention to others and encourage them to conform and so to participate in the cultural practice. Importantly for our understanding of collective intentionality, when adults teach children how to perform cultural practices, the children take it not as communication about the current episodic event but, rather, as something general about the world, applying to things and/or events like these in general (i.e., “Fishing takes place at midday”). Thus, an adult might communicate to a child, without teaching, that there is a fish right there in the water. But when he goes into teaching mode the message is something more like, “These kind of fish are found in places like this,” to facilitate the child’s fishing skills in general (Csibra and Gergely, 2009). The implication of this pedagogical mode of cooperative communication is that there is a kind of objective reality that works on general principles (these kinds of fish in general, and these kinds of places in general), and the current situation is merely one instance of this objective reality. Teaching is implicitly backed by the collective and objective perspective on things developed by our cultural group.

Modern human children are thus learning from adults that there are certain ways that things should work. This “should” implicitly undergirding adult teaching prompts children, in ways that we do not fully understand, to objectify and reify the generic facts they are being taught into an objective reality—a generic perspective that is the ultimate adjudicator among differing perspectives on the world. This process has many implications for human thinking, but a prominent one is the understanding of false beliefs (which great apes clearly do not do; see Tomasello and Moll, in press, for a review). Thus, we previously invoked something like Davidson’s notion of social triangulation to explain how it is that early humans came to understand that others have perspectives that differ from their own. But to get to an understanding of beliefs, including false beliefs, we must have some notion of a generalized perspective on an objective reality that is independent of any particular perspective. Something like this is needed to make the judgment not just that a belief is different from mine, but that it is wrong—since objective reality is the final arbiter. It is likely that young children begin to think in terms of multiple different perspectives on things from as soon as they participate in joint attention with its two perspectives during late infancy (Onishi and Baillargeon, 2005; Buttelmann et al., 2009), and we may hypothesize that this was the case for early humans as well. But it is not for several more years that children come to a full-blown understanding of beliefs, including false beliefs, because they (and so presumably all humans before modern humans) do not yet understand “objective reality.”1

Social Norms and Normative Self-Monitoring

In the small-scale collaborative interactions of early humans, individuals actively chose some collaborative partners and shunned others, and in some cases even rewarded and punished partners. But this was all done in second-personal mode, that is, as one individual evaluating another individual. What happened with modern group-minded humans was that these evaluations became conventionalized and so applied in agent-neutral, transpersonal mode, that is, applied by all to all (even by and to those not directly involved in an interaction) and with respect to objective, transpersonal standards. Although great apes retaliate for harm done to them, they do not punish other individuals for acts toward third parties (Riedl et al., 2012). In contrast, three-year-old children enforce social norms on others even when they are not personally involved or affected in any way, often using normative language about what one should or should not do in general (Rakoczy et al., 2008; see Schmidt and Tomasello, 2012, for a review).

Social norms are thus mutual expectations in the cultural common ground of the group that people behave in certain ways, where the mutual expectations are not just statistical but, rather, socially normative, as in you are expected to do your part (or else!). The force of the expectations derives from the fact that individuals who do not conform to our group’s way of doing things often create disruptions, which should not be tolerated, and indeed, if individuals behave too differently it signals that they are not one of us (or do not want to be one of us) and so cannot be trusted. Group-minded individuals thus view nonconformity in general as potentially harmful to group life in general. The result is that humans conform to social norms for instrumental reasons (to coordinate successfully), for prudential reasons (to avoid the group’s opprobrium), and in order to benefit of the group’s functioning since nonconformity potentially disrupts this functioning (a group-minded reason).

Like conventions in general, social norms operate not in second-personal mode but rather in agent-neutral, transpersonal, generic mode. First, and most basic, social norms are generic in that they imply an objective standard against which an individual’s behavior is evaluated and judged. In early humans’ social evaluations, individuals only knew who did things ineffectively or noncooperatively, but now the roles have specific agent-neutral standards (that can be taught as such). These objective standards come from the mutual understanding of how the different functions in particular conventionalized cultural practices are effected if everyone is to reap the anticipated benefit. Thus, if it is cultural common ground in the group that when collecting honey the person smoking out the bees must do so in this particular way, and that if she does not do it in this way we will all go home empty-handed, then her behavior may be evaluated relative to this objective standard for job performance.

Social norms are also generic in terms of their source. Social norms emanate not from an individual’s personal preferences and evaluations but, rather, from the group’s agreed-upon evaluations for these kinds of things. Thus, when an individual enforces a social norm, she is doing so, in effect, as an emissary of the group as a whole—knowing that the group will back her up. Group-minded individuals thus enforce social norms because their collective commitment to a social norm means that they commit not only to following it themselves but also to seeing that others do, too—for the benefit both of ourselves and of other group members with whom we are interdependent (Gilbert, 1983). The typical formulation of individuals enforcing social norms would be something like, “One cannot do it like that; one must do it like this,” which is of course very similar to the generic mode used in teaching. (Indeed, norm enforcement and teaching may be two versions of the same phenomenon: enculturating individuals to our group’s ways of doing and thinking about things.) The presupposition of norm enforcement is that this is the group’s collective perspective and evaluation, or possibly even something generalized beyond that to some deity or some external normative facts of the universe: it is just the case in our world that this is the right way and that is the wrong way to do it.

And finally, social norms are also generic in their target: group disapproval is aimed in an agent-neutral way at in principle anyone, that is, anyone who identifies with our group’s lifeways and so mutually knows and accepts in cultural common ground our social norms. (The common ground assumption exempts from the force of our social norms individuals from another social group, young children, and mentally incompetent persons [Schmidt et al., 2012].) This agent neutrality in application is nowhere more evident than in the fact that people apply social norms to themselves in acts of guilt and shame. Thus, if I take some honey that is needed by others, I will feel guilty under the force of the norm against stealing (perhaps undergirded by sympathy for the victim). More telling, if some illicit avocation of mine is made public, I will feel ashamed under the eyes of the group norm, even though I may not feel that it is wrong at all. Guilt and shame thus demonstrate with special clarity that the judgment being made is not my personal feeling about things (I wanted the honey and that avocation), but rather, it is the group’s—which is especially clear in the case of shame, in which I may not even agree with the group. Nevertheless, I, as an emissary of the group, am sanctioning myself. Guilt and shame may thus have some second-personal bases in me feeling bad that I harmed another individual or that I did not conform to the expectations of valued others, but the full versions require in addition that I know I violated a collective norm. It is not just that my victim feels bad, or that I offended others’ sense of decency but, more important, that the group—which includes me—disapproves.

Because I know that things work in this way, I self-monitor and self-regulate my actions via group norms so as to coordinate with the expectations of the group. In this normative self-monitoring, as we may call it, what one is often trying to protect is a public reputation, one’s status as a cooperative member of the group (Boehm, 2012). That is to say, with modern human collaboration taking place on the level of the entire cultural collectivity, my behavior in various contexts might be known to some degree in the cultural common ground of the group as a whole (e.g., because of the pervasiveness of gossip, done best in a conventional language; see below). This means that early humans’ concern with being judged is transformed by modern humans into a concern for one’s public reputation and social status. And, critically, reputational status is more than just a sum of many social evaluations; it is nothing less than a Searlian status function (see next section) in which my public persona is a reified cultural product created by the collectivity, who can take it away in a second, as any scandalized modern politician can attest.

Institutional Reality

In the limit, some conventional cultural practices turn into full-blown institutions. Obviously, the dividing line is fuzzy, but a basic prerequisite is that the cultural practice is not a solo activity but is in some sense collaborative, with well-defined, complementary roles. But the key feature distinguishing cultural institutions is that they comprise social norms that do not just regulate existing activities but, rather, create new cultural entities (the norms are not regulative but constitutive). For example, a human group might tend to make decisions about such things as where to travel next, how to set up defenses against potential predators, and so forth, by simply arguing among themselves. But if there are difficulties in making decisions, or infighting among several coalitions, then the group could institutionalize the process into some kind of governing council. Creating this council would give otherwise normal individuals abnormal status and powers. The council might then designate a chief, whom they would empower to do still other abnormal things, like banish people from the group. The council and chief are thus cultural creations, and their entitlements and obligations are bestowed upon them by the members of the group, who can, in theory, take them away and so turn the council members and chief back into everyday people again. The roles in institutions are explicitly agent neutral because, in theory if not in practice, anyone may play any role.

Searle (1995) has been most explicit about how this process works. First, obviously, is some kind of mutual agreement or joint acceptance among group members to designate, for example, an individual as chief. Second, there must be some kind of symbolizing capacity so as to enact Searle’s well-known formula “X counts as Y in context C” (X counts as chief in the context of group decision making). Related to this, there should be some actual physical symbol—something like a leader’s headdress or scepter or presidential seal—to help in marking the new status in a public way. The fact that institutions are public means that everyone knows that everyone knows about them and cannot plead ignorance in the face of overt symbolic marking. This is one reason why new institutions and officials are anointed with their new obligations and entitlements not just implicitly but explicitly and publicly. One could not do something bad to a chief wearing his official headdress, right after his official inauguration, and then plead ignorance of his status. Similarly with formally written rules and laws: their public nature essentially means that one cannot break them and expect to be excused by pleading ignorance.

Rakoczy and Tomasello (2007) argue that a simple model for understanding cultural institutions is rule games. Of course one may move a piece of wood shaped like a horse all around a checkered board in any way one likes. But if one wants to play chess, then one acknowledges that this horse-shaped piece is a “knight,” and one moves a knight only in certain ways, and the other pieces in other ways, toward the goal of winning the game, where winning is defined by certain agreed-upon configurations of pieces. The pieces are given their statuses by the norms or rules, whose existence comes from, and only from, the explicit agreement of the players. Thus, we would argue that the ontogenetic cradle of such cultural status functions is young children’s joint pretense when they, for example, designate together a stick to be a snake. In doing this they are engaging in the fundamental act that creates new statuses since, we would claim, this designation is a social, public agreement with one’s play partner (see Wyman et al., 2009). Importantly, although the ability to pretend derives evolutionarily, as we have argued, from the way that early humans created pretend realities by pantomiming situations for others communicatively, the normative dimension comes only with the group-mindedness and collectivity characteristic of modern human cultures.

The most important point for current purposes is that there are in the modern human world social or institutional facts. These are objective facts about the world: Barack Obama is president of the United States, the piece of paper in my pocket is a €20 note, and one wins at chess by checkmating one’s opponent. At the same time that they are objective, however, these facts are observer relative; that is, they are created by individuals in social groups, and so they may be just as easily dissolved (Searle, 1995). Barack Obama is president but only as long as we say so, euros are legal tender but only as long as we act so, and, in theory, the rules of chess may be changed at any time. What is absolutely extraordinary about social facts, then, is that they are both objectively real and socially created, speaking once again to the power of the objectification/reification process. Indeed, if one gives five-year-old children some objects with almost no instruction, they very quickly create their own rules for how to play with them, and they then apply these rules both to themselves and to new players as objective facts: “One must do this first,” “It works like this,” and so forth (Goeckeritz et al., unpublished manuscript). As in adult teaching and norm enforcement, the “must” here implies the guiding hand of an objective reality, independent of the perspective or wishes of any particular individual.

Summary: Group-Mindedness and Objectivity

The social interactions of early humans were wholly second-personal. The social interactions of modern humans added on top of this a group-minded layer, starting with identification with one’s group. Individuals in a particular cultural group know that everyone knows certain things, knows that everyone else knows them, and so on, in the cultural common ground of the group. There are collectively accepted perspectives on things (e.g., how we classify the animals of the forest, how we constitute our governing council) and collectively known standards for how particular roles in particular cultural practices should be performed—indeed, must be performed—if one is to be a member of the group. The group has its perspective and evaluations and I accept them; indeed, I myself help to constitute the group’s perspective and evaluations, even if the target is myself.

Crucially, the generality involved in this new group-mindedness is not just schematicity. We are not talking here about an individual perspective somehow generalized or made large, or some kind of simple adding up of many perspectives. Rather, what we are talking about is a generalization from the existence of many perspectives into something like “any possible perspective,” which means, essentially, “objective.” This “any possible” or “objective” perspective combines with a normative stance to encourage the inference that such things as social norms and institutional arrangements are objective parts of an external reality. The generic nature of the communicative intention in both norm enforcement and pedagogy derives from the inherently generic group-mindedness and social normativity governing the way that “we” expect “us” to do things, which is then objectified into this is the way things are, or ought to be, in the world at large.

Early humans’ dual-level cognitive model of jointness and individuality is thus scaled up by modern humans into a group-minded cognitive model of objectivity and individuality. Human group-mindedness thus reflects a profound shift in ways of both knowing and doing. Everything is genericized to fit anyone in the group in an agent-neutral manner, and this results in a kind of collective perspective on things, experienced as a sense of the “objectivity” of things, even those we have created. Thus is human joint intentionality “collectivized.”

The Emergence of Conventional Communication

In addition to conventionalizing their social lives in general into collective cultural practices, norms, and institutions, modern humans also conventionalized their natural gestures into collective linguistic conventions. Early humans’ spontaneous, natural gestures of the moment were important for coordinating their many collaborative activities, but conventionalized gestures and vocalizations—collectively known by all of those who have grown up in our cultural group, past and present, but not by others—enabled much more decontextualized and flexible forms of communication and social coordination with all members of the cultural group, even those with whom one had never before interacted.

Given a naive view of the nature of linguistic communication, one might assume that the use of language obviates the need for thinking in the communicative coordination of intentional states: I encode my “meaning” in language, and you decode it—the way telegraph operators used to work with Morse code. But, in fact, this is not the way linguistic communication works (Sperber and Wilson, 1996). For example, a large proportion of words in everyday spoken discourse are pronouns (hesheit), indexicals (herenow), or proper names (JohnMary), whose referents cannot be determined from a codebook but rather must be determined by accessing some kind of nonlinguistically constituted common conceptual ground. Moreover, our everyday discourse is liberally peppered with seemingly incoherent sequences, for example, Me: “Want to go to a movie tonight?” You: “I have a test in the morning.” For me to interpret your answer as a “No,” we must share in common ground an understanding that tests require studying beforehand, one cannot study and watch a movie at the same time, and so forth. This common ground then makes possible my abductive leap that you will not be coming with me to the movie.

And so the most basic thinking processes in linguistic communication are the same as in the account of pointing and pantomiming in chapter 3. In informative linguistic communication, I intend that you know something, and so I refer your attention or imagination to some situation (my referential act) in the hope that you will figure out what I intend you to know (my communicative intention). Then you, relying on our common ground (both personal and cultural), hypothesize abductively what my communicative intention might be, given that I want you to attend to this referential situation. Thus, I might come into your office and say, “The city of Leipzig is running a tennis camp for two months this summer.” You comprehend my referential act perfectly well, but you have no idea why I am informing you of this situation. But then the abductive light bulb comes on: oh, if he intended to make a suggestion about what my children might do this summer (since we were talking about that last week), then referring my attention to this fact would make perfect sense. Anticipating this process, and in order to be an effective communicator, I attempt to simulate your potential abductive inference ahead of time and so formulate my referential act to guide it in the intended direction—just as in pointing and pantomiming. For example, I might anticipate that if I refer to “tennis camp” without referring to “this summer,” she will think I mean a tennis camp for herself and not for her kids’ summer activities—a process that is essentially me thinking about her thinking about my communicative act (which is itself an intention toward her intentional states).

Needless to say, communicative conventions add much more articulate semantic content to the communicative act than do pointing and pantomiming, perhaps making things easier, but they do not thereby obviate all of the simulating, inferring, and thinking that must be done for two language users to communicate about complex situations successfully. In addition, despite this commonality of process with the use of spontaneous gestures at some basic level, linguistic communication also provides some powerful new resources for human thinking. We may focus on four of these, each elaborated in its own section below: (1) communicative conventions as inherited conceptualizations, (2) linguistic constructions as complex representational formats, (3) discourse and reflective thinking, and (4) shared decision making and the giving of reasons.

Communicative Conventions as Inherited Conceptualizations

Early humans’ use of spontaneous iconic gestures as symbols to direct one another’s attention and imagination to relevant situations now became, with modern humans, conventionalized in the group. This meant not only that interpreting a gesture depended on some personal common ground between communicative partners in the moment, as before, but also that it now depended on some cultural common ground about how we in this group expect others in this group to use and interpret this gesture (and expect others to expect us to, etc.). Thus, for instance, we all know in cultural common ground that when one wants to direct the attention or imagination of a partner to a snake-danger situation, one produces a wavy hand gesture in the direction of the potential danger. Such conventions are coordination devices in the sense that individuals want to use them only if everyone else uses them as well (Lewis, 1969; Clark, 1996). Communicative conventions thus come to be governed by constitutive norms in the sense that if I do not use them in the conventional way, I am just not in the game. As Wittgenstein (1955) argued so trenchantly, the criteria for conventional use are determined not by the individual but by the community of users. I can rebel, but to what effect?

This cultural dimension to communicative conventions—everyone in the group is expected in cultural common ground to know them and to conform to them—means that we can now consider human communicative acts to be fully explicit. Early humans made overt the way they were perspectivizing situations for others, for example, by pointing out relevant situations with a gesture, but the recipient could easily misunderstand, or even feign misunderstanding, and that would be the end of it. But now, if a modern human uses a communicative convention such as a snake-danger gesture, the partner cannot claim not to know it or, under normal circumstances, not to comprehend it. Since we all know the convention in cultural common ground, it is explicit, and so one must respond. The modern human was thus now not only under a second-personal pressure from the communicative partner to comprehend but also, in a sense, under normative pressure from the entire community: if you are one of us, then you know how to operate with this convention. Anyone not comprehending this communicative convention is just not one of us—which makes it culturally normative.

Iconic communicative conventions can quickly become noniconic. This happens reliably in the birth of sign languages, as isolated deaf individuals, who have been communicating their whole lives with hearing parents using spontaneous iconic gestures, come together and conventionalize some of these “home signs.” This often leads to a kind of stylization or shortening of signs (see Senghas et al., 2004). Thus, the wavy hand for snake danger might become abbreviated to the point of almost no waving. Typically this is because the recipient can predict what is coming in the communicative situation; for example, if she is about to turn over a rock, as soon as someone sticks out his hand she can anticipate the snake-danger gesture. Children and other newcomers would then just imitate or conform to the abbreviated hand-out (with no waving) gesture to direct attention to snake-danger situations.2 Powerful skills of imitation and conformity thus undermine iconicity in communication, as iconicity is not necessary in a group with cultural common ground about what gesture to use for communicating about certain situations conventionally. Communicative conventions can thus become “arbitrary.”

The implications of this conventional-arbitrary way of doing things for the individual and her processes of thinking were, needless to say, momentous. For one thing, children were now born into a group of people using a set of communicative conventions that their ancestors had previously found useful in coordinating their referential acts, and everyone was expected to acquire and use exactly these conventions. Individuals thus did not have to invent their own ways of conceptualizing things; they just had to learn those of others, which embodied, as it were, the entire collective intelligence of the entire cultural group over much historical time. Individuals thus “inherited” myriad ways of conceptualizing and perspectivizing the world for others, which created the possibility of viewing one and the same situation or entity simultaneously under different construals as, for example, berry, fruit, food, or trading resource. The mode of construal was not due to reality, or even to the communicator’s goals, but rather to the communicator’s thinking about how best to construe a situation or entity so that a recipient would most effectively discern his communicative intention.

In addition to this fundamentally new form of conventional/normative and perspectival cognitive representation, communicating conventionally with arbitrary devices also creates, or at least facilitates, two other new processes of cognitive representation. The first is that the arbitrariness leads to a higher level of abstractness. Thus, when gestures are purely iconic, the level of abstraction is typically low and local. For example, with spontaneous iconic signs, opening a door is pantomimed in one way whereas opening a jar is pantomimed in another. This pattern is typical of the individually created home signs of deaf children, which are under pressure to remain iconic since there is no community of users with whom to conventionalize them. However, in a community, as the iconicity fades for new learners, and arbitrary conventions arise, there comes a more stylized depiction of open that is highly abstract and resembles no particular manner of opening. This abstractness is characteristic of many signs in conventional sign languages, and of course vocal languages as well. Conventionalization, after a drift to the arbitrary, breeds abstractness. One can imagine that acquiring large numbers of arbitrary communicative conventions could lead to some kind of general insight that most of the communicative signs we use have only arbitrary connections to their intended referents, and so, voilà, we can make up new ones as needed.

The second new process of cognitive representation created, or at least facilitated, by arbitrary communicative conventions also involves abstractness, but of a different type. Many of the most abstract conceptualizations in contemporary languages are single items for highly complex situations involving multiple agents doing things over time; for example, to define a term like justice, one would most naturally proceed with a kind of narrative: justice is when someone … and then someone.… It is difficult to imagine how to indicate for others complex situations and events such as justice in pantomime, except by acting out a kind of full narrative, and this is true even for more concrete narrative events like a celebration or a funeral, for which one would also have to pantomime whole sequences. But with arbitrary signs one may simply designate these complex situations with a single sign. This means that, in essence, arbitrary signs open up the novel possibility of symbolizing aspects of the relational, thematic, or narrative organization of human cognition—in addition to its straightforward categorical or schematized structure as symbolized by tree or eat—which expands the range and complexity of human thinking immensely. As argued in box 1 (chapter 3), the source of human conceptual access to this relational-thematic-narrative organization—which may now be designated with simple signs—is complex collaborative activities with joint goals and various constitutive roles. Markman and Stillwell (2001) refer to “role-based concepts” for the role slots (e.g., tracker in a hunt) and “schema-based concepts” for the overall activity (e.g., the hunting trip itself), and it is unlikely that any other organisms conceptualize this thematic dimension of experience.

Arbitrary communicative conventions—after there is a “critical mass”—also create two new processes of inference. First, because humans communicate for different purposes and at different levels of abstraction on different occasions, the individual in a community of conventional communicators inherits a large inventory of communicative conventions that relate to one another in complex ways. For example, one can imagine that in some contexts individuals would conventionalize a gesture or vocalization for indicating a gazelle, whereas on other occasions they would conventionalize a gesture or vocalization for indicating an animal in general (or maybe a potential animal prey of whatever species). Children in this culture would learn, in different contexts, both of these expressions. This now opens the possibility for not just causal inferences but formal inferences. If I indicate for you that a gazelle is over the hill, you may infer, based on your knowledge of things, that there is a potential animal prey over the hill, but you cannot make a similar inference in the opposite direction from animal to gazelle. Although in theory one could have spontaneous pantomimes at different levels of generality, it is only with collectively known conventionalized signs that communicators can be certain that recipients have the conventional means necessary to make such formal inferences—and so count on these inferences in formulating their communicative acts.

Second, arbitrary communicative conventions come to form a kind of “system” such that, precisely because of their arbitrariness, the referential range of one is constrained by the referential range of others in the same “semantic field” (Saussure, 1916). It is thus in our cultural common ground that I am making a choice between certain conventional expressions that we both know together I have available to me. For example, if I report to a friend that I saw his brother foraging with “a woman,” the inference is that this was not his wife, even though his wife is a woman, too, because if it had been his wife I would have said “wife” and not “woman.” Or, if I say that our child ate “some of the meat,” the inference is that he did not eat all of it because—at least in the context where we are both quite hungry—if he had eaten all of it I would have said so. These kinds of pragmatic implicatures permeate discourse among contemporary language users based on the common cultural ground that we both have a certain inventory of conventional linguistic expressions from which we choose for communicative purposes. (Some of these inferences are recurrent and thus become what are often called, contentiously, conventional implicatures [Grice, 1975; Levinson, 2000].) Inferences of this type are not generated in the same way by spontaneous pantomimes or any other kinds of unconventionalized signs because in these cases it is not in the cultural common ground of the group that everyone knows all of the alternatives and so will be making inferences about a communicator’s choice among them.

And so, with the advent of communicative conventions, we now have some new forms of conceptualization. Modern humans “inherit” a set of communicative conventions in their cultural common ground with others in the group, and the use of these conventions is normatively governed, in the sense that deviance from them puts one outside the cultural practice. The arbitrariness of communicative conventions means that they may be used to conceptualize situations and entities of almost unlimited abstractness, including role-relational, thematic, and narrative schemas. And with communicative conventions, we now have collectively known inferential connections among conceptualizations, both formal and pragmatic, which were not possible in the same way with natural gestures.

Linguistic Constructions as Complex Representational Formats

If we imagine early modern humans with a small inventory of single-unit (holophrastic) communicative conventions, along with general cognitive abilities for creating novel mental combinations (possessed by all apes), we can easily imagine them creating multiunit linguistic combinations. And so, for example, perhaps there was a communicative convention for requesting eating by moving the hand to the open mouth. And perhaps there was an unrelated communicative convention for requesting going foraging for berries (miming a picking motion). It would not take a genius, in a situation in which someone offered something unpalatable to eat, to gesture eating followed by berries. Then, given an existing ability for schematization (possessed by apes and early humans, but now applied to conventions), one could imagine this individual generalizing the conventional eating gesture to other conventional food gestures, much like human toddlers do as they first say things like “more juice,” followed soon by “more milk,” “more berries,” and so forth in a “more X” pattern (so-called item-based schemas; Tomasello, 2003a).

Linguistic constructions begin with such simple item-based schemas but then are elaborated and made more abstract through discourse interactions. The key aspect of the process for current purposes is the communicative pressure—the demand for sufficient information—coming from the recipient. This forces the communicator to make things explicit that he might otherwise have left implicit. Communicators stutter out a string of different utterances, pidgin style, and recipients have to fill in the gaps inferentially. But communicative breakdowns occur, and recipients demand more information about how to relate the bits to one another, and so communicators must be more explicit about their communicative intentions. This process—in combination with skills for integrating and automatizing sequences—turns things like “I spear antelope … he dead” into “I speared the antelope dead.” When other, similar schemas are created (e.g., “I drank the gourd empty”), the result is a conventionalized linguistic construction, in this example the resultative construction (Langacker, 2000; Tomasello, 1998, 2003b, 2008). In the words of Givón (1995), today’s syntax is yesterday’s discourse.3

And so are born fully abstract linguistic constructions, which become gestalt-like symbolic conventions in their own right, with their own abstract communicative significance indicating different types of situations. For example, young English-speaking children learn early abstract constructions for (1) straightforward causal situations (e.g., the transitive construction: X VERBed Y); (2) causal situations viewed from the perspective of the affected object (e.g., the passive construction: Y got VERBed by X); (3) situations of object movement (the intransitive locative construction: X VERBed to/into/onto Y); (4) situations of transfer of possession (e.g., the ditransitive construction: X VERBed Y a Z); (5) situations in which an agent acts with no affected object (e.g., the unergative intransitive construction: X smiled/cried/swam); (6) situations in which objects change states with no specification of the agent or cause (e.g., the unaccusative intransitive construction: X broke/died); and so forth (Goldberg, 1995). Importantly, the communicative function of these abstract patterns is independent of the particular words used in them, as these schematic depictions illustrate.

In using a construction, communicators invite recipients to view or imagine situations from a particular perspective. Thus, in a conventional language there are various ways of designating subject, as perspectival topic, no matter who is performing or receiving the action. Thus, one and the same action may be referred to as “John broke the window,” “The window got broken by John,” “John’s throwing of the rock broke the window,” “The window was broken by John’s throwing of the rock,” “The rock broke the window,” “The window was broken by the rock,” and so on, depending of how the speaker wishes to conceptualize the situation for the listener on a particular occasion of use. Other constructions perspectivize the situation based on the communicator’s judgment of the recipient’s knowledge and expectations. For example, the English cleft construction, as in “It was John who broke the window,” is used to indicate that John did the breaking when the recipient currently believes that someone else did; that is, it is used as a correction to a mistaken belief (e.g., You: “Bill broke the window.” Me: “No, it was John who broke the window”). MacWhinney (1977) argues that these different construals derive from what communicators choose as their “starting point” or “perspective” for entering the event cognitively—which is conventionalized into the grammatical topic or subject.

From a cognitive point of view, abstract constructions give humans a new kind of abstract, syntagmatically organized conventional format for cognitive representation. These abstract constructions enable linguistic items to be used, and then reused, in a wide variety of different constructions, playing different roles on different occasions. Importantly, this flexibility in how items are used creates the need to explicitly mark the roles being played by different items. If I gesture or vocalize mantigereat, it is important to know who is the agent and who is the patient of the eating activity. Modern-day languages have a variety of means for doing this, such as case marking and contrastive word order. Markers used to indicate participant roles may be seen as kind of second-order symbols, because they are about the role the participant is playing in the larger construction (Tomasello, 1992).4 Importantly for current purposes, Croft (2001) argues that the linguistic items in an utterance gain their communicative functions not via their syntactic relations to other items but, rather, from the collaborative syntactic role they play in the utterance/construction as a whole. A linguistic construction may thus be seen, in a way, as a symbolic collaboration.

Abstract constructions are thus the main source of conventional linguistic productivity, and thus conceptual productivity in thinking. Individuals schematize and analogize to create abstract constructions, and they then can readily put new items into the slots of these constructions based on a more or less good fit with the communicative role of that slot. Indeed, using an item in a construction’s slot coerces a construal that may be atypical for that item; for example, we quite often say such things as “he treed the cat,” “he ate his pride,” “he coughed his age,” and so forth, which force atypical construals on some of the items. Such metaphorical or analogical thinking is testament to the fact that the construction itself has its own communicative function—from the top down, as it were—into which, up to some limit, items must be forced (Goldberg, 2006). In all, this system of conventionalized abstract constructions—and items that can be used and reused as needed in different constructions—enables the kind of creative conceptual combination by which humans may think, or at least try to think, of everything from flying toasters to colorless green ideas sleeping furiously.

All of this is about how the communicator specifies the referential situation that he and the recipient are communicating about during their ongoing second-personal communicative interaction. In addition, modern human communicators often use language as well to specify things about their communicative motive and their modal or epistemic relation to this referential material in their ongoing second-personal communicative interaction. Critically, this is something almost wholly new in the communicative process. Early humans made reference to things in the world in various ways, but their own relation to that referential material was left implicit—perhaps expressed unintentionally (procedurally) in a facial expression or vocalization, but not an intentional part of the communicative act under the communicator’s voluntary control in decision making.

But now modern human communicators explicitly indicate their communicative motive in communicative conventions. Thus, most languages have different constructions for speech acts such as requestives and informatives (assertives). Philosophers of mind and language believe it to be very important that the “same” fact-like propositional content may be used in different constructions with different communicative motives, as in “She is going to the lake,” “Is she going to the lake?,” “Go to the lake!,” “Oh, that she could go to the lake,” and so forth. The idea is that this independence of the propositional content from any particular “illocutionary force” makes the propositional content into a kind of quasi-independent, fact-like entity, free of particular instantiations in particular linguistic utterances (e.g., Searle, 2001). As speech act function came to be expressed conventionally, in either linguistic items or constructions as a whole (as above), both the communicative motive and the propositional content were now conventionalized into the same representational format of words and constructions. In this totally new communicative move, the communicator’s motive in the communicative interaction is now itself referred to and so conventionally conceptualized. This fact, among others, is what prompted Wittgenstein’s (1955, #11) observation that a major difficulty in understanding how language works is that very different functions are all expressed in the same basic way in words and constructions: “Of course, what confuses us is the uniform appearance of words when we hear them spoken.… For their application is not presented to us so clearly.”

In addition, communicators indicate through various linguistic devices their modal or epistemic “attitude” toward some propositional content in an utterance. Thus, a communicator might opine, modally, that “She must go to the lake” or that “She can go to the lake,” or, epistemically, that “I believe she is going to the lake” or that “I doubt she is going to the lake.” Again, the evolutionary raw material for the conventionalization of modal and epistemic attitudes is presumably facial expressions and prosody as expressed simultaneously with some utterance, for example, uncertainty, surprise, or indignation. But then these became conventionalized.5 Communicators thus now encase propositional content also in a kind of “modal-epistemic envelope” (Givón, 1995)—again, with all items in the same conventionalized representational format of words and constructions—that further encourages us to conceptualize them as quasi-independent mental entities. In this case, the independence is not only from speaker motive but also from how speakers feel or think about them. This distinction between content and attitude is also foundational to the idea of some kind of timeless, objective, propositionally structured facts that are independent of how anyone thinks or feels about them and, therefore, also to the general idea of an independent, “objective” reality.

If we now combine all of the distinctions we have made—that is, those that the communicator actively controls (and for which he must choose from among alternatives)—we have the basic structure of a conventional linguistic utterance: the force-content distinction, within that the attitude-content distinction, and within that the topic-focus (subject-predicate) distinction, as shown in Figure 4.1.

image

FIGURE 4.1  The basic structure of a conventional linguistic utterance

Overall, then, we may say that linguistic constructions are conventionalized and automatized segments of discourse that organize human experience into abstract patterns of various sorts, as individuals conceptualize things for others in communication. Constructions contain abstract roles—such as agent, recipient, location—marked via second-order symbols such as case markers, adpositions, or contrastive word order. The possibility of placing an almost unlimited inventory of linguistic items into these role slots is a major source of creative conceptual combination (see Clark’s [1996] famous “The newspaper boy porched the newspaper”). Topic-focus (subject-predicate) organization within specific constructions serves to conceptualize situations from the perspective of one or the other of the roles. Communicative devices for indicating speaker motives, along with the modal-epistemic envelope, serve to partition out fact-like propositional content as indicating some kind of timeless, objective facts about an objective world, independent of how anyone thinks or feels about them. These are all aspects of human linguistic communication unique to the species (see box 3).

Discourse and Reflective Thinking

Once we have linguistic communication, we have discourse. And what happens in discourse quite often is that the recipient responds to an utterance by signaling noncomprehension, requesting clarification, and so forth. The communicator then does his best to provide the needed information explicitly in further discourse. The key point for human thinking is that explicating conceptual content in conventional linguistic format (content that was only implicit in some original communicative act) makes this content ripe for self-reflection. That is to say, invoking again the analysis of Mead (1934; and to some degree that of Karmiloff-Smith, 1992), the collaborative nature of human communication means that the communicator can perceive and comprehend his own communicative act as if he were the recipient, which enables him to think about his own thinking, from an external perspective, as it were (see also Bermudez, 2003). Although the pointing and pantomiming of early humans enabled them to engage in some degree of reflection on their own overtly expressed thoughts, with modern humans and conventional linguistic communication, some new types of thoughts could now be expressed. Moreover, now the self-monitoring process came not just from the perspective of the recipient, but from the normative perspective of all users of the conventions. Three especially important examples are as follows.

First, one important piece of information often needing explication is the communicator’s intentional states (or propositional attitudes). For example, let us suppose that on my way back from a hunting trip I see gazelles drinking at watering hole no. 2, leading me to infer that their preferred watering hole no. 1 is currently dry (given the recent dry weather). Back at the home base, you inform me that you are headed to watering hole no. 1 to fetch water. I want to inform you that it very likely does not have water, but I do not want to just state as fact “It does not have water” since I am not certain. Presumably the first marker of speaker uncertainty used in such situations was some involuntary facial expression (see above). But then humans conventionalized ways of indicating doubt, for example, by saying something like, “Maybe it does not have water” or “I think it does not have water.” Interestingly, young English- and German-speaking children first use words for thinking not to indicate a specific mental act of thinking but rather to express their uncertainty in the same way as maybe (thus, “I think it does not have water” means maybe it does not; Diessel and Tomasello, 2001). Only later do they make explicit reference to third-person mental happenings. And so one hypothesis is that it was the demands of discourse that led humans to begin talking explicitly about mental states, and they did not do this across-the-board initially, but only for their own epistemic attitudes toward propositional contents. Later, they came to refer to the mental states of anyone and everyone, including both others and the self, with the exact same set of communicative conventions. Once humans could make explicit reference to intentional states, they could think reflectively about them in some new ways.


BOX 3. The “Language” of Kanzi et al.

Over the past few decades, a handful of great apes have been raised by humans and taught some form of human-like communication. They end up doing very interesting things, but it is not clear in which ways they are human-like and in which ways they are not. With particular regard to linguistic constructions, there is no doubt that apes can combine their signs, sometimes creatively, but they do not seem to have anything resembling human constructions (even though they are perfectly capable of schematizing conceptual content in general). Why is this? To help answer this question, here are some examples of the kinds of utterances they produce, with either manual gestures or human-provided visual symbols (not resembling their referents):

BITE BALL—wanting to do this

GUM HURRY—wanting to have some

CHEESE EAT—wanting to

You(point) CHASE me(point)—requesting it from other

The first thing to note is that these are all requests, reflecting the fact that systematic studies have found that over 95% of the communicative acts produced by these individuals are some form of imperative (and the other 5% are questionable; Greenfield and Savage-Rumbaugh, 1990, 1991; Rivas, 2005). This is because no matter how they are trained by humans, great apes will not acquire a motive to simply inform others of things or share information with them (Tomasello, 2008). And in strictly imperative communication, there is little functional need for all the complexities of human linguistic communication (prototypically, no subject, no tense, etc.).

Nevertheless, many of the communicative acts produced by these individuals are clearly complex, structured by a kind of event-participant structure, reflecting a partitioning of situations into the participants involved in the relations or actions indicated. But despite this complexity, some key things relative to human linguistic communication are missing. Basically what is missing is all of those aspects of human grammar that conceptually structure constructions for others and their knowledge, expectations, and perspective. Over and above events and participants (and perhaps locations), the linguistic apes have learned items that indicate their own desire (e.g., although not apparent on the human gloss, their use of the item “hurry” indicates that they want it now). But what is missing is all of those aspects of syntax that are geared at making the utterance comprehensible to the recipient—a key part of the cooperative motive. For example:

• They do not “ground” their acts of reference for the listener to help them identify the referent. That is to say, they do not have noun phrases with things like articles and adjectives that help to specify which ball or cheese is wanted, for example. Nor do they have any kind of markers of tense that would indicate which event, as indicated by when it occurred, they intend to indicate.

• They do not use second-order symbols such as case markers or word order to mark semantic roles and so to indicate who is doing what to whom in the utterance. Communicators do not need this information; it is provided to make sure that the listener understands the role of each participant in the larger situation or event being communicated about.

• They do not have constructions or other devices for indicating for listeners what is old versus new versus contrasting information. For example, if you adamantly expressed that Bill broke the window, I probably would correct you by using a cleft construction and say “No, it was FRED that broke the window.” Apes do not have such constructions.

• They do not choose constructions based on perspective. For example, I might describe the same event either as “I broke the vase” or “The vase broke,” based on your knowledge and expectations and my communicative intentions, whereas linguistic apes have not learned constructional alternatives of this type.

• They do not specifically indicate in their utterances their communicative motive (why should they, since it is always requestive) or anything of their epistemic or modal attitudes toward the referential situation.

The key theoretical point is that, beyond just supplying ordering preferences for utterances, human linguistic constructions are created with adaptations for the recipients’ knowledge, expectations, and perspective in mind. And even very simple constructions like noun phrases require adaptations to the recipient’s knowledge, expectations, and perspective. Humans also conventionalize expressions of motives and epistemic and modal attitudes in their constructions. Call all of this the pragmatic dimension of grammar, and call it uniquely human.


A second set of cognitive processes often in need of explication is the communicator’s logical inferring processes. These include most prominently those indicated by and and or, various kinds of negation (e.g., not), and implication (if … then …). For example, in response to pressure from the recipient in argumentative discourse, the speaker requires terms such as these to make explicit his reasoning processes. And so, analogous to communicative pressure in normal discourse, “logical pressure” in argumentative discourse forces disputants to make explicit in language the logical operations that until that time were only procedural and not representational at all. One can imagine a first gestural/iconic step in which, for example, or is expressed in some kind of pantomiming in which one offers someone either this object (held out with one hand) or that object (held out with the other). An “if … then …” implication could be acted out in pantomiming such everyday social interactions as threats and warnings (if X … then Y). But, as always, symbolizing these logical operators in linguistic conventions would make them much more abstract and powerful and, once again, much more readily available for self-monitoring and self-reflection.

Third, speakers are often forced to make explicit some of the background assumptions and/or common ground to help the recipient to comprehend. For example, assume we are foraging together for honey, a cultural practice with which we are both very familiar from our cultural common ground. The knowledge we share about this practice—what kind of hive we are looking for, the height in the tree we should scan, the tools we will need, the container we will need for transport, and so forth—directs many of our activities. Thus, if you go off and start picking and weaving together leaves, I wait for you patiently as we both know that a vessel will be needed for transport. But this shared knowledge is all implicit in our (cultural) common ground. An early human might make this knowledge overt by pointing out to his partner the presence of some appropriate leaves. But now imagine that I, as modern human, express my intention that you notice the leaves’ presence by means of some shared communicative conventions: “Look, there are some good leaves over there.” This draws your attention to the leaves in a much more explicit way, but there is still room for misunderstanding (good for what?). So perhaps you look over at the leaves but draw a blank. Depending on my assessment of what you are not comprehending, I might say, “Its banyan leaves,” or “We are going to need a vessel,” or “We need banyan leaves to make the vessel,” or whatever. I am making explicit for you the reason I am directing your attention to the leaves’ presence (which I erroneously thought you could infer from our common ground), and, in the process, make explicit the bases for my own thinking for communicating. Once more, this makes it possible for me to reflect on my thoughts and their connections in a way that I could not when they were only an implicit part of our common ground.

And so with modern humans such things as intentional states, logical operations, and background assumptions could be expressed explicitly in a relatively abstract and normatively governed set of collectively known linguistic conventions. Because of the conventional and normative nature of language, new processes of reflection now took place not just as when apes monitor their own uncertainty in making a decision, and not as when early humans monitor recipient comprehension, but rather as an “objectively” and normatively thinking communicator evaluating his own linguistic conceptualization as if it were coming from some other “objectively” and normatively thinking person. The outcome is that modern humans engage not just in individual self-monitoring or second-personal social evaluation but, rather, in fully normative self-reflection.

Shared Decision Making and the Giving of Reasons

We must single out, finally, a very special discourse context in human communication with world-changing implications for the process of human thinking: shared decision making. Prototypically, we may imagine as an example collaborative partners—or even a council of elders—attempting to choose a course of action, given that they know together in common ground that multiple courses of action are possible. Given their equal power in their interdependent situation, they cannot just tell the other or others what to do; rather, they must suggest a possible course of action and back it up with reasons.

Let us start with early humans. Because early human collaborators typically had much in common ground, they could point and pantomime in ways suggesting reasons implicitly. Thus, we may imagine two early humans following an antelope. They lose sight of the beast, and so pause in a clearing to make a joint decision about which way to go. One individual might in this context point to some tracks on the ground. They are relevant to the hunters because it is in their common ground that these are antelope tracks, possibly of the animal they were following. Also relevant is the direction of the tracks, which again is significant for the two of them because they know in common ground what this means for the antelope’s likely travel direction. The point is that in directing his partner’s attention to the tracks, the communicator’s goal is that his partner travel with him in a certain direction. But he is not pointing in that direction; he is only pointing to the ground. The communicator’s act is thus providing a kind of implicit reason for the recipient, which we may gloss as: see the tracks; given our common ground about what they mean for our prey’s likely travel direction, they give us a reason for going in this direction. The recipient might counter by pointing in a different direction, where they spy the antelope’s offspring next to some bushes, which is a better reason for traveling in this other direction. None of these reasons is explicit, of course, and so it does not constitute what we might call fully reasoned thinking. But it is a start.

With modern humans and their skills of conventional linguistic communication, we get to full-blooded reasoning, where “reasoning” means not just to think about something but to explicate in conventional form—for others or oneself—the reasons why one is thinking what one is thinking. This conflicts with the traditional view that human reasoning is a private affair. Most articulate on this point are Mercier and Sperber (2011) who recast the reasoning process in terms of communication and discourse, specifically argumentative discourse in which individuals make explicit to others their reasons for believing something to be the case. The basic idea is this: When a communicator informs a recipient of something, she wants to be believed, and often is (based on mutual assumptions of cooperation). But sometimes there is not enough trust on the recipient’s part (for whatever reason), and so the communicator gives reasons for her informative statement. In reason-giving discourse of this kind, individuals are attempting to convince others. Many lines of evidence suggest that the main function of reasoning is to convince others, for example, people’s tendency to look for supporting rather than disconfirming evidence (the confirmation bias). In this view, convincing others is good for individual fitness, and so humans evolved reasoning abilities not for getting at the truth but for convincing others of their views.

The proposal that human reasoning, including individual human reasoning, has a social-communicative origin is almost certainly correct. But Mercier and Sperber’s account tends to background the cooperative processes involved, and so here is an alternative account that foregrounds these processes: The key social context is joint or collective decision making, as it occurred regularly in collaborative activities. Thus, on a hunting trip, perhaps you think we should hunt for antelopes in this direction, and I think we would be better off going in that direction. To make your case, you make your reasoning more explicit in our conventional language by, for instance, noting that there is a watering hole to the south. I counter, also in language, by making explicit my reasoning that at this time of day it is likely that lions will be at the watering hole and so no antelopes will be—and besides, here are some antelope tracks going to the north. You say these tracks look old, but I think that is because they were in the direct sunlight this morning and actually they are from around dawn or so. And on and on. The key point is that arguing in this way assumes a cooperative context. As Darwall (2006, p. 14) puts it: “It is only in certain contexts, say, when you and I are trying to work out what to believe together, that either of us has any standing to demand that one another reason logically.”

Such cooperative argumentation, as we may call it, may be modeled in game theory as a battle of the sexes: our highest goals are collaborative—we will hunt together under all circumstances because otherwise there is zero hope of success—but within that cooperative framework we each argue our case. Critically, in this context, neither of us wants to convince the other if we are in fact wrong about the location of antelopes; each would rather lose the argument and eat tonight than win the argument and go hungry. And so a key dimension of our cooperativeness is that we both have agreed ahead of time, implicitly, that we will go in the direction for which there are the “best” reasons. That is what being reasonable is all about.

An appeal to “best” reasons invokes what Sellars (1963) calls “common standards of correctness and relevance, which relate what I do think to what anyone ought to think.” Our cooperative argumentation in the context of joint or collective decision making is thus premised on a shared metric that we both use in determining which reasons are indeed “best.” There have thus arisen social norms that govern cooperative argumentation in group decision making specifying, for example, that reasons based on direct observation trump reasons based on indirect evidence or hearsay. An even deeper, conceptual point is that to be in an argument in the first place means to accept as infrastructure certain “rules of the game,” namely, the group’s social norms for arguing cooperatively. This is the difference between a street fight and a boxing match. The early Greeks made explicit some of the most important of these norms of argumentation in Western culture, for example, the law of noncontradiction (a disputant cannot hold the same statement to be both true and false at the same time), and the law of identity (a disputant cannot change the identity of A during the course of the argument). Even before the Greeks, we can imagine that individuals who, for example, held the same statement to be both true and false at the same time were either ignored by others or else exhorted to argue rationally. The cooperative infrastructure was thus decisive in determining what it means to reason at all. The natural world itself may be totally “is”—the antelopes are where they are. However, the culturally embedded discourse processes by which we determine what that “is” in fact is—in the space of reasons, to use Sellars’s evocative phrase—are fraught with ought.

Cooperative argumentation would thus have been the birthplace of “assertive” speech acts. Assertions go beyond the informative speech acts from which they derive in that the asserter commits himself to the truth of a statement (i.e., I commit not just to honesty but to the objective truth of the statement) and, crucially, to backing it up with reasons and justifications as necessary. Reasons and justifications are intended to make explicit to others the bases on which I believe something, which, because they share these bases, might give them reason to believe it as well (e.g., we all know and accept ahead of time that if there are lions at the watering hole, then there will be no antelopes). One may also reject an argument because it violates the norms of argumentation (e.g., you just contradicted yourself) or implies something that we both know is not true. Overall, this ability to connect thoughts to other thoughts (both those of others and one’s own) by various inferential relations (prototypically by providing reasons and justifications) is key to human reason in general, and it leads to a kind of interconnection among all of an individual’s potential thoughts in a kind of holistic “web of beliefs.”

The capstone of all of this—recognized by all modern thinkers who take a sociocultural view of human thinking—is the internalization of these various interpersonal processes of making things explicit into individual rational thinking or reasoning. Making things explicit to facilitate the comprehension of a recipient leads the communicator to simulate, before actually producing an utterance, how his planned communicative act might be comprehended—perhaps in a kind of inner dialogue. Making things explicit to persuade someone in an argument leads the disputant to simulate ahead of time how a potential opponent might counter his argument, and so to make ready, in thought, an interconnected set of reasons and justifications—again, perhaps, in a kind of inner dialogue. As Brandom (1994, pp. 590–591) describes the process: “The conceptual contents employed in monological reasoning … are parasitic on and intelligible only in terms of the sort of content conferred by dialogical reasoning, in which the issue of what follows from what essentially involves assessments from the different social perspectives of scorekeeping interlocutors with different background commitments.”

The norms of human reasoning are thus at least implicitly agreed upon in the community, and individuals provide reasons and justifications as ways of convincing “any rational person.” Human reasoning, even when it is done internally with the self, is therefore shot through and through with a kind of collective normativity in which the individual regulates her actions and thinking based on the group’s normative conventions and standards—what some have called “normative self-governance” (e.g., Korsgaard, 2009).

Agent-Neutral Thinking

The second-personal thinking of early humans was aimed at solving coordination problems presented by direct collaborative and communicative interactions with specific others. Modern humans faced different kinds of coordination problems, namely, those involving unknown others, with whom one had little or no personal common ground. The solution on the behavioral level was the creation of group-wide, agent-neutral conventions, norms, and institutions, to which everyone expected everyone, in cultural common ground, to conform. To coordinate with others communicatively in such a world, human communication had to be conventional as well, based again not on personal but rather on cultural common ground. And to be a good communicative partner in conventional communication—especially, to be a cooperative participant in shared decision making—modern humans needed to express their reasons for thinking in certain ways explicitly in language and then simulate the cultural group’s normative judgments of the intelligibility and rationality of those linguistic acts and reasons. Modern humans participate not only in joint intentionality with other individuals but also in collective intentionality with the entire cultural group.

Representing “Objectively”

Early humans cognitively represented to themselves various situations and entities simultaneously from more than one perspective, and they then indicated or symbolized particular perspectives on those situations and entities for others in their deictic and iconic communicative acts. Modern humans then began collaborating and communicating with sometimes unfamiliar others structured by agent-neutral conventions, norms, and institutions, so that the cognitive models they were building and the perspectives they were simulating concerned not just particular others but, rather, some kind of generic other or, perhaps, the group at large. The linguistic conventions individuals were born into embodied the way that the group as a whole, from many years past, perspectivized and schematized experience, so that this way seemed inevitable. This new way of operating socially led to cognitive representations with three important new features.

CONVENTIONAL.  For the first time in the history of life, modern human individuals “inherited” a culturally constructed representational system in the form of a conventional language, comprising a structured inventory of conceptualizations that forebears in the culture had previously found useful in communicating with others. The uses of linguistic conventions were shared within the cultural common ground of the group, and this meant, given human group-mindedness and conformity, that they became normatively grounded in “community standards” governing their proper use. This made it seem, especially to language-acquiring children, that the way that one’s linguistic conventions carved up the world was somehow natural.

In addition, the arbitrariness of linguistic conventions created, or at least facilitated, the ability to operate with highly abstract conceptualizations such as justice or blackmail that schematize not a taxonomic class but, rather, a thematically or narratively defined entity. The arbitrariness of linguistic symbols also led to more abstract conceptualizations of relatively concrete terms, such as open or break, across different particular situations. Most important, because of their conventional nature, linguistic conventions and their interrelations made possible conceptualizations with explicitly contrasting aspectual shapes—gazelleanimaldinner—known to everyone in the cultural common ground of the group. These contrasting aspectual shapes created a gap for the individual between the world as she conceptualized it for her own instrumental actions and the world as conceptualized in various contrastive ways in her conventional language, a gap that humans have been pondering from the early Greeks to Benjamin Lee Whorf.

PROPOSITIONAL.  Modern humans began to use linguistic conventions together in combination in patterned ways that led to the creation of abstract linguistic constructions as kinds of linguistic gestalts. Many linguistic constructions conceptualize whole propositions, and they do this with internal constituents that are marked with second-order symbols as playing specified roles in the construction. Proposition-level linguistic constructions are perspectival (e.g., active vs. passive), and one of the elements (subject) provides a perspectival entry point into the conceptualized situation. The abstractness of linguistic constructions makes possible especially productive conceptual combinations, so that we may represent to ourselves all kinds of imaginary entities and situations, from a happy sun to a man in the moon. Linguistic constructions thus create the possibility of various kinds of metaphorical representations in which structural analogies provide a new framework for thinking, from one idea “undermining” another to activities “eating up” my free time. Making explicit in linguistic constructions various kinds of communicative motives and attitudes contributed to an objectified view of experience, as it suggested a fact-like propositional content independent of the desires or attitudes of any particular individual.

With their linguistic constructions, modern humans also began to make assertions—to whose objective truth they were committed—that could be either about particular episodic events or else, importantly, about generic events or facts of a type. They did this especially in norm enforcement (“One does not do that in public”) and teaching (“It works this way”). This genericness presumably originated from the normative “group voice” that lay behind the assertive expression and gave it an objectivity that transcended the individual.

“OBJECTIVE.”  Early humans lived in a world of different individual perspectives. Modern humans live in this world too, but in addition, in the context of group-minded cultures, there arose a kind of public world comprising collectively created entities such as conventions, norms, and institutions from marriage to money to governments. These entities existed before the individual arrived on the scene, and they existed independent of the thoughts and wishes of any single individual, giving them the same kind of “always already there” status as the physical world. In addition, these collective entities had preestablished roles into which, in theory, any agent could seamlessly fit; and, indeed, in some cases these roles created new realities, such as presidents and money, whose very real deontic powers were readily observable. Operating in this public world required that individuals be able to take a kind of agent-neutral perspective on things, a kind privileged, “transcendental” perspective that constituted the world “objectively” and that then justified personal judgments of true and false, right and wrong.

As modern human individuals were building their cognitive models of the world, the use of simple causal and intentional relations was not enough. To explain such things as chiefs and marriage, not to mention language and culture, they needed some understanding of things created by collective agreement and maintained by collective normative judgment. Said another way, they needed some new conceptualizations of collective realities that transcended the thoughts and attitudes of single individuals, even multiple individuals. Constructing such models would lead naturally to judgments such as realtrue, and right that come not from the individual herself but rather from her appropriation of the transpersonal, “objective” perspective engendered by her cultural world. Linguistic representations—especially assertions with a distinction between the second-personal attitudes of the communicator and some timeless, generic propositional content (e.g., “I think its raining”)—only added additional force to these objectifying and reifying tendencies. Modern humans thus “collectivized” early humans’ ways of life, and so “objectified” their cognitive models of the world.

Reasoning Reflectively

The inferences of the common ancestor to humans and great apes were simple causal and intentional inferences. The inferences of early humans were recursively structured, enabling them to produce and interpret cooperative communicative acts comprising nothing but a protruding finger. But now, the linguistic communication of modern humans opened up whole new vistas of inference and reasoning. We now have such things as formal and pragmatic inferences, and external communicative vehicles can be reflected upon by the communicator from an objective and normative perspective. And the giving of reasons and justifications to others—and to the self in internal reasoning—now serves to connect up an individual’s various conceptualizations into a single inferential web.

LINGUISTIC INFERENCES.  The hierarchical relationship between the referents of different linguistic conventions is part of the conventionalization process. Thus, everyone knows collectively that everyone uses gazelle only for a particular type of animal, and animal for all kinds of animals, of which gazelle is one type, and so we now have the possibility of formal inferences: if we know that a gazelle is over the hill, then we know that an animal is over the hill (but not the reverse). Much of the early development of formal logic was built on inferences of this type, and in contemporary conceptual role semantics, inferences of this type play an important role as well. Also crucial is the fact that we all know collectively that we all know the linguistic options available to a communicator, which leads to the kind of pragmatic inferences that Grice (1975) made famous: if I refer to someone as an “acquaintance,” that almost certainly means that we are not friends—because if we were friends I would have used the word friend. These implicatures and corresponding inferences are possible because, and only because, the options available are part of the group’s cultural common ground, so that we can wonder together why I made the choice that I did. Conventional linguistic communication thus makes possible powerful new kinds of inferences.

In addition, linguistic communication, and the arbitrary nature of linguistic conventions, enabled modern humans to express explicitly in language some conceptualizations that could not be expressed easily, if at all, in the natural gestures of early humans, for example, intentional states and logical operations. Based on the hypothesis that one can reflect on one’s thinking only as it is expressed in external behavior directed at another—because only then can one play the other’s role and attempt to comprehend it from her perspective—linguistic communication now made available to modern humans many new conceptualizations about which they could think reflectively. Importantly, as modern humans thought about their own thinking reflectively, at least in some situations, they did not do so merely from their own perspective, or even that of the other, but from a more “objective” perspective.

REFLECTIVE INFERENCES.  A special discourse situation is cooperative argumentation, in which we attempt to come to a group decision about either action or beliefs. We do this not only by making assertions, committed to the truth, but also by backing up those assertions with reasons and justifications, which means making connections to things that are collectively agreed to be true and reliable. The outcome of this process is that the various conceptualizations and propositionally structured thoughts of modern humans as expressed in language become ever more inferentially interconnected in a vast “web of beliefs,” such that each element in the web gains significance from its inferential relations with others. This interconnectedness is a key component in being a fully rational creature who “knows his way about” an entire conceptual system in which propositionally structured thoughts provide reasons and justifications for one another (i.e., they can be used as premises and conclusions for one another in argumentation; Brandom, 2009).

In addition, as modern humans began engaging in such cooperative argumentation, implicitly accepted norms emerged. These norms operated such that individuals who contradicted themselves from one assertion to the next, or changed the meaning of their terms in the middle of an argument, or held a single assertion to be both true and false, were basically ignored or excluded from the group decision-making process. The process of cooperative argumentation, then, was the special language game within which human norms of rationality came to govern all those who wanted a voice in collective decisions (such as political, judicial, and epistemic decisions).

All of this may be internalized. Internalization means simply that one directs a communicative act, as communicator, to oneself, as recipient, including holding the “other” to “objective” normative criteria of intelligibility, cooperative participation, and so on. The resulting internal dialogue is one especially salient type of human thinking (Vygotsky, 1978). When the communicative context is cooperative argumentation, what now get internalized are whole lines of argumentation and justifications for arguments. Now, an individual can give to himself a normatively justified reason for why he is thinking what he is thinking, and so his conceptualizations become defined, in large part, by their normatively sanctioned inferential relations with other conceptualizations. The resulting web of beliefs, and humans’ ability to navigate this web facilely, is foundational for the ability to engage in individual reasoning.

At this point, then, the inferring of modern humans is not just imagining causal and intentional sequences, as in apes, or even just perspectivizing and recursivizing them, as in early humans; rather the inferring of modern humans now includes new kind of inferences made possible by a conventional language, and new forms of reflecting on their own thinking, including in a kind of inner dialogue. When these processes operate in the special context of cooperative argumentation, the result is something we might call reasoning. Modern humans thus are on occasion able to engage in a kind of reasoned, or reflective, inferring in the context of the normative standards of the cultural group.

Normative Self-Monitoring

Early humans engaged in what we have called cooperative self-monitoring—regulating their collaborative activities by the evaluative reactions of specific partners; and communicative self-monitoring—regulating their communicative acts by the anticipated interpretations of specific partners. Scaling up these processes to the cultural way of life characteristic of modern humans means that individuals now regulate their behavioral decision making instead via the collectively known and collectively accepted norms of the cultural group. Modern humans thus came to feel not only a second-personal pressure in their decision making but also, on top of this, as it were, a group-level normative pressure to conform to the group. Thus, I do not renege on my commitments, first of all, because I do not want to disappoint my partner, and second of all, because “we” in this group do not treat others like that. This more generalized normativity thus ends up back at group identity: if I want to be a member of this group, I must behave as they do, that is, follow the norms to which we all together (including me) have committed ourselves.

Modern human thinking and reasoning become normatively structured and governed in multiple ways. When one communicates with others via communicative conventions, one needs to do it in the way that they do it to participate successfully. In addition, in the context of group decision making and cooperative argumentation, one must agree to certain norms of argumentation. Others in the group decision making have a stake in me participating in useful ways, and so everyone has a stake in everyone else making true assertions, following the norms of inference and argumentation, justifying by connecting to propositions and arguments already collectively accepted, and so forth. Internalized, this communicative process becomes individual reason.

NORMATIVE SELF-GOVERNANCE.  Normative self-governance results from an internalization of the processes of collective normativity, as the individual self-monitors and, indeed, self-regulates her actions by taking into account the social norms of the group, both cooperative and communicative. Modern humans communicate with themselves and so reflect on and evaluate their own thinking with group-held normative standards. This reflection means that humans know what they are thinking and can provide to themselves normatively sanctioned justifications and reasons for thinking in this way—thus connecting up their many and diverse thoughts in an intricate inferential web that is governed, to some extent, by “community standards.” The thinking subject also uses this reflection in exercising executive control over her own thinking and reasoning. Korsgaard (2009), in particular, has emphasized that humans not only have goals and make decisions and reason in particular ways but also attempt to assess ahead of time whether those are good goals to pursue or good decisions to make or good reasons to have—a clearly extra layer of reflection and evaluation. And the normative judgment here is not simply mine alone, nor that of a specific other partner, but rather a judgment about whether that would be a good goal or decision or line of reasoning for any rational person, that is, for anyone from our group who does things the way that we do them.

Modern humans thus operate with the social norms of the group as internalized guides to both action and thinking. This means that in their collaborative interactions modern humans conform to the collectively accepted ways of doing things, based on norms of cooperation, and in their communicative interactions they conform to the collectively accepted ways of using language and also linguistically formulated arguments, based on the group’s norms of reason.

Objectivity: The View from Nowhere

Unlike other great apes, who all live in the general vicinity of the equator, modern humans have migrated all over the globe. They have done this not as individuals but as cultural groups; in none of their local habitats could a modern human individual survive for very long on his own. Instead, in each specific environment, modern human cultural groups have developed collectively a set of specialized and cognitively complex cultural practices to accommodate the local conditions, from seal hunting and igloo building to tuber gathering and bow-and-arrow making—not to mention science and mathematics. What we have attempted to do here is to specify the skills of cognition and thinking that enable modern human individuals to coordinate with those around them, both collaboratively and communicatively, in their efforts to adapt together to the novel exigencies they encounter in their specific corner of the world.

An image for the advent of modern humans is this: Earlier humans are living quite nicely by collaborating and communicating with others in various ways for a variety of cooperative purposes. Then, in the face of some serious demographic challenges, a great wave of group-mindedness and conformity washes over everyone. Humans who were spontaneously coordinating with partners to hunt or gather their daily meals then began to develop for foraging a number of conventionalized cultural practices. Humans who were spontaneously communicating with their partners in coordinating their complex collaborative activities using ad hoc gestures then began developing skills of conventional linguistic communication. And humans who were spontaneously exhorting or admonishing one another second-personally in various cooperative directions then began to develop collectively known and applied social norms of morality and rationality. Early humans lived together and interacted with others jointly; modern humans lived together and interacted with others collectively.

One effect of this great wave of group-mindedness and conformity was cultural group selection accompanied by cumulative cultural evolution. Cultural group selection takes place when individuals conform within their group—and differentiate themselves from other groups—to the extent that the group itself becomes a unit of natural selection (Richerson and Boyd, 2006). In this way, successful cultural adjustments to local conditions stay around, and unsuccessful attempts die out. Cumulative cultural evolution takes place when the inventions in a cultural group are passed on with such fidelity that they remain stable in the group until a new and improved invention comes along (the so-called ratchet effect; Tomasello et al., 1993). Modern humans had a stronger ratchet than early humans and apes because they had—in addition to powerful skills of imitation—proclivities both to teach things to others and also to conform to others when they themselves were being taught. And so it is with this wave of group-mindedness and conformity that we get the possibility of cultural groups creating and constantly improving their own cognitive artifacts—from procedures for whale hunting to procedures for solving differential equations—that help them both to adapt to local conditions and to mark themselves as distinct from other cultural groups.

The subterranean effect of this wave of group-mindedness and conformity, as it were, was new and culturally collective forms of cognitive representation, inference, and self-monitoring for use in thinking. Modern humans began representing the world “objectively,” reflecting a kind of generic, agent-neutral perspective possible by any rational person. Further, humans’ new skills of conventional linguistic communication enabled them to talk about many things that they previously could not (e.g., mental states and logical operations), and this enabled reflective inferences—thinking about one’s own thinking—with much greater depth and breadth. In the context of cooperative argumentation, modern humans made explicit the reasons for their assertions, thus connecting them in an inferential web to their other knowledge, and then this social practice of reason-giving was internalized into fully reflective reason. And the self-monitoring of modern humans for the first time reflected not just their expectations about the second-personal evaluations of specific others but, rather, their expectations about the normative evaluations of “us” as a cultural group. Given all of these new ways of behaving and thinking, the crack in the human experiential egg now became a veritable chasm: the individual no longer contrasted her own perspective with that of a specific other—the view from here and there; rather, she contrasted her own perspective with some kind of generic perspective of anyone and everyone about things that were objectively real, true, and right from any perspective whatsoever—a perspectiveless view from nowhere.

And so, if from a moral point of view, cooperation always entails some kind of effacing of one’s own interests in deference to those of others or the group, then, from a cognitive point of view, cooperative thinking always entails some kind of effacing of one’s own perspective in deference to the more “objective” perspective of others or the group (Piaget, 1928). Thus, in cooperative communication I must always honor the perspective of my recipient, and in cooperative argumentation I must be committed to accept the reasons and arguments of others if they are better than my own—by the yardstick of our agreed upon normative criteria of rationality, which include our agreed upon objective reality—and so to abandon mine for theirs. In the words of Nagel (1986. p. 4): “Objectivity is a method of understanding … To acquire a more objective understanding of some aspect of life or the world, we step back from our initial view of it and form a new conception which has that view and its relation to the world as its object.… The process can be repeated, using a still more objective conception.” In this formulation, “objectivity” is the result of being able to think of things from ever wider perspectives and also recursively, as one embeds one’s perspective within another, more encompassing perspective. In the current view, more encompassing means simply from the perspective of an ever wider, more transpersonally constituted generic individual or social group—the view from anyone.

The monumental second step on the way to modern humans thus took the already cooperativized and perspectival thinking of early humans and collectivized and objectified it. Whereas early humans internalized and referenced the perspective of what Mead (1934) calls the “significant other”, modern humans internalized and referenced the perspective of the group as a whole, or any group member, Mead’s “generalized other.” Human thinking at this point is no longer a solely individual process, or even a second-personal social process; rather, it is an internalized dialogue between “what I do think” and “what anyone ought to think” (Sellars, 1963). Human thinking has now become collective, objective, reflective, and normative; that is to say, it has now become full-blown human reasoning.

If you find an error or have any questions, please email us at admin@erenow.org. Thank you!