Embodied and Oral · Representing meaning

Vocalisation, prosody and spoken language

Vocalisation, prosody and spoken language covers the human vocal-auditory system through which breath, voice, articulated sound, rhythm, pitch, timing and shared linguistic conventions are used to communicate meaning.

When it emerged
Deep prehistory; exact origin unknown
What changed
Carries compositional language and prosodic meaning through sound
Reading time
32 minutes
The essential questions

Vocalisation, prosody and spoken language, clearly explained

Vocalisation, prosody and spoken language covers the human vocal-auditory system through which breath, voice, articulated sound, rhythm, pitch, timing and shared linguistic conventions are used to communicate meaning.

What is it?

Vocalisation, prosody and spoken language form a biological and cultural communication system in which respiratory energy is converted into voice, shaped into patterned sound, organised by learned linguistic conventions and interpreted by listeners as emotional, social and symbolic meaning.

What problem did it solve?

Spoken language reduced the complexity constraint on immediate communication by making flexible symbolic description possible through rapidly produced, culturally shared sound.

How did it work?

The central historical breakthrough was not merely the ability to make sound. The breakthrough was the integration of voluntary vocal control, learned conventions, symbolic reference, combinatorial structure, shared attention, memory and rapid interpretation into a flexible communication system. Spoken language greatly reduced the complexity constraint on face-to-face communication.

What came before?

It built on Embodied non-verbal communication.

What did it make possible?

It helped make possible Conversation and dialogue, Oral tradition and storytelling, Song, music and chant, Machine Translation and Speech Recognition and Automated Transcription.

What survived?

Modern technologies repeatedly return to speech because typing and visual interfaces impose their own costs. Voice notes, podcasts, telephone calls, video meetings, smart speakers and conversational assistants are not departures from the topic. They are technological prostheses attached to one of humanity’s oldest communication systems.

Why does it still matter?

Marks the point at which information transmission became generative rather than merely reactive. The topic did not defeat distance, time or memory. It defeated a more fundamental limitation: the inability to package complex imagined relations into a signal another mind could reconstruct.

Deep dive

The deeper story

Vocalisation, prosody and spoken language covers the human vocal-auditory system through which breath, voice, articulated sound, rhythm, pitch, timing and shared linguistic conventions are used to communicate meaning.

The topic joins three related but non-identical phenomena:

  1. Vocalisation: audible signals produced by the vocal system, including cries, laughter, shouts and other non-lexical sounds.
  2. Prosody: variations in pitch, duration, intensity, rhythm, pausing and voice quality that communicate emotion, emphasis, grouping, stance and aspects of grammatical structure.
  3. Spoken language: a socially learned symbolic system in which conventional sound patterns are combined to express objects, actions, relationships, abstractions, possibilities and events beyond the immediate moment.

The central historical breakthrough was not merely the ability to make sound. Many animals do that. The breakthrough was the integration of voluntary vocal control, learned conventions, symbolic reference, combinatorial structure, shared attention, memory and rapid interpretation into a flexible communication system.

Spoken language greatly reduced the complexity constraint on face-to-face communication. It allowed humans to communicate things that could not be conveyed efficiently through posture, pointing, facial expression or undifferentiated calls alone. It made detailed coordination, explanation, negotiation, teaching and collective planning possible without requiring an external tool.

Its advantages were extraordinary:

  • negligible material cost;
  • immediate production;
  • rapid feedback;
  • high expressive flexibility;
  • hands-free use;
  • emotional and pragmatic richness;
  • adaptation to an enormous range of social situations.

Its weaknesses were equally consequential:

  • ordinary speech is transient;
  • unaided range is short;
  • accurate repetition depends on attention and memory;
  • meaning depends on shared language and context;
  • the signal is vulnerable to noise, hearing differences and ambiguity;
  • speakers can lie, deny, manipulate or exploit the absence of a durable record.

Those limitations created opportunities for later topics. Conversation formalised rapid feedback. Oral tradition developed repetition and social memory. Writing externalised language into durable marks. Recording preserved actual voices. Telephony, radio and digital networks extended speech across distance. Speech recognition, synthesis and conversational AI later made spoken language machine-readable, reproducible and transformable.

The big idea

| Dimension | Finding | |---|---| | Primary category | Encoding and expression | | Primary breakthrough | Flexible symbolic meaning encoded in controlled vocal sound | | Main problem addressed | Complexity of immediate communication | | Emergence | Deep prehistory; no defensible single invention date | | Origin confidence | Low | | Mechanism confidence | High | | Civilisational significance | Foundational | | Unaided persistence | Very low | | Unaided reach | Low | | Interactivity | Very high | | Material cost | Very low | | Principal descendants | Conversation, oral tradition, writing, recorded speech, telephony, broadcasting, voice interfaces | | Present status | Active, ubiquitous and foundational |

Main problem addressed

Carries compositional language and prosodic meaning through sound

Connections

What came before and what followed

Start with the key connections, then reveal the wider network when you need more context.

Extended or built upon
Telephone

Transmits interactive voice beyond ordinary acoustic range.

Specialised form
Song, music and chant

Strengthens rhythm, memorability, emotional synchronisation and ritual force.

Timeline

Key moments

How Vocalisation, prosody and spoken language emerged

This marks the broad emergence and development of Vocalisation, prosody and spoken language. Why it mattered: Carries compositional language and prosodic meaning through sound.

People and organisations

Who helped shape it?

No individually named contributors are listed for this topic yet. That does not mean it developed without human involvement.

Research notes

Open the full research notes

These expandable sections preserve the detailed research behind the public explanation.

1. Executive Summary

Vocalisation, prosody and spoken language covers the human vocal-auditory system through which breath, voice, articulated sound, rhythm, pitch, timing and shared linguistic conventions are used to communicate meaning.

The topic joins three related but non-identical phenomena:

  1. Vocalisation: audible signals produced by the vocal system, including cries, laughter, shouts and other non-lexical sounds.
  2. Prosody: variations in pitch, duration, intensity, rhythm, pausing and voice quality that communicate emotion, emphasis, grouping, stance and aspects of grammatical structure.
  3. Spoken language: a socially learned symbolic system in which conventional sound patterns are combined to express objects, actions, relationships, abstractions, possibilities and events beyond the immediate moment.

The central historical breakthrough was not merely the ability to make sound. Many animals do that. The breakthrough was the integration of voluntary vocal control, learned conventions, symbolic reference, combinatorial structure, shared attention, memory and rapid interpretation into a flexible communication system.

Spoken language greatly reduced the complexity constraint on face-to-face communication. It allowed humans to communicate things that could not be conveyed efficiently through posture, pointing, facial expression or undifferentiated calls alone. It made detailed coordination, explanation, negotiation, teaching and collective planning possible without requiring an external tool.

Its advantages were extraordinary:

  • negligible material cost;
  • immediate production;
  • rapid feedback;
  • high expressive flexibility;
  • hands-free use;
  • emotional and pragmatic richness;
  • adaptation to an enormous range of social situations.

Its weaknesses were equally consequential:

  • ordinary speech is transient;
  • unaided range is short;
  • accurate repetition depends on attention and memory;
  • meaning depends on shared language and context;
  • the signal is vulnerable to noise, hearing differences and ambiguity;
  • speakers can lie, deny, manipulate or exploit the absence of a durable record.

Those limitations created opportunities for later topics. Conversation formalised rapid feedback. Oral tradition developed repetition and social memory. Writing externalised language into durable marks. Recording preserved actual voices. Telephony, radio and digital networks extended speech across distance. Speech recognition, synthesis and conversational AI later made spoken language machine-readable, reproducible and transformable.

The big idea

| Dimension | Finding | |---|---| | Primary category | Encoding and expression | | Primary breakthrough | Flexible symbolic meaning encoded in controlled vocal sound | | Main problem addressed | Complexity of immediate communication | | Emergence | Deep prehistory; no defensible single invention date | | Origin confidence | Low | | Mechanism confidence | High | | Civilisational significance | Foundational | | Unaided persistence | Very low | | Unaided reach | Low | | Interactivity | Very high | | Material cost | Very low | | Principal descendants | Conversation, oral tradition, writing, recorded speech, telephony, broadcasting, voice interfaces | | Present status | Active, ubiquitous and foundational |

2. Identification

| Field | Value | |---|---| | Subject | Vocalisation, prosody and spoken language | | Register name | Vocalisation, prosody and spoken language | | Topic type | Method | | Primary category | Encoding and expression | | Secondary categories | Transmission; interpretation; feedback; distribution | | Era | Era I: Embodied and Oral Communication | | Suggested confidence label | Mixed: established mechanism, uncertain origin |

2.1 Recommended topic boundary

This research notes treats Vocalisation, prosody and spoken language as an umbrella topic for the evolution and operation of the vocal-auditory channel. It does not absorb every activity performed through speech.

The following remain separate topics:

  • Conversation and dialogue Conversation and dialogue, which concerns reciprocal exchange and feedback;
  • Oral tradition and storytelling Oral tradition and storytelling, which concerns intergenerational distribution and collective memory;
  • Song, music and chant Song, music and chant, which concerns heightened rhythmic and melodic encoding;
  • Neural and cognitive memory Biological memory, which preserves speech after the sound itself disappears;
  • Imitation, repetition and apprenticeship Imitation, repetition and apprenticeship, which reproduces behaviour and knowledge;
  • Long-distance acoustic and visual signals Long-distance acoustic and visual signals, which prioritise range over semantic complexity.

2.2 Important terminology correction

Voice, speech and language are not synonyms.

  • Voice is sound generated by airflow and vibration of the vocal folds.
  • Speech is the coordinated shaping of voice into recognisable sounds.
  • Language is a structured system for communicating and representing meaning.[1][2]

Language is also not inherently vocal. Natural sign languages demonstrate that complex language can operate through the visual-manual channel. Vocalisation, prosody and spoken language therefore concerns spoken language, not the total human language faculty.

2.3 Possible future split

A later register may split this umbrella into:

  • ENC-002A Non-lexical vocalisation and affective voice;
  • ENC-002B Prosodic vocal communication;
  • ENC-002C Spoken linguistic encoding.

For v0.1, retaining one umbrella topic preserves a readable master narrative. The internal distinctions should nevertheless remain visible in research and graphics.

3. Operational Definition

Vocalisation, prosody and spoken language form a biological and cultural communication system in which respiratory energy is converted into voice, shaped into patterned sound, organised by learned linguistic conventions and interpreted by listeners as emotional, social and symbolic meaning.

A minimal spoken-language event requires:

  1. a speaker with an intention or internal state;
  2. a shared or partly shared code;
  3. motor planning and control;
  4. airflow and phonation;
  5. articulation by the vocal tract;
  6. an acoustic channel;
  7. a listener capable of hearing or otherwise receiving the signal;
  8. perceptual decoding;
  9. linguistic and contextual interpretation.

Spoken language can express:

  • immediate observations;
  • requests and commands;
  • names and categories;
  • spatial and temporal relations;
  • causes and consequences;
  • hypothetical situations;
  • remembered events;
  • plans;
  • rules;
  • social commitments;
  • abstractions;
  • fiction and deception.

This expressive range distinguishes spoken language from a fixed inventory of calls tied narrowly to immediate states.

4. Communication Pattern

| Attribute | Classification | |---|---| | Participant structure | One-to-one, one-to-many, many-to-many | | Direction | Unidirectional or bidirectional | | Time relationship | Primarily synchronous; asynchronous only through memory, intermediaries or recording | | Spatial relationship | Co-located or within acoustic range unless technologically extended | | Persistence | Ephemeral in its unaided form | | Feedback | Potentially immediate | | Primary senses | Hearing, supplemented by vision and touch | | Production mode | Biological motor action guided by culturally learned conventions | | Typical authentication | Voice, face, context, social familiarity and turn-taking behaviour | | Typical noise | Physical noise, poor articulation, accent mismatch, hearing loss, ambiguity, distraction and hostile interpretation |

4.1 Multimodal reality

Ordinary spoken communication is rarely voice alone. Face-to-face language normally combines words with gaze, facial expression, mouth movement, gesture, posture and shared environmental context. Research on language processing and evolution increasingly treats human communication as multimodal rather than as an isolated stream of disembodied words.[3][4][5]

The topic is therefore vocal-centred but not voice-exclusive.

5. Expanded Communication Model
Internal state or communicative intention ↓
Conceptual selection ↓
Lexical and grammatical encoding ↓
Motor planning ↓
Respiration → phonation → articulation ↓
Acoustic speech signal ↓
Air and surrounding acoustic environment ↓
Auditory perception ↓
Segmentation and linguistic decoding ↓
Semantic and pragmatic interpretation ↓
Belief, emotion, decision or action ↓
Optional verbal or non-verbal feedback

5.1 Physical production

Speech uses several coordinated bodily systems. Airflow from the lungs powers vibration at the vocal folds. The resulting voice is shaped by coordinated movement in the larynx, mouth and nasal tract, particularly the tongue and lips.[2]

5.2 Segmental encoding

Languages divide and organise the speech stream using learned sound contrasts, syllables, morphemes, words and larger constructions. The exact units and structures vary across languages.

5.3 Prosodic encoding

Prosody is carried through combinations of:

  • fundamental frequency and pitch movement;
  • duration;
  • intensity;
  • rhythm;
  • pausing;
  • speech rate;
  • voice quality.

These cues can signal emphasis, grouping, emotion, attitude, lexical contrasts, grammatical boundaries and pragmatic intent.[6][7]

The phrase “I never said she stole the money” can change implication depending on which word is stressed. The words remain the same; the accusation performs a small costume change.

5.4 Contextual decoding

Listeners do not reconstruct meaning from sound alone. They combine the signal with:

  • knowledge of the language;
  • expectations about the speaker;
  • the physical situation;
  • previous turns;
  • cultural conventions;
  • inferred motives;
  • tone and gesture.

This makes spoken language powerful, efficient and dangerously capable of being misunderstood.

6. Historical Emergence

6.1 Dating conclusion

There is no defensible discovery date for speech or spoken language.

Language does not fossilise. Researchers must infer possible capacities from indirect evidence such as:

  • comparative animal communication;
  • fossil anatomy;
  • hearing and vocal-tract reconstructions;
  • genetics;
  • brain organisation;
  • tool production;
  • symbolic artefacts;
  • patterns of social learning;
  • present-day language development and variation.

These sources constrain hypotheses but do not reveal the first sentence, first grammar or first fully linguistic community. Major reviews continue to emphasise that the origin and evolutionary pathway of language remain deeply uncertain.[8]

6.2 Evidence streams

A. Comparative vocal communication

Vocal signalling is evolutionarily ancient and widespread. Humans retained older affective vocal behaviours such as crying, laughter, screaming and non-verbal calls while developing much greater voluntary control and learned vocal flexibility.[9]

Comparative work with primates, songbirds, bats and other species helps isolate possible component capacities. It does not demonstrate that another species possesses human language.

B. Vocal anatomy

The lungs, larynx, vocal folds, tongue, lips, mouth and nasal tract form a highly coordinated production system. Fossil anatomy can offer clues about possible speech capacities, but soft tissues do not preserve and reconstructed vocal tracts depend on disputed assumptions.

Earlier claims that Neanderthal anatomy prevented a broad vowel range have been substantially challenged. Current evidence is compatible with significant articulatory ability, but anatomy alone cannot establish grammar, vocabulary or actual language use.[10][11]

C. Neural control

Speech depends on distributed brain systems supporting auditory processing, motor sequencing, memory, conceptual representation and social cognition. Evolution may have depended as much on changes in neural control and learning as on gross changes in vocal anatomy.[12]

D. Genetics

Mutations affecting FOXP2 can cause severe speech-motor and wider language difficulties. The gene influences networks involved in neural development and sensory-guided motor learning. It is not a single “language gene,” nor does its presence prove that an extinct population spoke a modern language.[13][14]

The human capacity is polygenic and depends on development, anatomy, learning and culture.

E. Archaeology and technology

Stone tools preserve; speech does not. Toolmaking has therefore been used as an indirect window into teaching, sequencing and social learning.

In a modern transmission experiment, verbal teaching improved the reproduction of Oldowan-style toolmaking skills more than imitation or emulation alone. This supports the usefulness of language for high-fidelity instruction. It does not prove that Oldowan toolmakers possessed spoken language.[15]

F. Symbolic behaviour

Pigments, ornaments, burials, engravings and other symbolic traces may be compatible with complex language. They cannot by themselves reveal the modality, grammar or origin date of the communication system behind them.

6.3 Provisional evolutionary timeline

| Stage | Provisional description | Confidence | |---|---|---| | Pre-human vocal signalling | Ancestral primates used calls and affective vocal signals before the human lineage. | Established at broad level | | Increased voluntary and learned vocal control | Hominin evolution gradually expanded control, sequencing and flexibility. Exact timing unknown. | Broadly plausible | | Proto-linguistic communication | Conventional vocal symbols may have accumulated before modern grammar. This is a theoretical reconstruction, not an observed stage. | Speculative | | Structured spoken languages | Fully expressive languages existed before writing and almost certainly before surviving historical records. Exact antiquity unresolved. | Established in broad conclusion; unknown in date | | Oral-dominant societies | Spoken language, memory and performance carried most everyday knowledge and governance. | Established | | Speech alongside writing | Writing preserved selected linguistic content but did not replace speech. | Established | | Mechanically recorded and transmitted speech | Recording, telephony, radio, film and television separated voice from immediate co-presence. | Established | | Machine-readable and synthetic speech | Digital systems transcribe, translate, synthesise and generate spoken language. | Established |

6.4 Neanderthal question

The evidence permits at least three responsible positions:

  1. Neanderthals may have had substantial speech and language capacities.
  2. They may have possessed a system less flexible than that of later Homo sapiens.
  3. The surviving evidence may be incapable of deciding between these possibilities.

The research notes adopts the third position as its publication baseline. Claims that Neanderthals definitely spoke a modern language, or definitely could not speak, exceed the evidence.

7. Prerequisites

Spoken language was not one invention sitting alone on a rock waiting for somebody to patent it. It required a stack of biological, cognitive, social and cultural capacities.

7.1 Biological prerequisites

  • controlled breathing;
  • phonation;
  • flexible articulation;
  • auditory perception;
  • neural coordination of rapid motor sequences;
  • capacity for vocal learning and adjustment;
  • developmental plasticity.

7.2 Cognitive prerequisites

  • attention;
  • working and long-term memory;
  • categorisation;
  • sequence processing;
  • symbolic association;
  • inference;
  • perspective-taking;
  • planning;
  • ability to learn patterns from social input.

7.3 Social prerequisites

  • repeated interaction;
  • shared attention;
  • motivation to inform, request, persuade or coordinate;
  • tolerance of turn-taking;
  • conventions linking forms to meanings;
  • communities capable of maintaining those conventions;
  • children exposed to usable language during development.[1]

7.4 Cultural prerequisites

A spoken language exists because a community repeatedly reproduces it. No individual designs the whole system. Vocabulary, pronunciation, constructions and pragmatic norms emerge and change through many acts of learning and use.

This means the method is simultaneously:

  • biological in capacity;
  • cultural in form;
  • social in maintenance;
  • individual in performance.
8. Periodisation

8.1 Phase I: Affective vocal signalling

Calls, cries, laughter and tone communicated urgency, distress, affiliation and emotional state.

8.2 Phase II: Conventional vocal reference

Communities increasingly linked repeatable sounds to repeatable meanings. The date and sequence are unknown.

8.3 Phase III: Combinatorial spoken language

Conventional units could be combined to express relations, events and propositions beyond isolated signals.

8.4 Phase IV: Oral social systems

Speech became embedded in teaching, law, ritual, trade, negotiation, identity and collective memory.

8.5 Phase V: Coexistence with external records

Writing preserved some linguistic content while speech remained dominant in ordinary interaction.

8.6 Phase VI: Technological extension

Recording and telecommunications separated speech from immediate place and time.

8.7 Phase VII: Computational mediation

Machines began converting speech to text, text to speech, one spoken language to another and prompts into synthetic voices.

9. Primary Problem Solved

9.1 Before complex spoken language

Embodied signals could efficiently communicate:

  • emotion;
  • attention;
  • threat;
  • direction;
  • approval or rejection;
  • simple immediate intentions.

They were weaker at expressing:

  • absent objects;
  • detailed sequences;
  • causal explanations;
  • social rules;
  • counterfactuals;
  • remote events;
  • plans involving many steps;
  • abstract categories.

9.2 Constraint reduced

Spoken language reduced the complexity constraint on immediate communication by making flexible symbolic description possible through rapidly produced, culturally shared sound.

It also reduced:

  • the need to demonstrate every action physically;
  • dependence on visible line-of-sight gestures;
  • the cost of correcting misunderstandings;
  • the difficulty of coordinating unseen or future actions;
  • the burden of inventing a new signal for every message.

9.3 Constraints not solved

Speech did not solve:

  • persistence;
  • long-distance transmission;
  • cross-language understanding;
  • reliable large-scale copying;
  • searchability;
  • objective authentication;
  • permanent accountability.

Those unfinished jobs became invitations to later media.

10. Evaluation Matrix

The ratings below apply to unaided, live spoken language. Recording, telephony and broadcasting are evaluated in their own topics.

| Dimension | Rating | Rationale | |---|---|---| | Reach | Low | Normally limited by acoustic range and environment. | | Latency | Very high performance | Production and reception are effectively immediate. | | Bandwidth | Moderate | Carries complex symbolic and emotional information, but sequentially. A cross-linguistic study estimated speech information rates clustering around roughly 39 bits per second; this is an empirical estimate, not a timeless physiological constant.[16] | | Fidelity | High within a clear single exchange; declining across chains | Immediate clarification helps, but memory and repetition introduce distortion. | | Persistence | Very low | The acoustic signal disappears almost immediately. | | Replication cost | Very low | Repetition requires time and human effort but little material infrastructure. | | Distribution cost | Very low locally; high at scale | Addressing nearby listeners is cheap. Reaching large or distant audiences requires organisations or technology. | | Accessibility | High within a shared language community; uneven overall | Requires language knowledge, auditory access or accommodation, and sufficient cognitive development. | | Portability | Very high | The production system travels with the speaker. | | Interactivity | Very high | Questions, interruptions, correction and negotiation can occur immediately. | | Searchability | None | Unrecorded speech cannot be searched after it disappears. | | Editability | High before or during delivery; none after disappearance | Speakers can reformulate, but cannot alter what listeners already heard. | | Scalability | Low unaided | Audience size is constrained by volume, acoustics and attention. | | Authentication | Moderate | Familiar voices and co-presence help, but impersonation, hearsay and misattribution remain possible. | | Privacy | Low to moderate | Anyone within earshot may receive the message. Whispering improves privacy but reduces range. | | Censorship resistance | Moderate in private local use | Hard to eliminate completely, but speakers can be silenced, punished or excluded. | | Infrastructure dependence | Very low | Requires bodies, air, a language community and a usable acoustic setting. | | Energy dependence | Low | Powered biologically, though sustained speaking imposes physical effort. | | Interpretive burden | Moderate to high | Shared vocabulary, grammar, culture, context and pragmatic inference are required. | | Attention demand | Moderate | Listeners can receive speech without looking, but comprehension declines with divided attention and noise. |

11. Technical and Social Advantages

11.1 Flexible compositionality

A finite repertoire of learned distinctions can be recombined to generate an open-ended range of messages. Speakers do not need a unique cry for every tax rule, hunting route, family dispute or bad guitar solo.

11.2 Speed

Speech can be planned, produced and corrected during the same interaction. It is much faster than carving, painting or handwriting an equivalent ordinary exchange.

11.3 Low material cost

No external surface, ink, device, electricity or network is required.

11.4 Hands-free operation

Speech can accompany walking, carrying, tool use, cooking, driving, fighting, childcare and shared work.

11.5 Communication without visual contact

Unlike gesture, speech can operate in darkness, around modest obstructions and while participants look at a shared task rather than at one another.

11.6 Emotional richness

Prosody carries information beyond lexical content. Pitch, loudness, rate, timing and voice quality can signal urgency, confidence, affection, irony, hostility, fatigue and uncertainty.[6][7]

11.7 Immediate repair

Listeners can ask:

  • “What do you mean?”
  • “Which one?”
  • “Did you say fifteen or fifty?”

This repair loop makes spoken interaction robust despite ambiguity.

11.8 Personal identity

Voices carry information about individual identity, age, emotional state, community membership and social presentation. These cues support recognition while also enabling stereotyping.

11.9 Adaptability

Speakers can alter vocabulary, speed, volume, register and complexity for:

  • children;
  • experts;
  • strangers;
  • intimate partners;
  • public audiences;
  • emergencies;
  • ceremonial settings.

11.10 Support for high-fidelity teaching

Verbal instruction can reveal invisible decisions, sequences and causal reasoning that observation alone may not expose. Experimental work on stone-tool teaching shows how language can improve transmission for complex tasks.[15]

12. Contribution to Human Advancement

The claims below describe capabilities that spoken language enabled or strengthened. They do not imply that speech acted alone.

12.1 Survival

Spoken language improved the ability to communicate:

  • danger;
  • food locations;
  • animal behaviour;
  • weather observations;
  • injuries;
  • medicinal practices;
  • safe and unsafe routes;
  • plans for hunting or defence.

A shout can warn. A sentence can explain why, identify where, specify when and tell someone what to do next.

12.2 Cooperation

Speech supports:

  • task allocation;
  • collective planning;
  • negotiation;
  • promises;
  • correction;
  • instruction;
  • conflict mediation;
  • shared goals.

Human linguistic communication both supports cooperation and operates within social relationships that mix cooperation, competition and strategic ambiguity.[17]

12.3 Family and social organisation

Spoken language enables:

  • names and kinship categories;
  • courtship;
  • caregiving instruction;
  • moral correction;
  • emotional reassurance;
  • reputation sharing;
  • boundary negotiation;
  • group identity.

12.4 Governance

Before written administration, rules and decisions could be:

  • announced;
  • debated;
  • witnessed;
  • remembered;
  • repeated;
  • interpreted by recognised authorities.

Speech enabled councils, testimony, adjudication and public persuasion. Writing later increased persistence and administrative reach.

12.5 Trade and exchange

Spoken language allowed parties to negotiate:

  • quantities;
  • quality;
  • ownership;
  • credit;
  • obligations;
  • delivery;
  • trust;
  • remedies for breach.

Its weakness was that agreements could later be disputed, encouraging witnesses, tokens and written contracts.

12.6 Science, engineering and technical skill

Speech enables instructors to communicate:

  • sequences;
  • causal models;
  • warnings;
  • exceptions;
  • diagnostic reasoning;
  • corrections during practice.

It allows learners to ask questions about invisible intentions rather than merely copy visible movements. This likely strengthened cumulative technical culture, although the precise evolutionary relationship remains debated.[15][18]

12.7 Education

Spoken explanation remains central to:

  • demonstration;
  • questioning;
  • tutoring;
  • recitation;
  • discussion;
  • apprenticeship;
  • lectures;
  • feedback.

Writing expanded storage. Speech retained responsiveness.

12.8 Religion and philosophy

Speech enabled ritual instruction, prayer, preaching, dialogue, argument and interpretation. Repetition, rhythm and prosody strengthened authority and memorability, often converging with Song, music and chant Song, music and chant.

12.9 Warfare

Speech allowed immediate command, warning, morale-building, interrogation and tactical coordination. Its acoustic range, detectability and ambiguity limited performance, helping drive visual signals, drums, messengers, radio and secure communications.

12.10 Culture and entertainment

Spoken language enabled jokes, stories, dramatic performance, verbal games, praise, insult and commentary. Oral tradition and song later developed specialised ways to preserve and amplify these forms.

12.11 Collective memory

Speech can transmit memories to other people, but unstructured speech is a poor archive. Collective preservation required repetition, specialists, formulaic structures, ritual and social correction. Those functions belong primarily to Oral tradition and storytelling Oral tradition and storytelling.

13. Organisations, Roles and Power

Spoken language is biologically available but socially organised.

13.1 Roles created or strengthened

  • elder;
  • teacher;
  • storyteller;
  • herald;
  • negotiator;
  • interpreter;
  • priest or ritual specialist;
  • advocate;
  • judge;
  • diplomat;
  • commander;
  • public speaker;
  • performer.

These are not all prehistoric organisations. They illustrate how control of speech, memory and interpretation can become specialised.

13.2 Authority through speech

Power may derive from:

  • recognised knowledge;
  • persuasive skill;
  • fluency in a prestige language;
  • command of ritual forms;
  • control of who may speak;
  • ability to define acceptable vocabulary;
  • access to assemblies;
  • social status attached to accent or register.

13.3 Inclusion and exclusion

Spoken language can include those who share the code and exclude those who do not. Barriers may arise from:

  • different languages;
  • specialised jargon;
  • hearing impairment;
  • speech impairment;
  • accent prejudice;
  • gender or status rules about public speaking;
  • restricted access to interpreters;
  • punishment for forbidden speech.

13.4 Language prestige

Communities often rank ways of speaking. A dialect may be wrongly treated as defective because it lacks institutional prestige, while the dominant variety is dressed up as neutral correctness. This affects education, employment, justice and public credibility.

13.5 Control of the floor

Speech is scarce in time. Only a limited number of people can be understood at once. Power therefore includes the ability to:

  • interrupt;
  • dominate turn-taking;
  • set the agenda;
  • refuse questions;
  • declare a discussion closed;
  • decide whose testimony counts.

This becomes more explicit in Conversation and dialogue Conversation and dialogue and GOV topics.

14. Limitations

14.1 Ephemerality

The acoustic signal vanishes. Without memory, repetition or recording, the message leaves no durable object.

14.2 Limited range

Ordinary voices do not travel far. Wind, walls, terrain, crowds and competing sounds reduce intelligibility.

14.3 Sequential delivery

Speech unfolds over time. A listener cannot instantly scan the whole message as they might scan a page, table or diagram.

14.4 Memory dependence

Long or complex speech taxes working memory. Details, numbers and unfamiliar names are easily lost.

14.5 Fidelity decay

Repeated retelling can alter:

  • wording;
  • order;
  • emphasis;
  • causal links;
  • certainty;
  • identity of participants.

Oral traditions develop mechanisms to resist this decay, but ordinary hearsay does not arrive with checksum validation.

14.6 Language dependence

A signal is useful only when participants share enough vocabulary, grammar and context.

14.7 Ambiguity

The same sentence may support multiple interpretations. Prosody can clarify ambiguity, but can also create it.

14.8 Environmental vulnerability

Speech performance declines under:

  • noise;
  • poor acoustics;
  • distance;
  • masks or obstructions;
  • fatigue;
  • illness;
  • damaged hearing;
  • divided attention.

14.9 Weak accountability

Unrecorded speech enables denial:

  • “I never said that.”
  • “You misunderstood.”
  • “That is not what I meant.”

Witnesses and reputation partially compensate. Writing and recording provide stronger persistence but create new surveillance and privacy risks.

14.10 Limited large-scale distribution

One unaided speaker cannot efficiently address a continent. Architecture, heralds, repeated messengers, print and broadcast technologies were required to scale the audience.

14.11 Dependence on human presence or relay

A speaker cannot address future generations directly without memory, representation or recording.

15. Harms and Trade-Offs

15.1 Deception

Flexible symbolic communication makes it possible to describe things that are absent, hypothetical or false. The same capacity that enables planning also enables fabrication.

15.2 Rumour

Low-cost repetition and weak source tracking allow claims to move through communities faster than verification.

15.3 Manipulation

Prosody, charisma, repetition and emotional framing can influence recipients independently of evidential quality.

15.4 Coercion and intimidation

Voice can threaten, humiliate, command and dominate. Volume and interruption can become crude instruments of hierarchy.

15.5 Misunderstanding

Words can be interpreted through different cultural assumptions, personal histories or emotional states.

15.6 Accent and language discrimination

Listeners may infer competence, morality, class or intelligence from speech patterns that do not justify those conclusions.

15.7 Information overload

Speech can flood attention, particularly when listeners cannot pause, skim or retrieve earlier passages.

15.8 Social surveillance

Because speech can be overheard, private communication may expose speakers to punishment or reputational harm.

15.9 Loss through silence

Knowledge held only in speech disappears when speakers die, forget or are prevented from transmitting it.

15.10 Standardisation pressure

Prestige languages may displace minority languages, weakening local identity and access to cultural knowledge. This is principally a governance and cultural-distribution issue, but it travels through the spoken channel.

16. Predecessors and Prerequisites

| Topic | Relationship | Contribution to Vocalisation, prosody and spoken language | |---|---|---| | Embodied non-verbal communication Embodied non-verbal communication | Predecessor and complement | Provided gaze, gesture, expression, posture, shared attention and immediate social cues. | | Neural and cognitive memory Biological memory | Prerequisite | Retained sound patterns, meanings, people, events and conversational context. | | Pre-human biological vocal signalling | Evolutionary substrate | Supplied affective calls and vocal production systems. | | Shared cognition and social learning | Enabling dependency | Allowed communities to stabilise conventions and children to acquire them. | | Auditory perception | Biological prerequisite | Enabled discrimination and learning of speech patterns. | | Voluntary motor control | Biological prerequisite | Enabled flexible sequencing and modification of vocal output. |

17. Successors, Descendants and Extensions

| Topic | Relationship | Advantage gained | |---|---|---| | Conversation and dialogue Conversation and dialogue | Specialisation | Organises speech into reciprocal exchange, repair and negotiation. | | Oral tradition and storytelling Oral tradition and storytelling | Extension | Improves intergenerational persistence through repetition, structure and social custodianship. | | Imitation, repetition and apprenticeship Imitation, repetition and apprenticeship | Complement | Combines verbal explanation with demonstration and practice. | | Song, music and chant Song, music and chant | Specialisation and convergence | Strengthens rhythm, memorability, emotional synchronisation and ritual force. | | Long-distance acoustic and visual signals Long-distance acoustic signals | Specialisation | Trades semantic complexity for range and urgency. | | Writing systems Writing systems | Remediation and extension | Encodes linguistic information in durable visual form. | | Sound recording | Persistence extension | Preserves actual vocal performance beyond the event. | | Telephone | Distance extension | Transmits interactive voice beyond ordinary acoustic range. | | Radio | Scale extension | Distributes speech simultaneously to mass audiences. | | Film and television | Multimodal persistence and distribution | Reunites voice with visible performance across time and distance. | | Machine Translation Machine translation | Interpretive extension | Converts between languages. | | Speech Recognition and Automated Transcription Speech recognition | Processing extension | Converts speech into searchable and editable text. | | Text-to-Speech and Voice Synthesis Text-to-speech | Synthetic reproduction | Generates spoken output from text. | | Conversational AI Assistants and Retrieval-Augmented Generation Conversational AI | Machine mediation | Produces and interprets spoken or written dialogue computationally. |

17.1 Successor logic

Writing did not replace speech. Recording did not replace live conversation. Messaging did not kill the telephone. Spoken language survives because later systems solve only particular limitations while often sacrificing immediacy, embodiment, privacy or emotional detail.

18. What Remained Valuable

Spoken language remains valuable because it is:

  • fast;
  • embodied;
  • cheap;
  • interactive;
  • portable;
  • emotionally expressive;
  • adaptable to audience response;
  • usable without literacy;
  • usable without electricity;
  • capable of accompanying physical work;
  • effective for trust-building and relationship maintenance.

Modern technologies repeatedly return to speech because typing and visual interfaces impose their own costs. Voice notes, podcasts, telephone calls, video meetings, smart speakers and conversational assistants are not departures from the topic. They are technological prostheses attached to one of humanity’s oldest communication systems.

19. Representative Scenarios and Historical Moments

19.1 Evidence limitation

The first spoken-language events are unrecoverable. Any vivid reconstruction of “the first conversation” would be fiction wearing an archaeological hat.

This research notes therefore separates illustrative scenarios from directly documented historical events.

19.2 Illustrative scenario: teaching a hidden decision

A learner can watch someone strike a stone. Speech can additionally explain:

  • where to strike;
  • why that angle matters;
  • what flaw to avoid;
  • which result is intended;
  • how to correct an error.

Modern experiments show that verbal teaching can improve transmission of stone-tool production skills, but they do not establish the language of ancient toolmakers.[15]

19.3 Illustrative scenario: coordinating a future action

A group can agree to meet at a location after sunset, assign roles and discuss contingencies. Gesture can point to the hill. Spoken language can describe the plan that does not yet exist.

19.4 Illustrative scenario: repairing misunderstanding

A recipient misidentifies a plant. The speaker can contrast two similar plants, describe symptoms and ask the listener to repeat the distinction.

19.5 Historically visible consequence: public deliberation

Once written records appear, they reveal societies in which spoken testimony, debate, proclamation and teaching remain central. The surviving text is evidence of a speech-based institution, but not necessarily a verbatim transcript.

19.6 Historically visible consequence: preserved oratory

Recording technologies eventually allow later generations to hear cadence, timing, pronunciation and emotional force rather than merely read transcribed words. That transition belongs primarily to recorded sound and broadcasting, but it reveals how much information writing omits from speech.

20. Comparison with Adjacent Era I Topics

| Topic | What it primarily contributes | What distinguishes it from Vocalisation, prosody and spoken language | |---|---|---| | Embodied non-verbal communication Embodied non-verbal communication | Visible affect, attention, gesture and immediate reference | Does not require conventional vocal symbols or acoustic sequencing. | | Conversation and dialogue Conversation and dialogue | Feedback, turn-taking, repair and negotiation | Describes an interaction pattern rather than the encoding channel itself. | | Neural and cognitive memory Biological memory | Retention | Stores what speech no longer physically contains. | | Imitation, repetition and apprenticeship Imitation, repetition and apprenticeship | Reproduction of behaviour and skill | Can operate without language; speech improves explanation. | | Oral tradition and storytelling Oral tradition and storytelling | Social persistence and intergenerational distribution | Adds structured repetition, custodianship and cultural memory. | | Song, music and chant Song, music and chant | Rhythm, melody, affect and synchronisation | Can communicate without lexical language and often prioritises musical pattern. | | Long-distance acoustic and visual signals Long-distance acoustic and visual signals | Range and urgency | Uses a smaller signal inventory with less semantic flexibility. |

21. Claim Register

| Claim ID | Claim | Confidence | Publication note | |---|---|---|---| | ENC002-C01 | Voice, speech and language are related but distinct phenomena. | Established | Safe with definitions. | | ENC002-C02 | Spoken communication is normally multimodal, combining voice with gesture, gaze and facial cues. | Established | Avoid describing speech as an isolated audio stream. | | ENC002-C03 | Prosody uses pitch, duration, intensity, rhythm and voice quality to contribute meaning and structure. | Established | Safe. | | ENC002-C04 | The first emergence of spoken language cannot currently be dated precisely. | Established | Central uncertainty statement. | | ENC002-C05 | Language does not fossilise directly. | Established | Safe, but avoid implying archaeology is irrelevant. | | ENC002-C06 | Fossil anatomy alone cannot prove the presence of modern language. | Established | Safe. | | ENC002-C07 | Neanderthal speech and language capacity remains unresolved. | Broadly accepted as unresolved | Present multiple credible positions. | | ENC002-C08 | FOXP2 contributes to speech-motor and language-related development but is not a single language gene. | Established | Avoid “speech gene” shorthand except to correct it. | | ENC002-C09 | Verbal teaching can improve transmission of complex toolmaking skills in modern experiments. | Established | Do not infer ancient language directly. | | ENC002-C10 | Spoken language reduces the complexity constraint on immediate communication. | Analytical interpretation | Project thesis, not a direct source claim. | | ENC002-C11 | Unaided speech has low persistence and limited range but very high interactivity. | Established by mechanism | Safe. | | ENC002-C12 | Cross-linguistic speech information rates may cluster around roughly 39 bits per second. | Supported within one study design | Present as an estimate, not a universal constant. | | ENC002-C13 | Spoken language strengthened cumulative teaching and cooperation. | Broadly supported | Avoid monocausal claims. | | ENC002-C14 | Writing extends spoken language by adding persistence but does not replace it. | Established | Safe. | | ENC002-C15 | Speech technologies repeatedly remediate the vocal channel by extending distance, scale, persistence or machine interpretation. | Analytical synthesis | Safe when linked to successor topics. |

22. Open Questions

22.1 Origin and chronology

  • When did conventional vocal symbols first become productive language?
  • Did fully structured language emerge gradually or through one major cognitive transition?
  • Which capacities were present in the common ancestors of humans and Neanderthals?
  • Can archaeological proxies ever distinguish complex language from rich non-linguistic teaching?

22.2 Modality

  • Did gesture precede speech, speech precede gesture, or did language emerge from an integrated multimodal system?
  • How should the project represent sign languages without treating vocal speech as the default definition of language?

22.3 Prosody

  • Which prosodic functions predate lexical language?
  • Did language and music diverge from a shared prosodic communication system?
  • How much emotional prosody is biologically constrained versus culturally learned?

22.4 Cognition and sociality

  • Did language primarily evolve for teaching, cooperation, social bonding, planning, persuasion or several interacting pressures?
  • Which forms of shared intentionality are prerequisites and which are consequences?

22.5 Genetics

  • Which gene networks matter for speech-motor learning, auditory processing and language development?
  • How should genetic evidence be communicated without reviving simplistic “language gene” stories?

22.6 Information theory

  • How should spoken-language bandwidth be compared fairly with writing, images, gesture and digital media?
  • Should the master dataset distinguish acoustic bit rate, linguistic information rate and meaningful information rate?

22.7 Power and access

  • How should accent prejudice, language suppression and prestige varieties be represented across later governance topics?
  • At what point should interpreting and translation become distinct institutional topics?
23. Research Gaps Before v1.0

The following questions would benefit from further research.

  1. Add a dedicated sign-language comparison source and determine whether a separate early encoding topic is needed.
  2. Review recent palaeoanthropological work on hearing, vocal-tract reconstruction and Neanderthal communication.
  3. Add African linguistic and anthropological scholarship to counterbalance the current source base.
  4. Add evidence on speech in hunter-gatherer teaching, childcare and cooperative work without using one society as a fossil model of all prehistory.
  5. Develop a formal method for scoring bandwidth and fidelity across unlike media.
  6. Build a source table separating direct evidence, experimental analogy and theoretical inference.
  7. Review disability and accessibility framing with Deaf and speech-language scholarship.
  8. Verify all bibliography metadata and DOI records.
  9. Conduct an independent contradiction review focusing on the antiquity of language.
  10. Decide whether Vocalisation, prosody and spoken language remains an umbrella topic after the first Era I visualisation.
24. Visual Opportunities

24.1 The speech stack

A vertical anatomical and cognitive diagram:

Intent
↓
Concepts
↓
Words and grammar
↓
Motor plan
↓
Lungs + larynx + vocal tract
↓
Acoustic signal
↓
Ear + brain
↓
Interpretation

24.2 Words are only one layer

Show one spoken sentence as parallel channels:

  • lexical content;
  • grammar;
  • pitch;
  • rhythm;
  • loudness;
  • facial expression;
  • gesture;
  • context.

24.3 Constraint-relaxation card

Before: emotion, pointing, immediate reference.
After: explanation, plans, abstractions, absent events and hypotheticals.

24.4 Speech versus writing

| Speech | Writing | |---|---| | Immediate | Persistent | | Interactive | Searchable | | Prosodically rich | Visually scannable | | Short range unaided | Transportable as object | | Low material cost | Stronger replication and accountability |

24.5 The disappearing signal

Animate a waveform appearing and vanishing while biological memory attempts to retain fragments. Then introduce oral repetition, writing and recording as different persistence solutions.

24.6 The multimodal human

A speaker at the centre with channels radiating from:

  • mouth;
  • eyes;
  • hands;
  • face;
  • posture;
  • distance;
  • shared surroundings.

24.7 Evolutionary evidence fan

Place “origin of language” at the centre, surrounded by incomplete evidence streams:

  • fossils;
  • genes;
  • tools;
  • comparative animals;
  • child development;
  • present languages;
  • neuroscience.

No single stream reaches the centre alone.

24.8 Successor tree

Spoken language
├── conversation
├── oral tradition
├── song and chant
├── writing
├── recorded speech
├── telephone
├── broadcasting
├── speech recognition
├── machine translation
└── conversational AI
26. Proposed Register Update

The following register entry is recommended for v0.2:

|---|---|---|---|---|---|---|---|---|---|---| | Vocalisation, prosody and spoken language | Vocalisation, prosody and spoken language | Method | Encoding & expression | Transmission; interpretation; feedback | Deep prehistory; no reliable single origin date | Encodes complex, flexible and abstract meaning in rapidly produced vocal sound | Biological vocal signalling; embodied communication; memory; shared cognition | Conversation; oral tradition; writing; recorded speech; telephony; voice interfaces | Core | Researched |

26.1 Register note

Add:

Vocalisation, prosody and spoken language is an umbrella topic for the vocal-auditory channel. Language is not synonymous with speech; signed languages demonstrate that linguistic structure can operate in other modalities. A future split may separate non-lexical vocalisation, prosody and spoken linguistic encoding.

28. Key Analytical Conclusion

Vocalisation, prosody and spoken language marks the point at which information transmission became generative rather than merely reactive.

A fixed call can warn of a predator. Spoken language can describe:

  • the predator that was seen yesterday;
  • the route it may take tomorrow;
  • the reason the group should avoid the river;
  • the possibility that the report is wrong;
  • a plan for what to do if it returns.

The topic did not defeat distance, time or memory. It defeated a more fundamental limitation: the inability to package complex imagined relations into a signal another mind could reconstruct.

That achievement became the substrate upon which nearly every later human information system was built.

29. Sources

The sources below support factual and scholarly claims in this research notes. Interpretive claims specific to the Information Transmission Evolution Map are labelled as analysis in the text.

Core references

[1] National Institute on Deafness and Other Communication Disorders. “Voice, Speech, and Language: What Are They?” Updated 2023.

[2] National Institute on Deafness and Other Communication Disorders. “How Does the Human Body Produce Voice and Speech?” Updated 13 March 2023.

[3] Vigliocco, G., Perniss, P., and Vinson, D. “Language as a Multimodal Phenomenon: Implications for Language Learning, Processing and Evolution.” Philosophical Transactions of the Royal Society B 369 (2014): 20130292.

[4] Levinson, S. C., and Holler, J. “The Origin of Human Multi-Modal Communication.” Philosophical Transactions of the Royal Society B 369 (2014): 20130302.

[5] Zhang, Y. et al. “More Than Words: Word Predictability, Prosody, Gesture and Mouth Movements in Natural Language Comprehension.” Proceedings of the Royal Society B 288 (2021): 20210500.

[6] Wagner, M., and Watson, D. G. “Experimental and Theoretical Advances in Prosody: A Review.” Language and Cognitive Processes 25 (2010): 905–945.

[7] Matzinger, T. et al. “Voice-Modulatory Cues to Structure across Languages and Species.” Philosophical Transactions of the Royal Society B 376 (2021): 20200393.

[8] Hauser, M. D. et al. “The Mystery of Language Evolution.” Frontiers in Psychology 5 (2014): 401. DOI: 10.3389/fpsyg.2014.00401.

[9] Anikin, A. et al. “Beyond Speech: Exploring Diversity in the Human Voice.” iScience 26, no. 11 (2023): 108204. DOI: 10.1016/j.isci.2023.108204.

[10] Dediu, D., and Levinson, S. C. “On the Antiquity of Language: The Reinterpretation of Neandertal Linguistic Capacities and Its Consequences.” Frontiers in Psychology 4 (2013): 397.

[11] Dediu, D., Moisik, S. R., Baetsen, W. A., Bosman, A. M., and Waters-Rist, A. L. “The Vocal Tract as a Time Machine: Inferences about Past Speech and Language from the Anatomy of the Speech Organs.” Philosophical Transactions of the Royal Society B 376, no. 1824 (2021): 20200192.

[12] Boë, L.-J. et al. “Which Way to the Dawn of Speech? Reanalyzing Half a Century of Debates and Data in Light of Speech Science.” Science Advances 5, no. 12 (2019): eaaw3916. DOI: 10.1126/sciadv.aaw3916.

[13] Fisher, S. E., and Scharff, C. “FOXP2 as a Molecular Window into Speech and Language.” Trends in Genetics 25 (2009): 166–177.

[14] Nudel, R., and Newbury, D. F. “FOXP2.” WIREs Cognitive Science 4 (2013): 547–560. DOI: 10.1002/wcs.1247.

[15] Morgan, T. J. H. et al. “Experimental Evidence for the Co-Evolution of Hominin Tool-Making Teaching and Language.” Nature Communications 6 (2015): 6029. DOI: 10.1038/ncomms7029.

[16] Coupé, C., Oh, Y. M., Dediu, D., and Pellegrino, F. “Different Languages, Similar Encoding Efficiency: Comparable Information Rates across the Human Communicative Niche.” Science Advances 5 (2019): eaaw2594.

[17] Pinker, S., Nowak, M. A., and Lee, J. J. “The Logic of Indirect Speech.” Proceedings of the National Academy of Sciences 105 (2008): 833–838.

[18] Laland, K. N. “The Origins of Language in Teaching.” Psychonomic Bulletin & Review 24 (2017): 225–231.

Supplementary references for later review

  • Brown, S. “A Joint Prosodic Origin of Language and Music.” Frontiers in Psychology 8 (2017).
  • Burkhardt-Reed, M. M., Long, H. L., Bowman, D. D., Bene, E. R., and Oller, D. K. “The Origin of Language and Relative Roles of Voice and Gesture in Early Communication Development.” Infant Behavior and Development 65 (2021): 101648. DOI: 10.1016/j.infbeh.2021.101648.
  • Perlman, M. et al. “People Can Create Iconic Vocalizations to Communicate Various Meanings to Naïve Listeners.” Scientific Reports 8 (2018).
  • Blasi, D. E. et al. “Human Sound Systems Are Shaped by Post-Neolithic Changes in Bite Configuration.” Science 363 (2019).
  • Barney, A., Martelli, S., Serrurier, A., and Steele, J. “Articulatory Capacity of Neanderthals, a Very Recent and Human-Like Fossil Hominin.” Philosophical Transactions of the Royal Society B 367, no. 1585 (2012): 88-102. DOI: 10.1098/rstb.2011.0259.
30. Version Notes

v0.1: 28 July 2026

  • Created the first full topic research notes.
  • Preserved Vocalisation, prosody and spoken language as an umbrella topic.
  • Distinguished voice, speech, language and prosody.
  • Added explicit uncertainty around language origins and Neanderthal capacity.
  • Added communication model, evaluation matrix, civilisational contribution, harms, graph relationships, claim register and content opportunities.
  • Recommended update from Scoped to Researched.
  • Flagged source, accessibility and geographic coverage work required before v1.0.
Evidence

Sources and further reading

  1. National Institute on Deafness and Other Communication Disorders. “Voice, Speech, and Language: What Are They?” Updated 2023.

  2. National Institute on Deafness and Other Communication Disorders. “How Does the Human Body Produce Voice and Speech?” Updated 13 March 2023.

  3. Vigliocco, G., Perniss, P., and Vinson, D. “Language as a Multimodal Phenomenon: Implications for Language Learning, Processing and Evolution.” *Philosophical Transactions of the Royal Society B* 369 (2014): 20130292.

  4. Levinson, S. C., and Holler, J. “The Origin of Human Multi-Modal Communication.” *Philosophical Transactions of the Royal Society B* 369 (2014): 20130302.

  5. Zhang, Y. et al. “More Than Words: Word Predictability, Prosody, Gesture and Mouth Movements in Natural Language Comprehension.” *Proceedings of the Royal Society B* 288 (2021): 20210500.

  6. Wagner, M., and Watson, D. G. “Experimental and Theoretical Advances in Prosody: A Review.” *Language and Cognitive Processes* 25 (2010): 905–945.

  7. Matzinger, T. et al. “Voice-Modulatory Cues to Structure across Languages and Species.” *Philosophical Transactions of the Royal Society B* 376 (2021): 20200393.

  8. Hauser, M. D. et al. “The Mystery of Language Evolution.” *Frontiers in Psychology* 5 (2014): 401. DOI: 10.3389/fpsyg.2014.00401.

  9. Anikin, A. et al. “Beyond Speech: Exploring Diversity in the Human Voice.” *iScience* 26, no. 11 (2023): 108204. DOI: 10.1016/j.isci.2023.108204.

  10. Dediu, D., and Levinson, S. C. “On the Antiquity of Language: The Reinterpretation of Neandertal Linguistic Capacities and Its Consequences.” *Frontiers in Psychology* 4 (2013): 397.

  11. Dediu, D., Moisik, S. R., Baetsen, W. A., Bosman, A. M., and Waters-Rist, A. L. “The Vocal Tract as a Time Machine: Inferences about Past Speech and Language from the Anatomy of the Speech Organs.” *Philosophical Transactions of the Royal Society B* 376, no. 1824 (2021): 20200192.

  12. Boë, L.-J. et al. “Which Way to the Dawn of Speech? Reanalyzing Half a Century of Debates and Data in Light of Speech Science.” *Science Advances* 5, no. 12 (2019): eaaw3916. DOI: 10.1126/sciadv.aaw3916.

  13. Fisher, S. E., and Scharff, C. “FOXP2 as a Molecular Window into Speech and Language.” *Trends in Genetics* 25 (2009): 166–177.

  14. Nudel, R., and Newbury, D. F. “FOXP2.” *WIREs Cognitive Science* 4 (2013): 547–560. DOI: 10.1002/wcs.1247.

  15. Morgan, T. J. H. et al. “Experimental Evidence for the Co-Evolution of Hominin Tool-Making Teaching and Language.” *Nature Communications* 6 (2015): 6029. DOI: 10.1038/ncomms7029.

  16. Coupé, C., Oh, Y. M., Dediu, D., and Pellegrino, F. “Different Languages, Similar Encoding Efficiency: Comparable Information Rates across the Human Communicative Niche.” *Science Advances* 5 (2019): eaaw2594.

  17. Pinker, S., Nowak, M. A., and Lee, J. J. “The Logic of Indirect Speech.” *Proceedings of the National Academy of Sciences* 105 (2008): 833–838.

  18. Laland, K. N. “The Origins of Language in Teaching.” *Psychonomic Bulletin & Review* 24 (2017): 225–231. - Brown, S. “A Joint Prosodic Origin of Language and Music.” *Frontiers in Psychology* 8 (2017). - Burkhardt-Reed, M. M., Long, H. L., Bowman, D. D., Bene, E. R., and Oller, D. K. “The Origin of Language and Relative Roles of Voice and Gesture in Early Communication Development.” *Infant Behavior and Development* 65 (2021): 101648. DOI: 10.1016/j.infbeh.2021.101648. - Perlman, M. et al. “People Can Create Iconic Vocalizations to Communicate Various Meanings to Naïve Listeners.” *Scientific Reports* 8 (2018). - Blasi, D. E. et al. “Human Sound Systems Are Shaped by Post-Neolithic Changes in Bite Configuration.” *Science* 363 (2019). - Barney, A., Martelli, S., Serrurier, A., and Steele, J. “Articulatory Capacity of Neanderthals, a Very Recent and Human-Like Fossil Hominin.” *Philosophical Transactions of the Royal Society B* 367, no. 1585 (2012): 88-102. DOI: 10.1098/rstb.2011.0259. - Created the first full topic research notes. - Preserved `Vocalisation, prosody and spoken language` as an umbrella topic. - Distinguished voice, speech, language and prosody. - Added explicit uncertainty around language origins and Neanderthal capacity. - Added communication model, evaluation matrix, civilisational contribution, harms, graph relationships, claim register and content opportunities. - Recommended update from `Scoped` to `Researched`. - Flagged source, accessibility and geographic coverage work required before v1.0.