8.1 Symbolic input layer
Input may include words, numerals, abbreviations, markup, punctuation, emojis or code. Each requires a speaking policy.
8.2 Normalisation layer
Written forms are expanded. “£12.50” can be read differently by locale and context.
8.3 Linguistic layer
The system predicts pronunciation, stress, phrase boundaries and sometimes semantic features.
8.4 Speaker and style layer
Voice identity, emotion, speaking rate and expressive style condition the performance.
8.5 Alignment and duration layer
Text tokens are mapped to acoustic time. Failures can repeat, skip or stretch words.
8.6 Acoustic representation layer
The system predicts spectral or other features describing how speech should sound.
8.7 Waveform layer
A vocoder or direct generator produces audio samples.
8.8 Playback layer
Devices, codecs and environments alter the listener's experience.
8.9 Identity interpretation layer
Listeners infer age, gender, region, emotion and personhood from the voice, often more strongly than the system can justify.
8.10 Governance layer
Consent, attribution, watermarking, access control and provenance determine whether the output is legitimate and auditable.