Machine-Mediated Meaning · Interpreting meaning

Generative Image, Audio and Video Models

Generative image, audio and video models produce perceptual media from learned distributions rather than requiring a conventional camera, microphone, instrument or manually constructed frame for every output. They can generate media from noise, text prompts, reference images, sketches, motion, audio, masks, depth maps, style examples or other conditioning signals. The resulting object may resemble a photograph.

When it emerged
Neural generative media research from the 2010s; widespread text-conditioned generation from the 2020s
What changed
Reduces the requirement to capture, stage or manually construct every media object before communication
Reading time
24 minutes
The essential questions

Generative Image, Audio and Video Models, clearly explained

Generative image, audio and video models produce perceptual media from learned distributions rather than requiring a conventional camera, microphone, instrument or manually constructed frame for every output. They can generate media from noise, text prompts, reference images, sketches, motion, audio, masks, depth maps, style examples or other conditioning signals. The resulting object may resemble a photograph, recording or filmed event even when no corresponding scene or performance occurred.

What is it?

Generative image, audio and video models are defined here as learned computational systems that produce new perceptual media samples from a modelled distribution, optionally conditioned on text, reference media, structure, identity, style, motion or other controls.

What problem did it solve?

The primary constraint reduced is the need for every image, sound or moving scene to be directly captured, physically staged or manually constructed before it can be communicated.

How did it work?

They can generate media from noise, text prompts, reference images, sketches, motion, audio, masks, depth maps, style examples or other conditioning signals. The resulting object may resemble a photograph, recording or filmed event even when no corresponding scene or performance occurred. Variational autoencoders learned latent spaces from which new samples could be drawn.

What came before?

It built on Online video and streaming platforms.

What did it make possible?

Its methods, infrastructure or conventions were absorbed into later information systems.

What survived?

Older methods continued where they remained cheaper, more trustworthy, more accessible or better suited to local needs.

Why does it still matter?

A plausible image, voice or scene can be produced without a camera or performer recording that exact content. Concept art, storyboards, prototypes and illustrations can be generated rapidly. Inpainting, extension, relighting and motion synthesis blur the line between altering and creating.

Deep dive

The deeper story

Generative image, audio and video models produce perceptual media from learned distributions rather than requiring a conventional camera, microphone, instrument or manually constructed frame for every output. They can generate media from noise, text prompts, reference images, sketches, motion, audio, masks, depth maps, style examples or other conditioning signals. The resulting object may resemble a photograph, recording or filmed event even when no corresponding scene or performance occurred.

The topic grows from several technical lineages. Variational autoencoders learned latent spaces from which new samples could be drawn. Generative adversarial networks trained a generator against a discriminator and produced increasingly realistic images. Neural style transfer separated aspects of content and visual style. WaveNet modelled raw audio, while diffusion models learned to reverse a gradual noising process. Contrastive text-image models linked linguistic descriptions with visual representations, and latent diffusion moved generation into compressed feature spaces, making high-resolution text-conditioned image synthesis more practical. [S01-S08]

Audio generation developed beyond speech synthesis into music, sound effects and long-form acoustic continuation. Systems such as AudioLM modelled audio through discrete representations at several timescales, while MusicLM conditioned hierarchical audio generation on text. Video models added a temporal problem: objects, identities, lighting and motion must remain coherent across frames. Video diffusion systems and cascaded text-to-video models improved fidelity, duration and resolution, but temporal consistency and physical plausibility remain distinct from visual beauty. [S09-S13]

Synthetic media reduces the cost of visualisation, prototyping, dubbing, design, sound production and animation. It can help creators explore ideas that would be expensive or impossible to capture. It also destabilises a historical shortcut in human reasoning: photographic or acoustic resemblance has often been treated as evidence that something happened. A generated image is not a photograph merely because it looks photographic. A generated voice is not testimony. A generated video is not a record of an event.

The topic must therefore separate synthesis from capture, resemblance from identity, perceptual realism from physical truth, generation from editing and file metadata from trustworthy provenance. Detection alone is not a durable solution because generators and detectors co-evolve, media can be transformed and authentic files can lose metadata. Provenance systems such as C2PA offer a complementary strategy by carrying signed assertions about origin and edits, but provenance proves only what the signed chain supports. It does not make the depicted claim true. [S14-S16]

Synthetic media is also a labour and cultural system. Models are trained on the work, voices, faces and styles of many people. The output may compete with those contributors without credit or consent. The same system can support accessibility and localisation while enabling impersonation, non-consensual sexual imagery, fraud, propaganda and evidence pollution.

The big idea

Generative media models turn learned perceptual distributions into programmable images, sounds and moving scenes. Their defining achievement is media production without conventional capture or frame-by-frame construction. Their recurring danger is that perceptual realism can manufacture false evidence, identity and attribution while obscuring the human works and institutional choices from which the model learned.

Main problem addressed

Reduces the requirement to capture, stage or manually construct every media object before communication

Connections

What came before and what followed

Start with the key connections, then reveal the wider network when you need more context.

Connections for Generative Image, Audio and Video ModelsOnline video andstreaming platformsDigital Provenanceand AuthenticitySystemsGenerative Image, Audioand Video Models
Timeline

Key moments

How Generative Image, Audio and Video Models emerged

This marks the broad emergence and development of Generative Image, Audio and Video Models. Why it mattered: Reduces the requirement to capture, stage or manually construct every media object before communication.

Phase 3 - Neural style, super-resolution and face synthesis, mid-to-late 2010s

Models transform and generate increasingly realistic visual media.

Generative Image, Audio and Video Models · practical implementation

Phase 2 - Neural latent and adversarial generation, 2013-2016

VAEs and GANs make learned image generation a major research field.

Generative Image, Audio and Video Models · practical implementation

Phase 4 - Neural waveform and audio generation, 2016 onward

Raw-audio and learned-codec models extend generation to speech, music and sound.

Generative Image, Audio and Video Models · practical implementation

Phase 5 - Diffusion and text-image alignment, 2020-2022

Denoising models and contrastive encoders enable high-fidelity text-conditioned image generation.

Generative Image, Audio and Video Models · practical implementation

Phase 8 - Integrated synthetic production and provenance, 2020s onward

Generation enters mainstream media tools while provenance, watermarking and consent become infrastructure questions.

Generative Image, Audio and Video Models · practical implementation

Phase 6 - Latent and controllable image generation, 2021 onward

Compressed-space generation, inpainting and structural controls make synthesis more accessible and editable.

Generative Image, Audio and Video Models · practical implementation

Phase 7 - Text-conditioned video and long-form audio, 2022 onward

Hierarchical and diffusion systems generate coherent media across longer temporal spans.

Generative Image, Audio and Video Models · practical implementation
People and organisations

Who helped shape it?

Ian Goodfellow

Ian Goodfellow is one of the people connected to this topic. Open the profile for the wider historical context.

NIST

NIST is one of the organisations connected to this topic. Open the profile for the wider historical context.

Research notes

Open the full research notes

These expandable sections preserve the detailed research behind the public explanation.

1. Executive Summary

Generative image, audio and video models produce perceptual media from learned distributions rather than requiring a conventional camera, microphone, instrument or manually constructed frame for every output. They can generate media from noise, text prompts, reference images, sketches, motion, audio, masks, depth maps, style examples or other conditioning signals. The resulting object may resemble a photograph, recording or filmed event even when no corresponding scene or performance occurred.

The topic grows from several technical lineages. Variational autoencoders learned latent spaces from which new samples could be drawn. Generative adversarial networks trained a generator against a discriminator and produced increasingly realistic images. Neural style transfer separated aspects of content and visual style. WaveNet modelled raw audio, while diffusion models learned to reverse a gradual noising process. Contrastive text-image models linked linguistic descriptions with visual representations, and latent diffusion moved generation into compressed feature spaces, making high-resolution text-conditioned image synthesis more practical. [S01-S08]

Audio generation developed beyond speech synthesis into music, sound effects and long-form acoustic continuation. Systems such as AudioLM modelled audio through discrete representations at several timescales, while MusicLM conditioned hierarchical audio generation on text. Video models added a temporal problem: objects, identities, lighting and motion must remain coherent across frames. Video diffusion systems and cascaded text-to-video models improved fidelity, duration and resolution, but temporal consistency and physical plausibility remain distinct from visual beauty. [S09-S13]

Synthetic media reduces the cost of visualisation, prototyping, dubbing, design, sound production and animation. It can help creators explore ideas that would be expensive or impossible to capture. It also destabilises a historical shortcut in human reasoning: photographic or acoustic resemblance has often been treated as evidence that something happened. A generated image is not a photograph merely because it looks photographic. A generated voice is not testimony. A generated video is not a record of an event.

The topic must therefore separate synthesis from capture, resemblance from identity, perceptual realism from physical truth, generation from editing and file metadata from trustworthy provenance. Detection alone is not a durable solution because generators and detectors co-evolve, media can be transformed and authentic files can lose metadata. Provenance systems such as C2PA offer a complementary strategy by carrying signed assertions about origin and edits, but provenance proves only what the signed chain supports. It does not make the depicted claim true. [S14-S16]

Synthetic media is also a labour and cultural system. Models are trained on the work, voices, faces and styles of many people. The output may compete with those contributors without credit or consent. The same system can support accessibility and localisation while enabling impersonation, non-consensual sexual imagery, fraud, propaganda and evidence pollution.

The big idea

Generative media models turn learned perceptual distributions into programmable images, sounds and moving scenes. Their defining achievement is media production without conventional capture or frame-by-frame construction. Their recurring danger is that perceptual realism can manufacture false evidence, identity and attribution while obscuring the human works and institutional choices from which the model learned.

2. Identification

| Field | Value | |---|---| | Public title | Generative Image, Audio and Video Models | | Analytical title | Learned Big-picture essays of Visual, Acoustic and Audiovisual Artefacts From Latent and Conditional Representations | | Recommended type | Generative perceptual-media model family | | Primary category | Interpretation & mediation | | Secondary categories | Encoding; reproduction; processing; distribution; identity; governance | | Emergence | Neural generative image and audio research in the 2010s; widespread text-conditioned media generation in the 2020s |

3. Operational Definition

Generative image, audio and video models are defined here as learned computational systems that produce new perceptual media samples from a modelled distribution, optionally conditioned on text, reference media, structure, identity, style, motion or other controls.

The topic includes variational autoencoders, generative adversarial networks, autoregressive media models, diffusion models, flow-based and masked generative models, latent-variable models, text-image alignment, text-to-image synthesis, image-to-image translation, inpainting, outpainting, super-resolution where generative inference creates detail, audio language models, music generation, sound-effect generation, video prediction, text-to-video, image-to-video, video editing, avatar generation, face reenactment, synthetic actors, multimodal conditioning and generated-media provenance.

It includes voice cloning and speech generation only where needed for comparison; those are treated in detail in Text-to-Speech and Voice Synthesis. It excludes ordinary photography, recording and animation; deterministic codecs; procedural graphics that follow explicit hand-authored rules without learned distributions; and conventional editing where all visible or audible source material was captured or manually created.

A generated object can contain elements derived from reference media and still be synthetic. Conversely, an authentic recording can be edited, composited or relabelled deceptively without generative AI. “Synthetic” and “misleading” are therefore overlapping but not identical categories.

4. Why the Topic Matters

1. Media no longer requires a corresponding captured event

A plausible image, voice or scene can be produced without a camera or performer recording that exact content.

2. Visualisation cost falls

Concept art, storyboards, prototypes and illustrations can be generated rapidly.

3. Production becomes conversational and iterative

Text and reference controls allow non-specialists to explore media ideas.

4. Editing and generation converge

Inpainting, extension, relighting and motion synthesis blur the line between altering and creating.

5. Identity becomes a controllable condition

Faces, voices and performance styles can be generated or transferred.

6. Evidence assumptions weaken

Photorealism and acoustic realism no longer guarantee capture.

7. Training data becomes cultural infrastructure

Model outputs reflect large collections of human art, photographs, music and performance.

8. Media abundance increases discovery and trust pressure

The bottleneck moves from producing an image to determining origin, relevance, legitimacy and responsibility.

5. Terminology
  • Generative model: Model producing new samples from a learned distribution.
  • Latent variable: Hidden representation used to explain or generate observed data.
  • Latent space: Learned representation space in which related media characteristics may become organised.
  • Variational autoencoder: Encoder-decoder model trained with a probabilistic latent distribution. [1]
  • Generator: Network producing candidate samples.
  • Discriminator: Network trained to distinguish generated from training samples in a GAN.
  • Generative adversarial network: Generator and discriminator trained in competition. [2]
  • Mode collapse: GAN failure where the generator produces limited varieties of output.
  • Diffusion model: Generative model learning to reverse a progressive noising process. [5]
  • Denoising: Prediction used to move a noisy sample towards a data-like sample.
  • Sampling step: One iterative update during diffusion generation.
  • Latent diffusion: Diffusion performed in a compressed learned representation rather than directly in pixel space. [8]
  • Conditioning: Additional signal controlling generation, such as text, image, mask, identity or pose.
  • Text encoder: Model converting a prompt into representations used to condition media generation.
  • Contrastive learning: Learning representations by bringing matched pairs closer and separating mismatched pairs.
  • CLIP: Contrastive text-image representation model used widely for conditioning and evaluation. [6]
  • Text-to-image: Generation of images from textual descriptions.
  • Image-to-image: Transformation conditioned on a source image.
  • Inpainting: Generation inside a masked region.
  • Outpainting: Extension beyond an original image boundary.
  • Super-resolution: Increasing resolution, sometimes by generating plausible fine detail not directly observed.
  • Style transfer: Recomposition of content according to visual features associated with another image or style. [3]
  • Autoregressive media model: Model generating discrete or continuous elements sequentially.
  • Neural codec: Learned encoder-decoder for compact audio or visual representations.
  • Audio token: Discrete learned unit representing acoustic content or detail.
  • Audio language model: Sequence model generating audio representations.
  • Video diffusion: Diffusion model extended across spatial and temporal dimensions. [11]
  • Temporal coherence: Stability of objects, motion and identity across frames.
  • Frame interpolation: Generation of intermediate frames between observed frames.
  • Image-to-video: Generation of motion and additional frames from a still image.
  • Text-to-video: Generation of a video conditioned on a prompt.
  • Deepfake: Synthetic or manipulated media that depicts a person saying or doing something not authentically performed; popular term with inconsistent scope.
  • Face swap: Replacement or synthesis of identity in facial imagery.
  • Reenactment: Transfer of expression, pose or motion from a source performance to a target identity.
  • Voice cloning: Generation of speech resembling a target speaker.
  • Synthetic actor: Generated or composited audiovisual identity capable of new performances.
  • Prompt: Textual conditioning input.
  • Negative prompt: Requested features to suppress; implementation dependent.
  • Guidance: Technique steering generation towards conditioning.
  • Seed: Initial random state that can support partial reproducibility.
  • Watermark: Signal embedded in media to indicate origin or generation, with varying robustness.
  • Provenance credential: Signed record describing origin, edits or claims associated with an asset.
  • C2PA: Standard for content provenance and authenticity assertions. [15]
  • Detection: Statistical inference that media is synthetic or manipulated.
  • Attribution: Identification of a likely model, source or creator.
  • Authenticity: Context-dependent property concerning origin and integrity, not merely absence of generation.
  • Photorealism: Resemblance to photographic imagery, not evidence of photography.
  • Physical plausibility: Consistency with physical structure and causation.
  • Semantic alignment: Agreement between prompt or intended description and output.
  • Aesthetic score: Model or human rating of visual appeal, distinct from truth or fidelity.
6. Boundary With Neighbouring Topics

1. Generation versus capture

Capture records signals from an event. Generation computes a new media object from a learned distribution and conditions.

2. Generation versus editing

Editing modifies existing media. Generative editing synthesises new regions, frames or sounds that were not captured.

3. Synthetic versus false

Synthetic media can be clearly labelled fiction or design. Authentic footage can be presented with a false caption.

4. Photorealism versus authenticity

Photorealism is a visual property. Authenticity concerns origin, integrity and context.

5. Resemblance versus identity

An output can resemble a person without their participation or consent.

6. Style similarity versus source copying

A model may imitate broad characteristics, reproduce memorised fragments or create a new combination. These require separate evidence.

7. Image generation versus image retrieval

Generation creates a sample. Retrieval selects an existing image from an index.

8. Image synthesis versus procedural graphics

Procedural graphics follow explicit rules. Generative models learn distributions from examples.

9. Speech synthesis versus general audio generation

Speech synthesis is constrained by linguistic content and identity. General audio models produce music, ambience and sound events.

10. Video generation versus animation

Animation can be manually or procedurally authored frame by frame. Generative video infers frames and motion from learned data.

11. Detection versus provenance

Detection estimates whether media appears synthetic. Provenance records signed claims about origin and processing.

12. Provenance versus truth

A valid provenance chain can prove who signed and transformed an asset, not that the depicted claim is accurate.

13. Watermark versus identity

A watermark may indicate a tool or provider. It does not identify every human responsible for use.

14. Generation model versus media platform

The model produces assets. A platform stores, distributes, recommends, moderates and monetises them.

7. Communication Pattern

The basic training pattern is:

Human-created or captured media → collection and metadata → filtering and normalisation → learned representation → generative objective → parameter updates → model checkpoint

The basic generation pattern is:

Prompt or reference condition → representation encoder → random or initial latent state → iterative or autoregressive generation → decoder or vocoder → image, audio or video asset → metadata and provenance → platform distribution → receiver interpretation

A generative editing path adds:

Source asset → mask, pose, depth, motion or style control → synthesis of changed regions or frames → composite output

A verification path adds:

Asset → metadata and signature validation → provenance display → detection model or forensic analysis → contextual verification → evidentiary judgment

No single step proves the depicted event occurred. Verification remains a multi-source process.

8. Expanded Communication Model

8.1 Source-data layer

Images, recordings, films, captions, metadata and identities become training material. Selection and licensing shape the model.

8.2 Representation layer

Encoders convert perceptual media and text into latent or token representations.

8.3 Objective layer

The model learns reconstruction, adversarial, likelihood, denoising or contrastive objectives.

8.4 Conditioning layer

Text, sketches, masks, identity, pose, timing or reference media specify desired attributes.

8.5 Sampling layer

Randomness and decoding produce one candidate among many possible outputs.

8.6 Rendering layer

Latent representations become pixels, waveforms or compressed media streams.

8.7 Post-processing layer

Upscaling, colour correction, editing, compression and compositing alter the generated asset.

8.8 Provenance layer

Metadata, signed claims and watermarks may record origin and edits. They can also be absent or stripped.

8.9 Distribution layer

Platforms, messaging and social feeds determine exposure and context.

8.10 Interpretation layer

Receivers infer event, identity, intention and reality from perceptual cues.

8.11 Verification layer

Journalists, investigators and organisations compare sources, provenance, timelines and physical evidence.

8.12 Feedback layer

Generated media enters public corpora and may influence future models, detection and cultural expectations.

9. Historical Emergence

Generative modelling has deep roots in statistical signal processing and pattern synthesis, but neural generative media accelerated in the 2010s. Variational autoencoders provided a principled latent-variable framework in which an encoder approximated a posterior distribution and a decoder generated samples from latent representations. [1]

Generative adversarial networks introduced a competitive training process: a generator tried to create samples that a discriminator could not distinguish from data. GANs produced rapid improvements in image realism and enabled image translation, super-resolution, face synthesis and deepfake techniques. Their adversarial objective also brought instability and mode collapse. [2]

Neural style transfer showed that convolutional representations could separate and recombine aspects of image content and style. It became an early public demonstration that learned visual features could transform an image according to another aesthetic reference. [3]

Audio generation advanced through waveform and representation modelling. WaveNet generated raw audio autoregressively and demonstrated high-quality speech and broader audio modelling. Later systems used learned codecs and hierarchical token sequences to represent long-form sound more efficiently. AudioLM generated coherent continuations while separating semantic structure from acoustic detail; MusicLM conditioned hierarchical audio generation on text descriptions. [4][9][10]

Diffusion models became a dominant family after denoising diffusion probabilistic models demonstrated high-quality image synthesis through iterative reversal of a noising process. Text-image representation models such as CLIP supplied strong language-conditioned features. DALL-E explored text-to-image generation with discrete visual tokens, while latent diffusion performed synthesis in compressed spaces and supported flexible text, image and mask conditioning. [S05-S08]

Video generation extended these methods across time. Video Diffusion Models jointly learned spatial and temporal structure and introduced methods for extending duration and resolution. Imagen Video used cascaded diffusion and super-resolution models. Make-A-Video combined text-image priors with unlabelled video to learn motion. These systems demonstrated increasingly coherent short clips while exposing persistent problems: identity drift, object disappearance, impossible motion and causal inconsistency. [S11-S13]

By the 2020s, generative media entered consumer tools, creative suites and platform workflows. Images, music and videos could be generated from text, while reference-conditioned systems made identity and style increasingly controllable. The technology moved from laboratory samples to a general media-production layer.

Governance followed. NIST analysed risks from synthetic content, including detection, provenance and authentication. C2PA developed a standard for cryptographically signed content credentials describing asset origin and edits. These approaches recognise that visual inspection alone cannot carry the entire burden of trust. [S14-S16]

10. Prerequisites
  • Large digitised image, audio and video corpora.
  • Metadata and text-media pairing.
  • Neural-network training and automatic differentiation.
  • GPUs and distributed computing.
  • Digital image and audio encoding.
  • Learned compression and latent representations.
  • Language models and text encoders.
  • Storage and content-delivery infrastructure.
  • Human preference and quality evaluation.
  • Identity and consent governance.
  • Copyright and licensing systems.
  • Media forensics and provenance standards.
  • Moderation and distribution platforms.
11. Periodisation

Phase 1 - Statistical and procedural synthesis before deep learning

Graphics, signal models and procedural systems generate media through explicit rules and fitted models.

Phase 2 - Neural latent and adversarial generation, 2013-2016

VAEs and GANs make learned image generation a major research field.

Phase 3 - Neural style, super-resolution and face synthesis, mid-to-late 2010s

Models transform and generate increasingly realistic visual media.

Phase 4 - Neural waveform and audio generation, 2016 onward

Raw-audio and learned-codec models extend generation to speech, music and sound.

Phase 5 - Diffusion and text-image alignment, 2020-2022

Denoising models and contrastive encoders enable high-fidelity text-conditioned image generation.

Phase 6 - Latent and controllable image generation, 2021 onward

Compressed-space generation, inpainting and structural controls make synthesis more accessible and editable.

Phase 7 - Text-conditioned video and long-form audio, 2022 onward

Hierarchical and diffusion systems generate coherent media across longer temporal spans.

Phase 8 - Integrated synthetic production and provenance, 2020s onward

Generation enters mainstream media tools while provenance, watermarking and consent become infrastructure questions.

12. Main Problem Addressed

The primary constraint reduced is the need for every image, sound or moving scene to be directly captured, physically staged or manually constructed before it can be communicated.

Secondary constraints reduced include:

  • cost of concept art and visual prototyping;
  • need for performers to record every line or variation;
  • difficulty localising media across languages and markets;
  • expense of sets, cameras, travel and special effects;
  • inability to visualise imaginary or inaccessible scenes;
  • slow iteration in design and animation;
  • limited access to production tools for non-specialists;
  • scarcity of customised media variants.

The reduction creates a new constraint: receivers and organisations need stronger provenance, consent and verification systems because perceptual realism is no longer scarce.

13. Evaluation Matrix

| Dimension | Assessment | |---|---| | Visual or acoustic fidelity | Rapidly improving, model and domain dependent | | Prompt alignment | Variable; broad intent often easier than exact composition | | Identity consistency | Uneven, especially across video time | | Temporal coherence | Major video challenge | | Physical plausibility | Often weaker than surface realism | | Editability | Strong with masks and structural controls, but not fully deterministic | | Reproducibility | Seed, model version and infrastructure dependent | | Diversity | High in principle; datasets and objectives can narrow output | | Attribution | Weak without provenance or retained generation records | | Watermark robustness | Variable under cropping, compression and transformation | | Detection reliability | Arms-race dependent and distribution sensitive | | Language support | Prompt understanding uneven across languages | | Computational cost | High for training and advanced video generation | | Consent and identity risk | High for face and voice conditioning | | Copyright uncertainty | Significant around training, outputs and style imitation | | Evidentiary trust | Requires external verification; realism alone is insufficient |

14. Advantages
  1. Rapid ideation: Many visual or acoustic concepts can be explored quickly.
  2. Lower production barriers: Individuals can create media without full studios or crews.
  3. Customisation: Assets can be adapted to audience, language and format.
  4. Accessibility: Descriptions can become images, voices or audiovisual explanations.
  5. Restoration and completion: Damaged or missing regions can be reconstructed, with disclosure.
  6. Creative variation: Style, composition and performance can be explored across many samples.
  7. Previsualisation: Film, game and design teams can test concepts before expensive production.
  8. Synthetic training environments: Simulated data can support robotics and perception research.
  9. Localisation: Lip movement, voice and visuals can be adapted across languages.
  10. Impossible scenes: Historical, fictional or microscopic environments can be visualised without literal capture.
15. Civilisational Contributions

Generative media completes a long transformation from representation as record to representation as computation. Painting and animation always allowed scenes that never occurred, but they visibly depended on skilled construction. Photographs, film and audio recordings acquired a stronger evidentiary aura because they originated in physical signals captured from events. Generative models produce that aura without the event.

This does not end photography or recording. It changes the receiver's default assumptions. The question moves from “Does this look real?” to “What chain of evidence connects this object to an event?”

The technology also democratises some forms of media experimentation. A person can produce a storyboard, soundtrack or explanatory image without specialised equipment. Yet access to generation is not the same as equality of control. Large model providers determine datasets, filters, pricing and allowed identities, while creators whose work shaped the model may lack attribution or bargaining power.

For the information map, the decisive contribution is that media becomes a queryable possibility space. The user describes or conditions an artefact, and the system computes one version. Creation begins to resemble retrieval from a latent distribution, although the returned object did not previously exist as one stored record.

16. Organisations, Access and Power

16.1 Artists, performers and rights holders

Their work, likeness and style may become training material or generation targets.

16.2 Dataset builders

They decide collection, captioning, filtering, deduplication and removal procedures.

16.3 Model developers

They control architecture, safety filters, supported modalities and release conditions.

16.4 Compute and cloud providers

Video and high-resolution generation rely on concentrated hardware infrastructure.

16.5 Creative industries

Studios, agencies, game developers and publishers integrate models into production and labour decisions.

16.6 Platforms

They distribute, recommend, label and monetise generated assets.

16.7 Subjects and identity holders

People may be depicted or imitated without participation.

16.8 Journalists and investigators

They need verification methods that combine provenance, source checking and physical evidence.

16.9 Standards bodies

C2PA and related organisations define interoperable provenance structures.

16.10 Regulators and courts

They interpret consent, publicity rights, copyright, fraud, evidence and consumer deception.

17. Limitations, Harms and Trade-Offs

17.1 False evidence

Generated media can depict events that never occurred.

17.2 Identity impersonation

Faces and voices can be used without consent for fraud, harassment or propaganda.

17.3 Non-consensual sexual imagery

Generative systems can produce abusive intimate depictions without any original photograph.

17.4 Style and labour appropriation

Outputs can compete with creators whose works trained the model without meaningful credit or bargaining.

17.5 Dataset bias

Representation, stereotypes and caption errors shape generation.

17.6 Physical incoherence

Media can look polished while violating object structure, causality or temporal continuity.

17.7 Provenance stripping

Metadata can be removed through screenshots, re-encoding or hostile distribution.

17.8 Detection arms race

Detectors often degrade as models, compression and editing change.

17.9 Authenticity reversal

People can dismiss genuine evidence as synthetic, creating a liar's dividend.

17.10 Cultural homogenisation

Dominant datasets and aesthetic scorers can narrow visual and musical variety.

17.11 Recursive training

Generated media may enter future datasets, obscuring lineage and reducing diversity.

17.12 Environmental cost

Training and serving large media models require substantial compute and energy.

17.13 Platform concentration

Access to high-quality models can depend on a few providers and opaque policies.

17.14 Attribution ambiguity

Prompt writer, reference creator, dataset contributors and model provider all influence output.

17.15 Accessibility versus deception

Voice or avatar generation can restore communication and simultaneously enable impersonation.

17.16 Evidentiary overload

Verification organisations face more media objects and less confidence in unauthenticated artefacts.

18. Relationship to Other Topics

Direct predecessors

  • Photography Photography: mechanically captured visual evidence.
  • Recorded sound and the phonograph Sound Recording: captured acoustic performance.
  • Motion-picture film and cinema Film and Video Recording: captured moving images.
  • Electronic digital computers Electronic Digital Computers: digital synthesis.
  • Cloud computing and cloud storage Cloud Computing and Cloud Storage: training and generation infrastructure.
  • Text-to-Speech and Voice Synthesis Text-to-Speech and Voice Synthesis: learned acoustic generation.
  • Generative Language Models Generative Language Models: text conditioning and shared sequence methods.

Strong supporting relationships

  • Automated Classification and Content Moderation Automated Classification and Content Moderation: synthetic-media detection and governance.
  • Online video and streaming platforms Online Video and Streaming Platforms: distribution of generated audiovisual media.
  • Recommendation Algorithms and Personalised Feeds Recommendation Algorithms and Personalised Feeds: amplification and attention allocation.
  • Digital Provenance and Authenticity Systems Digital Provenance and Authenticity Systems: origin and edit claims.

Direct successors

  • Synthetic actors and personalised media.
  • Generative design and simulation systems.
  • Multimodal assistants.
  • Immersive world generation.
  • Provenance-aware creative tools.

Important relationship

Generation weakens capture as a default proof of occurrence. Provenance can strengthen origin claims, but verification must still connect the media object to independent evidence about the world.

19. Representative Implementations and Milestones

19.1 Variational autoencoders

Probabilistic latent-variable learning made continuous generative spaces practical. [1]

19.2 Generative adversarial networks

Adversarial learning drove major advances in image synthesis and manipulation. [2]

19.3 Neural style transfer

Learned visual features enabled recombination of content and aesthetic structure. [3]

19.4 WaveNet

Autoregressive raw-audio modelling demonstrated high-quality neural waveform generation. [4]

19.5 Denoising diffusion probabilistic models

Iterative denoising established a high-quality generative image family. [5]

19.6 CLIP

Contrastive text-image learning connected natural-language descriptions with visual representations. [6]

19.7 DALL-E

Text-conditioned image generation demonstrated compositional synthesis from language prompts. [7]

19.8 Latent diffusion

Compressed-space diffusion enabled efficient high-resolution generation, inpainting and flexible conditioning. [8]

19.9 AudioLM and MusicLM

Hierarchical token modelling extended coherent generation to long-form audio and text-conditioned music. [9][10]

19.10 Video Diffusion Models

Diffusion architectures were extended across time for video generation and prediction. [11]

19.11 Imagen Video and Make-A-Video

Cascaded and weakly paired systems improved text-conditioned video generation. [12][13]

19.12 C2PA content credentials

Signed provenance assertions provide an interoperable complement to synthetic-media detection. [15]

20. Failure and Edge Cases

20.1 Photographic impossibility

A convincing image contains shadows, reflections or anatomy inconsistent with one physical scene.

20.2 Identity drift

A video character's face or clothing changes across frames.

20.3 Object permanence failure

Items disappear, merge or change count during motion.

20.4 Text rendering failure

Generated signs and labels resemble writing without preserving exact symbols.

20.5 Audio-semantic mismatch

A generated soundscape is plausible but does not correspond to the described event timing.

20.6 Voice resemblance without consent

A cloned voice produces words the person never approved.

20.7 Provenance gap

An authentic camera file loses metadata during reposting and appears no more trustworthy than a generated copy.

20.8 Watermark removal

Cropping, recompression or analogue rerecording weakens an embedded signal.

20.9 Detection false positive

Authentic compressed or unusual media is incorrectly labelled synthetic.

20.10 Detection false negative

A generated asset is edited enough to evade a detector.

20.11 Memorised fragment

A generated image reproduces a distinctive training composition or watermark.

20.12 Prompt under-specification

The system fills cultural or visual assumptions the user did not request.

21. Research Uncertainty and Open Questions
  • How should training consent and compensation work for visual, audio and performance data?
  • What evidence distinguishes broad style learning from reproduction of protected expression?
  • Which provenance signals survive ordinary editing and hostile transformation?
  • How should provenance be preserved through screenshots and social-platform recompression?
  • Can generation systems provide source-neighbour or influence records without exposing private data?
  • How should identity holders control face, voice and performance synthesis?
  • What counts as meaningful consent for posthumous or archival likeness use?
  • Can synthetic-media detectors remain reliable across unseen generators?
  • How should courts and journalists weigh unsigned but authentic media?
  • How can accessibility uses avoid creating permanent vendor dependence?
  • What metrics capture temporal coherence and physical causality rather than appearance alone?
  • How should generated training data be labelled and excluded from future corpora where necessary?
  • Can local or open models provide cultural sovereignty without enabling uncontrolled abuse?
  • When should generation be disclosed to the receiver, and at what level of editing?
  • How should platforms distinguish satire, fiction, simulation and deceptive impersonation?
22. Claim Register

| Claim | Type | Confidence | Evidence | |---|---|---:|---| | VAEs and GANs established major neural generative-media families in the 2010s | Historical/technical | High | [1][2] | | Diffusion models generate samples through learned reversal of a noising process | Technical | High | [5] | | Contrastive text-image representations support language-conditioned visual generation | Technical | High | [S06-S08] | | Latent diffusion performs generative modelling in compressed representations | Technical | High | [8] | | Hierarchical token models can generate long-form audio and text-conditioned music | Technical | High | [9][10] | | Video generation requires temporal as well as spatial coherence | Technical | High | [S11-S13] | | Perceptual realism is not proof of capture or event authenticity | Analytical | High | Research notes synthesis | | Detection and provenance are complementary rather than interchangeable | Governance/technical | High | [S14-S16] | | A valid provenance chain does not by itself prove the depicted proposition is true | Analytical | High | Provenance boundary analysis | | Synthetic and deceptive are overlapping but distinct categories | Analytical | High | Research notes boundary analysis |

23. Comparative Analysis

Against photography

Photography captures light from a scene. Image generation synthesises pixels from learned representations and controls.

Against sound recording

Recording captures an acoustic event. Audio generation creates a waveform without that event.

Against animation

Animation is explicitly authored through drawing, modelling or procedural control. Generative video infers much of the content and motion from data.

Against retrieval

Retrieval returns an existing media object. Generation creates a new sample.

Against conventional editing

Conventional editing rearranges captured material. Generative editing can invent missing regions and frames.

Against speech synthesis

Speech synthesis is constrained by linguistic content and voice. General audio generation includes music, ambience and arbitrary sound.

Against provenance systems

Generation creates media. Provenance records signed claims about how media originated and changed.

Comparative principle

When media can be generated without capture, resemblance must stop carrying the whole burden of evidence. Trust moves from appearance towards provenance, corroboration and accountable production chains.

28. Final perspective

Generative image, audio and video models separate perceptual media from the physical act of capture. A plausible scene can be computed from a prompt, a reference image and random state. No lens, microphone or performer needs to have encountered the depicted event.

This is not the first time humans have represented fictional worlds. Painting, theatre, animation and special effects have always done so. The novelty is the automation of perceptual construction at scale and the ability to produce media with the surface grammar of photography and recording.

That surface grammar once carried evidentiary weight. A photograph could be staged or manipulated, but it normally began with light from a scene. A recording normally began with sound. Generative systems break that default connection. Verification must therefore move upstream and sideways: towards signed origin claims, retained generation records, independent sources, physical evidence and accountable publishers.

For Era VII, synthetic media makes machine-generated meaning visible and audible. The model does not only classify or describe the world. It manufactures a perceptual candidate world. That expands creativity and accessibility while giving deception a new wardrobe, a soundtrack and, increasingly, a perfectly plausible tracking shot.

Evidence

Sources and further reading

  1. Diederik P. Kingma and Max Welling. “Auto-Encoding Variational Bayes.” 2013. https://arxiv.org/abs/1312.6114

    Open source ↗

  2. Ian Goodfellow et al. “Generative Adversarial Nets.” 2014. https://arxiv.org/abs/1406.2661

    Open source ↗

  3. Leon A. Gatys, Alexander S. Ecker and Matthias Bethge. “A Neural Algorithm of Artistic Style.” 2015. https://arxiv.org/abs/1508.06576

    Open source ↗

  4. Aäron van den Oord et al. “WaveNet: A Generative Model for Raw Audio.” 2016. https://research.google/pubs/wavenet-a-generative-model-for-raw-audio/

    Open source ↗

  5. Jonathan Ho, Ajay Jain and Pieter Abbeel. “Denoising Diffusion Probabilistic Models.” 2020. https://arxiv.org/abs/2006.11239

    Open source ↗

  6. Alec Radford et al. “Learning Transferable Visual Models From Natural Language Supervision.” 2021. https://arxiv.org/abs/2103.00020

    Open source ↗

  7. Aditya Ramesh et al. “Zero-Shot Text-to-Image Generation.” 2021. https://arxiv.org/abs/2102.12092

    Open source ↗

  8. Robin Rombach et al. “High-Resolution Image Big-picture essays with Latent Diffusion Models.” 2022. https://arxiv.org/abs/2112.10752

    Open source ↗

  9. Zalán Borsos et al. “AudioLM: A Language Modeling Approach to Audio Generation.” 2022. https://arxiv.org/abs/2209.03143

    Open source ↗

  10. Andrea Agostinelli et al. “MusicLM: Generating Music From Text.” 2023. https://arxiv.org/abs/2301.11325

    Open source ↗

  11. Jonathan Ho et al. “Video Diffusion Models.” 2022. https://arxiv.org/abs/2204.03458

    Open source ↗

  12. Jonathan Ho et al. “Imagen Video: High Definition Video Generation with Diffusion Models.” 2022. https://arxiv.org/abs/2210.02303

    Open source ↗

  13. Uriel Singer et al. “Make-A-Video: Text-to-Video Generation without Text-Video Data.” 2022. https://arxiv.org/abs/2209.14792

    Open source ↗

  14. NIST. *Reducing Risks Posed by Synthetic Content*. 2024. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-4.pdf

    Open source ↗

  15. Coalition for Content Provenance and Authenticity. C2PA Technical Specification. https://c2pa.org/specifications/specifications/

    Open source ↗

  16. NIST. *Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile*. 2024. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf

    Open source ↗

  17. Yuezun Li and Siwei Lyu. “Exposing DeepFake Videos By Detecting Face Warping Artifacts.” 2018. https://arxiv.org/abs/1811.00656

    Open source ↗

  18. Andreas Rössler et al. “FaceForensics++: Learning to Detect Manipulated Facial Images.” 2019. https://arxiv.org/abs/1901.08971 Generative image, audio and video models separate perceptual media from the physical act of capture. A plausible scene can be computed from a prompt, a reference image and random state. No lens, microphone or performer needs to have encountered the depicted event. This is not the first time humans have represented fictional worlds. Painting, theatre, animation and special effects have always done so. The novelty is the automation of perceptual construction at scale and the ability to produce media with the surface grammar of photography and recording. That surface grammar once carried evidentiary weight. A photograph could be staged or manipulated, but it normally began with light from a scene. A recording normally began with sound. Generative systems break that default connection. Verification must therefore move upstream and sideways: towards signed origin claims, retained generation records, independent sources, physical evidence and accountable publishers. For Era VII, synthetic media makes machine-generated meaning visible and audible. The model does not only classify or describe the world. It manufactures a perceptual candidate world. That expands creativity and accessibility while giving deception a new wardrobe, a soundtrack and, increasingly, a perfectly plausible tracking shot.

    Open source ↗