Generative image, audio and video models produce perceptual media from learned distributions rather than requiring a conventional camera, microphone, instrument or manually constructed frame for every output. They can generate media from noise, text prompts, reference images, sketches, motion, audio, masks, depth maps, style examples or other conditioning signals. The resulting object may resemble a photograph, recording or filmed event even when no corresponding scene or performance occurred.
The topic grows from several technical lineages. Variational autoencoders learned latent spaces from which new samples could be drawn. Generative adversarial networks trained a generator against a discriminator and produced increasingly realistic images. Neural style transfer separated aspects of content and visual style. WaveNet modelled raw audio, while diffusion models learned to reverse a gradual noising process. Contrastive text-image models linked linguistic descriptions with visual representations, and latent diffusion moved generation into compressed feature spaces, making high-resolution text-conditioned image synthesis more practical. [S01-S08]
Audio generation developed beyond speech synthesis into music, sound effects and long-form acoustic continuation. Systems such as AudioLM modelled audio through discrete representations at several timescales, while MusicLM conditioned hierarchical audio generation on text. Video models added a temporal problem: objects, identities, lighting and motion must remain coherent across frames. Video diffusion systems and cascaded text-to-video models improved fidelity, duration and resolution, but temporal consistency and physical plausibility remain distinct from visual beauty. [S09-S13]
Synthetic media reduces the cost of visualisation, prototyping, dubbing, design, sound production and animation. It can help creators explore ideas that would be expensive or impossible to capture. It also destabilises a historical shortcut in human reasoning: photographic or acoustic resemblance has often been treated as evidence that something happened. A generated image is not a photograph merely because it looks photographic. A generated voice is not testimony. A generated video is not a record of an event.
The topic must therefore separate synthesis from capture, resemblance from identity, perceptual realism from physical truth, generation from editing and file metadata from trustworthy provenance. Detection alone is not a durable solution because generators and detectors co-evolve, media can be transformed and authentic files can lose metadata. Provenance systems such as C2PA offer a complementary strategy by carrying signed assertions about origin and edits, but provenance proves only what the signed chain supports. It does not make the depicted claim true. [S14-S16]
Synthetic media is also a labour and cultural system. Models are trained on the work, voices, faces and styles of many people. The output may compete with those contributors without credit or consent. The same system can support accessibility and localisation while enabling impersonation, non-consensual sexual imagery, fraud, propaganda and evidence pollution.
The big idea
Generative media models turn learned perceptual distributions into programmable images, sounds and moving scenes. Their defining achievement is media production without conventional capture or frame-by-frame construction. Their recurring danger is that perceptual realism can manufacture false evidence, identity and attribution while obscuring the human works and institutional choices from which the model learned.