Technical Guide

How Does AI Generate Images, Sound and Video? Diffusion, Sora, Veo and Choosing the Right Tools

Every promo image, product shot, voiceover or short clip an AI makes comes out of one of just three routes. This guide explains all three without assuming a technical background, shows where each is strong and weak, lists what the popular tools cost, and sets out what a business should check before paying.

Short answer

A language model only predicts the next token, so AI media comes from one of three routes: the model hands the prompt to a separate image, voice or video model; it predicts the media itself as tokens that a decoder turns back into pixels or sound; or a diffusion model refines random noise, step by step, into a picture or clip. Video costs the most, because every second multiplies the work.

How does AI generate images, sound and video?

A language model such as ChatGPT, Gemini or Claude does exactly one thing: it predicts the next piece of text, called a token, one at a time. So every image, voice and video you get from an AI product has to come out of one of three routes built around that limit.

The first route is hand-off: the language model writes a detailed prompt and calls a separate image, voice or video model to do the work. The second is media as tokens: images and sound are turned into codes the model can predict just like words. The third is diffusion: start from random noise and sculpt it, a little at a time, into a picture. The diagram below sets the three side by side.

Real products mix them. A chat model often plans and writes the prompt while a diffusion model renders it, and a model that generates images natively may still use a diffusion-style decoder for the final pass.

Three routes from a prompt to media
Route A · hand-offThe language model delegates
  1. You: "a poster for our Songkran sale"
  2. The language model writes a detailed prompt
  3. A separate image, voice or video model renders it
  4. The result comes back in the chat
Route B · tokensMedia becomes more "words"
  1. An encoder turns images or audio into a sequence of codes
  2. The same transformer predicts the next code, just like the next word
  3. A decoder turns the codes back into pixels or a waveform
Route C · denoisingSculpting the result out of noise
  1. Start from pure random noise
  2. Remove a little noise 20 to 50 times, steered by the prompt
  3. Decode the clean result to full resolution

What is the hand-off route?

In the hand-off route the language model never draws anything itself. It writes a detailed instruction and calls another tool or model to produce the media. You type "make a poster for our Songkran sale"; the language model turns that into a long prompt specifying colours, layout and style, sends it to an image model, and shows the result back in the chat.

The well-known example is ChatGPT calling DALL·E 3 in 2023, and most agent tools still work this way today. The strength is that you can swap in the best specialist model for each job, test cheaply, and control exactly what is sent where.

The weakness is that the language model never sees the pixels. When you follow up with "same poster, just move the logo left", the image model tends to redraw everything, and the details you already liked drift.

How does AI turn images and sound into tokens?

AI turns images and sound into tokens with a learned compressor. Its encoder squeezes the media into a short list of codes picked from a fixed codebook, and its decoder rebuilds the media from those codes. Once media is a list of codes, the same transformer that predicts words can predict image or audio codes next.

Audio uses compressors called neural audio codecs, such as Google's SoundStream and Meta's EnCodec. Both cut sound into short slices and give each slice a stack of codes: the first layer captures the broad shape, each further layer the detail still missing. The technique is called residual vector quantization, or RVQ. The real numbers for EnCodec at 6 kbps are 75 slices a second with 8 codes each, so 600 codes for one second of audio.

Google's AudioLM research uses two kinds of token: semantic tokens for what is said, and acoustic tokens for how it sounds, the voice and tone. The model predicts the first kind, then fills in the second. The diagram below follows an image and a sound from the encoder, through the transformer, to the decoder.

  • Good at text inside images and long instructions, because the prompt and the picture sit in one sequence and the model reads them together.
  • Good at follow-up edits, because the model sees the image codes it produced and can change one area without redrawing the rest.
  • Slower, because it predicts one code at a time and an image takes thousands of codes, while a text reply takes hundreds.
  • OpenAI describes GPT-4o's image generation (2025) as native and autoregressive but has not published the architecture. Outside analyses suggest it predicts image latents and finishes them with a diffusion-style decoder; that is inference, not information from OpenAI.
How an image or a sound becomes tokens
InputImageA photo, 256 × 256 pixels
EncoderChop and compressEach small block is replaced by its nearest codebook entry
One transformerPredict the next codeText and image codes share one sequence, so the prompt and the picture are read together
DecoderRebuild pixelsCodes back to pixels, often finished with a diffusion-style pass
InputAudioSpeech or music, 24,000 samples a second
Neural codecLayered codesEvery 1/75 s gets a stack of codes: coarse on top, finer detail below
One transformerPredict the next codeOften two passes: what is said first, then how it sounds
DecoderRebuild the waveformCodes back to sound you can play

What is diffusion, and how does it make a picture out of noise?

Diffusion makes an image by starting from pure random noise and removing a little of it at a time, over 20 to 50 steps, with the prompt steering each step, until the intended picture is left. In training, the model sees real images mixed with noise in varying amounts and learns which part is noise, so at generation time it can run the process in reverse.

Newer models such as Stable Diffusion 3 use rectified flow: noise and image sit on one straight line, and the model learns the direction along it. A straighter path allows fewer steps.

To save compute, none of this runs on full-size pixels. A compressor first shrinks the image into a latent that is 8 times smaller on each side in Stable Diffusion, the denoising happens in that small space, and a decoder expands the result back to a full image at the end.

A diffusion model does not understand language on its own. It relies on text encoders to turn the prompt into numbers that steer it, and Stable Diffusion 3 uses three: CLIP-L, CLIP-G and T5-XXL. The diagram below shows the picture emerging from the noise, step by step.

The blend of noise and image that rectified-flow models train on
start: noisestart: noise
noise 75%noise 75%
noise 50%noise 50%
noise 25%noise 25%
donedone

When generating, the model walks left to right, removing noise a step at a time, with the text prompt setting the direction.

How does AI generate video, and how do Sora and Veo differ?

AI generates video with diffusion that runs across space and time at once. OpenAI's Sora technical report lays it out most clearly: compress the video into a latent, cut it into small blocks called spacetime patches that span a few pixels and a few frames, and let a diffusion transformer denoise all the patches together.

Because every patch can attend to every other one, a face in frame 1 can stay the same face in frame 200. That consistency is the hard part, and compute grows with width, height and length, which is why video is the most expensive medium of all. The diagram below shows a clip cut into blocks across space and time.

Google's Veo 3 added sound in the same pass: dialogue, effects and ambience generated together with the picture, so lips and footsteps line up. Veo 3's model card describes it as latent diffusion applied jointly to video and audio latents, though in less detail than the Sora report.

The practical difference today is that Sora is gone. OpenAI closed the Sora app and website on 26 April 2026 and shuts the Sora API on 24 September 2026. If you are starting video work now, look at options such as Veo 3.1 or Runway instead.

Video cut into small blocks across space and time
frame 1timeaudio generated in the same pass (Veo 3)one spacetime patcha few pixels × a few frames

How does AI generate voice and music?

Most voice generation uses the token route: text or a sample voice goes in, the model predicts audio codec codes, and a decoder turns the codes back into a waveform. Microsoft's VALL-E research used EnCodec codes this way and could imitate a voice from a short sample.

Real-time voice modes in chat apps work speech-to-speech: they take your voice in as tokens and answer in audio tokens directly, without converting to text first, which keeps replies fast and preserves tone. The hard part is natural timing and emotion without adding delay.

Music is harder than speech, because it needs structure over minutes: verse, chorus, and the return of a melody. Research such as Google's MusicLM uses tokens, but most commercial music apps do not publish their architecture, so nobody outside can say for certain whether they use tokens, diffusion or both.

Which route does each medium use, and what makes it hard?

Images use either diffusion or tokens, speech mostly uses tokens, music is largely unpublished, and video uses a diffusion transformer. The table puts them side by side.

Each medium, its main route and its hardest part, as of September 2026
MediumMain route todayExamplesHardest part
ImageDiffusion, or image tokens inside the language modelStable Diffusion 3, GPT-4o image generation, Gemini image modelsHands, text inside the image, exact layouts
VoiceAudio codec tokens, then a decoderVALL-E, AudioLM, real-time voice modesNatural timing and emotion, with low delay
MusicTokens or diffusion (most vendors don't publish)MusicLM, commercial music appsStructure over minutes: verse, chorus, return
VideoDiffusion transformer over spacetime patchesSora (discontinued), Veo 3 and 3.1The same face and physics in every frame; cost

Which AI image, video and voice tools are good, and what do they cost?

The popular tools come in two kinds. Pay-as-you-go APIs suit wiring generation into a system; monthly web plans suit a marketing team working by hand. This table includes only tools whose prices and commercial-use terms we could verify on the vendor's own pages on 23 September 2026. Prices are in US dollars as published, before tax.

On rights, check three things before real use. First, most free plans forbid commercial use. Second, some vendors set conditions by company size: Midjourney requires companies with more than USD 1 million in gross revenue a year to be on Pro or Mega. Third, every vendor makes you responsible for ensuring that what you put in and what comes out infringes no one's rights. Terms change, so read the current ones yourself before paying.

Prices and commercial-use rights, checked on the vendors' pricing and terms pages on 23 September 2026
ToolUsed forPublished price (USD)Commercial use
OpenAI GPT Image 2 (API)ImagesAbout 0.006 per 1024×1024 image at low quality, 0.053 at medium and 0.211 at highOpenAI's Services Agreement says the customer owns the output
Google Gemini 3.1 Flash Image (API)Images0.067 per 1K image and 0.151 per 4K image on the paid tierGoogle does not claim ownership; content sent through the free tier may be used to improve its products
MidjourneyImages and short videoBasic 10, Standard 30, Pro 60, Mega 120 per month; 20% off billed annuallyCommercial use on all plans; companies over USD 1M gross revenue a year need Pro or Mega, the only plans with Stealth Mode
Google Veo 3.1 (API)Video with audioPer second: Standard 0.40 (720p and 1080p), Fast 0.12 (1080p), Lite 0.08 (1080p)Same terms as the Gemini API
RunwayVideo and imagesStandard 15 per month (625 credits, about 52 s of Gen-4.5), Pro 35 (2,250 credits), Max 95; less billed annuallyTerms place no restriction on commercial use of outputs, subject to the agreement; the free plan carries a watermark
ElevenLabsVoice, sound effects and musicStarter 6, Creator 22, Pro 99 per month for 30,000, 121,000 and 600,000 credits; text to speech uses about 1 credit per characterCommercial licence from Starter up; not on the free plan
SunoMusicPro 8 per month (20 song downloads), Premier 24 per month (60), billed annuallyCommercial use on Pro and Premier only; the free plan is personal use

Which tool fits which business? Four example scenarios

These four are composite scenarios drawn from problems Thai businesses commonly bring, not real clients. The figures come from the price table above and already allow for regenerating, because real work rarely gets a usable result on the first try.

Composite examples, not real clients; costs from published prices as of 23 September 2026
Composite exampleRoute and tools that fitRough cost
A beauty clinic making 30 promotional images a month for social mediaDiffusion or tokens via Midjourney or GPT Image 2. If the image needs Thai text, test a native generator first. Avoid images that look like a real patient's treatment results.GPT Image 2 at medium, 90 images to pick 30: about USD 5. Or Midjourney Standard at USD 30 a month (Pro at USD 60 if revenue exceeds USD 1M a year)
An online shop making 20 product clips of 8 seconds eachDiffusion transformer via Veo 3.1 or Runway. Start from a real product photo and let the AI move the camera, rather than having it draw the product.Veo 3.1 Fast at 1080p, 160 seconds: USD 19.20 a round, about USD 58 allowing three rounds. Or Runway Pro at USD 35 a month for about 187 seconds of Gen-4.5
A LINE customer assistant that can also answer by voiceHand-off: a language model writes the reply and ElevenLabs reads it aloud. For live phone conversation, look at speech-to-speech modes that use audio tokens.ElevenLabs Creator at USD 22 a month reads about 121,000 characters; the language model is billed separately
A 10-minute onboarding video for new staff, with narrationRecord real screens or slides, narrate with ElevenLabs, and add short Veo 3.1 Lite clips only for scenes you cannot film.ElevenLabs Creator at USD 22 a month, plus ten 6-second inserts at 1080p for about USD 4.80 a round

What should a business think about before using AI for media?

Cost follows the token count or the number of passes. A text reply is a few hundred tokens, an image is thousands of codes or dozens of denoising passes, and video is all of that multiplied by the seconds. Budget most for video, and price by the usable piece, not by each click of generate.

If you are wiring AI media into your own product or system, letting a language model plan and call a specialist model (hand-off) is the practical default: easy to swap, cheap to test, simple to control. But if the main job is repeated edits to the same image, a model that generates natively keeps untouched areas intact far better.

Before using it on anything customers see, check these risks.

  • Brand consistency: colours, logos and product details often drift between images, so someone must review every piece before it goes out.
  • Rights to faces and voices: a face or voice that identifies a person is personal data under Thailand's PDPA. Cloning an employee's voice or using a customer's photo needs a lawful basis, and written consent is the sensible default.
  • Telling customers what is AI: YouTube and TikTok have rules requiring labels on realistic AI-generated content, and saying so plainly also protects customer trust.
  • Some advertising, such as for clinics, food and medicines, has its own rules in Thailand, and AI-made images must follow them just as photographs do.
  • What leaves the building: unreleased product photos or customers' voices should not go into free plans whose vendors may use the data to improve their products.

How can SyncEdge help with AI for media?

We start from the work your team actually does, not from a list of tools. We look at which image, voice or video work costs the most time or money, and assess whether AI is worth applying there at all. If hiring a photographer or changing the process is cheaper, we say so.

Where it is worth it, we train your team on its own work and data, and set PDPA-sound rules on which photos, voices and customer data may go into external tools and which may not. You finish with a written summary of where AI applies, where it does not, and the next steps. Sessions run on-site or online, at a fixed price agreed after scoping.

Check these before paying for an AI media tool

  • Does the plan you are buying include commercial rights, and is there a condition based on company revenue?
  • Are your outputs public by default, and which plan keeps them private?
  • Does the vendor use what you upload to improve its own products?
  • Have you tested it on your real work, including Thai text and Thai voice?
  • Have you priced the usable piece, including the rounds you will regenerate?
  • Do the monthly credits cover your volume, and do unused credits roll over?
  • Do you have consent from the people whose faces or voices you use?
  • Who reviews the output before it is published, and against what?

Frequently asked questions

How does AI image generation work?

Most tools use diffusion: they start from random noise and remove it over dozens of steps, steered by the prompt, until the intended image is left. The other approach has a language model predict image codes one by one, like words, and decode them back into pixels. The second is better at text inside images and at edits, but slower.

What is diffusion?

Diffusion is a way of generating images or video by removing noise step by step. In training, the model sees real images mixed with noise at many levels and learns which part is noise; to generate, it starts from pure noise and works backwards to a picture. Stable Diffusion, Veo and Sora are built on this idea.

What is the difference between Sora and Veo?

Both generate video with diffusion over blocks that span space and time. Sora's technical report explains the method in the most detail, while Veo 3 generates dialogue and sound effects in the same pass as the picture. The practical difference now is that OpenAI closed the Sora app in April 2026 and shuts the Sora API on 24 September 2026, while Veo 3.1 remains available through the Gemini API.

How much does AI video generation cost?

It is mostly priced per second. As of 23 September 2026, Google Veo 3.1 through the API costs USD 0.08 to 0.40 per second at 1080p with audio, depending on the Lite, Fast or Standard tier. Runway starts at USD 15 a month for credits worth about 52 seconds of Gen-4.5. Allow for regenerating, since a usable clip rarely comes from the first try.

Can we use AI-generated images or audio commercially?

Yes, on plans that allow it, but the terms differ by vendor. The free plans of ElevenLabs and Suno carry no commercial rights, Midjourney requires companies with more than USD 1 million in yearly revenue to use Pro or Mega, and every vendor makes you responsible if an output infringes someone's rights. Read the current terms before real use.

Does cloning an employee's voice or face with AI breach the PDPA?

A face or voice that identifies a person is personal data under Thailand's PDPA, so generating audio or images from it needs a lawful basis. The safe course is written consent that states what it will be used for and for how long, and a vendor that does not train on your data. For complex cases, ask a lawyer.