Short answer
A language model only predicts the next token, so AI media comes from one of three routes: the model hands the prompt to a separate image, voice or video model; it predicts the media itself as tokens that a decoder turns back into pixels or sound; or a diffusion model refines random noise, step by step, into a picture or clip. Video costs the most, because every second multiplies the work.
How does AI generate images, sound and video?
A language model such as ChatGPT, Gemini or Claude does exactly one thing: it predicts the next piece of text, called a token, one at a time. So every image, voice and video you get from an AI product has to come out of one of three routes built around that limit.
The first route is hand-off: the language model writes a detailed prompt and calls a separate image, voice or video model to do the work. The second is media as tokens: images and sound are turned into codes the model can predict just like words. The third is diffusion: start from random noise and sculpt it, a little at a time, into a picture. The diagram below sets the three side by side.
Real products mix them. A chat model often plans and writes the prompt while a diffusion model renders it, and a model that generates images natively may still use a diffusion-style decoder for the final pass.
- You: "a poster for our Songkran sale"
- The language model writes a detailed prompt
- A separate image, voice or video model renders it
- The result comes back in the chat
- An encoder turns images or audio into a sequence of codes
- The same transformer predicts the next code, just like the next word
- A decoder turns the codes back into pixels or a waveform
- Start from pure random noise
- Remove a little noise 20 to 50 times, steered by the prompt
- Decode the clean result to full resolution
What is the hand-off route?
In the hand-off route the language model never draws anything itself. It writes a detailed instruction and calls another tool or model to produce the media. You type "make a poster for our Songkran sale"; the language model turns that into a long prompt specifying colours, layout and style, sends it to an image model, and shows the result back in the chat.
The well-known example is ChatGPT calling DALL·E 3 in 2023, and most agent tools still work this way today. The strength is that you can swap in the best specialist model for each job, test cheaply, and control exactly what is sent where.
The weakness is that the language model never sees the pixels. When you follow up with "same poster, just move the logo left", the image model tends to redraw everything, and the details you already liked drift.
How does AI turn images and sound into tokens?
AI turns images and sound into tokens with a learned compressor. Its encoder squeezes the media into a short list of codes picked from a fixed codebook, and its decoder rebuilds the media from those codes. Once media is a list of codes, the same transformer that predicts words can predict image or audio codes next.
Audio uses compressors called neural audio codecs, such as Google's SoundStream and Meta's EnCodec. Both cut sound into short slices and give each slice a stack of codes: the first layer captures the broad shape, each further layer the detail still missing. The technique is called residual vector quantization, or RVQ. The real numbers for EnCodec at 6 kbps are 75 slices a second with 8 codes each, so 600 codes for one second of audio.
Google's AudioLM research uses two kinds of token: semantic tokens for what is said, and acoustic tokens for how it sounds, the voice and tone. The model predicts the first kind, then fills in the second. The diagram below follows an image and a sound from the encoder, through the transformer, to the decoder.
- Good at text inside images and long instructions, because the prompt and the picture sit in one sequence and the model reads them together.
- Good at follow-up edits, because the model sees the image codes it produced and can change one area without redrawing the rest.
- Slower, because it predicts one code at a time and an image takes thousands of codes, while a text reply takes hundreds.
- OpenAI describes GPT-4o's image generation (2025) as native and autoregressive but has not published the architecture. Outside analyses suggest it predicts image latents and finishes them with a diffusion-style decoder; that is inference, not information from OpenAI.
What is diffusion, and how does it make a picture out of noise?
Diffusion makes an image by starting from pure random noise and removing a little of it at a time, over 20 to 50 steps, with the prompt steering each step, until the intended picture is left. In training, the model sees real images mixed with noise in varying amounts and learns which part is noise, so at generation time it can run the process in reverse.
Newer models such as Stable Diffusion 3 use rectified flow: noise and image sit on one straight line, and the model learns the direction along it. A straighter path allows fewer steps.
To save compute, none of this runs on full-size pixels. A compressor first shrinks the image into a latent that is 8 times smaller on each side in Stable Diffusion, the denoising happens in that small space, and a decoder expands the result back to a full image at the end.
A diffusion model does not understand language on its own. It relies on text encoders to turn the prompt into numbers that steer it, and Stable Diffusion 3 uses three: CLIP-L, CLIP-G and T5-XXL. The diagram below shows the picture emerging from the noise, step by step.
When generating, the model walks left to right, removing noise a step at a time, with the text prompt setting the direction.
How does AI generate video, and how do Sora and Veo differ?
AI generates video with diffusion that runs across space and time at once. OpenAI's Sora technical report lays it out most clearly: compress the video into a latent, cut it into small blocks called spacetime patches that span a few pixels and a few frames, and let a diffusion transformer denoise all the patches together.
Because every patch can attend to every other one, a face in frame 1 can stay the same face in frame 200. That consistency is the hard part, and compute grows with width, height and length, which is why video is the most expensive medium of all. The diagram below shows a clip cut into blocks across space and time.
Google's Veo 3 added sound in the same pass: dialogue, effects and ambience generated together with the picture, so lips and footsteps line up. Veo 3's model card describes it as latent diffusion applied jointly to video and audio latents, though in less detail than the Sora report.
The practical difference today is that Sora is gone. OpenAI closed the Sora app and website on 26 April 2026 and shuts the Sora API on 24 September 2026. If you are starting video work now, look at options such as Veo 3.1 or Runway instead.
How does AI generate voice and music?
Most voice generation uses the token route: text or a sample voice goes in, the model predicts audio codec codes, and a decoder turns the codes back into a waveform. Microsoft's VALL-E research used EnCodec codes this way and could imitate a voice from a short sample.
Real-time voice modes in chat apps work speech-to-speech: they take your voice in as tokens and answer in audio tokens directly, without converting to text first, which keeps replies fast and preserves tone. The hard part is natural timing and emotion without adding delay.
Music is harder than speech, because it needs structure over minutes: verse, chorus, and the return of a melody. Research such as Google's MusicLM uses tokens, but most commercial music apps do not publish their architecture, so nobody outside can say for certain whether they use tokens, diffusion or both.
Which route does each medium use, and what makes it hard?
Images use either diffusion or tokens, speech mostly uses tokens, music is largely unpublished, and video uses a diffusion transformer. The table puts them side by side.
| Medium | Main route today | Examples | Hardest part |
|---|---|---|---|
| Image | Diffusion, or image tokens inside the language model | Stable Diffusion 3, GPT-4o image generation, Gemini image models | Hands, text inside the image, exact layouts |
| Voice | Audio codec tokens, then a decoder | VALL-E, AudioLM, real-time voice modes | Natural timing and emotion, with low delay |
| Music | Tokens or diffusion (most vendors don't publish) | MusicLM, commercial music apps | Structure over minutes: verse, chorus, return |
| Video | Diffusion transformer over spacetime patches | Sora (discontinued), Veo 3 and 3.1 | The same face and physics in every frame; cost |
Which AI image, video and voice tools are good, and what do they cost?
The popular tools come in two kinds. Pay-as-you-go APIs suit wiring generation into a system; monthly web plans suit a marketing team working by hand. This table includes only tools whose prices and commercial-use terms we could verify on the vendor's own pages on 23 September 2026. Prices are in US dollars as published, before tax.
On rights, check three things before real use. First, most free plans forbid commercial use. Second, some vendors set conditions by company size: Midjourney requires companies with more than USD 1 million in gross revenue a year to be on Pro or Mega. Third, every vendor makes you responsible for ensuring that what you put in and what comes out infringes no one's rights. Terms change, so read the current ones yourself before paying.
| Tool | Used for | Published price (USD) | Commercial use |
|---|---|---|---|
| OpenAI GPT Image 2 (API) | Images | About 0.006 per 1024×1024 image at low quality, 0.053 at medium and 0.211 at high | OpenAI's Services Agreement says the customer owns the output |
| Google Gemini 3.1 Flash Image (API) | Images | 0.067 per 1K image and 0.151 per 4K image on the paid tier | Google does not claim ownership; content sent through the free tier may be used to improve its products |
| Midjourney | Images and short video | Basic 10, Standard 30, Pro 60, Mega 120 per month; 20% off billed annually | Commercial use on all plans; companies over USD 1M gross revenue a year need Pro or Mega, the only plans with Stealth Mode |
| Google Veo 3.1 (API) | Video with audio | Per second: Standard 0.40 (720p and 1080p), Fast 0.12 (1080p), Lite 0.08 (1080p) | Same terms as the Gemini API |
| Runway | Video and images | Standard 15 per month (625 credits, about 52 s of Gen-4.5), Pro 35 (2,250 credits), Max 95; less billed annually | Terms place no restriction on commercial use of outputs, subject to the agreement; the free plan carries a watermark |
| ElevenLabs | Voice, sound effects and music | Starter 6, Creator 22, Pro 99 per month for 30,000, 121,000 and 600,000 credits; text to speech uses about 1 credit per character | Commercial licence from Starter up; not on the free plan |
| Suno | Music | Pro 8 per month (20 song downloads), Premier 24 per month (60), billed annually | Commercial use on Pro and Premier only; the free plan is personal use |
Which tool fits which business? Four example scenarios
These four are composite scenarios drawn from problems Thai businesses commonly bring, not real clients. The figures come from the price table above and already allow for regenerating, because real work rarely gets a usable result on the first try.
| Composite example | Route and tools that fit | Rough cost |
|---|---|---|
| A beauty clinic making 30 promotional images a month for social media | Diffusion or tokens via Midjourney or GPT Image 2. If the image needs Thai text, test a native generator first. Avoid images that look like a real patient's treatment results. | GPT Image 2 at medium, 90 images to pick 30: about USD 5. Or Midjourney Standard at USD 30 a month (Pro at USD 60 if revenue exceeds USD 1M a year) |
| An online shop making 20 product clips of 8 seconds each | Diffusion transformer via Veo 3.1 or Runway. Start from a real product photo and let the AI move the camera, rather than having it draw the product. | Veo 3.1 Fast at 1080p, 160 seconds: USD 19.20 a round, about USD 58 allowing three rounds. Or Runway Pro at USD 35 a month for about 187 seconds of Gen-4.5 |
| A LINE customer assistant that can also answer by voice | Hand-off: a language model writes the reply and ElevenLabs reads it aloud. For live phone conversation, look at speech-to-speech modes that use audio tokens. | ElevenLabs Creator at USD 22 a month reads about 121,000 characters; the language model is billed separately |
| A 10-minute onboarding video for new staff, with narration | Record real screens or slides, narrate with ElevenLabs, and add short Veo 3.1 Lite clips only for scenes you cannot film. | ElevenLabs Creator at USD 22 a month, plus ten 6-second inserts at 1080p for about USD 4.80 a round |
What should a business think about before using AI for media?
Cost follows the token count or the number of passes. A text reply is a few hundred tokens, an image is thousands of codes or dozens of denoising passes, and video is all of that multiplied by the seconds. Budget most for video, and price by the usable piece, not by each click of generate.
If you are wiring AI media into your own product or system, letting a language model plan and call a specialist model (hand-off) is the practical default: easy to swap, cheap to test, simple to control. But if the main job is repeated edits to the same image, a model that generates natively keeps untouched areas intact far better.
Before using it on anything customers see, check these risks.
- Brand consistency: colours, logos and product details often drift between images, so someone must review every piece before it goes out.
- Rights to faces and voices: a face or voice that identifies a person is personal data under Thailand's PDPA. Cloning an employee's voice or using a customer's photo needs a lawful basis, and written consent is the sensible default.
- Telling customers what is AI: YouTube and TikTok have rules requiring labels on realistic AI-generated content, and saying so plainly also protects customer trust.
- Some advertising, such as for clinics, food and medicines, has its own rules in Thailand, and AI-made images must follow them just as photographs do.
- What leaves the building: unreleased product photos or customers' voices should not go into free plans whose vendors may use the data to improve their products.
How can SyncEdge help with AI for media?
We start from the work your team actually does, not from a list of tools. We look at which image, voice or video work costs the most time or money, and assess whether AI is worth applying there at all. If hiring a photographer or changing the process is cheaper, we say so.
Where it is worth it, we train your team on its own work and data, and set PDPA-sound rules on which photos, voices and customer data may go into external tools and which may not. You finish with a written summary of where AI applies, where it does not, and the next steps. Sessions run on-site or online, at a fixed price agreed after scoping.
Check these before paying for an AI media tool
- Does the plan you are buying include commercial rights, and is there a condition based on company revenue?
- Are your outputs public by default, and which plan keeps them private?
- Does the vendor use what you upload to improve its own products?
- Have you tested it on your real work, including Thai text and Thai voice?
- Have you priced the usable piece, including the rounds you will regenerate?
- Do the monthly credits cover your volume, and do unused credits roll over?
- Do you have consent from the people whose faces or voices you use?
- Who reviews the output before it is published, and against what?
Frequently asked questions
How does AI image generation work?
Most tools use diffusion: they start from random noise and remove it over dozens of steps, steered by the prompt, until the intended image is left. The other approach has a language model predict image codes one by one, like words, and decode them back into pixels. The second is better at text inside images and at edits, but slower.
What is diffusion?
Diffusion is a way of generating images or video by removing noise step by step. In training, the model sees real images mixed with noise at many levels and learns which part is noise; to generate, it starts from pure noise and works backwards to a picture. Stable Diffusion, Veo and Sora are built on this idea.
What is the difference between Sora and Veo?
Both generate video with diffusion over blocks that span space and time. Sora's technical report explains the method in the most detail, while Veo 3 generates dialogue and sound effects in the same pass as the picture. The practical difference now is that OpenAI closed the Sora app in April 2026 and shuts the Sora API on 24 September 2026, while Veo 3.1 remains available through the Gemini API.
How much does AI video generation cost?
It is mostly priced per second. As of 23 September 2026, Google Veo 3.1 through the API costs USD 0.08 to 0.40 per second at 1080p with audio, depending on the Lite, Fast or Standard tier. Runway starts at USD 15 a month for credits worth about 52 seconds of Gen-4.5. Allow for regenerating, since a usable clip rarely comes from the first try.
Can we use AI-generated images or audio commercially?
Yes, on plans that allow it, but the terms differ by vendor. The free plans of ElevenLabs and Suno carry no commercial rights, Midjourney requires companies with more than USD 1 million in yearly revenue to use Pro or Mega, and every vendor makes you responsible if an output infringes someone's rights. Read the current terms before real use.
Does cloning an employee's voice or face with AI breach the PDPA?
A face or voice that identifies a person is personal data under Thailand's PDPA, so generating audio or images from it needs a lawful basis. The safe course is written consent that states what it will be used for and for how long, and a vendor that does not train on your data. For complex cases, ask a lawyer.