SourceCited / Image & video SEO
Technical · Media
Crawlers can't see images or watch videos — they read the text around them. Describe media in alt text, captions, and transcripts, keep key facts in HTML, and use YouTube as earned media.
AI crawlers don't look at pixels or watch video, so media earns visibility through its text scaffolding: accurate alt text, descriptive filenames, captions, and — for video — an on-page transcript plus VideoObject schema. Never trap a fact inside an image. And treat YouTube as a top AI-citation surface: a genuine demo or explainer there is earned-media consensus as much as it is video SEO.
Non-rendering crawlers do not look at your images and cannot watch your videos — and a fact that lives only inside a picture is a fact AI cannot cite. The rule from the rendering guide applies with force here: if it matters, it must exist as text the crawler can read. Media still earns visibility, but through its text scaffolding — alt text, filenames, captions, transcripts, and schema — not the pixels themselves.
Prices, specs, or key numbers baked into an image or infographic. To a non-rendering bot they do not exist. Put every fact that matters in HTML text; use the image to illustrate, not to carry, the information.
| Element | Do this |
|---|---|
| Alt text | Describe the image accurately and specifically for what it shows — not keyword-stuffed; this is also your accessibility and agent-parsing signal |
| Filename | Descriptive, hyphenated (email-header-analysis.png), not IMG_4821.png |
| Caption & surrounding text | Engines read the words around an image to understand it — give it real context |
| Format & size | Modern formats (WebP/AVIF), compressed, correctly dimensioned — this feeds Core Web Vitals too |
| Image sitemap | List important images so they can be discovered and attributed |
| Original imagery | Your own screenshots and photos double as first-hand experience (E-E-A-T) and can carry provenance credentials |
Video matters for AI visibility in two distinct ways. On your own pages, a video needs the same text scaffolding — a transcript on the page (so the content is readable), a descriptive title and summary, and VideoObject schema. But the bigger lever is YouTube itself: it is among the strongest off-page signals for AI visibility and a top citation source inside AI Overviews, so a genuine demo, tutorial, or explainer there is earned-media consensus (Phase 3) as much as it is video SEO.
| Surface | What earns visibility |
|---|---|
| Video on your site | On-page transcript, descriptive title/summary, VideoObject schema, a video sitemap |
| YouTube | A real demo/tutorial/explainer, a keyword-honest title and description, an accurate transcript/captions, timestamps |
Whether on your site or YouTube, the transcript is what makes a video's content readable and quotable by AI. A great video with no transcript is, to a text crawler, a blank box.
Yes, but through text, not pixels. Non-rendering crawlers can't see images, so alt text, descriptive filenames, captions, and surrounding text are how an image earns visibility — and any fact that lives only inside an image is invisible and uncitable. Original images also double as first-hand experience for E-E-A-T.
Give the video readable text: an on-page transcript (so its content is quotable), a descriptive title and summary, VideoObject schema, and a video sitemap. The transcript is the key — without it, a text crawler sees a blank box.
Very — it's among the strongest off-page signals for AI visibility and a top citation source inside AI Overviews. A genuine demo, tutorial, or explainer with an accurate title, description, and transcript works as earned-media consensus, not just video SEO.
Not the important facts. Crawlers can't read text baked into an image, so prices, specs, and key numbers must also exist as HTML text. Use images to illustrate, and keep the citable information in the page text.