Search in 2026 is no longer built around typed keywords alone. People can point a camera at an object, upload a screenshot, circle an item on a phone screen or combine an image with a written question, and Google can use both the visual and verbal signals to work out what they mean. For authors, this changes the role of an image on a page. A picture is not merely decoration beside the copy; it can become part of the query itself and part of the evidence that helps a search system understand the page. The practical task is therefore simple: make the image, the nearby text and the purpose of the page tell the same story.
Google’s visual search has moved well beyond finding visually similar pictures. Lens and Circle to Search can identify several objects in one scene, while AI Mode can interpret an image together with a question and return a response that links back to relevant web sources. In May 2026, Google said AI Mode had passed one billion monthly users, and its new Search box could accept text, images, files, videos and Chrome tabs as inputs. That matters for editorial work because the searcher may arrive with an object, colour, chart, product detail or screenshot rather than a neatly phrased keyword.
This shift means that the text surrounding an image needs to do more than repeat what is visible. A search system can often recognise basic objects on its own, so the useful editorial layer is context: what the image shows, why it matters, how it relates to the subject of the page and what a reader should understand from it. A photograph of a damaged roof, for example, becomes far more informative when the nearby copy explains the type of damage, the likely cause, the location on the roof and the signs that distinguish it from ordinary wear.
The same principle applies to charts, diagrams, screenshots and product images. If a chart shows a change in organic traffic, the surrounding paragraph should name the period, the metric and the reason the change matters. If a screenshot shows a setting in a piece of software, the copy should explain what the setting controls and when a user would change it. The aim is not to write for a machine. It is to give the image enough human-readable context that both a reader and a search system can connect it with the topic of the page.
When planning copy, think about the questions a person could ask while looking at the image. A user photographing a plant may want its name, care advice or an explanation for yellow leaves. Someone circling a pair of shoes in a photograph may be looking for the model, material, price range or similar designs. Those are different intentions, even though the same object appears in the picture. The nearby text should therefore address the most likely intent that fits the purpose of the page instead of trying to mention every possible keyword.
A strong paragraph around an image usually contains three useful elements: identification, context and value. Identification says what the reader is looking at. Context explains where, when or under what conditions it applies. Value gives the point the reader actually needs, such as a comparison, warning, measurement, cause or next step. This is especially useful for editorial images that would otherwise be ambiguous. It also keeps the prose natural because the writer is answering a real question rather than assembling a string of search terms.
Authors should also resist the temptation to force a popular phrase into every caption or sentence near a picture. Google’s own guidance continues to stress people-first content and warns against keyword stuffing. In multimodal search, relevance is broader than exact wording: the page title, headings, main copy, image, caption and alternative text can all contribute context. Repetition does not create stronger meaning. Clear wording, accurate detail and a close relationship between the image and the surrounding section are more useful signals.
Start with placement. Google advises placing images near relevant text and on pages that genuinely match the subject of the image. For an author, that means the picture should sit beside the paragraph that explains it, not several screens away under a loosely related heading. The heading above that section should also describe the topic clearly. When the page has several images, each one should support a distinct part of the narrative, so a reader can understand why it is there without having to infer the connection.
Next, write a caption only when it adds useful information. A caption can name a person or place, explain a comparison, identify a date, state the source of a chart or point out a detail that is easy to miss. It should not simply restate the alt text or describe the obvious. For example, “Quarterly revenue by region, Q1–Q4 2025, source: company annual report” is useful because it adds scope and provenance. “Bar chart with bars” is technically descriptive but editorially weak.
The main paragraph should carry the deeper explanation. If the image is evidence for a claim, state what the evidence shows and where it comes from. If it is an original test, explain the conditions or method in plain language. If it is a comparison, tell the reader what changed and why the difference matters. This supports the trust principles emphasised in Google’s guidance: readers should be able to see who created the material, how the information was produced and why it is included.
Alternative text remains important in 2026 because it supports accessibility and also helps Google understand an image alongside computer vision and page content. Good alt text is concise but specific enough to communicate the image’s function. “Dalmatian puppy playing fetch” is more useful than “puppy”, while a decorative flourish may need empty alt text rather than a forced description. The right wording depends on context: the same photograph can require different alt text if it serves a different purpose on another page.
Filenames are a smaller signal, but they are still worth getting right when authors or editors control them. A short, descriptive name such as london-office-solar-panels.jpg is more meaningful than IMG_4821.jpg. On multilingual sites, filenames should be localised where practical instead of leaving every market with the source-language wording. None of this replaces good copy. A carefully named file attached to a thin or irrelevant page will not solve the underlying content problem.
Google also clarified in March 2026 that site owners can indicate a preferred image for a page through schema.org markup or the og:image tag. This is usually handled by an editor, developer or SEO specialist, but authors still influence the choice: the preferred image should be relevant, representative, high quality and not a generic logo. For articles that use structured data, Google recommends representative, crawlable images and supports multiple high-resolution aspect ratios, including 16:9, 4:3 and 1:1. The editorial decision comes first; the metadata simply helps describe that decision consistently.

One of the most useful changes for publishers arrived on 24 September 2026, when Google announced multimodal search reporting in Search Console. The reporting covers searches made through Lens, Circle to Search on Android, image uploads to Google Search and Chrome’s right-click image search. Teams can now separate this type of visibility from ordinary text-led search more clearly. Authors do not need to become analysts, but they should know which pages and visual assets are actually being surfaced through image-led queries.
Use that data to look for patterns rather than chasing single impressions. A page that receives multimodal visits may have a particularly clear product photo, useful diagram, original chart or tightly written explanation beside an image. Compare it with similar pages that receive little visual search traffic. Check whether the successful page has better image placement, more specific captions, stronger topical focus or more original visual evidence. The goal is to identify repeatable editorial habits, not to rewrite every page around a new metric.
Quality still comes before optimisation. Google’s people-first guidance remains relevant because multimodal systems need trustworthy source material just as conventional search does. Original photographs, first-hand screenshots, properly sourced charts and accurate captions can strengthen a page when they genuinely help the reader. Stock imagery that adds no information, misleading composites or pictures that contradict the copy weaken the experience. If an image has been generated or materially altered, clear disclosure is sensible when a reader could otherwise misunderstand what is real.
A practical workflow can be built into normal editing. Before publication, ask whether each important image has a clear purpose, sits beside the relevant section and is explained in the body copy. Check that the caption adds context rather than repeating the obvious, that the alt text describes the image’s function, and that the filename is sensible if it can still be changed. For charts and screenshots, record the source, date or method where those details affect interpretation. These checks improve accessibility and reader confidence even before search visibility is considered.
After publication, revisit pages that earn meaningful impressions or clicks from visual and multimodal searches. If a page performs well, preserve the relationship between the image and its surrounding text when updating it. If performance is weak, first ask whether the image is actually useful for the query the page serves. Replacing an irrelevant picture with a more informative one can be more effective than adding extra keywords. Likewise, a strong image may need a clearer explanatory paragraph rather than a longer caption.
The central editorial rule for 2026 is that text and imagery should work as one unit. Search can interpret more of what appears inside a picture, but it still benefits from accurate language that explains meaning, evidence and purpose. Authors who write precise surrounding copy, choose images that genuinely support the subject and make their sources and methods clear are better prepared for image-led queries without turning every article into a technical SEO project. That approach also matches the wider direction of search: useful content first, with optimisation serving clarity rather than replacing it.