SEO
Google Search Console Can Now Track Multimodal Search: What It Means for SEO and AI Visibility
Search Console now separates searches that start with an image (Google Lens, Circle to Search, image uploads, Chrome's "Search this image") from typed queries. Here's what the new data can and can't tell you, and what it means for SEO, visual content and AI visibility.

When the search starts with a camera, not a keyword
A shopper in a café notices a ceramic vase on the next table. They like it, but they have no idea what it's called, who makes it or what to type. A few years ago, the next step would have been a guess in a search bar ("white speckled round vase") and a lot of scrolling. Now they can point their phone at it with Google Lens, circle it on screen with Circle to Search, or upload a photo they took earlier.
That search never produced a keyword in the usual sense. It can still end on a retailer's product page, a maker's website or a local shop's listing. Until this week, though, site owners had no dedicated view in Search Console that separated those image-led visits from everything else.
On September 24, 2026, Google changed that when it announced multimodal Search reporting in Search Console. This article covers what was actually announced, what the new data can and can't tell you, and why the more interesting story is what it says about how customers find businesses. Throughout, we separate what Google has documented from our own analysis and practical recommendations.
Google just made multimodal search visible to marketers
Here is what Google has confirmed. The announcement, published on the Google Search Central Blog on September 24, 2026, adds reporting for "web multimodal search" in two places: the Performance report for Search results and the Generative AI performance report. It was co-authored by the product manager leads for Google Lens and for Search Console. In our reading, that joint authorship points to a coordinated visual search and measurement effort rather than a minor interface tweak.
The mechanism is a search type filter. According to Search Console's Performance report documentation, the web search type is now split into two views: "Web: text-based", covering traditional queries typed into the Google search bar, and "Web: multimodal", covering web search results where an image was used as part of the search.
Google says the rollout is global, starting on the day of the announcement, and that the metrics appear in a property's Performance report once the site is receiving traffic from these searches. An empty multimodal view doesn't mean anything is broken. It may simply mean the site isn't yet receiving traffic from these searches. The data can also be exported like the rest of the Performance report.
- Announced: September 24, 2026, on the Google Search Central Blog.
- Where: the Performance report for Search results and the Generative AI performance report.
- How: the search type filter, where "Web: multimodal" now sits alongside "Web: text-based".
- Included: Google Lens, Circle to Search on Android, image uploads to Google Search, and Chrome's right-click "Search this image".
- Rollout: global, starting September 24, 2026.
- Availability: metrics appear only when a property receives traffic from multimodal searches.
What counts as multimodal search?
In Google's reporting, a multimodal search is a web search where an image is used as part of the search. The wording "as part of" is worth noticing. Our reading is that the image may be the entire query, or one element alongside other input. Google names four experiences that feed the new data, listed below with everyday examples.
Two clarifications help avoid misreading the report. First, this is not the same as the existing "Image" search type, which covers Google Images results. The multimodal filter covers web results reached from an image-led search. The two answer different questions: "How do my images perform in Google Images?" versus "Which of my pages do people reach when they search with an image?"
Second, Google's documentation describes a single multimodal search type. It doesn't document a breakdown by entry point, so as things stand you shouldn't expect to separate Lens traffic from Circle to Search or Chrome image searches inside the report.
- Google Lens: pointing a phone camera at a product, a plant, a landmark, a menu or a shop sign, and searching what the camera sees.
- Circle to Search on Android: circling or highlighting something on screen, such as a jacket in a social video or a dish in a friend's photo, without leaving the app.
- Image uploads to Google Search: searching with a saved photo or a screenshot, for example a product seen in an advert or a building from a property listing.
- Chrome's "Search this image": right-clicking an image on any web page to find where it comes from or what it shows.
Why this is more important than another Search Console filter
On the surface, this is a reporting update. Our view is that the filter matters less as a feature and more as a signal. Platforms tend to build dedicated measurement for behaviour they expect site owners to care about, and we read this launch as a sign that Google treats image-led search as a meaningful, ongoing route to websites rather than a novelty. Google didn't publish usage figures in this announcement, and we won't estimate any.
None of this means typed search is fading. Text queries remain central to how people use Google, and nothing in the announcement suggests otherwise. What's changing is the range of starting points. A search can now begin with a camera frame, a screenshot, an object on a shelf, a photo in a chat, or an image on someone else's website.
Each of those starting points carries context that a person might never manage to put into words: colour, shape, material, style, brand marks, setting. That is what makes this kind of contextual search different. A text query only works if the customer knows the vocabulary. An image-led search doesn't need them to. For a business, it means people who don't know your product's name, your category's terminology or even your brand can still reach you, provided your pages can be matched to what they're looking at.
There is also a measurement consequence. Visual discovery has long been easy to talk about and hard to verify from a site owner's side. With a dedicated segment, teams can start testing their assumptions about it with their own data instead of relying on intuition.
Search intent is becoming harder to represent with keywords alone
Most SEO workflows are built on a simple model: query, then results page, then click. Keyword research estimates demand, pages are aligned to queries, and Search Console reports which queries brought impressions and clicks. It works because the typed query is a readable stand-in for intent.
Multimodal discovery follows a different path: object, image or context, then the search system's interpretation of it, then results, then discovery. The intent is embedded in the image rather than stated. A photo of a sofa might mean "what is this?", "where can I buy it?", "is there a cheaper version?" or "does it come in another colour?". Same image, very different intentions.
The implication, in our analysis, is that keyword-only reporting becomes less complete. Keyword research tools model typed demand; they can't fully represent demand that begins with a photograph. A business could see flat keyword numbers while picking up discovery through image-led searches that no keyword report was ever going to capture.
That doesn't make keyword research obsolete. It makes it necessary but partial. The practical shift is from matching phrases towards making each page unambiguous about what it is: which product, which place, which service, for whom. That clarity serves both typed and visual searches.
What Search Console can, and cannot, tell you
Start with what's documented. In the Performance report for Search results, choosing "Web: multimodal" in the search type filter segments the report to web searches where an image was part of the search, separately from typed ones. The Performance report's standard metrics are clicks, impressions, click-through rate and average position. Google's documentation doesn't list any exceptions for the multimodal view, but confirm what your own property shows.
In the Generative AI performance report, the same filter applies. That report shows impressions in generative AI features on Google Search, currently AI Overviews and AI Mode, grouped by page, country, device and date. Per Google's help page for the report, the metric it shows is impressions.
Now for what isn't documented. Google's announcement and help pages don't describe a way to see the image a person searched with, or which of your images was matched. As noted above, they don't document a split by Lens, Circle to Search, uploads or Chrome. And they don't explain what, if anything, appears in the Queries dimension for multimodal searches. Check that in your own property rather than assuming. Some multimodal searches may contain no typed words at all, so there may be nothing resembling a traditional query to report.
In our view, that changes how the data should be analysed. For multimodal traffic, the page, not the query, becomes the most dependable unit of analysis. Instead of asking "which keywords brought these visits?", the more useful question is "which pages are image-led searches reaching, and what do those pages have in common?" Our suggested approach:
- Compare "Web: multimodal" with "Web: text-based" for the same pages to see which ones depend more on visual discovery.
- Use the Pages dimension to spot which templates (product, location, gallery, recipe, article) attract image-led impressions.
- Check the device split to see whether visual discovery is phone-led for your site, or whether desktop image searches in Chrome also play a role.
- Review countries if you serve more than one market.
- If the Search appearance tab shows data for the multimodal view in your property, check which result types are involved.
- Record a baseline now and read trends over months, not days. Low or no data for a small site isn't a verdict on its visual content.
Which businesses should pay the most attention?
Any website with pages that correspond to something a person could photograph or screenshot has a stake in this. The businesses most exposed are those whose products or places are visual and hard to describe precisely in words.
That doesn't leave other industries unaffected. An industrial supplier whose parts get photographed on a job site to find a replacement, or a B2B brand whose equipment appears at trade events, is just as open to image-led search. The deciding question is whether your offering can be seen, not which category you're in. The list below shows where the connection is most obvious:
- E-commerce and retail: products spotted in real life, in social posts or in someone's home. Furniture, fashion, homeware, beauty and electronics are obvious examples.
- Restaurants and cafés: dishes, storefronts and menus are natural starting points for a camera search.
- Hospitality and travel: hotels, resorts, landmarks and destinations recognised from photos shared online.
- Real estate: developments, buildings and neighbourhoods people see in person or in adverts.
- Automotive: vehicles, parts and accessories identified by sight.
- Local businesses with a distinctive physical presence: signage, shopfronts and branded vehicles.
- Publishers and brands with strong original imagery: recipes, interiors, architecture and travel content.
Why this may matter for businesses in Oman, the UAE and the GCC
To be clear about the facts first: Google's rollout is global, and nothing in the announcement is specific to Oman, the UAE or the wider GCC. We're also not aware of reliable public data on how widely visual search is used in the region, so we won't guess.
What we can say is that several of the region's most important sectors are strongly visual: hospitality and tourism, restaurants and cafés, retail and malls, real estate, and automotive. Think of a visitor photographing a restaurant's façade in Muscat, a guest screenshotting a hotel lobby in Dubai from a friend's story, or a buyer saving an image of a villa from a social media advert. Each of those is a plausible starting point for an image-led search.
There is also a language angle. An image has no language, but the page it leads to does. For bilingual businesses, that makes clear, well-structured content in both Arabic and English, with business details that match across both versions, more important, not less.
For local businesses, the same applies to Google Business Profile photos and listing details. Genuine, current images of the premises, products and team, alongside consistent name, address and category information, are already part of sound local SEO practice, and they give image-led searches something accurate to connect to.
Multimodal search changes the role of images in SEO
Many websites still treat images as decoration: stock photography, generic banners, text baked into graphics, filenames like IMG_4821.jpg. In an image-led search, the image and the page around it may be exactly what connects a customer's photo to your business. In our view, that turns visual assets into discovery assets.
Google didn't publish a Lens-specific optimisation checklist alongside this announcement. The most relevant official reference remains Google's image SEO best practices. Google describes them as common to its visual discovery features, such as Google Images and Discover. They aren't Lens-specific, but they are the closest official guidance. The points below combine that guidance with our own practical recommendations:
- Use original, high-quality images. Google notes that sharp images are more appealing to users. In our experience, original photos of your actual products, rooms, dishes and premises also represent what customers see far better than stock imagery.
- Place images near relevant text. Google says it extracts information about an image's subject from page content, including captions and image titles, and recommends putting images on pages relevant to what they show.
- Write alt text that describes the image. Google uses alt text alongside computer vision and page content, and warns against keyword stuffing. "Handmade speckled ceramic vase with olive branches" is better than a list of search terms.
- Use short, descriptive filenames. Google calls these "very light clues", but they cost nothing. Translate them for localised versions of a page.
- Embed important images with standard HTML img elements. Google doesn't index CSS background images.
- Help discovery with an image sitemap if important images might otherwise be missed, and make sure pages holding them are indexable and not blocked by robots.txt.
- Use supported formats (such as JPEG, PNG, WebP or AVIF) and responsive images with srcset or picture, always keeping a fallback src.
- Set explicit width and height so images don't cause layout shift, and compress them. Google points out that images are often the largest contributor to page size.
- Add structured data where it genuinely applies. For supported rich result types, Google requires the image property for eligibility in Google Images, and product markup can describe price and availability.
- Consider image metadata. Google Images can use structured data or IPTC photo metadata to show creator, credit and licensing details.
The connection between multimodal search and AI search visibility
The factual link is straightforward. Google put the multimodal filter inside the Generative AI performance report too, so site owners can see impressions in AI Overviews and AI Mode that came from image-led searches. Google is connecting the two in its own reporting.
Our broader reading is that visual search and generative AI search are parts of the same shift: search systems that interpret more context, handle more formats and increasingly compose answers instead of only listing links. The multimodal filter's presence in the Generative AI report indicates that an image-led search can lead to a generative AI feature, not only a list of results.
Google's generative AI optimization guide is explicit that these features are rooted in its core Search ranking and quality systems, and that SEO best practices remain relevant. It says that if you already follow its image SEO and video SEO guidance, you're already optimizing for generative AI search. It also states that structured data isn't required for generative AI search and that there's no special schema.org markup to add, while recommending structured data as part of an overall SEO strategy.
So there is no single "AI SEO ranking factor", and schema markup doesn't place a business in AI answers by itself. What tends to help, in our experience, is information that is machine-understandable and trustworthy: crawlable pages, clear descriptions of what a product, service or place is, consistent entity details, images that are clearly tied to that information, and credible independent evidence. We've explored the wider picture in our look at how AI is reshaping customer discovery.
We've also written about businesses ranking well but missing from AI answers. Multimodal search adds another layer to that gap: a business can have strong text visibility and still be poorly represented when the search begins with an image.
What businesses should audit now
The framework below is practical, not a promise. None of these steps guarantees visibility in Google Lens, AI Overviews, AI Mode or any other system. They improve the chances that your content can be found, understood and matched when someone searches with an image.
Several of the checks overlap with standard technical SEO work, which is good news: the effort isn't duplicated. The same groundwork serves typed search, visual search and AI features alike.
One caution on measurement: impressions and clicks from image-led searches only matter if they lead somewhere. As with any channel, more traffic isn't the same as more business, so connect what Search Console shows to enquiries, bookings or sales in your analytics.
- Visual asset quality: are your key products, rooms, dishes and premises shown in original, sharp, representative photos rather than stock or text-heavy graphics?
- Crawlability and indexability: are important images in standard img elements, on indexable pages, and not blocked by robots.txt?
- Image context: does each important image sit near text that explains it, with descriptive alt text, a sensible filename and a caption where useful?
- Product or service information: does each page state plainly what the item or service is, including model, material, size, price, availability or scope where relevant?
- Structured data: is relevant markup (Product, LocalBusiness, Organization, Article) valid, consistent with visible content, and including images where supported?
- Entity consistency: are your business name, category, address and contact details identical across the website, Google Business Profile, directories and social profiles?
- Local signals: is your Google Business Profile complete, with current and genuine photos? For retailers, is Merchant Center product data accurate and aligned with the site?
- Mobile experience: do image-heavy pages load and read well on a phone, where camera-based searches start?
- Performance: are images compressed, responsively sized, served in modern formats and given explicit dimensions, with Core Web Vitals in good shape?
- Search Console multimodal data: have you recorded a baseline of pages, devices, countries and CTR for "Web: multimodal" against "Web: text-based"?
- Generative AI visibility: which pages appear in the Generative AI performance report with the multimodal filter applied?
- Measurement over time: is there a monthly or quarterly review, with site changes annotated and linked to conversions in analytics?
From keyword visibility to discovery visibility
SEO isn't becoming irrelevant. Google's own guidance says the opposite, and nothing in this announcement suggests typed queries are going away. What is broadening is discovery. A business can now be found through a typed keyword, a conversational AI prompt, a photo, a screenshot or a circled object in a video.
We think a more useful goal than "ranking for keywords" is what we'd call discovery visibility. It's a working term, not an industry standard, and it means asking whether your information can be found, interpreted and trusted across search interfaces and formats, not only whether a page ranks for a phrase.
In practice, that pulls together work that often sits in separate teams: SEO, content, photography, product data, local listings and analytics. In our view, visual search favours businesses where all of those tell the same, accurate story.
The new filter won't answer every question. It does answer one that Search Console couldn't answer directly before: whether people who searched with an image ended up on your site. A sensible first step is a modest one. Open the Performance report, switch the search type to "Web: multimodal", note what's there, and start treating your images as part of how customers find you, not just how your pages look.
