Faceted Navigation and Crawl Budget: What Actually Stops Filters From Wasting Googlebot’s Time
Faceted navigation and crawl budget collide when every filter click creates a new crawlable URL. Googlebot can’t tell a useful filter page from a useless one before fetching it, so it burns requests on filter combinations while new products wait. The fix is stopping the crawl, not the indexing alone.
- What really saves crawl budget
- What doesn’t
- How to fix it
Google doesn’t hide how common this is. In its 2025 year-end review of crawling problems, faceted navigation made up about half of the issues and action parameters such as add-to-cart another quarter, as Gary Illyes explained on Search Off the Record. No other URL pattern comes close.
No other URL pattern comes close.
I’ve been auditing online stores since 2015, and this topic attracts plenty of confident advice that doesn’t match Google’s own documentation. Below, I separate what saves crawl budget from what only tidies the index, then cover the platform quirks in Shopify and Magento.
Magento filters, sort options and toolbar parameters can turn a few hundred categories into millions of crawlable URLs. I audit where Googlebot spends its time, decide which filters deserve to rank, and hand over the exact fixes.
Get Magento SEO HelpHow Faceted Navigation Turns One Category Into Thousands of URLs
Take one category with five colours, six sizes, four price bands and three extra sort options. If shoppers can pick one value per filter, that page can produce 840 URL variations. Allow multi-select on colour, size and price, and it can generate over 131,000.
- 5colours
- ×
- 6sizes
- ×
- 4price bands
- ×
- 3extra sort options
Parameter order then doubles the damage. If ?color=red&size=10 and ?size=10&color=red both load, Google sees two URLs for one product grid. Add tracking tags, session IDs or add-to-cart parameters, and the URL space has no practical ceiling at all.
?color=red&size=10?size=10&color=red- Tracking tags
- Session IDs
- Add-to-cart parameters
The real problem is how crawlers learn. Google’s faceted navigation documentation explains that filter URLs look new, so crawlers usually fetch a very large number of them before deciding they’re useless. By then your server has paid for every request, and discovery of new pages has slowed.
Botify documented this on a shoe retailer where faceted category pages made up almost 90% of the site. Out of roughly 427,000 facet URLs, only 403 received organic visits. The study dates from 2020, but the mechanics behind that ratio haven’t changed.
The study dates from 2020, but the mechanics behind that ratio haven’t changed.
Does Your Store Actually Have a Crawl Budget Problem?
Most small stores don’t. Google’s crawl budget guide is aimed at sites with over a million unique pages changing weekly, sites with 10,000 or more pages changing daily, and sites with a large share of URLs stuck in Discovered, currently not indexed. Google calls these rough estimates, not hard thresholds.
That third group is where filters catch mid-sized catalogues. A store with 3,000 products can still expose hundreds of thousands of filter URLs. If your product count looks modest but Search Console shows a growing pile of discovered, unindexed URLs, the filters are the first place to look.
- 1Sites with over a million unique pages changing weekly
- 2Sites with 10,000 or more pages changing daily
- 3Sites with a large share of URLs stuck in Discovered, currently not indexed
That third group is where filters catch mid-sized catalogues. A store with 3,000 products can still expose hundreds of thousands of filter URLs.
What to Check in Search Console
Open Settings, then Crawl Stats. Look at total crawl requests over time, the split by response code, and the sample URLs Google lists. If a big share of the sampled requests carry filter or sort parameters, you have your answer before paying for any crawling tool.
Then open the Page indexing report. Filter URLs sitting under Crawled, currently not indexed or Duplicate without user-selected canonical show that Google fetched them and threw them away. That’s crawl spent for nothing. Your XML sitemap should never list filter URLs you don’t want ranked.
- Total crawl requests over time
- Split by response code
- Sample URLs Google lists
- Crawled, currently not indexed
- Duplicate without user-selected canonical
Your XML sitemap should never list filter URLs you don’t want ranked.
Why Server Logs Settle the Argument
Crawl Stats gives you totals but only a sample of URLs. Server logs show every request, so you can group Googlebot hits by parameter name and compare them with hits on product pages. When a single sort parameter takes more requests than the whole product catalogue, the priority list writes itself.
Verify the bot before trusting those numbers. Plenty of scrapers fake the Googlebot user agent. Google publishes its crawler IP ranges and a reverse DNS check, so filter out anything that fails verification before drawing conclusions about crawl waste.
- 1Google’s crawler IP ranges
- 2Reverse DNS check
- 3Filter out anything that fails
What Actually Saves Crawl Budget, and What Only Looks Like It
This is where most advice on faceted navigation SEO goes wrong. Several tactics that clean up the index do nothing for crawl budget, because Google still has to fetch the page to read the instruction. Here’s how each control behaves, based on Google’s own documentation.
| Control | Stops crawling? | Stops indexing? | Best use |
|---|---|---|---|
| robots.txt disallow | Yes | MostlyA blocked URL with links pointing to it can still appear, without content | Filter and sort URLs you never want in search |
| Filters in URL fragments (#) | YesNo new crawlable URL exists | Yes | New builds and JavaScript filtering |
| 404 for empty combinations | Reduces itGoogle recrawls 404s less over time | Yes | Filter combinations with zero products |
| rel=”canonical” | NoMay reduce crawling slowly over time | Consolidates signalsIt is a hint | Close variants you want merged into the parent |
| noindex | NoThe page must be crawled to be read | Yes | Removing filter URLs that are already indexed |
| rel=”nofollow” on filter links | Only if every link to the URL uses it | No | A backup measure, never the main control |
Robots.txt Is the Real Crawl Control
Google recommends robots.txt disallow rules when you don’t need filter URLs in search. Its guidance says there’s often no good reason to allow crawling of filtered items, and suggests keeping only product pages and one unfiltered listing page crawlable. Google’s own example looks like this:
user-agent: Googlebot
disallow: /*?*products=
disallow: /*?*color=
disallow: /*?*size=
allow: /*?products=all$
One catch matters here. A disallowed URL can still get indexed if other pages link to it, just without its content. And Googlebot never reads a canonical or noindex on a page it can’t fetch. So give each URL pattern one job, and never stack controls that cancel each other.
Don’t Create the URL at All
The cleanest fix is a filter that never produces a crawlable address. Google generally ignores URL fragments, so filters written as #color=red have no effect on crawling, positive or negative. Client-side filtering that updates the grid without adding new links does the same job.
The trade-off is real, though. A filter state behind a fragment can’t rank on its own. So keep fragment or JavaScript filtering for sort orders and low-value combinations, and give the few high-demand filters proper static URLs that search engines can crawl.
- Sort orders
- Low-value combinations
A filter state behind a fragment can’t rank on its own
- The few high-demand filters
- Pages search engines can crawl
Canonical Tags Consolidate, They Don’t Block
Pointing filter pages at the parent category with a canonical is still worth doing, because it merges link signals. But Google says a canonical may decrease crawling of non-canonical URLs only over time, and rates it less effective in the long run than robots.txt or fragments.
Noindex Is Not a Crawl Budget Tool
This is the most repeated myth on the subject. Google’s crawl budget guide tells site owners not to use noindex for this purpose, because Google still requests the page and only drops it after seeing the tag. The crawl has already happened. Noindex controls the index, nothing more.
Noindex still has one useful job. If filter URLs are already indexed, add noindex first, let Google recrawl and drop them, then add the robots.txt block. Reverse that order and the URLs stay stuck in the index, because Google can no longer read the tag.
- 1Add noindex first
- 2Let Google recrawl and drop them
- 3Then add the robots.txt block
Reverse that order and the URLs stay stuck in the index, because Google can no longer read the tag.
Nofollow Needs Total Coverage
Nofollow on filter links can discourage crawling, but Google’s Crawling December post on faceted navigation stresses that every link to those URLs, internal and external, must carry it. One unmarked link in a footer, feed or email is enough for discovery, so treat nofollow as a backup only.
One unmarked link is enough for discovery, so treat nofollow as a backup only.
Return a 404 for Empty Filter Results
When a filter combination has no products, Google wants a real 404 status on that URL. Not a 200 page saying nothing was found, and not a redirect to a generic error page. The same rule covers duplicate filters, nonsense combinations and pagination past the last page.
This matters for crawl because Google’s crawl budget guide says soft 404 pages keep getting crawled, while a 404 is a strong signal not to crawl that URL again. An empty product grid returning 200 is a soft 404 waiting to happen.
- Duplicate filters
- Nonsense combinations
- Pagination past the last page
Blocking Filters Won’t Automatically Speed Up Your Other Pages
Here’s a nuance most guides skip. Google sets crawl budget from two things: crawl capacity, meaning how hard your server can be crawled, and crawl demand, meaning how much Google wants your content. Cutting filter crawl helps each of these in a different way.
If filter crawling overloads your server, blocking it frees capacity, and Google raises its limit as response times improve. But Google won’t shift freed budget to other pages unless your site already hits its capacity limit. On a healthy server, blocking cuts waste without boosting crawl elsewhere.
Blocking filter crawl frees capacity if it overloads your server. Google raises its limit as response times improve.
- Perceived inventory
- Popularity
- Staleness
Google won’t shift freed budget to other pages unless your site already hits its capacity limit.
So what raises demand? Google names three demand factors: perceived inventory, popularity and staleness. Perceived inventory is the one you control directly. Fewer junk URLs, more internal links to key categories and accurate lastmod dates in your sitemap are what pull crawling toward the pages that sell.
- Fewer junk URLs
- More internal links to key categories
- Accurate lastmod dates in your sitemap
Which Filter Pages Deserve to Be Crawled and Indexed
Blocking every filter is the opposite mistake. Some filters match real searches, like black leather boots or wide fit running shoes, and those pages can outrank the parent category. My test is simple: proven search demand, enough products to fill the page, and a stable URL.
Check demand with Search Console queries, keyword tools and your own site search data. If a filter passes, treat it as a landing page rather than a filter state. It’s the same category-level thinking I use in ecommerce SEO work, where each search intent gets one strong page.
If a filter passes, treat it as a landing page rather than a filter state.
For those pages, follow Google’s rules for crawlable facets. Use the standard & separator, keep filters in one fixed order, and never allow duplicate filters. Give each page a self-referencing canonical, a unique title and intro, a sitemap entry, and plain HTML links from the parent category.
- The standard & separator
- Filters in one fixed order
- Never allow duplicate filters
- A self-referencing canonical
- A unique title and intro
- A sitemap entry
- Plain HTML links from the parent category
Everything else needs a crawl decision, not only an index decision. Sort orders, view modes, price sliders, stacked filters and tracking parameters should be blocked or kept off crawlable links entirely. Single low-demand filters can stay crawlable with a canonical to the parent if their link signals are worth merging.
- Sort orders, view modes, price sliders, stacked filters, tracking parametersBlocked or kept off crawlable links entirely
- Single low-demand filtersCan stay crawlable with a canonical to the parent if their link signals are worth merging
Platform Notes for Shopify, Magento and Multi-Language Stores
Shopify
Shopify storefront filters add parameters such as filter.v.option.color and filter.p.vendor to collection URLs. The default robots.txt already blocks sort_by URLs, tag combinations joined with a plus sign, and collection URLs carrying two or more filter parameters. Check yours at yourstore.com/robots.txt.
The gap is single-filter URLs. One filter parameter stays crawlable by default, and many themes link every filter value as a plain anchor. On a large catalogue, that alone can mean thousands of URLs, which is a common thread in my Shopify SEO consulting audits.
sort_byURLs- Tag combinations joined with a plus sign
- Collection URLs carrying two or more filter parameters
- Single-filter URLs stay crawlable by default
- Many themes link every filter value as a plain anchor
filter.v.option.colorfilter.p.vendor
Shopify lets you edit robots.txt.liquid to add your own rules. Be careful here. A loose pattern can hide entire collections from Google, so test every new rule against real collection, filter and product URLs before publishing the theme change.
Magento and Adobe Commerce
Magento layered navigation builds query strings from attribute option IDs, so filter URLs look like ?color=49&size=170. The toolbar adds more: product_list_order, product_list_dir, product_list_limit and product_list_mode for sorting, direction, items per page and grid or list view, plus p for pagination.
?color=49&size=170Query strings from attribute option IDsproduct_list_orderSortingproduct_list_dirDirectionproduct_list_limitItems per pageproduct_list_modeGrid or list viewpPaginationTwo admin settings matter most. Adobe’s layered navigation documentation shows each filterable attribute is set to Filterable (with results) or Filterable (no results). The second option shows values with zero matching products, handing crawlers links to empty pages. I’d use with results unless there’s a strong UX reason not to.
Also enable Use Canonical Link Meta Tag for Categories, which Adobe recommends alongside the product canonical setting. It points filtered category states at the clean category URL. Then block the toolbar parameters in robots.txt, because a canonical alone won’t stop Googlebot fetching them.
Filterable (with results)I’d use with results unless there’s a strong UX reason not toFilterable (no results)Shows values with zero matching products, handing crawlers links to empty pagesUse Canonical Link Meta Tag for CategoriesPoints filtered category states at the clean category URLThen block the toolbar parameters in robots.txt, because a canonical alone won’t stop Googlebot fetching them.
Multi-Language Stores
Google treats each hostname as a separate site with its own crawl budget. A store on de.example.com and fr.example.com splits its budget by subdomain, while subfolder locales share one. Either way, every locale copies the full filter space, so multilingual SEO setups need facet rules applied per market.
de.example.comfr.example.comEither way, every locale copies the full filter space, so facet rules apply per market.
AI Crawlers Hit the Same Filter URLs
Googlebot isn’t the only bot walking your filters now. Cloudflare’s August 2025 analysis found training accounts for nearly 80% of AI bot crawling, and noted that some AI crawlers ignore robots.txt directives. Every open filter URL is available to all of them.
AI bots don’t draw from Google’s crawl budget, but they hit the same server. Google lowers its crawl capacity limit when response times rise or the server returns 5xx or 429 errors. If scrapers slow your store down on filter pages, Googlebot pays for it too.
- 1Scrapers slow your store down on filter pages
- 2Response times rise, or the server returns 5xx or 429 errors
- 3Google lowers its crawl capacity limit
AI bots don’t draw from Google’s crawl budget, but they hit the same server.
The practical answer is the same: shrink the crawlable filter space. Robots.txt handles well-behaved bots, and bot management or rate limiting at your CDN handles the rest. Keep product and category pages open if you want your store cited in AI answers.
- Robots.txt handles well-behaved bots
- Bot management or rate limiting at your CDN handles the rest
- Keep product and category pages open if you want your store cited in AI answers
Advice That’s Out of Date
Two pieces of old advice on faceted navigation and crawl budget still show up in guides published this year. The first is configuring parameter handling in Search Console. Google retired the URL Parameters tool in 2022, noting that only about 1% of its configurations were useful for crawling.
The second is relying on rel=”next” and rel=”prev” to manage paginated filter pages. Google confirmed in 2019 that it no longer uses them for indexing. They’re harmless markup, but they don’t control crawling. Link deep pages with normal anchors and return a 404 past the last page.
- 1
Configuring parameter handling in Search ConsoleGoogle retired the URL Parameters tool in 2022Only about 1% of its configurations were useful for crawling - 2
Relying on rel=”next” and rel=”prev” for paginated filter pagesGoogle confirmed in 2019 it no longer uses them for indexingLink deep pages with normal anchors and return a 404 past the last page
A Safe Order for Fixing Faceted Navigation and Crawl Budget Issues
Changing robots.txt on a live store without a plan is how categories vanish from Google. This is the sequence I follow, because each step protects the one after it and keeps every change easy to reverse if the data surprises you.
- 1Measure first. Pull at least 30 days of Crawl Stats and, if you can, server logs. Group requests by parameter so you know which filters take the most crawl.
- 2Classify every parameter. For each filter, sort and toolbar parameter, decide whether it becomes a landing page, gets consolidated with a canonical, or stays out of the crawl.
- 3Fix links before rules. Stop linking sort orders and low-value filters as plain anchors, and remove every filter URL from your XML sitemaps.
- 4Return real 404s. Serve a 404 status for empty combinations, duplicate filters and pagination beyond the last page.
- 5Deindex, then block. Add noindex to filter URLs that are already indexed, wait until they drop out, then add the robots.txt disallow rules.
- 6Test every pattern. Run each robots.txt rule against sample URLs, including different parameter orders, so you don’t block pages you want ranked.
- 7Monitor for at least two months. Track Crawl Stats, the Discovered, currently not indexed count and how fast new products get indexed.
If you’d like a second pair of eyes on the crawl data before touching robots.txt on a live store, a one-to-one SEO consultation is a cheap way to avoid blocking the wrong pages. It costs far less than recovering lost category rankings afterwards.
Faceted Navigation and Crawl Budget FAQs
Does faceted navigation hurt SEO?
Only when it creates crawlable URLs nobody searches for. Filters help shoppers find products faster. The damage comes from uncontrolled URL generation, which wastes crawling, creates near-duplicate pages and slows discovery of new products. Controlled facets with real demand can rank very well.
Should I use noindex or robots.txt for filter pages?
Use robots.txt when the goal is saving crawl budget, because noindex pages still get crawled. Use noindex only to remove filter URLs that are already indexed, then block them once they drop out. Applying both at once stops Google reading the noindex.
Do canonical tags fix crawl budget?
Not directly. Google still crawls canonicalised filter URLs, and it says canonicals may reduce that crawling only gradually. They’re useful for merging signals between near-duplicate pages, but they aren’t a substitute for blocking or never creating low-value URLs.
How many filter pages should I index?
Index the filters that have real search demand and enough products to be useful, and no more. For most stores that’s a short list per category. Remember that every indexed filter page also becomes a page you have to maintain and keep fresh.
Can a small store ignore crawl budget?
Usually, yes. Google says sites whose new pages get crawled the same day they’re published don’t need its crawl budget guide. The exception is a small catalogue whose filters expose huge numbers of URLs, which shows up as a growing Discovered, currently not indexed count.
On this page
Google doesn’t hide how common this is. In its 2025 year-end review of crawling problems, faceted navigation made up about half of the issues and action parameters such as add-to-cart another quarter, as Gary Illyes explained on Search Off the Record. No other URL pattern comes close.
No other URL pattern comes close.
I’ve been auditing online stores since 2015, and this topic attracts plenty of confident advice that doesn’t match Google’s own documentation. Below, I separate what saves crawl budget from what only tidies the index, then cover the platform quirks in Shopify and Magento.
Magento filters, sort options and toolbar parameters can turn a few hundred categories into millions of crawlable URLs. I audit where Googlebot spends its time, decide which filters deserve to rank, and hand over the exact fixes.
Get Magento SEO HelpHow Faceted Navigation Turns One Category Into Thousands of URLs
Take one category with five colours, six sizes, four price bands and three extra sort options. If shoppers can pick one value per filter, that page can produce 840 URL variations. Allow multi-select on colour, size and price, and it can generate over 131,000.
- 5colours
- ×
- 6sizes
- ×
- 4price bands
- ×
- 3extra sort options
Parameter order then doubles the damage. If ?color=red&size=10 and ?size=10&color=red both load, Google sees two URLs for one product grid. Add tracking tags, session IDs or add-to-cart parameters, and the URL space has no practical ceiling at all.
?color=red&size=10?size=10&color=red- Tracking tags
- Session IDs
- Add-to-cart parameters
The real problem is how crawlers learn. Google’s faceted navigation documentation explains that filter URLs look new, so crawlers usually fetch a very large number of them before deciding they’re useless. By then your server has paid for every request, and discovery of new pages has slowed.
Botify documented this on a shoe retailer where faceted category pages made up almost 90% of the site. Out of roughly 427,000 facet URLs, only 403 received organic visits. The study dates from 2020, but the mechanics behind that ratio haven’t changed.
The study dates from 2020, but the mechanics behind that ratio haven’t changed.
Does Your Store Actually Have a Crawl Budget Problem?
Most small stores don’t. Google’s crawl budget guide is aimed at sites with over a million unique pages changing weekly, sites with 10,000 or more pages changing daily, and sites with a large share of URLs stuck in Discovered, currently not indexed. Google calls these rough estimates, not hard thresholds.
That third group is where filters catch mid-sized catalogues. A store with 3,000 products can still expose hundreds of thousands of filter URLs. If your product count looks modest but Search Console shows a growing pile of discovered, unindexed URLs, the filters are the first place to look.
- 1Sites with over a million unique pages changing weekly
- 2Sites with 10,000 or more pages changing daily
- 3Sites with a large share of URLs stuck in Discovered, currently not indexed
That third group is where filters catch mid-sized catalogues. A store with 3,000 products can still expose hundreds of thousands of filter URLs.
What to Check in Search Console
Open Settings, then Crawl Stats. Look at total crawl requests over time, the split by response code, and the sample URLs Google lists. If a big share of the sampled requests carry filter or sort parameters, you have your answer before paying for any crawling tool.
Then open the Page indexing report. Filter URLs sitting under Crawled, currently not indexed or Duplicate without user-selected canonical show that Google fetched them and threw them away. That’s crawl spent for nothing. Your XML sitemap should never list filter URLs you don’t want ranked.
- Total crawl requests over time
- Split by response code
- Sample URLs Google lists
- Crawled, currently not indexed
- Duplicate without user-selected canonical
Your XML sitemap should never list filter URLs you don’t want ranked.
Why Server Logs Settle the Argument
Crawl Stats gives you totals but only a sample of URLs. Server logs show every request, so you can group Googlebot hits by parameter name and compare them with hits on product pages. When a single sort parameter takes more requests than the whole product catalogue, the priority list writes itself.
Verify the bot before trusting those numbers. Plenty of scrapers fake the Googlebot user agent. Google publishes its crawler IP ranges and a reverse DNS check, so filter out anything that fails verification before drawing conclusions about crawl waste.
- 1Google’s crawler IP ranges
- 2Reverse DNS check
- 3Filter out anything that fails
What Actually Saves Crawl Budget, and What Only Looks Like It
This is where most advice on faceted navigation SEO goes wrong. Several tactics that clean up the index do nothing for crawl budget, because Google still has to fetch the page to read the instruction. Here’s how each control behaves, based on Google’s own documentation.
| Control | Stops crawling? | Stops indexing? | Best use |
|---|---|---|---|
| robots.txt disallow | Yes | MostlyA blocked URL with links pointing to it can still appear, without content | Filter and sort URLs you never want in search |
| Filters in URL fragments (#) | YesNo new crawlable URL exists | Yes | New builds and JavaScript filtering |
| 404 for empty combinations | Reduces itGoogle recrawls 404s less over time | Yes | Filter combinations with zero products |
| rel=”canonical” | NoMay reduce crawling slowly over time | Consolidates signalsIt is a hint | Close variants you want merged into the parent |
| noindex | NoThe page must be crawled to be read | Yes | Removing filter URLs that are already indexed |
| rel=”nofollow” on filter links | Only if every link to the URL uses it | No | A backup measure, never the main control |
Robots.txt Is the Real Crawl Control
Google recommends robots.txt disallow rules when you don’t need filter URLs in search. Its guidance says there’s often no good reason to allow crawling of filtered items, and suggests keeping only product pages and one unfiltered listing page crawlable. Google’s own example looks like this:
user-agent: Googlebot
disallow: /*?*products=
disallow: /*?*color=
disallow: /*?*size=
allow: /*?products=all$
One catch matters here. A disallowed URL can still get indexed if other pages link to it, just without its content. And Googlebot never reads a canonical or noindex on a page it can’t fetch. So give each URL pattern one job, and never stack controls that cancel each other.
Don’t Create the URL at All
The cleanest fix is a filter that never produces a crawlable address. Google generally ignores URL fragments, so filters written as #color=red have no effect on crawling, positive or negative. Client-side filtering that updates the grid without adding new links does the same job.
The trade-off is real, though. A filter state behind a fragment can’t rank on its own. So keep fragment or JavaScript filtering for sort orders and low-value combinations, and give the few high-demand filters proper static URLs that search engines can crawl.
- Sort orders
- Low-value combinations
A filter state behind a fragment can’t rank on its own
- The few high-demand filters
- Pages search engines can crawl
Canonical Tags Consolidate, They Don’t Block
Pointing filter pages at the parent category with a canonical is still worth doing, because it merges link signals. But Google says a canonical may decrease crawling of non-canonical URLs only over time, and rates it less effective in the long run than robots.txt or fragments.
Noindex Is Not a Crawl Budget Tool
This is the most repeated myth on the subject. Google’s crawl budget guide tells site owners not to use noindex for this purpose, because Google still requests the page and only drops it after seeing the tag. The crawl has already happened. Noindex controls the index, nothing more.
Noindex still has one useful job. If filter URLs are already indexed, add noindex first, let Google recrawl and drop them, then add the robots.txt block. Reverse that order and the URLs stay stuck in the index, because Google can no longer read the tag.
- 1Add noindex first
- 2Let Google recrawl and drop them
- 3Then add the robots.txt block
Reverse that order and the URLs stay stuck in the index, because Google can no longer read the tag.
Nofollow Needs Total Coverage
Nofollow on filter links can discourage crawling, but Google’s Crawling December post on faceted navigation stresses that every link to those URLs, internal and external, must carry it. One unmarked link in a footer, feed or email is enough for discovery, so treat nofollow as a backup only.
One unmarked link is enough for discovery, so treat nofollow as a backup only.
Return a 404 for Empty Filter Results
When a filter combination has no products, Google wants a real 404 status on that URL. Not a 200 page saying nothing was found, and not a redirect to a generic error page. The same rule covers duplicate filters, nonsense combinations and pagination past the last page.
This matters for crawl because Google’s crawl budget guide says soft 404 pages keep getting crawled, while a 404 is a strong signal not to crawl that URL again. An empty product grid returning 200 is a soft 404 waiting to happen.
- Duplicate filters
- Nonsense combinations
- Pagination past the last page
Blocking Filters Won’t Automatically Speed Up Your Other Pages
Here’s a nuance most guides skip. Google sets crawl budget from two things: crawl capacity, meaning how hard your server can be crawled, and crawl demand, meaning how much Google wants your content. Cutting filter crawl helps each of these in a different way.
If filter crawling overloads your server, blocking it frees capacity, and Google raises its limit as response times improve. But Google won’t shift freed budget to other pages unless your site already hits its capacity limit. On a healthy server, blocking cuts waste without boosting crawl elsewhere.
Blocking filter crawl frees capacity if it overloads your server. Google raises its limit as response times improve.
- Perceived inventory
- Popularity
- Staleness
Google won’t shift freed budget to other pages unless your site already hits its capacity limit.
So what raises demand? Google names three demand factors: perceived inventory, popularity and staleness. Perceived inventory is the one you control directly. Fewer junk URLs, more internal links to key categories and accurate lastmod dates in your sitemap are what pull crawling toward the pages that sell.
- Fewer junk URLs
- More internal links to key categories
- Accurate lastmod dates in your sitemap
Which Filter Pages Deserve to Be Crawled and Indexed
Blocking every filter is the opposite mistake. Some filters match real searches, like black leather boots or wide fit running shoes, and those pages can outrank the parent category. My test is simple: proven search demand, enough products to fill the page, and a stable URL.
Check demand with Search Console queries, keyword tools and your own site search data. If a filter passes, treat it as a landing page rather than a filter state. It’s the same category-level thinking I use in ecommerce SEO work, where each search intent gets one strong page.
If a filter passes, treat it as a landing page rather than a filter state.
For those pages, follow Google’s rules for crawlable facets. Use the standard & separator, keep filters in one fixed order, and never allow duplicate filters. Give each page a self-referencing canonical, a unique title and intro, a sitemap entry, and plain HTML links from the parent category.
- The standard & separator
- Filters in one fixed order
- Never allow duplicate filters
- A self-referencing canonical
- A unique title and intro
- A sitemap entry
- Plain HTML links from the parent category
Everything else needs a crawl decision, not only an index decision. Sort orders, view modes, price sliders, stacked filters and tracking parameters should be blocked or kept off crawlable links entirely. Single low-demand filters can stay crawlable with a canonical to the parent if their link signals are worth merging.
- Sort orders, view modes, price sliders, stacked filters, tracking parametersBlocked or kept off crawlable links entirely
- Single low-demand filtersCan stay crawlable with a canonical to the parent if their link signals are worth merging
Platform Notes for Shopify, Magento and Multi-Language Stores
Shopify
Shopify storefront filters add parameters such as filter.v.option.color and filter.p.vendor to collection URLs. The default robots.txt already blocks sort_by URLs, tag combinations joined with a plus sign, and collection URLs carrying two or more filter parameters. Check yours at yourstore.com/robots.txt.
The gap is single-filter URLs. One filter parameter stays crawlable by default, and many themes link every filter value as a plain anchor. On a large catalogue, that alone can mean thousands of URLs, which is a common thread in my Shopify SEO consulting audits.
sort_byURLs- Tag combinations joined with a plus sign
- Collection URLs carrying two or more filter parameters
- Single-filter URLs stay crawlable by default
- Many themes link every filter value as a plain anchor
filter.v.option.colorfilter.p.vendor
Shopify lets you edit robots.txt.liquid to add your own rules. Be careful here. A loose pattern can hide entire collections from Google, so test every new rule against real collection, filter and product URLs before publishing the theme change.
Magento and Adobe Commerce
Magento layered navigation builds query strings from attribute option IDs, so filter URLs look like ?color=49&size=170. The toolbar adds more: product_list_order, product_list_dir, product_list_limit and product_list_mode for sorting, direction, items per page and grid or list view, plus p for pagination.
?color=49&size=170Query strings from attribute option IDsproduct_list_orderSortingproduct_list_dirDirectionproduct_list_limitItems per pageproduct_list_modeGrid or list viewpPaginationTwo admin settings matter most. Adobe’s layered navigation documentation shows each filterable attribute is set to Filterable (with results) or Filterable (no results). The second option shows values with zero matching products, handing crawlers links to empty pages. I’d use with results unless there’s a strong UX reason not to.
Also enable Use Canonical Link Meta Tag for Categories, which Adobe recommends alongside the product canonical setting. It points filtered category states at the clean category URL. Then block the toolbar parameters in robots.txt, because a canonical alone won’t stop Googlebot fetching them.
Filterable (with results)I’d use with results unless there’s a strong UX reason not toFilterable (no results)Shows values with zero matching products, handing crawlers links to empty pagesUse Canonical Link Meta Tag for CategoriesPoints filtered category states at the clean category URLThen block the toolbar parameters in robots.txt, because a canonical alone won’t stop Googlebot fetching them.
Multi-Language Stores
Google treats each hostname as a separate site with its own crawl budget. A store on de.example.com and fr.example.com splits its budget by subdomain, while subfolder locales share one. Either way, every locale copies the full filter space, so multilingual SEO setups need facet rules applied per market.
de.example.comfr.example.comEither way, every locale copies the full filter space, so facet rules apply per market.
AI Crawlers Hit the Same Filter URLs
Googlebot isn’t the only bot walking your filters now. Cloudflare’s August 2025 analysis found training accounts for nearly 80% of AI bot crawling, and noted that some AI crawlers ignore robots.txt directives. Every open filter URL is available to all of them.
AI bots don’t draw from Google’s crawl budget, but they hit the same server. Google lowers its crawl capacity limit when response times rise or the server returns 5xx or 429 errors. If scrapers slow your store down on filter pages, Googlebot pays for it too.
- 1Scrapers slow your store down on filter pages
- 2Response times rise, or the server returns 5xx or 429 errors
- 3Google lowers its crawl capacity limit
AI bots don’t draw from Google’s crawl budget, but they hit the same server.
The practical answer is the same: shrink the crawlable filter space. Robots.txt handles well-behaved bots, and bot management or rate limiting at your CDN handles the rest. Keep product and category pages open if you want your store cited in AI answers.
- Robots.txt handles well-behaved bots
- Bot management or rate limiting at your CDN handles the rest
- Keep product and category pages open if you want your store cited in AI answers
Advice That’s Out of Date
Two pieces of old advice on faceted navigation and crawl budget still show up in guides published this year. The first is configuring parameter handling in Search Console. Google retired the URL Parameters tool in 2022, noting that only about 1% of its configurations were useful for crawling.
The second is relying on rel=”next” and rel=”prev” to manage paginated filter pages. Google confirmed in 2019 that it no longer uses them for indexing. They’re harmless markup, but they don’t control crawling. Link deep pages with normal anchors and return a 404 past the last page.
- 1
Configuring parameter handling in Search ConsoleGoogle retired the URL Parameters tool in 2022Only about 1% of its configurations were useful for crawling - 2
Relying on rel=”next” and rel=”prev” for paginated filter pagesGoogle confirmed in 2019 it no longer uses them for indexingLink deep pages with normal anchors and return a 404 past the last page
A Safe Order for Fixing Faceted Navigation and Crawl Budget Issues
Changing robots.txt on a live store without a plan is how categories vanish from Google. This is the sequence I follow, because each step protects the one after it and keeps every change easy to reverse if the data surprises you.
- 1Measure first. Pull at least 30 days of Crawl Stats and, if you can, server logs. Group requests by parameter so you know which filters take the most crawl.
- 2Classify every parameter. For each filter, sort and toolbar parameter, decide whether it becomes a landing page, gets consolidated with a canonical, or stays out of the crawl.
- 3Fix links before rules. Stop linking sort orders and low-value filters as plain anchors, and remove every filter URL from your XML sitemaps.
- 4Return real 404s. Serve a 404 status for empty combinations, duplicate filters and pagination beyond the last page.
- 5Deindex, then block. Add noindex to filter URLs that are already indexed, wait until they drop out, then add the robots.txt disallow rules.
- 6Test every pattern. Run each robots.txt rule against sample URLs, including different parameter orders, so you don’t block pages you want ranked.
- 7Monitor for at least two months. Track Crawl Stats, the Discovered, currently not indexed count and how fast new products get indexed.
If you’d like a second pair of eyes on the crawl data before touching robots.txt on a live store, a one-to-one SEO consultation is a cheap way to avoid blocking the wrong pages. It costs far less than recovering lost category rankings afterwards.
Faceted Navigation and Crawl Budget FAQs
Does faceted navigation hurt SEO?
Only when it creates crawlable URLs nobody searches for. Filters help shoppers find products faster. The damage comes from uncontrolled URL generation, which wastes crawling, creates near-duplicate pages and slows discovery of new products. Controlled facets with real demand can rank very well.
Should I use noindex or robots.txt for filter pages?
Use robots.txt when the goal is saving crawl budget, because noindex pages still get crawled. Use noindex only to remove filter URLs that are already indexed, then block them once they drop out. Applying both at once stops Google reading the noindex.
Do canonical tags fix crawl budget?
Not directly. Google still crawls canonicalised filter URLs, and it says canonicals may reduce that crawling only gradually. They’re useful for merging signals between near-duplicate pages, but they aren’t a substitute for blocking or never creating low-value URLs.
How many filter pages should I index?
Index the filters that have real search demand and enough products to be useful, and no more. For most stores that’s a short list per category. Remember that every indexed filter page also becomes a page you have to maintain and keep fresh.
Can a small store ignore crawl budget?
Usually, yes. Google says sites whose new pages get crawled the same day they’re published don’t need its crawl budget guide. The exception is a small catalogue whose filters expose huge numbers of URLs, which shows up as a growing Discovered, currently not indexed count.
More SEO guides
SEO Trends for 2025: Top 10 Tips That Will Work
Want to know what’s in 2025 SEO Trends? Learn with us about the latest SEO trends and strategies to enhance your website’s visibility and ranking.
On-Page SEO vs Off-Page SEO
How on-page and off-page SEO differ, what each one covers, and where each one fits in your wider SEO work. See which parts you fix on the page itself, which come from signals outside your site, and how the two connect.
15 Best Link Building Tools for SEO in 2025 [Free & Paid]
Are you searching for the best and most accurate link-building tool? In this article, I’ve reviewed the 15 free and paid link building tools you can…
More SEO guides
SEO Trends for 2025: Top 10 Tips That Will Work
Want to know what’s in 2025 SEO Trends? Learn with us about the latest SEO trends and strategies to enhance your website’s visibility and ranking.
On-Page SEO vs Off-Page SEO
How on-page and off-page SEO differ, what each one covers, and where each one fits in your wider SEO work. See which parts you fix on the page itself, which come from signals outside your site, and how the two connect.
15 Best Link Building Tools for SEO in 2025 [Free & Paid]
Are you searching for the best and most accurate link-building tool? In this article, I’ve reviewed the 15 free and paid link building tools you can…
Want a second pair of eyes on the crawl data?
Before touching robots.txt on a live store, a one-to-one SEO consultation is a cheap way to avoid blocking the wrong pages.
Want a second pair of eyes on the crawl data?
Before touching robots.txt on a live store, a one-to-one SEO consultation is a cheap way to avoid blocking the wrong pages.

I’m an awarded SEO expert and seo consultant with 9 years of experience. I have generated over 0 to 100M sales by SEO for my clients like you 👉 See Case Study