In brief
- The “Request indexing” button in Search Console puts a page in a priority queue, with no guarantee of indexing.
- A noindex tag removes a page from Google’s results, provided Google can crawl the page to read the tag.
- A robots.txt file blocks crawling: a blocked page can still show up in Google if other sites link to it.
- Google’s crawl budget guide is aimed at very large sites.
- On our own site, 106 pages sat in a queue without ever being fetched. After three weeks of targeted requests, the number of indexed pages went from 18 to 90.
You can’t force page indexing on Google, only request it, in Search Console’s URL Inspection tool or through a sitemap, and Google then decides on its own. To keep a page out of the index, put a noindex on a page Google can crawl, without blocking it in robots.txt. What remains is choosing, page by page, what you want to see in the index.
On 27 August 2026, after redesigning dotsland.com, we resubmitted our sitemaps to Google, which hadn’t read them since 7 November 2025. Twelve days later, Search Console showed 18 pages indexed and 123 not indexed, 106 of them “Discovered – currently not indexed”, with no crawl date at all. On 29 September, it showed 90 indexed pages.
Can you really force Google to index a page?
No, and Google’s documentation says so plainly. Google first finds an address, through a link or a sitemap. It fetches it when it sees fit. Then it decides whether to keep it: according to its documentation, a crawled page gets evaluated first, and not all of them end up in the index. The Search Console help puts it in one line: “Don’t expect every URL on your site to be indexed.”
The “Request indexing” button in the URL Inspection tool has a daily limit per property. Google makes clear that a request doesn’t guarantee indexing and that asking again for the same URL won’t speed it up. Crawling can take anywhere from a few days to a few weeks, and a sitemap is only a hint. As for the shortcuts people pass around, Google’s Indexing API is reserved for job postings and livestream videos, and Google isn’t among the participating search engines listed by the IndexNow protocol.
Blocking, on the other hand, does work. A noindex tag in the page code, or the same rule in the HTTP header, drops the page from Google’s results, even if other sites link to it. But only once Google has crawled the page and read the rule.
Robots.txt doesn’t take a page out of Google
The mistake looks logical: if the crawler can’t get in, the page won’t be indexed. Google says the opposite. A robots.txt file tells crawlers which addresses they can visit, and Google states that “this is not a mechanism for keeping a web page out of Google”. A page disallowed in robots.txt can still be indexed if other sites link to it. It then appears without a description. Search Console even has a status for this case: “Indexed, though blocked by robots.txt”.
The trap closes when you combine the two. A page with a noindex that is also blocked in robots.txt will never be crawled, so Google will never read the noindex, and the page can stay in the results. For the test copies of our site, we therefore set a noindex header and deliberately added no robots.txt block: if an outside link ever led Google to them, it had to be able to read the rule.
The second trap lies in checking. When we switched our 11 category archives to noindex, the site’s SEO plugin removed them from the sitemap straight away. The pages themselves kept announcing “index, follow” in their code until each category had been saved a second time. The sitemap treated the job as done, while only the page code showed where things really stood.
Robots.txt still has its place. For catalogue filter URLs that have no reason to appear in search, Google recommends blocking them from crawling, because its crawlers visit a great many of them before working out they’re useless. As we read it, the order matters: if those URLs are already indexed, add the noindex first and only block them once they have dropped out.
Our 106 pages waiting to be crawled
In Search Console, “Crawled – currently not indexed” and “Discovered – currently not indexed” look alike and call for opposite treatment. In the first case, Google fetched the page and didn’t keep it, most often, as we read it, because of content or duplication. In the second, Google knows the address and has put off crawling it. All 106 of our pages were in the second group.
The reflex is to blame internal linking. We counted before touching anything: our contact, expertise and resources pages each received 78 internal links, as many as our “about” page, which was indexed. The technical side was sound too: 126 URLs out of 126 responded normally, with a median response time of 0.40 seconds, and the live test confirmed Google could index the page.
That left demand. The Crawl stats report settled it: 710 Google requests in 90 days, roughly 8 a day, 99% of them to refresh known pages and under 1% to discover new ones. HTML pages accounted for only 19% of those requests, about three pages a day. We weren’t hitting any ceiling: the crawler barely came, because a domain that had long been dormant, with hardly any links from elsewhere, generates no demand.
So we asked, about ten pages per 24-hour window, 100 requests between 8 and 25 September. On 15 September, the 49 pages requested in the first week were inspected one by one: all 49 were indexed, and one had been crawled within an hour of the request. A page we hadn’t requested yet, kept as a control, still hadn’t been crawled. The indexing report, meanwhile, still showed a last update dated 4 September, before the first request. It caught up on 17 September: 66 indexed pages, then 86 on the 22nd, then 90.
One caveat. By 25 September, pages we had never requested were getting indexed too: of 41 published addresses missing from our list, 40 were already indexed without a request. Our new articles now get into the index on their own within one to three days. As we read it, the manual requests got things moving without dealing with the cause, the lack of links from other sites, which the method further down comes back to.
Crawl budget is aimed at very large sites
The usual argument for keeping pages out of the index is crawl budget. Google aims its guide at sites with a million pages changing weekly, or over 10,000 pages changing daily. When we took our category archives out of the index, we recorded a budget gain of zero: we did it for quality.
For a small business, choosing is about deciding what Google sees of you. The pages to exclude are the ones nobody has a reason to find through a search: the thank-you page, your on-site search results, basket and customer account, filter combinations. Variants of the same product are more a job for the canonical, and that’s often where the visibility of product pages that don’t rank is decided. The pages to push are the ones that answer a question people type into Google: your service pages and your new content.
Our English pages made up more than half of the addresses waiting in the queue: we put the request quota on French first, the language that matters commercially. If you are opening a second country, the same logic of order applies.
Which option acts on which step?
Each option works on a different step. Here’s what it really does, and when to use it.
| Option | What it really does | When to use it |
|---|---|---|
| Request in URL Inspection | Puts a page in a priority queue, with a daily cap and no guarantee | A few important pages, new or updated |
| Sitemap | Tells Google about addresses; the last modified date is only used if it’s accurate | Always, for the pages you want indexed and only those |
| Internal links | Help Google find pages and understand where they sit | When ignored pages get fewer links than indexed ones |
| Links from other sites | Bring the crawler, and so create crawl demand | A rarely cited site, or one coming out of a long pause |
| noindex | Removes the page from results, once Google has crawled it | Pages useful to visitors but with no search value |
| robots.txt | Blocks crawling, not indexing | Filters and parameters that were never indexed |
| Canonical | A strong signal naming the reference page, which Google may not follow | Variants and near-duplicates |
| Temporary removal | Hides a URL from results for about six months | Emergencies, while you add a noindex or a 404 or 410 code |
Google notes that the signals naming the reference page reinforce each other when they say the same thing. Conversely, a noindex blocked by robots.txt or a noindexed page listed in the sitemap sends opposite signals, and Google decides.
This week’s work in Search Console
Here is the method we use, with Search Console and a spreadsheet.
- Read the date before the number. At the top of the page indexing report is a last updated date. On 31 August, ours said 21 August, like two other unrelated properties: the delay was on Google’s side. A frozen counter doesn’t prove nothing is moving.
- Separate the two queues. “Discovered” calls for crawl demand. “Crawled” calls for work on the page, or accepting it won’t be indexed: Google says there’s no need to resubmit it.
- Draw up a short list of pages to push, around twenty ranked by commercial value, and another of pages to remove, each with the option you chose from the table. For every page on both lists, note the last crawl date in URL Inspection and compare the Google-selected canonical with the one you declared.
- Remove cleanly. Add the noindex while checking robots.txt isn’t blocking those pages, then verify on the page itself, through the source code or the “Test live URL” button, since the sitemap can change before the page does.
- Request, then verify through inspection a few days later, without waiting for the report. A request interrupted before confirmation also gets inspected before it’s repeated.
- Deal with the cause. In Crawl stats, look at the share of “Discovery”. If it’s close to zero, the lasting fix is outside your site: links from sites Google already visits, such as partners, suppliers, trade directories, local press. Google says it uses links to find new pages to crawl. As we read it, the same goes for AI engines, which only cite what they’ve found. On our own site, their crawlers barely read our llms.txt file, as we measured in our piece on llms.txt v2.
- Track each URL over time. Our own tracking fits in a text file, one line per URL with its request date. A tracking tool should match the sitemap against each URL’s real status and HTTP header, keep each URL’s history and date its own data, since Google’s report can lag by several days.
If a page that matters for your sales is still out of the index after these checks, pass us its URL and we’ll look with you at what is holding it back. Indexing follow-up is part of our search, advertising and social expertise.
Sources
Pages read on 30 September 2026. The dotsland.com figures are from our own Search Console, recorded from 27 August to 29 September 2026.
- Google Search Central, “Ask Google to Recrawl Your Website”, updated 10 December 2025. View
- Search Console Help, “URL Inspection tool”. View
- Search Console Help, “Page indexing report”. View
- Google Search Central, “Block Search Indexing with noindex”, updated 10 December 2025. View
- Google Search Central, “Robots.txt Introduction and Guide”, updated 10 December 2025. View
- Google Search Central, “Build and Submit a Sitemap”, updated 8 July 2026. View
- Google Search Central, “How to Specify a Canonical with rel=”canonical” and Other Methods”, updated 10 July 2026. View
- Search Console Help, “Removals and SafeSearch reports tool”. View
- Google, “Crawl Budget Management”, updated 22 July 2026. View
- Search Console Help, “Crawl Stats report”. View
- Google, “Managing crawling of faceted navigation URLs”, updated 18 December 2025. View
- Google Search Central, “SEO Link Best Practices for Google”, updated 10 December 2025. View
- Google Search Central, “Indexing API Quickstart”, updated 16 July 2026. View
- IndexNow.org, “FAQ”. View
FAQ
How many indexing requests can you make per day?
Google doesn’t publish a figure, it only mentions a daily limit per property. On ours, we got about ten requests per 24-hour window, with each unit coming back roughly a day after it was used. A one-off generic error message wasn’t the quota: the same page went through on the next run. A quota refusal, by contrast, says so in plain words.
My page is “Crawled – currently not indexed”: should I request indexing again?
No. Google says it may or may not be indexed later and that there’s no need to resubmit it. Google has read it and not kept it: if it duplicates another page or doesn’t really answer any question, merge it into a fuller page and redirect.
How do you quickly remove a page published by mistake?
Use Search Console’s removals tool, which hides the URL for about six months, and at the same time put the permanent fix in place: noindex, password protection, or deleting the page with a 404 or 410 code. Without that second step, the page can come back when the period runs out.

