[Atomic Glue](atomicglue.co)
SEO
Home › Glossary › SEO· 110 ·

Crawling

\kraw-ling\n.
Filed underSEO
In brief · quick answer

Crawling is the process by which search engine bots discover and download web pages by following links across the internet.

§ 1 Definition

Crawling is the discovery phase of search. Bots (most commonly Googlebot) start from a seed set of known URLs and follow hyperlinks to find new pages. Each discovered URL is fetched, its content is read, and links on that page are added to the crawl queue. Crawling is the prerequisite for everything else: if a page is never crawled, it cannot be indexed, ranked, or served to users.

§ 2 How Crawling Works

Googlebot maintains a priority queue of URLs. It decides which to crawl first based on factors like PageRank, how often the page changes, and how important the site is. The bot respects directives in robots.txt, meta robots tags, and crawl delay settings.

§ 3 Common Crawl Problems

Pages hidden behind login walls, blocked by robots.txt, requiring JavaScript rendering without server-side fallback, or orphaned with no internal links pointing to them are effectively invisible to crawlers. If Googlebot cannot reach it, it does not exist for search.

§ 4 Note

A common misunderstanding: 'crawled' does not equal 'indexed.' Crawling is just the discovery step. A page can be crawled perfectly and still be excluded from the index due to quality filters, duplicate content signals, or explicit noindex directives.

§ 5 Common questions

Q. Can I block Googlebot from crawling certain pages?
A. Yes, using robots.txt or a noindex meta tag. robots.txt blocks crawling but may still allow indexing if the page is discovered another way. Noindex blocks indexing entirely.
Key takeaways
  • Crawling is the discovery phase, not the storage phase
  • Internal links are the primary crawl paths
  • Blocked or orphaned pages do not get crawled
How Atomic Glue helps

Atomic Glue builds SEO strategies that turn technical fundamentals into measurable rankings. Get in touch to discuss your SEO & GEO services needs.

Get in touch
# Crawling

Crawling is the process by which search engine bots discover and download web pages by following links across the internet.

Category: SEO

Author: Atomic Glue SEO & GEO Team

## Definition

Crawling is the discovery phase of search. Bots (most commonly Googlebot) start from a seed set of known URLs and follow hyperlinks to find new pages. Each discovered URL is fetched, its content is read, and links on that page are added to the crawl queue. Crawling is the prerequisite for everything else: if a page is never crawled, it cannot be indexed, ranked, or served to users.

## How Crawling Works

Googlebot maintains a priority queue of URLs. It decides which to crawl first based on factors like PageRank, how often the page changes, and how important the site is. The bot respects directives in robots.txt, meta robots tags, and crawl delay settings.

## Common Crawl Problems

Pages hidden behind login walls, blocked by robots.txt, requiring JavaScript rendering without server-side fallback, or orphaned with no internal links pointing to them are effectively invisible to crawlers. If Googlebot cannot reach it, it does not exist for search.

## Note

A common misunderstanding: 'crawled' does not equal 'indexed.' Crawling is just the discovery step. A page can be crawled perfectly and still be excluded from the index due to quality filters, duplicate content signals, or explicit noindex directives.

## Common questions

Q: Can I block Googlebot from crawling certain pages?

A: Yes, using robots.txt or a noindex meta tag. robots.txt blocks crawling but may still allow indexing if the page is discovered another way. Noindex blocks indexing entirely.

## Key takeaways

## Related entries


Last updated June 2026. Permalink: atomicglue.co/glossary/crawling

Schedule a call

30 min · Video call

1
Date
2
Time
3
Details