sitemap.xml checker
A sitemap is a machine-readable list of the URLs you want indexed. Search engines do not need it to crawl you, but on a new site with few backlinks it is the fastest way to get deep pages discovered. Enter a domain and we'll check whether the file is actually served and whether crawlers are being told where it is.
Across 955 websites analyzed
Measured across every site in our public database.
Why this matters
It decides how fast new pages get found, not whether they rank
A sitemap is a discovery aid. It will not push a weak page up the results, and submitting one does not guarantee indexing — Google routinely crawls sitemap URLs and then declines to index them. What it does reliably is shorten the gap between publishing a page and a crawler seeing it, which matters most on large sites and brand-new domains.
lastmod is the part most sites get wrong
Google uses the `<lastmod>` date to decide what to re-crawl — but only if it trusts it. Sites that stamp every URL with today's date whenever the sitemap regenerates teach crawlers to ignore the field entirely. Set it from the actual content-change time, and leave it alone when nothing changed.
How to fix it
1. Generate it, don't hand-write it
Needs a developerEvery major platform has this built in: WordPress ships one at `/wp-sitemap.xml` (Yoast and Rank Math use `/sitemap_index.xml`), Shopify at `/sitemap.xml`, Cafe24 and Imweb expose one in their SEO settings, and Next.js generates one from `app/sitemap.ts`. A hand-maintained file goes stale within a month.
2. Only include URLs you actually want indexed
Needs a developerDrop anything that is noindexed, canonicalised elsewhere, redirecting, or returning a 404. A sitemap full of URLs Google then refuses to index is a quality signal working against you. Same for thin or near-duplicate pages — leave them out until they're worth indexing.
3. Declare it in robots.txt and submit it once
Needs server or hosting accessAdd `Sitemap: https://yoursite.com/sitemap.xml` to robots.txt — that covers every crawler. Then submit it once in Google Search Console and Naver Search Advisor. After that you do not need to resubmit; crawlers re-fetch it on their own.
4. Split it if it's large
Needs a developerThe limits are 50,000 URLs and 50 MB uncompressed per file. Past that, publish a sitemap index that references child sitemaps. Splitting by section (products, posts, categories) is also useful diagnostically — Search Console reports indexing per sitemap, so you can see which section is being rejected.
Want this fixed rather than explained?
We do this as fixed-price work — no hourly billing, no retainer. You get a quote before anything starts.
See fixed prices →Frequently asked questions
- My sitemap is submitted but pages show 'Discovered — currently not indexed'. Why?
- That status means Google found the URL and chose not to fetch or keep it — a quality and crawl-budget decision, not a sitemap problem. Resubmitting will not change it. It usually resolves by making the pages substantially more useful, reducing near-duplicates, and earning links to the section.
- Do I need to resubmit after publishing a new page?
- No. Once a sitemap is submitted, crawlers re-fetch it periodically on their own. For faster pickup you can ping IndexNow (Bing, Naver, Yandex) or use the URL Inspection tool in Search Console for a single important page.
- Should I include images and videos?
- Only if media search matters to your business — image and video sitemap extensions help e-commerce and publishers, and do very little for a typical service site. The base URL list is what matters first.
- XML or plain text?
- A plain-text file with one URL per line is valid and accepted by Google and Bing. It just cannot carry lastmod or priority, so XML is the better default for anything that changes.
Real sites failing this check
Pulled live from our corpus — each links to its full report.
yale.eduYale UniversityC 59SEO 74AI 45Schema 10- example.comExample DomainF 35SEO 22AI 20Schema 0
- so.com360搜索,SO靠谱F 33SEO 50AI 22Schema 0
- tmon.co.kr티나는 쇼핑, 티몬F 26SEO 10AI 3Schema 0
- digitalocean.comAI-Native Cloud | DigitalOceanC 58SEO 78AI 45Schema 0
arxiv.orgarXiv.org e-Print archiveD 40SEO 40AI 25Schema 0
statcounter.comWeb Analytics Made Easy - StatcounterD 43SEO 46AI 35Schema 0
blogger.comBlogger.com - Create a unique and beautiful blog easily.D 48SEO 60AI 15Schema 0