cuyava ← IndexSweep home
Case study · the audit that became IndexSweep

38 junk URLs, zero real articles: anatomy of a "Crawled – currently not indexed" report

A 19-article niche site in the home-and-garden space had a Search Console page-indexing report nearly twice the size of its actual content. Here's what every flagged URL turned out to be, where each came from, and why this pattern is hiding in almost every WordPress site more than a few years old.

76
URLs known to Google
19
real articles on the site
38
"crawled – not indexed" URLs
0
real articles among them

The scare

If you own a WordPress site, you've seen this screen: Search Console reports dozens — sometimes thousands — of pages "Crawled – currently not indexed." The site in question had 38 of them against just 19 published articles. Was the site being penalized? Was the content thin? Was something broken?

[ Screenshot: GSC page-indexing report — domain redacted ]

The owner did what most people do: worried, read contradictory forum threads, and considered rewriting content that wasn't the problem. Then we audited every flagged URL individually. The answer was better and stranger than expected.

What the junk actually was

PatternExample shapeWhere it came from
Malformed numeric suffixes/article-slug/1000, //1000A years-old theme's broken relative links, still being re-crawled long after the theme was replaced
RSS feed URLs/feed/, /article/feed/WordPress generates one per page; crawlers fetch them all
Affiliate redirect endpoints/att/r/…, /dtt/r/…A link-cloaking plugin's internal redirects, crawled and correctly refused
Parameter variants?page_num=, ?s=, ?replytocom=Old pagination scripts, internal search, legacy comment links
Ad-tech leftovers?ez_*= URLsResidue from a managed ad platform the site had left years earlier

Thirty-eight flagged URLs. Every single one was junk. All 19 real articles were indexed and fine. The scary report was, in its entirety, the archaeological record of every theme, plugin, and ad platform the site had ever used.

The honest part

Google was already handling it. These URLs weren't being indexed — that's the point of the report — and on a site this small, the junk wasn't measurably hurting rankings. If a tool tells you otherwise, be suspicious of it. What the junk was costing: crawl budget spent re-fetching garbage instead of new content, duplicate-signal noise around real pages, permanent illegibility of the one report that's supposed to tell you how your site is doing — and hours of the owner's anxiety about a problem that didn't exist.

That last line is why IndexSweep exists. The fix for a report like this isn't panic-rewriting your content — it's classification (which URLs are junk and why), cleanup (reversible noindex/redirect/410 rules so the junk stops accumulating), and source repair (fixing the broken in-content links that keep minting ghost URLs). On bigger, older, messier sites — parameter-heavy stores, decade-old blogs with deep plugin archaeology — the same cleanup also has real mechanical value: consolidated signals and crawl budget flowing to pages that matter.

What IndexSweep does with a report like this

Upload the same Search Console export into IndexSweep and each URL is classified against a library of these exact patterns — with a recommended action and a plain-English "why." One click applies a reversible fix per pattern or per URL; the content scan finds and repairs the malformed links that created the ghosts in the first place; and every change is logged with one-click rollback.

[ Screenshot: IndexSweep GSC Import tab — classified results ]
IndexSweep is coming to the WordPress.org plugin directory.

Free forever for auditing and fixing. Nothing leaves your server.

See what's in Free vs Pro