38 junk URLs, zero real articles: anatomy of a "Crawled – currently not indexed" report
A 19-article niche site in the home-and-garden space had a Search Console page-indexing report nearly twice the size of its actual content. Here's what every flagged URL turned out to be, where each came from, and why this pattern is hiding in almost every WordPress site more than a few years old.
The scare
If you own a WordPress site, you've seen this screen: Search Console reports dozens — sometimes thousands — of pages "Crawled – currently not indexed." The site in question had 38 of them against just 19 published articles. Was the site being penalized? Was the content thin? Was something broken?
The owner did what most people do: worried, read contradictory forum threads, and considered rewriting content that wasn't the problem. Then we audited every flagged URL individually. The answer was better and stranger than expected.
What the junk actually was
| Pattern | Example shape | Where it came from |
|---|---|---|
| Malformed numeric suffixes | /article-slug/1000, //1000 | A years-old theme's broken relative links, still being re-crawled long after the theme was replaced |
| RSS feed URLs | /feed/, /article/feed/ | WordPress generates one per page; crawlers fetch them all |
| Affiliate redirect endpoints | /att/r/…, /dtt/r/… | A link-cloaking plugin's internal redirects, crawled and correctly refused |
| Parameter variants | ?page_num=, ?s=, ?replytocom= | Old pagination scripts, internal search, legacy comment links |
| Ad-tech leftovers | ?ez_*= URLs | Residue from a managed ad platform the site had left years earlier |
Thirty-eight flagged URLs. Every single one was junk. All 19 real articles were indexed and fine. The scary report was, in its entirety, the archaeological record of every theme, plugin, and ad platform the site had ever used.
The honest part
Google was already handling it. These URLs weren't being indexed — that's the point of the report — and on a site this small, the junk wasn't measurably hurting rankings. If a tool tells you otherwise, be suspicious of it. What the junk was costing: crawl budget spent re-fetching garbage instead of new content, duplicate-signal noise around real pages, permanent illegibility of the one report that's supposed to tell you how your site is doing — and hours of the owner's anxiety about a problem that didn't exist.
That last line is why IndexSweep exists. The fix for a report like this isn't panic-rewriting your content — it's classification (which URLs are junk and why), cleanup (reversible noindex/redirect/410 rules so the junk stops accumulating), and source repair (fixing the broken in-content links that keep minting ghost URLs). On bigger, older, messier sites — parameter-heavy stores, decade-old blogs with deep plugin archaeology — the same cleanup also has real mechanical value: consolidated signals and crawl budget flowing to pages that matter.
What IndexSweep does with a report like this
Upload the same Search Console export into IndexSweep and each URL is classified against a library of these exact patterns — with a recommended action and a plain-English "why." One click applies a reversible fix per pattern or per URL; the content scan finds and repairs the malformed links that created the ghosts in the first place; and every change is logged with one-click rollback.
Free forever for auditing and fixing. Nothing leaves your server.
See what's in Free vs Pro