Zubair Akhtar logo — Muhammad Zubair Akhtar, Software Engineer

Blog

How I Removed Hacked Spam URLs From Google (Case Study)

Case study: how I cleared 248 hacked spam URLs from Google on my Next.js site with 410 Gone + noindex, an allowlist proxy, and Search Console prefix removals.

Muhammad Zubair Akhtar

I build websites for a living, so finding spam indexed under my own domain was humbling. On 5 October 2026 I exported the indexed pages list from Google Search Console for zubairakhtar.com. It showed 249 indexed URLs. Exactly one of them was a real page: my homepage. The other 248 were spam I had never published.

They fell into three families. There were 212 homepage URLs carrying junk query parameters such as ?b=, ?m= and ?k=, 32 URLs under /shopdetail/, and 5 multilingual casino slugs: a Russian "vavada" casino page, a Dutch "betnjet" review and a French "casea" payments page. None of it was content I had written, and all of it was sitting in Google under my name.

This post walks through how I spotted it, why cleaning files is not enough, the allowlist + 410 Gone fix I shipped on my Next.js site, and how I used Search Console's Removals tool to remove URLs from Google search without hiding the real site.

End-to-end cleanup workflow used on zubairakhtar.com.

Figure 1. End-to-end cleanup workflow used on zubairakhtar.com.

How I spotted it

A site:zubairakhtar.com search was the first hint. It is never an exact count, but it showed URLs I did not recognise. The real evidence came from Search Console: Indexing → Pages, then exporting the indexed URL list.

My portfolio is small. It has seven routes, and those seven are the only ones in my sitemap:

  • /
  • /about
  • /experience
  • /skills
  • /projects
  • /education
  • /contact

In the 2026-10-05 export, the homepage was the only one of those routes in the indexed list. Every other indexed URL, all 248 of them, was junk. Once you have that export and a known-good list, you stop guessing. Anything not on the list is suspect.

If you are in the same position, check one more thing early: Settings → Users and permissions in Search Console. Remove any owner or user you did not add, along with their verification tokens. Spammers sometimes verify themselves so they can submit their own sitemaps.

What the spam looked like

1. /shopdetail/ pages (32 URLs)

A batch of shop-listing style URLs under /shopdetail/…. Nothing like that has ever existed on my site. Before the fix these returned 404, which is better than 200 but still a soft "not found, maybe later" signal.

2. Junk query parameters on the homepage (212 URLs)

This was the biggest group and the sneakiest. Google had indexed homepage variants like /?b=…, /?m=…, /?k=…, /?c=…, /?d=…, /?h=…, /?n=… and /?category/…. These are classic doorway-style URLs: the parameters mean nothing to my app, but they give a spammer endless unique URLs on a trusted domain.

The problem was that every one of them returned 200 and served my real homepage. My app ignored unknown parameters and rendered the page anyway. From Google's point of view, those were valid, working URLs.

3. Multilingual casino slugs (5 URLs)

The last group was casino spam in several languages:

  • /kak-zaregistrirovat-sia-v-vavada-kazino-onlain/ (Russian, "vavada" casino registration)
  • /betnjet-app-review-registratie-bonussen-betaalmethoden-en-veiligheid/ (Dutch, "betnjet" app review)
  • /casea-online-casino-methodes-de-paiement-depots-instantanes-et-retraits-rapides/ (French, "casea" payment methods)

There were also date-path casino posts under /2026/… in the classic /year/month/day/slug/ permalink shape. On my current site these slugs returned a 308 redirect (Next.js normalising the trailing slash) and then, most likely, a 404 at the end of the chain.

None of these were good signals. A 200 says "this page is real". A 404 says "not found right now". A redirect chain makes crawlers work harder for a weak answer.

Why cleaning files alone is not enough

If a site is hacked, the first job is still to remove malicious code, close the hole and rotate credentials. But once the files are clean, Google's index does not reset itself. Google keeps the URLs it has already discovered and recrawls them on its own schedule. If those URLs now return an ambiguous answer, or worse a 200 with your homepage, they can stay indexed for a long time.

So there are really two jobs:

  1. Stop the infection wherever it lives: the current server, an old host, plugins, the database.
  2. Tell search engines clearly and consistently that every junk URL is permanently gone.

The fix I shipped: allowlist, 410 Gone and noindex

My site runs on Next.js, so I handled this in the request layer instead of chasing individual URLs. The rule is simple: only known-good paths are served. Everything else is gone.

Two files do the work:

  • lib/index-policy.ts holds the allowlist and an evaluateIndexPolicy function that looks at each request and returns a decision: allow, gone or redirect.
  • proxy.ts runs before routing, calls that function and acts on the decision.

The allowlist

This is the actual list of indexable pages:

export const INDEXABLE_PAGES = [
  "/",
  "/about",
  "/experience",
  "/skills",
  "/projects",
  "/education",
  "/contact",
];

Instead of writing a rule for every spam pattern (and missing the next one), I list what is real. The policy is:

  • Allowlisted paths pass through and return 200 as normal.
  • The homepage only accepts tracking parameters: utm_*, gclid and fbclid. Any other query parameter, like ?b= or ?category/…, gets a 410.
  • Unknown paths, including /shopdetail/… and the casino slugs, get a 410 directly, with no 308 → 404 chain first.
  • Static assets and framework internals are excluded from the check so the site keeps working.

Returning 410 Gone

This is the branch in proxy.ts that handles the gone decision:

if (decision.action === "gone") {
  return new NextResponse(GONE_HTML, {
    status: 410,
    headers: {
      "Content-Type": "text/html; charset=utf-8",
      "X-Robots-Tag": "noindex",
      "Cache-Control": "public, max-age=86400",
    },
  });
}

GONE_HTML is a tiny page for any human who lands there by accident. The response carries three signals: the 410 status, an X-Robots-Tag: noindex header, and a cache header so a CDN can answer cheaply when bots keep hitting these URLs.

Next.js proxy + index-policy: only the seven portfolio routes return 200; everything else returns 410 Gone with noindex.

Figure 2. Next.js proxy + index-policy: only the seven portfolio routes return 200; everything else returns 410 Gone with noindex.

Why 410 instead of 404

Both codes tell Google the page is not there, and Google eventually drops both. The difference is intent. A 404 leaves room for the page to come back, so Google may recrawl it a few more times before letting go. A 410 says the removal was deliberate and permanent. For spam you never want back, 410 is the more honest answer.

The noindex header is a second, consistent signal. Because it is an HTTP header, it works on any response, not only on HTML pages with a meta tag.

Why I did not block the spam in robots.txt

It is tempting to add Disallow: /shopdetail/ to robots.txt. Don't. If Googlebot is not allowed to crawl a URL, it never sees your 410 or your noindex, and the URL can stay in the index with no description. Leave spam paths crawlable so bots can come back, get the 410 and drop them.

Temporary Removals in Search Console

The 410 is the lasting fix, but Google has to recrawl each URL to see it. To hide the junk from results sooner, on 5 October 2026 I submitted removal requests in Search Console → Indexing → Removals → New request.

Three "Remove all URLs with this prefix" requests:

  • https://zubairakhtar.com/shopdetail/
  • https://zubairakhtar.com/2026/
  • https://zubairakhtar.com/?

Plus the casino slugs as single-URL removals, including the vavada, betnjet and casea pages above.

Never submit your bare domain as a prefix

If you submit https://example.com/ as a prefix, every URL on your site matches, including your homepage. You would hide your whole site from Google for months.

The query-string prefix needs the same care. https://zubairakhtar.com/? (with the question mark) only matches homepage URLs that have a query string, such as /?b=123. It does not match the clean homepage or /about. Before submitting a prefix like that, read it character by character and compare it against your allowlist. One side effect I accepted: homepage URLs with utm_ parameters also fall under that prefix, which is fine because they should never show in results as separate pages.

Removals are temporary

A removal request hides URLs for roughly six months. It does not deindex them permanently. If a URL still returns 200 when the request expires, it can come back. That is why the order matters: ship the 410 first, then use Removals to clean up what searchers see in the meantime.

How to verify with curl

After deploying I checked the responses directly. curl -I fetches only the headers.

Spam examples should return 410 with noindex:

curl -sI "https://zubairakhtar.com/shopdetail/12345"
curl -sI "https://zubairakhtar.com/?b=12345"
curl -sI "https://zubairakhtar.com/kak-zaregistrirovat-sia-v-vavada-kazino-onlain/"

Expected output includes:

HTTP/2 410
x-robots-tag: noindex

Real pages and tracking-only homepage URLs should return 200:

curl -sI "https://zubairakhtar.com/"
curl -sI "https://zubairakhtar.com/about"
curl -sI "https://zubairakhtar.com/?utm_source=test"

Because hacked sites often cloak, repeat a spam check with a crawler user agent, for example curl -sI -A "Googlebot" plus the URL. If the answer differs from a normal request, something is still treating bots differently. In Search Console, URL Inspection → Test live URL shows what Google actually receives.

Post-deploy checklist

The proxy and the Search Console removals are done. These are the remaining items I am working through:

  • Sitemap contains only the seven real pages, resubmitted in Search Console.
  • Search Console users and owners reviewed, with no unknown verifications.
  • Bing Webmaster Tools: Bing has its own index and its own URL removal tool, so the same prefix care applies there.
  • IndexNow (optional): pinging the junk URLs can prompt Bing and other participating engines to recrawl them and see the 410.
  • Old host check: if the domain ever pointed at older hosting, check it for backdoors, unknown admin users and injected files, and make sure it no longer serves anything for the domain.
  • Re-run the curl checks after any deploy that touches routing.

What this does not fix alone

A 410 policy controls what your current site tells search engines. It does not remove malware. If spam was injected through an old CMS install, a compromised plugin or a forgotten hosting account, that infection still has to be found and cleaned, or it can keep generating pages and redirecting visitors wherever it still runs. Change passwords, revoke old keys and FTP accounts, and rebuild anything you no longer trust.

It also does not make Google forget instantly. Recrawling takes time.

What I'm watching next

I am not going to quote a "fixed in X days" number, because I don't have one yet. What I am watching:

  • The indexed count in Search Console's Pages report, week by week, heading toward the real pages only.
  • The "Not indexed" reasons, where the junk URLs should start appearing as excluded.
  • site: results for the spam patterns, as a rough cross-check.
  • Any new spam pattern. If one appears, it points to an infection still active somewhere.

Need help with something similar?

I'm a Senior Software Engineer in Lahore working across Laravel/PHP, React, Vue, Next.js, WordPress and APIs, and technical SEO keeps coming up in that work. This cleanup was not a plugin. It was understanding how crawlers read status codes and headers and building that into the app.

If your Search Console shows far more indexed pages than you have published, or junk URLs won't leave Google, have a look at my technical SEO services or get in touch.

FAQ

How do I remove a URL from Google search quickly?

Use Search Console → Removals for a fast, temporary hide (about six months), and make the URL return 410 or noindex so it is dropped permanently when Google recrawls it.

Should I use 404 or 410 for hacked URLs?

Use 410 Gone for spam you never want back. Both lead to removal eventually, but 410 says the removal is deliberate and permanent. An X-Robots-Tag: noindex header adds a second signal.

Should I block spam URLs in robots.txt?

No. Blocking them stops Googlebot from crawling, so it never sees your 410 or noindex, and the URLs can stay indexed.

Do Search Console removals permanently deindex pages?

No. They are temporary. To deindex pages from Google permanently, the URLs must return 410, 404 or noindex when recrawled. Never submit your bare domain as a prefix.

Is this the same as the Japanese keyword hack?

It is a related pattern. The Japanese keyword hack, which other sites run into, floods a domain with auto-generated Japanese pages. What I found here was different: /shopdetail/ pages, junk homepage query parameters and multilingual casino slugs. The cleanup approach (allowlist, 410, noindex, careful removals) works for both.

Questions about a similar cleanup? Email me at <[email protected]>, use the contact page, or message me on WhatsApp.

Discuss this with Muhammad

Message Muhammad Zubair Akhtar on WhatsApp or email [email protected].