FreeTechnical SEO tool

Log File Analyzer

Paste lines from your Apache or Nginx access log. The analyzer keeps the ones from Googlebot, pulls the requested path out of each, and shows how many Googlebot lines it found plus a table of the 25 most-requested paths.

Googlebot lines onlyTop 25 paths by hitsParsed in your browser

Updated · by the Linkstonic team

linkstonic.com/tool/log-file-analyzerLive
How it works

How lines are matched and counted

Log formats vary, so the parser keeps its rules short. Here is exactly what it does with each line you paste.

01

Googlebot filter

A line counts if the word "googlebot" appears anywhere in it, in any case. That includes Googlebot-Image and Googlebot-Video. Other Google crawlers like AdsBot-Google, Storebot-Google and GoogleOther are skipped, as is every non-Google bot and all human traffic.

02

No verification

The filter trusts the user agent text. Scrapers that pretend to be Googlebot are counted too. Google recommends confirming real Googlebot with a reverse DNS lookup on the IP, so check the IPs behind any path that looks suspiciously busy.

03

Path extraction

The path is taken from the quoted request, such as "GET /blog/post HTTP/1.1". Only GET, POST and HEAD are recognized. If a line has no quoted request in that shape, its hit is filed under "(unparsed)" rather than dropped.

04

Query strings are dropped

Everything from the question mark onward is cut, so /shop?sort=price and /shop?sort=name are both counted as /shop. That makes totals per page easy to read, but hides which parameter combinations Googlebot is spending time on.

05

What the table shows

Paths are sorted by hit count, highest first, and the table stops at 25 rows. Status codes, response sizes, timestamps and IPs are ignored, so a path returning 404 or 301 looks the same as one returning 200.

Worked example

Six log lines, worked through

We pasted six lines in the standard combined log format: five from Googlebot and one from Bingbot.

Input
66.249.66.1 - - [21/Sep/2026:06:14:02 +0000] "GET /product/blue-mug HTTP/1.1" 200 5123 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
66.249.66.1 - - [21/Sep/2026:06:14:09 +0000] "GET /shop?sort=price HTTP/1.1" 200 8840 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
66.249.66.7 - - [21/Sep/2026:06:15:31 +0000] "GET /shop?sort=name HTTP/1.1" 200 8812 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
40.77.167.2 - - [21/Sep/2026:06:16:40 +0000] "GET /blog/ HTTP/1.1" 200 6420 "-" "Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)"
66.249.66.7 - - [21/Sep/2026:06:17:03 +0000] "GET /product/blue-mug HTTP/1.1" 304 0 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
66.249.66.9 - - [21/Sep/2026:06:18:22 +0000] "GET /robots.txt HTTP/1.1" 200 312 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
Output
5Googlebot lines matched
3Distinct paths
2/product/blue-mug
2/shop
1/robots.txt

The two sorted shop URLs merge into a single /shop row because the query string is dropped. The Bingbot line never appears, and the 304 response is counted the same as the 200s.

Use cases

When to pull the logs

Log data answers questions Search Console can only hint at. These are the moments we reach for it.

After a migration

Paste a day of logs from launch week and check that Googlebot is hitting the new URLs, not hammering old ones. If old paths dominate, your redirects or internal links haven't been picked up yet.

Ecommerce SEOs

Stores leak crawl activity into filters and sort orders. Group by path here, then grep the raw log for the busiest path to see which parameters Googlebot keeps requesting and whether they return 200.

Developers debugging indexing

When a new section isn't getting indexed, the first question is whether Googlebot has requested it at all. A quick paste answers that in seconds, without setting up a log platform.

Site owners chasing slow indexing

If new blog posts take weeks to appear in Google, paste a week of logs and look for them in the table. No hits means discovery is the problem, so check internal links and the sitemap.

Reference

Google crawlers you'll see in logs

Not every Google user agent contains the word Googlebot. This is what the filter catches and what it leaves out.

User agent tokenWhat it fetchesCounted here
Googlebot (smartphone)Pages for mobile-first indexing, the bulk of Google's crawlingYes
Googlebot (desktop)Pages from a desktop user agent, now a small shareYes
Googlebot-ImageImages for Google Images and other image featuresYes
Googlebot-VideoVideo files and video-related pagesYes
Storebot-GoogleProduct and checkout pages for Google ShoppingNo
AdsBot-GoogleLanding pages to check Google Ads qualityNo
GoogleOtherOne-off fetches by Google teams, such as researchNo
Google-InspectionToolSearch Console URL Inspection and the Rich Results TestNo

Google-Extended never appears in logs. It's a robots.txt token only, controlling use of content for Gemini training, and crawling is done by the regular Googlebot.

Log File Analyzer illustration: a bar chart icon next to a sample result card showing 5 googlebot lines matched and 3 distinct paths
Log File Analyzer, sample result
Our take

Logs are the only first-hand crawl data you have

Search Console's Crawl Stats report is useful, but it's a sample and a summary. Your server logs record every request Googlebot actually made. When you want to know whether a page was crawled last Tuesday, or how often a section gets visited, the log is the only place with a definite answer.

This analyzer is a quick look, not a log platform. It suits a few thousand lines copied from a server or a CDN export. For months of data, or to break results down by status code and day, you'll want a proper log tool or a warehouse. I'd use this one to answer one question fast.

AI crawlers show up in the same logs. GPTBot, ClaudeBot and PerplexityBot all identify themselves in the user agent, but this tool only counts Googlebot. To see them, search the raw log for their names. Some AI requests arrive through CDNs or caches and never reach your origin server at all.

Practical tips

Getting clean results

A few minutes of prep makes the table far more useful.

  1. 01

    Filter to Googlebot before you paste, for example with grep -i googlebot access.log, so the text box isn't handling millions of irrelevant lines.

  2. 02

    Paste one day at a time. Crawl patterns change daily and a single day is easier to compare against a known deploy or outage.

  3. 03

    Check the (unparsed) row. If it's large, your log format doesn't quote the request line and the path can't be read.

  4. 04

    Look for 404 and 301 responses on your top paths separately, since this table counts every status the same.

  5. 05

    Verify a few of the busiest IPs with a reverse DNS lookup. Genuine Googlebot resolves to googlebot.com or google.com.

  6. 06

    Keep a copy of the top-25 table from each run. Comparing this week's list with last month's shows whether Googlebot's attention is moving toward the pages you care about.

Log File Analyzer tips illustration: a checklist card that starts with "Filter to Googlebot before you paste," and "Paste one day at a time"
Log File Analyzer, checklist
Our customers

How we help marketers win

01techventuresShare of voice
02northlineReporting time
03arclabsAI mentions
FAQ

Log File Analyzer questions

01What log formats does the analyzer accept?

Any text log where each request is on its own line with a quoted request like "GET /path HTTP/1.1" and the user agent somewhere on the same line. The Apache and Nginx combined formats both work. JSON logs work if they keep that quoted request string.

02Are my server logs uploaded anywhere?

No. The lines are parsed inside your browser tab and the tool makes no network request. That matters because access logs contain visitor IP addresses, which you usually shouldn't share with third parties.

03Can log analysis replace Search Console?

No, they answer different questions. Logs show what Googlebot requested from your server. Search Console shows what Google indexed, how pages perform in search and which errors it found. Use logs to explain what you see in Search Console.

04How do I know a Googlebot hit is real?

Run a reverse DNS lookup on the IP address. Real Googlebot resolves to a googlebot.com or google.com hostname, and a forward lookup of that hostname returns the same IP. Google also publishes its crawler IP ranges in JSON files.

05Why are my URLs with parameters grouped together?

The parser cuts everything after the question mark, so /shop?page=2 and /shop?page=3 are both counted as /shop. To see individual parameter URLs, search the raw log for that path instead.

Beyond one-off checks

See how your pages do in Google and AI answers

Linkstonic tracks rankings, audits your site and shows whether ChatGPT, Gemini and Perplexity mention you. The free plan needs no credit card.