ESLint for AI search. Lint your website for AI-search readiness — AI crawler access, llms.txt, structured data and citability.

🇩🇪 Deutsch · 🇪🇸 Español · 🇯🇵 日本語

No install, no config:

npx @iliasabk/geolint check yoursite.comgeolint fetches the page, its robots.txt and llms.txt, evaluates 51 known AI crawler tokens against your robots.txt, runs 52 audit rules, and prints a scored report with a concrete fix for every finding.

- AI answers are the new front page. ChatGPT, Perplexity, Claude, Copilot and Google AI Overviews send traffic — or don't — based on whether their crawlers can fetch and quote your pages.

- Most sites accidentally block or confuse AI crawlers. A stale

Disallow: /, anoindexleft over from staging, a client-rendered page that looks empty to a bot that doesn't run JavaScript.

- Existing tools are blocklists or score-only web apps. They tell you to block everything, or give you a number with no path to improve it. geolint is the linter: concrete findings, concrete fixes, runnable in CI on every PR.

52 rules across 5 categories — geolint rules lists them all, and

docs/rules.md documents what each rule checks, why it matters

and how to fix violations.

Real output, auditing the bundled demo site (examples/demo-site, which

deliberately blocks two bots) — trimmed for width:

$ geolint check localhost:4173 --ignore technical/https

geolint v0.2.1 — AI-search readiness

http://localhost:4173/

200 OK · text/html · TTFB 113ms · robots 200 · llms.txt 404

██████████████████████████░░░░ 86/100 Grade B

CATEGORIES

AI Crawler Access ███████░░░ 70 ✗ 2 errors

llms.txt █████████░ 92 ⚠ 1 warning · 1 hint

Structured Data █████████░ 88 ⚠ 1 warning · 3 hints

Citability ████████░░ 82 ⚠ 2 warnings · 3 hints

Technical Foundation ██████████ 100 ✓ clean

AI CRAWLER ACCESS — 49/51 allowed · 2 blocked

OpenAI

GPTBot ✓ training

OAI-SearchBot ✓ search

ChatGPT-User ✓ user-fetch

Perplexity

PerplexityBot ✗ search

Perplexity-User ✓ user-fetch

Google

Googlebot ✓ search

Google-Extended ✓ training

… 51 tokens total, grouped by vendor …

FINDINGS

AI Crawler Access

✗ ai-crawler/search-bots-blocked PerplexityBot is blocked by robots.txt — Perplexity cannot use your pages as AI answer sources

fix: Remove the Disallow covering PerplexityBot in robots.txt, or add an explicit "Allow: /" for it.

evidence: Disallow: / (matched by PerplexityBot)

llms.txt

⚠ llms-txt/missing No llms.txt found

fix: Create /llms.txt at the site root: an H1 title, a short blockquote summary, and ## sections linking to your key content.

evidence: http://localhost:4173/llms.txt → HTTP 404

────────────────────────────────────────────────────────────────────

2 errors · 4 warnings · 7 hints · 32/44 checks passed

Every finding carries a rule id, a severity, the evidence geolint matched, and a fix. Compare two pages or two competitors head-to-head:

geolint check a.com --compare b.comFull flag reference: docs/configuration.md.

- uses: iliasabk/geolint@v1

id: geolint

with:

url: https://example.com

fail-under: 80

- uses: github/codeql-action/upload-sarif@v3

if: always()

with:

sarif_file: ${{ steps.geolint.outputs.sarif-file }}The action produces score/grade step outputs, a SARIF report for GitHub code scanning, and a markdown report for job summaries and PR comments. Full recipes — SARIF upload, updating a single PR comment, baseline drift detection — in docs/github-action.md.

npx @iliasabk/geolint check https://example.com --fail-under 80Exit code is 1 when the score drops below the gate (or findings regress

against --baseline), 0 otherwise — works in GitLab CI, CircleCI, npm

scripts, pre-deploy hooks.

npx @iliasabk/geolint check https://example.com --badge

# → writes geolint-badge.svg + prints the markdown snippet to pasteCommit the SVG, or regenerate a shields endpoint JSON in CI

(--badge-endpoint) for a badge that never goes stale.

-f pretty (default) renders the terminal report above. The machine formats:

- -f json— the full- ScanReport: findings, per-category scores, bot access matrix

- -f sarif— SARIF 2.1.0, upload straight to GitHub code scanning

- -f markdown— PR-comment/job-summary-ready tables

- -f html— a self-contained interactive report (score ring, findings filter, bot matrix) you can share or host anywhere

Add -o report.json to write to a file; stdout stays clean for piping.

The repo dogfoods itself: a nightly workflow re-audits eight

well-known sites and commits the scores back, and the showcase

site publishes the full interactive

reports — github.com, anthropic.com, stripe.com and more, regenerated on every

push to main.

import { scan } from '@iliasabk/geolint';

const report = await scan('https://example.com', {

ignore: ['technical/https'],

timeout: 10_000,

});

console.log(report.score, report.grade); // e.g. 86 'B'

for (const f of report.findings) {

console.log(f.severity, f.ruleId, f.message, f.fix);

}scan(url, options) returns a typed ScanReport. Also exported: the bot

registry (AI_BOTS, botsByPurpose), the rule registry (allRules,

ruleById), robots.txt/llms.txt parsers, badge generators, scorers and all

four reporters.

geolint mcp speaks the Model Context Protocol

over stdio — Claude Desktop, Cursor, VS Code and Windsurf can audit sites,

generate llms.txt and compare URLs as native tools:

// claude_desktop_config.json / ~/.cursor/mcp.json

{

"mcpServers": {

"geolint": {

"command": "npx",

"args": ["-y", "@iliasabk/geolint", "mcp"]

}

}

}Five tools: audit_url, generate_llms_txt, compare_urls, list_rules,

list_ai_bots — all read-only, with structured output and per-call timeouts.

Setup for every client: docs/mcp.md.

geolint bots lists 51 AI crawler tokens with a purpose-aware impact

assessment — because "should I block this bot?" has a different answer for each:

And two nuances other tools miss:

- Some fetchers ignore robots.txt. OpenAI, Perplexity and Meta document that

their user-triggered fetchers (ChatGPT-User, Perplexity-User,

Meta-ExternalFetcher) may not honor robots.txt. ai-crawler/user-fetch-bypasstells you when aDisallowwon't work — enforce at the WAF/auth layer instead.

- Stale tokens. anthropic-ai,Claude-Web,FacebookBotare retired.ai-crawler/stale-tokensflags them and names the replacement token — aUser-agent: anthropic-airule does nothing today.

Control-only tokens like Google-Extended and Applebot-Extended never fetch

at all — they only set a preference — and geolint treats them accordingly.

- llms.txt is a proposal, not a standard. No major AI vendor has committed

to reading it — so llms-txt/*findings are weighted as warnings and hints, not errors. geolint still checks it (andgeolint initgenerates it) because adoption is growing and the cost is one file.

- Correlation ≠ causation. The citability rules are grounded in published

GEO research (quotations/statistics/citations measurably lift share-of-answer;

AI crawlers other than Googlebot and Applebot don't execute JavaScript), but

signals like question-shaped headings are hints, not facts — they're infoseverity and geolint says so.

- Every rule shows its reasoning. docs/rules.md documents why each rule exists; the research sources are in docs/research-notes.md, including the vendor docs behind every bot's robots.txt posture.

- The bot registry is a standalone reference.

docs/ai-crawlers.md lists every tracked token with

purpose, per-vendor robots.txt posture and vendor docs — the same data

geolint botsand thelist_ai_botsMCP tool expose.

Details and the reasoning behind each column: docs/comparison.md. geolint also ships an MCP server, a score badge and regression baselines.

Planned for v0.4+:

- geolint watch— re-audit on deploys/file changes

- Custom rule API for project-specific checks

- Deeper schema coverage (more @typevalidators)

- Homebrew formula

- Report localization beyond English

Issues and PRs welcome — see CONTRIBUTING.md. New rules are

the best contribution: each needs a check(ctx), findings with fix, a test

and a docs entry.

If geolint helped, a ⭐ helps others find it.