KEEL · 龙骨 · A CURRICULUM FOR THE AI ERA

SEO 速查清单 — keel 龙骨

SEO 收录与优化 的参考信息:SEO 速查清单

一页纸版本,按执行顺序排列。

A. 地基(上线前必做)

# 正文在 HTML 里(不依赖 JS)
curl -s https://www.xxxxxx.cn/learning/courses/tools/lessons/1/ | grep -c '本章目标'

# robots.txt(必须在域名根目录)
curl -s https://www.xxxxxx.cn/robots.txt
#   应含:User-agent: * / Allow: / / Sitemap: https://...

# sitemap
curl -s https://www.xxxxxx.cn/learning/sitemap.xml | grep -c '<loc>'

# 无 noindex 残留
curl -s https://www.xxxxxx.cn/learning/ | grep -ci noindex || echo OK

# 未知路径返回 404 而非 200
curl -s -o /dev/null -w "%{http_code}\n" https://www.xxxxxx.cn/learning/not-a-real-page/

B. 元数据

C. 提交(三家)

平台 地址 填什么 验证方式
Google search.google.com/search-console 网址前缀 https://www.xxxxxx.cn/ DNS TXT(推荐)/ HTML 文件
Bing bing.com/webmasters 从 GSC 导入 免二次验证
百度 ziyuan.baidu.com www.xxxxxx.cn CNAME(推荐)/ HTML / TXT

sitemap 统一填:https://www.xxxxxx.cn/learning/sitemap.xml

⚠️ 验证记录(TXT/CNAME/HTML 文件)永久保留,删了会失去所有权。
⚠️ 百度对未备案 + 境外服务器基本不收录,先判断再投入。

D. 诊断(日志)

LOG=/var/log/nginx/learning.access.log

# 谁来了
grep -iE 'googlebot|bingbot|baiduspider|bytespider|GPTBot|ClaudeBot|PerplexityBot' $LOG \
  | grep -oiE 'Googlebot|Bingbot|Baiduspider|Bytespider|GPTBot|ClaudeBot|PerplexityBot' \
  | sort | uniq -c | sort -rn

# 抓到了什么(非 2xx)
grep -iE 'googlebot|bingbot|baiduspider' $LOG | awk '$9>=400 {print $9, $7}' | sort | uniq -c | sort -rn

# 真假爬虫(双向 DNS)
dig -x <IP> +short        # 应为 *.googlebot.com / *.google.com

判据:计数 0 = 没提交过;有抓取但没收录 = 内容/canonical 问题。

E. 内容与内链

F. 性能

指标 良好 常见解法
LCP ≤ 2.5s 预渲染/SSR、首屏图片不加 lazy、内联关键 CSS
INP ≤ 200ms 减少长任务、缓存渲染结果、缩小 MutationObserver 范围
CLS ≤ 0.1 图片显式尺寸、骨架屏等尺寸占位、字体预加载
curl -sI --http2 https://www.xxxxxx.cn/learning/ | head -1           # HTTP/2
curl -sI -H 'Accept-Encoding: gzip' https://www.xxxxxx.cn/learning/ | grep -i content-encoding
curl -sI https://www.xxxxxx.cn/learning/assets/index-xxx.js | grep -i cache-control  # immutable

G. 月度

最重要的一条

判断收录只用 site: 查询,不要用"搜原文"。
搜原文搜不到 = 没收录 或 收录了但排名靠后,两者的处理方式完全不同。

进入 keel 阅读