KEEL · 龙骨 · A CURRICULUM FOR THE AI ERA

02 · 元数据与结构化数据:让页面会自我描述 — keel 龙骨

本章目标:每个页面都有唯一且准确的 title / description / canonical / og 标签,并用 JSON-LD 说明"这是什么"。 验收标准:抓任一页面能提取出完整且互不矛盾的元信息,JSON-LD 通过校验器。

本章目标:每个页面都有唯一且准确的 title / description / canonical / og 标签,并用 JSON-LD 说明"这是什么"。
验收标准:抓任一页面能提取出完整且互不矛盾的元信息,JSON-LD 通过校验器。


一、每个页面都必须有的四件事

<title>工具调用 Tool Calling · AI 全栈工程师学习资料</title>
<meta name="description" content="从 Registry、输入校验到幂等与审批,拆开一条可靠的工具执行链。">
<link rel="canonical" href="https://www.xxxxxx.cn/learning/courses/tools/lessons/1/">
<meta name="robots" content="index, follow">

title 的规则

description 的规则

description 不直接影响排名,但直接影响点击率。搜索结果里用户看到的就是这两行字,写得好比堆词有用得多。

meta robots 的三种状态

<meta name="robots" content="index, follow">      <!-- 默认:收录并跟踪链接 -->
<meta name="robots" content="noindex, nofollow">  <!-- 不收录,不跟踪 -->
<meta name="robots" content="noindex, follow">    <!-- 不收录但可以顺着链接爬 -->

上线前务必确认生产环境没有遗留 noindex——脚手架和预发环境经常带着它。这是又一个"其他都对但不收录"的经典原因:

curl -s https://www.xxxxxx.cn/learning/ | grep -i 'noindex' || echo "干净"

二、Open Graph 与 Twitter Card(社媒分享用)

不直接影响搜索排名,但影响分享出去的样子,间接影响外链获取:

<meta property="og:title" content="工具调用 Tool Calling">
<meta property="og:description" content="从 Registry、输入校验到幂等与审批,拆开一条可靠的工具执行链。">
<meta property="og:type" content="article">
<meta property="og:url" content="https://www.xxxxxx.cn/learning/courses/tools/lessons/1/">
<meta property="og:image" content="https://www.xxxxxx.cn/learning/og/cover.png">
<meta property="og:site_name" content="AI 全栈工程师学习资料">
<meta name="twitter:card" content="summary_large_image">

三条要点:

  1. og:url 必须与 canonical 完全一致
  2. og:image 必须是 https 绝对地址,尺寸建议 1200×630,文件 < 1MB
  3. og:type 区分 website(首页)与 article(内容页)

验证:把 URL 贴进微信/飞书/Twitter 看卡片渲染,或用各家调试工具(Facebook Sharing Debugger、Twitter Card Validator)。


三、JSON-LD 结构化数据

结构化数据是给机器看的"这页是什么"的说明书。用 JSON-LD 格式(推荐,不污染 DOM):

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Article",
  "headline": "工具调用 Tool Calling",
  "description": "从 Registry、输入校验到幂等与审批,拆开一条可靠的工具执行链。",
  "datePublished": "2026-09-20",
  "dateModified": "2026-09-30",
  "author":    { "@type": "Person", "name": "站点作者" },
  "publisher": { "@type": "Organization", "name": "xxxxxx.cn" },
  "mainEntityOfPage": {
    "@type": "WebPage",
    "@id": "https://www.xxxxxx.cn/learning/courses/tools/lessons/1/"
  }
}
</script>

按页面类型选 type

页面 推荐 type 额外字段
首页 WebSite potentialAction(站内搜索)、name、url
文章/章节 Article 或 TechArticle headline、datePublished、author
课程列表 ItemList itemListElement 数组
面包屑 BreadcrumbList itemListElement(带 position)
作者页 Person name、url、sameAs

课程类站点尤其值得加 BreadcrumbList

{
  "@context": "https://schema.org",
  "@type": "BreadcrumbList",
  "itemListElement": [
    { "@type": "ListItem", "position": 1, "name": "首页",
      "item": "https://www.xxxxxx.cn/learning/" },
    { "@type": "ListItem", "position": 2, "name": "基础",
      "item": "https://www.xxxxxx.cn/learning/tracks/foundation/" },
    { "@type": "ListItem", "position": 3, "name": "Tool Calling",
      "item": "https://www.xxxxxx.cn/learning/courses/tools/" }
  ]
}

面包屑在搜索结果里会直接显示成 xxxxxx.cn › 基础 › Tool Calling 这样的路径,点击率高出一截。

验证

# 提取页面里的 JSON-LD
curl -s https://www.xxxxxx.cn/learning/courses/tools/lessons/1/ \
  | grep -o '<script type="application/ld+json">.*</script>' | head -1

然后丢进 Google 的「富媒体搜索结果测试」或 Schema.org 校验器。JSON 语法错误会让整块结构化数据被忽略,而页面看起来毫无异常——所以必须校验。


四、生成策略:一处定义,多处复用

这些标签的共同特点是依赖页面上下文,手写必然出错。本站的做法:

内容源(Markdown + catalog.json)
   ↓  构建脚本
content/manifest.json  ← 唯一的元数据来源
   ↓
① 预渲染 route shell:把 title/description/canonical/JSON-LD 写进静态 HTML
② sitemap.xml:同一份 URL 与 lastmod
③ 前端运行时:SPA 切换路由时同步更新 document.title 与 meta

关键约束:三条产线读同一份数据。任何一条自己拼 URL 或自己造标题,就会出现不一致,而搜索引擎对不一致的惩罚是把决定权收归自己——结果不可控。

前端在 SPA 路由切换时也要更新 meta(用户可能直接分享当前页):

useEffect(() => {
  document.title = lesson.title;
  setMeta('description', lesson.excerpt);
  setMeta('og:url', canonical);
  setLink('canonical', canonical);
}, [lesson]);

五、本章验收

URL=https://www.xxxxxx.cn/learning/courses/tools/lessons/1/
curl -s $URL | grep -oE '<title>[^<]*</title>'
curl -s $URL | grep -oE 'name="description"[^>]*'
curl -s $URL | grep -oE 'rel="canonical"[^>]*'
curl -s $URL | grep -oE 'property="og:[a-z_]*"[^>]*'
curl -s $URL | grep -c 'application/ld+json'
curl -s $URL | grep -ci 'noindex' || echo "无 noindex"

元信息就绪后,就可以正式向搜索引擎提交站点了。

进入 keel 阅读