LLM Structured Output 入門:等 AI 齋出乾淨 JSON,唔再靠 regex 拆答案

JSON mode、function calling、json_schema 三層做法逐一講,附 Python 實戰代碼

AI 教學AI 生成

重點整理

  • 痛點:直接問 LLM 攞資料,佢會夠埋「好嘅,以下是…」呢啲廢話——寫程式想攞乾淨 JSON 就頭痛
  • 三層做法:prompt 講明(最弱)→ JSON mode(強制合法 JSON)→ json_schema(連欄位都幫你鎖死)
  • 實戰:用 Ollama 本地模型 + format schema,零 API 費由影評抽齣電影名/評分/一句總結
  • pitfalls:溫度調低、schema 唔好太深、數字欄位要驗證——LLM 出 JSON 都會出錯

點解要 Structured Output?

你叫 LLM「將呢篇影評嘅評分同重點列出嚟」,佢會答:

好嘅!以下是這篇影評的重點:\n1. 評分:8.5 分…

人類睇冇問題,但程式要 parse 就想死——前綴、編號、全形標點全部要執。Structured output 就係令模型齋出合法 JSON,唔好講嘢。

三層做法(由弱到強)

層級做法可靠度
1. Prompt 講明「請只用 JSON 回答」❌ 時靈時唔靈
2. JSON modeAPI 參數強制輸出合法 JSON✅ 一定合法,但欄位隨佢
3. json_schema畀埋 JSON Schema,欄位/類型都鎖死✅✅ 最穩

實戰:Ollama 本地模型 + Schema

import requests

schema = {
    "type": "object",
    "properties": {
        "movie": {"type": "string"},
        "score": {"type": "number"},
        "one_liner": {"type": "string"},
    },
    "required": ["movie", "score", "one_liner"],
}

r = requests.post("http://localhost:11434/api/chat", json={
    "model": "qwen3:8b",
    "messages": [{"role": "user", "content": "從呢段影評抽齣電影名、評分同一句總結:<評論全文>"}],
    "format": schema,   # Ollama 支援直接畀 JSON Schema
    "stream": False,
    "options": {"temperature": 0},
})
import json as J
print(J.loads(r.json()["message"]["content"]))
# {'movie': 'Dune: Part Two', 'score': 8.5, 'one_liner': '...'}

雲端 API 一樣有對應:OpenAI 用 response_format={"type": "json_schema", ...},Anthropic 就用 tool use 強制結構。

實用 tips

  • 溫度調低(0~0.3):結構化任務唔需要創意
  • Schema 唔好太深:巢狀超過三層,細模型好易漏欄位
  • 數字要驗證:出嚟都係要 pydantic / 手動 check,唔好盲信
  • enum 鎖選項:情感分類就用 "enum": ["pos", "neg", "neutral"],唔會出第四個答案

總結

Structured output 係「LLM 由玩具變工具」嗰一腳——出嘅嘢可以直接入資料庫、入 pipeline。記住三層做法:能唔能用 prompt 解決就 prompt,認真項目就一定上 json_schema。

分享畀朋友

相關文章

更多「AI 教學」

睇全部分類 →